{
  "id": 602435,
  "title": "3rd place solution ",
  "url": "/competitions/multi-class-object-detection-challenge/writeups/3rd-place-solution",
  "author_name": "",
  "post_date": "2025-08-27T10:09:55.973Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p><strong>Solution Overview</strong></p>\n<p>Many thanks to the organizers for hosting this challenge. It was a valuable experience, and I’d like to briefly share my approach and key takeaways for anyone interested in the solution.</p>\n<p><strong>Data sources</strong></p>\n<p>Here’s everything I used:<br>\n1 – Multi-class Object Detection Challenge – the official competition dataset.<br>\n2 – Synthetic data generated with FalconCloud – This formed the major portion of my dataset. The dataset was created in two modes: automatic generation (where I specified object attributes and FalconCloud produced the images) and manual mode (where I directly controlled the camera to capture scenes). At the beginning, I focused on producing a large volume of data to enrich the training set. Later, I shifted toward more controlled scenarios aimed at addressing weaknesses I observed during model evaluation:</p>\n<ul>\n<li>Inability to detect overlapping objects<br>\n→ Improved by creating scenarios with intentional overlaps and providing a variety of perspectives to help the model generalize.</li>\n<li>Difficulty in detecting distant soup cans<br>\n→ Addressed by capturing images at longer distances with varying backgrounds and lighting conditions.</li>\n<li>Limited variety in object placement<br>\n→ Addressed by frequently rearranging objects, changing orientations, and incorporating edge-case scenarios.</li>\n<li>Close-up captures of detection targets and similar objects <br>\n→ Objects and visually similar items were photographed up close from multiple angles and perspectives to improve the model’s ability to distinguish between them.<br>\n3 – Synthetic datasets generated via FalconCloud by <a href=\"https://www.kaggle.com/kadirkrtls\" target=\"_blank\">@kadirkrtls</a> </li>\n</ul>\n<p>During the training process, the dataset was iteratively refined by removing images in which the detection confidence of objects was below certain thresholds. Multiple dataset variations were created, including only images with detection probabilities above 0.6, 0.65, 0.7, and 0.8. The best model performance was achieved using the dataset filtered at a threshold of 0.65, balancing quality and diversity.<br>\nAdditionally, the validation set was increased to better reflect the overall data distribution. Originally comprising less than 10% of the total dataset, it was expanded to approximately 15% by randomly selecting additional images from the training data. This adjustment helped improve validation reliability and provided a more representative evaluation of model performance. However, the validation set remained somewhat irrelevant to the test data, primarily because the test images were real-world photos, whereas the training and validation images were synthetically generated. This mismatch limited the direct comparability between validation and test performance, but the refinements still helped improve model robustness.</p>\n<p>**Training Pipeline and Hyperparameter Experiments **<br>\nModels were trained using YOLO architectures with careful control over reproducibility by setting seed values. I experimented with models of various sizes, including 8l, 8m, 11m, 11l, 11x, and 12x, and observed significant performance differences between the architectures. Surprisingly, the 12x model performed significantly worse than the 11x variant, despite being larger.<br>\nTwo main sets of hyperparameters were tested the most. I varied and tested different parameters, eventually coming to test the models mainly by varying the flip probability, brightness adjustment, number of frozen layers, dropout, and learning rate (lr0 and lrf). Interestingly, the best performance was achieved with the 11x model with the second set of hyperparameters, and the surprisingly low number of epochs (20) outperformed the longer runs (50, 80, 100, and 200 epochs).</p>\n<p>The models were trained locally, but I will attach the code with the two main sets of hyperparameters in a notebook for reference.</p>\n<p><strong>Thoughts</strong><br>\nUndoubtedly, careful dataset preparation is crucial for achieving better results. Technical aspect is also important: it would be interesting to train models with larger batch sizes, which generally perform more reliably and achieve higher accuracy, but require a powerful GPU. Balancing dataset quality, parameters, and computational resources is key to optimizing performance.</p>",
  "messages": [
    {
      "id": "3277092",
      "postDate": "08/27/2025 10:09:29",
      "content": "<p><strong>Solution Overview</strong></p>\n<p>Many thanks to the organizers for hosting this challenge. It was a valuable experience, and I’d like to briefly share my approach and key takeaways for anyone interested in the solution.</p>\n<p><strong>Data sources</strong></p>\n<p>Here’s everything I used:<br>\n1 – Multi-class Object Detection Challenge – the official competition dataset.<br>\n2 – Synthetic data generated with FalconCloud – This formed the major portion of my dataset. The dataset was created in two modes: automatic generation (where I specified object attributes and FalconCloud produced the images) and manual mode (where I directly controlled the camera to capture scenes). At the beginning, I focused on producing a large volume of data to enrich the training set. Later, I shifted toward more controlled scenarios aimed at addressing weaknesses I observed during model evaluation:</p>\n<ul>\n<li>Inability to detect overlapping objects<br>\n→ Improved by creating scenarios with intentional overlaps and providing a variety of perspectives to help the model generalize.</li>\n<li>Difficulty in detecting distant soup cans<br>\n→ Addressed by capturing images at longer distances with varying backgrounds and lighting conditions.</li>\n<li>Limited variety in object placement<br>\n→ Addressed by frequently rearranging objects, changing orientations, and incorporating edge-case scenarios.</li>\n<li>Close-up captures of detection targets and similar objects <br>\n→ Objects and visually similar items were photographed up close from multiple angles and perspectives to improve the model’s ability to distinguish between them.<br>\n3 – Synthetic datasets generated via FalconCloud by <a href=\"https://www.kaggle.com/kadirkrtls\" target=\"_blank\">@kadirkrtls</a> </li>\n</ul>\n<p>During the training process, the dataset was iteratively refined by removing images in which the detection confidence of objects was below certain thresholds. Multiple dataset variations were created, including only images with detection probabilities above 0.6, 0.65, 0.7, and 0.8. The best model performance was achieved using the dataset filtered at a threshold of 0.65, balancing quality and diversity.<br>\nAdditionally, the validation set was increased to better reflect the overall data distribution. Originally comprising less than 10% of the total dataset, it was expanded to approximately 15% by randomly selecting additional images from the training data. This adjustment helped improve validation reliability and provided a more representative evaluation of model performance. However, the validation set remained somewhat irrelevant to the test data, primarily because the test images were real-world photos, whereas the training and validation images were synthetically generated. This mismatch limited the direct comparability between validation and test performance, but the refinements still helped improve model robustness.</p>\n<p>**Training Pipeline and Hyperparameter Experiments **<br>\nModels were trained using YOLO architectures with careful control over reproducibility by setting seed values. I experimented with models of various sizes, including 8l, 8m, 11m, 11l, 11x, and 12x, and observed significant performance differences between the architectures. Surprisingly, the 12x model performed significantly worse than the 11x variant, despite being larger.<br>\nTwo main sets of hyperparameters were tested the most. I varied and tested different parameters, eventually coming to test the models mainly by varying the flip probability, brightness adjustment, number of frozen layers, dropout, and learning rate (lr0 and lrf). Interestingly, the best performance was achieved with the 11x model with the second set of hyperparameters, and the surprisingly low number of epochs (20) outperformed the longer runs (50, 80, 100, and 200 epochs).</p>\n<p>The models were trained locally, but I will attach the code with the two main sets of hyperparameters in a notebook for reference.</p>\n<p><strong>Thoughts</strong><br>\nUndoubtedly, careful dataset preparation is crucial for achieving better results. Technical aspect is also important: it would be interesting to train models with larger batch sizes, which generally perform more reliably and achieve higher accuracy, but require a powerful GPU. Balancing dataset quality, parameters, and computational resources is key to optimizing performance.</p>",
      "rawMarkdown": "**Solution Overview**\n\nMany thanks to the organizers for hosting this challenge. It was a valuable experience, and I’d like to briefly share my approach and key takeaways for anyone interested in the solution.\n\n**Data sources**\n\nHere’s everything I used:\n1 – Multi-class Object Detection Challenge – the official competition dataset.\n2 – Synthetic data generated with FalconCloud – This formed the major portion of my dataset. The dataset was created in two modes: automatic generation (where I specified object attributes and FalconCloud produced the images) and manual mode (where I directly controlled the camera to capture scenes). At the beginning, I focused on producing a large volume of data to enrich the training set. Later, I shifted toward more controlled scenarios aimed at addressing weaknesses I observed during model evaluation:\n- Inability to detect overlapping objects\n   → Improved by creating scenarios with intentional overlaps and providing a variety of perspectives to help the model generalize.\n- Difficulty in detecting distant soup cans\n   → Addressed by capturing images at longer distances with varying backgrounds and lighting conditions.\n- Limited variety in object placement\n   → Addressed by frequently rearranging objects, changing orientations, and incorporating edge-case scenarios.\n- Close-up captures of detection targets and similar objects \n   → Objects and visually similar items were photographed up close from multiple angles and perspectives to improve the model’s ability to distinguish between them.\n3 – Synthetic datasets generated via FalconCloud by @kadirkrtls \n\nDuring the training process, the dataset was iteratively refined by removing images in which the detection confidence of objects was below certain thresholds. Multiple dataset variations were created, including only images with detection probabilities above 0.6, 0.65, 0.7, and 0.8. The best model performance was achieved using the dataset filtered at a threshold of 0.65, balancing quality and diversity.\nAdditionally, the validation set was increased to better reflect the overall data distribution. Originally comprising less than 10% of the total dataset, it was expanded to approximately 15% by randomly selecting additional images from the training data. This adjustment helped improve validation reliability and provided a more representative evaluation of model performance. However, the validation set remained somewhat irrelevant to the test data, primarily because the test images were real-world photos, whereas the training and validation images were synthetically generated. This mismatch limited the direct comparability between validation and test performance, but the refinements still helped improve model robustness.\n\n**Training Pipeline and Hyperparameter Experiments **\nModels were trained using YOLO architectures with careful control over reproducibility by setting seed values. I experimented with models of various sizes, including 8l, 8m, 11m, 11l, 11x, and 12x, and observed significant performance differences between the architectures. Surprisingly, the 12x model performed significantly worse than the 11x variant, despite being larger.\nTwo main sets of hyperparameters were tested the most. I varied and tested different parameters, eventually coming to test the models mainly by varying the flip probability, brightness adjustment, number of frozen layers, dropout, and learning rate (lr0 and lrf). Interestingly, the best performance was achieved with the 11x model with the second set of hyperparameters, and the surprisingly low number of epochs (20) outperformed the longer runs (50, 80, 100, and 200 epochs).\n\nThe models were trained locally, but I will attach the code with the two main sets of hyperparameters in a notebook for reference.\n\n**Thoughts**\nUndoubtedly, careful dataset preparation is crucial for achieving better results. Technical aspect is also important: it would be interesting to train models with larger batch sizes, which generally perform more reliably and achieve higher accuracy, but require a powerful GPU. Balancing dataset quality, parameters, and computational resources is key to optimizing performance.",
      "votes": null
    },
    {
      "id": "3281729",
      "postDate": "09/04/2025 19:11:36",
      "content": "<p>Thank you for sharing!<br>\nI was looking for your model weights, too, but didn't see them. Could you either tell me where they are or send them to me? Thank you!</p>",
      "rawMarkdown": "Thank you for sharing!\nI was looking for your model weights, too, but didn't see them. Could you either tell me where they are or send them to me? Thank you!",
      "votes": null
    },
    {
      "id": "3281747",
      "postDate": "09/04/2025 20:18:02",
      "content": "<p>I was also curious how many images you had for each dataset at the different filtered thresholds?</p>",
      "rawMarkdown": "I was also curious how many images you had for each dataset at the different filtered thresholds?",
      "votes": null
    },
    {
      "id": "3283324",
      "postDate": "09/07/2025 14:26:23",
      "content": "<p>Most objects in photos have a threshold of 0.95, so when filtering above 0.6, 0.65, 0.7, up to 150 photos out of 4100+ are removed</p>",
      "rawMarkdown": "Most objects in photos have a threshold of 0.95, so when filtering above 0.6, 0.65, 0.7, up to 150 photos out of 4100+ are removed",
      "votes": null
    },
    {
      "id": "3283326",
      "postDate": "09/07/2025 14:30:40",
      "content": "<p>In notebook called multi_class_obj_det1 there are planty of models, best results I achieved with best_143.pt</p>",
      "rawMarkdown": "In notebook called multi_class_obj_det1 there are planty of models, best results I achieved with best_143.pt",
      "votes": null
    },
    {
      "id": "3285936",
      "postDate": "09/08/2025 23:12:19",
      "content": "<p>Noted. I left you a comment there, it looks like we are unable to see models from that location. Are you able to either send it to us or post it in the models tab? Thanks! <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25370829%2F8b374bcce40f38da0f617ce1dc175d05%2FScreenshot%202025-09-08%20171051.png?generation=1757373138974742&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Noted. I left you a comment there, it looks like we are unable to see models from that location. Are you able to either send it to us or post it in the models tab? Thanks! ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25370829%2F8b374bcce40f38da0f617ce1dc175d05%2FScreenshot%202025-09-08%20171051.png?generation=1757373138974742&alt=media)",
      "votes": null
    },
    {
      "id": "3285978",
      "postDate": "09/09/2025 02:07:00",
      "content": "<p>can't seem to access the models in multi_class_obj_det1 notebook</p>",
      "rawMarkdown": "can't seem to access the models in multi_class_obj_det1 notebook",
      "votes": null
    },
    {
      "id": "3286101",
      "postDate": "09/09/2025 08:26:30",
      "content": "<p><a href=\"https://www.kaggle.com/models/kostya876/best_for_multi_class_obj_det_challenge\" target=\"_blank\">https://www.kaggle.com/models/kostya876/best_for_multi_class_obj_det_challenge</a></p>",
      "rawMarkdown": "https://www.kaggle.com/models/kostya876/best_for_multi_class_obj_det_challenge",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3281729,
      "author_name": "rebekahduality",
      "author_url": "",
      "post_date": "09/04/2025 19:11:36",
      "content": "<p>Thank you for sharing!<br>\nI was looking for your model weights, too, but didn't see them. Could you either tell me where they are or send them to me? Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3281747,
          "author_name": "rebekahduality",
          "author_url": "",
          "post_date": "09/04/2025 20:18:02",
          "content": "<p>I was also curious how many images you had for each dataset at the different filtered thresholds?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3283324,
              "author_name": "kostya876",
              "author_url": "",
              "post_date": "09/07/2025 14:26:23",
              "content": "<p>Most objects in photos have a threshold of 0.95, so when filtering above 0.6, 0.65, 0.7, up to 150 photos out of 4100+ are removed</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3283326,
          "author_name": "kostya876",
          "author_url": "",
          "post_date": "09/07/2025 14:30:40",
          "content": "<p>In notebook called multi_class_obj_det1 there are planty of models, best results I achieved with best_143.pt</p>",
          "votes": null,
          "replies": [
            {
              "id": 3285936,
              "author_name": "rebekahduality",
              "author_url": "",
              "post_date": "09/08/2025 23:12:19",
              "content": "<p>Noted. I left you a comment there, it looks like we are unable to see models from that location. Are you able to either send it to us or post it in the models tab? Thanks! <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25370829%2F8b374bcce40f38da0f617ce1dc175d05%2FScreenshot%202025-09-08%20171051.png?generation=1757373138974742&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3286101,
                  "author_name": "kostya876",
                  "author_url": "",
                  "post_date": "09/09/2025 08:26:30",
                  "content": "<p><a href=\"https://www.kaggle.com/models/kostya876/best_for_multi_class_obj_det_challenge\" target=\"_blank\">https://www.kaggle.com/models/kostya876/best_for_multi_class_obj_det_challenge</a></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3285978,
              "author_name": "muhammadhaaris27083",
              "author_url": "",
              "post_date": "09/09/2025 02:07:00",
              "content": "<p>can't seem to access the models in multi_class_obj_det1 notebook</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3277092": "**Solution Overview**\n\nMany thanks to the organizers for hosting this challenge. It was a valuable experience, and I’d like to briefly share my approach and key takeaways for anyone interested in the solution.\n\n**Data sources**\n\nHere’s everything I used:\n1 – Multi-class Object Detection Challenge – the official competition dataset.\n2 – Synthetic data generated with FalconCloud – This formed the major portion of my dataset. The dataset was created in two modes: automatic generation (where I specified object attributes and FalconCloud produced the images) and manual mode (where I directly controlled the camera to capture scenes). At the beginning, I focused on producing a large volume of data to enrich the training set. Later, I shifted toward more controlled scenarios aimed at addressing weaknesses I observed during model evaluation:\n- Inability to detect overlapping objects\n   → Improved by creating scenarios with intentional overlaps and providing a variety of perspectives to help the model generalize.\n- Difficulty in detecting distant soup cans\n   → Addressed by capturing images at longer distances with varying backgrounds and lighting conditions.\n- Limited variety in object placement\n   → Addressed by frequently rearranging objects, changing orientations, and incorporating edge-case scenarios.\n- Close-up captures of detection targets and similar objects \n   → Objects and visually similar items were photographed up close from multiple angles and perspectives to improve the model’s ability to distinguish between them.\n3 – Synthetic datasets generated via FalconCloud by @kadirkrtls \n\nDuring the training process, the dataset was iteratively refined by removing images in which the detection confidence of objects was below certain thresholds. Multiple dataset variations were created, including only images with detection probabilities above 0.6, 0.65, 0.7, and 0.8. The best model performance was achieved using the dataset filtered at a threshold of 0.65, balancing quality and diversity.\nAdditionally, the validation set was increased to better reflect the overall data distribution. Originally comprising less than 10% of the total dataset, it was expanded to approximately 15% by randomly selecting additional images from the training data. This adjustment helped improve validation reliability and provided a more representative evaluation of model performance. However, the validation set remained somewhat irrelevant to the test data, primarily because the test images were real-world photos, whereas the training and validation images were synthetically generated. This mismatch limited the direct comparability between validation and test performance, but the refinements still helped improve model robustness.\n\n**Training Pipeline and Hyperparameter Experiments **\nModels were trained using YOLO architectures with careful control over reproducibility by setting seed values. I experimented with models of various sizes, including 8l, 8m, 11m, 11l, 11x, and 12x, and observed significant performance differences between the architectures. Surprisingly, the 12x model performed significantly worse than the 11x variant, despite being larger.\nTwo main sets of hyperparameters were tested the most. I varied and tested different parameters, eventually coming to test the models mainly by varying the flip probability, brightness adjustment, number of frozen layers, dropout, and learning rate (lr0 and lrf). Interestingly, the best performance was achieved with the 11x model with the second set of hyperparameters, and the surprisingly low number of epochs (20) outperformed the longer runs (50, 80, 100, and 200 epochs).\n\nThe models were trained locally, but I will attach the code with the two main sets of hyperparameters in a notebook for reference.\n\n**Thoughts**\nUndoubtedly, careful dataset preparation is crucial for achieving better results. Technical aspect is also important: it would be interesting to train models with larger batch sizes, which generally perform more reliably and achieve higher accuracy, but require a powerful GPU. Balancing dataset quality, parameters, and computational resources is key to optimizing performance.",
    "3281729": "Thank you for sharing!\nI was looking for your model weights, too, but didn't see them. Could you either tell me where they are or send them to me? Thank you!",
    "3281747": "I was also curious how many images you had for each dataset at the different filtered thresholds?",
    "3283324": "Most objects in photos have a threshold of 0.95, so when filtering above 0.6, 0.65, 0.7, up to 150 photos out of 4100+ are removed",
    "3283326": "In notebook called multi_class_obj_det1 there are planty of models, best results I achieved with best_143.pt",
    "3285936": "Noted. I left you a comment there, it looks like we are unable to see models from that location. Are you able to either send it to us or post it in the models tab? Thanks! ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25370829%2F8b374bcce40f38da0f617ce1dc175d05%2FScreenshot%202025-09-08%20171051.png?generation=1757373138974742&alt=media)",
    "3285978": "can't seem to access the models in multi_class_obj_det1 notebook",
    "3286101": "https://www.kaggle.com/models/kostya876/best_for_multi_class_obj_det_challenge"
  },
  "source": "meta"
}