{
  "id": 583289,
  "title": "22th Place Solution - single yolov8m with Pseudo label",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/shake-game-and-good-luck-22th-place-solution-singl",
  "author_name": "",
  "post_date": "2025-06-06T07:26:22.117Z",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<h3>Introduction</h3>\n<p>First I want to thank Kaggle and the organizers of the competition for the baseline provided. </p>\n<p>Our final solution is a simple YOLOv8m model, the Place Solution—YOLOv8m. </p>\n<p>We experimented with multiple image sizes (640, 960, 1024) and YOLO versions, and we finally decided to use YOLOv8m. We adopted several data augmentation techniques shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , such as mosaic, mixup, perspective, shear, and auto augmentation. On top of that, CLAHE enhances local contrast and emphasizes fine-grained textures, thereby improving feature discrimination in low-contrast scenarios.</p>\n<p>Here's the full setting of our augmentations:</p>\n<pre><code>T = [\n         A.RandomBrightnessContrast(p=, brightness_limit=, contrast_limit=),            \n         A.RandomGamma(p=, gamma_limit=(, )),\n         A.HorizontalFlip(p=),\n         A.VerticalFlip(p=),\n         A.ShiftScaleRotate(shift_limit=, scale_limit=, rotate_limit=, border_mode=, p=),\n         A.ImageCompression(quality_lower=, quality_upper=, p=),\n         A.CLAHE(p=, clip_limit=),\n         A.ToGray(p=),\n       ]\n</code></pre>\n<p>Our final solution employs a single YOLOv8m model without any ensemble, resulting in a relatively simple inference pipeline. However, we believe that this competition requires participants to focus on in-depth data analysis, particularly in identifying and addressing domain shifts arising from different instruments and acquisition devices. Moreover, it is crucial to design solutions that are robust to threshold variations in order to mitigate the risk of performance degradation on the private leaderboard.</p>\n<p>According to the competition's evaluation metric, detections within a relatively large radius are considered correct, and the F𝛽 score used in the competition tolerates a higher number of false positives. This design choice is of particular significance and warrants deeper reflection.</p>\n<p>In the context of medical image object detection, missing true targets is often more detrimental than detecting spurious ones. Models in such applications serve the purpose of narrowing down massive datasets into candidate regions that can be manually reviewed and annotated with higher precision. Ultimately, this human-in-the-loop approach ensures that the final annotations are more accurate. From the perspective of medical practitioners, especially medical students involved in annotation, it is highly desirable for models to achieve high recall while maintaining a reasonably acceptable precision. Such a balance reduces their workload by minimizing missed detections while keeping the number of false alarms manageable.</p>\n<p>The key strategy come up from the simple fact I observed. We can iteratively use pseudo labelling to address the missing gt labels in the training data (I think it may be somehow just like the competition organizers they would do on the leaderborad data) . This is also the key reason my 2 yolov8m models can reach plb840 and lb822 (we selected a plb830;lb824 a single model though), which is possible to reach gold zone with only simple yolov8m model!! <a href=\"https://www.kaggle.com/code/shanzhong8/22nd-solution-tta-yolo-ensemblemodel-plb840/edit\" target=\"_blank\">840PLB model</a></p>\n<h3><strong>Key Improvements</strong></h3>\n<ul>\n<li><p><strong>Correction of data generation issues in the public notebook</strong>:<br>\nWe identified and resolved a problem where, in images containing multiple motors, the dataset generation process saved multiple YOLO-format annotations and corresponding <code>.txt</code> labels. This led to conflicting positive and negative samples for the same input during training, causing the model to become confused and prone to overfitting to noise.</p></li>\n<li><p><strong>Incorporation of negative samples</strong>:<br>\nWe added negative samples to balance the positive-to-negative sample ratio, which helps the model learn more robust decision boundaries.</p></li>\n<li><p><strong>Loss function enhancements</strong>:<br>\nWe introduced additional loss functions including <strong>Normalized Wasserstein Distance (NWD) Loss</strong> and <strong>IoU-Shape Loss</strong> to improve the model’s ability to detect small objects effectively.</p></li>\n<li><p><strong>Utilization of additional annotated data</strong>:<br>\nWe leveraged 36-box additional data shared by <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> and a subset of 24-box training data. Despite the higher noise level in this supplementary dataset, it contributed positively to model generalization.</p></li>\n<li><p><strong>TTA</strong>:<br>\nRotation during Inference makes stable performance.</p></li>\n</ul>\n<h3><strong>Limitations and Unresolved Issues</strong></h3>\n<ul>\n<li><p><strong>Lack of ensemble strategy</strong>:<br>\nDue to time constraints, we did not implement an ensemble strategy. Our final submission was based on a single YOLOv8m model.</p></li>\n<li><p><strong>2.5D DEiM model underperformance</strong>:<br>\nAlthough we experimented with a 2.5D DEiM variant that achieved nearly perfect recall (~1.0), its performance on the leaderboard was suboptimal, potentially due to a high false positive rate or limited generalization across test domains.</p></li>\n</ul>\n<h3>Summary</h3>\n<p>Despite the inherent variability in leaderboard rankings, innovative techniques—such as the rank-based thresholding approach in Bartley’s solution—demonstrate how thoughtful design can yield robust results. I am truly grateful for the insights and knowledge shared by the community throughout this competition.</p>",
  "messages": [
    {
      "id": "3218176",
      "postDate": "06/05/2025 23:58:00",
      "content": "<h3>Introduction</h3>\n<p>First I want to thank Kaggle and the organizers of the competition for the baseline provided. </p>\n<p>Our final solution is a simple YOLOv8m model, the Place Solution—YOLOv8m. </p>\n<p>We experimented with multiple image sizes (640, 960, 1024) and YOLO versions, and we finally decided to use YOLOv8m. We adopted several data augmentation techniques shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , such as mosaic, mixup, perspective, shear, and auto augmentation. On top of that, CLAHE enhances local contrast and emphasizes fine-grained textures, thereby improving feature discrimination in low-contrast scenarios.</p>\n<p>Here's the full setting of our augmentations:</p>\n<pre><code>T = [\n         A.RandomBrightnessContrast(p=, brightness_limit=, contrast_limit=),            \n         A.RandomGamma(p=, gamma_limit=(, )),\n         A.HorizontalFlip(p=),\n         A.VerticalFlip(p=),\n         A.ShiftScaleRotate(shift_limit=, scale_limit=, rotate_limit=, border_mode=, p=),\n         A.ImageCompression(quality_lower=, quality_upper=, p=),\n         A.CLAHE(p=, clip_limit=),\n         A.ToGray(p=),\n       ]\n</code></pre>\n<p>Our final solution employs a single YOLOv8m model without any ensemble, resulting in a relatively simple inference pipeline. However, we believe that this competition requires participants to focus on in-depth data analysis, particularly in identifying and addressing domain shifts arising from different instruments and acquisition devices. Moreover, it is crucial to design solutions that are robust to threshold variations in order to mitigate the risk of performance degradation on the private leaderboard.</p>\n<p>According to the competition's evaluation metric, detections within a relatively large radius are considered correct, and the F𝛽 score used in the competition tolerates a higher number of false positives. This design choice is of particular significance and warrants deeper reflection.</p>\n<p>In the context of medical image object detection, missing true targets is often more detrimental than detecting spurious ones. Models in such applications serve the purpose of narrowing down massive datasets into candidate regions that can be manually reviewed and annotated with higher precision. Ultimately, this human-in-the-loop approach ensures that the final annotations are more accurate. From the perspective of medical practitioners, especially medical students involved in annotation, it is highly desirable for models to achieve high recall while maintaining a reasonably acceptable precision. Such a balance reduces their workload by minimizing missed detections while keeping the number of false alarms manageable.</p>\n<p>The key strategy come up from the simple fact I observed. We can iteratively use pseudo labelling to address the missing gt labels in the training data (I think it may be somehow just like the competition organizers they would do on the leaderborad data) . This is also the key reason my 2 yolov8m models can reach plb840 and lb822 (we selected a plb830;lb824 a single model though), which is possible to reach gold zone with only simple yolov8m model!! <a href=\"https://www.kaggle.com/code/shanzhong8/22nd-solution-tta-yolo-ensemblemodel-plb840/edit\" target=\"_blank\">840PLB model</a></p>\n<h3><strong>Key Improvements</strong></h3>\n<ul>\n<li><p><strong>Correction of data generation issues in the public notebook</strong>:<br>\nWe identified and resolved a problem where, in images containing multiple motors, the dataset generation process saved multiple YOLO-format annotations and corresponding <code>.txt</code> labels. This led to conflicting positive and negative samples for the same input during training, causing the model to become confused and prone to overfitting to noise.</p></li>\n<li><p><strong>Incorporation of negative samples</strong>:<br>\nWe added negative samples to balance the positive-to-negative sample ratio, which helps the model learn more robust decision boundaries.</p></li>\n<li><p><strong>Loss function enhancements</strong>:<br>\nWe introduced additional loss functions including <strong>Normalized Wasserstein Distance (NWD) Loss</strong> and <strong>IoU-Shape Loss</strong> to improve the model’s ability to detect small objects effectively.</p></li>\n<li><p><strong>Utilization of additional annotated data</strong>:<br>\nWe leveraged 36-box additional data shared by <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> and a subset of 24-box training data. Despite the higher noise level in this supplementary dataset, it contributed positively to model generalization.</p></li>\n<li><p><strong>TTA</strong>:<br>\nRotation during Inference makes stable performance.</p></li>\n</ul>\n<h3><strong>Limitations and Unresolved Issues</strong></h3>\n<ul>\n<li><p><strong>Lack of ensemble strategy</strong>:<br>\nDue to time constraints, we did not implement an ensemble strategy. Our final submission was based on a single YOLOv8m model.</p></li>\n<li><p><strong>2.5D DEiM model underperformance</strong>:<br>\nAlthough we experimented with a 2.5D DEiM variant that achieved nearly perfect recall (~1.0), its performance on the leaderboard was suboptimal, potentially due to a high false positive rate or limited generalization across test domains.</p></li>\n</ul>\n<h3>Summary</h3>\n<p>Despite the inherent variability in leaderboard rankings, innovative techniques—such as the rank-based thresholding approach in Bartley’s solution—demonstrate how thoughtful design can yield robust results. I am truly grateful for the insights and knowledge shared by the community throughout this competition.</p>",
      "rawMarkdown": "### Introduction \nFirst I want to thank Kaggle and the organizers of the competition for the baseline provided. \n\nOur final solution is a simple YOLOv8m model, the Place Solution—YOLOv8m. \n\nWe experimented with multiple image sizes (640, 960, 1024) and YOLO versions, and we finally decided to use YOLOv8m. We adopted several data augmentation techniques shared by @hengck23 , such as mosaic, mixup, perspective, shear, and auto augmentation. On top of that, CLAHE enhances local contrast and emphasizes fine-grained textures, thereby improving feature discrimination in low-contrast scenarios.\n\nHere's the full setting of our augmentations:\n```python\nT = [\n         A.RandomBrightnessContrast(p=0.2, brightness_limit=0.1, contrast_limit=0.1),            \n         A.RandomGamma(p=0.2, gamma_limit=(90, 110)),\n         A.HorizontalFlip(p=0.5),\n         A.VerticalFlip(p=0.2),\n         A.ShiftScaleRotate(shift_limit=0.02, scale_limit=0.1, rotate_limit=10, border_mode=0, p=0.3),\n         A.ImageCompression(quality_lower=90, quality_upper=100, p=0.1),\n         A.CLAHE(p=0.1, clip_limit=2.0),\n         A.ToGray(p=0.05),\n       ]\n```\n\nOur final solution employs a single YOLOv8m model without any ensemble, resulting in a relatively simple inference pipeline. However, we believe that this competition requires participants to focus on in-depth data analysis, particularly in identifying and addressing domain shifts arising from different instruments and acquisition devices. Moreover, it is crucial to design solutions that are robust to threshold variations in order to mitigate the risk of performance degradation on the private leaderboard.\n\nAccording to the competition's evaluation metric, detections within a relatively large radius are considered correct, and the F𝛽 score used in the competition tolerates a higher number of false positives. This design choice is of particular significance and warrants deeper reflection.\n\n In the context of medical image object detection, missing true targets is often more detrimental than detecting spurious ones. Models in such applications serve the purpose of narrowing down massive datasets into candidate regions that can be manually reviewed and annotated with higher precision. Ultimately, this human-in-the-loop approach ensures that the final annotations are more accurate. From the perspective of medical practitioners, especially medical students involved in annotation, it is highly desirable for models to achieve high recall while maintaining a reasonably acceptable precision. Such a balance reduces their workload by minimizing missed detections while keeping the number of false alarms manageable.\n\nThe key strategy come up from the simple fact I observed. We can iteratively use pseudo labelling to address the missing gt labels in the training data (I think it may be somehow just like the competition organizers they would do on the leaderborad data) . This is also the key reason my 2 yolov8m models can reach plb840 and lb822 (we selected a plb830;lb824 a single model though), which is possible to reach gold zone with only simple yolov8m model!! [840PLB model](https://www.kaggle.com/code/shanzhong8/22nd-solution-tta-yolo-ensemblemodel-plb840/edit)\n\n\n\n### **Key Improvements**\n\n* **Correction of data generation issues in the public notebook**:\n  We identified and resolved a problem where, in images containing multiple motors, the dataset generation process saved multiple YOLO-format annotations and corresponding `.txt` labels. This led to conflicting positive and negative samples for the same input during training, causing the model to become confused and prone to overfitting to noise.\n\n* **Incorporation of negative samples**:\n  We added negative samples to balance the positive-to-negative sample ratio, which helps the model learn more robust decision boundaries.\n\n* **Loss function enhancements**:\n  We introduced additional loss functions including **Normalized Wasserstein Distance (NWD) Loss** and **IoU-Shape Loss** to improve the model’s ability to detect small objects effectively.\n\n* **Utilization of additional annotated data**:\n  We leveraged 36-box additional data shared by @brendanartley and a subset of 24-box training data. Despite the higher noise level in this supplementary dataset, it contributed positively to model generalization.\n\n* **TTA**:\n  Rotation during Inference makes stable performance.\n\n### **Limitations and Unresolved Issues**\n\n* **Lack of ensemble strategy**:\n  Due to time constraints, we did not implement an ensemble strategy. Our final submission was based on a single YOLOv8m model.\n\n* **2.5D DEiM model underperformance**:\n  Although we experimented with a 2.5D DEiM variant that achieved nearly perfect recall (\\~1.0), its performance on the leaderboard was suboptimal, potentially due to a high false positive rate or limited generalization across test domains.\n\n\n### Summary\nDespite the inherent variability in leaderboard rankings, innovative techniques—such as the rank-based thresholding approach in Bartley’s solution—demonstrate how thoughtful design can yield robust results. I am truly grateful for the insights and knowledge shared by the community throughout this competition.",
      "votes": null
    },
    {
      "id": "3218184",
      "postDate": "06/06/2025 00:40:29",
      "content": "<p>may i know how you handle the data generation issue?</p>",
      "rawMarkdown": "may i know how you handle the data generation issue?",
      "votes": null
    },
    {
      "id": "3218185",
      "postDate": "06/06/2025 00:46:05",
      "content": "<p>Inference on the training dataset with multiple models.</p>",
      "rawMarkdown": "Inference on the training dataset with multiple models.",
      "votes": null
    },
    {
      "id": "3220290",
      "postDate": "06/09/2025 05:43:21",
      "content": "<blockquote>\n  <p>We identified and resolved a problem where, in images containing multiple motors, the dataset generation process saved multiple YOLO-format annotations and corresponding .txt labels.</p>\n</blockquote>\n<p>I noticed this early on, but put off dealing with it and forgot about it. What a surprise! But thanks for bringing it to my attention.</p>",
      "rawMarkdown": ">We identified and resolved a problem where, in images containing multiple motors, the dataset generation process saved multiple YOLO-format annotations and corresponding .txt labels.\n\nI noticed this early on, but put off dealing with it and forgot about it. What a surprise! But thanks for bringing it to my attention.",
      "votes": null
    },
    {
      "id": "3220292",
      "postDate": "06/09/2025 05:48:07",
      "content": "<p>Yeah, one of the keys to make YOLO training stable</p>",
      "rawMarkdown": "Yeah, one of the keys to make YOLO training stable",
      "votes": null
    },
    {
      "id": "3235424",
      "postDate": "06/29/2025 08:21:29",
      "content": "<p>Hello, where to change the yolo code to apply the augmentations and how to change the loss functions?</p>",
      "rawMarkdown": "Hello, where to change the yolo code to apply the augmentations and how to change the loss functions?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218184,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "06/06/2025 00:40:29",
      "content": "<p>may i know how you handle the data generation issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218185,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/06/2025 00:46:05",
          "content": "<p>Inference on the training dataset with multiple models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3220290,
      "author_name": "minfuka",
      "author_url": "",
      "post_date": "06/09/2025 05:43:21",
      "content": "<blockquote>\n  <p>We identified and resolved a problem where, in images containing multiple motors, the dataset generation process saved multiple YOLO-format annotations and corresponding .txt labels.</p>\n</blockquote>\n<p>I noticed this early on, but put off dealing with it and forgot about it. What a surprise! But thanks for bringing it to my attention.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3220292,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/09/2025 05:48:07",
          "content": "<p>Yeah, one of the keys to make YOLO training stable</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3235424,
      "author_name": "xyzxyzxyzxyzxz",
      "author_url": "",
      "post_date": "06/29/2025 08:21:29",
      "content": "<p>Hello, where to change the yolo code to apply the augmentations and how to change the loss functions?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218176": "### Introduction \nFirst I want to thank Kaggle and the organizers of the competition for the baseline provided. \n\nOur final solution is a simple YOLOv8m model, the Place Solution—YOLOv8m. \n\nWe experimented with multiple image sizes (640, 960, 1024) and YOLO versions, and we finally decided to use YOLOv8m. We adopted several data augmentation techniques shared by @hengck23 , such as mosaic, mixup, perspective, shear, and auto augmentation. On top of that, CLAHE enhances local contrast and emphasizes fine-grained textures, thereby improving feature discrimination in low-contrast scenarios.\n\nHere's the full setting of our augmentations:\n```python\nT = [\n         A.RandomBrightnessContrast(p=0.2, brightness_limit=0.1, contrast_limit=0.1),            \n         A.RandomGamma(p=0.2, gamma_limit=(90, 110)),\n         A.HorizontalFlip(p=0.5),\n         A.VerticalFlip(p=0.2),\n         A.ShiftScaleRotate(shift_limit=0.02, scale_limit=0.1, rotate_limit=10, border_mode=0, p=0.3),\n         A.ImageCompression(quality_lower=90, quality_upper=100, p=0.1),\n         A.CLAHE(p=0.1, clip_limit=2.0),\n         A.ToGray(p=0.05),\n       ]\n```\n\nOur final solution employs a single YOLOv8m model without any ensemble, resulting in a relatively simple inference pipeline. However, we believe that this competition requires participants to focus on in-depth data analysis, particularly in identifying and addressing domain shifts arising from different instruments and acquisition devices. Moreover, it is crucial to design solutions that are robust to threshold variations in order to mitigate the risk of performance degradation on the private leaderboard.\n\nAccording to the competition's evaluation metric, detections within a relatively large radius are considered correct, and the F𝛽 score used in the competition tolerates a higher number of false positives. This design choice is of particular significance and warrants deeper reflection.\n\n In the context of medical image object detection, missing true targets is often more detrimental than detecting spurious ones. Models in such applications serve the purpose of narrowing down massive datasets into candidate regions that can be manually reviewed and annotated with higher precision. Ultimately, this human-in-the-loop approach ensures that the final annotations are more accurate. From the perspective of medical practitioners, especially medical students involved in annotation, it is highly desirable for models to achieve high recall while maintaining a reasonably acceptable precision. Such a balance reduces their workload by minimizing missed detections while keeping the number of false alarms manageable.\n\nThe key strategy come up from the simple fact I observed. We can iteratively use pseudo labelling to address the missing gt labels in the training data (I think it may be somehow just like the competition organizers they would do on the leaderborad data) . This is also the key reason my 2 yolov8m models can reach plb840 and lb822 (we selected a plb830;lb824 a single model though), which is possible to reach gold zone with only simple yolov8m model!! [840PLB model](https://www.kaggle.com/code/shanzhong8/22nd-solution-tta-yolo-ensemblemodel-plb840/edit)\n\n\n\n### **Key Improvements**\n\n* **Correction of data generation issues in the public notebook**:\n  We identified and resolved a problem where, in images containing multiple motors, the dataset generation process saved multiple YOLO-format annotations and corresponding `.txt` labels. This led to conflicting positive and negative samples for the same input during training, causing the model to become confused and prone to overfitting to noise.\n\n* **Incorporation of negative samples**:\n  We added negative samples to balance the positive-to-negative sample ratio, which helps the model learn more robust decision boundaries.\n\n* **Loss function enhancements**:\n  We introduced additional loss functions including **Normalized Wasserstein Distance (NWD) Loss** and **IoU-Shape Loss** to improve the model’s ability to detect small objects effectively.\n\n* **Utilization of additional annotated data**:\n  We leveraged 36-box additional data shared by @brendanartley and a subset of 24-box training data. Despite the higher noise level in this supplementary dataset, it contributed positively to model generalization.\n\n* **TTA**:\n  Rotation during Inference makes stable performance.\n\n### **Limitations and Unresolved Issues**\n\n* **Lack of ensemble strategy**:\n  Due to time constraints, we did not implement an ensemble strategy. Our final submission was based on a single YOLOv8m model.\n\n* **2.5D DEiM model underperformance**:\n  Although we experimented with a 2.5D DEiM variant that achieved nearly perfect recall (\\~1.0), its performance on the leaderboard was suboptimal, potentially due to a high false positive rate or limited generalization across test domains.\n\n\n### Summary\nDespite the inherent variability in leaderboard rankings, innovative techniques—such as the rank-based thresholding approach in Bartley’s solution—demonstrate how thoughtful design can yield robust results. I am truly grateful for the insights and knowledge shared by the community throughout this competition.",
    "3218184": "may i know how you handle the data generation issue?",
    "3218185": "Inference on the training dataset with multiple models.",
    "3220290": ">We identified and resolved a problem where, in images containing multiple motors, the dataset generation process saved multiple YOLO-format annotations and corresponding .txt labels.\n\nI noticed this early on, but put off dealing with it and forgot about it. What a surprise! But thanks for bringing it to my attention.",
    "3220292": "Yeah, one of the keys to make YOLO training stable",
    "3235424": "Hello, where to change the yolo code to apply the augmentations and how to change the loss functions?"
  },
  "source": "meta"
}