{
  "id": 583242,
  "title": "[9 Place] Recall, Rotate and Zoom in",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/forcewithdonut-9-place-recall-rotate-and-zoom-in",
  "author_name": "",
  "post_date": "2025-06-05T15:51:07.096029800Z",
  "votes": 28,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Overview</h1>\n<p>We are thrilled to win a gold medal in this shake-up competition. Firstly, we'd like to express our gratitude to the organizers for hosting such an great game and for open-sourcing the entire workflow—from EDA to data processing, training, and submission. This significantly lowered the barrier to entry for us. We also want to thank all of the open-sourced authors for their contributions. Finally, a special thanks to my teammate, <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> . We poured much effort into this competition, and we're so happy about the rewarding outcome</p>\n<p>Overall, our inference pipeline consists of two stages. The first stage uses a very low threshold to recall as many candidates as possible, while the second stage applies an appropriate threshold to make the final decision.</p>\n<p>Regarding training, we coudn't find a stable way throughout the entire process. In the end, all models we used were either YOLOv8 or YOLO11 models trained simultaneously on both external data <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> and competition data, sharing the same training configuration.</p>\n<h2>Inference Pipeline</h2>\n<ol>\n<li><strong>Two-Stage Detection Pipeline:</strong> <ul>\n<li><strong>Stage 1 (Candidate Generation):</strong> Identifies initial potential motor locations. This stage uses standard model ensembles or SAHI-based ensembles.</li>\n<li><strong>Stage 2 (Candidate filter):</strong> Further processes and validates candidates from Stage 1.</li></ul></li>\n<li><strong>Ensemble Models:</strong> Multiple models or different configurations of the same model are used in combination to enhance detection accuracy and robustness, which applied in both Stage 1 and Stage 2.</li>\n<li><strong>Test-Time Augmentation (TTA):</strong> Since the test dataset has large variance in image size, we use multiple resolutions for inference to capture features at different scales.</li>\n<li><strong>Bypass Logic &amp; Midpoint Reasoning:</strong><ul>\n<li><strong>Bypass Logic:</strong> Skip Stage 2 if Stage 1 generates highly confident detections, which improving efficiency with no cost. The bypass threshold is set to 0.6. </li>\n<li><strong>Midpoint Reasoning:</strong> If top2 detections are very confident and close(either in Stage 1 or after Stage 2), a new detection point at their geometric midpoint with a slightly boosted confidence will be created returned as prediction point.</li></ul></li>\n<li><strong>SAHI:</strong> In some slides in Stage 1(if Stage1 doesn't bypass), we use SAHI method to devide large tomogram slices into non-overlapping patches for inference, then merges results to enhance multi-scale detection.</li>\n<li><strong>Rotation-based Refinement with zoom in in stage 2:</strong> Instead of rotation around the candidate, we crop a rotated zoomed-in square around the target, and pad the original image slice with the mean pixel value. </li>\n</ol>\n<hr>\n<h2>Model training</h2>\n<ol>\n<li>We trained yolov8l or yolov11l with ultralytics.</li>\n<li>All of the models are trained with a mixture of external data('trust' 3) and competition data('trust' 4).</li>\n<li>The external dataset is fixed by <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> .</li>\n<li>We trained some 'local' model by random cropping around target while training. These models are trained for sahi or stage2.</li>\n</ol>\n<h2>Submission Comparison</h2>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>A</th>\n<th>B</th>\n<th>C   (Selected)</th>\n<th>D(Selected)</th>\n<th>E</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Overall Scheme</strong></td>\n<td>Std+S2</td>\n<td>Parallel[Std+SAHI]+S2</td>\n<td>STD+S2</td>\n<td>Parallel[Std+SAHI]+S2</td>\n<td>STD+S2</td>\n</tr>\n<tr>\n<td><strong>Stage 1A Config</strong></td>\n<td>yolo8l <br> res:960/1280/832</td>\n<td>yolov8l no sam<br>yolo11l sam:1/4 res:960</td>\n<td>yolov8 no sam res:960 <br>yolov11 sam:1/2 res 960<br></td>\n<td>yolo8l no sam res:960</td>\n<td>yolov8l res:960</td>\n</tr>\n<tr>\n<td><strong>Stage 1B(SAHI)</strong></td>\n<td>N/A</td>\n<td>yolo8l<br>sam:1/3<br>patch:768</td>\n<td>N/A</td>\n<td>yolo11-cz model<br>sam:1/4<br>patch:640</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td><strong>Stage 2</strong></td>\n<td>yolo8l z1.5 res512 <br> yolo11-cz z1.5 res640 <br> yolo8l-cz z2 res640</td>\n<td>yolo8l z1.5 res512 <br> yolo11-cz z3 res640 <br> yolo8l-cz z2 res640</td>\n<td>yolo8l z1.5 res512 <br> yolo11-cz z1.5 res512 <br> yolo8l-cz z2 res640</td>\n<td>yolo8l z1.5 res512</td>\n<td>yolo8l z1.5 res512</td>\n</tr>\n<tr>\n<td><strong>Public Score</strong></td>\n<td>0.837</td>\n<td>0.853</td>\n<td>0.856</td>\n<td>0.856</td>\n<td>0.860</td>\n</tr>\n<tr>\n<td><strong>Private Score</strong></td>\n<td>0.853</td>\n<td>0.855</td>\n<td><strong>0.852</strong></td>\n<td>0.832</td>\n<td>0.826</td>\n</tr>\n</tbody>\n</table>\n<h3>Configuration Notes</h3>\n<ul>\n<li><code>Standard</code>: <a href=\"https://www.kaggle.com/code/andrewjdarley/submission-notebook\" target=\"_blank\">Public inferece pipeline</a>.</li>\n<li><code>sahi</code>: <a href=\"https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook\" target=\"_blank\">Sahi inference pipeline</a>.</li>\n<li><code>S2</code>: Stage2 inference. Infernce on zoomin+rotate+crop images, with almost same code as <code>Standard</code>.</li>\n<li><code>sam{}</code>: The sample rate applied to the SAHI processing</li>\n<li><code>no sam</code>: No sampling, meaning taking all of the slices.</li>\n<li><code>max_image_side</code>: Maximum dimension allowed for input images before slicing</li>\n<li><code>patches_per_side</code>: Number of grid divisions (n x n) for each slice</li>\n<li><code>patch_size</code>: Target dimension for each patch (e.g., 768x768 pixels)</li>\n<li><code>res</code>: image resolution.</li>\n<li><code>-cz z{1.5/2/3}</code>: Means this is a 'local' model, 'cz' means crop and zoomin. And <code>z{}</code> indicates the zoom-in scale while inference.  </li>\n</ul>\n<h2>End</h2>\n<p>Any questions or suggestions are welcome. Happy kaggling!</p>",
  "messages": [
    {
      "id": "3217969",
      "postDate": "06/05/2025 15:51:07",
      "content": "<h1>Overview</h1>\n<p>We are thrilled to win a gold medal in this shake-up competition. Firstly, we'd like to express our gratitude to the organizers for hosting such an great game and for open-sourcing the entire workflow—from EDA to data processing, training, and submission. This significantly lowered the barrier to entry for us. We also want to thank all of the open-sourced authors for their contributions. Finally, a special thanks to my teammate, <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> . We poured much effort into this competition, and we're so happy about the rewarding outcome</p>\n<p>Overall, our inference pipeline consists of two stages. The first stage uses a very low threshold to recall as many candidates as possible, while the second stage applies an appropriate threshold to make the final decision.</p>\n<p>Regarding training, we coudn't find a stable way throughout the entire process. In the end, all models we used were either YOLOv8 or YOLO11 models trained simultaneously on both external data <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> and competition data, sharing the same training configuration.</p>\n<h2>Inference Pipeline</h2>\n<ol>\n<li><strong>Two-Stage Detection Pipeline:</strong> <ul>\n<li><strong>Stage 1 (Candidate Generation):</strong> Identifies initial potential motor locations. This stage uses standard model ensembles or SAHI-based ensembles.</li>\n<li><strong>Stage 2 (Candidate filter):</strong> Further processes and validates candidates from Stage 1.</li></ul></li>\n<li><strong>Ensemble Models:</strong> Multiple models or different configurations of the same model are used in combination to enhance detection accuracy and robustness, which applied in both Stage 1 and Stage 2.</li>\n<li><strong>Test-Time Augmentation (TTA):</strong> Since the test dataset has large variance in image size, we use multiple resolutions for inference to capture features at different scales.</li>\n<li><strong>Bypass Logic &amp; Midpoint Reasoning:</strong><ul>\n<li><strong>Bypass Logic:</strong> Skip Stage 2 if Stage 1 generates highly confident detections, which improving efficiency with no cost. The bypass threshold is set to 0.6. </li>\n<li><strong>Midpoint Reasoning:</strong> If top2 detections are very confident and close(either in Stage 1 or after Stage 2), a new detection point at their geometric midpoint with a slightly boosted confidence will be created returned as prediction point.</li></ul></li>\n<li><strong>SAHI:</strong> In some slides in Stage 1(if Stage1 doesn't bypass), we use SAHI method to devide large tomogram slices into non-overlapping patches for inference, then merges results to enhance multi-scale detection.</li>\n<li><strong>Rotation-based Refinement with zoom in in stage 2:</strong> Instead of rotation around the candidate, we crop a rotated zoomed-in square around the target, and pad the original image slice with the mean pixel value. </li>\n</ol>\n<hr>\n<h2>Model training</h2>\n<ol>\n<li>We trained yolov8l or yolov11l with ultralytics.</li>\n<li>All of the models are trained with a mixture of external data('trust' 3) and competition data('trust' 4).</li>\n<li>The external dataset is fixed by <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> .</li>\n<li>We trained some 'local' model by random cropping around target while training. These models are trained for sahi or stage2.</li>\n</ol>\n<h2>Submission Comparison</h2>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>A</th>\n<th>B</th>\n<th>C   (Selected)</th>\n<th>D(Selected)</th>\n<th>E</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Overall Scheme</strong></td>\n<td>Std+S2</td>\n<td>Parallel[Std+SAHI]+S2</td>\n<td>STD+S2</td>\n<td>Parallel[Std+SAHI]+S2</td>\n<td>STD+S2</td>\n</tr>\n<tr>\n<td><strong>Stage 1A Config</strong></td>\n<td>yolo8l <br> res:960/1280/832</td>\n<td>yolov8l no sam<br>yolo11l sam:1/4 res:960</td>\n<td>yolov8 no sam res:960 <br>yolov11 sam:1/2 res 960<br></td>\n<td>yolo8l no sam res:960</td>\n<td>yolov8l res:960</td>\n</tr>\n<tr>\n<td><strong>Stage 1B(SAHI)</strong></td>\n<td>N/A</td>\n<td>yolo8l<br>sam:1/3<br>patch:768</td>\n<td>N/A</td>\n<td>yolo11-cz model<br>sam:1/4<br>patch:640</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td><strong>Stage 2</strong></td>\n<td>yolo8l z1.5 res512 <br> yolo11-cz z1.5 res640 <br> yolo8l-cz z2 res640</td>\n<td>yolo8l z1.5 res512 <br> yolo11-cz z3 res640 <br> yolo8l-cz z2 res640</td>\n<td>yolo8l z1.5 res512 <br> yolo11-cz z1.5 res512 <br> yolo8l-cz z2 res640</td>\n<td>yolo8l z1.5 res512</td>\n<td>yolo8l z1.5 res512</td>\n</tr>\n<tr>\n<td><strong>Public Score</strong></td>\n<td>0.837</td>\n<td>0.853</td>\n<td>0.856</td>\n<td>0.856</td>\n<td>0.860</td>\n</tr>\n<tr>\n<td><strong>Private Score</strong></td>\n<td>0.853</td>\n<td>0.855</td>\n<td><strong>0.852</strong></td>\n<td>0.832</td>\n<td>0.826</td>\n</tr>\n</tbody>\n</table>\n<h3>Configuration Notes</h3>\n<ul>\n<li><code>Standard</code>: <a href=\"https://www.kaggle.com/code/andrewjdarley/submission-notebook\" target=\"_blank\">Public inferece pipeline</a>.</li>\n<li><code>sahi</code>: <a href=\"https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook\" target=\"_blank\">Sahi inference pipeline</a>.</li>\n<li><code>S2</code>: Stage2 inference. Infernce on zoomin+rotate+crop images, with almost same code as <code>Standard</code>.</li>\n<li><code>sam{}</code>: The sample rate applied to the SAHI processing</li>\n<li><code>no sam</code>: No sampling, meaning taking all of the slices.</li>\n<li><code>max_image_side</code>: Maximum dimension allowed for input images before slicing</li>\n<li><code>patches_per_side</code>: Number of grid divisions (n x n) for each slice</li>\n<li><code>patch_size</code>: Target dimension for each patch (e.g., 768x768 pixels)</li>\n<li><code>res</code>: image resolution.</li>\n<li><code>-cz z{1.5/2/3}</code>: Means this is a 'local' model, 'cz' means crop and zoomin. And <code>z{}</code> indicates the zoom-in scale while inference.  </li>\n</ul>\n<h2>End</h2>\n<p>Any questions or suggestions are welcome. Happy kaggling!</p>",
      "rawMarkdown": "# Overview\nWe are thrilled to win a gold medal in this shake-up competition. Firstly, we'd like to express our gratitude to the organizers for hosting such an great game and for open-sourcing the entire workflow—from EDA to data processing, training, and submission. This significantly lowered the barrier to entry for us. We also want to thank all of the open-sourced authors for their contributions. Finally, a special thanks to my teammate, @forcewithme . We poured much effort into this competition, and we're so happy about the rewarding outcome\n\nOverall, our inference pipeline consists of two stages. The first stage uses a very low threshold to recall as many candidates as possible, while the second stage applies an appropriate threshold to make the final decision.\n\nRegarding training, we coudn't find a stable way throughout the entire process. In the end, all models we used were either YOLOv8 or YOLO11 models trained simultaneously on both external data @brendanartley and competition data, sharing the same training configuration.\n\n## Inference Pipeline\n\n1.  **Two-Stage Detection Pipeline:** \n    *   **Stage 1 (Candidate Generation):** Identifies initial potential motor locations. This stage uses standard model ensembles or SAHI-based ensembles.\n    *   **Stage 2 (Candidate filter):** Further processes and validates candidates from Stage 1.\n2.  **Ensemble Models:** Multiple models or different configurations of the same model are used in combination to enhance detection accuracy and robustness, which applied in both Stage 1 and Stage 2.\n3.  **Test-Time Augmentation (TTA):** Since the test dataset has large variance in image size, we use multiple resolutions for inference to capture features at different scales.\n4.  **Bypass Logic & Midpoint Reasoning:**\n    *   **Bypass Logic:** Skip Stage 2 if Stage 1 generates highly confident detections, which improving efficiency with no cost. The bypass threshold is set to 0.6. \n    *   **Midpoint Reasoning:** If top2 detections are very confident and close(either in Stage 1 or after Stage 2), a new detection point at their geometric midpoint with a slightly boosted confidence will be created returned as prediction point.\n5.  **SAHI:** In some slides in Stage 1(if Stage1 doesn't bypass), we use SAHI method to devide large tomogram slices into non-overlapping patches for inference, then merges results to enhance multi-scale detection.\n6.  **Rotation-based Refinement with zoom in in stage 2:** Instead of rotation around the candidate, we crop a rotated zoomed-in square around the target, and pad the original image slice with the mean pixel value. \n\n---\n\n## Model training\n\n1. We trained yolov8l or yolov11l with ultralytics.\n2. All of the models are trained with a mixture of external data('trust' 3) and competition data('trust' 4).\n3. The external dataset is fixed by @tatamikenn .\n4. We trained some 'local' model by random cropping around target while training. These models are trained for sahi or stage2.\n\n## Submission Comparison\n\n| Submission           | A  | B   | C   (Selected)         | D(Selected)            |E|\n|:-------------------|:------------|:-------------|:-------------|:-------------|:-------------|\n| **Overall Scheme** | Std+S2 | Parallel[Std+SAHI]+S2 | STD+S2 | Parallel[Std+SAHI]+S2 |STD+S2\n| **Stage 1A Config**| yolo8l <br/> res:960/1280/832 | yolov8l no sam<br/>yolo11l sam:1/4 res:960 | yolov8 no sam res:960 <br/>yolov11 sam:1/2 res 960<br/>  | yolo8l no sam res:960 | yolov8l res:960 |\n| **Stage 1B(SAHI)** | N/A | yolo8l<br/>sam:1/3<br/>patch:768 | N/A | yolo11-cz model<br/>sam:1/4<br/>patch:640 | N/A |\n| **Stage 2** |  yolo8l z1.5 res512 <br/> yolo11-cz z1.5 res640 <br/> yolo8l-cz z2 res640 | yolo8l z1.5 res512 <br/> yolo11-cz z3 res640 <br/> yolo8l-cz z2 res640 | yolo8l z1.5 res512 <br/> yolo11-cz z1.5 res512 <br/> yolo8l-cz z2 res640| yolo8l z1.5 res512 | yolo8l z1.5 res512 |\n| **Public Score**   | 0.837 | 0.853 | 0.856 | 0.856 | 0.860 |\n| **Private Score**  | 0.853 | 0.855 | **0.852** | 0.832 | 0.826 |\n\n\n### Configuration Notes\n  - `Standard`: [Public inferece pipeline](https://www.kaggle.com/code/andrewjdarley/submission-notebook).\n  - `sahi`: [Sahi inference pipeline](https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook).\n  - `S2`: Stage2 inference. Infernce on zoomin+rotate+crop images, with almost same code as `Standard`.\n  - `sam{}`: The sample rate applied to the SAHI processing\n  - `no sam`: No sampling, meaning taking all of the slices.\n  - `max_image_side`: Maximum dimension allowed for input images before slicing\n  - `patches_per_side`: Number of grid divisions (n x n) for each slice\n  - `patch_size`: Target dimension for each patch (e.g., 768x768 pixels)\n  - `res`: image resolution.\n  - `-cz z{1.5/2/3}`: Means this is a 'local' model, 'cz' means crop and zoomin. And `z{}` indicates the zoom-in scale while inference.  \n\n## End\nAny questions or suggestions are welcome. Happy kaggling!",
      "votes": null
    },
    {
      "id": "3218407",
      "postDate": "06/06/2025 06:33:58",
      "content": "<p><a href=\"https://www.kaggle.com/doonut\" target=\"_blank\">@doonut</a> Congratulations on the gold medal, and thanks for the mention!</p>\n<p>I have a couple of questions:</p>\n<ol>\n<li>How many candidate detections are passed to stage 2?</li>\n<li>How is the \"trust\" level of the dataset utilized in your pipeline?</li>\n</ol>",
      "rawMarkdown": "doonut Congratulations on the gold medal, and thanks for the mention!\n\nI have a couple of questions:\n1. How many candidate detections are passed to stage 2?\n2. How is the \"trust\" level of the dataset utilized in your pipeline?",
      "votes": null
    },
    {
      "id": "3218514",
      "postDate": "06/06/2025 09:41:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> ,</p>\n<blockquote>\n  <p>Q: How many candidate detections are passed to stage 2?<br>\n  A: The number of candidates varies in different submission. But overall it is in range(3,6).&gt;</p>\n  <p>Q: How is the \"trust\" level of the dataset utilized in your pipeline?<br>\n  A: Do you mean the \"trust\" level in the stage 2 of inference pipeline？ We didn't set a <code>trust</code> in inference pipeline.</p>\n</blockquote>",
      "rawMarkdown": "Hi @tatamikenn ,\n\n> Q: How many candidate detections are passed to stage 2?\nA: The number of candidates varies in different submission. But overall it is in range(3,6).>\n\n> Q: How is the \"trust\" level of the dataset utilized in your pipeline?\nA: Do you mean the \"trust\" level in the stage 2 of inference pipeline？ We didn't set a `trust` in inference pipeline.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218407,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "06/06/2025 06:33:58",
      "content": "<p><a href=\"https://www.kaggle.com/doonut\" target=\"_blank\">@doonut</a> Congratulations on the gold medal, and thanks for the mention!</p>\n<p>I have a couple of questions:</p>\n<ol>\n<li>How many candidate detections are passed to stage 2?</li>\n<li>How is the \"trust\" level of the dataset utilized in your pipeline?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 3218514,
          "author_name": "forcewithme",
          "author_url": "",
          "post_date": "06/06/2025 09:41:55",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> ,</p>\n<blockquote>\n  <p>Q: How many candidate detections are passed to stage 2?<br>\n  A: The number of candidates varies in different submission. But overall it is in range(3,6).&gt;</p>\n  <p>Q: How is the \"trust\" level of the dataset utilized in your pipeline?<br>\n  A: Do you mean the \"trust\" level in the stage 2 of inference pipeline？ We didn't set a <code>trust</code> in inference pipeline.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3217969": "# Overview\nWe are thrilled to win a gold medal in this shake-up competition. Firstly, we'd like to express our gratitude to the organizers for hosting such an great game and for open-sourcing the entire workflow—from EDA to data processing, training, and submission. This significantly lowered the barrier to entry for us. We also want to thank all of the open-sourced authors for their contributions. Finally, a special thanks to my teammate, @forcewithme . We poured much effort into this competition, and we're so happy about the rewarding outcome\n\nOverall, our inference pipeline consists of two stages. The first stage uses a very low threshold to recall as many candidates as possible, while the second stage applies an appropriate threshold to make the final decision.\n\nRegarding training, we coudn't find a stable way throughout the entire process. In the end, all models we used were either YOLOv8 or YOLO11 models trained simultaneously on both external data @brendanartley and competition data, sharing the same training configuration.\n\n## Inference Pipeline\n\n1.  **Two-Stage Detection Pipeline:** \n    *   **Stage 1 (Candidate Generation):** Identifies initial potential motor locations. This stage uses standard model ensembles or SAHI-based ensembles.\n    *   **Stage 2 (Candidate filter):** Further processes and validates candidates from Stage 1.\n2.  **Ensemble Models:** Multiple models or different configurations of the same model are used in combination to enhance detection accuracy and robustness, which applied in both Stage 1 and Stage 2.\n3.  **Test-Time Augmentation (TTA):** Since the test dataset has large variance in image size, we use multiple resolutions for inference to capture features at different scales.\n4.  **Bypass Logic & Midpoint Reasoning:**\n    *   **Bypass Logic:** Skip Stage 2 if Stage 1 generates highly confident detections, which improving efficiency with no cost. The bypass threshold is set to 0.6. \n    *   **Midpoint Reasoning:** If top2 detections are very confident and close(either in Stage 1 or after Stage 2), a new detection point at their geometric midpoint with a slightly boosted confidence will be created returned as prediction point.\n5.  **SAHI:** In some slides in Stage 1(if Stage1 doesn't bypass), we use SAHI method to devide large tomogram slices into non-overlapping patches for inference, then merges results to enhance multi-scale detection.\n6.  **Rotation-based Refinement with zoom in in stage 2:** Instead of rotation around the candidate, we crop a rotated zoomed-in square around the target, and pad the original image slice with the mean pixel value. \n\n---\n\n## Model training\n\n1. We trained yolov8l or yolov11l with ultralytics.\n2. All of the models are trained with a mixture of external data('trust' 3) and competition data('trust' 4).\n3. The external dataset is fixed by @tatamikenn .\n4. We trained some 'local' model by random cropping around target while training. These models are trained for sahi or stage2.\n\n## Submission Comparison\n\n| Submission           | A  | B   | C   (Selected)         | D(Selected)            |E|\n|:-------------------|:------------|:-------------|:-------------|:-------------|:-------------|\n| **Overall Scheme** | Std+S2 | Parallel[Std+SAHI]+S2 | STD+S2 | Parallel[Std+SAHI]+S2 |STD+S2\n| **Stage 1A Config**| yolo8l <br/> res:960/1280/832 | yolov8l no sam<br/>yolo11l sam:1/4 res:960 | yolov8 no sam res:960 <br/>yolov11 sam:1/2 res 960<br/>  | yolo8l no sam res:960 | yolov8l res:960 |\n| **Stage 1B(SAHI)** | N/A | yolo8l<br/>sam:1/3<br/>patch:768 | N/A | yolo11-cz model<br/>sam:1/4<br/>patch:640 | N/A |\n| **Stage 2** |  yolo8l z1.5 res512 <br/> yolo11-cz z1.5 res640 <br/> yolo8l-cz z2 res640 | yolo8l z1.5 res512 <br/> yolo11-cz z3 res640 <br/> yolo8l-cz z2 res640 | yolo8l z1.5 res512 <br/> yolo11-cz z1.5 res512 <br/> yolo8l-cz z2 res640| yolo8l z1.5 res512 | yolo8l z1.5 res512 |\n| **Public Score**   | 0.837 | 0.853 | 0.856 | 0.856 | 0.860 |\n| **Private Score**  | 0.853 | 0.855 | **0.852** | 0.832 | 0.826 |\n\n\n### Configuration Notes\n  - `Standard`: [Public inferece pipeline](https://www.kaggle.com/code/andrewjdarley/submission-notebook).\n  - `sahi`: [Sahi inference pipeline](https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook).\n  - `S2`: Stage2 inference. Infernce on zoomin+rotate+crop images, with almost same code as `Standard`.\n  - `sam{}`: The sample rate applied to the SAHI processing\n  - `no sam`: No sampling, meaning taking all of the slices.\n  - `max_image_side`: Maximum dimension allowed for input images before slicing\n  - `patches_per_side`: Number of grid divisions (n x n) for each slice\n  - `patch_size`: Target dimension for each patch (e.g., 768x768 pixels)\n  - `res`: image resolution.\n  - `-cz z{1.5/2/3}`: Means this is a 'local' model, 'cz' means crop and zoomin. And `z{}` indicates the zoom-in scale while inference.  \n\n## End\nAny questions or suggestions are welcome. Happy kaggling!",
    "3218407": "doonut Congratulations on the gold medal, and thanks for the mention!\n\nI have a couple of questions:\n1. How many candidate detections are passed to stage 2?\n2. How is the \"trust\" level of the dataset utilized in your pipeline?",
    "3218514": "Hi @tatamikenn ,\n\n> Q: How many candidate detections are passed to stage 2?\nA: The number of candidates varies in different submission. But overall it is in range(3,6).>\n\n> Q: How is the \"trust\" level of the dataset utilized in your pipeline?\nA: Do you mean the \"trust\" level in the stage 2 of inference pipeline？ We didn't set a `trust` in inference pipeline."
  },
  "source": "meta"
}