{
  "id": 697809,
  "title": "2nd Place Solution/SIFT Maching",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/697809",
  "author_name": "shiba-inu",
  "post_date": "2026-05-07T13:25:40.543000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Recod.ai/LUC - Scientific Image Forgery Detection 2nd Place Solution</h1>\n<p>First of all, thank you to the organizers for hosting such a wonderful competition. It was a highly meaningful theme—protecting the integrity of scientific research—and I learned a great deal from it.</p>\n<p>Result: <strong>Public 3rd → Private 2nd</strong> (out of 1564 teams)</p>\n<h2>Solution Overview</h2>\n<p>Rather than a deep learning-based approach, my solution is built around <strong>classical feature matching (SIFT + RANSAC)</strong>, combined with YOLO-based preprocessing and custom post-processing in a 3-stage pipeline.</p>\n<pre><code>Stage 1: Valid Region Detection (Panel/Text detection with YOLOv8)\n    ↓\nStage 2: Feature Matching (SIFT + G2NN + RANSAC)\n    ↓\nStage 3: Post-processing (Cluster merging + High-precision mask generation)\n</code></pre>\n<p>The overall algorithm overview is shown in the image below.</p>\n<hr>\n<h2>Stage 1: Valid Region Detection</h2>\n<p>In the subsequent feature matching stage, <strong>graph regions</strong> produce a large number of false positives, and <strong>text regions</strong> cause erroneous matching on caption parts (e.g., matching on the \"mm\" portion of \"10 mm\"). These lead to significant accuracy degradation, making their exclusion essential.</p>\n<p>I trained a <strong>Panel detection model</strong> and a <strong>Text detection model</strong> using YOLOv8-m to address this problem.</p>\n<h3>Panel Detection</h3>\n<p>To exclude irrelevant regions such as graphs, only relevant areas (image regions) are detected. Simultaneously, YOLO classifies detections into 3 classes: <strong>corn images</strong>, <strong>Western blot protein images</strong>, and <strong>other biological images</strong>. The reason for separating classes is that the optimal RANSAC inlier threshold differs by image type (described below).</p>\n<p>Since no suitable training data existed, I addressed this by <strong>automatically generating synthetic data</strong>.</p>\n<ul>\n<li>Used a subset of 2,000 images from external data (BioFors dataset) due to time constraints</li>\n<li>Classified and extracted cell/plant images from the competition's authentic images</li>\n<li>Auto-generated synthetic graphs (scatter plots, line charts, heatmaps, violin plots)</li>\n<li>Randomly arranged the above in various layout patterns (1×1, 2×2, 3×1, etc.) to create training data</li>\n<li>Ensembled 5 YOLOv8-m models trained with different seeds using <strong>Weighted Boxes Fusion (WBF)</strong></li>\n</ul>\n<h3>Text Detection</h3>\n<p>Existing OCR solutions (EasyOCR, etc.) were insufficiently accurate—for example, misrecognizing round cells as the letter \"O\". Since text in scientific paper images has minimal distortion, I determined that OCR-specific heads were unnecessary and <strong>trained YOLOv8-m as an object detector</strong>.</p>\n<ul>\n<li>Manually corrected low-threshold EasyOCR detection results to create annotations (~800 images)</li>\n<li>Added synthetic data (~2,000 images) with randomly placed text on text-free images</li>\n<li>Applied diversity in background color, text color, font, rotation, and size for augmentation</li>\n<li>At inference, performed TTA with two image sizes (640, 1280) and merged both results</li>\n</ul>\n<hr>\n<h2>Stage 2: Feature Matching</h2>\n<h3>Aggressive SIFT Keypoint Extraction</h3>\n<p>Biological images have very little texture, and default settings fail to extract sufficient keypoints.</p>\n<ul>\n<li>Used <code>contrast_threshold=0.001</code> (approximately 1/40 of the default 0.04) to extract a large number of keypoints</li>\n<li>For small images (&lt;1024px), upscaled 4× before feature extraction, then mapped coordinates back to original scale</li>\n</ul>\n<h3>G2NN Matching</h3>\n<p>The absolute L2 distance values vary greatly across images, making fixed thresholds impractical.</p>\n<p>Adopted <strong>G2NN (Good-to-Next Neighbor)</strong>: sort matching distances and accept only matches up to the breakpoint where $T_{n+1}/T_n &gt; \\alpha$ ($\\alpha=0.7$). This achieves threshold-independent matching.</p>\n<h3>Speedup: KDTree</h3>\n<p>Due to the enormous number of keypoints, used FLANN KDTree approximate nearest neighbor search to reduce complexity from $O(N^2)$ to $O(N \\log N)$.</p>\n<h3>RANSAC + Geometric Filtering</h3>\n<p>Here, matching features are detected based on keypoints within the valid regions detected in Stage 1. Two sampling strategies were employed for detection:</p>\n<ul>\n<li><strong>Local sampling</strong>: Running RANSAC on all keypoints extracted from the entire image would require enormous computation time. Instead, keypoints within local regions are sampled from all extracted keypoints, and RANSAC is executed on these subsets. Regions are sampled in order of keypoint density, repeating until the number of keypoints falls below a certain threshold. The H matrix is then analyzed to filter out unrealistic transformations (scale &gt;4×, shear deformation, etc.)</li>\n<li><strong>Global sampling</strong>: Execute RANSAC once on all matches, then apply HDBSCAN clustering to remove false matches</li>\n</ul>\n<p>Additionally, RANSAC inlier thresholds were adjusted per image type based on the 3 classes detected in Stage 1:</p>\n<ul>\n<li><strong>Corn images</strong>: Fine black-and-white line patterns cause frequent false matches, so the inlier threshold was set <strong>high</strong> for strict filtering</li>\n<li><strong>Protein images (Western blot)</strong>: Very few features make matching inherently difficult, so the inlier threshold was set <strong>low</strong> to enable detection even with few matches</li>\n<li><strong>Other biological images</strong>: An intermediate threshold was used</li>\n</ul>\n<hr>\n<h2>Stage 3: Post-processing</h2>\n<p>After feature matching, I discovered that <strong>careful clustering is essential</strong>. Without it, the F1 score (the competition metric) drops significantly. To achieve high-precision clustering, I implemented the following processing.</p>\n<h3>Two-stage Cluster Merging</h3>\n<ol>\n<li><strong>Stage 1 (within same pair)</strong>: Merge nearby clusters with similar H matrices (normalized difference of rotation angle, scale, and translation &lt; 0.02) using Union-Find</li>\n<li><strong>Stage 2 (across different pairs)</strong>: Merge as the same instance when IoMin (Intersection / min(area1, area2)) between source/dest bboxes ≥ 0.3</li>\n</ol>\n<h3>High-precision Mask Generation</h3>\n<ol>\n<li>Recompute H matrix from all corresponding points of merged clusters</li>\n<li>Generate source/dest masks via convex hull + dilation</li>\n<li>Refine masks precisely using H matrix transformation error (within mean + 2σ)</li>\n<li>Apply bidirectional refinement using inverse transformation</li>\n</ol>\n<hr>\n<h2>What Didn't Work</h2>\n<ul>\n<li><strong>Other feature descriptors</strong> (SURF, ORB, AKAZE, SuperPoint): Could not extract sufficient keypoints in the biological image domain, resulting in poor accuracy</li>\n<li><strong>Deep learning-based approaches</strong>: Tried CMFD models combining DINOv2/SegFormer backbones with self-correlation modules, but accuracy was insufficient</li>\n</ul>\n<hr>\n<h2>Summary</h2>\n<table>\n<thead>\n<tr>\n<th>Component</th>\n<th>Details</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Panel Detection</td>\n<td>YOLOv8-m × 5 (WBF ensemble), trained on synthetic data</td>\n</tr>\n<tr>\n<td>Text Detection</td>\n<td>YOLOv8-m, manual annotation + synthetic data</td>\n</tr>\n<tr>\n<td>Features</td>\n<td>SIFT (contrast_threshold=0.001)</td>\n</tr>\n<tr>\n<td>Matching</td>\n<td>G2NN (α=0.7) + FLANN KDTree</td>\n</tr>\n<tr>\n<td>Geometric Estimation</td>\n<td>Affine RANSAC + H matrix filtering</td>\n</tr>\n<tr>\n<td>Clustering</td>\n<td>HDBSCAN + 2-stage merging (H similarity → IoMin)</td>\n</tr>\n<tr>\n<td>Mask Refinement</td>\n<td>Convex hull + H transformation error-based refinement</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 3454572,
      "postDate": "2026-05-07T13:25:40.543Z",
      "content": "<h1>Recod.ai/LUC - Scientific Image Forgery Detection 2nd Place Solution</h1>\n<p>First of all, thank you to the organizers for hosting such a wonderful competition. It was a highly meaningful theme—protecting the integrity of scientific research—and I learned a great deal from it.</p>\n<p>Result: <strong>Public 3rd → Private 2nd</strong> (out of 1564 teams)</p>\n<h2>Solution Overview</h2>\n<p>Rather than a deep learning-based approach, my solution is built around <strong>classical feature matching (SIFT + RANSAC)</strong>, combined with YOLO-based preprocessing and custom post-processing in a 3-stage pipeline.</p>\n<pre><code>Stage 1: Valid Region Detection (Panel/Text detection with YOLOv8)\n    ↓\nStage 2: Feature Matching (SIFT + G2NN + RANSAC)\n    ↓\nStage 3: Post-processing (Cluster merging + High-precision mask generation)\n</code></pre>\n<p>The overall algorithm overview is shown in the image below.</p>\n<hr>\n<h2>Stage 1: Valid Region Detection</h2>\n<p>In the subsequent feature matching stage, <strong>graph regions</strong> produce a large number of false positives, and <strong>text regions</strong> cause erroneous matching on caption parts (e.g., matching on the \"mm\" portion of \"10 mm\"). These lead to significant accuracy degradation, making their exclusion essential.</p>\n<p>I trained a <strong>Panel detection model</strong> and a <strong>Text detection model</strong> using YOLOv8-m to address this problem.</p>\n<h3>Panel Detection</h3>\n<p>To exclude irrelevant regions such as graphs, only relevant areas (image regions) are detected. Simultaneously, YOLO classifies detections into 3 classes: <strong>corn images</strong>, <strong>Western blot protein images</strong>, and <strong>other biological images</strong>. The reason for separating classes is that the optimal RANSAC inlier threshold differs by image type (described below).</p>\n<p>Since no suitable training data existed, I addressed this by <strong>automatically generating synthetic data</strong>.</p>\n<ul>\n<li>Used a subset of 2,000 images from external data (BioFors dataset) due to time constraints</li>\n<li>Classified and extracted cell/plant images from the competition's authentic images</li>\n<li>Auto-generated synthetic graphs (scatter plots, line charts, heatmaps, violin plots)</li>\n<li>Randomly arranged the above in various layout patterns (1×1, 2×2, 3×1, etc.) to create training data</li>\n<li>Ensembled 5 YOLOv8-m models trained with different seeds using <strong>Weighted Boxes Fusion (WBF)</strong></li>\n</ul>\n<h3>Text Detection</h3>\n<p>Existing OCR solutions (EasyOCR, etc.) were insufficiently accurate—for example, misrecognizing round cells as the letter \"O\". Since text in scientific paper images has minimal distortion, I determined that OCR-specific heads were unnecessary and <strong>trained YOLOv8-m as an object detector</strong>.</p>\n<ul>\n<li>Manually corrected low-threshold EasyOCR detection results to create annotations (~800 images)</li>\n<li>Added synthetic data (~2,000 images) with randomly placed text on text-free images</li>\n<li>Applied diversity in background color, text color, font, rotation, and size for augmentation</li>\n<li>At inference, performed TTA with two image sizes (640, 1280) and merged both results</li>\n</ul>\n<hr>\n<h2>Stage 2: Feature Matching</h2>\n<h3>Aggressive SIFT Keypoint Extraction</h3>\n<p>Biological images have very little texture, and default settings fail to extract sufficient keypoints.</p>\n<ul>\n<li>Used <code>contrast_threshold=0.001</code> (approximately 1/40 of the default 0.04) to extract a large number of keypoints</li>\n<li>For small images (&lt;1024px), upscaled 4× before feature extraction, then mapped coordinates back to original scale</li>\n</ul>\n<h3>G2NN Matching</h3>\n<p>The absolute L2 distance values vary greatly across images, making fixed thresholds impractical.</p>\n<p>Adopted <strong>G2NN (Good-to-Next Neighbor)</strong>: sort matching distances and accept only matches up to the breakpoint where $T_{n+1}/T_n &gt; \\alpha$ ($\\alpha=0.7$). This achieves threshold-independent matching.</p>\n<h3>Speedup: KDTree</h3>\n<p>Due to the enormous number of keypoints, used FLANN KDTree approximate nearest neighbor search to reduce complexity from $O(N^2)$ to $O(N \\log N)$.</p>\n<h3>RANSAC + Geometric Filtering</h3>\n<p>Here, matching features are detected based on keypoints within the valid regions detected in Stage 1. Two sampling strategies were employed for detection:</p>\n<ul>\n<li><strong>Local sampling</strong>: Running RANSAC on all keypoints extracted from the entire image would require enormous computation time. Instead, keypoints within local regions are sampled from all extracted keypoints, and RANSAC is executed on these subsets. Regions are sampled in order of keypoint density, repeating until the number of keypoints falls below a certain threshold. The H matrix is then analyzed to filter out unrealistic transformations (scale &gt;4×, shear deformation, etc.)</li>\n<li><strong>Global sampling</strong>: Execute RANSAC once on all matches, then apply HDBSCAN clustering to remove false matches</li>\n</ul>\n<p>Additionally, RANSAC inlier thresholds were adjusted per image type based on the 3 classes detected in Stage 1:</p>\n<ul>\n<li><strong>Corn images</strong>: Fine black-and-white line patterns cause frequent false matches, so the inlier threshold was set <strong>high</strong> for strict filtering</li>\n<li><strong>Protein images (Western blot)</strong>: Very few features make matching inherently difficult, so the inlier threshold was set <strong>low</strong> to enable detection even with few matches</li>\n<li><strong>Other biological images</strong>: An intermediate threshold was used</li>\n</ul>\n<hr>\n<h2>Stage 3: Post-processing</h2>\n<p>After feature matching, I discovered that <strong>careful clustering is essential</strong>. Without it, the F1 score (the competition metric) drops significantly. To achieve high-precision clustering, I implemented the following processing.</p>\n<h3>Two-stage Cluster Merging</h3>\n<ol>\n<li><strong>Stage 1 (within same pair)</strong>: Merge nearby clusters with similar H matrices (normalized difference of rotation angle, scale, and translation &lt; 0.02) using Union-Find</li>\n<li><strong>Stage 2 (across different pairs)</strong>: Merge as the same instance when IoMin (Intersection / min(area1, area2)) between source/dest bboxes ≥ 0.3</li>\n</ol>\n<h3>High-precision Mask Generation</h3>\n<ol>\n<li>Recompute H matrix from all corresponding points of merged clusters</li>\n<li>Generate source/dest masks via convex hull + dilation</li>\n<li>Refine masks precisely using H matrix transformation error (within mean + 2σ)</li>\n<li>Apply bidirectional refinement using inverse transformation</li>\n</ol>\n<hr>\n<h2>What Didn't Work</h2>\n<ul>\n<li><strong>Other feature descriptors</strong> (SURF, ORB, AKAZE, SuperPoint): Could not extract sufficient keypoints in the biological image domain, resulting in poor accuracy</li>\n<li><strong>Deep learning-based approaches</strong>: Tried CMFD models combining DINOv2/SegFormer backbones with self-correlation modules, but accuracy was insufficient</li>\n</ul>\n<hr>\n<h2>Summary</h2>\n<table>\n<thead>\n<tr>\n<th>Component</th>\n<th>Details</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Panel Detection</td>\n<td>YOLOv8-m × 5 (WBF ensemble), trained on synthetic data</td>\n</tr>\n<tr>\n<td>Text Detection</td>\n<td>YOLOv8-m, manual annotation + synthetic data</td>\n</tr>\n<tr>\n<td>Features</td>\n<td>SIFT (contrast_threshold=0.001)</td>\n</tr>\n<tr>\n<td>Matching</td>\n<td>G2NN (α=0.7) + FLANN KDTree</td>\n</tr>\n<tr>\n<td>Geometric Estimation</td>\n<td>Affine RANSAC + H matrix filtering</td>\n</tr>\n<tr>\n<td>Clustering</td>\n<td>HDBSCAN + 2-stage merging (H similarity → IoMin)</td>\n</tr>\n<tr>\n<td>Mask Refinement</td>\n<td>Convex hull + H transformation error-based refinement</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# Recod.ai/LUC - Scientific Image Forgery Detection 2nd Place Solution\n\nFirst of all, thank you to the organizers for hosting such a wonderful competition. It was a highly meaningful theme—protecting the integrity of scientific research—and I learned a great deal from it.\n\nResult: **Public 3rd → Private 2nd** (out of 1564 teams)\n\n## Solution Overview\n\nRather than a deep learning-based approach, my solution is built around **classical feature matching (SIFT + RANSAC)**, combined with YOLO-based preprocessing and custom post-processing in a 3-stage pipeline.\n\n```\nStage 1: Valid Region Detection (Panel/Text detection with YOLOv8)\n    ↓\nStage 2: Feature Matching (SIFT + G2NN + RANSAC)\n    ↓\nStage 3: Post-processing (Cluster merging + High-precision mask generation)\n```\nThe overall algorithm overview is shown in the image below.\n\n---\n\n## Stage 1: Valid Region Detection\n\nIn the subsequent feature matching stage, **graph regions** produce a large number of false positives, and **text regions** cause erroneous matching on caption parts (e.g., matching on the \"mm\" portion of \"10 mm\"). These lead to significant accuracy degradation, making their exclusion essential.\n\nI trained a **Panel detection model** and a **Text detection model** using YOLOv8-m to address this problem.\n\n### Panel Detection\n\nTo exclude irrelevant regions such as graphs, only relevant areas (image regions) are detected. Simultaneously, YOLO classifies detections into 3 classes: **corn images**, **Western blot protein images**, and **other biological images**. The reason for separating classes is that the optimal RANSAC inlier threshold differs by image type (described below).\n\nSince no suitable training data existed, I addressed this by **automatically generating synthetic data**.\n\n- Used a subset of 2,000 images from external data (BioFors dataset) due to time constraints\n- Classified and extracted cell/plant images from the competition's authentic images\n- Auto-generated synthetic graphs (scatter plots, line charts, heatmaps, violin plots)\n- Randomly arranged the above in various layout patterns (1×1, 2×2, 3×1, etc.) to create training data\n- Ensembled 5 YOLOv8-m models trained with different seeds using **Weighted Boxes Fusion (WBF)**\n\n\n\n### Text Detection\n\nExisting OCR solutions (EasyOCR, etc.) were insufficiently accurate—for example, misrecognizing round cells as the letter \"O\". Since text in scientific paper images has minimal distortion, I determined that OCR-specific heads were unnecessary and **trained YOLOv8-m as an object detector**.\n\n- Manually corrected low-threshold EasyOCR detection results to create annotations (~800 images)\n- Added synthetic data (~2,000 images) with randomly placed text on text-free images\n- Applied diversity in background color, text color, font, rotation, and size for augmentation\n- At inference, performed TTA with two image sizes (640, 1280) and merged both results\n\n---\n\n## Stage 2: Feature Matching\n\n### Aggressive SIFT Keypoint Extraction\n\nBiological images have very little texture, and default settings fail to extract sufficient keypoints.\n\n- Used `contrast_threshold=0.001` (approximately 1/40 of the default 0.04) to extract a large number of keypoints\n- For small images (<1024px), upscaled 4× before feature extraction, then mapped coordinates back to original scale\n\n### G2NN Matching\n\nThe absolute L2 distance values vary greatly across images, making fixed thresholds impractical.\n\nAdopted **G2NN (Good-to-Next Neighbor)**: sort matching distances and accept only matches up to the breakpoint where $T_{n+1}/T_n > \\alpha$ ($\\alpha=0.7$). This achieves threshold-independent matching.\n\n### Speedup: KDTree\n\nDue to the enormous number of keypoints, used FLANN KDTree approximate nearest neighbor search to reduce complexity from $O(N^2)$ to $O(N \\log N)$.\n\n### RANSAC + Geometric Filtering\n\nHere, matching features are detected based on keypoints within the valid regions detected in Stage 1. Two sampling strategies were employed for detection:\n\n- **Local sampling**: Running RANSAC on all keypoints extracted from the entire image would require enormous computation time. Instead, keypoints within local regions are sampled from all extracted keypoints, and RANSAC is executed on these subsets. Regions are sampled in order of keypoint density, repeating until the number of keypoints falls below a certain threshold. The H matrix is then analyzed to filter out unrealistic transformations (scale >4×, shear deformation, etc.)\n- **Global sampling**: Execute RANSAC once on all matches, then apply HDBSCAN clustering to remove false matches\n\nAdditionally, RANSAC inlier thresholds were adjusted per image type based on the 3 classes detected in Stage 1:\n\n- **Corn images**: Fine black-and-white line patterns cause frequent false matches, so the inlier threshold was set **high** for strict filtering\n- **Protein images (Western blot)**: Very few features make matching inherently difficult, so the inlier threshold was set **low** to enable detection even with few matches\n- **Other biological images**: An intermediate threshold was used\n\n---\n\n## Stage 3: Post-processing\n\nAfter feature matching, I discovered that **careful clustering is essential**. Without it, the F1 score (the competition metric) drops significantly. To achieve high-precision clustering, I implemented the following processing.\n\n### Two-stage Cluster Merging\n\n1. **Stage 1 (within same pair)**: Merge nearby clusters with similar H matrices (normalized difference of rotation angle, scale, and translation < 0.02) using Union-Find\n2. **Stage 2 (across different pairs)**: Merge as the same instance when IoMin (Intersection / min(area1, area2)) between source/dest bboxes ≥ 0.3\n\n### High-precision Mask Generation\n\n1. Recompute H matrix from all corresponding points of merged clusters\n2. Generate source/dest masks via convex hull + dilation\n3. Refine masks precisely using H matrix transformation error (within mean + 2σ)\n4. Apply bidirectional refinement using inverse transformation\n\n---\n\n## What Didn't Work\n\n- **Other feature descriptors** (SURF, ORB, AKAZE, SuperPoint): Could not extract sufficient keypoints in the biological image domain, resulting in poor accuracy\n- **Deep learning-based approaches**: Tried CMFD models combining DINOv2/SegFormer backbones with self-correlation modules, but accuracy was insufficient\n\n---\n\n## Summary\n\n| Component | Details |\n|-----------|---------|\n| Panel Detection | YOLOv8-m × 5 (WBF ensemble), trained on synthetic data |\n| Text Detection | YOLOv8-m, manual annotation + synthetic data |\n| Features | SIFT (contrast_threshold=0.001) |\n| Matching | G2NN (α=0.7) + FLANN KDTree |\n| Geometric Estimation | Affine RANSAC + H matrix filtering |\n| Clustering | HDBSCAN + 2-stage merging (H similarity → IoMin) |\n| Mask Refinement | Convex hull + H transformation error-based refinement |\n\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3454572": "# Recod.ai/LUC - Scientific Image Forgery Detection 2nd Place Solution\n\nFirst of all, thank you to the organizers for hosting such a wonderful competition. It was a highly meaningful theme—protecting the integrity of scientific research—and I learned a great deal from it.\n\nResult: **Public 3rd → Private 2nd** (out of 1564 teams)\n\n## Solution Overview\n\nRather than a deep learning-based approach, my solution is built around **classical feature matching (SIFT + RANSAC)**, combined with YOLO-based preprocessing and custom post-processing in a 3-stage pipeline.\n\n```\nStage 1: Valid Region Detection (Panel/Text detection with YOLOv8)\n    ↓\nStage 2: Feature Matching (SIFT + G2NN + RANSAC)\n    ↓\nStage 3: Post-processing (Cluster merging + High-precision mask generation)\n```\nThe overall algorithm overview is shown in the image below.\n\n---\n\n## Stage 1: Valid Region Detection\n\nIn the subsequent feature matching stage, **graph regions** produce a large number of false positives, and **text regions** cause erroneous matching on caption parts (e.g., matching on the \"mm\" portion of \"10 mm\"). These lead to significant accuracy degradation, making their exclusion essential.\n\nI trained a **Panel detection model** and a **Text detection model** using YOLOv8-m to address this problem.\n\n### Panel Detection\n\nTo exclude irrelevant regions such as graphs, only relevant areas (image regions) are detected. Simultaneously, YOLO classifies detections into 3 classes: **corn images**, **Western blot protein images**, and **other biological images**. The reason for separating classes is that the optimal RANSAC inlier threshold differs by image type (described below).\n\nSince no suitable training data existed, I addressed this by **automatically generating synthetic data**.\n\n- Used a subset of 2,000 images from external data (BioFors dataset) due to time constraints\n- Classified and extracted cell/plant images from the competition's authentic images\n- Auto-generated synthetic graphs (scatter plots, line charts, heatmaps, violin plots)\n- Randomly arranged the above in various layout patterns (1×1, 2×2, 3×1, etc.) to create training data\n- Ensembled 5 YOLOv8-m models trained with different seeds using **Weighted Boxes Fusion (WBF)**\n\n\n\n### Text Detection\n\nExisting OCR solutions (EasyOCR, etc.) were insufficiently accurate—for example, misrecognizing round cells as the letter \"O\". Since text in scientific paper images has minimal distortion, I determined that OCR-specific heads were unnecessary and **trained YOLOv8-m as an object detector**.\n\n- Manually corrected low-threshold EasyOCR detection results to create annotations (~800 images)\n- Added synthetic data (~2,000 images) with randomly placed text on text-free images\n- Applied diversity in background color, text color, font, rotation, and size for augmentation\n- At inference, performed TTA with two image sizes (640, 1280) and merged both results\n\n---\n\n## Stage 2: Feature Matching\n\n### Aggressive SIFT Keypoint Extraction\n\nBiological images have very little texture, and default settings fail to extract sufficient keypoints.\n\n- Used `contrast_threshold=0.001` (approximately 1/40 of the default 0.04) to extract a large number of keypoints\n- For small images (<1024px), upscaled 4× before feature extraction, then mapped coordinates back to original scale\n\n### G2NN Matching\n\nThe absolute L2 distance values vary greatly across images, making fixed thresholds impractical.\n\nAdopted **G2NN (Good-to-Next Neighbor)**: sort matching distances and accept only matches up to the breakpoint where $T_{n+1}/T_n > \\alpha$ ($\\alpha=0.7$). This achieves threshold-independent matching.\n\n### Speedup: KDTree\n\nDue to the enormous number of keypoints, used FLANN KDTree approximate nearest neighbor search to reduce complexity from $O(N^2)$ to $O(N \\log N)$.\n\n### RANSAC + Geometric Filtering\n\nHere, matching features are detected based on keypoints within the valid regions detected in Stage 1. Two sampling strategies were employed for detection:\n\n- **Local sampling**: Running RANSAC on all keypoints extracted from the entire image would require enormous computation time. Instead, keypoints within local regions are sampled from all extracted keypoints, and RANSAC is executed on these subsets. Regions are sampled in order of keypoint density, repeating until the number of keypoints falls below a certain threshold. The H matrix is then analyzed to filter out unrealistic transformations (scale >4×, shear deformation, etc.)\n- **Global sampling**: Execute RANSAC once on all matches, then apply HDBSCAN clustering to remove false matches\n\nAdditionally, RANSAC inlier thresholds were adjusted per image type based on the 3 classes detected in Stage 1:\n\n- **Corn images**: Fine black-and-white line patterns cause frequent false matches, so the inlier threshold was set **high** for strict filtering\n- **Protein images (Western blot)**: Very few features make matching inherently difficult, so the inlier threshold was set **low** to enable detection even with few matches\n- **Other biological images**: An intermediate threshold was used\n\n---\n\n## Stage 3: Post-processing\n\nAfter feature matching, I discovered that **careful clustering is essential**. Without it, the F1 score (the competition metric) drops significantly. To achieve high-precision clustering, I implemented the following processing.\n\n### Two-stage Cluster Merging\n\n1. **Stage 1 (within same pair)**: Merge nearby clusters with similar H matrices (normalized difference of rotation angle, scale, and translation < 0.02) using Union-Find\n2. **Stage 2 (across different pairs)**: Merge as the same instance when IoMin (Intersection / min(area1, area2)) between source/dest bboxes ≥ 0.3\n\n### High-precision Mask Generation\n\n1. Recompute H matrix from all corresponding points of merged clusters\n2. Generate source/dest masks via convex hull + dilation\n3. Refine masks precisely using H matrix transformation error (within mean + 2σ)\n4. Apply bidirectional refinement using inverse transformation\n\n---\n\n## What Didn't Work\n\n- **Other feature descriptors** (SURF, ORB, AKAZE, SuperPoint): Could not extract sufficient keypoints in the biological image domain, resulting in poor accuracy\n- **Deep learning-based approaches**: Tried CMFD models combining DINOv2/SegFormer backbones with self-correlation modules, but accuracy was insufficient\n\n---\n\n## Summary\n\n| Component | Details |\n|-----------|---------|\n| Panel Detection | YOLOv8-m × 5 (WBF ensemble), trained on synthetic data |\n| Text Detection | YOLOv8-m, manual annotation + synthetic data |\n| Features | SIFT (contrast_threshold=0.001) |\n| Matching | G2NN (α=0.7) + FLANN KDTree |\n| Geometric Estimation | Affine RANSAC + H matrix filtering |\n| Clustering | HDBSCAN + 2-stage merging (H similarity → IoMin) |\n| Mask Refinement | Convex hull + H transformation error-based refinement |\n\n"
  }
}