{
  "id": 697675,
  "title": "Solution: Public LB 8th place | Private LB 3rd place (updated 2026-04-28)",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/697675",
  "author_name": "CoreyJamesLevinson",
  "post_date": "2026-05-06T21:01:41.683000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Please read the attached pdfs below</p>\n<p><em>New update after the competition ended</em></p>\n<h1>Segmentation of image into panels</h1>\n<ol>\n<li>Segment Anything 3 model is prompted with domain-specific text cues to detect individual panels within a composite scientific figure. Multiple prompts are tried in sequence:\n[\"outlined scientific images\", \"microscopic images\", \"bordered rectangles\", …]</li>\n<li>Accept the first prompt that returns any segmentation masks.</li>\n</ol>\n<h1>Pre-processing of panels</h1>\n<ol>\n<li>Apply CLAHE, which corrects for brightness/exposure manipulations</li>\n<li>Detect text and arrows which are overlaid on top of the images, and blur them out.</li>\n</ol>\n<h1>SIFT + Matching</h1>\n<ol>\n<li>Extract keypoints from each panel using SIFT, which is scale and rotation invariant.</li>\n<li>Candidate pairs are matched using FLANN approximate nearest-neighbor matcher using the SIFT keypoints.</li>\n<li>Filter out low-probability candidates by using Lowe's ratio test</li>\n<li>Use RANSAC (via homography) to filter out bad matches and enforce geometric consistency between panels</li>\n</ol>\n<h1>Tightening the regions</h1>\n<ol>\n<li>instead of matching the entire panel, only segment around the tight bounding box of the RANSAC inlier matched keypoints - this handles cases where only a <em>partial</em> region of one image is duplicated, rather than the full crop.</li>\n<li>Remove candidates where there is high spatial overlap between the images (SAM3 detection error of the same region), or the two are vastly different areas (geometric inconsistency)</li>\n</ol>\n<h1>Rerun segmentation</h1>\n<ol>\n<li>Re-run SAM3, but now having prompts targeting image context rather than panel structure - seeking biological objects rather than borders. The prompts are:\n[\"biological cells\", \"scientific blobs\", \"item of interest\", …]</li>\n<li>If SAM3 finds nothing, then inject deterministic crops of the overall image. This ensures every figure has at least some regions to compare. For example: split image into halves, thirds, quadrants, etc.</li>\n<li>Re-run previous steps with relaxed parameters:\nCLAHE -&gt; text/arrow suppression -&gt; SIFT -&gt; FLANN -&gt; Lowe's ratio test -&gt; RANSAC -&gt; tighten bounding boxes -&gt; remove rejections</li>\n</ol>\n<h1>Interesting findings</h1>\n<p>No model training was used. Things that did not work:</p>\n<ul>\n<li>training a ViT UNET / segformer / DinoV3 / mask2former to predict duplicates. For that matter, extending training on other copy-forge datasets like Casia, DefactoCopyMove, figshare_wb, kaggle-dsbowl, etc.</li>\n<li>approaches in academic literature did not work for me: BCMNet, Beit Base, Busternet, LBRT</li>\n<li>initially I trained my own custom DETR model to identify where panels were (using noisy training data by saying anything non-white background was part of a panel). However, SAM3 model from Facebook was slightly better. I even tried some unique ideas like Comic Panel Detection NN and DocLayout YOLO DoyLayNet.</li>\n<li>using cellpose to segment cells, then performing matching based on that (too slow)</li>\n<li>other keypoint methods (ROMA V2) are promising but too slow</li>\n<li>sherloq (<a href=\"https://github.com/GuidoBartoli/sherloq\" target=\"_blank\">https://github.com/GuidoBartoli/sherloq</a>) which was only able to match duplicate text</li>\n<li>photoholmes (<a href=\"https://github.com/photoholmes/photoholmes\" target=\"_blank\">https://github.com/photoholmes/photoholmes</a>) wasn't good at copy-forge detection</li>\n<li>Forensically (<a href=\"https://29a.ch/photo-forensics/#forensic-magnifier\" target=\"_blank\">https://29a.ch/photo-forensics/#forensic-magnifier</a>) - i recreated the algorithm in Python, but the hyperparameters are too fiddly to work with. Would probably be valuable as an uncorrelated approach to find more duplicates.</li>\n<li>imagetwin, imachek, proofig, figcheck - these all cost money and essentially the competition is to create free open-source versions of these</li>\n</ul>\n<p>I also handlabeled an extra dataset of a few hundred duplicate pairs from PubPeer/RetractionWatch. However, I did not have enough time to use it in my validation set beyond trying out different SAM3 prompts on it.</p>\n<p>The key insight to this competition was to look at your model performance on the supplemental test data. Empirically you can see that training a NN simply on (synthetic) individual images does not translate well to the supplemental data which was \"figure\"-based (multiple panels). Therefore a completely different approach was required</p>",
  "messages": [
    {
      "id": 3454299,
      "postDate": "2026-05-06T21:01:41.683Z",
      "content": "<p>Please read the attached pdfs below</p>\n<p><em>New update after the competition ended</em></p>\n<h1>Segmentation of image into panels</h1>\n<ol>\n<li>Segment Anything 3 model is prompted with domain-specific text cues to detect individual panels within a composite scientific figure. Multiple prompts are tried in sequence:\n[\"outlined scientific images\", \"microscopic images\", \"bordered rectangles\", …]</li>\n<li>Accept the first prompt that returns any segmentation masks.</li>\n</ol>\n<h1>Pre-processing of panels</h1>\n<ol>\n<li>Apply CLAHE, which corrects for brightness/exposure manipulations</li>\n<li>Detect text and arrows which are overlaid on top of the images, and blur them out.</li>\n</ol>\n<h1>SIFT + Matching</h1>\n<ol>\n<li>Extract keypoints from each panel using SIFT, which is scale and rotation invariant.</li>\n<li>Candidate pairs are matched using FLANN approximate nearest-neighbor matcher using the SIFT keypoints.</li>\n<li>Filter out low-probability candidates by using Lowe's ratio test</li>\n<li>Use RANSAC (via homography) to filter out bad matches and enforce geometric consistency between panels</li>\n</ol>\n<h1>Tightening the regions</h1>\n<ol>\n<li>instead of matching the entire panel, only segment around the tight bounding box of the RANSAC inlier matched keypoints - this handles cases where only a <em>partial</em> region of one image is duplicated, rather than the full crop.</li>\n<li>Remove candidates where there is high spatial overlap between the images (SAM3 detection error of the same region), or the two are vastly different areas (geometric inconsistency)</li>\n</ol>\n<h1>Rerun segmentation</h1>\n<ol>\n<li>Re-run SAM3, but now having prompts targeting image context rather than panel structure - seeking biological objects rather than borders. The prompts are:\n[\"biological cells\", \"scientific blobs\", \"item of interest\", …]</li>\n<li>If SAM3 finds nothing, then inject deterministic crops of the overall image. This ensures every figure has at least some regions to compare. For example: split image into halves, thirds, quadrants, etc.</li>\n<li>Re-run previous steps with relaxed parameters:\nCLAHE -&gt; text/arrow suppression -&gt; SIFT -&gt; FLANN -&gt; Lowe's ratio test -&gt; RANSAC -&gt; tighten bounding boxes -&gt; remove rejections</li>\n</ol>\n<h1>Interesting findings</h1>\n<p>No model training was used. Things that did not work:</p>\n<ul>\n<li>training a ViT UNET / segformer / DinoV3 / mask2former to predict duplicates. For that matter, extending training on other copy-forge datasets like Casia, DefactoCopyMove, figshare_wb, kaggle-dsbowl, etc.</li>\n<li>approaches in academic literature did not work for me: BCMNet, Beit Base, Busternet, LBRT</li>\n<li>initially I trained my own custom DETR model to identify where panels were (using noisy training data by saying anything non-white background was part of a panel). However, SAM3 model from Facebook was slightly better. I even tried some unique ideas like Comic Panel Detection NN and DocLayout YOLO DoyLayNet.</li>\n<li>using cellpose to segment cells, then performing matching based on that (too slow)</li>\n<li>other keypoint methods (ROMA V2) are promising but too slow</li>\n<li>sherloq (<a href=\"https://github.com/GuidoBartoli/sherloq\" target=\"_blank\">https://github.com/GuidoBartoli/sherloq</a>) which was only able to match duplicate text</li>\n<li>photoholmes (<a href=\"https://github.com/photoholmes/photoholmes\" target=\"_blank\">https://github.com/photoholmes/photoholmes</a>) wasn't good at copy-forge detection</li>\n<li>Forensically (<a href=\"https://29a.ch/photo-forensics/#forensic-magnifier\" target=\"_blank\">https://29a.ch/photo-forensics/#forensic-magnifier</a>) - i recreated the algorithm in Python, but the hyperparameters are too fiddly to work with. Would probably be valuable as an uncorrelated approach to find more duplicates.</li>\n<li>imagetwin, imachek, proofig, figcheck - these all cost money and essentially the competition is to create free open-source versions of these</li>\n</ul>\n<p>I also handlabeled an extra dataset of a few hundred duplicate pairs from PubPeer/RetractionWatch. However, I did not have enough time to use it in my validation set beyond trying out different SAM3 prompts on it.</p>\n<p>The key insight to this competition was to look at your model performance on the supplemental test data. Empirically you can see that training a NN simply on (synthetic) individual images does not translate well to the supplemental data which was \"figure\"-based (multiple panels). Therefore a completely different approach was required</p>",
      "rawMarkdown": "Please read the attached pdfs below\n\n*New update after the competition ended*\n\n# Segmentation of image into panels\n1. Segment Anything 3 model is prompted with domain-specific text cues to detect individual panels within a composite scientific figure. Multiple prompts are tried in sequence:\n   [\"outlined scientific images\", \"microscopic images\", \"bordered rectangles\", ...]\n2. Accept the first prompt that returns any segmentation masks.\n\n# Pre-processing of panels\n1. Apply CLAHE, which corrects for brightness/exposure manipulations\n2. Detect text and arrows which are overlaid on top of the images, and blur them out.\n\n# SIFT + Matching\n1. Extract keypoints from each panel using SIFT, which is scale and rotation invariant.\n2. Candidate pairs are matched using FLANN approximate nearest-neighbor matcher using the SIFT keypoints.\n3. Filter out low-probability candidates by using Lowe's ratio test\n4. Use RANSAC (via homography) to filter out bad matches and enforce geometric consistency between panels\n\n# Tightening the regions\n1. instead of matching the entire panel, only segment around the tight bounding box of the RANSAC inlier matched keypoints - this handles cases where only a *partial* region of one image is duplicated, rather than the full crop.\n2. Remove candidates where there is high spatial overlap between the images (SAM3 detection error of the same region), or the two are vastly different areas (geometric inconsistency)\n\n# Rerun segmentation\n1. Re-run SAM3, but now having prompts targeting image context rather than panel structure - seeking biological objects rather than borders. The prompts are:\n   [\"biological cells\", \"scientific blobs\", \"item of interest\", ...]\n2. If SAM3 finds nothing, then inject deterministic crops of the overall image. This ensures every figure has at least some regions to compare. For example: split image into halves, thirds, quadrants, etc.\n3. Re-run previous steps with relaxed parameters:\n   CLAHE -> text/arrow suppression -> SIFT -> FLANN -> Lowe's ratio test -> RANSAC -> tighten bounding boxes -> remove rejections\n\n# Interesting findings\n\nNo model training was used. Things that did not work:\n\n- training a ViT UNET / segformer / DinoV3 / mask2former to predict duplicates. For that matter, extending training on other copy-forge datasets like Casia, DefactoCopyMove, figshare_wb, kaggle-dsbowl, etc.\n- approaches in academic literature did not work for me: BCMNet, Beit Base, Busternet, LBRT\n- initially I trained my own custom DETR model to identify where panels were (using noisy training data by saying anything non-white background was part of a panel). However, SAM3 model from Facebook was slightly better. I even tried some unique ideas like Comic Panel Detection NN and DocLayout YOLO DoyLayNet.\n- using cellpose to segment cells, then performing matching based on that (too slow)\n- other keypoint methods (ROMA V2) are promising but too slow\n- sherloq (https://github.com/GuidoBartoli/sherloq) which was only able to match duplicate text\n- photoholmes (https://github.com/photoholmes/photoholmes) wasn't good at copy-forge detection\n- Forensically (https://29a.ch/photo-forensics/#forensic-magnifier) - i recreated the algorithm in Python, but the hyperparameters are too fiddly to work with. Would probably be valuable as an uncorrelated approach to find more duplicates.\n- imagetwin, imachek, proofig, figcheck - these all cost money and essentially the competition is to create free open-source versions of these\n\nI also handlabeled an extra dataset of a few hundred duplicate pairs from PubPeer/RetractionWatch. However, I did not have enough time to use it in my validation set beyond trying out different SAM3 prompts on it.\n\nThe key insight to this competition was to look at your model performance on the supplemental test data. Empirically you can see that training a NN simply on (synthetic) individual images does not translate well to the supplemental data which was \"figure\"-based (multiple panels). Therefore a completely different approach was required",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3454299": "Please read the attached pdfs below\n\n*New update after the competition ended*\n\n# Segmentation of image into panels\n1. Segment Anything 3 model is prompted with domain-specific text cues to detect individual panels within a composite scientific figure. Multiple prompts are tried in sequence:\n   [\"outlined scientific images\", \"microscopic images\", \"bordered rectangles\", ...]\n2. Accept the first prompt that returns any segmentation masks.\n\n# Pre-processing of panels\n1. Apply CLAHE, which corrects for brightness/exposure manipulations\n2. Detect text and arrows which are overlaid on top of the images, and blur them out.\n\n# SIFT + Matching\n1. Extract keypoints from each panel using SIFT, which is scale and rotation invariant.\n2. Candidate pairs are matched using FLANN approximate nearest-neighbor matcher using the SIFT keypoints.\n3. Filter out low-probability candidates by using Lowe's ratio test\n4. Use RANSAC (via homography) to filter out bad matches and enforce geometric consistency between panels\n\n# Tightening the regions\n1. instead of matching the entire panel, only segment around the tight bounding box of the RANSAC inlier matched keypoints - this handles cases where only a *partial* region of one image is duplicated, rather than the full crop.\n2. Remove candidates where there is high spatial overlap between the images (SAM3 detection error of the same region), or the two are vastly different areas (geometric inconsistency)\n\n# Rerun segmentation\n1. Re-run SAM3, but now having prompts targeting image context rather than panel structure - seeking biological objects rather than borders. The prompts are:\n   [\"biological cells\", \"scientific blobs\", \"item of interest\", ...]\n2. If SAM3 finds nothing, then inject deterministic crops of the overall image. This ensures every figure has at least some regions to compare. For example: split image into halves, thirds, quadrants, etc.\n3. Re-run previous steps with relaxed parameters:\n   CLAHE -> text/arrow suppression -> SIFT -> FLANN -> Lowe's ratio test -> RANSAC -> tighten bounding boxes -> remove rejections\n\n# Interesting findings\n\nNo model training was used. Things that did not work:\n\n- training a ViT UNET / segformer / DinoV3 / mask2former to predict duplicates. For that matter, extending training on other copy-forge datasets like Casia, DefactoCopyMove, figshare_wb, kaggle-dsbowl, etc.\n- approaches in academic literature did not work for me: BCMNet, Beit Base, Busternet, LBRT\n- initially I trained my own custom DETR model to identify where panels were (using noisy training data by saying anything non-white background was part of a panel). However, SAM3 model from Facebook was slightly better. I even tried some unique ideas like Comic Panel Detection NN and DocLayout YOLO DoyLayNet.\n- using cellpose to segment cells, then performing matching based on that (too slow)\n- other keypoint methods (ROMA V2) are promising but too slow\n- sherloq (https://github.com/GuidoBartoli/sherloq) which was only able to match duplicate text\n- photoholmes (https://github.com/photoholmes/photoholmes) wasn't good at copy-forge detection\n- Forensically (https://29a.ch/photo-forensics/#forensic-magnifier) - i recreated the algorithm in Python, but the hyperparameters are too fiddly to work with. Would probably be valuable as an uncorrelated approach to find more duplicates.\n- imagetwin, imachek, proofig, figcheck - these all cost money and essentially the competition is to create free open-source versions of these\n\nI also handlabeled an extra dataset of a few hundred duplicate pairs from PubPeer/RetractionWatch. However, I did not have enough time to use it in my validation set beyond trying out different SAM3 prompts on it.\n\nThe key insight to this competition was to look at your model performance on the supplemental test data. Empirically you can see that training a NN simply on (synthetic) individual images does not translate well to the supplemental data which was \"figure\"-based (multiple panels). Therefore a completely different approach was required"
  }
}