{
  "id": 700586,
  "title": "6th place solution: SuperPoint + LightGlue",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/700586",
  "author_name": "Pavel Kazlou",
  "post_date": "2026-05-17T18:47:04.686000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<ul>\n<li>Public LB: 0.322 (398th)</li>\n<li>Private LB: 0.336 (6th)</li>\n<li><a href=\"https://www.kaggle.com/code/paulik/recod-ai-sifd-no-dino-superpoint-lightglue\" target=\"_blank\">Submission notebook</a></li>\n</ul>\n<h2>Overview</h2>\n<p>The pipeline for detecting matching regions was pretty simple:</p>\n<ol>\n<li><strong>Tile Detection</strong>: segmenting the input image into rectangular tiles to be compared in pairs.</li>\n<li><strong>Feature Extraction</strong>: using SuperPoint (2048 keypoints) to extract local features from each tile.</li>\n<li><strong>Feature Matching</strong>: using LightGlue to match SuperPoint features between all tile pairs.</li>\n<li><strong>RANSAC Affine Verification</strong>: validating matches geometrically</li>\n<li><strong>Mask Generation &amp; Merging</strong>: producing masks out of matching tile pairs</li>\n</ol>\n<p>Let me dive into more details for each step.</p>\n<h2>Tile Detection</h2>\n<ol>\n<li>The image is first binarised with a threshold of 245 over grayscale version of image and then inverted.</li>\n<li>Morphological open+close 3x3 kernels are run on top to remove white noise on black (open) as well as black noise on white (close).</li>\n<li>Continuous regions are found with contours,  representing objects on the background. </li>\n<li>Regions are fit into bounding boxes - these are initial candidates for tiles</li>\n<li>Candidate tiles are then filtered: should be &gt;= 0.5% of image size,  width and height &gt;= 50px. </li>\n</ol>\n<p>There were two issues with this algorithm: </p>\n<ol>\n<li>Sometime actual tiles were organised in tables with no white background separators, and thus such tables were detected as one large tile. To deal with such cases I've introduced separate post-processing to try and detect grid lines. To do so I've used Canny edge detection and required &gt; 50% of pixels be an edge in a given row or column. Upon detecting grid lines, the tile was split into sub-tiles. Unfortunately, while this algo was correctly handling tables, it produced false splits in other cases and dropped the public LB score, despite improving the score on local validation. So it was turned off.</li>\n<li>Generated images like plots and diagrams were detected as tiles. They contained texts, axes and icons that were detected as matches resulting in false positives regions. To filter out such tiles, I've used heuristics based on grayscale histogram. The idea is simple: plots and diagrams use small amount of colors and thus their histograms will contains sharp spikes, while biological images will have smoother histograms.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F044312420e72aac43d133fe1c7128ce3%2FScreenshot%202026-05-17%20at%2021.00.09.png?generation=1779041695038957&amp;alt=media\" alt=\"\"> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F3ff60df7cc682dd9a87c0b8f709acd46%2FScreenshot%202026-05-17%20at%2021.00.52.png?generation=1779041718636493&amp;alt=media\" alt=\"\"></p>\n<h2>Feature Extraction</h2>\n<p>Superpoint was used to detect 2048 most interesting points per tile and extract their features as 256-dim vectors.</p>\n<p>Any texts were producing points and resulting in false positives matches. So a separate post-processing was introduced to detect texts and remove any points associated with them. Text was detected with DB50 model.</p>\n<h2>Feature Matching</h2>\n<p>LightGlue was used to match points from one tile with points from another tile. For some tile pairs the LightGlue was stopping matching early due to low match scores, resulting in under-detection. This is why I had to disable early stop explicitly, increased inference time was not a problem. The points were then filtered by match score threshold, leaving only the pairs of points with decent match confidence.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F94872de5e19f235aa306672a9e6b3a73%2FScreenshot%202026-05-17%20at%2021.03.50.png?generation=1779041639040268&amp;alt=media\" alt=\"\"></p>\n<h2>RANSAC Affine Verification</h2>\n<p>At this stage there were pairs of matching points potentially containing some noisy matches from different parts of the image. So to further narrow down the matching region, an affine transform was fit to the points, requiring a group of points to behave as a whole: the position of matching points on tile T1 and T2 should be described with the same/similar affine transform. For this RANSAC method from OpenCV library was used, with requirements to have at least 20 inliers and at least 20% inlier ratio among points.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F86a52cd3d9feb63f4a999b879f1d0fb6%2FScreenshot%202026-05-17%20at%2021.33.26.png?generation=1779042836917640&amp;alt=media\" alt=\"\"></p>\n<p>Then minimum area rotated bounding box was computed around the inlier keypoints in each tile, followed by axis-aligned bounding boxes calculation.</p>\n<h2>Mask Generation &amp; Merging</h2>\n<p>Bounding boxes for matched regions were converted to masks on the original image. </p>\n<p>Sometimes mask were intersecting (for example the matching regions were A - B and A - C), in such case masks were merged into the same layer (requiring intersection-over-union to be at least 0.1).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2Fa4b936dfc5664b01df86aed2a72c55f3%2FScreenshot%202026-05-17%20at%2020.47.08.png?generation=1779042269764739&amp;alt=media\" alt=\"\"></p>\n<h2>Additional thoughts</h2>\n<p>I want to thank the organizers and Kaggle for hosting this competition. </p>\n<p>After the submission phase ended, I still had hope for my solution despite being far outside the medal range on the public leaderboard. Unlike most public solutions based on DINO training, I used a point-matching approach, which I expected to be more robust to distribution shifts in new data. Still, finishing in 6th place came as a big surprise to me.</p>",
  "messages": [
    {
      "id": 3459395,
      "postDate": "2026-05-17T18:47:04.687Z",
      "content": "<ul>\n<li>Public LB: 0.322 (398th)</li>\n<li>Private LB: 0.336 (6th)</li>\n<li><a href=\"https://www.kaggle.com/code/paulik/recod-ai-sifd-no-dino-superpoint-lightglue\" target=\"_blank\">Submission notebook</a></li>\n</ul>\n<h2>Overview</h2>\n<p>The pipeline for detecting matching regions was pretty simple:</p>\n<ol>\n<li><strong>Tile Detection</strong>: segmenting the input image into rectangular tiles to be compared in pairs.</li>\n<li><strong>Feature Extraction</strong>: using SuperPoint (2048 keypoints) to extract local features from each tile.</li>\n<li><strong>Feature Matching</strong>: using LightGlue to match SuperPoint features between all tile pairs.</li>\n<li><strong>RANSAC Affine Verification</strong>: validating matches geometrically</li>\n<li><strong>Mask Generation &amp; Merging</strong>: producing masks out of matching tile pairs</li>\n</ol>\n<p>Let me dive into more details for each step.</p>\n<h2>Tile Detection</h2>\n<ol>\n<li>The image is first binarised with a threshold of 245 over grayscale version of image and then inverted.</li>\n<li>Morphological open+close 3x3 kernels are run on top to remove white noise on black (open) as well as black noise on white (close).</li>\n<li>Continuous regions are found with contours,  representing objects on the background. </li>\n<li>Regions are fit into bounding boxes - these are initial candidates for tiles</li>\n<li>Candidate tiles are then filtered: should be &gt;= 0.5% of image size,  width and height &gt;= 50px. </li>\n</ol>\n<p>There were two issues with this algorithm: </p>\n<ol>\n<li>Sometime actual tiles were organised in tables with no white background separators, and thus such tables were detected as one large tile. To deal with such cases I've introduced separate post-processing to try and detect grid lines. To do so I've used Canny edge detection and required &gt; 50% of pixels be an edge in a given row or column. Upon detecting grid lines, the tile was split into sub-tiles. Unfortunately, while this algo was correctly handling tables, it produced false splits in other cases and dropped the public LB score, despite improving the score on local validation. So it was turned off.</li>\n<li>Generated images like plots and diagrams were detected as tiles. They contained texts, axes and icons that were detected as matches resulting in false positives regions. To filter out such tiles, I've used heuristics based on grayscale histogram. The idea is simple: plots and diagrams use small amount of colors and thus their histograms will contains sharp spikes, while biological images will have smoother histograms.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F044312420e72aac43d133fe1c7128ce3%2FScreenshot%202026-05-17%20at%2021.00.09.png?generation=1779041695038957&amp;alt=media\" alt=\"\"> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F3ff60df7cc682dd9a87c0b8f709acd46%2FScreenshot%202026-05-17%20at%2021.00.52.png?generation=1779041718636493&amp;alt=media\" alt=\"\"></p>\n<h2>Feature Extraction</h2>\n<p>Superpoint was used to detect 2048 most interesting points per tile and extract their features as 256-dim vectors.</p>\n<p>Any texts were producing points and resulting in false positives matches. So a separate post-processing was introduced to detect texts and remove any points associated with them. Text was detected with DB50 model.</p>\n<h2>Feature Matching</h2>\n<p>LightGlue was used to match points from one tile with points from another tile. For some tile pairs the LightGlue was stopping matching early due to low match scores, resulting in under-detection. This is why I had to disable early stop explicitly, increased inference time was not a problem. The points were then filtered by match score threshold, leaving only the pairs of points with decent match confidence.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F94872de5e19f235aa306672a9e6b3a73%2FScreenshot%202026-05-17%20at%2021.03.50.png?generation=1779041639040268&amp;alt=media\" alt=\"\"></p>\n<h2>RANSAC Affine Verification</h2>\n<p>At this stage there were pairs of matching points potentially containing some noisy matches from different parts of the image. So to further narrow down the matching region, an affine transform was fit to the points, requiring a group of points to behave as a whole: the position of matching points on tile T1 and T2 should be described with the same/similar affine transform. For this RANSAC method from OpenCV library was used, with requirements to have at least 20 inliers and at least 20% inlier ratio among points.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F86a52cd3d9feb63f4a999b879f1d0fb6%2FScreenshot%202026-05-17%20at%2021.33.26.png?generation=1779042836917640&amp;alt=media\" alt=\"\"></p>\n<p>Then minimum area rotated bounding box was computed around the inlier keypoints in each tile, followed by axis-aligned bounding boxes calculation.</p>\n<h2>Mask Generation &amp; Merging</h2>\n<p>Bounding boxes for matched regions were converted to masks on the original image. </p>\n<p>Sometimes mask were intersecting (for example the matching regions were A - B and A - C), in such case masks were merged into the same layer (requiring intersection-over-union to be at least 0.1).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2Fa4b936dfc5664b01df86aed2a72c55f3%2FScreenshot%202026-05-17%20at%2020.47.08.png?generation=1779042269764739&amp;alt=media\" alt=\"\"></p>\n<h2>Additional thoughts</h2>\n<p>I want to thank the organizers and Kaggle for hosting this competition. </p>\n<p>After the submission phase ended, I still had hope for my solution despite being far outside the medal range on the public leaderboard. Unlike most public solutions based on DINO training, I used a point-matching approach, which I expected to be more robust to distribution shifts in new data. Still, finishing in 6th place came as a big surprise to me.</p>",
      "rawMarkdown": "- Public LB: 0.322 (398th)\n- Private LB: 0.336 (6th)\n- [Submission notebook](https://www.kaggle.com/code/paulik/recod-ai-sifd-no-dino-superpoint-lightglue)\n\n## Overview ##\n\nThe pipeline for detecting matching regions was pretty simple:\n1. **Tile Detection**: segmenting the input image into rectangular tiles to be compared in pairs.\n1. **Feature Extraction**: using SuperPoint (2048 keypoints) to extract local features from each tile.\n1. **Feature Matching**: using LightGlue to match SuperPoint features between all tile pairs.\n1. **RANSAC Affine Verification**: validating matches geometrically\n1. **Mask Generation & Merging**: producing masks out of matching tile pairs\n\nLet me dive into more details for each step.\n\n## Tile Detection ##\n1. The image is first binarised with a threshold of 245 over grayscale version of image and then inverted.\n1. Morphological open+close 3x3 kernels are run on top to remove white noise on black (open) as well as black noise on white (close).\n1. Continuous regions are found with contours,  representing objects on the background. \n1. Regions are fit into bounding boxes - these are initial candidates for tiles\n1. Candidate tiles are then filtered: should be >= 0.5% of image size,  width and height >= 50px. \n\nThere were two issues with this algorithm: \n1. Sometime actual tiles were organised in tables with no white background separators, and thus such tables were detected as one large tile. To deal with such cases I've introduced separate post-processing to try and detect grid lines. To do so I've used Canny edge detection and required > 50% of pixels be an edge in a given row or column. Upon detecting grid lines, the tile was split into sub-tiles. Unfortunately, while this algo was correctly handling tables, it produced false splits in other cases and dropped the public LB score, despite improving the score on local validation. So it was turned off.\n2. Generated images like plots and diagrams were detected as tiles. They contained texts, axes and icons that were detected as matches resulting in false positives regions. To filter out such tiles, I've used heuristics based on grayscale histogram. The idea is simple: plots and diagrams use small amount of colors and thus their histograms will contains sharp spikes, while biological images will have smoother histograms.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F044312420e72aac43d133fe1c7128ce3%2FScreenshot%202026-05-17%20at%2021.00.09.png?generation=1779041695038957&alt=media =500x200) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F3ff60df7cc682dd9a87c0b8f709acd46%2FScreenshot%202026-05-17%20at%2021.00.52.png?generation=1779041718636493&alt=media =500x200)\n\n## Feature Extraction ##\nSuperpoint was used to detect 2048 most interesting points per tile and extract their features as 256-dim vectors.\n\nAny texts were producing points and resulting in false positives matches. So a separate post-processing was introduced to detect texts and remove any points associated with them. Text was detected with DB50 model.\n\n## Feature Matching ##\nLightGlue was used to match points from one tile with points from another tile. For some tile pairs the LightGlue was stopping matching early due to low match scores, resulting in under-detection. This is why I had to disable early stop explicitly, increased inference time was not a problem. The points were then filtered by match score threshold, leaving only the pairs of points with decent match confidence.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F94872de5e19f235aa306672a9e6b3a73%2FScreenshot%202026-05-17%20at%2021.03.50.png?generation=1779041639040268&alt=media =700x300)\n\n## RANSAC Affine Verification ##\nAt this stage there were pairs of matching points potentially containing some noisy matches from different parts of the image. So to further narrow down the matching region, an affine transform was fit to the points, requiring a group of points to behave as a whole: the position of matching points on tile T1 and T2 should be described with the same/similar affine transform. For this RANSAC method from OpenCV library was used, with requirements to have at least 20 inliers and at least 20% inlier ratio among points.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F86a52cd3d9feb63f4a999b879f1d0fb6%2FScreenshot%202026-05-17%20at%2021.33.26.png?generation=1779042836917640&alt=media =700x300)\n\nThen minimum area rotated bounding box was computed around the inlier keypoints in each tile, followed by axis-aligned bounding boxes calculation.\n\n## Mask Generation & Merging ##\nBounding boxes for matched regions were converted to masks on the original image. \n\nSometimes mask were intersecting (for example the matching regions were A - B and A - C), in such case masks were merged into the same layer (requiring intersection-over-union to be at least 0.1).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2Fa4b936dfc5664b01df86aed2a72c55f3%2FScreenshot%202026-05-17%20at%2020.47.08.png?generation=1779042269764739&alt=media =300x300)\n\n## Additional thoughts ##\nI want to thank the organizers and Kaggle for hosting this competition. \n\nAfter the submission phase ended, I still had hope for my solution despite being far outside the medal range on the public leaderboard. Unlike most public solutions based on DINO training, I used a point-matching approach, which I expected to be more robust to distribution shifts in new data. Still, finishing in 6th place came as a big surprise to me.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3459395": "- Public LB: 0.322 (398th)\n- Private LB: 0.336 (6th)\n- [Submission notebook](https://www.kaggle.com/code/paulik/recod-ai-sifd-no-dino-superpoint-lightglue)\n\n## Overview ##\n\nThe pipeline for detecting matching regions was pretty simple:\n1. **Tile Detection**: segmenting the input image into rectangular tiles to be compared in pairs.\n1. **Feature Extraction**: using SuperPoint (2048 keypoints) to extract local features from each tile.\n1. **Feature Matching**: using LightGlue to match SuperPoint features between all tile pairs.\n1. **RANSAC Affine Verification**: validating matches geometrically\n1. **Mask Generation & Merging**: producing masks out of matching tile pairs\n\nLet me dive into more details for each step.\n\n## Tile Detection ##\n1. The image is first binarised with a threshold of 245 over grayscale version of image and then inverted.\n1. Morphological open+close 3x3 kernels are run on top to remove white noise on black (open) as well as black noise on white (close).\n1. Continuous regions are found with contours,  representing objects on the background. \n1. Regions are fit into bounding boxes - these are initial candidates for tiles\n1. Candidate tiles are then filtered: should be >= 0.5% of image size,  width and height >= 50px. \n\nThere were two issues with this algorithm: \n1. Sometime actual tiles were organised in tables with no white background separators, and thus such tables were detected as one large tile. To deal with such cases I've introduced separate post-processing to try and detect grid lines. To do so I've used Canny edge detection and required > 50% of pixels be an edge in a given row or column. Upon detecting grid lines, the tile was split into sub-tiles. Unfortunately, while this algo was correctly handling tables, it produced false splits in other cases and dropped the public LB score, despite improving the score on local validation. So it was turned off.\n2. Generated images like plots and diagrams were detected as tiles. They contained texts, axes and icons that were detected as matches resulting in false positives regions. To filter out such tiles, I've used heuristics based on grayscale histogram. The idea is simple: plots and diagrams use small amount of colors and thus their histograms will contains sharp spikes, while biological images will have smoother histograms.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F044312420e72aac43d133fe1c7128ce3%2FScreenshot%202026-05-17%20at%2021.00.09.png?generation=1779041695038957&alt=media =500x200) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F3ff60df7cc682dd9a87c0b8f709acd46%2FScreenshot%202026-05-17%20at%2021.00.52.png?generation=1779041718636493&alt=media =500x200)\n\n## Feature Extraction ##\nSuperpoint was used to detect 2048 most interesting points per tile and extract their features as 256-dim vectors.\n\nAny texts were producing points and resulting in false positives matches. So a separate post-processing was introduced to detect texts and remove any points associated with them. Text was detected with DB50 model.\n\n## Feature Matching ##\nLightGlue was used to match points from one tile with points from another tile. For some tile pairs the LightGlue was stopping matching early due to low match scores, resulting in under-detection. This is why I had to disable early stop explicitly, increased inference time was not a problem. The points were then filtered by match score threshold, leaving only the pairs of points with decent match confidence.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F94872de5e19f235aa306672a9e6b3a73%2FScreenshot%202026-05-17%20at%2021.03.50.png?generation=1779041639040268&alt=media =700x300)\n\n## RANSAC Affine Verification ##\nAt this stage there were pairs of matching points potentially containing some noisy matches from different parts of the image. So to further narrow down the matching region, an affine transform was fit to the points, requiring a group of points to behave as a whole: the position of matching points on tile T1 and T2 should be described with the same/similar affine transform. For this RANSAC method from OpenCV library was used, with requirements to have at least 20 inliers and at least 20% inlier ratio among points.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2F86a52cd3d9feb63f4a999b879f1d0fb6%2FScreenshot%202026-05-17%20at%2021.33.26.png?generation=1779042836917640&alt=media =700x300)\n\nThen minimum area rotated bounding box was computed around the inlier keypoints in each tile, followed by axis-aligned bounding boxes calculation.\n\n## Mask Generation & Merging ##\nBounding boxes for matched regions were converted to masks on the original image. \n\nSometimes mask were intersecting (for example the matching regions were A - B and A - C), in such case masks were merged into the same layer (requiring intersection-over-union to be at least 0.1).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582676%2Fa4b936dfc5664b01df86aed2a72c55f3%2FScreenshot%202026-05-17%20at%2020.47.08.png?generation=1779042269764739&alt=media =300x300)\n\n## Additional thoughts ##\nI want to thank the organizers and Kaggle for hosting this competition. \n\nAfter the submission phase ended, I still had hope for my solution despite being far outside the medal range on the public leaderboard. Unlike most public solutions based on DINO training, I used a point-matching approach, which I expected to be more robust to distribution shifts in new data. Still, finishing in 6th place came as a big surprise to me."
  }
}