{
  "id": 694442,
  "title": "65th Place Solution🥈|  DINOv2 with a Tiny Convolutional Decoder",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/writeups/65th-place-solution-dinov2-with-a-tiny-convolu",
  "author_name": "",
  "post_date": "2026-04-24T20:54:34.917Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First, a big thank you to the competition organizers and to the Recod.ai/LUC lab for putting this challenge together. <a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a> , <a href=\"https://www.kaggle.com/ashleyoldacre\" target=\"_blank\">@ashleyoldacre</a> </p>\n<p><strong>Scientific image forensics is exactly the kind of problem I enjoy working on:</strong> research-grounded, socially useful, and technically open-ended. Detecting image manipulation in biomedical figures has real implications for the integrity of published science, and building models in this space felt meaningful beyond the leaderboard. Thank you for a well-run competition and a thoughtfully curated dataset.</p>\n<h2>Summary</h2>\n<table>\n<thead>\n<tr>\n<th>Item</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Final rank</td>\n<td>65 / 1564 teams</td>\n</tr>\n<tr>\n<td>Medal</td>\n<td>Silver</td>\n</tr>\n<tr>\n<td>Architecture</td>\n<td>DINOv2 base (frozen + partial unfreeze) with a tiny convolutional decoder</td>\n</tr>\n<tr>\n<td>Input size</td>\n<td>518 × 518</td>\n</tr>\n<tr>\n<td>Training strategy</td>\n<td>Two-stage: decoder warmup, then joint fine-tuning</td>\n</tr>\n<tr>\n<td>Inference tricks</td>\n<td>Flip TTA, gradient-enhanced adaptive mask, grid-searched thresholds</td>\n</tr>\n</tbody>\n</table>\n<h2>What I tried, and what worked</h2>\n<p>Early on I explored a few directions: pure CNN segmenters (U-Net variants on EfficientNet and ResNet backbones), standard ViT features with a linear probe, and classical copy-move detectors based on keypoint matching. None of these matched the performance of a DINOv2-based approach, and in hindsight the reason is intuitive.</p>\n<p>Copy-move forgery is defined by two regions in an image being <em>semantically and texturally identical</em> to each other. You are not looking for an out-of-distribution object or a rendering artifact. You are looking for self-similarity. The features you want are ones that (a) are sensitive to fine-grained texture and (b) are stable enough that the same texture produces the same feature vector wherever it appears in the image. DINOv2 was trained with a self-distillation objective specifically designed to produce representations with this property. Its patch features are dense, locality-preserving, and strong out of the box, even on domains it was not trained on.</p>\n<p>So the setup that worked best is straightforward: use DINOv2 as a frozen feature extractor, put a small decoder on top, and train only the decoder (then carefully unfreeze the upper layers of DINOv2 in a second stage).</p>\n<h2>Architecture</h2>\n<p>DINOv2 base produces 768-dimensional features on a 37 × 37 grid for a 518 × 518 input. That feature map already carries most of the spatial information needed to localize a forgery. A heavy decoder would mostly add parameters without adding signal, so I kept the decoder small.</p>\n<p>The decoder has three convolutional blocks that progressively halve the channel width: 768 → 384 → 192 → 96, each followed by ReLU, with dropout 0.1 in the first two blocks. A final 1 × 1 convolution projects to a single logit channel. Between blocks the feature map is bilinearly upsampled (37 → 74 → 148 → 296 → 518) so the output is at full input resolution.</p>\n<p>If anyone would like a clearer picture of the flow, I can add an architecture diagram, but the core idea is: frozen DINOv2 on the bottom, small conv decoder with progressive upsampling on top, BCE loss on the full-resolution logit map.</p>\n<h2>Two-stage training</h2>\n<p>DINOv2 starts fully frozen. Training runs in two stages.</p>\n<p><strong>Stage 1 is decoder warmup.</strong> Only the decoder parameters go to the optimizer. The backbone is kept in eval mode during the forward pass, and its <code>requires_grad</code> flags are all false, so nothing in DINOv2 gets updated. This stage converges quickly. Validation loss drops in the first few epochs and plateaus around epoch 12 to 15. Early stopping with patience 3 handles the plateau.</p>\n<p><strong>Stage 2 is joint fine-tuning.</strong> The last 12 transformer blocks of DINOv2 are unfrozen. The optimizer now has two parameter groups: the decoder keeps its stage 1 learning rate (1e-5), and the backbone gets a much smaller one (5e-7). The reason for the gap is that the pretrained DINOv2 features are already good. I want to nudge the upper layers toward biomedical imagery, not rewrite them. A higher backbone learning rate would overwhelm the pretraining signal and collapse the features.</p>\n<p>Both stages use AdamW with weight decay 1e-4, cosine learning rate scheduling, and gradient accumulation over 8 micro-batches (effective batch size 16). The best checkpoint on validation loss is saved to disk, and stage 2 only overwrites it if it beats the stage 1 best. If stage 2 regresses, reloading at the end recovers the stage 1 model automatically, a small safety net that matters when fine-tuning large pretrained backbones, because it is easy to make things worse.</p>\n<h2>Inference pipeline</h2>\n<p>The inference side of the pipeline is where a fair amount of the score comes from, and it is worth going through in detail.</p>\n<p><strong>Test-time augmentation.</strong> The test image is passed through the model three times: as-is, horizontally flipped, and vertically flipped. The predictions from the flipped versions are flipped back before averaging. The final probability map is the mean of the three. This is essentially free at inference time and reliably smooths out spurious activations.</p>\n<p><strong>Adaptive mask from the probability map.</strong> Rather than thresholding the probability map directly, I enhance it with its own spatial gradient. Sobel gradients are computed on the probability map, their magnitude is normalized to [0, 1], and the enhanced map is a weighted blend:</p>\n<pre><code>enhanced = (1 - alpha) * prob + alpha * gradient_magnitude\n</code></pre>\n<p>with <code>alpha = 0.45</code>. This sharpens the boundaries of high-confidence regions. The gradient is largest where probability changes most, so blending it in amplifies edges. The enhanced map is then lightly Gaussian-blurred and thresholded at <code>mean + 0.3 * std</code>. Morphological close (5 × 5) then open (3 × 3) removes speckle and fills small gaps.</p>\n<p><strong>Area and probability thresholds.</strong> A raw mask is only returned as a prediction if two conditions hold: the mask has at least <code>AREA_MIN</code> foreground pixels, and the mean probability inside the mask is at least <code>PROB_MIN</code>. Otherwise the image is classified as authentic and the submission writes the string <code>\"authentic\"</code> for that row. These two thresholds are grid-searched over the entire validation split (both forged and authentic images) to maximize the competition F1. The tuned values landed in the range <code>AREA_MIN ≈ 200</code> and <code>PROB_MIN ≈ 0.20</code> to <code>0.22</code>, though they shift slightly with the <code>alpha</code> choice above.</p>\n<p><strong>Submission format.</strong> Forged predictions are run-length encoded in the exact format the competition expects; authentic predictions are the literal string <code>\"authentic\"</code>.</p>\n<h2>Reproducing the result</h2>\n<p>Everything needed to reproduce the submission is on Kaggle and GitHub. The fastest path is to open the inference notebook and use Copy and Edit. The data, the DINOv2 base model, and the trained weights are already attached.</p>\n<h2>Thanks</h2>\n<p>Thanks again to the organizers and to the Recod.ai/LUC team for hosting. Thanks to Dr. Elisabeth Bik for the early guidance that shaped the dataset, and to the Fapesp Horus and CNPq Aletheia teams for the technical support behind it.</p>\n<p>A particular thank you to <a href=\"https://www.kaggle.com/pankajiitr\" target=\"_blank\">@pankajiitr</a> , whose public work and shared notebook were genuinely helpful throughout this competition. Open sharing in the Kaggle community is what makes competitions like this valuable beyond the leaderboard, and I benefited from it directly.</p>\n<p>Congratulations to everyone who competed.</p>",
  "messages": [
    {
      "id": "3448187",
      "postDate": "04/24/2026 20:53:32",
      "content": "<p>First, a big thank you to the competition organizers and to the Recod.ai/LUC lab for putting this challenge together. <a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a> , <a href=\"https://www.kaggle.com/ashleyoldacre\" target=\"_blank\">@ashleyoldacre</a> </p>\n<p><strong>Scientific image forensics is exactly the kind of problem I enjoy working on:</strong> research-grounded, socially useful, and technically open-ended. Detecting image manipulation in biomedical figures has real implications for the integrity of published science, and building models in this space felt meaningful beyond the leaderboard. Thank you for a well-run competition and a thoughtfully curated dataset.</p>\n<h2>Summary</h2>\n<table>\n<thead>\n<tr>\n<th>Item</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Final rank</td>\n<td>65 / 1564 teams</td>\n</tr>\n<tr>\n<td>Medal</td>\n<td>Silver</td>\n</tr>\n<tr>\n<td>Architecture</td>\n<td>DINOv2 base (frozen + partial unfreeze) with a tiny convolutional decoder</td>\n</tr>\n<tr>\n<td>Input size</td>\n<td>518 × 518</td>\n</tr>\n<tr>\n<td>Training strategy</td>\n<td>Two-stage: decoder warmup, then joint fine-tuning</td>\n</tr>\n<tr>\n<td>Inference tricks</td>\n<td>Flip TTA, gradient-enhanced adaptive mask, grid-searched thresholds</td>\n</tr>\n</tbody>\n</table>\n<h2>What I tried, and what worked</h2>\n<p>Early on I explored a few directions: pure CNN segmenters (U-Net variants on EfficientNet and ResNet backbones), standard ViT features with a linear probe, and classical copy-move detectors based on keypoint matching. None of these matched the performance of a DINOv2-based approach, and in hindsight the reason is intuitive.</p>\n<p>Copy-move forgery is defined by two regions in an image being <em>semantically and texturally identical</em> to each other. You are not looking for an out-of-distribution object or a rendering artifact. You are looking for self-similarity. The features you want are ones that (a) are sensitive to fine-grained texture and (b) are stable enough that the same texture produces the same feature vector wherever it appears in the image. DINOv2 was trained with a self-distillation objective specifically designed to produce representations with this property. Its patch features are dense, locality-preserving, and strong out of the box, even on domains it was not trained on.</p>\n<p>So the setup that worked best is straightforward: use DINOv2 as a frozen feature extractor, put a small decoder on top, and train only the decoder (then carefully unfreeze the upper layers of DINOv2 in a second stage).</p>\n<h2>Architecture</h2>\n<p>DINOv2 base produces 768-dimensional features on a 37 × 37 grid for a 518 × 518 input. That feature map already carries most of the spatial information needed to localize a forgery. A heavy decoder would mostly add parameters without adding signal, so I kept the decoder small.</p>\n<p>The decoder has three convolutional blocks that progressively halve the channel width: 768 → 384 → 192 → 96, each followed by ReLU, with dropout 0.1 in the first two blocks. A final 1 × 1 convolution projects to a single logit channel. Between blocks the feature map is bilinearly upsampled (37 → 74 → 148 → 296 → 518) so the output is at full input resolution.</p>\n<p>If anyone would like a clearer picture of the flow, I can add an architecture diagram, but the core idea is: frozen DINOv2 on the bottom, small conv decoder with progressive upsampling on top, BCE loss on the full-resolution logit map.</p>\n<h2>Two-stage training</h2>\n<p>DINOv2 starts fully frozen. Training runs in two stages.</p>\n<p><strong>Stage 1 is decoder warmup.</strong> Only the decoder parameters go to the optimizer. The backbone is kept in eval mode during the forward pass, and its <code>requires_grad</code> flags are all false, so nothing in DINOv2 gets updated. This stage converges quickly. Validation loss drops in the first few epochs and plateaus around epoch 12 to 15. Early stopping with patience 3 handles the plateau.</p>\n<p><strong>Stage 2 is joint fine-tuning.</strong> The last 12 transformer blocks of DINOv2 are unfrozen. The optimizer now has two parameter groups: the decoder keeps its stage 1 learning rate (1e-5), and the backbone gets a much smaller one (5e-7). The reason for the gap is that the pretrained DINOv2 features are already good. I want to nudge the upper layers toward biomedical imagery, not rewrite them. A higher backbone learning rate would overwhelm the pretraining signal and collapse the features.</p>\n<p>Both stages use AdamW with weight decay 1e-4, cosine learning rate scheduling, and gradient accumulation over 8 micro-batches (effective batch size 16). The best checkpoint on validation loss is saved to disk, and stage 2 only overwrites it if it beats the stage 1 best. If stage 2 regresses, reloading at the end recovers the stage 1 model automatically, a small safety net that matters when fine-tuning large pretrained backbones, because it is easy to make things worse.</p>\n<h2>Inference pipeline</h2>\n<p>The inference side of the pipeline is where a fair amount of the score comes from, and it is worth going through in detail.</p>\n<p><strong>Test-time augmentation.</strong> The test image is passed through the model three times: as-is, horizontally flipped, and vertically flipped. The predictions from the flipped versions are flipped back before averaging. The final probability map is the mean of the three. This is essentially free at inference time and reliably smooths out spurious activations.</p>\n<p><strong>Adaptive mask from the probability map.</strong> Rather than thresholding the probability map directly, I enhance it with its own spatial gradient. Sobel gradients are computed on the probability map, their magnitude is normalized to [0, 1], and the enhanced map is a weighted blend:</p>\n<pre><code>enhanced = (1 - alpha) * prob + alpha * gradient_magnitude\n</code></pre>\n<p>with <code>alpha = 0.45</code>. This sharpens the boundaries of high-confidence regions. The gradient is largest where probability changes most, so blending it in amplifies edges. The enhanced map is then lightly Gaussian-blurred and thresholded at <code>mean + 0.3 * std</code>. Morphological close (5 × 5) then open (3 × 3) removes speckle and fills small gaps.</p>\n<p><strong>Area and probability thresholds.</strong> A raw mask is only returned as a prediction if two conditions hold: the mask has at least <code>AREA_MIN</code> foreground pixels, and the mean probability inside the mask is at least <code>PROB_MIN</code>. Otherwise the image is classified as authentic and the submission writes the string <code>\"authentic\"</code> for that row. These two thresholds are grid-searched over the entire validation split (both forged and authentic images) to maximize the competition F1. The tuned values landed in the range <code>AREA_MIN ≈ 200</code> and <code>PROB_MIN ≈ 0.20</code> to <code>0.22</code>, though they shift slightly with the <code>alpha</code> choice above.</p>\n<p><strong>Submission format.</strong> Forged predictions are run-length encoded in the exact format the competition expects; authentic predictions are the literal string <code>\"authentic\"</code>.</p>\n<h2>Reproducing the result</h2>\n<p>Everything needed to reproduce the submission is on Kaggle and GitHub. The fastest path is to open the inference notebook and use Copy and Edit. The data, the DINOv2 base model, and the trained weights are already attached.</p>\n<h2>Thanks</h2>\n<p>Thanks again to the organizers and to the Recod.ai/LUC team for hosting. Thanks to Dr. Elisabeth Bik for the early guidance that shaped the dataset, and to the Fapesp Horus and CNPq Aletheia teams for the technical support behind it.</p>\n<p>A particular thank you to <a href=\"https://www.kaggle.com/pankajiitr\" target=\"_blank\">@pankajiitr</a> , whose public work and shared notebook were genuinely helpful throughout this competition. Open sharing in the Kaggle community is what makes competitions like this valuable beyond the leaderboard, and I benefited from it directly.</p>\n<p>Congratulations to everyone who competed.</p>",
      "rawMarkdown": "First, a big thank you to the competition organizers and to the Recod.ai/LUC lab for putting this challenge together. @joophillipecardenuto , @ashleyoldacre \n\n**Scientific image forensics is exactly the kind of problem I enjoy working on:** research-grounded, socially useful, and technically open-ended. Detecting image manipulation in biomedical figures has real implications for the integrity of published science, and building models in this space felt meaningful beyond the leaderboard. Thank you for a well-run competition and a thoughtfully curated dataset.\n\n\n## Summary\n\n| Item | Value |\n|---|---|\n| Final rank | 65 / 1564 teams |\n| Medal | Silver |\n| Architecture | DINOv2 base (frozen + partial unfreeze) with a tiny convolutional decoder |\n| Input size | 518 × 518 |\n| Training strategy | Two-stage: decoder warmup, then joint fine-tuning |\n| Inference tricks | Flip TTA, gradient-enhanced adaptive mask, grid-searched thresholds |\n\n\n\n## What I tried, and what worked\n\nEarly on I explored a few directions: pure CNN segmenters (U-Net variants on EfficientNet and ResNet backbones), standard ViT features with a linear probe, and classical copy-move detectors based on keypoint matching. None of these matched the performance of a DINOv2-based approach, and in hindsight the reason is intuitive.\n\nCopy-move forgery is defined by two regions in an image being *semantically and texturally identical* to each other. You are not looking for an out-of-distribution object or a rendering artifact. You are looking for self-similarity. The features you want are ones that (a) are sensitive to fine-grained texture and (b) are stable enough that the same texture produces the same feature vector wherever it appears in the image. DINOv2 was trained with a self-distillation objective specifically designed to produce representations with this property. Its patch features are dense, locality-preserving, and strong out of the box, even on domains it was not trained on.\n\nSo the setup that worked best is straightforward: use DINOv2 as a frozen feature extractor, put a small decoder on top, and train only the decoder (then carefully unfreeze the upper layers of DINOv2 in a second stage).\n\n\n\n## Architecture\n\nDINOv2 base produces 768-dimensional features on a 37 × 37 grid for a 518 × 518 input. That feature map already carries most of the spatial information needed to localize a forgery. A heavy decoder would mostly add parameters without adding signal, so I kept the decoder small.\n\nThe decoder has three convolutional blocks that progressively halve the channel width: 768 → 384 → 192 → 96, each followed by ReLU, with dropout 0.1 in the first two blocks. A final 1 × 1 convolution projects to a single logit channel. Between blocks the feature map is bilinearly upsampled (37 → 74 → 148 → 296 → 518) so the output is at full input resolution.\n\nIf anyone would like a clearer picture of the flow, I can add an architecture diagram, but the core idea is: frozen DINOv2 on the bottom, small conv decoder with progressive upsampling on top, BCE loss on the full-resolution logit map.\n\n\n## Two-stage training\n\nDINOv2 starts fully frozen. Training runs in two stages.\n\n**Stage 1 is decoder warmup.** Only the decoder parameters go to the optimizer. The backbone is kept in eval mode during the forward pass, and its `requires_grad` flags are all false, so nothing in DINOv2 gets updated. This stage converges quickly. Validation loss drops in the first few epochs and plateaus around epoch 12 to 15. Early stopping with patience 3 handles the plateau.\n\n**Stage 2 is joint fine-tuning.** The last 12 transformer blocks of DINOv2 are unfrozen. The optimizer now has two parameter groups: the decoder keeps its stage 1 learning rate (1e-5), and the backbone gets a much smaller one (5e-7). The reason for the gap is that the pretrained DINOv2 features are already good. I want to nudge the upper layers toward biomedical imagery, not rewrite them. A higher backbone learning rate would overwhelm the pretraining signal and collapse the features.\n\nBoth stages use AdamW with weight decay 1e-4, cosine learning rate scheduling, and gradient accumulation over 8 micro-batches (effective batch size 16). The best checkpoint on validation loss is saved to disk, and stage 2 only overwrites it if it beats the stage 1 best. If stage 2 regresses, reloading at the end recovers the stage 1 model automatically, a small safety net that matters when fine-tuning large pretrained backbones, because it is easy to make things worse.\n\n\n## Inference pipeline\n\nThe inference side of the pipeline is where a fair amount of the score comes from, and it is worth going through in detail.\n\n**Test-time augmentation.** The test image is passed through the model three times: as-is, horizontally flipped, and vertically flipped. The predictions from the flipped versions are flipped back before averaging. The final probability map is the mean of the three. This is essentially free at inference time and reliably smooths out spurious activations.\n\n**Adaptive mask from the probability map.** Rather than thresholding the probability map directly, I enhance it with its own spatial gradient. Sobel gradients are computed on the probability map, their magnitude is normalized to [0, 1], and the enhanced map is a weighted blend:\n\n```\nenhanced = (1 - alpha) * prob + alpha * gradient_magnitude\n```\n\nwith `alpha = 0.45`. This sharpens the boundaries of high-confidence regions. The gradient is largest where probability changes most, so blending it in amplifies edges. The enhanced map is then lightly Gaussian-blurred and thresholded at `mean + 0.3 * std`. Morphological close (5 × 5) then open (3 × 3) removes speckle and fills small gaps.\n\n**Area and probability thresholds.** A raw mask is only returned as a prediction if two conditions hold: the mask has at least `AREA_MIN` foreground pixels, and the mean probability inside the mask is at least `PROB_MIN`. Otherwise the image is classified as authentic and the submission writes the string `\"authentic\"` for that row. These two thresholds are grid-searched over the entire validation split (both forged and authentic images) to maximize the competition F1. The tuned values landed in the range `AREA_MIN ≈ 200` and `PROB_MIN ≈ 0.20` to `0.22`, though they shift slightly with the `alpha` choice above.\n\n**Submission format.** Forged predictions are run-length encoded in the exact format the competition expects; authentic predictions are the literal string `\"authentic\"`.\n\n\n\n## Reproducing the result\n\nEverything needed to reproduce the submission is on Kaggle and GitHub. The fastest path is to open the inference notebook and use Copy and Edit. The data, the DINOv2 base model, and the trained weights are already attached.\n\n\n## Thanks\n\nThanks again to the organizers and to the Recod.ai/LUC team for hosting. Thanks to Dr. Elisabeth Bik for the early guidance that shaped the dataset, and to the Fapesp Horus and CNPq Aletheia teams for the technical support behind it.\n\nA particular thank you to @pankajiitr , whose public work and shared notebook were genuinely helpful throughout this competition. Open sharing in the Kaggle community is what makes competitions like this valuable beyond the leaderboard, and I benefited from it directly.\n\nCongratulations to everyone who competed.",
      "votes": null
    },
    {
      "id": "3448607",
      "postDate": "04/26/2026 05:56:46",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "3450176",
      "postDate": "04/29/2026 12:49:31",
      "content": "<p>Thanks for the kind words! It’s awesome to see that my shared work was helpful for your approach. Congrats again on the silver!</p>",
      "rawMarkdown": "Thanks for the kind words! It’s awesome to see that my shared work was helpful for your approach. Congrats again on the silver!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3448607,
      "author_name": "",
      "author_url": "",
      "post_date": "04/26/2026 05:56:46",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3450176,
      "author_name": "pankajiitr",
      "author_url": "",
      "post_date": "04/29/2026 12:49:31",
      "content": "<p>Thanks for the kind words! It’s awesome to see that my shared work was helpful for your approach. Congrats again on the silver!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3448187": "First, a big thank you to the competition organizers and to the Recod.ai/LUC lab for putting this challenge together. @joophillipecardenuto , @ashleyoldacre \n\n**Scientific image forensics is exactly the kind of problem I enjoy working on:** research-grounded, socially useful, and technically open-ended. Detecting image manipulation in biomedical figures has real implications for the integrity of published science, and building models in this space felt meaningful beyond the leaderboard. Thank you for a well-run competition and a thoughtfully curated dataset.\n\n\n## Summary\n\n| Item | Value |\n|---|---|\n| Final rank | 65 / 1564 teams |\n| Medal | Silver |\n| Architecture | DINOv2 base (frozen + partial unfreeze) with a tiny convolutional decoder |\n| Input size | 518 × 518 |\n| Training strategy | Two-stage: decoder warmup, then joint fine-tuning |\n| Inference tricks | Flip TTA, gradient-enhanced adaptive mask, grid-searched thresholds |\n\n\n\n## What I tried, and what worked\n\nEarly on I explored a few directions: pure CNN segmenters (U-Net variants on EfficientNet and ResNet backbones), standard ViT features with a linear probe, and classical copy-move detectors based on keypoint matching. None of these matched the performance of a DINOv2-based approach, and in hindsight the reason is intuitive.\n\nCopy-move forgery is defined by two regions in an image being *semantically and texturally identical* to each other. You are not looking for an out-of-distribution object or a rendering artifact. You are looking for self-similarity. The features you want are ones that (a) are sensitive to fine-grained texture and (b) are stable enough that the same texture produces the same feature vector wherever it appears in the image. DINOv2 was trained with a self-distillation objective specifically designed to produce representations with this property. Its patch features are dense, locality-preserving, and strong out of the box, even on domains it was not trained on.\n\nSo the setup that worked best is straightforward: use DINOv2 as a frozen feature extractor, put a small decoder on top, and train only the decoder (then carefully unfreeze the upper layers of DINOv2 in a second stage).\n\n\n\n## Architecture\n\nDINOv2 base produces 768-dimensional features on a 37 × 37 grid for a 518 × 518 input. That feature map already carries most of the spatial information needed to localize a forgery. A heavy decoder would mostly add parameters without adding signal, so I kept the decoder small.\n\nThe decoder has three convolutional blocks that progressively halve the channel width: 768 → 384 → 192 → 96, each followed by ReLU, with dropout 0.1 in the first two blocks. A final 1 × 1 convolution projects to a single logit channel. Between blocks the feature map is bilinearly upsampled (37 → 74 → 148 → 296 → 518) so the output is at full input resolution.\n\nIf anyone would like a clearer picture of the flow, I can add an architecture diagram, but the core idea is: frozen DINOv2 on the bottom, small conv decoder with progressive upsampling on top, BCE loss on the full-resolution logit map.\n\n\n## Two-stage training\n\nDINOv2 starts fully frozen. Training runs in two stages.\n\n**Stage 1 is decoder warmup.** Only the decoder parameters go to the optimizer. The backbone is kept in eval mode during the forward pass, and its `requires_grad` flags are all false, so nothing in DINOv2 gets updated. This stage converges quickly. Validation loss drops in the first few epochs and plateaus around epoch 12 to 15. Early stopping with patience 3 handles the plateau.\n\n**Stage 2 is joint fine-tuning.** The last 12 transformer blocks of DINOv2 are unfrozen. The optimizer now has two parameter groups: the decoder keeps its stage 1 learning rate (1e-5), and the backbone gets a much smaller one (5e-7). The reason for the gap is that the pretrained DINOv2 features are already good. I want to nudge the upper layers toward biomedical imagery, not rewrite them. A higher backbone learning rate would overwhelm the pretraining signal and collapse the features.\n\nBoth stages use AdamW with weight decay 1e-4, cosine learning rate scheduling, and gradient accumulation over 8 micro-batches (effective batch size 16). The best checkpoint on validation loss is saved to disk, and stage 2 only overwrites it if it beats the stage 1 best. If stage 2 regresses, reloading at the end recovers the stage 1 model automatically, a small safety net that matters when fine-tuning large pretrained backbones, because it is easy to make things worse.\n\n\n## Inference pipeline\n\nThe inference side of the pipeline is where a fair amount of the score comes from, and it is worth going through in detail.\n\n**Test-time augmentation.** The test image is passed through the model three times: as-is, horizontally flipped, and vertically flipped. The predictions from the flipped versions are flipped back before averaging. The final probability map is the mean of the three. This is essentially free at inference time and reliably smooths out spurious activations.\n\n**Adaptive mask from the probability map.** Rather than thresholding the probability map directly, I enhance it with its own spatial gradient. Sobel gradients are computed on the probability map, their magnitude is normalized to [0, 1], and the enhanced map is a weighted blend:\n\n```\nenhanced = (1 - alpha) * prob + alpha * gradient_magnitude\n```\n\nwith `alpha = 0.45`. This sharpens the boundaries of high-confidence regions. The gradient is largest where probability changes most, so blending it in amplifies edges. The enhanced map is then lightly Gaussian-blurred and thresholded at `mean + 0.3 * std`. Morphological close (5 × 5) then open (3 × 3) removes speckle and fills small gaps.\n\n**Area and probability thresholds.** A raw mask is only returned as a prediction if two conditions hold: the mask has at least `AREA_MIN` foreground pixels, and the mean probability inside the mask is at least `PROB_MIN`. Otherwise the image is classified as authentic and the submission writes the string `\"authentic\"` for that row. These two thresholds are grid-searched over the entire validation split (both forged and authentic images) to maximize the competition F1. The tuned values landed in the range `AREA_MIN ≈ 200` and `PROB_MIN ≈ 0.20` to `0.22`, though they shift slightly with the `alpha` choice above.\n\n**Submission format.** Forged predictions are run-length encoded in the exact format the competition expects; authentic predictions are the literal string `\"authentic\"`.\n\n\n\n## Reproducing the result\n\nEverything needed to reproduce the submission is on Kaggle and GitHub. The fastest path is to open the inference notebook and use Copy and Edit. The data, the DINOv2 base model, and the trained weights are already attached.\n\n\n## Thanks\n\nThanks again to the organizers and to the Recod.ai/LUC team for hosting. Thanks to Dr. Elisabeth Bik for the early guidance that shaped the dataset, and to the Fapesp Horus and CNPq Aletheia teams for the technical support behind it.\n\nA particular thank you to @pankajiitr , whose public work and shared notebook were genuinely helpful throughout this competition. Open sharing in the Kaggle community is what makes competitions like this valuable beyond the leaderboard, and I benefited from it directly.\n\nCongratulations to everyone who competed.",
    "3448607": "Congratulations!",
    "3450176": "Thanks for the kind words! It’s awesome to see that my shared work was helpful for your approach. Congrats again on the silver!"
  },
  "source": "meta"
}