{
  "id": 697699,
  "title": "5th Place Solution — DINOv3 Semantic Segmentation",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/697699",
  "author_name": "Guanshuo Xu",
  "post_date": "2026-05-06T23:20:47.997000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I would like to thank Kaggle and Recod.ai/LUC for hosting this interesting competition. I finished 5th on the private leaderboard.</p>\n<p><strong>Final scores:</strong></p>\n<ul>\n<li>Public LB: <strong>0.422</strong></li>\n<li>Private LB: <strong>0.376</strong></li>\n<li><a href=\"https://www.kaggle.com/code/wowfattie/recodfinal\" target=\"_blank\"><strong>inference code</strong></a></li>\n<li><strong>training code attached at the end</strong></li>\n</ul>\n<h2>Overview</h2>\n<p>My solution focused on simplicity and robustness. I treated the task primarily as a semantic segmentation problem.</p>\n<p>For training mask generation, when an image contained multiple copy-move instances, I used the union of all manipulated regions as the training target. In other words, the model was trained as a binary manipulated-region segmenter rather than as an instance segmentation model.</p>\n<p>At inference time, I also did not separate multiple copy-move instances. This was a deliberate simplification.</p>\n<p>Two factors contributed most significantly to the final score:</p>\n<ol>\n<li>Using external datasets to improve generalization.</li>\n<li>Using the strongest available pretrained DINOv3 backbone.</li>\n</ol>\n<h2>Validation</h2>\n<p>I mainly validated my models on the 48 supplemental images provided by the host, because these multi-panel images were the closest match to the actual test set.</p>\n<h2>Backbone and Segmentation Head</h2>\n<p>I used pretrained DINOv2/DINOv3 backbones with a very simple segmentation head.</p>\n<p>The model takes the final patch tokens from the ViT backbone, reshapes them into a 2D feature grid, and applies a single 1×1 Conv2d head to produce one logit map.</p>\n<p>During training, I downsampled the ground-truth mask to match the patch-grid resolution:</p>\n<ul>\n<li>14×14 for DINOv2</li>\n<li>16×16 for DINOv3</li>\n</ul>\n<p>During inference, after predicting the low-resolution mask, I simply resized it back to the original image size using bilinear interpolation.</p>\n<h2>Augmentations</h2>\n<p>The most useful augmentations were simple image-level augmentations:</p>\n<pre><code>A.RandomRotate90(p=1.0)\nA.HorizontalFlip(p=0.5)\nA.GaussianBlur(p=0.15)\nA.ToGray(...)\nA.HueSaturationValue(...)\nA.RandomBrightnessContrast(p=0.9)\nA.ImageCompression(p=0.25)\n</code></pre>\n<p>I used fairly aggressive color augmentation because the manipulation signal is not always color-specific. The model needed to learn structural and local consistency cues rather than memorize color distributions.</p>\n<h2>Important Training Choices</h2>\n<ul>\n<li>BF16 autocast</li>\n<li>Full fine-tuning</li>\n<li>FlashAttention-2</li>\n<li>Input image size: 1024 × 1024</li>\n<li>Vanilla BCE loss</li>\n</ul>\n<h2>Important Inference Choices</h2>\n<ul>\n<li>8-bit inference to fit within the 16GB memory limit of a T4 GPU</li>\n<li>SDPA attention</li>\n<li>Input image size: 1024 × 1024</li>\n<li>Four-way <code>rotate90</code> test-time augmentation</li>\n</ul>\n<h2>Post-processing</h2>\n<p>The post-processing pipeline was simple:</p>\n<ol>\n<li>Create a high-confidence binary mask using a fixed threshold.</li>\n<li>Compute the area of this high-confidence mask by summing its active pixels.</li>\n<li>If the high-confidence area is smaller than <code>min_area</code>, classify the sample as authentic and stop.</li>\n<li>Create the final binary mask using the configured threshold <code>mask_thr</code>.</li>\n<li>If the final mask contains no active pixels, classify the sample as authentic and stop.</li>\n</ol>\n<h2>Early Results</h2>\n<p>I started with a DINOv2 model. Its validation score on the supplemental images was only <strong>0.092</strong>, although the validation score on a split of the provided training set was much higher.</p>\n<p>After inspecting the training data, I found that although the dataset was not small, it contained only a limited variety of image types. As a result, the fine-tuned model did not generalize well to the supplemental set, which was visually quite different from the training set.</p>\n<h2>External Datasets</h2>\n<p>To improve generalization, I searched for external datasets that could increase the diversity of the training data. I eventually added the following datasets:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/divg07/casia-20-image-tampering-detection-dataset\" target=\"_blank\">CASIA</a></li>\n<li><a href=\"https://www.grip.unina.it/download/prog/CMFD/\" target=\"_blank\">GRIP</a></li>\n<li><a href=\"https://www.cs1.tf.fau.de/research/multimedia-security/code/image-manipulation-dataset/#collapse_1\" target=\"_blank\">FAU</a></li>\n</ul>\n<p>Validation results on the supplemental images were:</p>\n<table>\n<thead>\n<tr>\n<th>Training data</th>\n<th>Validation score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Competition training set only</td>\n<td>0.092</td>\n</tr>\n<tr>\n<td>Competition training set + CASIA</td>\n<td>0.185</td>\n</tr>\n<tr>\n<td>Competition training set + GRIP</td>\n<td>0.095</td>\n</tr>\n<tr>\n<td>Competition training set + FAU</td>\n<td>0.129</td>\n</tr>\n<tr>\n<td>Competition training set + CASIA + GRIP + FAU</td>\n<td>0.240</td>\n</tr>\n</tbody>\n</table>\n<p>Although these external datasets are not biomedical, their diversity helped the model learn more general copy-move detection cues.</p>\n<h2>Better Backbones</h2>\n<p>Because copy-move detection is a difficult task, stronger pretrained backbones made a large difference.</p>\n<p>Upgrading the backbone from <strong>DINOv2 Giant</strong> to <strong>DINOv3 Huge+</strong> improved the validation score from the earlier range to around <strong>0.35</strong>. Since these two models are relatively close in size, I believe the improvement mainly came from the stronger pretraining of DINOv3.</p>\n<p>I then used the original <strong>DINOv3 7B</strong> model, which further improved the validation score to <strong>0.51</strong>. With additional hyperparameter tuning, the best validation score exceeded <strong>0.56</strong>.</p>\n<p>Overall, the final solution remained simple: a strong pretrained ViT backbone, a minimal segmentation head, diverse training data, and lightweight post-processing.</p>",
  "messages": [
    {
      "id": 3454335,
      "postDate": "2026-05-06T23:20:47.997Z",
      "content": "<p>I would like to thank Kaggle and Recod.ai/LUC for hosting this interesting competition. I finished 5th on the private leaderboard.</p>\n<p><strong>Final scores:</strong></p>\n<ul>\n<li>Public LB: <strong>0.422</strong></li>\n<li>Private LB: <strong>0.376</strong></li>\n<li><a href=\"https://www.kaggle.com/code/wowfattie/recodfinal\" target=\"_blank\"><strong>inference code</strong></a></li>\n<li><strong>training code attached at the end</strong></li>\n</ul>\n<h2>Overview</h2>\n<p>My solution focused on simplicity and robustness. I treated the task primarily as a semantic segmentation problem.</p>\n<p>For training mask generation, when an image contained multiple copy-move instances, I used the union of all manipulated regions as the training target. In other words, the model was trained as a binary manipulated-region segmenter rather than as an instance segmentation model.</p>\n<p>At inference time, I also did not separate multiple copy-move instances. This was a deliberate simplification.</p>\n<p>Two factors contributed most significantly to the final score:</p>\n<ol>\n<li>Using external datasets to improve generalization.</li>\n<li>Using the strongest available pretrained DINOv3 backbone.</li>\n</ol>\n<h2>Validation</h2>\n<p>I mainly validated my models on the 48 supplemental images provided by the host, because these multi-panel images were the closest match to the actual test set.</p>\n<h2>Backbone and Segmentation Head</h2>\n<p>I used pretrained DINOv2/DINOv3 backbones with a very simple segmentation head.</p>\n<p>The model takes the final patch tokens from the ViT backbone, reshapes them into a 2D feature grid, and applies a single 1×1 Conv2d head to produce one logit map.</p>\n<p>During training, I downsampled the ground-truth mask to match the patch-grid resolution:</p>\n<ul>\n<li>14×14 for DINOv2</li>\n<li>16×16 for DINOv3</li>\n</ul>\n<p>During inference, after predicting the low-resolution mask, I simply resized it back to the original image size using bilinear interpolation.</p>\n<h2>Augmentations</h2>\n<p>The most useful augmentations were simple image-level augmentations:</p>\n<pre><code>A.RandomRotate90(p=1.0)\nA.HorizontalFlip(p=0.5)\nA.GaussianBlur(p=0.15)\nA.ToGray(...)\nA.HueSaturationValue(...)\nA.RandomBrightnessContrast(p=0.9)\nA.ImageCompression(p=0.25)\n</code></pre>\n<p>I used fairly aggressive color augmentation because the manipulation signal is not always color-specific. The model needed to learn structural and local consistency cues rather than memorize color distributions.</p>\n<h2>Important Training Choices</h2>\n<ul>\n<li>BF16 autocast</li>\n<li>Full fine-tuning</li>\n<li>FlashAttention-2</li>\n<li>Input image size: 1024 × 1024</li>\n<li>Vanilla BCE loss</li>\n</ul>\n<h2>Important Inference Choices</h2>\n<ul>\n<li>8-bit inference to fit within the 16GB memory limit of a T4 GPU</li>\n<li>SDPA attention</li>\n<li>Input image size: 1024 × 1024</li>\n<li>Four-way <code>rotate90</code> test-time augmentation</li>\n</ul>\n<h2>Post-processing</h2>\n<p>The post-processing pipeline was simple:</p>\n<ol>\n<li>Create a high-confidence binary mask using a fixed threshold.</li>\n<li>Compute the area of this high-confidence mask by summing its active pixels.</li>\n<li>If the high-confidence area is smaller than <code>min_area</code>, classify the sample as authentic and stop.</li>\n<li>Create the final binary mask using the configured threshold <code>mask_thr</code>.</li>\n<li>If the final mask contains no active pixels, classify the sample as authentic and stop.</li>\n</ol>\n<h2>Early Results</h2>\n<p>I started with a DINOv2 model. Its validation score on the supplemental images was only <strong>0.092</strong>, although the validation score on a split of the provided training set was much higher.</p>\n<p>After inspecting the training data, I found that although the dataset was not small, it contained only a limited variety of image types. As a result, the fine-tuned model did not generalize well to the supplemental set, which was visually quite different from the training set.</p>\n<h2>External Datasets</h2>\n<p>To improve generalization, I searched for external datasets that could increase the diversity of the training data. I eventually added the following datasets:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/divg07/casia-20-image-tampering-detection-dataset\" target=\"_blank\">CASIA</a></li>\n<li><a href=\"https://www.grip.unina.it/download/prog/CMFD/\" target=\"_blank\">GRIP</a></li>\n<li><a href=\"https://www.cs1.tf.fau.de/research/multimedia-security/code/image-manipulation-dataset/#collapse_1\" target=\"_blank\">FAU</a></li>\n</ul>\n<p>Validation results on the supplemental images were:</p>\n<table>\n<thead>\n<tr>\n<th>Training data</th>\n<th>Validation score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Competition training set only</td>\n<td>0.092</td>\n</tr>\n<tr>\n<td>Competition training set + CASIA</td>\n<td>0.185</td>\n</tr>\n<tr>\n<td>Competition training set + GRIP</td>\n<td>0.095</td>\n</tr>\n<tr>\n<td>Competition training set + FAU</td>\n<td>0.129</td>\n</tr>\n<tr>\n<td>Competition training set + CASIA + GRIP + FAU</td>\n<td>0.240</td>\n</tr>\n</tbody>\n</table>\n<p>Although these external datasets are not biomedical, their diversity helped the model learn more general copy-move detection cues.</p>\n<h2>Better Backbones</h2>\n<p>Because copy-move detection is a difficult task, stronger pretrained backbones made a large difference.</p>\n<p>Upgrading the backbone from <strong>DINOv2 Giant</strong> to <strong>DINOv3 Huge+</strong> improved the validation score from the earlier range to around <strong>0.35</strong>. Since these two models are relatively close in size, I believe the improvement mainly came from the stronger pretraining of DINOv3.</p>\n<p>I then used the original <strong>DINOv3 7B</strong> model, which further improved the validation score to <strong>0.51</strong>. With additional hyperparameter tuning, the best validation score exceeded <strong>0.56</strong>.</p>\n<p>Overall, the final solution remained simple: a strong pretrained ViT backbone, a minimal segmentation head, diverse training data, and lightweight post-processing.</p>",
      "rawMarkdown": "I would like to thank Kaggle and Recod.ai/LUC for hosting this interesting competition. I finished 5th on the private leaderboard.\n\n**Final scores:**\n\n* Public LB: **0.422**\n* Private LB: **0.376**\n* [**inference code**](https://www.kaggle.com/code/wowfattie/recodfinal)\n* **training code attached at the end**\n\n## Overview\n\nMy solution focused on simplicity and robustness. I treated the task primarily as a semantic segmentation problem.\n\nFor training mask generation, when an image contained multiple copy-move instances, I used the union of all manipulated regions as the training target. In other words, the model was trained as a binary manipulated-region segmenter rather than as an instance segmentation model.\n\nAt inference time, I also did not separate multiple copy-move instances. This was a deliberate simplification.\n\nTwo factors contributed most significantly to the final score:\n\n1. Using external datasets to improve generalization.\n2. Using the strongest available pretrained DINOv3 backbone.\n\n## Validation\n\nI mainly validated my models on the 48 supplemental images provided by the host, because these multi-panel images were the closest match to the actual test set.\n\n## Backbone and Segmentation Head\n\nI used pretrained DINOv2/DINOv3 backbones with a very simple segmentation head.\n\nThe model takes the final patch tokens from the ViT backbone, reshapes them into a 2D feature grid, and applies a single 1×1 Conv2d head to produce one logit map.\n\nDuring training, I downsampled the ground-truth mask to match the patch-grid resolution:\n\n* 14×14 for DINOv2\n* 16×16 for DINOv3\n\nDuring inference, after predicting the low-resolution mask, I simply resized it back to the original image size using bilinear interpolation.\n\n## Augmentations\n\nThe most useful augmentations were simple image-level augmentations:\n\n```python\nA.RandomRotate90(p=1.0)\nA.HorizontalFlip(p=0.5)\nA.GaussianBlur(p=0.15)\nA.ToGray(...)\nA.HueSaturationValue(...)\nA.RandomBrightnessContrast(p=0.9)\nA.ImageCompression(p=0.25)\n```\n\nI used fairly aggressive color augmentation because the manipulation signal is not always color-specific. The model needed to learn structural and local consistency cues rather than memorize color distributions.\n\n## Important Training Choices\n\n* BF16 autocast\n* Full fine-tuning\n* FlashAttention-2\n* Input image size: 1024 × 1024\n* Vanilla BCE loss\n\n## Important Inference Choices\n\n* 8-bit inference to fit within the 16GB memory limit of a T4 GPU\n* SDPA attention\n* Input image size: 1024 × 1024\n* Four-way `rotate90` test-time augmentation\n\n## Post-processing\n\nThe post-processing pipeline was simple:\n\n1. Create a high-confidence binary mask using a fixed threshold.\n2. Compute the area of this high-confidence mask by summing its active pixels.\n3. If the high-confidence area is smaller than `min_area`, classify the sample as authentic and stop.\n4. Create the final binary mask using the configured threshold `mask_thr`.\n5. If the final mask contains no active pixels, classify the sample as authentic and stop.\n\n## Early Results\n\nI started with a DINOv2 model. Its validation score on the supplemental images was only **0.092**, although the validation score on a split of the provided training set was much higher.\n\nAfter inspecting the training data, I found that although the dataset was not small, it contained only a limited variety of image types. As a result, the fine-tuned model did not generalize well to the supplemental set, which was visually quite different from the training set.\n\n## External Datasets\n\nTo improve generalization, I searched for external datasets that could increase the diversity of the training data. I eventually added the following datasets:\n\n* [CASIA](https://www.kaggle.com/datasets/divg07/casia-20-image-tampering-detection-dataset)\n* [GRIP](https://www.grip.unina.it/download/prog/CMFD/)\n* [FAU](https://www.cs1.tf.fau.de/research/multimedia-security/code/image-manipulation-dataset/#collapse_1)\n\nValidation results on the supplemental images were:\n\n| Training data                                 | Validation score |\n| --------------------------------------------- | ---------------: |\n| Competition training set only                 |            0.092 |\n| Competition training set + CASIA              |            0.185 |\n| Competition training set + GRIP               |            0.095 |\n| Competition training set + FAU                |            0.129 |\n| Competition training set + CASIA + GRIP + FAU |            0.240 |\n\nAlthough these external datasets are not biomedical, their diversity helped the model learn more general copy-move detection cues.\n\n## Better Backbones\n\nBecause copy-move detection is a difficult task, stronger pretrained backbones made a large difference.\n\nUpgrading the backbone from **DINOv2 Giant** to **DINOv3 Huge+** improved the validation score from the earlier range to around **0.35**. Since these two models are relatively close in size, I believe the improvement mainly came from the stronger pretraining of DINOv3.\n\nI then used the original **DINOv3 7B** model, which further improved the validation score to **0.51**. With additional hyperparameter tuning, the best validation score exceeded **0.56**.\n\nOverall, the final solution remained simple: a strong pretrained ViT backbone, a minimal segmentation head, diverse training data, and lightweight post-processing.\n",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3454335": "I would like to thank Kaggle and Recod.ai/LUC for hosting this interesting competition. I finished 5th on the private leaderboard.\n\n**Final scores:**\n\n* Public LB: **0.422**\n* Private LB: **0.376**\n* [**inference code**](https://www.kaggle.com/code/wowfattie/recodfinal)\n* **training code attached at the end**\n\n## Overview\n\nMy solution focused on simplicity and robustness. I treated the task primarily as a semantic segmentation problem.\n\nFor training mask generation, when an image contained multiple copy-move instances, I used the union of all manipulated regions as the training target. In other words, the model was trained as a binary manipulated-region segmenter rather than as an instance segmentation model.\n\nAt inference time, I also did not separate multiple copy-move instances. This was a deliberate simplification.\n\nTwo factors contributed most significantly to the final score:\n\n1. Using external datasets to improve generalization.\n2. Using the strongest available pretrained DINOv3 backbone.\n\n## Validation\n\nI mainly validated my models on the 48 supplemental images provided by the host, because these multi-panel images were the closest match to the actual test set.\n\n## Backbone and Segmentation Head\n\nI used pretrained DINOv2/DINOv3 backbones with a very simple segmentation head.\n\nThe model takes the final patch tokens from the ViT backbone, reshapes them into a 2D feature grid, and applies a single 1×1 Conv2d head to produce one logit map.\n\nDuring training, I downsampled the ground-truth mask to match the patch-grid resolution:\n\n* 14×14 for DINOv2\n* 16×16 for DINOv3\n\nDuring inference, after predicting the low-resolution mask, I simply resized it back to the original image size using bilinear interpolation.\n\n## Augmentations\n\nThe most useful augmentations were simple image-level augmentations:\n\n```python\nA.RandomRotate90(p=1.0)\nA.HorizontalFlip(p=0.5)\nA.GaussianBlur(p=0.15)\nA.ToGray(...)\nA.HueSaturationValue(...)\nA.RandomBrightnessContrast(p=0.9)\nA.ImageCompression(p=0.25)\n```\n\nI used fairly aggressive color augmentation because the manipulation signal is not always color-specific. The model needed to learn structural and local consistency cues rather than memorize color distributions.\n\n## Important Training Choices\n\n* BF16 autocast\n* Full fine-tuning\n* FlashAttention-2\n* Input image size: 1024 × 1024\n* Vanilla BCE loss\n\n## Important Inference Choices\n\n* 8-bit inference to fit within the 16GB memory limit of a T4 GPU\n* SDPA attention\n* Input image size: 1024 × 1024\n* Four-way `rotate90` test-time augmentation\n\n## Post-processing\n\nThe post-processing pipeline was simple:\n\n1. Create a high-confidence binary mask using a fixed threshold.\n2. Compute the area of this high-confidence mask by summing its active pixels.\n3. If the high-confidence area is smaller than `min_area`, classify the sample as authentic and stop.\n4. Create the final binary mask using the configured threshold `mask_thr`.\n5. If the final mask contains no active pixels, classify the sample as authentic and stop.\n\n## Early Results\n\nI started with a DINOv2 model. Its validation score on the supplemental images was only **0.092**, although the validation score on a split of the provided training set was much higher.\n\nAfter inspecting the training data, I found that although the dataset was not small, it contained only a limited variety of image types. As a result, the fine-tuned model did not generalize well to the supplemental set, which was visually quite different from the training set.\n\n## External Datasets\n\nTo improve generalization, I searched for external datasets that could increase the diversity of the training data. I eventually added the following datasets:\n\n* [CASIA](https://www.kaggle.com/datasets/divg07/casia-20-image-tampering-detection-dataset)\n* [GRIP](https://www.grip.unina.it/download/prog/CMFD/)\n* [FAU](https://www.cs1.tf.fau.de/research/multimedia-security/code/image-manipulation-dataset/#collapse_1)\n\nValidation results on the supplemental images were:\n\n| Training data                                 | Validation score |\n| --------------------------------------------- | ---------------: |\n| Competition training set only                 |            0.092 |\n| Competition training set + CASIA              |            0.185 |\n| Competition training set + GRIP               |            0.095 |\n| Competition training set + FAU                |            0.129 |\n| Competition training set + CASIA + GRIP + FAU |            0.240 |\n\nAlthough these external datasets are not biomedical, their diversity helped the model learn more general copy-move detection cues.\n\n## Better Backbones\n\nBecause copy-move detection is a difficult task, stronger pretrained backbones made a large difference.\n\nUpgrading the backbone from **DINOv2 Giant** to **DINOv3 Huge+** improved the validation score from the earlier range to around **0.35**. Since these two models are relatively close in size, I believe the improvement mainly came from the stronger pretraining of DINOv3.\n\nI then used the original **DINOv3 7B** model, which further improved the validation score to **0.51**. With additional hyperparameter tuning, the best validation score exceeded **0.56**.\n\nOverall, the final solution remained simple: a strong pretrained ViT backbone, a minimal segmentation head, diverse training data, and lightweight post-processing.\n"
  }
}