{
  "id": 694168,
  "title": "RECOD.AI - LUC Scientific Image Forgery Detection - 29th Place Solution",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/694168",
  "author_name": "Gunes Evitan",
  "post_date": "2026-04-23T17:02:40.716000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>RECOD.AI - LUC Scientific Image Forgery Detection - 29th Place Solution</h1>\n<h2>Dataset</h2>\n<p>All samples from the competition dataset are utilized: the training data containing authentic and forged scientific figures with pixel-level instance masks, and the supplemental data of additional forged samples.</p>\n<p>The <a href=\"https://zenodo.org/records/15095089\" target=\"_blank\">RSIID (Recod.ai Scientific Image Integrity Dataset)</a> is used as external training data. It is a benchmark dataset of 39,423 synthetically tampered scientific figures derived from 2,923 pristine images.</p>\n<p>From RSIID, only the copy-move duplication subset is used, at both simple and compound levels.</p>\n<p><strong>Simple copy-move</strong> cases come with a paired pristine image (<code>figure_pristine.png</code>) that is the original figure before the forgery was applied. The ground-truth mask is computed as the pixel-wise union of the forgery mask and the pristine mask, covering both the source and destination copy regions. Only the forged image is added to training; the dataset also includes standalone authentic figures from a separate pristine split.</p>\n<pre><code># union of forgery and pristine masks covers both source and destination regions\nmask = np.maximum(forgery_mask, pristine_mask)\nmask_objects = np.unique(mask)[1:]  # skip background\nmask = np.stack([(mask == i).astype(np.uint8) for i in mask_objects], axis=0)\n</code></pre>\n<p><strong>Compound copy-move</strong> cases require counterfactual reconstruction to generate a paired authentic sample, and come in two sub-types:</p>\n<ul>\n<li><strong>Intra-panel</strong>: Each case directory contains a clean version of the forged panel (<code>panel_pristine.png</code>). A pristine image is reconstructed by locating the forged region bounding box from <code>annotations.json</code>, resizing <code>panel_pristine.png</code> to match it, and compositing it back into the forged figure:</li>\n</ul>\n<pre><code>pristine_panel = cv2.resize(pristine_panel, (w, h), interpolation=cv2.INTER_LINEAR)\npristine_image = forgery_image.copy()\npristine_image[y0:y1, x0:x1] = pristine_panel\n</code></pre>\n<ul>\n<li><strong>Inter-panel</strong>: No pre-existing clean panel is available. A replacement image is sampled from the RSIID source image pool (<code>datasetSrc.csv</code>), matched to the same <code>subset_tag</code> as the forged panel and excluding images already present in the figure, then resized and composited into the forgery region:</li>\n</ul>\n<pre><code># sample a same-category image that was not already used in the figure\ncandidate_ids = df_rsiid_src.loc[\n    (df_rsiid_src['subset_tag'] == forgery_image_subset_tag) &amp;\n    ~df_rsiid_src['dataPath'].isin(used_image_ids),\n    'dataPath'\n]\nreplace_image = cv2.imread(rsiid_src_directory / np.random.choice(candidate_ids))\nreplace_image = cv2.resize(replace_image, (w, h), interpolation=cv2.INTER_LINEAR)\npristine_image = forgery_image.copy()\npristine_image[y0:y1, x0:x1] = replace_image\n</code></pre>\n<p>In both compound sub-types every forged image produces a paired authentic counterpart with the same image ID. All sources are merged into a single dataset:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Authentic</th>\n<th>Forged</th>\n<th>Total</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Competition train</td>\n<td>2,377</td>\n<td>2,751</td>\n<td>5,128</td>\n</tr>\n<tr>\n<td>Competition supplemental</td>\n<td>—</td>\n<td>48</td>\n<td>48</td>\n</tr>\n<tr>\n<td>RSIID simple copy-move</td>\n<td>2,923</td>\n<td>5,390</td>\n<td>8,313</td>\n</tr>\n<tr>\n<td>RSIID compound copy-move</td>\n<td>19,000</td>\n<td>19,000</td>\n<td>38,000</td>\n</tr>\n<tr>\n<td><strong>Total</strong></td>\n<td><strong>24,300</strong></td>\n<td><strong>27,189</strong></td>\n<td><strong>51,489</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Validation</h2>\n<p>The dataset is split into 5 folds using <code>StratifiedGroupKFold</code>, stratified on the joint combination of class and dataset source so that each fold preserves the original distribution across all dataset-class groups. Samples are grouped by <code>image_id</code>, which ensures that forged/authentic counterfactual pairs from the compound copy-move reconstruction always land in the same fold and never leak across the train/validation boundary.</p>\n<h2>Model</h2>\n<p>The segmentation model is <code>DinoV2Segmenter</code>, a <code>facebook/dinov2-base</code> Vision Transformer encoder paired with a lightweight convolutional decoder. The encoder is fully frozen except for the last 5 transformer blocks and the final LayerNorm; patch embeddings remain frozen throughout.</p>\n<p>At inference, the encoder's output token sequence is stripped of the CLS token, reshaped into a 2D spatial feature map, and passed to <code>DinoDecoder</code>, which upsamples it back to the input resolution:</p>\n<pre><code>def forward(self, x: torch.Tensor) -&gt; torch.Tensor:\n    H, W = x.shape[-2:]\n    out = self.encoder(pixel_values=x).last_hidden_state  # (B, N+1, 768)\n    feature_map = self._tokens_to_map(out)                # (B, 768, H/14, W/14)\n    return self.decoder(feature_map, (H, W))              # (B, 1, H, W)\n\ndef _tokens_to_map(self, tokens: torch.Tensor) -&gt; torch.Tensor:\n    B, N, C = tokens.shape\n    num_reg = int(getattr(self.encoder.config, 'num_register_tokens', 0) or 0)\n    patch_tokens = tokens[:, 1 + num_reg:, :]  # drop CLS and register tokens\n    s = int(math.isqrt(patch_tokens.shape[1]))\n    return patch_tokens.permute(0, 2, 1).reshape(B, C, s, s)\n</code></pre>\n<p><code>DinoDecoder</code> is a three-block convolutional head that halves the channel width at each stage and inserts a 2× bilinear upsample between the first two blocks. A final interpolation brings the output to exactly the input image size:</p>\n<pre><code>class DinoDecoder(nn.Module):\n    def __init__(self, in_ch, mid_ch=256, out_ch=1, dropout=0.1):\n        self.block1 = nn.Sequential(Conv2d(in_ch,       mid_ch,     3, padding=1), ReLU(), Dropout2d(dropout))\n        self.block2 = nn.Sequential(Conv2d(mid_ch,      mid_ch//2,  3, padding=1), ReLU(), Dropout2d(dropout))\n        self.block3 = nn.Sequential(Conv2d(mid_ch//2,   mid_ch//4,  3, padding=1), ReLU())\n        self.out_conv = Conv2d(mid_ch//4, out_ch, kernel_size=1)\n\n    def forward(self, feature_map, out_size):\n        x = self.block1(feature_map)\n        x = F.interpolate(x, scale_factor=2.0, mode='bilinear', align_corners=False)\n        x = self.block2(x)\n        x = F.interpolate(x, scale_factor=2.0, mode='bilinear', align_corners=False)\n        x = self.block3(x)\n        x = self.out_conv(x)\n        return F.interpolate(x, size=out_size, mode='bilinear', align_corners=False)\n</code></pre>\n<h2>Training</h2>\n<p>AdamW is used with differential learning rates: the decoder is trained at <code>lr = 1e-3</code> while the unfrozen encoder blocks receive a 10× smaller rate <code>lr_encoder = 1e-4</code>, preventing the pretrained representations from being overwritten too quickly. Both parameter groups share <code>weight_decay = 1e-3</code>.</p>\n<p><code>BCEWithLogitsLoss</code> is used for binary pixel-level supervision. Training runs with Automatic Mixed Precision (<code>bfloat16</code>) and gradient norm clipping at 1.0. Models are trained for 50 epochs with a batch size of 64, and checkpoints are saved whenever validation loss or mean image score improves so both type of \"best\" model can be explored.</p>\n<p>Training images are resized to 512×512 and augmented with transpose, vertical flip, and horizontal flip (each at p=0.5). Light color jitter is applied with low probability: random brightness/contrast (±0.15, p=0.1), hue/saturation/value shift (±5, p=0.1), and JPEG compression (quality 95–100, p=0.1) to simulate mild compression artifacts present in real scientific figures.</p>\n<table>\n<thead>\n<tr>\n<th>Transform</th>\n<th>Parameters</th>\n<th>p</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Transpose</td>\n<td>—</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>Vertical flip</td>\n<td>—</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>Horizontal flip</td>\n<td>—</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>RandomBrightnessContrast</td>\n<td>limit ±0.15</td>\n<td>0.1</td>\n</tr>\n<tr>\n<td>HueSaturationValue</td>\n<td>shift ±5</td>\n<td>0.1</td>\n</tr>\n<tr>\n<td>ImageCompression</td>\n<td>JPEG, quality 95–100</td>\n<td>0.1</td>\n</tr>\n</tbody>\n</table>\n<h2>Inference</h2>\n<p>At test-time, 8 augmented versions of each image are averaged: 4 rotations (0°, 90°, 180°, 270°) × 2 horizontal flip states. Each augmented prediction is inverted back to the original orientation before averaging.</p>\n<h2>Post-processing</h2>\n<p>Each image is scored independently. Authentic images receive 1.0 for a correct <code>\"authentic\"</code> prediction and 0.0 for any mask. Forged images are scored by the optimal instance-level F1 computed via the Hungarian algorithm matching predicted masks to ground-truth masks. Predicting more mask instances than the ground truth carries a penalty:</p>\n<pre><code>excess_predictions_penalty = len(gt_masks) / max(len(pred_masks), len(gt_masks))\nscore = np.mean(f1_score_matrix[row_ind, col_ind]) * excess_predictions_penalty\n</code></pre>\n<p>The naive baseline of predicting everything as <code>\"authentic\"</code> achieves a score equal to the fraction of authentic images in the dataset. Any useful submission must exceed this.</p>\n<p>These two properties; the cost of a false positive and the penalty for excess mask instances, directly shape every post-processing decision: the forgery/authentic threshold must be conservative enough to avoid tagging authentic images as forged, and when a forgery is detected the entire probability map is merged into a single mask rather than split into separate instances to avoid the excess predictions penalty.</p>\n<p>OOF probability maps are resized back to the original image resolution with cubic interpolation. A two-step rule decides the final prediction for each image: if the 99.9th percentile (<code>q999</code>) of the probability map is at least 0.95, the image is predicted as forged and the mask is thresholded at 0.3; otherwise the image is predicted as <code>\"authentic\"</code>.</p>\n<pre><code>if np.quantile(probability_map, 0.999) &gt;= 0.95:\n    probability_map = cv2.resize(probability_map, (orig_w, orig_h), interpolation=cv2.INTER_CUBIC)\n    mask = (probability_map &gt; THRESHOLD).astype(np.uint8)\n    prediction = rle_encode([mask])\nelse:\n    prediction = 'authentic'\n</code></pre>\n<p>The intuition is that a copy-move forgery creates a localized region the model responds to with near-certain confidence. Even if that region is small, it drives the extreme tail of the probability distribution sharply upward. On authentic images the model stays uncertain across the board, so the tail never breaks above the threshold.</p>\n<p>The <code>q999</code> gate is intentionally conservative: genuinely forged images produce a tight cluster of very high-confidence pixels that pushes the extreme tail well above 0.95, while authentic images stay flat near zero throughout. The thresholded map is always submitted as a single mask instance. Predicting multiple instances when the ground truth has fewer triggers the excess predictions penalty and scales the score down by <code>len(gt_masks) / len(pred_masks)</code>, so submitting one merged mask is safer than attempting to separate individual regions.</p>\n<h2>Submission Selection</h2>\n<p>This function computes the same set of metrics globally and separately for each dataset split. For each split it reports: the mean image score, the baseline score (fraction of authentic images, what predicting everything as <code>\"authentic\"</code> would achieve), lift over that baseline, TP/FP/TN/FN counts, precision, recall, accuracy, and the two impact terms <code>tp_impact</code> and <code>fp_impact</code>.</p>\n<pre><code>def get_detection_stats(df: pd.DataFrame) -&gt; dict[str, int | float]:\n\n    \"\"\"\n    Calculates detection statistics globally and per dataset.\n\n    Parameters\n    ----------\n    df: pd.DataFrame\n        DataFrame with 'target', 'prediction', 'score', and optionally 'dataset'.\n\n    Returns\n    -------\n    dict[str, int | float]\n        A dictionary containing global stats and per-dataset stats.\n    \"\"\"\n\n    def _compute_metrics(sub_df: pd.DataFrame) -&gt; dict[str, int | float]:\n\n        is_forged_gt = sub_df['target'] != 'authentic'\n        has_mask_pred = sub_df['prediction'] != 'authentic'\n\n        tp_mask = (is_forged_gt) &amp; (has_mask_pred)\n        fp_mask = (~is_forged_gt) &amp; (has_mask_pred)\n        tn_mask = (~is_forged_gt) &amp; (~has_mask_pred)\n        fn_mask = (is_forged_gt) &amp; (~has_mask_pred)\n\n        tp = tp_mask.sum()\n        fp = fp_mask.sum()\n        tn = tn_mask.sum()\n        fn = fn_mask.sum()\n\n        total = len(sub_df)\n\n        total_tp_gain = sub_df.loc[tp_mask, 'score'].sum()\n        total_fp_loss = float(fp) * 1.0\n        fp_impact = total_fp_loss / total\n        tp_impact = total_tp_gain / total\n\n        current_mean_score = float(sub_df['score'].mean()) if not sub_df.empty else 0.0\n        count_authentic_gt = int((~is_forged_gt).sum())\n        baseline_score = float(count_authentic_gt / total)\n        lift = current_mean_score - baseline_score\n\n        return {\n            'total_samples': int(total),\n            'accuracy': float((tp + tn) / total) if total &gt; 0 else 0.0,\n            'precision': float(tp / (tp + fp)) if (tp + fp) &gt; 0 else 0.0,\n            'recall': float(tp / (tp + fn)) if (tp + fn) &gt; 0 else 0.0,\n            'count_forged_correct (TP)': int(tp),\n            'count_authentic_correct (TN)': int(tn),\n            'count_false_positives (FP)': int(fp),\n            'count_missed_forgeries (FN)': int(fn),\n            'tp_impact': float(tp_impact),\n            'fp_impact': float(fp_impact),\n            'mean_image_score': current_mean_score,\n            'baseline_score': baseline_score,\n            'lift_over_baseline': lift,\n        }\n\n    stats = _compute_metrics(df)\n    if 'dataset' in df.columns:\n        for ds_name, ds_df in df.groupby('dataset'):\n            ds_stats = _compute_metrics(ds_df)\n            for key, value in ds_stats.items():\n                stats[f'{ds_name}_{key}'] = value\n\n    return stats\n</code></pre>\n<p>Each dataset split has a different authentic/forged ratio, so the baseline score differs across splits and a threshold that improves the global score can still hurt an individual split if it generates too many false positives there. Tracking lift per split catches this: a submission is only considered safe if every split shows a positive lift, not just the aggregate.</p>\n<p>Within each split, <code>tp_impact</code> and <code>fp_impact</code> make the cost/benefit of the threshold explicit. <code>tp_impact</code> measures how much the correctly detected forgeries contribute to the score; <code>fp_impact</code> measures how much is being lost to false positives since each FP on an authentic image converts a guaranteed 1.0 into 0.0. The selected threshold is the most conservative one where every dataset split still shows a positive lift, meaning the TP gain outweighs the FP loss in each split independently rather than the threshold that maximises raw TP recall.</p>\n<p>This framework made submission selection straightforward: rather than guessing at a threshold and hoping it generalised, the per-split breakdown showed exactly where false positives were appearing and whether the lift was real or an artefact of one dominant split. The final submission was chosen as the most conservative configuration that demonstrated consistent positive lift across all splits with minimal fp_impact. The exact CV and public leaderboard scores were not recorded at the time and are no longer available.</p>",
  "messages": [
    {
      "id": 3447664,
      "postDate": "2026-04-23T17:02:40.717Z",
      "content": "<h1>RECOD.AI - LUC Scientific Image Forgery Detection - 29th Place Solution</h1>\n<h2>Dataset</h2>\n<p>All samples from the competition dataset are utilized: the training data containing authentic and forged scientific figures with pixel-level instance masks, and the supplemental data of additional forged samples.</p>\n<p>The <a href=\"https://zenodo.org/records/15095089\" target=\"_blank\">RSIID (Recod.ai Scientific Image Integrity Dataset)</a> is used as external training data. It is a benchmark dataset of 39,423 synthetically tampered scientific figures derived from 2,923 pristine images.</p>\n<p>From RSIID, only the copy-move duplication subset is used, at both simple and compound levels.</p>\n<p><strong>Simple copy-move</strong> cases come with a paired pristine image (<code>figure_pristine.png</code>) that is the original figure before the forgery was applied. The ground-truth mask is computed as the pixel-wise union of the forgery mask and the pristine mask, covering both the source and destination copy regions. Only the forged image is added to training; the dataset also includes standalone authentic figures from a separate pristine split.</p>\n<pre><code># union of forgery and pristine masks covers both source and destination regions\nmask = np.maximum(forgery_mask, pristine_mask)\nmask_objects = np.unique(mask)[1:]  # skip background\nmask = np.stack([(mask == i).astype(np.uint8) for i in mask_objects], axis=0)\n</code></pre>\n<p><strong>Compound copy-move</strong> cases require counterfactual reconstruction to generate a paired authentic sample, and come in two sub-types:</p>\n<ul>\n<li><strong>Intra-panel</strong>: Each case directory contains a clean version of the forged panel (<code>panel_pristine.png</code>). A pristine image is reconstructed by locating the forged region bounding box from <code>annotations.json</code>, resizing <code>panel_pristine.png</code> to match it, and compositing it back into the forged figure:</li>\n</ul>\n<pre><code>pristine_panel = cv2.resize(pristine_panel, (w, h), interpolation=cv2.INTER_LINEAR)\npristine_image = forgery_image.copy()\npristine_image[y0:y1, x0:x1] = pristine_panel\n</code></pre>\n<ul>\n<li><strong>Inter-panel</strong>: No pre-existing clean panel is available. A replacement image is sampled from the RSIID source image pool (<code>datasetSrc.csv</code>), matched to the same <code>subset_tag</code> as the forged panel and excluding images already present in the figure, then resized and composited into the forgery region:</li>\n</ul>\n<pre><code># sample a same-category image that was not already used in the figure\ncandidate_ids = df_rsiid_src.loc[\n    (df_rsiid_src['subset_tag'] == forgery_image_subset_tag) &amp;\n    ~df_rsiid_src['dataPath'].isin(used_image_ids),\n    'dataPath'\n]\nreplace_image = cv2.imread(rsiid_src_directory / np.random.choice(candidate_ids))\nreplace_image = cv2.resize(replace_image, (w, h), interpolation=cv2.INTER_LINEAR)\npristine_image = forgery_image.copy()\npristine_image[y0:y1, x0:x1] = replace_image\n</code></pre>\n<p>In both compound sub-types every forged image produces a paired authentic counterpart with the same image ID. All sources are merged into a single dataset:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Authentic</th>\n<th>Forged</th>\n<th>Total</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Competition train</td>\n<td>2,377</td>\n<td>2,751</td>\n<td>5,128</td>\n</tr>\n<tr>\n<td>Competition supplemental</td>\n<td>—</td>\n<td>48</td>\n<td>48</td>\n</tr>\n<tr>\n<td>RSIID simple copy-move</td>\n<td>2,923</td>\n<td>5,390</td>\n<td>8,313</td>\n</tr>\n<tr>\n<td>RSIID compound copy-move</td>\n<td>19,000</td>\n<td>19,000</td>\n<td>38,000</td>\n</tr>\n<tr>\n<td><strong>Total</strong></td>\n<td><strong>24,300</strong></td>\n<td><strong>27,189</strong></td>\n<td><strong>51,489</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Validation</h2>\n<p>The dataset is split into 5 folds using <code>StratifiedGroupKFold</code>, stratified on the joint combination of class and dataset source so that each fold preserves the original distribution across all dataset-class groups. Samples are grouped by <code>image_id</code>, which ensures that forged/authentic counterfactual pairs from the compound copy-move reconstruction always land in the same fold and never leak across the train/validation boundary.</p>\n<h2>Model</h2>\n<p>The segmentation model is <code>DinoV2Segmenter</code>, a <code>facebook/dinov2-base</code> Vision Transformer encoder paired with a lightweight convolutional decoder. The encoder is fully frozen except for the last 5 transformer blocks and the final LayerNorm; patch embeddings remain frozen throughout.</p>\n<p>At inference, the encoder's output token sequence is stripped of the CLS token, reshaped into a 2D spatial feature map, and passed to <code>DinoDecoder</code>, which upsamples it back to the input resolution:</p>\n<pre><code>def forward(self, x: torch.Tensor) -&gt; torch.Tensor:\n    H, W = x.shape[-2:]\n    out = self.encoder(pixel_values=x).last_hidden_state  # (B, N+1, 768)\n    feature_map = self._tokens_to_map(out)                # (B, 768, H/14, W/14)\n    return self.decoder(feature_map, (H, W))              # (B, 1, H, W)\n\ndef _tokens_to_map(self, tokens: torch.Tensor) -&gt; torch.Tensor:\n    B, N, C = tokens.shape\n    num_reg = int(getattr(self.encoder.config, 'num_register_tokens', 0) or 0)\n    patch_tokens = tokens[:, 1 + num_reg:, :]  # drop CLS and register tokens\n    s = int(math.isqrt(patch_tokens.shape[1]))\n    return patch_tokens.permute(0, 2, 1).reshape(B, C, s, s)\n</code></pre>\n<p><code>DinoDecoder</code> is a three-block convolutional head that halves the channel width at each stage and inserts a 2× bilinear upsample between the first two blocks. A final interpolation brings the output to exactly the input image size:</p>\n<pre><code>class DinoDecoder(nn.Module):\n    def __init__(self, in_ch, mid_ch=256, out_ch=1, dropout=0.1):\n        self.block1 = nn.Sequential(Conv2d(in_ch,       mid_ch,     3, padding=1), ReLU(), Dropout2d(dropout))\n        self.block2 = nn.Sequential(Conv2d(mid_ch,      mid_ch//2,  3, padding=1), ReLU(), Dropout2d(dropout))\n        self.block3 = nn.Sequential(Conv2d(mid_ch//2,   mid_ch//4,  3, padding=1), ReLU())\n        self.out_conv = Conv2d(mid_ch//4, out_ch, kernel_size=1)\n\n    def forward(self, feature_map, out_size):\n        x = self.block1(feature_map)\n        x = F.interpolate(x, scale_factor=2.0, mode='bilinear', align_corners=False)\n        x = self.block2(x)\n        x = F.interpolate(x, scale_factor=2.0, mode='bilinear', align_corners=False)\n        x = self.block3(x)\n        x = self.out_conv(x)\n        return F.interpolate(x, size=out_size, mode='bilinear', align_corners=False)\n</code></pre>\n<h2>Training</h2>\n<p>AdamW is used with differential learning rates: the decoder is trained at <code>lr = 1e-3</code> while the unfrozen encoder blocks receive a 10× smaller rate <code>lr_encoder = 1e-4</code>, preventing the pretrained representations from being overwritten too quickly. Both parameter groups share <code>weight_decay = 1e-3</code>.</p>\n<p><code>BCEWithLogitsLoss</code> is used for binary pixel-level supervision. Training runs with Automatic Mixed Precision (<code>bfloat16</code>) and gradient norm clipping at 1.0. Models are trained for 50 epochs with a batch size of 64, and checkpoints are saved whenever validation loss or mean image score improves so both type of \"best\" model can be explored.</p>\n<p>Training images are resized to 512×512 and augmented with transpose, vertical flip, and horizontal flip (each at p=0.5). Light color jitter is applied with low probability: random brightness/contrast (±0.15, p=0.1), hue/saturation/value shift (±5, p=0.1), and JPEG compression (quality 95–100, p=0.1) to simulate mild compression artifacts present in real scientific figures.</p>\n<table>\n<thead>\n<tr>\n<th>Transform</th>\n<th>Parameters</th>\n<th>p</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Transpose</td>\n<td>—</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>Vertical flip</td>\n<td>—</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>Horizontal flip</td>\n<td>—</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>RandomBrightnessContrast</td>\n<td>limit ±0.15</td>\n<td>0.1</td>\n</tr>\n<tr>\n<td>HueSaturationValue</td>\n<td>shift ±5</td>\n<td>0.1</td>\n</tr>\n<tr>\n<td>ImageCompression</td>\n<td>JPEG, quality 95–100</td>\n<td>0.1</td>\n</tr>\n</tbody>\n</table>\n<h2>Inference</h2>\n<p>At test-time, 8 augmented versions of each image are averaged: 4 rotations (0°, 90°, 180°, 270°) × 2 horizontal flip states. Each augmented prediction is inverted back to the original orientation before averaging.</p>\n<h2>Post-processing</h2>\n<p>Each image is scored independently. Authentic images receive 1.0 for a correct <code>\"authentic\"</code> prediction and 0.0 for any mask. Forged images are scored by the optimal instance-level F1 computed via the Hungarian algorithm matching predicted masks to ground-truth masks. Predicting more mask instances than the ground truth carries a penalty:</p>\n<pre><code>excess_predictions_penalty = len(gt_masks) / max(len(pred_masks), len(gt_masks))\nscore = np.mean(f1_score_matrix[row_ind, col_ind]) * excess_predictions_penalty\n</code></pre>\n<p>The naive baseline of predicting everything as <code>\"authentic\"</code> achieves a score equal to the fraction of authentic images in the dataset. Any useful submission must exceed this.</p>\n<p>These two properties; the cost of a false positive and the penalty for excess mask instances, directly shape every post-processing decision: the forgery/authentic threshold must be conservative enough to avoid tagging authentic images as forged, and when a forgery is detected the entire probability map is merged into a single mask rather than split into separate instances to avoid the excess predictions penalty.</p>\n<p>OOF probability maps are resized back to the original image resolution with cubic interpolation. A two-step rule decides the final prediction for each image: if the 99.9th percentile (<code>q999</code>) of the probability map is at least 0.95, the image is predicted as forged and the mask is thresholded at 0.3; otherwise the image is predicted as <code>\"authentic\"</code>.</p>\n<pre><code>if np.quantile(probability_map, 0.999) &gt;= 0.95:\n    probability_map = cv2.resize(probability_map, (orig_w, orig_h), interpolation=cv2.INTER_CUBIC)\n    mask = (probability_map &gt; THRESHOLD).astype(np.uint8)\n    prediction = rle_encode([mask])\nelse:\n    prediction = 'authentic'\n</code></pre>\n<p>The intuition is that a copy-move forgery creates a localized region the model responds to with near-certain confidence. Even if that region is small, it drives the extreme tail of the probability distribution sharply upward. On authentic images the model stays uncertain across the board, so the tail never breaks above the threshold.</p>\n<p>The <code>q999</code> gate is intentionally conservative: genuinely forged images produce a tight cluster of very high-confidence pixels that pushes the extreme tail well above 0.95, while authentic images stay flat near zero throughout. The thresholded map is always submitted as a single mask instance. Predicting multiple instances when the ground truth has fewer triggers the excess predictions penalty and scales the score down by <code>len(gt_masks) / len(pred_masks)</code>, so submitting one merged mask is safer than attempting to separate individual regions.</p>\n<h2>Submission Selection</h2>\n<p>This function computes the same set of metrics globally and separately for each dataset split. For each split it reports: the mean image score, the baseline score (fraction of authentic images, what predicting everything as <code>\"authentic\"</code> would achieve), lift over that baseline, TP/FP/TN/FN counts, precision, recall, accuracy, and the two impact terms <code>tp_impact</code> and <code>fp_impact</code>.</p>\n<pre><code>def get_detection_stats(df: pd.DataFrame) -&gt; dict[str, int | float]:\n\n    \"\"\"\n    Calculates detection statistics globally and per dataset.\n\n    Parameters\n    ----------\n    df: pd.DataFrame\n        DataFrame with 'target', 'prediction', 'score', and optionally 'dataset'.\n\n    Returns\n    -------\n    dict[str, int | float]\n        A dictionary containing global stats and per-dataset stats.\n    \"\"\"\n\n    def _compute_metrics(sub_df: pd.DataFrame) -&gt; dict[str, int | float]:\n\n        is_forged_gt = sub_df['target'] != 'authentic'\n        has_mask_pred = sub_df['prediction'] != 'authentic'\n\n        tp_mask = (is_forged_gt) &amp; (has_mask_pred)\n        fp_mask = (~is_forged_gt) &amp; (has_mask_pred)\n        tn_mask = (~is_forged_gt) &amp; (~has_mask_pred)\n        fn_mask = (is_forged_gt) &amp; (~has_mask_pred)\n\n        tp = tp_mask.sum()\n        fp = fp_mask.sum()\n        tn = tn_mask.sum()\n        fn = fn_mask.sum()\n\n        total = len(sub_df)\n\n        total_tp_gain = sub_df.loc[tp_mask, 'score'].sum()\n        total_fp_loss = float(fp) * 1.0\n        fp_impact = total_fp_loss / total\n        tp_impact = total_tp_gain / total\n\n        current_mean_score = float(sub_df['score'].mean()) if not sub_df.empty else 0.0\n        count_authentic_gt = int((~is_forged_gt).sum())\n        baseline_score = float(count_authentic_gt / total)\n        lift = current_mean_score - baseline_score\n\n        return {\n            'total_samples': int(total),\n            'accuracy': float((tp + tn) / total) if total &gt; 0 else 0.0,\n            'precision': float(tp / (tp + fp)) if (tp + fp) &gt; 0 else 0.0,\n            'recall': float(tp / (tp + fn)) if (tp + fn) &gt; 0 else 0.0,\n            'count_forged_correct (TP)': int(tp),\n            'count_authentic_correct (TN)': int(tn),\n            'count_false_positives (FP)': int(fp),\n            'count_missed_forgeries (FN)': int(fn),\n            'tp_impact': float(tp_impact),\n            'fp_impact': float(fp_impact),\n            'mean_image_score': current_mean_score,\n            'baseline_score': baseline_score,\n            'lift_over_baseline': lift,\n        }\n\n    stats = _compute_metrics(df)\n    if 'dataset' in df.columns:\n        for ds_name, ds_df in df.groupby('dataset'):\n            ds_stats = _compute_metrics(ds_df)\n            for key, value in ds_stats.items():\n                stats[f'{ds_name}_{key}'] = value\n\n    return stats\n</code></pre>\n<p>Each dataset split has a different authentic/forged ratio, so the baseline score differs across splits and a threshold that improves the global score can still hurt an individual split if it generates too many false positives there. Tracking lift per split catches this: a submission is only considered safe if every split shows a positive lift, not just the aggregate.</p>\n<p>Within each split, <code>tp_impact</code> and <code>fp_impact</code> make the cost/benefit of the threshold explicit. <code>tp_impact</code> measures how much the correctly detected forgeries contribute to the score; <code>fp_impact</code> measures how much is being lost to false positives since each FP on an authentic image converts a guaranteed 1.0 into 0.0. The selected threshold is the most conservative one where every dataset split still shows a positive lift, meaning the TP gain outweighs the FP loss in each split independently rather than the threshold that maximises raw TP recall.</p>\n<p>This framework made submission selection straightforward: rather than guessing at a threshold and hoping it generalised, the per-split breakdown showed exactly where false positives were appearing and whether the lift was real or an artefact of one dominant split. The final submission was chosen as the most conservative configuration that demonstrated consistent positive lift across all splits with minimal fp_impact. The exact CV and public leaderboard scores were not recorded at the time and are no longer available.</p>",
      "rawMarkdown": "# RECOD.AI - LUC Scientific Image Forgery Detection - 29th Place Solution\n\n## Dataset\n\nAll samples from the competition dataset are utilized: the training data containing authentic and forged scientific figures with pixel-level instance masks, and the supplemental data of additional forged samples.\n\nThe [RSIID (Recod.ai Scientific Image Integrity Dataset)](https://zenodo.org/records/15095089) is used as external training data. It is a benchmark dataset of 39,423 synthetically tampered scientific figures derived from 2,923 pristine images.\n\nFrom RSIID, only the copy-move duplication subset is used, at both simple and compound levels.\n\n**Simple copy-move** cases come with a paired pristine image (`figure_pristine.png`) that is the original figure before the forgery was applied. The ground-truth mask is computed as the pixel-wise union of the forgery mask and the pristine mask, covering both the source and destination copy regions. Only the forged image is added to training; the dataset also includes standalone authentic figures from a separate pristine split.\n\n```python\n# union of forgery and pristine masks covers both source and destination regions\nmask = np.maximum(forgery_mask, pristine_mask)\nmask_objects = np.unique(mask)[1:]  # skip background\nmask = np.stack([(mask == i).astype(np.uint8) for i in mask_objects], axis=0)\n```\n\n**Compound copy-move** cases require counterfactual reconstruction to generate a paired authentic sample, and come in two sub-types:\n\n- **Intra-panel**: Each case directory contains a clean version of the forged panel (`panel_pristine.png`). A pristine image is reconstructed by locating the forged region bounding box from `annotations.json`, resizing `panel_pristine.png` to match it, and compositing it back into the forged figure:\n\n```python\npristine_panel = cv2.resize(pristine_panel, (w, h), interpolation=cv2.INTER_LINEAR)\npristine_image = forgery_image.copy()\npristine_image[y0:y1, x0:x1] = pristine_panel\n```\n\n- **Inter-panel**: No pre-existing clean panel is available. A replacement image is sampled from the RSIID source image pool (`datasetSrc.csv`), matched to the same `subset_tag` as the forged panel and excluding images already present in the figure, then resized and composited into the forgery region:\n\n```python\n# sample a same-category image that was not already used in the figure\ncandidate_ids = df_rsiid_src.loc[\n    (df_rsiid_src['subset_tag'] == forgery_image_subset_tag) &\n    ~df_rsiid_src['dataPath'].isin(used_image_ids),\n    'dataPath'\n]\nreplace_image = cv2.imread(rsiid_src_directory / np.random.choice(candidate_ids))\nreplace_image = cv2.resize(replace_image, (w, h), interpolation=cv2.INTER_LINEAR)\npristine_image = forgery_image.copy()\npristine_image[y0:y1, x0:x1] = replace_image\n```\n\nIn both compound sub-types every forged image produces a paired authentic counterpart with the same image ID. All sources are merged into a single dataset:\n\n| Dataset | Authentic | Forged | Total |\n|---|---|---|---|\n| Competition train | 2,377 | 2,751 | 5,128 |\n| Competition supplemental | — | 48 | 48 |\n| RSIID simple copy-move | 2,923 | 5,390 | 8,313 |\n| RSIID compound copy-move | 19,000 | 19,000 | 38,000 |\n| **Total** | **24,300** | **27,189** | **51,489** |\n\n## Validation\n\nThe dataset is split into 5 folds using `StratifiedGroupKFold`, stratified on the joint combination of class and dataset source so that each fold preserves the original distribution across all dataset-class groups. Samples are grouped by `image_id`, which ensures that forged/authentic counterfactual pairs from the compound copy-move reconstruction always land in the same fold and never leak across the train/validation boundary.\n\n## Model\n\nThe segmentation model is `DinoV2Segmenter`, a `facebook/dinov2-base` Vision Transformer encoder paired with a lightweight convolutional decoder. The encoder is fully frozen except for the last 5 transformer blocks and the final LayerNorm; patch embeddings remain frozen throughout.\n\nAt inference, the encoder's output token sequence is stripped of the CLS token, reshaped into a 2D spatial feature map, and passed to `DinoDecoder`, which upsamples it back to the input resolution:\n\n```python\ndef forward(self, x: torch.Tensor) -> torch.Tensor:\n    H, W = x.shape[-2:]\n    out = self.encoder(pixel_values=x).last_hidden_state  # (B, N+1, 768)\n    feature_map = self._tokens_to_map(out)                # (B, 768, H/14, W/14)\n    return self.decoder(feature_map, (H, W))              # (B, 1, H, W)\n\ndef _tokens_to_map(self, tokens: torch.Tensor) -> torch.Tensor:\n    B, N, C = tokens.shape\n    num_reg = int(getattr(self.encoder.config, 'num_register_tokens', 0) or 0)\n    patch_tokens = tokens[:, 1 + num_reg:, :]  # drop CLS and register tokens\n    s = int(math.isqrt(patch_tokens.shape[1]))\n    return patch_tokens.permute(0, 2, 1).reshape(B, C, s, s)\n```\n\n`DinoDecoder` is a three-block convolutional head that halves the channel width at each stage and inserts a 2× bilinear upsample between the first two blocks. A final interpolation brings the output to exactly the input image size:\n\n```python\nclass DinoDecoder(nn.Module):\n    def __init__(self, in_ch, mid_ch=256, out_ch=1, dropout=0.1):\n        self.block1 = nn.Sequential(Conv2d(in_ch,       mid_ch,     3, padding=1), ReLU(), Dropout2d(dropout))\n        self.block2 = nn.Sequential(Conv2d(mid_ch,      mid_ch//2,  3, padding=1), ReLU(), Dropout2d(dropout))\n        self.block3 = nn.Sequential(Conv2d(mid_ch//2,   mid_ch//4,  3, padding=1), ReLU())\n        self.out_conv = Conv2d(mid_ch//4, out_ch, kernel_size=1)\n\n    def forward(self, feature_map, out_size):\n        x = self.block1(feature_map)\n        x = F.interpolate(x, scale_factor=2.0, mode='bilinear', align_corners=False)\n        x = self.block2(x)\n        x = F.interpolate(x, scale_factor=2.0, mode='bilinear', align_corners=False)\n        x = self.block3(x)\n        x = self.out_conv(x)\n        return F.interpolate(x, size=out_size, mode='bilinear', align_corners=False)\n```\n\n## Training\n\nAdamW is used with differential learning rates: the decoder is trained at `lr = 1e-3` while the unfrozen encoder blocks receive a 10× smaller rate `lr_encoder = 1e-4`, preventing the pretrained representations from being overwritten too quickly. Both parameter groups share `weight_decay = 1e-3`.\n\n`BCEWithLogitsLoss` is used for binary pixel-level supervision. Training runs with Automatic Mixed Precision (`bfloat16`) and gradient norm clipping at 1.0. Models are trained for 50 epochs with a batch size of 64, and checkpoints are saved whenever validation loss or mean image score improves so both type of \"best\" model can be explored.\n\nTraining images are resized to 512×512 and augmented with transpose, vertical flip, and horizontal flip (each at p=0.5). Light color jitter is applied with low probability: random brightness/contrast (±0.15, p=0.1), hue/saturation/value shift (±5, p=0.1), and JPEG compression (quality 95–100, p=0.1) to simulate mild compression artifacts present in real scientific figures.\n\n| Transform | Parameters | p |\n|---|---|---|\n| Transpose | — | 0.5 |\n| Vertical flip | — | 0.5 |\n| Horizontal flip | — | 0.5 |\n| RandomBrightnessContrast | limit ±0.15 | 0.1 |\n| HueSaturationValue | shift ±5 | 0.1 |\n| ImageCompression | JPEG, quality 95–100 | 0.1 |\n\n## Inference\nAt test-time, 8 augmented versions of each image are averaged: 4 rotations (0°, 90°, 180°, 270°) × 2 horizontal flip states. Each augmented prediction is inverted back to the original orientation before averaging.\n\n## Post-processing\n\nEach image is scored independently. Authentic images receive 1.0 for a correct `\"authentic\"` prediction and 0.0 for any mask. Forged images are scored by the optimal instance-level F1 computed via the Hungarian algorithm matching predicted masks to ground-truth masks. Predicting more mask instances than the ground truth carries a penalty:\n\n```python\nexcess_predictions_penalty = len(gt_masks) / max(len(pred_masks), len(gt_masks))\nscore = np.mean(f1_score_matrix[row_ind, col_ind]) * excess_predictions_penalty\n```\n\nThe naive baseline of predicting everything as `\"authentic\"` achieves a score equal to the fraction of authentic images in the dataset. Any useful submission must exceed this.\n\nThese two properties; the cost of a false positive and the penalty for excess mask instances, directly shape every post-processing decision: the forgery/authentic threshold must be conservative enough to avoid tagging authentic images as forged, and when a forgery is detected the entire probability map is merged into a single mask rather than split into separate instances to avoid the excess predictions penalty.\n\nOOF probability maps are resized back to the original image resolution with cubic interpolation. A two-step rule decides the final prediction for each image: if the 99.9th percentile (`q999`) of the probability map is at least 0.95, the image is predicted as forged and the mask is thresholded at 0.3; otherwise the image is predicted as `\"authentic\"`.\n\n```python\nif np.quantile(probability_map, 0.999) >= 0.95:\n    probability_map = cv2.resize(probability_map, (orig_w, orig_h), interpolation=cv2.INTER_CUBIC)\n    mask = (probability_map > THRESHOLD).astype(np.uint8)\n    prediction = rle_encode([mask])\nelse:\n    prediction = 'authentic'\n```\n\nThe intuition is that a copy-move forgery creates a localized region the model responds to with near-certain confidence. Even if that region is small, it drives the extreme tail of the probability distribution sharply upward. On authentic images the model stays uncertain across the board, so the tail never breaks above the threshold.\n\nThe `q999` gate is intentionally conservative: genuinely forged images produce a tight cluster of very high-confidence pixels that pushes the extreme tail well above 0.95, while authentic images stay flat near zero throughout. The thresholded map is always submitted as a single mask instance. Predicting multiple instances when the ground truth has fewer triggers the excess predictions penalty and scales the score down by `len(gt_masks) / len(pred_masks)`, so submitting one merged mask is safer than attempting to separate individual regions.\n\n## Submission Selection\n\nThis function computes the same set of metrics globally and separately for each dataset split. For each split it reports: the mean image score, the baseline score (fraction of authentic images, what predicting everything as `\"authentic\"` would achieve), lift over that baseline, TP/FP/TN/FN counts, precision, recall, accuracy, and the two impact terms `tp_impact` and `fp_impact`.\n\n```\ndef get_detection_stats(df: pd.DataFrame) -> dict[str, int | float]:\n\n    \"\"\"\n    Calculates detection statistics globally and per dataset.\n\n    Parameters\n    ----------\n    df: pd.DataFrame\n        DataFrame with 'target', 'prediction', 'score', and optionally 'dataset'.\n\n    Returns\n    -------\n    dict[str, int | float]\n        A dictionary containing global stats and per-dataset stats.\n    \"\"\"\n\n    def _compute_metrics(sub_df: pd.DataFrame) -> dict[str, int | float]:\n\n        is_forged_gt = sub_df['target'] != 'authentic'\n        has_mask_pred = sub_df['prediction'] != 'authentic'\n\n        tp_mask = (is_forged_gt) & (has_mask_pred)\n        fp_mask = (~is_forged_gt) & (has_mask_pred)\n        tn_mask = (~is_forged_gt) & (~has_mask_pred)\n        fn_mask = (is_forged_gt) & (~has_mask_pred)\n\n        tp = tp_mask.sum()\n        fp = fp_mask.sum()\n        tn = tn_mask.sum()\n        fn = fn_mask.sum()\n        \n        total = len(sub_df)\n\n        total_tp_gain = sub_df.loc[tp_mask, 'score'].sum()\n        total_fp_loss = float(fp) * 1.0\n        fp_impact = total_fp_loss / total\n        tp_impact = total_tp_gain / total\n\n        current_mean_score = float(sub_df['score'].mean()) if not sub_df.empty else 0.0\n        count_authentic_gt = int((~is_forged_gt).sum())\n        baseline_score = float(count_authentic_gt / total)\n        lift = current_mean_score - baseline_score\n\n        return {\n            'total_samples': int(total),\n            'accuracy': float((tp + tn) / total) if total > 0 else 0.0,\n            'precision': float(tp / (tp + fp)) if (tp + fp) > 0 else 0.0,\n            'recall': float(tp / (tp + fn)) if (tp + fn) > 0 else 0.0,\n            'count_forged_correct (TP)': int(tp),\n            'count_authentic_correct (TN)': int(tn),\n            'count_false_positives (FP)': int(fp),\n            'count_missed_forgeries (FN)': int(fn),\n            'tp_impact': float(tp_impact),\n            'fp_impact': float(fp_impact),\n            'mean_image_score': current_mean_score,\n            'baseline_score': baseline_score,\n            'lift_over_baseline': lift,\n        }\n\n    stats = _compute_metrics(df)\n    if 'dataset' in df.columns:\n        for ds_name, ds_df in df.groupby('dataset'):\n            ds_stats = _compute_metrics(ds_df)\n            for key, value in ds_stats.items():\n                stats[f'{ds_name}_{key}'] = value\n    \n    return stats\n```\n\nEach dataset split has a different authentic/forged ratio, so the baseline score differs across splits and a threshold that improves the global score can still hurt an individual split if it generates too many false positives there. Tracking lift per split catches this: a submission is only considered safe if every split shows a positive lift, not just the aggregate.\n\nWithin each split, `tp_impact` and `fp_impact` make the cost/benefit of the threshold explicit. `tp_impact` measures how much the correctly detected forgeries contribute to the score; `fp_impact` measures how much is being lost to false positives since each FP on an authentic image converts a guaranteed 1.0 into 0.0. The selected threshold is the most conservative one where every dataset split still shows a positive lift, meaning the TP gain outweighs the FP loss in each split independently rather than the threshold that maximises raw TP recall.\n\nThis framework made submission selection straightforward: rather than guessing at a threshold and hoping it generalised, the per-split breakdown showed exactly where false positives were appearing and whether the lift was real or an artefact of one dominant split. The final submission was chosen as the most conservative configuration that demonstrated consistent positive lift across all splits with minimal fp_impact. The exact CV and public leaderboard scores were not recorded at the time and are no longer available.\n\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3447664": "# RECOD.AI - LUC Scientific Image Forgery Detection - 29th Place Solution\n\n## Dataset\n\nAll samples from the competition dataset are utilized: the training data containing authentic and forged scientific figures with pixel-level instance masks, and the supplemental data of additional forged samples.\n\nThe [RSIID (Recod.ai Scientific Image Integrity Dataset)](https://zenodo.org/records/15095089) is used as external training data. It is a benchmark dataset of 39,423 synthetically tampered scientific figures derived from 2,923 pristine images.\n\nFrom RSIID, only the copy-move duplication subset is used, at both simple and compound levels.\n\n**Simple copy-move** cases come with a paired pristine image (`figure_pristine.png`) that is the original figure before the forgery was applied. The ground-truth mask is computed as the pixel-wise union of the forgery mask and the pristine mask, covering both the source and destination copy regions. Only the forged image is added to training; the dataset also includes standalone authentic figures from a separate pristine split.\n\n```python\n# union of forgery and pristine masks covers both source and destination regions\nmask = np.maximum(forgery_mask, pristine_mask)\nmask_objects = np.unique(mask)[1:]  # skip background\nmask = np.stack([(mask == i).astype(np.uint8) for i in mask_objects], axis=0)\n```\n\n**Compound copy-move** cases require counterfactual reconstruction to generate a paired authentic sample, and come in two sub-types:\n\n- **Intra-panel**: Each case directory contains a clean version of the forged panel (`panel_pristine.png`). A pristine image is reconstructed by locating the forged region bounding box from `annotations.json`, resizing `panel_pristine.png` to match it, and compositing it back into the forged figure:\n\n```python\npristine_panel = cv2.resize(pristine_panel, (w, h), interpolation=cv2.INTER_LINEAR)\npristine_image = forgery_image.copy()\npristine_image[y0:y1, x0:x1] = pristine_panel\n```\n\n- **Inter-panel**: No pre-existing clean panel is available. A replacement image is sampled from the RSIID source image pool (`datasetSrc.csv`), matched to the same `subset_tag` as the forged panel and excluding images already present in the figure, then resized and composited into the forgery region:\n\n```python\n# sample a same-category image that was not already used in the figure\ncandidate_ids = df_rsiid_src.loc[\n    (df_rsiid_src['subset_tag'] == forgery_image_subset_tag) &\n    ~df_rsiid_src['dataPath'].isin(used_image_ids),\n    'dataPath'\n]\nreplace_image = cv2.imread(rsiid_src_directory / np.random.choice(candidate_ids))\nreplace_image = cv2.resize(replace_image, (w, h), interpolation=cv2.INTER_LINEAR)\npristine_image = forgery_image.copy()\npristine_image[y0:y1, x0:x1] = replace_image\n```\n\nIn both compound sub-types every forged image produces a paired authentic counterpart with the same image ID. All sources are merged into a single dataset:\n\n| Dataset | Authentic | Forged | Total |\n|---|---|---|---|\n| Competition train | 2,377 | 2,751 | 5,128 |\n| Competition supplemental | — | 48 | 48 |\n| RSIID simple copy-move | 2,923 | 5,390 | 8,313 |\n| RSIID compound copy-move | 19,000 | 19,000 | 38,000 |\n| **Total** | **24,300** | **27,189** | **51,489** |\n\n## Validation\n\nThe dataset is split into 5 folds using `StratifiedGroupKFold`, stratified on the joint combination of class and dataset source so that each fold preserves the original distribution across all dataset-class groups. Samples are grouped by `image_id`, which ensures that forged/authentic counterfactual pairs from the compound copy-move reconstruction always land in the same fold and never leak across the train/validation boundary.\n\n## Model\n\nThe segmentation model is `DinoV2Segmenter`, a `facebook/dinov2-base` Vision Transformer encoder paired with a lightweight convolutional decoder. The encoder is fully frozen except for the last 5 transformer blocks and the final LayerNorm; patch embeddings remain frozen throughout.\n\nAt inference, the encoder's output token sequence is stripped of the CLS token, reshaped into a 2D spatial feature map, and passed to `DinoDecoder`, which upsamples it back to the input resolution:\n\n```python\ndef forward(self, x: torch.Tensor) -> torch.Tensor:\n    H, W = x.shape[-2:]\n    out = self.encoder(pixel_values=x).last_hidden_state  # (B, N+1, 768)\n    feature_map = self._tokens_to_map(out)                # (B, 768, H/14, W/14)\n    return self.decoder(feature_map, (H, W))              # (B, 1, H, W)\n\ndef _tokens_to_map(self, tokens: torch.Tensor) -> torch.Tensor:\n    B, N, C = tokens.shape\n    num_reg = int(getattr(self.encoder.config, 'num_register_tokens', 0) or 0)\n    patch_tokens = tokens[:, 1 + num_reg:, :]  # drop CLS and register tokens\n    s = int(math.isqrt(patch_tokens.shape[1]))\n    return patch_tokens.permute(0, 2, 1).reshape(B, C, s, s)\n```\n\n`DinoDecoder` is a three-block convolutional head that halves the channel width at each stage and inserts a 2× bilinear upsample between the first two blocks. A final interpolation brings the output to exactly the input image size:\n\n```python\nclass DinoDecoder(nn.Module):\n    def __init__(self, in_ch, mid_ch=256, out_ch=1, dropout=0.1):\n        self.block1 = nn.Sequential(Conv2d(in_ch,       mid_ch,     3, padding=1), ReLU(), Dropout2d(dropout))\n        self.block2 = nn.Sequential(Conv2d(mid_ch,      mid_ch//2,  3, padding=1), ReLU(), Dropout2d(dropout))\n        self.block3 = nn.Sequential(Conv2d(mid_ch//2,   mid_ch//4,  3, padding=1), ReLU())\n        self.out_conv = Conv2d(mid_ch//4, out_ch, kernel_size=1)\n\n    def forward(self, feature_map, out_size):\n        x = self.block1(feature_map)\n        x = F.interpolate(x, scale_factor=2.0, mode='bilinear', align_corners=False)\n        x = self.block2(x)\n        x = F.interpolate(x, scale_factor=2.0, mode='bilinear', align_corners=False)\n        x = self.block3(x)\n        x = self.out_conv(x)\n        return F.interpolate(x, size=out_size, mode='bilinear', align_corners=False)\n```\n\n## Training\n\nAdamW is used with differential learning rates: the decoder is trained at `lr = 1e-3` while the unfrozen encoder blocks receive a 10× smaller rate `lr_encoder = 1e-4`, preventing the pretrained representations from being overwritten too quickly. Both parameter groups share `weight_decay = 1e-3`.\n\n`BCEWithLogitsLoss` is used for binary pixel-level supervision. Training runs with Automatic Mixed Precision (`bfloat16`) and gradient norm clipping at 1.0. Models are trained for 50 epochs with a batch size of 64, and checkpoints are saved whenever validation loss or mean image score improves so both type of \"best\" model can be explored.\n\nTraining images are resized to 512×512 and augmented with transpose, vertical flip, and horizontal flip (each at p=0.5). Light color jitter is applied with low probability: random brightness/contrast (±0.15, p=0.1), hue/saturation/value shift (±5, p=0.1), and JPEG compression (quality 95–100, p=0.1) to simulate mild compression artifacts present in real scientific figures.\n\n| Transform | Parameters | p |\n|---|---|---|\n| Transpose | — | 0.5 |\n| Vertical flip | — | 0.5 |\n| Horizontal flip | — | 0.5 |\n| RandomBrightnessContrast | limit ±0.15 | 0.1 |\n| HueSaturationValue | shift ±5 | 0.1 |\n| ImageCompression | JPEG, quality 95–100 | 0.1 |\n\n## Inference\nAt test-time, 8 augmented versions of each image are averaged: 4 rotations (0°, 90°, 180°, 270°) × 2 horizontal flip states. Each augmented prediction is inverted back to the original orientation before averaging.\n\n## Post-processing\n\nEach image is scored independently. Authentic images receive 1.0 for a correct `\"authentic\"` prediction and 0.0 for any mask. Forged images are scored by the optimal instance-level F1 computed via the Hungarian algorithm matching predicted masks to ground-truth masks. Predicting more mask instances than the ground truth carries a penalty:\n\n```python\nexcess_predictions_penalty = len(gt_masks) / max(len(pred_masks), len(gt_masks))\nscore = np.mean(f1_score_matrix[row_ind, col_ind]) * excess_predictions_penalty\n```\n\nThe naive baseline of predicting everything as `\"authentic\"` achieves a score equal to the fraction of authentic images in the dataset. Any useful submission must exceed this.\n\nThese two properties; the cost of a false positive and the penalty for excess mask instances, directly shape every post-processing decision: the forgery/authentic threshold must be conservative enough to avoid tagging authentic images as forged, and when a forgery is detected the entire probability map is merged into a single mask rather than split into separate instances to avoid the excess predictions penalty.\n\nOOF probability maps are resized back to the original image resolution with cubic interpolation. A two-step rule decides the final prediction for each image: if the 99.9th percentile (`q999`) of the probability map is at least 0.95, the image is predicted as forged and the mask is thresholded at 0.3; otherwise the image is predicted as `\"authentic\"`.\n\n```python\nif np.quantile(probability_map, 0.999) >= 0.95:\n    probability_map = cv2.resize(probability_map, (orig_w, orig_h), interpolation=cv2.INTER_CUBIC)\n    mask = (probability_map > THRESHOLD).astype(np.uint8)\n    prediction = rle_encode([mask])\nelse:\n    prediction = 'authentic'\n```\n\nThe intuition is that a copy-move forgery creates a localized region the model responds to with near-certain confidence. Even if that region is small, it drives the extreme tail of the probability distribution sharply upward. On authentic images the model stays uncertain across the board, so the tail never breaks above the threshold.\n\nThe `q999` gate is intentionally conservative: genuinely forged images produce a tight cluster of very high-confidence pixels that pushes the extreme tail well above 0.95, while authentic images stay flat near zero throughout. The thresholded map is always submitted as a single mask instance. Predicting multiple instances when the ground truth has fewer triggers the excess predictions penalty and scales the score down by `len(gt_masks) / len(pred_masks)`, so submitting one merged mask is safer than attempting to separate individual regions.\n\n## Submission Selection\n\nThis function computes the same set of metrics globally and separately for each dataset split. For each split it reports: the mean image score, the baseline score (fraction of authentic images, what predicting everything as `\"authentic\"` would achieve), lift over that baseline, TP/FP/TN/FN counts, precision, recall, accuracy, and the two impact terms `tp_impact` and `fp_impact`.\n\n```\ndef get_detection_stats(df: pd.DataFrame) -> dict[str, int | float]:\n\n    \"\"\"\n    Calculates detection statistics globally and per dataset.\n\n    Parameters\n    ----------\n    df: pd.DataFrame\n        DataFrame with 'target', 'prediction', 'score', and optionally 'dataset'.\n\n    Returns\n    -------\n    dict[str, int | float]\n        A dictionary containing global stats and per-dataset stats.\n    \"\"\"\n\n    def _compute_metrics(sub_df: pd.DataFrame) -> dict[str, int | float]:\n\n        is_forged_gt = sub_df['target'] != 'authentic'\n        has_mask_pred = sub_df['prediction'] != 'authentic'\n\n        tp_mask = (is_forged_gt) & (has_mask_pred)\n        fp_mask = (~is_forged_gt) & (has_mask_pred)\n        tn_mask = (~is_forged_gt) & (~has_mask_pred)\n        fn_mask = (is_forged_gt) & (~has_mask_pred)\n\n        tp = tp_mask.sum()\n        fp = fp_mask.sum()\n        tn = tn_mask.sum()\n        fn = fn_mask.sum()\n        \n        total = len(sub_df)\n\n        total_tp_gain = sub_df.loc[tp_mask, 'score'].sum()\n        total_fp_loss = float(fp) * 1.0\n        fp_impact = total_fp_loss / total\n        tp_impact = total_tp_gain / total\n\n        current_mean_score = float(sub_df['score'].mean()) if not sub_df.empty else 0.0\n        count_authentic_gt = int((~is_forged_gt).sum())\n        baseline_score = float(count_authentic_gt / total)\n        lift = current_mean_score - baseline_score\n\n        return {\n            'total_samples': int(total),\n            'accuracy': float((tp + tn) / total) if total > 0 else 0.0,\n            'precision': float(tp / (tp + fp)) if (tp + fp) > 0 else 0.0,\n            'recall': float(tp / (tp + fn)) if (tp + fn) > 0 else 0.0,\n            'count_forged_correct (TP)': int(tp),\n            'count_authentic_correct (TN)': int(tn),\n            'count_false_positives (FP)': int(fp),\n            'count_missed_forgeries (FN)': int(fn),\n            'tp_impact': float(tp_impact),\n            'fp_impact': float(fp_impact),\n            'mean_image_score': current_mean_score,\n            'baseline_score': baseline_score,\n            'lift_over_baseline': lift,\n        }\n\n    stats = _compute_metrics(df)\n    if 'dataset' in df.columns:\n        for ds_name, ds_df in df.groupby('dataset'):\n            ds_stats = _compute_metrics(ds_df)\n            for key, value in ds_stats.items():\n                stats[f'{ds_name}_{key}'] = value\n    \n    return stats\n```\n\nEach dataset split has a different authentic/forged ratio, so the baseline score differs across splits and a threshold that improves the global score can still hurt an individual split if it generates too many false positives there. Tracking lift per split catches this: a submission is only considered safe if every split shows a positive lift, not just the aggregate.\n\nWithin each split, `tp_impact` and `fp_impact` make the cost/benefit of the threshold explicit. `tp_impact` measures how much the correctly detected forgeries contribute to the score; `fp_impact` measures how much is being lost to false positives since each FP on an authentic image converts a guaranteed 1.0 into 0.0. The selected threshold is the most conservative one where every dataset split still shows a positive lift, meaning the TP gain outweighs the FP loss in each split independently rather than the threshold that maximises raw TP recall.\n\nThis framework made submission selection straightforward: rather than guessing at a threshold and hoping it generalised, the per-split breakdown showed exactly where false positives were appearing and whether the lift was real or an artefact of one dominant split. The final submission was chosen as the most conservative configuration that demonstrated consistent positive lift across all splits with minimal fp_impact. The exact CV and public leaderboard scores were not recorded at the time and are no longer available.\n\n"
  }
}