{
  "id": 704254,
  "title": "🥈 2nd Place Solution - Synthetic Image Attribution",
  "url": "/competitions/dlmmdd-workshop-synthetic-source-attribution-challenge/discussion/704254",
  "author_name": "Maxim",
  "post_date": "2026-06-04T02:11:48.675000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Thank you!</h1>\n<p>Before diving into the solution, we would like to thank the organizers for designing such an educational competition.</p>\n<p>What initially appeared to be a straightforward image classification task turned out to be a comprehensive exploration of modern computer vision techniques: strong CNN backbones, data augmentation, distribution shift, test-time augmentation, pseudo-labeling, model ensembling, frequency-domain analysis, and validation strategy design.</p>\n<p>The competition was particularly valuable because leaderboard improvements rarely came from simply using larger models. Instead, success depended on understanding the underlying data distribution and carefully matching the training process to the hidden test conditions.</p>\n<p>Many of the ideas explored in this competition directly reflect techniques commonly used in state-of-the-art Kaggle solutions and real-world machine learning systems, making it one of the most educational computer vision competitions we have participated in.</p>\n<h1>Summary</h1>\n<p>Our solution achieved 0.998 Public / 0.9973 Private LB (2nd place) using EfficientNetV2-L/XL with a key insight: simulating test-time post-processing operations during training closes the distribution gap between clean training data and post-processed test data. Pseudo-labeling further improved results by exposing the model to real JPEG AI and Super Resolution artifacts that cannot be simulated.</p>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public LB</td>\n<td>0.9980</td>\n</tr>\n<tr>\n<td>Private LB</td>\n<td>0.9973</td>\n</tr>\n<tr>\n<td>Final Rank</td>\n<td>2 / 92</td>\n</tr>\n</tbody>\n</table>\n<h1>Problem Understanding</h1>\n<p>The key challenge is a distribution mismatch: training images are clean, while every test image has 1–3 post-processing operations applied:\nJPEG compression, WebP compression, random central crop, resizing, small rotation, contrast/brightness adjustment, Gaussian blur, grayscale conversion, AI super-resolution, JPEG AI compression. Without addressing this gap, models trained on clean images struggle to generalize to post-processed test images.</p>\n<h1>Model Architecture</h1>\n<p>Backbone: EfficientNetV2-L/XL (tf_efficientnetv2_l/xl, pretrained on ImageNet-21k)\nInput size: 384×384/512x512\nTraining: 5-fold StratifiedKFold, EMA (decay=0.99)\nLoss: CrossEntropy with label smoothing=0.1\nAugmentations: MixUp (α=0.2) + CutMix (α=1.0)\nOptimizer: AdamW (lr=3e-4)\nScheduler: CosineAnnealingLR</p>\n<h1>Training Pipeline</h1>\n<p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F32676712%2F34a94f8bec22d87e9e947226f07ad309%2FFigure%201%20-%20Pipeline.jpg?generation=1780532487226529&amp;alt=media\" style=\"border-radius: 18px\">\n</p>\n<p>The overall training procedure consisted of:</p>\n<ol>\n<li>Test-op simulation</li>\n<li>EfficientNetV2 training</li>\n<li>Teacher ensemble prediction</li>\n<li>Pseudo-label generation</li>\n<li>Retraining from scratch</li>\n<li>Final ensemble</li>\n</ol>\n<h1>Key Insight: Test-Time Operations Simulation</h1>\n<p>The most impactful improvement was simulating test-time post-processing during training:</p>\n<pre><code>def apply_random_test_ops(img, num_ops=None):\n    ops = [\n        op_jpeg,               # JPEG compression, quality=30-95\n        op_webp,               # WebP compression, quality=30-95\n        op_random_central_crop, # center crop, scale=0.6-0.95\n        op_resize,             # resize, scale=0.4-1.5\n        op_small_rotation,     # rotation ±10° with aspect-preserving crop\n        op_contrast_brightness, # contrast/brightness, factor=0.6-1.6\n        op_gaussian_blur,      # Gaussian blur, radius=0.5-2.5\n        op_grayscale,          # grayscale → RGB\n    ]\n    num_ops = random.randint(1, 3)\n    chosen  = random.sample(ops, num_ops)\n    for op in chosen:\n        img = op(img)\n    return img\n\n# Applied with probability 0.7 during training\nif random.random() &lt; 0.7:\n    img = apply_random_test_ops(img)\n</code></pre>\n<h1>Multi-Stage Pseudo-Labeling</h1>\n<p>We found that standard fine-tuning with low pseudo-label weight gave minimal improvement. The most effective configuration was training from scratch with full pseudo-label weight.</p>\n<pre><code>Stage 1: Train EfficientNetV2-L on clean data + test ops simulation\n         → Public LB: 0.996666\n\nStage 2: Generate pseudo-labels from Stage 1 models (threshold=0.90)\n         Train EfficientNetV2-L from scratch with pseudo-labels\n         → Public LB: 0.998000\n\nStage 3: Generate pseudo-labels from Stage 2 models (threshold=0.80)  \n         Train EfficientNetV2-XL (512px) from scratch with pseudo-labels\n         → Private LB: 0.997333\n</code></pre>\n<p>Key parameters:</p>\n<pre><code>threshold     = 0.80-0.90  # confidence threshold\npseudo_weight = 1.0         # full trust in pseudo-labels\nepochs        = 20          # full training from scratch\nlr            = 3e-4        # full learning rate (not fine-tuning lr)\n</code></pre>\n<p>Why training from scratch matters: fine-tuning with weight=0.3 gave negligible improvement because the pseudo-label signal was too weak relative to the pretrained weights. Training from scratch with weight=1.0 forces the model to fully incorporate real test-set artifacts (JPEG AI, Super Resolution) that cannot be simulated.\nWhy it works: test images contain real JPEG AI and AI Super Resolution artifacts. Pseudo-labels provide the only way to expose the model to these real artifacts during training.\nPseudo-labeled images (already post-processed) do NOT receive additional augmentation.</p>\n<h1>TTA</h1>\n<p>Only horizontal flip was beneficial. All other TTA variants hurt performance:</p>\n<table>\n<thead>\n<tr>\n<th>✅ Horizontal flip</th>\n<th>neutral to slightly positive</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>❌ Multiscale TTA</td>\n<td>-0.001 LB</td>\n</tr>\n<tr>\n<td>❌ Test ops TTA</td>\n<td>-0.003 LB</td>\n</tr>\n<tr>\n<td>❌ Grayscale/blur TTA</td>\n<td>-0.003 LB</td>\n</tr>\n</tbody>\n</table>\n<p>Interesting observation: multiscale TTA (256/384/480px) was beneficial in early experiments but became harmful after correcting the rotation augmentation. Our hypothesis is that the incorrect rotation (with black corners) inadvertently regularized the model to be more robust to input variations, making multiscale TTA helpful. After fixing rotation, the model became more precise but also more sensitive to input size changes.</p>\n<h1>What Didn't Work</h1>\n<p>Several approaches were investigated but did not improve leaderboard performance.</p>\n<h2>Specialized Pairwise Classifiers</h2>\n<p>Error analysis revealed that most remaining mistakes were concentrated in two confusion pairs:</p>\n<ul>\n<li>SD3 ↔ SD3.5</li>\n<li>Pixart ↔ Hunyuan</li>\n</ul>\n<p>Specialized binary classifiers were trained on top of the EfficientNetV2-L backbone.</p>\n<pre><code>❌ Binary classifier (SD3 vs SD3.5)\n   OOF accuracy: 0.9993\n   LB: 0.993\n\n❌ Binary classifier (Pixart vs Hunyuan)\n   OOF accuracy: 0.9993\n   LB: 0.993\n</code></pre>\n<p>Despite near-perfect validation accuracy, both models reduced leaderboard performance, suggesting severe overfitting to clean training images.</p>\n<h2>Frequency-Domain Features</h2>\n<pre><code>❌ FFT magnitude spectrum\n❌ Wavelet decomposition (db4, 3 levels)\n❌ FFT/Wavelet features added to the model\n</code></pre>\n<p>Frequency-domain representations were evaluated both as standalone features and as additional inputs to the neural network. None of the investigated approaches improved validation accuracy or leaderboard scores.</p>\n<h2>Alternative Architectures</h2>\n<pre><code>❌ ConvNeXt\n❌ Vision Transformer (ViT)\n</code></pre>\n<p>Using the same training pipeline and hyperparameters, both architectures underperformed EfficientNetV2-L and were not explored further.</p>\n<h2>Other Negative Results</h2>\n<pre><code>❌ EfficientNetV2-XL without pseudo-labeling\n❌ Multiscale TTA after fixing rotation augmentation\n❌ Pseudo-labeling with weight = 0.3\n</code></pre>\n<p>A consistent pattern throughout the competition was that methods improving clean validation metrics often failed to improve leaderboard performance. In contrast, approaches explicitly targeting the train-test distribution gap consistently produced the largest gains.</p>\n<h1>Error Analysis</h1>\n<p>To better approximate the hidden test distribution, validation folds were evaluated under simulated test-time post-processing and predictions were averaged over 5 stochastic runs.</p>\n<p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F32676712%2F91499ce9bb97124af4e6a224d36952c8%2Fconfusion_matrix_aug_oof.png?generation=1780538086145976&amp;alt=media\" style=\"border-radius: 18px\">\n</p>\n<p>The dominant confusion pairs were:</p>\n<table>\n<thead>\n<tr>\n<th>True class</th>\n<th>Predicted class</th>\n<th>Count</th>\n<th>Error rate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SD3.5</td>\n<td>SD3</td>\n<td>5</td>\n<td>0.0071</td>\n</tr>\n<tr>\n<td>Pixart</td>\n<td>Hunyuan</td>\n<td>4</td>\n<td>0.0057</td>\n</tr>\n<tr>\n<td>SD3</td>\n<td>SD3.5</td>\n<td>3</td>\n<td>0.0043</td>\n</tr>\n<tr>\n<td>Hunyuan</td>\n<td>Pixart</td>\n<td>2</td>\n<td>0.0029</td>\n</tr>\n<tr>\n<td>SDXL-Turbo</td>\n<td>Photon</td>\n<td>1</td>\n<td>0.0014</td>\n</tr>\n</tbody>\n</table>\n<p>Most errors came from SD3 ↔ SD3.5 and Pixart ↔ Hunyuan. This matches the intuition that attribution becomes hardest when generators share similar architectures or visual statistics.</p>\n<h1>Final Takeaways</h1>\n<p>The most important lesson from this competition was that train-test distribution alignment mattered more than architecture scaling or handcrafted features.</p>\n<p>Simulating post-processing operations during training produced the largest single improvement, while pseudo-labeling provided exposure to real artifacts that could not be reproduced synthetically.</p>\n<p>Many techniques that improved clean validation accuracy failed to improve leaderboard performance, highlighting the importance of matching the hidden test distribution rather than optimizing for clean validation metrics alone.</p>",
  "messages": [
    {
      "id": 3466362,
      "postDate": "2026-06-04T02:11:48.677Z",
      "content": "<h1>Thank you!</h1>\n<p>Before diving into the solution, we would like to thank the organizers for designing such an educational competition.</p>\n<p>What initially appeared to be a straightforward image classification task turned out to be a comprehensive exploration of modern computer vision techniques: strong CNN backbones, data augmentation, distribution shift, test-time augmentation, pseudo-labeling, model ensembling, frequency-domain analysis, and validation strategy design.</p>\n<p>The competition was particularly valuable because leaderboard improvements rarely came from simply using larger models. Instead, success depended on understanding the underlying data distribution and carefully matching the training process to the hidden test conditions.</p>\n<p>Many of the ideas explored in this competition directly reflect techniques commonly used in state-of-the-art Kaggle solutions and real-world machine learning systems, making it one of the most educational computer vision competitions we have participated in.</p>\n<h1>Summary</h1>\n<p>Our solution achieved 0.998 Public / 0.9973 Private LB (2nd place) using EfficientNetV2-L/XL with a key insight: simulating test-time post-processing operations during training closes the distribution gap between clean training data and post-processed test data. Pseudo-labeling further improved results by exposing the model to real JPEG AI and Super Resolution artifacts that cannot be simulated.</p>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public LB</td>\n<td>0.9980</td>\n</tr>\n<tr>\n<td>Private LB</td>\n<td>0.9973</td>\n</tr>\n<tr>\n<td>Final Rank</td>\n<td>2 / 92</td>\n</tr>\n</tbody>\n</table>\n<h1>Problem Understanding</h1>\n<p>The key challenge is a distribution mismatch: training images are clean, while every test image has 1–3 post-processing operations applied:\nJPEG compression, WebP compression, random central crop, resizing, small rotation, contrast/brightness adjustment, Gaussian blur, grayscale conversion, AI super-resolution, JPEG AI compression. Without addressing this gap, models trained on clean images struggle to generalize to post-processed test images.</p>\n<h1>Model Architecture</h1>\n<p>Backbone: EfficientNetV2-L/XL (tf_efficientnetv2_l/xl, pretrained on ImageNet-21k)\nInput size: 384×384/512x512\nTraining: 5-fold StratifiedKFold, EMA (decay=0.99)\nLoss: CrossEntropy with label smoothing=0.1\nAugmentations: MixUp (α=0.2) + CutMix (α=1.0)\nOptimizer: AdamW (lr=3e-4)\nScheduler: CosineAnnealingLR</p>\n<h1>Training Pipeline</h1>\n<p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F32676712%2F34a94f8bec22d87e9e947226f07ad309%2FFigure%201%20-%20Pipeline.jpg?generation=1780532487226529&amp;alt=media\" style=\"border-radius: 18px\">\n</p>\n<p>The overall training procedure consisted of:</p>\n<ol>\n<li>Test-op simulation</li>\n<li>EfficientNetV2 training</li>\n<li>Teacher ensemble prediction</li>\n<li>Pseudo-label generation</li>\n<li>Retraining from scratch</li>\n<li>Final ensemble</li>\n</ol>\n<h1>Key Insight: Test-Time Operations Simulation</h1>\n<p>The most impactful improvement was simulating test-time post-processing during training:</p>\n<pre><code>def apply_random_test_ops(img, num_ops=None):\n    ops = [\n        op_jpeg,               # JPEG compression, quality=30-95\n        op_webp,               # WebP compression, quality=30-95\n        op_random_central_crop, # center crop, scale=0.6-0.95\n        op_resize,             # resize, scale=0.4-1.5\n        op_small_rotation,     # rotation ±10° with aspect-preserving crop\n        op_contrast_brightness, # contrast/brightness, factor=0.6-1.6\n        op_gaussian_blur,      # Gaussian blur, radius=0.5-2.5\n        op_grayscale,          # grayscale → RGB\n    ]\n    num_ops = random.randint(1, 3)\n    chosen  = random.sample(ops, num_ops)\n    for op in chosen:\n        img = op(img)\n    return img\n\n# Applied with probability 0.7 during training\nif random.random() &lt; 0.7:\n    img = apply_random_test_ops(img)\n</code></pre>\n<h1>Multi-Stage Pseudo-Labeling</h1>\n<p>We found that standard fine-tuning with low pseudo-label weight gave minimal improvement. The most effective configuration was training from scratch with full pseudo-label weight.</p>\n<pre><code>Stage 1: Train EfficientNetV2-L on clean data + test ops simulation\n         → Public LB: 0.996666\n\nStage 2: Generate pseudo-labels from Stage 1 models (threshold=0.90)\n         Train EfficientNetV2-L from scratch with pseudo-labels\n         → Public LB: 0.998000\n\nStage 3: Generate pseudo-labels from Stage 2 models (threshold=0.80)  \n         Train EfficientNetV2-XL (512px) from scratch with pseudo-labels\n         → Private LB: 0.997333\n</code></pre>\n<p>Key parameters:</p>\n<pre><code>threshold     = 0.80-0.90  # confidence threshold\npseudo_weight = 1.0         # full trust in pseudo-labels\nepochs        = 20          # full training from scratch\nlr            = 3e-4        # full learning rate (not fine-tuning lr)\n</code></pre>\n<p>Why training from scratch matters: fine-tuning with weight=0.3 gave negligible improvement because the pseudo-label signal was too weak relative to the pretrained weights. Training from scratch with weight=1.0 forces the model to fully incorporate real test-set artifacts (JPEG AI, Super Resolution) that cannot be simulated.\nWhy it works: test images contain real JPEG AI and AI Super Resolution artifacts. Pseudo-labels provide the only way to expose the model to these real artifacts during training.\nPseudo-labeled images (already post-processed) do NOT receive additional augmentation.</p>\n<h1>TTA</h1>\n<p>Only horizontal flip was beneficial. All other TTA variants hurt performance:</p>\n<table>\n<thead>\n<tr>\n<th>✅ Horizontal flip</th>\n<th>neutral to slightly positive</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>❌ Multiscale TTA</td>\n<td>-0.001 LB</td>\n</tr>\n<tr>\n<td>❌ Test ops TTA</td>\n<td>-0.003 LB</td>\n</tr>\n<tr>\n<td>❌ Grayscale/blur TTA</td>\n<td>-0.003 LB</td>\n</tr>\n</tbody>\n</table>\n<p>Interesting observation: multiscale TTA (256/384/480px) was beneficial in early experiments but became harmful after correcting the rotation augmentation. Our hypothesis is that the incorrect rotation (with black corners) inadvertently regularized the model to be more robust to input variations, making multiscale TTA helpful. After fixing rotation, the model became more precise but also more sensitive to input size changes.</p>\n<h1>What Didn't Work</h1>\n<p>Several approaches were investigated but did not improve leaderboard performance.</p>\n<h2>Specialized Pairwise Classifiers</h2>\n<p>Error analysis revealed that most remaining mistakes were concentrated in two confusion pairs:</p>\n<ul>\n<li>SD3 ↔ SD3.5</li>\n<li>Pixart ↔ Hunyuan</li>\n</ul>\n<p>Specialized binary classifiers were trained on top of the EfficientNetV2-L backbone.</p>\n<pre><code>❌ Binary classifier (SD3 vs SD3.5)\n   OOF accuracy: 0.9993\n   LB: 0.993\n\n❌ Binary classifier (Pixart vs Hunyuan)\n   OOF accuracy: 0.9993\n   LB: 0.993\n</code></pre>\n<p>Despite near-perfect validation accuracy, both models reduced leaderboard performance, suggesting severe overfitting to clean training images.</p>\n<h2>Frequency-Domain Features</h2>\n<pre><code>❌ FFT magnitude spectrum\n❌ Wavelet decomposition (db4, 3 levels)\n❌ FFT/Wavelet features added to the model\n</code></pre>\n<p>Frequency-domain representations were evaluated both as standalone features and as additional inputs to the neural network. None of the investigated approaches improved validation accuracy or leaderboard scores.</p>\n<h2>Alternative Architectures</h2>\n<pre><code>❌ ConvNeXt\n❌ Vision Transformer (ViT)\n</code></pre>\n<p>Using the same training pipeline and hyperparameters, both architectures underperformed EfficientNetV2-L and were not explored further.</p>\n<h2>Other Negative Results</h2>\n<pre><code>❌ EfficientNetV2-XL without pseudo-labeling\n❌ Multiscale TTA after fixing rotation augmentation\n❌ Pseudo-labeling with weight = 0.3\n</code></pre>\n<p>A consistent pattern throughout the competition was that methods improving clean validation metrics often failed to improve leaderboard performance. In contrast, approaches explicitly targeting the train-test distribution gap consistently produced the largest gains.</p>\n<h1>Error Analysis</h1>\n<p>To better approximate the hidden test distribution, validation folds were evaluated under simulated test-time post-processing and predictions were averaged over 5 stochastic runs.</p>\n<p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F32676712%2F91499ce9bb97124af4e6a224d36952c8%2Fconfusion_matrix_aug_oof.png?generation=1780538086145976&amp;alt=media\" style=\"border-radius: 18px\">\n</p>\n<p>The dominant confusion pairs were:</p>\n<table>\n<thead>\n<tr>\n<th>True class</th>\n<th>Predicted class</th>\n<th>Count</th>\n<th>Error rate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SD3.5</td>\n<td>SD3</td>\n<td>5</td>\n<td>0.0071</td>\n</tr>\n<tr>\n<td>Pixart</td>\n<td>Hunyuan</td>\n<td>4</td>\n<td>0.0057</td>\n</tr>\n<tr>\n<td>SD3</td>\n<td>SD3.5</td>\n<td>3</td>\n<td>0.0043</td>\n</tr>\n<tr>\n<td>Hunyuan</td>\n<td>Pixart</td>\n<td>2</td>\n<td>0.0029</td>\n</tr>\n<tr>\n<td>SDXL-Turbo</td>\n<td>Photon</td>\n<td>1</td>\n<td>0.0014</td>\n</tr>\n</tbody>\n</table>\n<p>Most errors came from SD3 ↔ SD3.5 and Pixart ↔ Hunyuan. This matches the intuition that attribution becomes hardest when generators share similar architectures or visual statistics.</p>\n<h1>Final Takeaways</h1>\n<p>The most important lesson from this competition was that train-test distribution alignment mattered more than architecture scaling or handcrafted features.</p>\n<p>Simulating post-processing operations during training produced the largest single improvement, while pseudo-labeling provided exposure to real artifacts that could not be reproduced synthetically.</p>\n<p>Many techniques that improved clean validation accuracy failed to improve leaderboard performance, highlighting the importance of matching the hidden test distribution rather than optimizing for clean validation metrics alone.</p>",
      "rawMarkdown": "# Thank you!\n\nBefore diving into the solution, we would like to thank the organizers for designing such an educational competition.\n\nWhat initially appeared to be a straightforward image classification task turned out to be a comprehensive exploration of modern computer vision techniques: strong CNN backbones, data augmentation, distribution shift, test-time augmentation, pseudo-labeling, model ensembling, frequency-domain analysis, and validation strategy design.\n\nThe competition was particularly valuable because leaderboard improvements rarely came from simply using larger models. Instead, success depended on understanding the underlying data distribution and carefully matching the training process to the hidden test conditions.\n\nMany of the ideas explored in this competition directly reflect techniques commonly used in state-of-the-art Kaggle solutions and real-world machine learning systems, making it one of the most educational computer vision competitions we have participated in.\n\n# Summary\n\nOur solution achieved 0.998 Public / 0.9973 Private LB (2nd place) using EfficientNetV2-L/XL with a key insight: simulating test-time post-processing operations during training closes the distribution gap between clean training data and post-processed test data. Pseudo-labeling further improved results by exposing the model to real JPEG AI and Super Resolution artifacts that cannot be simulated.\n\n| Metric | Score |\n|----------|----------|\n| Public LB | 0.9980 |\n| Private LB | 0.9973 |\n| Final Rank | 2 / 92 |\n\n# Problem Understanding\n\nThe key challenge is a distribution mismatch: training images are clean, while every test image has 1–3 post-processing operations applied:\nJPEG compression, WebP compression, random central crop, resizing, small rotation, contrast/brightness adjustment, Gaussian blur, grayscale conversion, AI super-resolution, JPEG AI compression. Without addressing this gap, models trained on clean images struggle to generalize to post-processed test images.\n\n# Model Architecture\n\nBackbone: EfficientNetV2-L/XL (tf_efficientnetv2_l/xl, pretrained on ImageNet-21k)\nInput size: 384×384/512x512\nTraining: 5-fold StratifiedKFold, EMA (decay=0.99)\nLoss: CrossEntropy with label smoothing=0.1\nAugmentations: MixUp (α=0.2) + CutMix (α=1.0)\nOptimizer: AdamW (lr=3e-4)\nScheduler: CosineAnnealingLR\n\n# Training Pipeline\n\n<p align=\"center\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F32676712%2F34a94f8bec22d87e9e947226f07ad309%2FFigure%201%20-%20Pipeline.jpg?generation=1780532487226529&alt=media\"\n       width=\"700\" style=\"border-radius: 18px;\">\n</p>\n\nThe overall training procedure consisted of:\n\n1. Test-op simulation\n2. EfficientNetV2 training\n3. Teacher ensemble prediction\n4. Pseudo-label generation\n5. Retraining from scratch\n6. Final ensemble\n\n# Key Insight: Test-Time Operations Simulation\n\nThe most impactful improvement was simulating test-time post-processing during training:\n```\ndef apply_random_test_ops(img, num_ops=None):\n    ops = [\n        op_jpeg,               # JPEG compression, quality=30-95\n        op_webp,               # WebP compression, quality=30-95\n        op_random_central_crop, # center crop, scale=0.6-0.95\n        op_resize,             # resize, scale=0.4-1.5\n        op_small_rotation,     # rotation ±10° with aspect-preserving crop\n        op_contrast_brightness, # contrast/brightness, factor=0.6-1.6\n        op_gaussian_blur,      # Gaussian blur, radius=0.5-2.5\n        op_grayscale,          # grayscale → RGB\n    ]\n    num_ops = random.randint(1, 3)\n    chosen  = random.sample(ops, num_ops)\n    for op in chosen:\n        img = op(img)\n    return img\n\n# Applied with probability 0.7 during training\nif random.random() < 0.7:\n    img = apply_random_test_ops(img)\n```\n\n# Multi-Stage Pseudo-Labeling\nWe found that standard fine-tuning with low pseudo-label weight gave minimal improvement. The most effective configuration was training from scratch with full pseudo-label weight.\n```\nStage 1: Train EfficientNetV2-L on clean data + test ops simulation\n         → Public LB: 0.996666\n\nStage 2: Generate pseudo-labels from Stage 1 models (threshold=0.90)\n         Train EfficientNetV2-L from scratch with pseudo-labels\n         → Public LB: 0.998000\n\nStage 3: Generate pseudo-labels from Stage 2 models (threshold=0.80)  \n         Train EfficientNetV2-XL (512px) from scratch with pseudo-labels\n         → Private LB: 0.997333\n```\nKey parameters:\n```\nthreshold     = 0.80-0.90  # confidence threshold\npseudo_weight = 1.0         # full trust in pseudo-labels\nepochs        = 20          # full training from scratch\nlr            = 3e-4        # full learning rate (not fine-tuning lr)\n```\nWhy training from scratch matters: fine-tuning with weight=0.3 gave negligible improvement because the pseudo-label signal was too weak relative to the pretrained weights. Training from scratch with weight=1.0 forces the model to fully incorporate real test-set artifacts (JPEG AI, Super Resolution) that cannot be simulated.\nWhy it works: test images contain real JPEG AI and AI Super Resolution artifacts. Pseudo-labels provide the only way to expose the model to these real artifacts during training.\nPseudo-labeled images (already post-processed) do NOT receive additional augmentation.\n\n# TTA\nOnly horizontal flip was beneficial. All other TTA variants hurt performance:\n\n|✅ Horizontal flip | neutral to slightly positive |\n| --- | --- |\n| ❌ Multiscale TTA | -0.001 LB |\n|❌ Test ops TTA | -0.003 LB |\n|❌ Grayscale/blur TTA | -0.003 LB |\n\nInteresting observation: multiscale TTA (256/384/480px) was beneficial in early experiments but became harmful after correcting the rotation augmentation. Our hypothesis is that the incorrect rotation (with black corners) inadvertently regularized the model to be more robust to input variations, making multiscale TTA helpful. After fixing rotation, the model became more precise but also more sensitive to input size changes.\n\n# What Didn't Work\n\nSeveral approaches were investigated but did not improve leaderboard performance.\n\n## Specialized Pairwise Classifiers\n\nError analysis revealed that most remaining mistakes were concentrated in two confusion pairs:\n\n- SD3 ↔ SD3.5\n- Pixart ↔ Hunyuan\n\nSpecialized binary classifiers were trained on top of the EfficientNetV2-L backbone.\n\n```text\n❌ Binary classifier (SD3 vs SD3.5)\n   OOF accuracy: 0.9993\n   LB: 0.993\n\n❌ Binary classifier (Pixart vs Hunyuan)\n   OOF accuracy: 0.9993\n   LB: 0.993\n```\n\nDespite near-perfect validation accuracy, both models reduced leaderboard performance, suggesting severe overfitting to clean training images.\n\n## Frequency-Domain Features\n\n```text\n❌ FFT magnitude spectrum\n❌ Wavelet decomposition (db4, 3 levels)\n❌ FFT/Wavelet features added to the model\n```\n\nFrequency-domain representations were evaluated both as standalone features and as additional inputs to the neural network. None of the investigated approaches improved validation accuracy or leaderboard scores.\n\n## Alternative Architectures\n\n```text\n❌ ConvNeXt\n❌ Vision Transformer (ViT)\n```\n\nUsing the same training pipeline and hyperparameters, both architectures underperformed EfficientNetV2-L and were not explored further.\n\n## Other Negative Results\n\n```text\n❌ EfficientNetV2-XL without pseudo-labeling\n❌ Multiscale TTA after fixing rotation augmentation\n❌ Pseudo-labeling with weight = 0.3\n```\n\nA consistent pattern throughout the competition was that methods improving clean validation metrics often failed to improve leaderboard performance. In contrast, approaches explicitly targeting the train-test distribution gap consistently produced the largest gains.\n\n# Error Analysis\n\nTo better approximate the hidden test distribution, validation folds were evaluated under simulated test-time post-processing and predictions were averaged over 5 stochastic runs.\n\n<p align=\"center\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F32676712%2F91499ce9bb97124af4e6a224d36952c8%2Fconfusion_matrix_aug_oof.png?generation=1780538086145976&alt=media\" width=\"85%\" style=\"border-radius: 18px;\">\n</p>\n\nThe dominant confusion pairs were:\n\n| True class | Predicted class | Count | Error rate |\n|---|---:|---:|---:|\n| SD3.5 | SD3 | 5 | 0.0071 |\n| Pixart | Hunyuan | 4 | 0.0057 |\n| SD3 | SD3.5 | 3 | 0.0043 |\n| Hunyuan | Pixart | 2 | 0.0029 |\n| SDXL-Turbo | Photon | 1 | 0.0014 |\n\nMost errors came from SD3 ↔ SD3.5 and Pixart ↔ Hunyuan. This matches the intuition that attribution becomes hardest when generators share similar architectures or visual statistics.\n\n# Final Takeaways\n\nThe most important lesson from this competition was that train-test distribution alignment mattered more than architecture scaling or handcrafted features.\n\nSimulating post-processing operations during training produced the largest single improvement, while pseudo-labeling provided exposure to real artifacts that could not be reproduced synthetically.\n\nMany techniques that improved clean validation accuracy failed to improve leaderboard performance, highlighting the importance of matching the hidden test distribution rather than optimizing for clean validation metrics alone.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3466362": "# Thank you!\n\nBefore diving into the solution, we would like to thank the organizers for designing such an educational competition.\n\nWhat initially appeared to be a straightforward image classification task turned out to be a comprehensive exploration of modern computer vision techniques: strong CNN backbones, data augmentation, distribution shift, test-time augmentation, pseudo-labeling, model ensembling, frequency-domain analysis, and validation strategy design.\n\nThe competition was particularly valuable because leaderboard improvements rarely came from simply using larger models. Instead, success depended on understanding the underlying data distribution and carefully matching the training process to the hidden test conditions.\n\nMany of the ideas explored in this competition directly reflect techniques commonly used in state-of-the-art Kaggle solutions and real-world machine learning systems, making it one of the most educational computer vision competitions we have participated in.\n\n# Summary\n\nOur solution achieved 0.998 Public / 0.9973 Private LB (2nd place) using EfficientNetV2-L/XL with a key insight: simulating test-time post-processing operations during training closes the distribution gap between clean training data and post-processed test data. Pseudo-labeling further improved results by exposing the model to real JPEG AI and Super Resolution artifacts that cannot be simulated.\n\n| Metric | Score |\n|----------|----------|\n| Public LB | 0.9980 |\n| Private LB | 0.9973 |\n| Final Rank | 2 / 92 |\n\n# Problem Understanding\n\nThe key challenge is a distribution mismatch: training images are clean, while every test image has 1–3 post-processing operations applied:\nJPEG compression, WebP compression, random central crop, resizing, small rotation, contrast/brightness adjustment, Gaussian blur, grayscale conversion, AI super-resolution, JPEG AI compression. Without addressing this gap, models trained on clean images struggle to generalize to post-processed test images.\n\n# Model Architecture\n\nBackbone: EfficientNetV2-L/XL (tf_efficientnetv2_l/xl, pretrained on ImageNet-21k)\nInput size: 384×384/512x512\nTraining: 5-fold StratifiedKFold, EMA (decay=0.99)\nLoss: CrossEntropy with label smoothing=0.1\nAugmentations: MixUp (α=0.2) + CutMix (α=1.0)\nOptimizer: AdamW (lr=3e-4)\nScheduler: CosineAnnealingLR\n\n# Training Pipeline\n\n<p align=\"center\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F32676712%2F34a94f8bec22d87e9e947226f07ad309%2FFigure%201%20-%20Pipeline.jpg?generation=1780532487226529&alt=media\"\n       width=\"700\" style=\"border-radius: 18px;\">\n</p>\n\nThe overall training procedure consisted of:\n\n1. Test-op simulation\n2. EfficientNetV2 training\n3. Teacher ensemble prediction\n4. Pseudo-label generation\n5. Retraining from scratch\n6. Final ensemble\n\n# Key Insight: Test-Time Operations Simulation\n\nThe most impactful improvement was simulating test-time post-processing during training:\n```\ndef apply_random_test_ops(img, num_ops=None):\n    ops = [\n        op_jpeg,               # JPEG compression, quality=30-95\n        op_webp,               # WebP compression, quality=30-95\n        op_random_central_crop, # center crop, scale=0.6-0.95\n        op_resize,             # resize, scale=0.4-1.5\n        op_small_rotation,     # rotation ±10° with aspect-preserving crop\n        op_contrast_brightness, # contrast/brightness, factor=0.6-1.6\n        op_gaussian_blur,      # Gaussian blur, radius=0.5-2.5\n        op_grayscale,          # grayscale → RGB\n    ]\n    num_ops = random.randint(1, 3)\n    chosen  = random.sample(ops, num_ops)\n    for op in chosen:\n        img = op(img)\n    return img\n\n# Applied with probability 0.7 during training\nif random.random() < 0.7:\n    img = apply_random_test_ops(img)\n```\n\n# Multi-Stage Pseudo-Labeling\nWe found that standard fine-tuning with low pseudo-label weight gave minimal improvement. The most effective configuration was training from scratch with full pseudo-label weight.\n```\nStage 1: Train EfficientNetV2-L on clean data + test ops simulation\n         → Public LB: 0.996666\n\nStage 2: Generate pseudo-labels from Stage 1 models (threshold=0.90)\n         Train EfficientNetV2-L from scratch with pseudo-labels\n         → Public LB: 0.998000\n\nStage 3: Generate pseudo-labels from Stage 2 models (threshold=0.80)  \n         Train EfficientNetV2-XL (512px) from scratch with pseudo-labels\n         → Private LB: 0.997333\n```\nKey parameters:\n```\nthreshold     = 0.80-0.90  # confidence threshold\npseudo_weight = 1.0         # full trust in pseudo-labels\nepochs        = 20          # full training from scratch\nlr            = 3e-4        # full learning rate (not fine-tuning lr)\n```\nWhy training from scratch matters: fine-tuning with weight=0.3 gave negligible improvement because the pseudo-label signal was too weak relative to the pretrained weights. Training from scratch with weight=1.0 forces the model to fully incorporate real test-set artifacts (JPEG AI, Super Resolution) that cannot be simulated.\nWhy it works: test images contain real JPEG AI and AI Super Resolution artifacts. Pseudo-labels provide the only way to expose the model to these real artifacts during training.\nPseudo-labeled images (already post-processed) do NOT receive additional augmentation.\n\n# TTA\nOnly horizontal flip was beneficial. All other TTA variants hurt performance:\n\n|✅ Horizontal flip | neutral to slightly positive |\n| --- | --- |\n| ❌ Multiscale TTA | -0.001 LB |\n|❌ Test ops TTA | -0.003 LB |\n|❌ Grayscale/blur TTA | -0.003 LB |\n\nInteresting observation: multiscale TTA (256/384/480px) was beneficial in early experiments but became harmful after correcting the rotation augmentation. Our hypothesis is that the incorrect rotation (with black corners) inadvertently regularized the model to be more robust to input variations, making multiscale TTA helpful. After fixing rotation, the model became more precise but also more sensitive to input size changes.\n\n# What Didn't Work\n\nSeveral approaches were investigated but did not improve leaderboard performance.\n\n## Specialized Pairwise Classifiers\n\nError analysis revealed that most remaining mistakes were concentrated in two confusion pairs:\n\n- SD3 ↔ SD3.5\n- Pixart ↔ Hunyuan\n\nSpecialized binary classifiers were trained on top of the EfficientNetV2-L backbone.\n\n```text\n❌ Binary classifier (SD3 vs SD3.5)\n   OOF accuracy: 0.9993\n   LB: 0.993\n\n❌ Binary classifier (Pixart vs Hunyuan)\n   OOF accuracy: 0.9993\n   LB: 0.993\n```\n\nDespite near-perfect validation accuracy, both models reduced leaderboard performance, suggesting severe overfitting to clean training images.\n\n## Frequency-Domain Features\n\n```text\n❌ FFT magnitude spectrum\n❌ Wavelet decomposition (db4, 3 levels)\n❌ FFT/Wavelet features added to the model\n```\n\nFrequency-domain representations were evaluated both as standalone features and as additional inputs to the neural network. None of the investigated approaches improved validation accuracy or leaderboard scores.\n\n## Alternative Architectures\n\n```text\n❌ ConvNeXt\n❌ Vision Transformer (ViT)\n```\n\nUsing the same training pipeline and hyperparameters, both architectures underperformed EfficientNetV2-L and were not explored further.\n\n## Other Negative Results\n\n```text\n❌ EfficientNetV2-XL without pseudo-labeling\n❌ Multiscale TTA after fixing rotation augmentation\n❌ Pseudo-labeling with weight = 0.3\n```\n\nA consistent pattern throughout the competition was that methods improving clean validation metrics often failed to improve leaderboard performance. In contrast, approaches explicitly targeting the train-test distribution gap consistently produced the largest gains.\n\n# Error Analysis\n\nTo better approximate the hidden test distribution, validation folds were evaluated under simulated test-time post-processing and predictions were averaged over 5 stochastic runs.\n\n<p align=\"center\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F32676712%2F91499ce9bb97124af4e6a224d36952c8%2Fconfusion_matrix_aug_oof.png?generation=1780538086145976&alt=media\" width=\"85%\" style=\"border-radius: 18px;\">\n</p>\n\nThe dominant confusion pairs were:\n\n| True class | Predicted class | Count | Error rate |\n|---|---:|---:|---:|\n| SD3.5 | SD3 | 5 | 0.0071 |\n| Pixart | Hunyuan | 4 | 0.0057 |\n| SD3 | SD3.5 | 3 | 0.0043 |\n| Hunyuan | Pixart | 2 | 0.0029 |\n| SDXL-Turbo | Photon | 1 | 0.0014 |\n\nMost errors came from SD3 ↔ SD3.5 and Pixart ↔ Hunyuan. This matches the intuition that attribution becomes hardest when generators share similar architectures or visual statistics.\n\n# Final Takeaways\n\nThe most important lesson from this competition was that train-test distribution alignment mattered more than architecture scaling or handcrafted features.\n\nSimulating post-processing operations during training produced the largest single improvement, while pseudo-labeling provided exposure to real artifacts that could not be reproduced synthetically.\n\nMany techniques that improved clean validation accuracy failed to improve leaderboard performance, highlighting the importance of matching the hidden test distribution rather than optimizing for clean validation metrics alone."
  }
}