{
  "id": 704738,
  "title": "🏆 9th Place Solution - Synthetic Image Attribution",
  "url": "/competitions/dlmmdd-workshop-synthetic-source-attribution-challenge/discussion/704738",
  "author_name": "nicole quilang",
  "post_date": "2026-06-05T21:35:58.638000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Thank You!</h2>\n<p>First, I would like to thank the organizers for creating such an enjoyable and educational competition.</p>\n<p>At the beginning, I approached this challenge as a standard image classification problem. My initial assumption was that selecting a strong architecture and training it well would be enough to achieve competitive performance.</p>\n<p>However, after studying the competition setup more closely, it became clear that the real challenge was not simply recognizing generator-specific patterns. The challenge was recognizing those patterns after the images had been compressed, resized, cropped, blurred, rotated, or otherwise modified.</p>\n<p>This changed how I approached the competition.</p>\n<p>Instead of focusing primarily on model architecture, I focused on making the training data resemble the hidden test conditions as closely as possible.</p>\n<p>That decision ultimately became the foundation of my entire solution.</p>\n<hr>\n<h1>Summary</h1>\n<p>My final solution achieved:</p>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public LB</td>\n<td>0.987333</td>\n</tr>\n<tr>\n<td>Private LB</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td>Final Rank</td>\n<td>9 / 92</td>\n</tr>\n</tbody>\n</table>\n<h2>The solution combined:</h2>\n<ul>\n<li>EfficientNet-B4</li>\n<li>ConvNeXt-Tiny</li>\n<li>Post-processing simulation</li>\n<li>MixUp augmentation</li>\n<li>Label smoothing</li>\n<li>Horizontal Flip TTA</li>\n<li>Weighted soft-voting</li>\n</ul>\n<p>One practical constraint throughout the competition was hardware availability. Rather than relying on extremely large models, I focused on maximizing performance through robust training strategies, realistic post-processing simulation, and complementary CNN architectures.</p>\n<p>This ultimately proved more valuable than simply increasing model size and allowed me to achieve a Top 10 finish using computationally efficient backbones.</p>\n<hr>\n<h1>Solution Pipeline</h1>\n<pre><code>Training Images\n        ↓\nPost-Processing Simulation\n        ↓\nEfficientNet-B4 Training\n        ↓\nConvNeXt-Tiny Training\n        ↓\nHorizontal Flip TTA\n        ↓\nWeighted Soft Voting\n        ↓\nFinal Predictions\n</code></pre>\n<h1>Understanding the Challenge</h1>\n<p>One observation quickly became apparent:</p>\n<p>The training images were relatively clean, but the hidden test images were intentionally degraded through various post-processing operations.</p>\n<p>These included:</p>\n<ul>\n<li>JPEG compression</li>\n<li>WebP compression</li>\n<li>Cropping</li>\n<li>Resizing</li>\n<li>Rotation</li>\n<li>Brightness and contrast adjustment</li>\n<li>Gaussian blur</li>\n<li>Grayscale conversion</li>\n<li>Super-resolution</li>\n</ul>\n<p>A model trained exclusively on clean images could easily learn shortcuts that disappear after these transformations.</p>\n<p>Instead of learning fragile pixel-level patterns, the model needed to learn generator-specific characteristics that remained stable after repeated image manipulations.</p>\n<p>This realization ultimately shaped the entire training pipeline.</p>\n<h1>EfficientNet-B4 Baseline</h1>\n<p>My first competitive model was <strong>EfficientNet-B4</strong>.</p>\n<p>The model was trained using:</p>\n<ul>\n<li>Input size: 384×384</li>\n<li>AdamW</li>\n<li>Label smoothing (0.1)</li>\n<li>MixUp (α = 0.2)</li>\n<li>Cosine Annealing Warm Restarts</li>\n</ul>\n<h3>Results</h3>\n<ul>\n<li>Validation Accuracy: <strong>98.36%</strong></li>\n<li>Test Accuracy: <strong>98.07%</strong></li>\n<li>Public LB: <strong>0.981333</strong></li>\n<li>Private LB: <strong>0.979333</strong></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F127ff9aa2a2f75f73c6d2f2a0a29296c%2FEFFCIENT%20TRAIN%20AND%20VAL%20LOS.png?generation=1780696288263203&amp;alt=media\" alt=\"\"></p>\n<p>One encouraging sign was the small gap between validation and testing performance. This suggested that the augmentation strategy was helping the model generalize rather than simply memorize training samples.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F531da2ac95cec084a3db2b593db00201%2FEFFICIENT%20NET.png?generation=1780696299175320&amp;alt=media\" alt=\"\"></p>\n<p>Most mistakes occurred between <strong>Stable Diffusion 3</strong> and <strong>Stable Diffusion 3.5</strong>, which was not particularly surprising given their architectural similarity and overlapping visual characteristics.</p>\n<h1>Why I Added ConvNeXt</h1>\n<p>Although EfficientNet-B4 was already performing well, I wanted to determine whether a different architecture could learn complementary representations from the same data.</p>\n<p>Rather than changing the augmentation pipeline, I kept everything identical and replaced only the backbone.</p>\n<p>This allowed me to isolate the effect of architecture.</p>\n<h3>ConvNeXt-Tiny Results</h3>\n<ul>\n<li>Validation Accuracy: <strong>99.21%</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F262846bbcd71790c68610f215cb6bbef%2Fcovnextt.png?generation=1780696332534394&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>Interestingly, ConvNeXt-Tiny achieved a higher validation accuracy than EfficientNet-B4 despite using the exact same preprocessing and augmentation pipeline.</p>\n<p>This suggested that the training strategy generalized effectively across architectures and that ConvNeXt was capturing information not fully represented by EfficientNet-B4.</p>\n<h1>What Contributed Most to Performance?</h1>\n<p>Throughout the competition, I experimented with different architectures and ensemble configurations.</p>\n<p>One of the most important observations was that the largest improvements did not come from increasing model complexity.</p>\n<p>They came from improving robustness.</p>\n<h2>1. Post-Processing Simulation (Largest Contribution)</h2>\n<p>The biggest breakthrough came when I stopped treating the problem as a standard image classification task and started treating it as a robustness problem.</p>\n<p>The hidden test images were intentionally modified through compression, resizing, blurring, and other transformations. As a result, models trained only on clean images risked learning shortcuts that would disappear during evaluation.</p>\n<p>To bridge this gap, I created a post-processing simulator that exposed the models to realistic image degradations throughout training:</p>\n<pre><code>operations = [\n    ('jpeg', lambda x: self.apply_jpeg_compression(x)),\n    ('webp', lambda x: self.apply_webp_compression(x)),\n    ('crop', lambda x: self.apply_random_crop(x)),\n    ('resize', lambda x: self.apply_resizing(x)),\n    ('rotate', lambda x: self.apply_rotation_crop(x)),\n    ('blur', lambda x: self.apply_blur(x)),\n    ('brightness', lambda x: self.apply_brightness_contrast(x)),\n    ('grayscale', lambda x: self.apply_grayscale(x)),\n    ('superres', lambda x: self.apply_super_resolution(x)),\n]\n\nnum_ops = np.random.randint(1, 3)\nselected = np.random.choice(len(operations), num_ops, replace=False)\n\nfor idx in selected:\n    _, op_func = operations[idx]\n    img = op_func(img)\n</code></pre>\n<p>By repeatedly exposing the models to transformed versions of the same image, they were forced to learn generator-specific fingerprints that remained stable across image degradations.</p>\n<p>Looking back, this ended up being the most important part of the solution.</p>\n<h2>2. MixUp and Label Smoothing</h2>\n<p>To further improve generalization, I combined MixUp augmentation and label smoothing.</p>\n<h3>MixUp</h3>\n<pre><code>def mixup_data(x, y, alpha=0.2):\n    lam = np.random.beta(alpha, alpha)\n\n    index = torch.randperm(x.size(0)).to(x.device)\n\n    mixed_x = lam * x + (1 - lam) * x[index]\n\n    y_a, y_b = y, y[index]\n\n    return mixed_x, y_a, y_b, lam\n</code></pre>\n<h3>Label Smoothing</h3>\n<pre><code>criterion = nn.CrossEntropyLoss(\n    label_smoothing=0.1\n)\n</code></pre>\n<p>These techniques improved generalization, reduced overconfidence, and helped the models distinguish visually similar generators.</p>\n<h2>3. Test-Time Augmentation (TTA)</h2>\n<p>During inference, I used horizontal flip test-time augmentation.</p>\n<p>Predictions from the original image and its horizontally flipped version were averaged before generating the final prediction.</p>\n<pre><code>def predict_with_tta(model, images):\n    outputs = model(images)\n\n    outputs_flip = model(\n        torch.flip(images, dims=[3])\n    )\n\n    return (outputs + outputs_flip) / 2\n</code></pre>\n<p>While the improvement was modest compared to post-processing simulation, TTA consistently produced more stable predictions and provided a useful boost to the final ensemble.</p>\n<h1>Ensemble Strategy</h1>\n<p>After establishing strong individual models, I explored combining EfficientNet-B4 and ConvNeXt-Tiny through weighted soft-voting.</p>\n<p>Although both models were CNN-based architectures, they learned different representations of the data and produced slightly different prediction patterns.</p>\n<p>Several weight combinations were evaluated:</p>\n<table>\n<thead>\n<tr>\n<th>EfficientNet-B4</th>\n<th>ConvNeXt-Tiny</th>\n<th>Validation Accuracy</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.7</td>\n<td>0.3</td>\n<td>99.07%</td>\n</tr>\n<tr>\n<td>0.6</td>\n<td>0.4</td>\n<td>99.21%</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.5</td>\n<td>99.29%</td>\n</tr>\n<tr>\n<td>0.4</td>\n<td>0.6</td>\n<td>99.36%</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>0.7</td>\n<td>99.36%</td>\n</tr>\n</tbody>\n</table>\n<p>The final ensemble used:</p>\n<ul>\n<li>EfficientNet-B4: 30%</li>\n<li>ConvNeXt-Tiny: 70%</li>\n</ul>\n<pre><code>ensemble_out = (\n    0.30 * efficientnet_output +\n    0.70 * convnext_output\n)\n</code></pre>\n<h3>Final Ensemble Results</h3>\n<ul>\n<li>Validation Accuracy: <strong>99.36%</strong></li>\n<li>Public LB: <strong>0.987333</strong></li>\n<li>Private LB: <strong>0.990000</strong></li>\n</ul>\n<p>The ensemble consistently outperformed both individual models and proved more stable across validation and leaderboard evaluations.</p>\n<h1>Leaderboard Progression</h1>\n<table>\n<thead>\n<tr>\n<th>Version</th>\n<th>Model</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>V4</td>\n<td>EfficientNet-B4</td>\n<td>0.981333</td>\n<td>0.979333</td>\n</tr>\n<tr>\n<td>V5</td>\n<td>Initial Ensemble</td>\n<td>0.982666</td>\n<td>0.988000</td>\n</tr>\n<tr>\n<td>V6</td>\n<td>EfficientNet-B4 + ConvNeXt-Tiny</td>\n<td>0.987333</td>\n<td>0.990000</td>\n</tr>\n</tbody>\n</table>\n<p>Each iteration built upon the previous one, with the final ensemble delivering the strongest overall performance.</p>\n<h1>Key Observations</h1>\n<p>Several observations stood out during the competition.</p>\n<h3>Post-processing simulation was the biggest contributor</h3>\n<p>The largest gains came from exposing the models to realistic image transformations during training rather than increasing model complexity.</p>\n<p>Training on compressed, resized, cropped, blurred, and grayscale images forced the models to learn generator-specific fingerprints that remained detectable even after heavy image manipulation.</p>\n<h3>Most errors were concentrated in a few generator pairs</h3>\n<p>The confusion matrix showed that errors were not evenly distributed across all classes.</p>\n<p>Most misclassifications occurred between Stable Diffusion 3 and Stable Diffusion 3.5, while several classes achieved near-perfect or perfect classification accuracy.</p>\n<p>This suggests that the model successfully learned distinctive generator fingerprints but still struggled when generators shared highly similar architectural characteristics.</p>\n<h3>ConvNeXt and EfficientNet learned different patterns</h3>\n<p>ConvNeXt-Tiny achieved higher standalone validation accuracy, but EfficientNet-B4 continued to contribute useful information during ensembling.</p>\n<p>The improvement from V4 (0.979333 Private LB) to V6 (0.990000 Private LB) suggests that the two models were making different mistakes and provided complementary predictions.</p>\n<h3>Simple inference techniques were enough</h3>\n<p>Horizontal flip TTA consistently improved prediction stability without introducing significant computational overhead.</p>\n<p>More importantly, most of the performance gains came from the training pipeline itself rather than complex inference tricks.</p>\n<h3>Efficient architectures remained competitive</h3>\n<p>A major goal throughout the competition was achieving strong performance without relying on extremely large models.</p>\n<p>EfficientNet-B4 and ConvNeXt-Tiny provided an effective balance between computational cost and accuracy, ultimately achieving a Top 10 finish while remaining practical to train on limited hardware.</p>\n<h1>Final Takeaways</h1>\n<p>The most valuable lesson from this competition was that strong solutions are not always the result of larger models.</p>\n<p>In this challenge, the largest improvements came from understanding the evaluation conditions and designing a training pipeline that reflected those conditions as closely as possible.</p>\n<p>Post-processing simulation proved far more impactful than increasing model complexity, while ConvNeXt-Tiny and EfficientNet-B4 demonstrated that efficient CNN architectures can still achieve highly competitive results when paired with robust augmentation strategies.</p>\n<p>By combining realistic image transformations, MixUp, label smoothing, horizontal-flip TTA, and weighted soft-voting, I was able to improve from a <strong>0.979333 Private Leaderboard baseline</strong> to a final score of <strong>0.990000</strong>, securing <strong>9th place out of 92 teams</strong>.</p>\n<p>This competition was a great reminder that understanding the data distribution often matters more than simply scaling the model.</p>\n<p>Thank you again to the organizers for creating such a challenging and educational competition.</p>",
  "messages": [
    {
      "id": 3467310,
      "postDate": "2026-06-05T21:35:58.640Z",
      "content": "<h2>Thank You!</h2>\n<p>First, I would like to thank the organizers for creating such an enjoyable and educational competition.</p>\n<p>At the beginning, I approached this challenge as a standard image classification problem. My initial assumption was that selecting a strong architecture and training it well would be enough to achieve competitive performance.</p>\n<p>However, after studying the competition setup more closely, it became clear that the real challenge was not simply recognizing generator-specific patterns. The challenge was recognizing those patterns after the images had been compressed, resized, cropped, blurred, rotated, or otherwise modified.</p>\n<p>This changed how I approached the competition.</p>\n<p>Instead of focusing primarily on model architecture, I focused on making the training data resemble the hidden test conditions as closely as possible.</p>\n<p>That decision ultimately became the foundation of my entire solution.</p>\n<hr>\n<h1>Summary</h1>\n<p>My final solution achieved:</p>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public LB</td>\n<td>0.987333</td>\n</tr>\n<tr>\n<td>Private LB</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td>Final Rank</td>\n<td>9 / 92</td>\n</tr>\n</tbody>\n</table>\n<h2>The solution combined:</h2>\n<ul>\n<li>EfficientNet-B4</li>\n<li>ConvNeXt-Tiny</li>\n<li>Post-processing simulation</li>\n<li>MixUp augmentation</li>\n<li>Label smoothing</li>\n<li>Horizontal Flip TTA</li>\n<li>Weighted soft-voting</li>\n</ul>\n<p>One practical constraint throughout the competition was hardware availability. Rather than relying on extremely large models, I focused on maximizing performance through robust training strategies, realistic post-processing simulation, and complementary CNN architectures.</p>\n<p>This ultimately proved more valuable than simply increasing model size and allowed me to achieve a Top 10 finish using computationally efficient backbones.</p>\n<hr>\n<h1>Solution Pipeline</h1>\n<pre><code>Training Images\n        ↓\nPost-Processing Simulation\n        ↓\nEfficientNet-B4 Training\n        ↓\nConvNeXt-Tiny Training\n        ↓\nHorizontal Flip TTA\n        ↓\nWeighted Soft Voting\n        ↓\nFinal Predictions\n</code></pre>\n<h1>Understanding the Challenge</h1>\n<p>One observation quickly became apparent:</p>\n<p>The training images were relatively clean, but the hidden test images were intentionally degraded through various post-processing operations.</p>\n<p>These included:</p>\n<ul>\n<li>JPEG compression</li>\n<li>WebP compression</li>\n<li>Cropping</li>\n<li>Resizing</li>\n<li>Rotation</li>\n<li>Brightness and contrast adjustment</li>\n<li>Gaussian blur</li>\n<li>Grayscale conversion</li>\n<li>Super-resolution</li>\n</ul>\n<p>A model trained exclusively on clean images could easily learn shortcuts that disappear after these transformations.</p>\n<p>Instead of learning fragile pixel-level patterns, the model needed to learn generator-specific characteristics that remained stable after repeated image manipulations.</p>\n<p>This realization ultimately shaped the entire training pipeline.</p>\n<h1>EfficientNet-B4 Baseline</h1>\n<p>My first competitive model was <strong>EfficientNet-B4</strong>.</p>\n<p>The model was trained using:</p>\n<ul>\n<li>Input size: 384×384</li>\n<li>AdamW</li>\n<li>Label smoothing (0.1)</li>\n<li>MixUp (α = 0.2)</li>\n<li>Cosine Annealing Warm Restarts</li>\n</ul>\n<h3>Results</h3>\n<ul>\n<li>Validation Accuracy: <strong>98.36%</strong></li>\n<li>Test Accuracy: <strong>98.07%</strong></li>\n<li>Public LB: <strong>0.981333</strong></li>\n<li>Private LB: <strong>0.979333</strong></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F127ff9aa2a2f75f73c6d2f2a0a29296c%2FEFFCIENT%20TRAIN%20AND%20VAL%20LOS.png?generation=1780696288263203&amp;alt=media\" alt=\"\"></p>\n<p>One encouraging sign was the small gap between validation and testing performance. This suggested that the augmentation strategy was helping the model generalize rather than simply memorize training samples.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F531da2ac95cec084a3db2b593db00201%2FEFFICIENT%20NET.png?generation=1780696299175320&amp;alt=media\" alt=\"\"></p>\n<p>Most mistakes occurred between <strong>Stable Diffusion 3</strong> and <strong>Stable Diffusion 3.5</strong>, which was not particularly surprising given their architectural similarity and overlapping visual characteristics.</p>\n<h1>Why I Added ConvNeXt</h1>\n<p>Although EfficientNet-B4 was already performing well, I wanted to determine whether a different architecture could learn complementary representations from the same data.</p>\n<p>Rather than changing the augmentation pipeline, I kept everything identical and replaced only the backbone.</p>\n<p>This allowed me to isolate the effect of architecture.</p>\n<h3>ConvNeXt-Tiny Results</h3>\n<ul>\n<li>Validation Accuracy: <strong>99.21%</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F262846bbcd71790c68610f215cb6bbef%2Fcovnextt.png?generation=1780696332534394&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>Interestingly, ConvNeXt-Tiny achieved a higher validation accuracy than EfficientNet-B4 despite using the exact same preprocessing and augmentation pipeline.</p>\n<p>This suggested that the training strategy generalized effectively across architectures and that ConvNeXt was capturing information not fully represented by EfficientNet-B4.</p>\n<h1>What Contributed Most to Performance?</h1>\n<p>Throughout the competition, I experimented with different architectures and ensemble configurations.</p>\n<p>One of the most important observations was that the largest improvements did not come from increasing model complexity.</p>\n<p>They came from improving robustness.</p>\n<h2>1. Post-Processing Simulation (Largest Contribution)</h2>\n<p>The biggest breakthrough came when I stopped treating the problem as a standard image classification task and started treating it as a robustness problem.</p>\n<p>The hidden test images were intentionally modified through compression, resizing, blurring, and other transformations. As a result, models trained only on clean images risked learning shortcuts that would disappear during evaluation.</p>\n<p>To bridge this gap, I created a post-processing simulator that exposed the models to realistic image degradations throughout training:</p>\n<pre><code>operations = [\n    ('jpeg', lambda x: self.apply_jpeg_compression(x)),\n    ('webp', lambda x: self.apply_webp_compression(x)),\n    ('crop', lambda x: self.apply_random_crop(x)),\n    ('resize', lambda x: self.apply_resizing(x)),\n    ('rotate', lambda x: self.apply_rotation_crop(x)),\n    ('blur', lambda x: self.apply_blur(x)),\n    ('brightness', lambda x: self.apply_brightness_contrast(x)),\n    ('grayscale', lambda x: self.apply_grayscale(x)),\n    ('superres', lambda x: self.apply_super_resolution(x)),\n]\n\nnum_ops = np.random.randint(1, 3)\nselected = np.random.choice(len(operations), num_ops, replace=False)\n\nfor idx in selected:\n    _, op_func = operations[idx]\n    img = op_func(img)\n</code></pre>\n<p>By repeatedly exposing the models to transformed versions of the same image, they were forced to learn generator-specific fingerprints that remained stable across image degradations.</p>\n<p>Looking back, this ended up being the most important part of the solution.</p>\n<h2>2. MixUp and Label Smoothing</h2>\n<p>To further improve generalization, I combined MixUp augmentation and label smoothing.</p>\n<h3>MixUp</h3>\n<pre><code>def mixup_data(x, y, alpha=0.2):\n    lam = np.random.beta(alpha, alpha)\n\n    index = torch.randperm(x.size(0)).to(x.device)\n\n    mixed_x = lam * x + (1 - lam) * x[index]\n\n    y_a, y_b = y, y[index]\n\n    return mixed_x, y_a, y_b, lam\n</code></pre>\n<h3>Label Smoothing</h3>\n<pre><code>criterion = nn.CrossEntropyLoss(\n    label_smoothing=0.1\n)\n</code></pre>\n<p>These techniques improved generalization, reduced overconfidence, and helped the models distinguish visually similar generators.</p>\n<h2>3. Test-Time Augmentation (TTA)</h2>\n<p>During inference, I used horizontal flip test-time augmentation.</p>\n<p>Predictions from the original image and its horizontally flipped version were averaged before generating the final prediction.</p>\n<pre><code>def predict_with_tta(model, images):\n    outputs = model(images)\n\n    outputs_flip = model(\n        torch.flip(images, dims=[3])\n    )\n\n    return (outputs + outputs_flip) / 2\n</code></pre>\n<p>While the improvement was modest compared to post-processing simulation, TTA consistently produced more stable predictions and provided a useful boost to the final ensemble.</p>\n<h1>Ensemble Strategy</h1>\n<p>After establishing strong individual models, I explored combining EfficientNet-B4 and ConvNeXt-Tiny through weighted soft-voting.</p>\n<p>Although both models were CNN-based architectures, they learned different representations of the data and produced slightly different prediction patterns.</p>\n<p>Several weight combinations were evaluated:</p>\n<table>\n<thead>\n<tr>\n<th>EfficientNet-B4</th>\n<th>ConvNeXt-Tiny</th>\n<th>Validation Accuracy</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.7</td>\n<td>0.3</td>\n<td>99.07%</td>\n</tr>\n<tr>\n<td>0.6</td>\n<td>0.4</td>\n<td>99.21%</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.5</td>\n<td>99.29%</td>\n</tr>\n<tr>\n<td>0.4</td>\n<td>0.6</td>\n<td>99.36%</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>0.7</td>\n<td>99.36%</td>\n</tr>\n</tbody>\n</table>\n<p>The final ensemble used:</p>\n<ul>\n<li>EfficientNet-B4: 30%</li>\n<li>ConvNeXt-Tiny: 70%</li>\n</ul>\n<pre><code>ensemble_out = (\n    0.30 * efficientnet_output +\n    0.70 * convnext_output\n)\n</code></pre>\n<h3>Final Ensemble Results</h3>\n<ul>\n<li>Validation Accuracy: <strong>99.36%</strong></li>\n<li>Public LB: <strong>0.987333</strong></li>\n<li>Private LB: <strong>0.990000</strong></li>\n</ul>\n<p>The ensemble consistently outperformed both individual models and proved more stable across validation and leaderboard evaluations.</p>\n<h1>Leaderboard Progression</h1>\n<table>\n<thead>\n<tr>\n<th>Version</th>\n<th>Model</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>V4</td>\n<td>EfficientNet-B4</td>\n<td>0.981333</td>\n<td>0.979333</td>\n</tr>\n<tr>\n<td>V5</td>\n<td>Initial Ensemble</td>\n<td>0.982666</td>\n<td>0.988000</td>\n</tr>\n<tr>\n<td>V6</td>\n<td>EfficientNet-B4 + ConvNeXt-Tiny</td>\n<td>0.987333</td>\n<td>0.990000</td>\n</tr>\n</tbody>\n</table>\n<p>Each iteration built upon the previous one, with the final ensemble delivering the strongest overall performance.</p>\n<h1>Key Observations</h1>\n<p>Several observations stood out during the competition.</p>\n<h3>Post-processing simulation was the biggest contributor</h3>\n<p>The largest gains came from exposing the models to realistic image transformations during training rather than increasing model complexity.</p>\n<p>Training on compressed, resized, cropped, blurred, and grayscale images forced the models to learn generator-specific fingerprints that remained detectable even after heavy image manipulation.</p>\n<h3>Most errors were concentrated in a few generator pairs</h3>\n<p>The confusion matrix showed that errors were not evenly distributed across all classes.</p>\n<p>Most misclassifications occurred between Stable Diffusion 3 and Stable Diffusion 3.5, while several classes achieved near-perfect or perfect classification accuracy.</p>\n<p>This suggests that the model successfully learned distinctive generator fingerprints but still struggled when generators shared highly similar architectural characteristics.</p>\n<h3>ConvNeXt and EfficientNet learned different patterns</h3>\n<p>ConvNeXt-Tiny achieved higher standalone validation accuracy, but EfficientNet-B4 continued to contribute useful information during ensembling.</p>\n<p>The improvement from V4 (0.979333 Private LB) to V6 (0.990000 Private LB) suggests that the two models were making different mistakes and provided complementary predictions.</p>\n<h3>Simple inference techniques were enough</h3>\n<p>Horizontal flip TTA consistently improved prediction stability without introducing significant computational overhead.</p>\n<p>More importantly, most of the performance gains came from the training pipeline itself rather than complex inference tricks.</p>\n<h3>Efficient architectures remained competitive</h3>\n<p>A major goal throughout the competition was achieving strong performance without relying on extremely large models.</p>\n<p>EfficientNet-B4 and ConvNeXt-Tiny provided an effective balance between computational cost and accuracy, ultimately achieving a Top 10 finish while remaining practical to train on limited hardware.</p>\n<h1>Final Takeaways</h1>\n<p>The most valuable lesson from this competition was that strong solutions are not always the result of larger models.</p>\n<p>In this challenge, the largest improvements came from understanding the evaluation conditions and designing a training pipeline that reflected those conditions as closely as possible.</p>\n<p>Post-processing simulation proved far more impactful than increasing model complexity, while ConvNeXt-Tiny and EfficientNet-B4 demonstrated that efficient CNN architectures can still achieve highly competitive results when paired with robust augmentation strategies.</p>\n<p>By combining realistic image transformations, MixUp, label smoothing, horizontal-flip TTA, and weighted soft-voting, I was able to improve from a <strong>0.979333 Private Leaderboard baseline</strong> to a final score of <strong>0.990000</strong>, securing <strong>9th place out of 92 teams</strong>.</p>\n<p>This competition was a great reminder that understanding the data distribution often matters more than simply scaling the model.</p>\n<p>Thank you again to the organizers for creating such a challenging and educational competition.</p>",
      "rawMarkdown": "## Thank You!\n\nFirst, I would like to thank the organizers for creating such an enjoyable and educational competition.\n\nAt the beginning, I approached this challenge as a standard image classification problem. My initial assumption was that selecting a strong architecture and training it well would be enough to achieve competitive performance.\n\nHowever, after studying the competition setup more closely, it became clear that the real challenge was not simply recognizing generator-specific patterns. The challenge was recognizing those patterns after the images had been compressed, resized, cropped, blurred, rotated, or otherwise modified.\n\nThis changed how I approached the competition.\n\nInstead of focusing primarily on model architecture, I focused on making the training data resemble the hidden test conditions as closely as possible.\n\nThat decision ultimately became the foundation of my entire solution.\n\n---\n\n# Summary\n\nMy final solution achieved:\n\n| Metric     | Score    |\n| ---------- | -------- |\n| Public LB  | 0.987333 |\n| Private LB | 0.990000 |\n| Final Rank | 9 / 92   |\n\n## The solution combined:\n\n* EfficientNet-B4\n* ConvNeXt-Tiny\n* Post-processing simulation\n* MixUp augmentation\n* Label smoothing\n* Horizontal Flip TTA\n* Weighted soft-voting\n\nOne practical constraint throughout the competition was hardware availability. Rather than relying on extremely large models, I focused on maximizing performance through robust training strategies, realistic post-processing simulation, and complementary CNN architectures.\n\nThis ultimately proved more valuable than simply increasing model size and allowed me to achieve a Top 10 finish using computationally efficient backbones.\n\n---\n\n# Solution Pipeline\n\n```text\nTraining Images\n        ↓\nPost-Processing Simulation\n        ↓\nEfficientNet-B4 Training\n        ↓\nConvNeXt-Tiny Training\n        ↓\nHorizontal Flip TTA\n        ↓\nWeighted Soft Voting\n        ↓\nFinal Predictions\n```\n\n# Understanding the Challenge\n\nOne observation quickly became apparent:\n\nThe training images were relatively clean, but the hidden test images were intentionally degraded through various post-processing operations.\n\nThese included:\n\n* JPEG compression\n* WebP compression\n* Cropping\n* Resizing\n* Rotation\n* Brightness and contrast adjustment\n* Gaussian blur\n* Grayscale conversion\n* Super-resolution\n\nA model trained exclusively on clean images could easily learn shortcuts that disappear after these transformations.\n\nInstead of learning fragile pixel-level patterns, the model needed to learn generator-specific characteristics that remained stable after repeated image manipulations.\n\nThis realization ultimately shaped the entire training pipeline.\n\n# EfficientNet-B4 Baseline\n\nMy first competitive model was **EfficientNet-B4**.\n\nThe model was trained using:\n\n* Input size: 384×384\n* AdamW\n* Label smoothing (0.1)\n* MixUp (α = 0.2)\n* Cosine Annealing Warm Restarts\n\n### Results\n\n* Validation Accuracy: **98.36%**\n* Test Accuracy: **98.07%**\n* Public LB: **0.981333**\n* Private LB: **0.979333**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F127ff9aa2a2f75f73c6d2f2a0a29296c%2FEFFCIENT%20TRAIN%20AND%20VAL%20LOS.png?generation=1780696288263203&alt=media)\n\nOne encouraging sign was the small gap between validation and testing performance. This suggested that the augmentation strategy was helping the model generalize rather than simply memorize training samples.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F531da2ac95cec084a3db2b593db00201%2FEFFICIENT%20NET.png?generation=1780696299175320&alt=media)\n\nMost mistakes occurred between **Stable Diffusion 3** and **Stable Diffusion 3.5**, which was not particularly surprising given their architectural similarity and overlapping visual characteristics.\n\n\n# Why I Added ConvNeXt\n\nAlthough EfficientNet-B4 was already performing well, I wanted to determine whether a different architecture could learn complementary representations from the same data.\n\nRather than changing the augmentation pipeline, I kept everything identical and replaced only the backbone.\n\nThis allowed me to isolate the effect of architecture.\n\n### ConvNeXt-Tiny Results\n\n* Validation Accuracy: **99.21%**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F262846bbcd71790c68610f215cb6bbef%2Fcovnextt.png?generation=1780696332534394&alt=media)\n\nInterestingly, ConvNeXt-Tiny achieved a higher validation accuracy than EfficientNet-B4 despite using the exact same preprocessing and augmentation pipeline.\n\nThis suggested that the training strategy generalized effectively across architectures and that ConvNeXt was capturing information not fully represented by EfficientNet-B4.\n\n\n\n# What Contributed Most to Performance?\n\nThroughout the competition, I experimented with different architectures and ensemble configurations.\n\nOne of the most important observations was that the largest improvements did not come from increasing model complexity.\n\nThey came from improving robustness.\n\n## 1. Post-Processing Simulation (Largest Contribution)\n\nThe biggest breakthrough came when I stopped treating the problem as a standard image classification task and started treating it as a robustness problem.\n\nThe hidden test images were intentionally modified through compression, resizing, blurring, and other transformations. As a result, models trained only on clean images risked learning shortcuts that would disappear during evaluation.\n\nTo bridge this gap, I created a post-processing simulator that exposed the models to realistic image degradations throughout training:\n\n```python\noperations = [\n    ('jpeg', lambda x: self.apply_jpeg_compression(x)),\n    ('webp', lambda x: self.apply_webp_compression(x)),\n    ('crop', lambda x: self.apply_random_crop(x)),\n    ('resize', lambda x: self.apply_resizing(x)),\n    ('rotate', lambda x: self.apply_rotation_crop(x)),\n    ('blur', lambda x: self.apply_blur(x)),\n    ('brightness', lambda x: self.apply_brightness_contrast(x)),\n    ('grayscale', lambda x: self.apply_grayscale(x)),\n    ('superres', lambda x: self.apply_super_resolution(x)),\n]\n\nnum_ops = np.random.randint(1, 3)\nselected = np.random.choice(len(operations), num_ops, replace=False)\n\nfor idx in selected:\n    _, op_func = operations[idx]\n    img = op_func(img)\n```\n\nBy repeatedly exposing the models to transformed versions of the same image, they were forced to learn generator-specific fingerprints that remained stable across image degradations.\n\nLooking back, this ended up being the most important part of the solution.\n\n## 2. MixUp and Label Smoothing\n\nTo further improve generalization, I combined MixUp augmentation and label smoothing.\n\n### MixUp\n\n```python\ndef mixup_data(x, y, alpha=0.2):\n    lam = np.random.beta(alpha, alpha)\n\n    index = torch.randperm(x.size(0)).to(x.device)\n\n    mixed_x = lam * x + (1 - lam) * x[index]\n\n    y_a, y_b = y, y[index]\n\n    return mixed_x, y_a, y_b, lam\n```\n\n### Label Smoothing\n\n```python\ncriterion = nn.CrossEntropyLoss(\n    label_smoothing=0.1\n)\n```\n\nThese techniques improved generalization, reduced overconfidence, and helped the models distinguish visually similar generators.\n\n\n## 3. Test-Time Augmentation (TTA)\n\nDuring inference, I used horizontal flip test-time augmentation.\n\nPredictions from the original image and its horizontally flipped version were averaged before generating the final prediction.\n\n```python\ndef predict_with_tta(model, images):\n    outputs = model(images)\n\n    outputs_flip = model(\n        torch.flip(images, dims=[3])\n    )\n\n    return (outputs + outputs_flip) / 2\n```\n\nWhile the improvement was modest compared to post-processing simulation, TTA consistently produced more stable predictions and provided a useful boost to the final ensemble.\n\n\n# Ensemble Strategy\n\nAfter establishing strong individual models, I explored combining EfficientNet-B4 and ConvNeXt-Tiny through weighted soft-voting.\n\nAlthough both models were CNN-based architectures, they learned different representations of the data and produced slightly different prediction patterns.\n\nSeveral weight combinations were evaluated:\n\n| EfficientNet-B4 | ConvNeXt-Tiny | Validation Accuracy |\n| --------------- | ------------- | ------------------- |\n| 0.7             | 0.3           | 99.07%              |\n| 0.6             | 0.4           | 99.21%              |\n| 0.5             | 0.5           | 99.29%              |\n| 0.4             | 0.6           | 99.36%              |\n| 0.3             | 0.7           | 99.36%              |\n\nThe final ensemble used:\n\n* EfficientNet-B4: 30%\n* ConvNeXt-Tiny: 70%\n\n```python\nensemble_out = (\n    0.30 * efficientnet_output +\n    0.70 * convnext_output\n)\n```\n\n### Final Ensemble Results\n\n* Validation Accuracy: **99.36%**\n* Public LB: **0.987333**\n* Private LB: **0.990000**\n\nThe ensemble consistently outperformed both individual models and proved more stable across validation and leaderboard evaluations.\n\n\n# Leaderboard Progression\n\n| Version | Model                           | Public LB | Private LB |\n| ------- | ------------------------------- | --------- | ---------- |\n| V4      | EfficientNet-B4                 | 0.981333  | 0.979333   |\n| V5      | Initial Ensemble                | 0.982666  | 0.988000   |\n| V6      | EfficientNet-B4 + ConvNeXt-Tiny | 0.987333  | 0.990000   |\n\nEach iteration built upon the previous one, with the final ensemble delivering the strongest overall performance.\n\n\n# Key Observations\n\nSeveral observations stood out during the competition.\n\n### Post-processing simulation was the biggest contributor\n\nThe largest gains came from exposing the models to realistic image transformations during training rather than increasing model complexity.\n\nTraining on compressed, resized, cropped, blurred, and grayscale images forced the models to learn generator-specific fingerprints that remained detectable even after heavy image manipulation.\n\n### Most errors were concentrated in a few generator pairs\n\nThe confusion matrix showed that errors were not evenly distributed across all classes.\n\nMost misclassifications occurred between Stable Diffusion 3 and Stable Diffusion 3.5, while several classes achieved near-perfect or perfect classification accuracy.\n\nThis suggests that the model successfully learned distinctive generator fingerprints but still struggled when generators shared highly similar architectural characteristics.\n\n### ConvNeXt and EfficientNet learned different patterns\n\nConvNeXt-Tiny achieved higher standalone validation accuracy, but EfficientNet-B4 continued to contribute useful information during ensembling.\n\nThe improvement from V4 (0.979333 Private LB) to V6 (0.990000 Private LB) suggests that the two models were making different mistakes and provided complementary predictions.\n\n### Simple inference techniques were enough\n\nHorizontal flip TTA consistently improved prediction stability without introducing significant computational overhead.\n\nMore importantly, most of the performance gains came from the training pipeline itself rather than complex inference tricks.\n\n### Efficient architectures remained competitive\n\nA major goal throughout the competition was achieving strong performance without relying on extremely large models.\n\nEfficientNet-B4 and ConvNeXt-Tiny provided an effective balance between computational cost and accuracy, ultimately achieving a Top 10 finish while remaining practical to train on limited hardware.\n\n# Final Takeaways\n\nThe most valuable lesson from this competition was that strong solutions are not always the result of larger models.\n\nIn this challenge, the largest improvements came from understanding the evaluation conditions and designing a training pipeline that reflected those conditions as closely as possible.\n\nPost-processing simulation proved far more impactful than increasing model complexity, while ConvNeXt-Tiny and EfficientNet-B4 demonstrated that efficient CNN architectures can still achieve highly competitive results when paired with robust augmentation strategies.\n\nBy combining realistic image transformations, MixUp, label smoothing, horizontal-flip TTA, and weighted soft-voting, I was able to improve from a **0.979333 Private Leaderboard baseline** to a final score of **0.990000**, securing **9th place out of 92 teams**.\n\nThis competition was a great reminder that understanding the data distribution often matters more than simply scaling the model.\n\nThank you again to the organizers for creating such a challenging and educational competition.\n",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3467310": "## Thank You!\n\nFirst, I would like to thank the organizers for creating such an enjoyable and educational competition.\n\nAt the beginning, I approached this challenge as a standard image classification problem. My initial assumption was that selecting a strong architecture and training it well would be enough to achieve competitive performance.\n\nHowever, after studying the competition setup more closely, it became clear that the real challenge was not simply recognizing generator-specific patterns. The challenge was recognizing those patterns after the images had been compressed, resized, cropped, blurred, rotated, or otherwise modified.\n\nThis changed how I approached the competition.\n\nInstead of focusing primarily on model architecture, I focused on making the training data resemble the hidden test conditions as closely as possible.\n\nThat decision ultimately became the foundation of my entire solution.\n\n---\n\n# Summary\n\nMy final solution achieved:\n\n| Metric     | Score    |\n| ---------- | -------- |\n| Public LB  | 0.987333 |\n| Private LB | 0.990000 |\n| Final Rank | 9 / 92   |\n\n## The solution combined:\n\n* EfficientNet-B4\n* ConvNeXt-Tiny\n* Post-processing simulation\n* MixUp augmentation\n* Label smoothing\n* Horizontal Flip TTA\n* Weighted soft-voting\n\nOne practical constraint throughout the competition was hardware availability. Rather than relying on extremely large models, I focused on maximizing performance through robust training strategies, realistic post-processing simulation, and complementary CNN architectures.\n\nThis ultimately proved more valuable than simply increasing model size and allowed me to achieve a Top 10 finish using computationally efficient backbones.\n\n---\n\n# Solution Pipeline\n\n```text\nTraining Images\n        ↓\nPost-Processing Simulation\n        ↓\nEfficientNet-B4 Training\n        ↓\nConvNeXt-Tiny Training\n        ↓\nHorizontal Flip TTA\n        ↓\nWeighted Soft Voting\n        ↓\nFinal Predictions\n```\n\n# Understanding the Challenge\n\nOne observation quickly became apparent:\n\nThe training images were relatively clean, but the hidden test images were intentionally degraded through various post-processing operations.\n\nThese included:\n\n* JPEG compression\n* WebP compression\n* Cropping\n* Resizing\n* Rotation\n* Brightness and contrast adjustment\n* Gaussian blur\n* Grayscale conversion\n* Super-resolution\n\nA model trained exclusively on clean images could easily learn shortcuts that disappear after these transformations.\n\nInstead of learning fragile pixel-level patterns, the model needed to learn generator-specific characteristics that remained stable after repeated image manipulations.\n\nThis realization ultimately shaped the entire training pipeline.\n\n# EfficientNet-B4 Baseline\n\nMy first competitive model was **EfficientNet-B4**.\n\nThe model was trained using:\n\n* Input size: 384×384\n* AdamW\n* Label smoothing (0.1)\n* MixUp (α = 0.2)\n* Cosine Annealing Warm Restarts\n\n### Results\n\n* Validation Accuracy: **98.36%**\n* Test Accuracy: **98.07%**\n* Public LB: **0.981333**\n* Private LB: **0.979333**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F127ff9aa2a2f75f73c6d2f2a0a29296c%2FEFFCIENT%20TRAIN%20AND%20VAL%20LOS.png?generation=1780696288263203&alt=media)\n\nOne encouraging sign was the small gap between validation and testing performance. This suggested that the augmentation strategy was helping the model generalize rather than simply memorize training samples.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F531da2ac95cec084a3db2b593db00201%2FEFFICIENT%20NET.png?generation=1780696299175320&alt=media)\n\nMost mistakes occurred between **Stable Diffusion 3** and **Stable Diffusion 3.5**, which was not particularly surprising given their architectural similarity and overlapping visual characteristics.\n\n\n# Why I Added ConvNeXt\n\nAlthough EfficientNet-B4 was already performing well, I wanted to determine whether a different architecture could learn complementary representations from the same data.\n\nRather than changing the augmentation pipeline, I kept everything identical and replaced only the backbone.\n\nThis allowed me to isolate the effect of architecture.\n\n### ConvNeXt-Tiny Results\n\n* Validation Accuracy: **99.21%**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F23409343%2F262846bbcd71790c68610f215cb6bbef%2Fcovnextt.png?generation=1780696332534394&alt=media)\n\nInterestingly, ConvNeXt-Tiny achieved a higher validation accuracy than EfficientNet-B4 despite using the exact same preprocessing and augmentation pipeline.\n\nThis suggested that the training strategy generalized effectively across architectures and that ConvNeXt was capturing information not fully represented by EfficientNet-B4.\n\n\n\n# What Contributed Most to Performance?\n\nThroughout the competition, I experimented with different architectures and ensemble configurations.\n\nOne of the most important observations was that the largest improvements did not come from increasing model complexity.\n\nThey came from improving robustness.\n\n## 1. Post-Processing Simulation (Largest Contribution)\n\nThe biggest breakthrough came when I stopped treating the problem as a standard image classification task and started treating it as a robustness problem.\n\nThe hidden test images were intentionally modified through compression, resizing, blurring, and other transformations. As a result, models trained only on clean images risked learning shortcuts that would disappear during evaluation.\n\nTo bridge this gap, I created a post-processing simulator that exposed the models to realistic image degradations throughout training:\n\n```python\noperations = [\n    ('jpeg', lambda x: self.apply_jpeg_compression(x)),\n    ('webp', lambda x: self.apply_webp_compression(x)),\n    ('crop', lambda x: self.apply_random_crop(x)),\n    ('resize', lambda x: self.apply_resizing(x)),\n    ('rotate', lambda x: self.apply_rotation_crop(x)),\n    ('blur', lambda x: self.apply_blur(x)),\n    ('brightness', lambda x: self.apply_brightness_contrast(x)),\n    ('grayscale', lambda x: self.apply_grayscale(x)),\n    ('superres', lambda x: self.apply_super_resolution(x)),\n]\n\nnum_ops = np.random.randint(1, 3)\nselected = np.random.choice(len(operations), num_ops, replace=False)\n\nfor idx in selected:\n    _, op_func = operations[idx]\n    img = op_func(img)\n```\n\nBy repeatedly exposing the models to transformed versions of the same image, they were forced to learn generator-specific fingerprints that remained stable across image degradations.\n\nLooking back, this ended up being the most important part of the solution.\n\n## 2. MixUp and Label Smoothing\n\nTo further improve generalization, I combined MixUp augmentation and label smoothing.\n\n### MixUp\n\n```python\ndef mixup_data(x, y, alpha=0.2):\n    lam = np.random.beta(alpha, alpha)\n\n    index = torch.randperm(x.size(0)).to(x.device)\n\n    mixed_x = lam * x + (1 - lam) * x[index]\n\n    y_a, y_b = y, y[index]\n\n    return mixed_x, y_a, y_b, lam\n```\n\n### Label Smoothing\n\n```python\ncriterion = nn.CrossEntropyLoss(\n    label_smoothing=0.1\n)\n```\n\nThese techniques improved generalization, reduced overconfidence, and helped the models distinguish visually similar generators.\n\n\n## 3. Test-Time Augmentation (TTA)\n\nDuring inference, I used horizontal flip test-time augmentation.\n\nPredictions from the original image and its horizontally flipped version were averaged before generating the final prediction.\n\n```python\ndef predict_with_tta(model, images):\n    outputs = model(images)\n\n    outputs_flip = model(\n        torch.flip(images, dims=[3])\n    )\n\n    return (outputs + outputs_flip) / 2\n```\n\nWhile the improvement was modest compared to post-processing simulation, TTA consistently produced more stable predictions and provided a useful boost to the final ensemble.\n\n\n# Ensemble Strategy\n\nAfter establishing strong individual models, I explored combining EfficientNet-B4 and ConvNeXt-Tiny through weighted soft-voting.\n\nAlthough both models were CNN-based architectures, they learned different representations of the data and produced slightly different prediction patterns.\n\nSeveral weight combinations were evaluated:\n\n| EfficientNet-B4 | ConvNeXt-Tiny | Validation Accuracy |\n| --------------- | ------------- | ------------------- |\n| 0.7             | 0.3           | 99.07%              |\n| 0.6             | 0.4           | 99.21%              |\n| 0.5             | 0.5           | 99.29%              |\n| 0.4             | 0.6           | 99.36%              |\n| 0.3             | 0.7           | 99.36%              |\n\nThe final ensemble used:\n\n* EfficientNet-B4: 30%\n* ConvNeXt-Tiny: 70%\n\n```python\nensemble_out = (\n    0.30 * efficientnet_output +\n    0.70 * convnext_output\n)\n```\n\n### Final Ensemble Results\n\n* Validation Accuracy: **99.36%**\n* Public LB: **0.987333**\n* Private LB: **0.990000**\n\nThe ensemble consistently outperformed both individual models and proved more stable across validation and leaderboard evaluations.\n\n\n# Leaderboard Progression\n\n| Version | Model                           | Public LB | Private LB |\n| ------- | ------------------------------- | --------- | ---------- |\n| V4      | EfficientNet-B4                 | 0.981333  | 0.979333   |\n| V5      | Initial Ensemble                | 0.982666  | 0.988000   |\n| V6      | EfficientNet-B4 + ConvNeXt-Tiny | 0.987333  | 0.990000   |\n\nEach iteration built upon the previous one, with the final ensemble delivering the strongest overall performance.\n\n\n# Key Observations\n\nSeveral observations stood out during the competition.\n\n### Post-processing simulation was the biggest contributor\n\nThe largest gains came from exposing the models to realistic image transformations during training rather than increasing model complexity.\n\nTraining on compressed, resized, cropped, blurred, and grayscale images forced the models to learn generator-specific fingerprints that remained detectable even after heavy image manipulation.\n\n### Most errors were concentrated in a few generator pairs\n\nThe confusion matrix showed that errors were not evenly distributed across all classes.\n\nMost misclassifications occurred between Stable Diffusion 3 and Stable Diffusion 3.5, while several classes achieved near-perfect or perfect classification accuracy.\n\nThis suggests that the model successfully learned distinctive generator fingerprints but still struggled when generators shared highly similar architectural characteristics.\n\n### ConvNeXt and EfficientNet learned different patterns\n\nConvNeXt-Tiny achieved higher standalone validation accuracy, but EfficientNet-B4 continued to contribute useful information during ensembling.\n\nThe improvement from V4 (0.979333 Private LB) to V6 (0.990000 Private LB) suggests that the two models were making different mistakes and provided complementary predictions.\n\n### Simple inference techniques were enough\n\nHorizontal flip TTA consistently improved prediction stability without introducing significant computational overhead.\n\nMore importantly, most of the performance gains came from the training pipeline itself rather than complex inference tricks.\n\n### Efficient architectures remained competitive\n\nA major goal throughout the competition was achieving strong performance without relying on extremely large models.\n\nEfficientNet-B4 and ConvNeXt-Tiny provided an effective balance between computational cost and accuracy, ultimately achieving a Top 10 finish while remaining practical to train on limited hardware.\n\n# Final Takeaways\n\nThe most valuable lesson from this competition was that strong solutions are not always the result of larger models.\n\nIn this challenge, the largest improvements came from understanding the evaluation conditions and designing a training pipeline that reflected those conditions as closely as possible.\n\nPost-processing simulation proved far more impactful than increasing model complexity, while ConvNeXt-Tiny and EfficientNet-B4 demonstrated that efficient CNN architectures can still achieve highly competitive results when paired with robust augmentation strategies.\n\nBy combining realistic image transformations, MixUp, label smoothing, horizontal-flip TTA, and weighted soft-voting, I was able to improve from a **0.979333 Private Leaderboard baseline** to a final score of **0.990000**, securing **9th place out of 92 teams**.\n\nThis competition was a great reminder that understanding the data distribution often matters more than simply scaling the model.\n\nThank you again to the organizers for creating such a challenging and educational competition.\n"
  }
}