{
  "id": 587553,
  "title": "45th Place Solution",
  "url": "/competitions/waveform-inversion/writeups/iwa-iwa-45th-place-solution",
  "author_name": "",
  "post_date": "2025-07-01T14:39:25.541383800Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<h2>Introduction</h2>\n<p>First of all, I would like to express my deep gratitude to the organizers Yale, UNC-CH, Kaggle. This competition was interesting and educational, and I'm glad I participated.<br>\nI would also like to deeply thank <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for providing the public code. My solution is developed based on Bartley's public code.</p>\n<hr>\n<h2>1.Overview</h2>\n<ul>\n<li>Trained several models using ConvNeXt and CAFormer models with different model-size configurations</li>\n<li>Best single model was ConvNeXt-Base (CV=23.87, LB=25.7)</li>\n<li>Final submission was a weighted median ensemble of 9 submissions achieving Public MAE = 23.2, Private MAE = 23.1</li>\n<li>Due to the large dataset size, cross-validation was challenging, so I adopted an extended hold-out approach, which maintained high correlation with Public LB</li>\n</ul>\n<h2>2. Dataset</h2>\n<ul>\n<li><strong>OpenFWI (all data)</strong> only  </li>\n<li>No resizing, no augmentation, no synthetic generation  </li>\n</ul>\n<h2>3. Validation Strategy</h2>\n<ul>\n<li><strong>Expanded hold-out</strong> split: <strong>Train ≈ 88 % / Val ≈ 12 %</strong> from the public code. </li>\n</ul>\n<h2>4. Single-Model Results</h2>\n<p><strong>Table&nbsp;1 – Single-model performance (lower is better)</strong>  </p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Epochs</th>\n<th>CV MAE</th>\n<th>Public MAE</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ConvNeXt-Base</td>\n<td>200</td>\n<td><strong>23.87</strong></td>\n<td><strong>25.7</strong></td>\n<td>Best single model</td>\n</tr>\n<tr>\n<td>CAFormer-m36</td>\n<td>180</td>\n<td>24.08</td>\n<td>26.0</td>\n<td>StarReLU</td>\n</tr>\n<tr>\n<td>CAFormer-b36</td>\n<td>180</td>\n<td>24.96</td>\n<td>29.5</td>\n<td>StarReLU</td>\n</tr>\n<tr>\n<td>ConvNeXt-Base</td>\n<td>150</td>\n<td>25.58</td>\n<td>26.7</td>\n<td>Checkpoint fine-tuned twice</td>\n</tr>\n<tr>\n<td>ConvNeXt-Large</td>\n<td>150</td>\n<td>26.08</td>\n<td>27.8</td>\n<td>—</td>\n</tr>\n<tr>\n<td>Flash InternImage-b</td>\n<td>50</td>\n<td>34.60</td>\n<td>37.6</td>\n<td>FP32, ≈ 210 min / epoch</td>\n</tr>\n<tr>\n<td>ConvNeXt-Base (full data)</td>\n<td>218</td>\n<td>—</td>\n<td>27.7</td>\n<td>Trained on 100 % data</td>\n</tr>\n<tr>\n<td>ConvNeXt-Base (full data)</td>\n<td>243</td>\n<td>—</td>\n<td>26.7</td>\n<td>Trained on 100 % data</td>\n</tr>\n<tr>\n<td>ConvNeXt-Base (full data)</td>\n<td>258</td>\n<td>—</td>\n<td>26.6</td>\n<td>Trained on 100 % data</td>\n</tr>\n</tbody>\n</table>\n<p>Hyper-parameters tuned: <code>weight_decay</code>, <code>dropout_rate</code>, <code>decoder_channels</code>, <code>clip_grad_norm</code>  <br>\nLR restarts: resumed from the previous terminal LR (<code>1 e-5</code>) with <code>CosineAnnealingLR(T_max = 50)</code>, which proved more stable than warm restarts.</p>\n<h2>5. Ensembling &amp; Post-processing</h2>\n<p><strong>Table&nbsp;2 – Ensemble performance (Public MAE)</strong>  </p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Submissions</th>\n<th>Public MAE</th>\n<th>Δ vs. Best Single</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Simple median</td>\n<td>9</td>\n<td>23.4</td>\n<td>−2.3</td>\n</tr>\n<tr>\n<td><strong>Weighted median (final)</strong></td>\n<td><strong>9</strong></td>\n<td><strong>23.2</strong></td>\n<td><strong>−2.5</strong></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Median-based methods clearly outperform means for MAE.  </li>\n<li>Weighted median improved MAE by <strong>0.2</strong> over the simple median.  </li>\n<li>Type-wise weighted means (after predicting sample type with 99.5 % accuracy) did <strong>not</strong> beat median.  </li>\n<li>Post-processing is limited to <strong><code>clip(1500, 4500)</code></strong> (additional –0.002 CV MAE).</li>\n</ul>\n<h2>6. Speed-up Techniques</h2>\n<table>\n<thead>\n<tr>\n<th>Category</th>\n<th>Key settings</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Compiler / Optimizer</strong></td>\n<td><code>torch.compile(mode=\"max-autotune\")</code>, <strong>Fused AdamW</strong></td>\n</tr>\n<tr>\n<td><strong>Mixed Precision</strong></td>\n<td>AMP (<strong>bfloat16</strong>) for all ConvNeXt / CAFormer runs</td>\n</tr>\n<tr>\n<td><strong>DataLoader</strong></td>\n<td><code>num_workers</code> tuned, <code>prefetch_factor</code> tuned, <code>pin_memory=True</code>, <code>persistent_workers=True</code></td>\n</tr>\n<tr>\n<td><strong>Hardware</strong></td>\n<td>Experiments run on <strong>RTX 3090</strong> &amp; <strong>A100</strong><br>Flash InternImage remained FP32 due to DCNv4 → <strong>≈ 210 min / epoch (A100)</strong></td>\n</tr>\n</tbody>\n</table>\n<p>“num_workers was experimentally tuned—the value is a sweet spot: too many or too few threads slowed training, so we benchmark-benchmarked different settings and used the fastest one for all runs.”</p>\n<h2>7. What Didn’t Work</h2>\n<table>\n<thead>\n<tr>\n<th>What failed</th>\n<th>Likely cause / lesson</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Type-wise weights, heavy post-processing</td>\n<td>Median already robust; extra weighting brought no gain</td>\n</tr>\n<tr>\n<td>Flash InternImage-b</td>\n<td>AMP unstable; FP32 made training prohibitively slow, limiting epochs</td>\n</tr>\n<tr>\n<td>EVA family</td>\n<td>Mishandled window / resize parameters – never converged</td>\n</tr>\n<tr>\n<td>Full-data training (100 %)</td>\n<td>Slightly worse Public MAE vs. 88 % split, much longer runtime</td>\n</tr>\n</tbody>\n</table>\n<h2>8. Results &amp; Lessons Learned</h2>\n<ul>\n<li><strong>Final scores:</strong> Public MAE 23.2 / Private MAE 23.1 (weighted median)  </li>\n<li><strong>Ensembling gain</strong> saturated at ≈ 2.5 MAE → future effort should target single-model quality.  </li>\n<li>LR scheduling and restart strategy had greater impact than backbone size beyond <em>ConvNeXt-Base</em>.  </li>\n<li>Further gains likely lie in synthetic data and higher-resolution training, as hinted by top teams.</li>\n</ul>\n<hr>\n<h2>References</h2>\n<ul>\n<li>Bartley Baseline Notebook  <br>\n<a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline</a>  <br>\n<a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved</a></li>\n<li>Training / Speed-up Discussion  <br>\n<a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/583896\" target=\"_blank\">https://www.kaggle.com/competitions/waveform-inversion/discussion/583896</a></li>\n</ul>\n<hr>",
  "messages": [
    {
      "id": "3238099",
      "postDate": "07/01/2025 14:39:25",
      "content": "<h2>Introduction</h2>\n<p>First of all, I would like to express my deep gratitude to the organizers Yale, UNC-CH, Kaggle. This competition was interesting and educational, and I'm glad I participated.<br>\nI would also like to deeply thank <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for providing the public code. My solution is developed based on Bartley's public code.</p>\n<hr>\n<h2>1.Overview</h2>\n<ul>\n<li>Trained several models using ConvNeXt and CAFormer models with different model-size configurations</li>\n<li>Best single model was ConvNeXt-Base (CV=23.87, LB=25.7)</li>\n<li>Final submission was a weighted median ensemble of 9 submissions achieving Public MAE = 23.2, Private MAE = 23.1</li>\n<li>Due to the large dataset size, cross-validation was challenging, so I adopted an extended hold-out approach, which maintained high correlation with Public LB</li>\n</ul>\n<h2>2. Dataset</h2>\n<ul>\n<li><strong>OpenFWI (all data)</strong> only  </li>\n<li>No resizing, no augmentation, no synthetic generation  </li>\n</ul>\n<h2>3. Validation Strategy</h2>\n<ul>\n<li><strong>Expanded hold-out</strong> split: <strong>Train ≈ 88 % / Val ≈ 12 %</strong> from the public code. </li>\n</ul>\n<h2>4. Single-Model Results</h2>\n<p><strong>Table&nbsp;1 – Single-model performance (lower is better)</strong>  </p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Epochs</th>\n<th>CV MAE</th>\n<th>Public MAE</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ConvNeXt-Base</td>\n<td>200</td>\n<td><strong>23.87</strong></td>\n<td><strong>25.7</strong></td>\n<td>Best single model</td>\n</tr>\n<tr>\n<td>CAFormer-m36</td>\n<td>180</td>\n<td>24.08</td>\n<td>26.0</td>\n<td>StarReLU</td>\n</tr>\n<tr>\n<td>CAFormer-b36</td>\n<td>180</td>\n<td>24.96</td>\n<td>29.5</td>\n<td>StarReLU</td>\n</tr>\n<tr>\n<td>ConvNeXt-Base</td>\n<td>150</td>\n<td>25.58</td>\n<td>26.7</td>\n<td>Checkpoint fine-tuned twice</td>\n</tr>\n<tr>\n<td>ConvNeXt-Large</td>\n<td>150</td>\n<td>26.08</td>\n<td>27.8</td>\n<td>—</td>\n</tr>\n<tr>\n<td>Flash InternImage-b</td>\n<td>50</td>\n<td>34.60</td>\n<td>37.6</td>\n<td>FP32, ≈ 210 min / epoch</td>\n</tr>\n<tr>\n<td>ConvNeXt-Base (full data)</td>\n<td>218</td>\n<td>—</td>\n<td>27.7</td>\n<td>Trained on 100 % data</td>\n</tr>\n<tr>\n<td>ConvNeXt-Base (full data)</td>\n<td>243</td>\n<td>—</td>\n<td>26.7</td>\n<td>Trained on 100 % data</td>\n</tr>\n<tr>\n<td>ConvNeXt-Base (full data)</td>\n<td>258</td>\n<td>—</td>\n<td>26.6</td>\n<td>Trained on 100 % data</td>\n</tr>\n</tbody>\n</table>\n<p>Hyper-parameters tuned: <code>weight_decay</code>, <code>dropout_rate</code>, <code>decoder_channels</code>, <code>clip_grad_norm</code>  <br>\nLR restarts: resumed from the previous terminal LR (<code>1 e-5</code>) with <code>CosineAnnealingLR(T_max = 50)</code>, which proved more stable than warm restarts.</p>\n<h2>5. Ensembling &amp; Post-processing</h2>\n<p><strong>Table&nbsp;2 – Ensemble performance (Public MAE)</strong>  </p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Submissions</th>\n<th>Public MAE</th>\n<th>Δ vs. Best Single</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Simple median</td>\n<td>9</td>\n<td>23.4</td>\n<td>−2.3</td>\n</tr>\n<tr>\n<td><strong>Weighted median (final)</strong></td>\n<td><strong>9</strong></td>\n<td><strong>23.2</strong></td>\n<td><strong>−2.5</strong></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Median-based methods clearly outperform means for MAE.  </li>\n<li>Weighted median improved MAE by <strong>0.2</strong> over the simple median.  </li>\n<li>Type-wise weighted means (after predicting sample type with 99.5 % accuracy) did <strong>not</strong> beat median.  </li>\n<li>Post-processing is limited to <strong><code>clip(1500, 4500)</code></strong> (additional –0.002 CV MAE).</li>\n</ul>\n<h2>6. Speed-up Techniques</h2>\n<table>\n<thead>\n<tr>\n<th>Category</th>\n<th>Key settings</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Compiler / Optimizer</strong></td>\n<td><code>torch.compile(mode=\"max-autotune\")</code>, <strong>Fused AdamW</strong></td>\n</tr>\n<tr>\n<td><strong>Mixed Precision</strong></td>\n<td>AMP (<strong>bfloat16</strong>) for all ConvNeXt / CAFormer runs</td>\n</tr>\n<tr>\n<td><strong>DataLoader</strong></td>\n<td><code>num_workers</code> tuned, <code>prefetch_factor</code> tuned, <code>pin_memory=True</code>, <code>persistent_workers=True</code></td>\n</tr>\n<tr>\n<td><strong>Hardware</strong></td>\n<td>Experiments run on <strong>RTX 3090</strong> &amp; <strong>A100</strong><br>Flash InternImage remained FP32 due to DCNv4 → <strong>≈ 210 min / epoch (A100)</strong></td>\n</tr>\n</tbody>\n</table>\n<p>“num_workers was experimentally tuned—the value is a sweet spot: too many or too few threads slowed training, so we benchmark-benchmarked different settings and used the fastest one for all runs.”</p>\n<h2>7. What Didn’t Work</h2>\n<table>\n<thead>\n<tr>\n<th>What failed</th>\n<th>Likely cause / lesson</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Type-wise weights, heavy post-processing</td>\n<td>Median already robust; extra weighting brought no gain</td>\n</tr>\n<tr>\n<td>Flash InternImage-b</td>\n<td>AMP unstable; FP32 made training prohibitively slow, limiting epochs</td>\n</tr>\n<tr>\n<td>EVA family</td>\n<td>Mishandled window / resize parameters – never converged</td>\n</tr>\n<tr>\n<td>Full-data training (100 %)</td>\n<td>Slightly worse Public MAE vs. 88 % split, much longer runtime</td>\n</tr>\n</tbody>\n</table>\n<h2>8. Results &amp; Lessons Learned</h2>\n<ul>\n<li><strong>Final scores:</strong> Public MAE 23.2 / Private MAE 23.1 (weighted median)  </li>\n<li><strong>Ensembling gain</strong> saturated at ≈ 2.5 MAE → future effort should target single-model quality.  </li>\n<li>LR scheduling and restart strategy had greater impact than backbone size beyond <em>ConvNeXt-Base</em>.  </li>\n<li>Further gains likely lie in synthetic data and higher-resolution training, as hinted by top teams.</li>\n</ul>\n<hr>\n<h2>References</h2>\n<ul>\n<li>Bartley Baseline Notebook  <br>\n<a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline</a>  <br>\n<a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved</a></li>\n<li>Training / Speed-up Discussion  <br>\n<a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/583896\" target=\"_blank\">https://www.kaggle.com/competitions/waveform-inversion/discussion/583896</a></li>\n</ul>\n<hr>",
      "rawMarkdown": "## Introduction\nFirst of all, I would like to express my deep gratitude to the organizers Yale, UNC-CH, Kaggle. This competition was interesting and educational, and I'm glad I participated.\nI would also like to deeply thank @brendanartley for providing the public code. My solution is developed based on Bartley's public code.\n\n---\n## 1.Overview\n- Trained several models using ConvNeXt and CAFormer models with different model-size configurations\n- Best single model was ConvNeXt-Base (CV=23.87, LB=25.7)\n- Final submission was a weighted median ensemble of 9 submissions achieving Public MAE = 23.2, Private MAE = 23.1\n- Due to the large dataset size, cross-validation was challenging, so I adopted an extended hold-out approach, which maintained high correlation with Public LB\n\n## 2. Dataset\n- **OpenFWI (all data)** only  \n- No resizing, no augmentation, no synthetic generation  \n\n## 3. Validation Strategy\n- **Expanded hold-out** split: **Train ≈ 88 % / Val ≈ 12 %** from the public code. \n\n## 4. Single-Model Results  \n\n**Table&nbsp;1 – Single-model performance (lower is better)**  \n\n| Backbone | Epochs | CV MAE | Public MAE | Notes |\n|----------|-------:|-------:|-----------:|-------|\n| ConvNeXt-Base | 200 | **23.87** | **25.7** | Best single model |\n| CAFormer-m36  | 180 | 24.08 | 26.0 | StarReLU |\n| CAFormer-b36  | 180 | 24.96 | 29.5 | StarReLU |\n| ConvNeXt-Base | 150 | 25.58 | 26.7 | Checkpoint fine-tuned twice |\n| ConvNeXt-Large | 150 | 26.08 | 27.8 | — |\n| Flash InternImage-b | 50 | 34.60 | 37.6 | FP32, ≈ 210 min / epoch |\n| ConvNeXt-Base (full data) | 218 | — | 27.7 | Trained on 100 % data |\n| ConvNeXt-Base (full data) | 243 | — | 26.7 | Trained on 100 % data |\n| ConvNeXt-Base (full data) | 258 | — | 26.6 | Trained on 100 % data |\n\nHyper-parameters tuned: `weight_decay`, `dropout_rate`, `decoder_channels`, `clip_grad_norm`  \nLR restarts: resumed from the previous terminal LR (`1 e-5`) with `CosineAnnealingLR(T_max = 50)`, which proved more stable than warm restarts.\n\n## 5. Ensembling & Post-processing  \n\n**Table&nbsp;2 – Ensemble performance (Public MAE)**  \n\n| Method | Submissions | Public MAE | Δ vs. Best Single |\n|--------|-----------:|----:|------------------:|\n| Simple median | 9 | 23.4 | −2.3 |\n| **Weighted median (final)** | **9** | **23.2** | **−2.5** |\n\n- Median-based methods clearly outperform means for MAE.  \n- Weighted median improved MAE by **0.2** over the simple median.  \n- Type-wise weighted means (after predicting sample type with 99.5 % accuracy) did **not** beat median.  \n- Post-processing is limited to **`clip(1500, 4500)`** (additional –0.002 CV MAE).\n\n## 6. Speed-up Techniques\n| Category | Key settings |\n|----------|--------------|\n| **Compiler / Optimizer** | `torch.compile(mode=\"max-autotune\")`, **Fused AdamW** |\n| **Mixed Precision** | AMP (**bfloat16**) for all ConvNeXt / CAFormer runs |\n| **DataLoader** | `num_workers` tuned, `prefetch_factor` tuned, `pin_memory=True`, `persistent_workers=True` |\n| **Hardware** | Experiments run on **RTX 3090** & **A100**<br>Flash InternImage remained FP32 due to DCNv4 → **≈ 210 min / epoch (A100)** |\n\n“num_workers was experimentally tuned—the value is a sweet spot: too many or too few threads slowed training, so we benchmark-benchmarked different settings and used the fastest one for all runs.”\n\n\n## 7. What Didn’t Work\n\n| What failed | Likely cause / lesson |\n|-------------|-----------------------|\n| Type-wise weights, heavy post-processing | Median already robust; extra weighting brought no gain |\n| Flash InternImage-b | AMP unstable; FP32 made training prohibitively slow, limiting epochs |\n| EVA family | Mishandled window / resize parameters – never converged |\n| Full-data training (100 %) | Slightly worse Public MAE vs. 88 % split, much longer runtime |\n\n## 8. Results & Lessons Learned\n- **Final scores:** Public MAE 23.2 / Private MAE 23.1 (weighted median)  \n- **Ensembling gain** saturated at ≈ 2.5 MAE → future effort should target single-model quality.  \n- LR scheduling and restart strategy had greater impact than backbone size beyond *ConvNeXt-Base*.  \n- Further gains likely lie in synthetic data and higher-resolution training, as hinted by top teams.\n\n---\n\n## References\n- Bartley Baseline Notebook  \n  <https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline>  \n  <https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved>\n- Training / Speed-up Discussion  \n  <https://www.kaggle.com/competitions/waveform-inversion/discussion/583896>\n\n---",
      "votes": null
    },
    {
      "id": "3239435",
      "postDate": "07/02/2025 20:18:10",
      "content": "<p>Thanks for many experimental results! They are very useful, but some results are not obvious. Do you have some idea why?</p>\n<ul>\n<li>Why ConvNeXt-Base (full data) runs are worse than the first ConvNeXt-Base although there ~10% more data? Difference in learning rates?</li>\n<li>Why ConvNeXt-Large is worse than Base for same 150 epochs? Do you see larger train-val loss gap (overfit)? Some people said larger models were always better in this competition</li>\n</ul>\n<p>Weighted median is great! I only thought about median and weighted mean. Is it implemented in well known libraries or in public notebook?</p>",
      "rawMarkdown": "Thanks for many experimental results! They are very useful, but some results are not obvious. Do you have some idea why?\n\n* Why ConvNeXt-Base (full data) runs are worse than the first ConvNeXt-Base although there ~10% more data? Difference in learning rates?\n* Why ConvNeXt-Large is worse than Base for same 150 epochs? Do you see larger train-val loss gap (overfit)? Some people said larger models were always better in this competition\n\nWeighted median is great! I only thought about median and weighted mean. Is it implemented in well known libraries or in public notebook?",
      "votes": null
    },
    {
      "id": "3240376",
      "postDate": "07/03/2025 18:23:58",
      "content": "<p>Thank you for your comment!</p>\n<blockquote>\n  <p>Why ConvNeXt-Base (full data) runs are worse than the first ConvNeXt-Base although there ~10% more data? Difference in learning rates?</p>\n</blockquote>\n<p>I realize my earlier conclusion that \"full-data didn't help\" was premature because the two runs weren't strictly comparable.</p>\n<p>The key differences we've identified are:<br>\n100% DATA: batch_size = 64, epochs = 300, drop_last_train = True<br>\n88% DATA: batch_size = 32, epochs = 200, drop_last_train = False</p>\n<p>Because the CosineAnnealingLR schedule was tied to <em>epochs</em>, the LR in the 100 % run fell much sooner (in step terms) than in the 88 % run—likely the key factor. The bigger batch and <code>drop_last_train=True</code> probably played a role as well.</p>\n<blockquote>\n  <p>Why ConvNeXt-Large is worse than Base for same 150 epochs? Do you see larger train-val loss gap (overfit)? Some people said larger models were always better in this competition</p>\n</blockquote>\n<p>I’m not sure of the reason, and I wonder why the Large model stagnated earlier than expected.<br>\nIn my experiments, the Large tends to overfit the training data more than the Base, and its validation MAE stops improving at an earlier stage. To counter this, I adjusted the hyper-parameters for the Large model as shown below, but even with these tweaks the validation MAE still plateaued sooner.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train MAE</th>\n<th>Validation MAE</th>\n<th>Key Hyper-parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>ConvNeXt-Base</strong></td>\n<td>19.83</td>\n<td>25.58</td>\n<td><code>drop_path_rate = 0.2</code>, <code>weight_decay = 0.03</code>, <code>layer_decay = 1.0</code> (disabled)</td>\n</tr>\n<tr>\n<td><strong>ConvNeXt-Large</strong></td>\n<td>18.62</td>\n<td>26.08 <br>(validation MAE plateaus earlier than Base)</td>\n<td><code>drop_path_rate = 0.4</code>, <code>weight_decay = 0.05</code>, <code>layer_decay = 0.7</code></td>\n</tr>\n</tbody>\n</table>\n<p>I intend to study other participants’ solutions to better understand how to improve the Large model. <br>\nIf you have any ideas, I’d like to hear them.</p>",
      "rawMarkdown": "Thank you for your comment!\n\n>Why ConvNeXt-Base (full data) runs are worse than the first ConvNeXt-Base although there ~10% more data? Difference in learning rates?\n\nI realize my earlier conclusion that \"full-data didn't help\" was premature because the two runs weren't strictly comparable.\n\nThe key differences we've identified are:\n100% DATA: batch_size = 64, epochs = 300, drop_last_train = True\n88% DATA: batch_size = 32, epochs = 200, drop_last_train = False\n\nBecause the CosineAnnealingLR schedule was tied to *epochs*, the LR in the 100 % run fell much sooner (in step terms) than in the 88 % run—likely the key factor. The bigger batch and `drop_last_train=True` probably played a role as well.\n\n>Why ConvNeXt-Large is worse than Base for same 150 epochs? Do you see larger train-val loss gap (overfit)? Some people said larger models were always better in this competition\n\nI’m not sure of the reason, and I wonder why the Large model stagnated earlier than expected.\nIn my experiments, the Large tends to overfit the training data more than the Base, and its validation MAE stops improving at an earlier stage. To counter this, I adjusted the hyper-parameters for the Large model as shown below, but even with these tweaks the validation MAE still plateaued sooner.\n\n| Model              | Train MAE | Validation MAE                                            | Key Hyper-parameters                                                          |\n| ------------------ | --------- | --------------------------------------------------------- | ----------------------------------------------------------------------------- |\n| **ConvNeXt-Base**  | 19.83   | 25.58                                                   | `drop_path_rate = 0.2`, `weight_decay = 0.03`, `layer_decay = 1.0` (disabled) |\n| **ConvNeXt-Large** | 18.62   | 26.08 <br />(validation MAE plateaus earlier than Base) | `drop_path_rate = 0.4`, `weight_decay = 0.05`, `layer_decay = 0.7`            |\n\n\nI intend to study other participants’ solutions to better understand how to improve the Large model. \nIf you have any ideas, I’d like to hear them.",
      "votes": null
    },
    {
      "id": "3240434",
      "postDate": "07/03/2025 19:40:13",
      "content": "<p>Thanks for more investigation!</p>\n<p>I see, I know some strong Kagglers use full-data training, but I was worrying training without val score; e.g. should we change batch size to have same number steps or keep batch size and change learning rate? Your result confirms that full-data training can be difficult and requires experience. In this competition, we have lots of data, so I keep one file (500 data) per family as a \"test set,\" in addition to K-fold validation set. This helped me having \"test score\" in the almost full-data training.</p>",
      "rawMarkdown": "Thanks for more investigation!\n\nI see, I know some strong Kagglers use full-data training, but I was worrying training without val score; e.g. should we change batch size to have same number steps or keep batch size and change learning rate? Your result confirms that full-data training can be difficult and requires experience. In this competition, we have lots of data, so I keep one file (500 data) per family as a \"test set,\" in addition to K-fold validation set. This helped me having \"test score\" in the almost full-data training.",
      "votes": null
    },
    {
      "id": "3241187",
      "postDate": "07/04/2025 16:06:51",
      "content": "<p>I think your point is spot-on. In fact, bumping the training set from 88% to 95%  would probably be enough to see whether the extra data boosts score. Putting too much faith in the Public Leaderboard is risky in many competitions—although in this one we were lucky that the public and private scores were closely aligned.</p>\n<blockquote>\n  <p>Weighted median is great! I only thought about median and weighted mean. Is it implemented in well known libraries or in public notebook?</p>\n</blockquote>\n<p>I missed your last question. Since Polars doesn't have a built-in weighted median, I implemented my own. As far as I know, this approach hasn't been widely used for ensembling in Kaggle notebooks. Median was already giving a much better score than the mean, so I experimented with variations on the median, and the weighted median happened to work well in this case.</p>",
      "rawMarkdown": "I think your point is spot-on. In fact, bumping the training set from 88% to 95%  would probably be enough to see whether the extra data boosts score. Putting too much faith in the Public Leaderboard is risky in many competitions—although in this one we were lucky that the public and private scores were closely aligned.\n\n>Weighted median is great! I only thought about median and weighted mean. Is it implemented in well known libraries or in public notebook?\n\nI missed your last question. Since Polars doesn't have a built-in weighted median, I implemented my own. As far as I know, this approach hasn't been widely used for ensembling in Kaggle notebooks. Median was already giving a much better score than the mean, so I experimented with variations on the median, and the weighted median happened to work well in this case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3239435,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "07/02/2025 20:18:10",
      "content": "<p>Thanks for many experimental results! They are very useful, but some results are not obvious. Do you have some idea why?</p>\n<ul>\n<li>Why ConvNeXt-Base (full data) runs are worse than the first ConvNeXt-Base although there ~10% more data? Difference in learning rates?</li>\n<li>Why ConvNeXt-Large is worse than Base for same 150 epochs? Do you see larger train-val loss gap (overfit)? Some people said larger models were always better in this competition</li>\n</ul>\n<p>Weighted median is great! I only thought about median and weighted mean. Is it implemented in well known libraries or in public notebook?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3240376,
          "author_name": "iwaiwa",
          "author_url": "",
          "post_date": "07/03/2025 18:23:58",
          "content": "<p>Thank you for your comment!</p>\n<blockquote>\n  <p>Why ConvNeXt-Base (full data) runs are worse than the first ConvNeXt-Base although there ~10% more data? Difference in learning rates?</p>\n</blockquote>\n<p>I realize my earlier conclusion that \"full-data didn't help\" was premature because the two runs weren't strictly comparable.</p>\n<p>The key differences we've identified are:<br>\n100% DATA: batch_size = 64, epochs = 300, drop_last_train = True<br>\n88% DATA: batch_size = 32, epochs = 200, drop_last_train = False</p>\n<p>Because the CosineAnnealingLR schedule was tied to <em>epochs</em>, the LR in the 100 % run fell much sooner (in step terms) than in the 88 % run—likely the key factor. The bigger batch and <code>drop_last_train=True</code> probably played a role as well.</p>\n<blockquote>\n  <p>Why ConvNeXt-Large is worse than Base for same 150 epochs? Do you see larger train-val loss gap (overfit)? Some people said larger models were always better in this competition</p>\n</blockquote>\n<p>I’m not sure of the reason, and I wonder why the Large model stagnated earlier than expected.<br>\nIn my experiments, the Large tends to overfit the training data more than the Base, and its validation MAE stops improving at an earlier stage. To counter this, I adjusted the hyper-parameters for the Large model as shown below, but even with these tweaks the validation MAE still plateaued sooner.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train MAE</th>\n<th>Validation MAE</th>\n<th>Key Hyper-parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>ConvNeXt-Base</strong></td>\n<td>19.83</td>\n<td>25.58</td>\n<td><code>drop_path_rate = 0.2</code>, <code>weight_decay = 0.03</code>, <code>layer_decay = 1.0</code> (disabled)</td>\n</tr>\n<tr>\n<td><strong>ConvNeXt-Large</strong></td>\n<td>18.62</td>\n<td>26.08 <br>(validation MAE plateaus earlier than Base)</td>\n<td><code>drop_path_rate = 0.4</code>, <code>weight_decay = 0.05</code>, <code>layer_decay = 0.7</code></td>\n</tr>\n</tbody>\n</table>\n<p>I intend to study other participants’ solutions to better understand how to improve the Large model. <br>\nIf you have any ideas, I’d like to hear them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3240434,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "07/03/2025 19:40:13",
          "content": "<p>Thanks for more investigation!</p>\n<p>I see, I know some strong Kagglers use full-data training, but I was worrying training without val score; e.g. should we change batch size to have same number steps or keep batch size and change learning rate? Your result confirms that full-data training can be difficult and requires experience. In this competition, we have lots of data, so I keep one file (500 data) per family as a \"test set,\" in addition to K-fold validation set. This helped me having \"test score\" in the almost full-data training.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3241187,
              "author_name": "iwaiwa",
              "author_url": "",
              "post_date": "07/04/2025 16:06:51",
              "content": "<p>I think your point is spot-on. In fact, bumping the training set from 88% to 95%  would probably be enough to see whether the extra data boosts score. Putting too much faith in the Public Leaderboard is risky in many competitions—although in this one we were lucky that the public and private scores were closely aligned.</p>\n<blockquote>\n  <p>Weighted median is great! I only thought about median and weighted mean. Is it implemented in well known libraries or in public notebook?</p>\n</blockquote>\n<p>I missed your last question. Since Polars doesn't have a built-in weighted median, I implemented my own. As far as I know, this approach hasn't been widely used for ensembling in Kaggle notebooks. Median was already giving a much better score than the mean, so I experimented with variations on the median, and the weighted median happened to work well in this case.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3238099": "## Introduction\nFirst of all, I would like to express my deep gratitude to the organizers Yale, UNC-CH, Kaggle. This competition was interesting and educational, and I'm glad I participated.\nI would also like to deeply thank @brendanartley for providing the public code. My solution is developed based on Bartley's public code.\n\n---\n## 1.Overview\n- Trained several models using ConvNeXt and CAFormer models with different model-size configurations\n- Best single model was ConvNeXt-Base (CV=23.87, LB=25.7)\n- Final submission was a weighted median ensemble of 9 submissions achieving Public MAE = 23.2, Private MAE = 23.1\n- Due to the large dataset size, cross-validation was challenging, so I adopted an extended hold-out approach, which maintained high correlation with Public LB\n\n## 2. Dataset\n- **OpenFWI (all data)** only  \n- No resizing, no augmentation, no synthetic generation  \n\n## 3. Validation Strategy\n- **Expanded hold-out** split: **Train ≈ 88 % / Val ≈ 12 %** from the public code. \n\n## 4. Single-Model Results  \n\n**Table&nbsp;1 – Single-model performance (lower is better)**  \n\n| Backbone | Epochs | CV MAE | Public MAE | Notes |\n|----------|-------:|-------:|-----------:|-------|\n| ConvNeXt-Base | 200 | **23.87** | **25.7** | Best single model |\n| CAFormer-m36  | 180 | 24.08 | 26.0 | StarReLU |\n| CAFormer-b36  | 180 | 24.96 | 29.5 | StarReLU |\n| ConvNeXt-Base | 150 | 25.58 | 26.7 | Checkpoint fine-tuned twice |\n| ConvNeXt-Large | 150 | 26.08 | 27.8 | — |\n| Flash InternImage-b | 50 | 34.60 | 37.6 | FP32, ≈ 210 min / epoch |\n| ConvNeXt-Base (full data) | 218 | — | 27.7 | Trained on 100 % data |\n| ConvNeXt-Base (full data) | 243 | — | 26.7 | Trained on 100 % data |\n| ConvNeXt-Base (full data) | 258 | — | 26.6 | Trained on 100 % data |\n\nHyper-parameters tuned: `weight_decay`, `dropout_rate`, `decoder_channels`, `clip_grad_norm`  \nLR restarts: resumed from the previous terminal LR (`1 e-5`) with `CosineAnnealingLR(T_max = 50)`, which proved more stable than warm restarts.\n\n## 5. Ensembling & Post-processing  \n\n**Table&nbsp;2 – Ensemble performance (Public MAE)**  \n\n| Method | Submissions | Public MAE | Δ vs. Best Single |\n|--------|-----------:|----:|------------------:|\n| Simple median | 9 | 23.4 | −2.3 |\n| **Weighted median (final)** | **9** | **23.2** | **−2.5** |\n\n- Median-based methods clearly outperform means for MAE.  \n- Weighted median improved MAE by **0.2** over the simple median.  \n- Type-wise weighted means (after predicting sample type with 99.5 % accuracy) did **not** beat median.  \n- Post-processing is limited to **`clip(1500, 4500)`** (additional –0.002 CV MAE).\n\n## 6. Speed-up Techniques\n| Category | Key settings |\n|----------|--------------|\n| **Compiler / Optimizer** | `torch.compile(mode=\"max-autotune\")`, **Fused AdamW** |\n| **Mixed Precision** | AMP (**bfloat16**) for all ConvNeXt / CAFormer runs |\n| **DataLoader** | `num_workers` tuned, `prefetch_factor` tuned, `pin_memory=True`, `persistent_workers=True` |\n| **Hardware** | Experiments run on **RTX 3090** & **A100**<br>Flash InternImage remained FP32 due to DCNv4 → **≈ 210 min / epoch (A100)** |\n\n“num_workers was experimentally tuned—the value is a sweet spot: too many or too few threads slowed training, so we benchmark-benchmarked different settings and used the fastest one for all runs.”\n\n\n## 7. What Didn’t Work\n\n| What failed | Likely cause / lesson |\n|-------------|-----------------------|\n| Type-wise weights, heavy post-processing | Median already robust; extra weighting brought no gain |\n| Flash InternImage-b | AMP unstable; FP32 made training prohibitively slow, limiting epochs |\n| EVA family | Mishandled window / resize parameters – never converged |\n| Full-data training (100 %) | Slightly worse Public MAE vs. 88 % split, much longer runtime |\n\n## 8. Results & Lessons Learned\n- **Final scores:** Public MAE 23.2 / Private MAE 23.1 (weighted median)  \n- **Ensembling gain** saturated at ≈ 2.5 MAE → future effort should target single-model quality.  \n- LR scheduling and restart strategy had greater impact than backbone size beyond *ConvNeXt-Base*.  \n- Further gains likely lie in synthetic data and higher-resolution training, as hinted by top teams.\n\n---\n\n## References\n- Bartley Baseline Notebook  \n  <https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline>  \n  <https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved>\n- Training / Speed-up Discussion  \n  <https://www.kaggle.com/competitions/waveform-inversion/discussion/583896>\n\n---",
    "3239435": "Thanks for many experimental results! They are very useful, but some results are not obvious. Do you have some idea why?\n\n* Why ConvNeXt-Base (full data) runs are worse than the first ConvNeXt-Base although there ~10% more data? Difference in learning rates?\n* Why ConvNeXt-Large is worse than Base for same 150 epochs? Do you see larger train-val loss gap (overfit)? Some people said larger models were always better in this competition\n\nWeighted median is great! I only thought about median and weighted mean. Is it implemented in well known libraries or in public notebook?",
    "3240376": "Thank you for your comment!\n\n>Why ConvNeXt-Base (full data) runs are worse than the first ConvNeXt-Base although there ~10% more data? Difference in learning rates?\n\nI realize my earlier conclusion that \"full-data didn't help\" was premature because the two runs weren't strictly comparable.\n\nThe key differences we've identified are:\n100% DATA: batch_size = 64, epochs = 300, drop_last_train = True\n88% DATA: batch_size = 32, epochs = 200, drop_last_train = False\n\nBecause the CosineAnnealingLR schedule was tied to *epochs*, the LR in the 100 % run fell much sooner (in step terms) than in the 88 % run—likely the key factor. The bigger batch and `drop_last_train=True` probably played a role as well.\n\n>Why ConvNeXt-Large is worse than Base for same 150 epochs? Do you see larger train-val loss gap (overfit)? Some people said larger models were always better in this competition\n\nI’m not sure of the reason, and I wonder why the Large model stagnated earlier than expected.\nIn my experiments, the Large tends to overfit the training data more than the Base, and its validation MAE stops improving at an earlier stage. To counter this, I adjusted the hyper-parameters for the Large model as shown below, but even with these tweaks the validation MAE still plateaued sooner.\n\n| Model              | Train MAE | Validation MAE                                            | Key Hyper-parameters                                                          |\n| ------------------ | --------- | --------------------------------------------------------- | ----------------------------------------------------------------------------- |\n| **ConvNeXt-Base**  | 19.83   | 25.58                                                   | `drop_path_rate = 0.2`, `weight_decay = 0.03`, `layer_decay = 1.0` (disabled) |\n| **ConvNeXt-Large** | 18.62   | 26.08 <br />(validation MAE plateaus earlier than Base) | `drop_path_rate = 0.4`, `weight_decay = 0.05`, `layer_decay = 0.7`            |\n\n\nI intend to study other participants’ solutions to better understand how to improve the Large model. \nIf you have any ideas, I’d like to hear them.",
    "3240434": "Thanks for more investigation!\n\nI see, I know some strong Kagglers use full-data training, but I was worrying training without val score; e.g. should we change batch size to have same number steps or keep batch size and change learning rate? Your result confirms that full-data training can be difficult and requires experience. In this competition, we have lots of data, so I keep one file (500 data) per family as a \"test set,\" in addition to K-fold validation set. This helped me having \"test score\" in the almost full-data training.",
    "3241187": "I think your point is spot-on. In fact, bumping the training set from 88% to 95%  would probably be enough to see whether the extra data boosts score. Putting too much faith in the Public Leaderboard is risky in many competitions—although in this one we were lucky that the public and private scores were closely aligned.\n\n>Weighted median is great! I only thought about median and weighted mean. Is it implemented in well known libraries or in public notebook?\n\nI missed your last question. Since Polars doesn't have a built-in weighted median, I implemented my own. As far as I know, this approach hasn't been widely used for ensembling in Kaggle notebooks. Median was already giving a much better score than the mean, so I experimented with variations on the median, and the weighted median happened to work well in this case."
  },
  "source": "meta"
}