{
  "id": 579040,
  "title": "The Power of Auxiliary Loss Functions in Minimizing MAE Loss",
  "url": "/competitions/waveform-inversion/discussion/579040",
  "author_name": "",
  "post_date": "2025-05-14T19:57:44.103667800Z",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<p>A key insight from my recent notebook was that adding some extra loss functions can really help the model learn what it was missing.  My next series of experiments all involve dealing with the insights from this notebook: ( 👉 <a href=\"https://www.kaggle.com/code/tpmeli/beyond-mae-depth-curves-spectra-and-residuals\" target=\"_blank\">Beyond MAE: Depth Curves, Spectra, and Residuals</a>.).  <strong>But I was curious - why do we even need extra losses if we are optimizing for MAE to begin with?</strong>  Won't MAE just \"get to the same conclusion\" if we run it long enough?  Well,  not necessarily.  And here's why.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640743%2Febab63930c091f266da06c68309851f7%2FCurveFault_B_diagnostic_grid.png?generation=1747252586518151&amp;alt=media\" alt=\"Some diagnostic plots of Curvefault-B\"></p>\n<h3>Core Intuition for FWI on OpenFWI</h3>\n<p>In the context of Full Waveform Inversion (FWI) using the OpenFWI dataset, optimizing purely for Mean Absolute Error (MAE) is necessary but not sufficient. MAE focuses on minimizing the average prediction error across all depths but does not prioritize where errors critically impact geological interpretations—such as deeper layers, fault structures, or frequency-specific features. Auxiliary loss functions can help address these critical structured errors, improving overall predictive performance and geological fidelity.</p>\n<h3>Relevant Loss Function Enhancements</h3>\n<p>In FWI modeling with the OpenFWI dataset, structured errors frequently emerge at greater depths and specific frequency bands. Using targeted auxiliary losses can mitigate these issues:</p>\n<h4>Depth-weighted MAE</h4>\n<ul>\n<li><strong>Purpose</strong>: Reduces large residuals commonly observed at deeper geological layers.</li>\n<li><strong>Mechanism</strong>: Errors at greater depths receive higher penalty weights, guiding the model to prioritize accuracy in critical deeper sections.</li>\n</ul>\n<h4>Spectral (FFT) Loss</h4>\n<ul>\n<li><strong>Purpose</strong>: Corrects systematic frequency-domain inaccuracies seen in spectral analyses (e.g., bright lines or cross-shaped artifacts in FFT plots).</li>\n<li><strong>Mechanism</strong>: Penalizes amplitude mismatches between predicted and true waveforms in frequency space, improving alignment with real seismic signals.</li>\n</ul>\n<h4>Edge / Gradient (Charbonnier) Loss</h4>\n<ul>\n<li><strong>Purpose</strong>: Enhances the sharpness of geological discontinuities such as fault lines.</li>\n<li><strong>Mechanism</strong>: Penalizes overly smooth gradients, helping the model produce clearer fault definitions and reducing blurring.</li>\n</ul>\n<h4>Huber Loss</h4>\n<ul>\n<li><strong>Purpose</strong>: Manages error distribution by pulling in heavy-tailed prediction errors common in challenging \"Style\" geological families.</li>\n<li><strong>Mechanism</strong>: Combines advantages of L1 and L2 losses—robust near-zero errors and aggressive on larger residuals—effectively handling outlier predictions.</li>\n</ul>\n<h3>Formal Loss Structure</h3>\n<p>The refined loss function tailored for FWI tasks on OpenFWI:</p>\n<p>$$<br>\nL_{\\text{total}} = 1.0 \\times L_{\\text{MAE}} + 0.50 \\times L_{\\text{depth-weighted}} + 0.10 \\times L_{\\text{spectral}} + 0.01 \\times L_{\\text{gradient/edge}}<br>\n$$</p>\n<p><strong>Plain-English Translation:</strong><br>\nThe final loss combines MAE fully with depth-specific, spectral, and gradient-based penalties, each scaled to ensure targeted improvement without overwhelming basic accuracy metrics.</p>\n<h3>Trade-offs Specific to OpenFWI</h3>\n<table>\n<thead>\n<tr>\n<th>Loss Component</th>\n<th>Pros for FWI on OpenFWI dataset</th>\n<th>Cons / Risks</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Depth-weighted MAE</strong></td>\n<td>Improves accuracy in deeper, structurally important geological regions.</td>\n<td>Needs careful tuning; risks worsening shallow-layer accuracy.</td>\n</tr>\n<tr>\n<td><strong>Spectral (FFT) Loss</strong></td>\n<td>Aligns predicted waveform frequencies closely with true seismic signals.</td>\n<td>Computationally intensive; careful selection of penalty scale needed to avoid over-penalizing minor frequency shifts.</td>\n</tr>\n<tr>\n<td><strong>Edge (Gradient) Loss</strong></td>\n<td>Clarifies geological discontinuities like faults and stratigraphic boundaries.</td>\n<td>May amplify noise or small-scale artifacts if weighted too strongly.</td>\n</tr>\n<tr>\n<td><strong>Huber Loss</strong></td>\n<td>Robustly reduces large residual tails common in complex geological styles.</td>\n<td>Adds complexity to gradient calculation; tuning threshold (\\$\\delta\\$) required.</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Explicit Assumption:</strong><br>\nImproved geological accuracy and fidelity are more valuable for real-world utility than merely achieving minimal MAE, especially when evaluating deeper layers and complex geological structures.</p>\n<h3>Practical Recommendations</h3>\n<ol>\n<li><p><strong>Begin with Depth-weighted MAE</strong>:</p>\n<ul>\n<li>Use moderate initial depth weighting to immediately address common deep-layer errors.</li></ul></li>\n<li><p><strong>Introduce Spectral Loss Cautiously</strong>:</p>\n<ul>\n<li>Gradually incorporate spectral loss after validating that frequency-domain errors significantly impact prediction quality.</li></ul></li>\n<li><p><strong>Edge Loss for Fault Clarity</strong>:</p>\n<ul>\n<li>Implement a conservative edge loss initially to enhance geological feature clarity, tuning upward as model stability permits.</li></ul></li>\n<li><p><strong>Monitor Continuously</strong>:</p>\n<ul>\n<li>Regularly evaluate depth-specific error curves, spectral domain analyses, and fault sharpness to iteratively refine and balance loss components.</li></ul></li>\n</ol>\n<p>By strategically applying these auxiliary losses, FWI models trained on the OpenFWI dataset can achieve more accurate, geologically meaningful predictions, addressing the critical structured errors inherent in purely MAE-driven optimization.</p>\n<h2>Reading List</h2>\n<p>Here’s a starter reading list (with plain-language notes) on <strong>auxiliary or alternative loss functions that push optimisation “faster and farther” than a plain MAE / L1 objective</strong>. I’ve grouped them by theme so you can skim to what’s most relevant.</p>\n<hr>\n<h3>1. Frequency-aware / spectral losses — sharper details, quicker convergence</h3>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Core idea (plain words)</th>\n<th>Why it matters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Guided Frequency Loss (GFL)</strong>, ICCV 2023 (<a href=\"https://arxiv.org/abs/2309.15563?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Combines a Charbonnier (robust-L1), Laplacian-pyramid, and gradual-frequency term so the network balances low- and high-frequency energy while it learns.</td>\n<td>Training stabilises sooner and PSNR jumps in image-restoration tasks; the same notion of “teach the model where in the spectrum it’s wrong” maps cleanly to FWI spectra.</td>\n</tr>\n<tr>\n<td><strong>Focal Frequency Loss (FFL)</strong>, ICCV 2021 (<a href=\"https://openaccess.thecvf.com/content/ICCV2021/papers/Jiang_Focal_Frequency_Loss_for_Image_Reconstruction_and_Synthesis_ICCV_2021_paper.pdf?utm_source=chatgpt.com\" target=\"_blank\">CVF Open Access</a>)</td>\n<td>Lets the model <em>adaptively</em> up-weight frequency bins it currently struggles with, down-weighting the easy ones.</td>\n<td>Empirically speeds convergence and yields crisper outputs; conceptually similar to putting a moving spotlight on the hardest frequencies in seismic waveforms.</td>\n</tr>\n<tr>\n<td><strong>ω-FWI: Fourier-metric Full Waveform Inversion</strong>, arXiv 2022 (<a href=\"https://arxiv.org/abs/2205.09234?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Replaces the sample-by-sample misfit with a Fourier-domain power-spectrum distance, giving the optimiser smoother gradients and less cycle-skipping.</td>\n<td>Demonstrates faster recovery of low-wavenumber structure from poor starting models—exactly the “go farther” property we want.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>2. Structure-aware losses — reward “looks right”, not just “averages out”</h3>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Core idea</th>\n<th>Why it matters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>MS_ATpV-FWI</strong>: Multi-scale SSIM + p-Variation, 2025 (<a href=\"https://arxiv.org/html/2504.01695\" target=\"_blank\">arXiv</a>)</td>\n<td>Uses <strong>multi-scale SSIM</strong> (captures structural similarity) and <strong>anisotropic total p-variation</strong> as auxiliary terms.</td>\n<td>Multiscale SSIM gives gradients that focus on phase + amplitude coherence, reducing cycle-skipping; p-Variation keeps faults sharp. Reported to converge where plain L2 or MAE stalls.</td>\n</tr>\n<tr>\n<td><strong>ML-misfit</strong>, arXiv 2020 (<a href=\"https://arxiv.org/abs/2002.03163?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td><em>Learns</em> the misfit itself (meta-learning) so it becomes convex around realistic waveform shifts.</td>\n<td>Shows that a data-driven auxiliary loss can widen the basin of convergence in FWI—fewer restarts, deeper minima.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>3. Optimal-transport &amp; Wasserstein misfits — align whole waveforms, not points</h3>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Core idea</th>\n<th>Why it matters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>W2-FWI (Quadratic Wasserstein)</strong>, Engquist et al. 2016 (<a href=\"https://arxiv.org/abs/1612.05075?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Treats traces as mass distributions and measures the <em>cost to morph one into the other</em>; inherently accounts for phase shifts.</td>\n<td>Strong convexity properties give larger, smoother descent steps—practically fewer iterations to reach a usable model.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>4. Expert-diversity / load-balancing losses (if you use MoE routers)</h3>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Core idea</th>\n<th>Take-away</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Auxiliary-Loss-Free Load Balancing</strong>, 2024 (<a href=\"https://arxiv.org/abs/2408.15664?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Shows classic load-balance penalties can inject noisy gradients; proposes bias-based routing that <em>keeps balance</em> <strong>without</strong> an extra loss term.</td>\n<td>Useful caution: auxiliary terms help, but ill-designed ones can hinder optimisation speed.</td>\n</tr>\n<tr>\n<td><strong>Demons in the Detail: Revisiting Load-Balancing Loss</strong>, 2025 (<a href=\"https://arxiv.org/abs/2501.11873?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Demonstrates that <em>how</em> you compute the balance loss (micro- vs global-batch) dramatically affects expert specialisation and final quality.</td>\n<td>Highlights that even tiny auxiliary losses need thoughtful implementation to really “go farther.”</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>5. Honorable mentions &amp; cross-links</h3>\n<ul>\n<li><strong>Guided Frequency-aware SSIM</strong>, <strong>Amplitude-based misfit</strong> models, and other <strong>multi-scale FWI</strong> variants extend the same philosophy: combine <em>where</em> it’s wrong (depth, frequency, edges) with <em>how much</em> it’s wrong (MAE/L2) for faster, deeper convergence.</li>\n</ul>\n<hr>\n<h4>TL;DR</h4>\n<p>Yes—there’s a growing body of work showing that <em>well-chosen</em> auxiliary losses (frequency, structural, transport, or balance-oriented) can <strong>speed up training, escape shallow minima, and land on solutions MAE alone cannot reach</strong>.  These papers offer concrete design patterns you can port straight into your OpenFWI experiments.</p>",
  "messages": [
    {
      "id": "3202071",
      "postDate": "05/14/2025 19:57:44",
      "content": "<p>A key insight from my recent notebook was that adding some extra loss functions can really help the model learn what it was missing.  My next series of experiments all involve dealing with the insights from this notebook: ( 👉 <a href=\"https://www.kaggle.com/code/tpmeli/beyond-mae-depth-curves-spectra-and-residuals\" target=\"_blank\">Beyond MAE: Depth Curves, Spectra, and Residuals</a>.).  <strong>But I was curious - why do we even need extra losses if we are optimizing for MAE to begin with?</strong>  Won't MAE just \"get to the same conclusion\" if we run it long enough?  Well,  not necessarily.  And here's why.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640743%2Febab63930c091f266da06c68309851f7%2FCurveFault_B_diagnostic_grid.png?generation=1747252586518151&amp;alt=media\" alt=\"Some diagnostic plots of Curvefault-B\"></p>\n<h3>Core Intuition for FWI on OpenFWI</h3>\n<p>In the context of Full Waveform Inversion (FWI) using the OpenFWI dataset, optimizing purely for Mean Absolute Error (MAE) is necessary but not sufficient. MAE focuses on minimizing the average prediction error across all depths but does not prioritize where errors critically impact geological interpretations—such as deeper layers, fault structures, or frequency-specific features. Auxiliary loss functions can help address these critical structured errors, improving overall predictive performance and geological fidelity.</p>\n<h3>Relevant Loss Function Enhancements</h3>\n<p>In FWI modeling with the OpenFWI dataset, structured errors frequently emerge at greater depths and specific frequency bands. Using targeted auxiliary losses can mitigate these issues:</p>\n<h4>Depth-weighted MAE</h4>\n<ul>\n<li><strong>Purpose</strong>: Reduces large residuals commonly observed at deeper geological layers.</li>\n<li><strong>Mechanism</strong>: Errors at greater depths receive higher penalty weights, guiding the model to prioritize accuracy in critical deeper sections.</li>\n</ul>\n<h4>Spectral (FFT) Loss</h4>\n<ul>\n<li><strong>Purpose</strong>: Corrects systematic frequency-domain inaccuracies seen in spectral analyses (e.g., bright lines or cross-shaped artifacts in FFT plots).</li>\n<li><strong>Mechanism</strong>: Penalizes amplitude mismatches between predicted and true waveforms in frequency space, improving alignment with real seismic signals.</li>\n</ul>\n<h4>Edge / Gradient (Charbonnier) Loss</h4>\n<ul>\n<li><strong>Purpose</strong>: Enhances the sharpness of geological discontinuities such as fault lines.</li>\n<li><strong>Mechanism</strong>: Penalizes overly smooth gradients, helping the model produce clearer fault definitions and reducing blurring.</li>\n</ul>\n<h4>Huber Loss</h4>\n<ul>\n<li><strong>Purpose</strong>: Manages error distribution by pulling in heavy-tailed prediction errors common in challenging \"Style\" geological families.</li>\n<li><strong>Mechanism</strong>: Combines advantages of L1 and L2 losses—robust near-zero errors and aggressive on larger residuals—effectively handling outlier predictions.</li>\n</ul>\n<h3>Formal Loss Structure</h3>\n<p>The refined loss function tailored for FWI tasks on OpenFWI:</p>\n<p>$$<br>\nL_{\\text{total}} = 1.0 \\times L_{\\text{MAE}} + 0.50 \\times L_{\\text{depth-weighted}} + 0.10 \\times L_{\\text{spectral}} + 0.01 \\times L_{\\text{gradient/edge}}<br>\n$$</p>\n<p><strong>Plain-English Translation:</strong><br>\nThe final loss combines MAE fully with depth-specific, spectral, and gradient-based penalties, each scaled to ensure targeted improvement without overwhelming basic accuracy metrics.</p>\n<h3>Trade-offs Specific to OpenFWI</h3>\n<table>\n<thead>\n<tr>\n<th>Loss Component</th>\n<th>Pros for FWI on OpenFWI dataset</th>\n<th>Cons / Risks</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Depth-weighted MAE</strong></td>\n<td>Improves accuracy in deeper, structurally important geological regions.</td>\n<td>Needs careful tuning; risks worsening shallow-layer accuracy.</td>\n</tr>\n<tr>\n<td><strong>Spectral (FFT) Loss</strong></td>\n<td>Aligns predicted waveform frequencies closely with true seismic signals.</td>\n<td>Computationally intensive; careful selection of penalty scale needed to avoid over-penalizing minor frequency shifts.</td>\n</tr>\n<tr>\n<td><strong>Edge (Gradient) Loss</strong></td>\n<td>Clarifies geological discontinuities like faults and stratigraphic boundaries.</td>\n<td>May amplify noise or small-scale artifacts if weighted too strongly.</td>\n</tr>\n<tr>\n<td><strong>Huber Loss</strong></td>\n<td>Robustly reduces large residual tails common in complex geological styles.</td>\n<td>Adds complexity to gradient calculation; tuning threshold (\\$\\delta\\$) required.</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Explicit Assumption:</strong><br>\nImproved geological accuracy and fidelity are more valuable for real-world utility than merely achieving minimal MAE, especially when evaluating deeper layers and complex geological structures.</p>\n<h3>Practical Recommendations</h3>\n<ol>\n<li><p><strong>Begin with Depth-weighted MAE</strong>:</p>\n<ul>\n<li>Use moderate initial depth weighting to immediately address common deep-layer errors.</li></ul></li>\n<li><p><strong>Introduce Spectral Loss Cautiously</strong>:</p>\n<ul>\n<li>Gradually incorporate spectral loss after validating that frequency-domain errors significantly impact prediction quality.</li></ul></li>\n<li><p><strong>Edge Loss for Fault Clarity</strong>:</p>\n<ul>\n<li>Implement a conservative edge loss initially to enhance geological feature clarity, tuning upward as model stability permits.</li></ul></li>\n<li><p><strong>Monitor Continuously</strong>:</p>\n<ul>\n<li>Regularly evaluate depth-specific error curves, spectral domain analyses, and fault sharpness to iteratively refine and balance loss components.</li></ul></li>\n</ol>\n<p>By strategically applying these auxiliary losses, FWI models trained on the OpenFWI dataset can achieve more accurate, geologically meaningful predictions, addressing the critical structured errors inherent in purely MAE-driven optimization.</p>\n<h2>Reading List</h2>\n<p>Here’s a starter reading list (with plain-language notes) on <strong>auxiliary or alternative loss functions that push optimisation “faster and farther” than a plain MAE / L1 objective</strong>. I’ve grouped them by theme so you can skim to what’s most relevant.</p>\n<hr>\n<h3>1. Frequency-aware / spectral losses — sharper details, quicker convergence</h3>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Core idea (plain words)</th>\n<th>Why it matters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Guided Frequency Loss (GFL)</strong>, ICCV 2023 (<a href=\"https://arxiv.org/abs/2309.15563?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Combines a Charbonnier (robust-L1), Laplacian-pyramid, and gradual-frequency term so the network balances low- and high-frequency energy while it learns.</td>\n<td>Training stabilises sooner and PSNR jumps in image-restoration tasks; the same notion of “teach the model where in the spectrum it’s wrong” maps cleanly to FWI spectra.</td>\n</tr>\n<tr>\n<td><strong>Focal Frequency Loss (FFL)</strong>, ICCV 2021 (<a href=\"https://openaccess.thecvf.com/content/ICCV2021/papers/Jiang_Focal_Frequency_Loss_for_Image_Reconstruction_and_Synthesis_ICCV_2021_paper.pdf?utm_source=chatgpt.com\" target=\"_blank\">CVF Open Access</a>)</td>\n<td>Lets the model <em>adaptively</em> up-weight frequency bins it currently struggles with, down-weighting the easy ones.</td>\n<td>Empirically speeds convergence and yields crisper outputs; conceptually similar to putting a moving spotlight on the hardest frequencies in seismic waveforms.</td>\n</tr>\n<tr>\n<td><strong>ω-FWI: Fourier-metric Full Waveform Inversion</strong>, arXiv 2022 (<a href=\"https://arxiv.org/abs/2205.09234?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Replaces the sample-by-sample misfit with a Fourier-domain power-spectrum distance, giving the optimiser smoother gradients and less cycle-skipping.</td>\n<td>Demonstrates faster recovery of low-wavenumber structure from poor starting models—exactly the “go farther” property we want.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>2. Structure-aware losses — reward “looks right”, not just “averages out”</h3>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Core idea</th>\n<th>Why it matters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>MS_ATpV-FWI</strong>: Multi-scale SSIM + p-Variation, 2025 (<a href=\"https://arxiv.org/html/2504.01695\" target=\"_blank\">arXiv</a>)</td>\n<td>Uses <strong>multi-scale SSIM</strong> (captures structural similarity) and <strong>anisotropic total p-variation</strong> as auxiliary terms.</td>\n<td>Multiscale SSIM gives gradients that focus on phase + amplitude coherence, reducing cycle-skipping; p-Variation keeps faults sharp. Reported to converge where plain L2 or MAE stalls.</td>\n</tr>\n<tr>\n<td><strong>ML-misfit</strong>, arXiv 2020 (<a href=\"https://arxiv.org/abs/2002.03163?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td><em>Learns</em> the misfit itself (meta-learning) so it becomes convex around realistic waveform shifts.</td>\n<td>Shows that a data-driven auxiliary loss can widen the basin of convergence in FWI—fewer restarts, deeper minima.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>3. Optimal-transport &amp; Wasserstein misfits — align whole waveforms, not points</h3>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Core idea</th>\n<th>Why it matters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>W2-FWI (Quadratic Wasserstein)</strong>, Engquist et al. 2016 (<a href=\"https://arxiv.org/abs/1612.05075?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Treats traces as mass distributions and measures the <em>cost to morph one into the other</em>; inherently accounts for phase shifts.</td>\n<td>Strong convexity properties give larger, smoother descent steps—practically fewer iterations to reach a usable model.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>4. Expert-diversity / load-balancing losses (if you use MoE routers)</h3>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Core idea</th>\n<th>Take-away</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Auxiliary-Loss-Free Load Balancing</strong>, 2024 (<a href=\"https://arxiv.org/abs/2408.15664?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Shows classic load-balance penalties can inject noisy gradients; proposes bias-based routing that <em>keeps balance</em> <strong>without</strong> an extra loss term.</td>\n<td>Useful caution: auxiliary terms help, but ill-designed ones can hinder optimisation speed.</td>\n</tr>\n<tr>\n<td><strong>Demons in the Detail: Revisiting Load-Balancing Loss</strong>, 2025 (<a href=\"https://arxiv.org/abs/2501.11873?utm_source=chatgpt.com\" target=\"_blank\">arXiv</a>)</td>\n<td>Demonstrates that <em>how</em> you compute the balance loss (micro- vs global-batch) dramatically affects expert specialisation and final quality.</td>\n<td>Highlights that even tiny auxiliary losses need thoughtful implementation to really “go farther.”</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>5. Honorable mentions &amp; cross-links</h3>\n<ul>\n<li><strong>Guided Frequency-aware SSIM</strong>, <strong>Amplitude-based misfit</strong> models, and other <strong>multi-scale FWI</strong> variants extend the same philosophy: combine <em>where</em> it’s wrong (depth, frequency, edges) with <em>how much</em> it’s wrong (MAE/L2) for faster, deeper convergence.</li>\n</ul>\n<hr>\n<h4>TL;DR</h4>\n<p>Yes—there’s a growing body of work showing that <em>well-chosen</em> auxiliary losses (frequency, structural, transport, or balance-oriented) can <strong>speed up training, escape shallow minima, and land on solutions MAE alone cannot reach</strong>.  These papers offer concrete design patterns you can port straight into your OpenFWI experiments.</p>",
      "rawMarkdown": "A key insight from my recent notebook was that adding some extra loss functions can really help the model learn what it was missing.  My next series of experiments all involve dealing with the insights from this notebook: ( 👉 [Beyond MAE: Depth Curves, Spectra, and Residuals](https://www.kaggle.com/code/tpmeli/beyond-mae-depth-curves-spectra-and-residuals).).  **But I was curious - why do we even need extra losses if we are optimizing for MAE to begin with?**  Won't MAE just \"get to the same conclusion\" if we run it long enough?  Well,  not necessarily.  And here's why.\n\n![Some diagnostic plots of Curvefault-B](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640743%2Febab63930c091f266da06c68309851f7%2FCurveFault_B_diagnostic_grid.png?generation=1747252586518151&alt=media)\n\n### Core Intuition for FWI on OpenFWI\n\nIn the context of Full Waveform Inversion (FWI) using the OpenFWI dataset, optimizing purely for Mean Absolute Error (MAE) is necessary but not sufficient. MAE focuses on minimizing the average prediction error across all depths but does not prioritize where errors critically impact geological interpretations—such as deeper layers, fault structures, or frequency-specific features. Auxiliary loss functions can help address these critical structured errors, improving overall predictive performance and geological fidelity.\n\n### Relevant Loss Function Enhancements\n\nIn FWI modeling with the OpenFWI dataset, structured errors frequently emerge at greater depths and specific frequency bands. Using targeted auxiliary losses can mitigate these issues:\n\n#### Depth-weighted MAE\n\n* **Purpose**: Reduces large residuals commonly observed at deeper geological layers.\n* **Mechanism**: Errors at greater depths receive higher penalty weights, guiding the model to prioritize accuracy in critical deeper sections.\n\n#### Spectral (FFT) Loss\n\n* **Purpose**: Corrects systematic frequency-domain inaccuracies seen in spectral analyses (e.g., bright lines or cross-shaped artifacts in FFT plots).\n* **Mechanism**: Penalizes amplitude mismatches between predicted and true waveforms in frequency space, improving alignment with real seismic signals.\n\n#### Edge / Gradient (Charbonnier) Loss\n\n* **Purpose**: Enhances the sharpness of geological discontinuities such as fault lines.\n* **Mechanism**: Penalizes overly smooth gradients, helping the model produce clearer fault definitions and reducing blurring.\n\n#### Huber Loss\n\n* **Purpose**: Manages error distribution by pulling in heavy-tailed prediction errors common in challenging \"Style\" geological families.\n* **Mechanism**: Combines advantages of L1 and L2 losses—robust near-zero errors and aggressive on larger residuals—effectively handling outlier predictions.\n\n### Formal Loss Structure\n\nThe refined loss function tailored for FWI tasks on OpenFWI:\n\n$$\nL_{\\text{total}} = 1.0 \\times L_{\\text{MAE}} + 0.50 \\times L_{\\text{depth-weighted}} + 0.10 \\times L_{\\text{spectral}} + 0.01 \\times L_{\\text{gradient/edge}}\n$$\n\n**Plain-English Translation:**\nThe final loss combines MAE fully with depth-specific, spectral, and gradient-based penalties, each scaled to ensure targeted improvement without overwhelming basic accuracy metrics.\n\n### Trade-offs Specific to OpenFWI\n\n| Loss Component           | Pros for FWI on OpenFWI dataset                                                | Cons / Risks                                                                                                          |\n| ------------------------ | ------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------- |\n| **Depth-weighted MAE**   | Improves accuracy in deeper, structurally important geological regions.        | Needs careful tuning; risks worsening shallow-layer accuracy.                                                         |\n| **Spectral (FFT) Loss**  | Aligns predicted waveform frequencies closely with true seismic signals.       | Computationally intensive; careful selection of penalty scale needed to avoid over-penalizing minor frequency shifts. |\n| **Edge (Gradient) Loss** | Clarifies geological discontinuities like faults and stratigraphic boundaries. | May amplify noise or small-scale artifacts if weighted too strongly.                                                  |\n| **Huber Loss**           | Robustly reduces large residual tails common in complex geological styles.     | Adds complexity to gradient calculation; tuning threshold (\\$\\delta\\$) required.                                      |\n\n**Explicit Assumption:**\nImproved geological accuracy and fidelity are more valuable for real-world utility than merely achieving minimal MAE, especially when evaluating deeper layers and complex geological structures.\n\n### Practical Recommendations\n\n1. **Begin with Depth-weighted MAE**:\n\n   * Use moderate initial depth weighting to immediately address common deep-layer errors.\n\n2. **Introduce Spectral Loss Cautiously**:\n\n   * Gradually incorporate spectral loss after validating that frequency-domain errors significantly impact prediction quality.\n\n3. **Edge Loss for Fault Clarity**:\n\n   * Implement a conservative edge loss initially to enhance geological feature clarity, tuning upward as model stability permits.\n\n4. **Monitor Continuously**:\n\n   * Regularly evaluate depth-specific error curves, spectral domain analyses, and fault sharpness to iteratively refine and balance loss components.\n\nBy strategically applying these auxiliary losses, FWI models trained on the OpenFWI dataset can achieve more accurate, geologically meaningful predictions, addressing the critical structured errors inherent in purely MAE-driven optimization.\n\n## Reading List\n\nHere’s a starter reading list (with plain-language notes) on **auxiliary or alternative loss functions that push optimisation “faster and farther” than a plain MAE / L1 objective**. I’ve grouped them by theme so you can skim to what’s most relevant.\n\n---\n\n### 1. Frequency-aware / spectral losses — sharper details, quicker convergence\n\n| Paper                                                                      | Core idea (plain words)                                                                                                                                   | Why it matters                                                                                                                                                           |\n| -------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| **Guided Frequency Loss (GFL)**, ICCV 2023 ([arXiv][1])                    | Combines a Charbonnier (robust-L1), Laplacian-pyramid, and gradual-frequency term so the network balances low- and high-frequency energy while it learns. | Training stabilises sooner and PSNR jumps in image-restoration tasks; the same notion of “teach the model where in the spectrum it’s wrong” maps cleanly to FWI spectra. |\n| **Focal Frequency Loss (FFL)**, ICCV 2021 ([CVF Open Access][2])           | Lets the model *adaptively* up-weight frequency bins it currently struggles with, down-weighting the easy ones.                                           | Empirically speeds convergence and yields crisper outputs; conceptually similar to putting a moving spotlight on the hardest frequencies in seismic waveforms.           |\n| **ω-FWI: Fourier-metric Full Waveform Inversion**, arXiv 2022 ([arXiv][3]) | Replaces the sample-by-sample misfit with a Fourier-domain power-spectrum distance, giving the optimiser smoother gradients and less cycle-skipping.      | Demonstrates faster recovery of low-wavenumber structure from poor starting models—exactly the “go farther” property we want.                                            |\n\n---\n\n### 2. Structure-aware losses — reward “looks right”, not just “averages out”\n\n| Paper                                                               | Core idea                                                                                                            | Why it matters                                                                                                                                                                         |\n| ------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| **MS\\_ATpV-FWI**: Multi-scale SSIM + p-Variation, 2025 ([arXiv][4]) | Uses **multi-scale SSIM** (captures structural similarity) and **anisotropic total p-variation** as auxiliary terms. | Multiscale SSIM gives gradients that focus on phase + amplitude coherence, reducing cycle-skipping; p-Variation keeps faults sharp. Reported to converge where plain L2 or MAE stalls. |\n| **ML-misfit**, arXiv 2020 ([arXiv][5])                              | *Learns* the misfit itself (meta-learning) so it becomes convex around realistic waveform shifts.                    | Shows that a data-driven auxiliary loss can widen the basin of convergence in FWI—fewer restarts, deeper minima.                                                                       |\n\n---\n\n### 3. Optimal-transport & Wasserstein misfits — align whole waveforms, not points\n\n| Paper                                                                 | Core idea                                                                                                                      | Why it matters                                                                                                        |\n| --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------- |\n| **W2-FWI (Quadratic Wasserstein)**, Engquist et al. 2016 ([arXiv][6]) | Treats traces as mass distributions and measures the *cost to morph one into the other*; inherently accounts for phase shifts. | Strong convexity properties give larger, smoother descent steps—practically fewer iterations to reach a usable model. |\n\n---\n\n### 4. Expert-diversity / load-balancing losses (if you use MoE routers)\n\n| Paper                                                                       | Core idea                                                                                                                                         | Take-away                                                                                         |\n| --------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- |\n| **Auxiliary-Loss-Free Load Balancing**, 2024 ([arXiv][7])                   | Shows classic load-balance penalties can inject noisy gradients; proposes bias-based routing that *keeps balance* **without** an extra loss term. | Useful caution: auxiliary terms help, but ill-designed ones can hinder optimisation speed.        |\n| **Demons in the Detail: Revisiting Load-Balancing Loss**, 2025 ([arXiv][8]) | Demonstrates that *how* you compute the balance loss (micro- vs global-batch) dramatically affects expert specialisation and final quality.       | Highlights that even tiny auxiliary losses need thoughtful implementation to really “go farther.” |\n\n---\n\n### 5. Honorable mentions & cross-links\n\n* **Guided Frequency-aware SSIM**, **Amplitude-based misfit** models, and other **multi-scale FWI** variants extend the same philosophy: combine *where* it’s wrong (depth, frequency, edges) with *how much* it’s wrong (MAE/L2) for faster, deeper convergence.\n\n---\n\n#### TL;DR\n\nYes—there’s a growing body of work showing that *well-chosen* auxiliary losses (frequency, structural, transport, or balance-oriented) can **speed up training, escape shallow minima, and land on solutions MAE alone cannot reach**.  These papers offer concrete design patterns you can port straight into your OpenFWI experiments.\n\n[1]: https://arxiv.org/abs/2309.15563?utm_source=chatgpt.com \"Guided Frequency Loss for Image Restoration\"\n[2]: https://openaccess.thecvf.com/content/ICCV2021/papers/Jiang_Focal_Frequency_Loss_for_Image_Reconstruction_and_Synthesis_ICCV_2021_paper.pdf?utm_source=chatgpt.com \"[PDF] Focal Frequency Loss for Image Reconstruction and Synthesis\"\n[3]: https://arxiv.org/abs/2205.09234?utm_source=chatgpt.com \"$ω$-FWI: Robust full-waveform inversion with Fourier-based metric\"\n[4]: https://arxiv.org/html/2504.01695 \"MS_ATpV-FWI: Full Waveform Inversion based on Multi-scale Structural Similarity Index Measure and Anisotropic Total p-Variation Regularization\"\n[5]: https://arxiv.org/abs/2002.03163?utm_source=chatgpt.com \"ML-misfit: Learning a robust misfit function for full-waveform inversion using machine learning\"\n[6]: https://arxiv.org/abs/1612.05075?utm_source=chatgpt.com \"Application of Optimal Transport and the Quadratic Wasserstein Metric to Full-Waveform Inversion\"\n[7]: https://arxiv.org/abs/2408.15664?utm_source=chatgpt.com \"Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts\"\n[8]: https://arxiv.org/abs/2501.11873?utm_source=chatgpt.com \"Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models\"",
      "votes": null
    },
    {
      "id": "3202129",
      "postDate": "05/14/2025 21:52:24",
      "content": "<p>There is a lot of text and figures both in this post and the notebook. However I did not find the most important thing- ablation study scores or something similar that proves:</p>\n<ol>\n<li>That auxiliary loss helps.  </li>\n<li>By how much compared to baseline…</li>\n</ol>",
      "rawMarkdown": "There is a lot of text and figures both in this post and the notebook. However I did not find the most important thing- ablation study scores or something similar that proves:\n1. That auxiliary loss helps.  \n2. By how much compared to baseline...",
      "votes": null
    },
    {
      "id": "3202160",
      "postDate": "05/15/2025 00:17:19",
      "content": "<p>Update:   <strong><em>The notebook goes over 100+ possible experimental combinations that I could not possibly do alone.</em></strong>  This will take time and is better performed by the community as a whole.   Your comment misses the purpose of this discussion and my notebook.</p>\n<p>On my own, so far, I have not had any luck with auxiliary loss functions on any of the models I've used… but I have not extensively tested them, nor do I have the means to.  <strong><em>I only have one 3060, and running any one experiment takes 2-5 days</em></strong>, and that tends to only get me to epoch 30-40 with the whole dataset… It really isn't enough to come to any conclusion at all.  </p>\n<p>Auxiliary loss functions also seem to make the loss surface potentially more complicated, but <strong><em>don't overfit on this very small result of n=1</em></strong>.  Results are super sensitive the hyper-parameters chosen, the weight values given to the loss functions, and the learning rate isn't the same when you use auxiliary loss functions etc.  <strong><em>This is precisely why I was hoping we'd have many people experimenting instead of just me.   I hope others play with these ideas and share their results too.</em></strong>  The loss surface can become a bit more complicated with auxiliary loss functions, but the real problem is just low compute.</p>\n<p>The real question is whether there is an area of the loss surface that is lower with auxiliary losses, and we just don't know that.  The research I cite suggests this may be the case.</p>\n<p>Current ideas are to realize that the depth based loss might need to be non-linear.  Spectral loss does correlate well with val loss, so that seems more promising.</p>\n<p>--Older message--</p>\n<p>My current experiments are with different NN architectures, so results from them won't necessarily transfer to an ablation study of the NN architectures I wrote about above (which the community seems to be mostly using).  But this is besides the point.</p>\n<p>The purpose of my posts is to clarify possible high-yield directions and experimental priorities to <strong><em>empower others (hundreds of Kagglers) to test more things than I could possibly test alone</em></strong>.  </p>\n<p>The notebook I've shared represents many hours of detailed research.  I wouldn't have done that unless I thought it was very valuable.  </p>\n<p>So I disagree with the claim that the results alone are the \"most important thing.\"   <strong><em>Just like a data generating process is richer than the specific data points it outputs, I believe a series of research directions are far richer and deeper than any of the specific and temporary results they generate.</em></strong></p>\n<p>Results will come, but be patient.  It turns out that I do have a life outside of Kaggle ;).</p>\n<p>If you're eager for quicker results from an ablation study, I warmly encourage you to perform one yourself and contribute your findings. That would be a meaningful addition to our collective effort, and clearly <strong><em>you</em></strong> would find it more meaningful than my \"text and figures\" 🤣</p>",
      "rawMarkdown": "Update:   ***The notebook goes over 100+ possible experimental combinations that I could not possibly do alone.***  This will take time and is better performed by the community as a whole.   Your comment misses the purpose of this discussion and my notebook.\n\nOn my own, so far, I have not had any luck with auxiliary loss functions on any of the models I've used... but I have not extensively tested them, nor do I have the means to.  ***I only have one 3060, and running any one experiment takes 2-5 days***, and that tends to only get me to epoch 30-40 with the whole dataset... It really isn't enough to come to any conclusion at all.  \n\nAuxiliary loss functions also seem to make the loss surface potentially more complicated, but ***don't overfit on this very small result of n=1***.  Results are super sensitive the hyper-parameters chosen, the weight values given to the loss functions, and the learning rate isn't the same when you use auxiliary loss functions etc.  ***This is precisely why I was hoping we'd have many people experimenting instead of just me.   I hope others play with these ideas and share their results too.***  The loss surface can become a bit more complicated with auxiliary loss functions, but the real problem is just low compute.\n\nThe real question is whether there is an area of the loss surface that is lower with auxiliary losses, and we just don't know that.  The research I cite suggests this may be the case.\n\nCurrent ideas are to realize that the depth based loss might need to be non-linear.  Spectral loss does correlate well with val loss, so that seems more promising.\n\n--Older message--\n\nMy current experiments are with different NN architectures, so results from them won't necessarily transfer to an ablation study of the NN architectures I wrote about above (which the community seems to be mostly using).  But this is besides the point.\n\nThe purpose of my posts is to clarify possible high-yield directions and experimental priorities to ***empower others (hundreds of Kagglers) to test more things than I could possibly test alone***.  \n\nThe notebook I've shared represents many hours of detailed research.  I wouldn't have done that unless I thought it was very valuable.  \n\nSo I disagree with the claim that the results alone are the \"most important thing.\"   ***Just like a data generating process is richer than the specific data points it outputs, I believe a series of research directions are far richer and deeper than any of the specific and temporary results they generate.***\n\nResults will come, but be patient.  It turns out that I do have a life outside of Kaggle ;).\n\nIf you're eager for quicker results from an ablation study, I warmly encourage you to perform one yourself and contribute your findings. That would be a meaningful addition to our collective effort, and clearly ***you*** would find it more meaningful than my \"text and figures\" 🤣",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3202129,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "05/14/2025 21:52:24",
      "content": "<p>There is a lot of text and figures both in this post and the notebook. However I did not find the most important thing- ablation study scores or something similar that proves:</p>\n<ol>\n<li>That auxiliary loss helps.  </li>\n<li>By how much compared to baseline…</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 3202160,
          "author_name": "tpmeli",
          "author_url": "",
          "post_date": "05/15/2025 00:17:19",
          "content": "<p>Update:   <strong><em>The notebook goes over 100+ possible experimental combinations that I could not possibly do alone.</em></strong>  This will take time and is better performed by the community as a whole.   Your comment misses the purpose of this discussion and my notebook.</p>\n<p>On my own, so far, I have not had any luck with auxiliary loss functions on any of the models I've used… but I have not extensively tested them, nor do I have the means to.  <strong><em>I only have one 3060, and running any one experiment takes 2-5 days</em></strong>, and that tends to only get me to epoch 30-40 with the whole dataset… It really isn't enough to come to any conclusion at all.  </p>\n<p>Auxiliary loss functions also seem to make the loss surface potentially more complicated, but <strong><em>don't overfit on this very small result of n=1</em></strong>.  Results are super sensitive the hyper-parameters chosen, the weight values given to the loss functions, and the learning rate isn't the same when you use auxiliary loss functions etc.  <strong><em>This is precisely why I was hoping we'd have many people experimenting instead of just me.   I hope others play with these ideas and share their results too.</em></strong>  The loss surface can become a bit more complicated with auxiliary loss functions, but the real problem is just low compute.</p>\n<p>The real question is whether there is an area of the loss surface that is lower with auxiliary losses, and we just don't know that.  The research I cite suggests this may be the case.</p>\n<p>Current ideas are to realize that the depth based loss might need to be non-linear.  Spectral loss does correlate well with val loss, so that seems more promising.</p>\n<p>--Older message--</p>\n<p>My current experiments are with different NN architectures, so results from them won't necessarily transfer to an ablation study of the NN architectures I wrote about above (which the community seems to be mostly using).  But this is besides the point.</p>\n<p>The purpose of my posts is to clarify possible high-yield directions and experimental priorities to <strong><em>empower others (hundreds of Kagglers) to test more things than I could possibly test alone</em></strong>.  </p>\n<p>The notebook I've shared represents many hours of detailed research.  I wouldn't have done that unless I thought it was very valuable.  </p>\n<p>So I disagree with the claim that the results alone are the \"most important thing.\"   <strong><em>Just like a data generating process is richer than the specific data points it outputs, I believe a series of research directions are far richer and deeper than any of the specific and temporary results they generate.</em></strong></p>\n<p>Results will come, but be patient.  It turns out that I do have a life outside of Kaggle ;).</p>\n<p>If you're eager for quicker results from an ablation study, I warmly encourage you to perform one yourself and contribute your findings. That would be a meaningful addition to our collective effort, and clearly <strong><em>you</em></strong> would find it more meaningful than my \"text and figures\" 🤣</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3202071": "A key insight from my recent notebook was that adding some extra loss functions can really help the model learn what it was missing.  My next series of experiments all involve dealing with the insights from this notebook: ( 👉 [Beyond MAE: Depth Curves, Spectra, and Residuals](https://www.kaggle.com/code/tpmeli/beyond-mae-depth-curves-spectra-and-residuals).).  **But I was curious - why do we even need extra losses if we are optimizing for MAE to begin with?**  Won't MAE just \"get to the same conclusion\" if we run it long enough?  Well,  not necessarily.  And here's why.\n\n![Some diagnostic plots of Curvefault-B](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640743%2Febab63930c091f266da06c68309851f7%2FCurveFault_B_diagnostic_grid.png?generation=1747252586518151&alt=media)\n\n### Core Intuition for FWI on OpenFWI\n\nIn the context of Full Waveform Inversion (FWI) using the OpenFWI dataset, optimizing purely for Mean Absolute Error (MAE) is necessary but not sufficient. MAE focuses on minimizing the average prediction error across all depths but does not prioritize where errors critically impact geological interpretations—such as deeper layers, fault structures, or frequency-specific features. Auxiliary loss functions can help address these critical structured errors, improving overall predictive performance and geological fidelity.\n\n### Relevant Loss Function Enhancements\n\nIn FWI modeling with the OpenFWI dataset, structured errors frequently emerge at greater depths and specific frequency bands. Using targeted auxiliary losses can mitigate these issues:\n\n#### Depth-weighted MAE\n\n* **Purpose**: Reduces large residuals commonly observed at deeper geological layers.\n* **Mechanism**: Errors at greater depths receive higher penalty weights, guiding the model to prioritize accuracy in critical deeper sections.\n\n#### Spectral (FFT) Loss\n\n* **Purpose**: Corrects systematic frequency-domain inaccuracies seen in spectral analyses (e.g., bright lines or cross-shaped artifacts in FFT plots).\n* **Mechanism**: Penalizes amplitude mismatches between predicted and true waveforms in frequency space, improving alignment with real seismic signals.\n\n#### Edge / Gradient (Charbonnier) Loss\n\n* **Purpose**: Enhances the sharpness of geological discontinuities such as fault lines.\n* **Mechanism**: Penalizes overly smooth gradients, helping the model produce clearer fault definitions and reducing blurring.\n\n#### Huber Loss\n\n* **Purpose**: Manages error distribution by pulling in heavy-tailed prediction errors common in challenging \"Style\" geological families.\n* **Mechanism**: Combines advantages of L1 and L2 losses—robust near-zero errors and aggressive on larger residuals—effectively handling outlier predictions.\n\n### Formal Loss Structure\n\nThe refined loss function tailored for FWI tasks on OpenFWI:\n\n$$\nL_{\\text{total}} = 1.0 \\times L_{\\text{MAE}} + 0.50 \\times L_{\\text{depth-weighted}} + 0.10 \\times L_{\\text{spectral}} + 0.01 \\times L_{\\text{gradient/edge}}\n$$\n\n**Plain-English Translation:**\nThe final loss combines MAE fully with depth-specific, spectral, and gradient-based penalties, each scaled to ensure targeted improvement without overwhelming basic accuracy metrics.\n\n### Trade-offs Specific to OpenFWI\n\n| Loss Component           | Pros for FWI on OpenFWI dataset                                                | Cons / Risks                                                                                                          |\n| ------------------------ | ------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------- |\n| **Depth-weighted MAE**   | Improves accuracy in deeper, structurally important geological regions.        | Needs careful tuning; risks worsening shallow-layer accuracy.                                                         |\n| **Spectral (FFT) Loss**  | Aligns predicted waveform frequencies closely with true seismic signals.       | Computationally intensive; careful selection of penalty scale needed to avoid over-penalizing minor frequency shifts. |\n| **Edge (Gradient) Loss** | Clarifies geological discontinuities like faults and stratigraphic boundaries. | May amplify noise or small-scale artifacts if weighted too strongly.                                                  |\n| **Huber Loss**           | Robustly reduces large residual tails common in complex geological styles.     | Adds complexity to gradient calculation; tuning threshold (\\$\\delta\\$) required.                                      |\n\n**Explicit Assumption:**\nImproved geological accuracy and fidelity are more valuable for real-world utility than merely achieving minimal MAE, especially when evaluating deeper layers and complex geological structures.\n\n### Practical Recommendations\n\n1. **Begin with Depth-weighted MAE**:\n\n   * Use moderate initial depth weighting to immediately address common deep-layer errors.\n\n2. **Introduce Spectral Loss Cautiously**:\n\n   * Gradually incorporate spectral loss after validating that frequency-domain errors significantly impact prediction quality.\n\n3. **Edge Loss for Fault Clarity**:\n\n   * Implement a conservative edge loss initially to enhance geological feature clarity, tuning upward as model stability permits.\n\n4. **Monitor Continuously**:\n\n   * Regularly evaluate depth-specific error curves, spectral domain analyses, and fault sharpness to iteratively refine and balance loss components.\n\nBy strategically applying these auxiliary losses, FWI models trained on the OpenFWI dataset can achieve more accurate, geologically meaningful predictions, addressing the critical structured errors inherent in purely MAE-driven optimization.\n\n## Reading List\n\nHere’s a starter reading list (with plain-language notes) on **auxiliary or alternative loss functions that push optimisation “faster and farther” than a plain MAE / L1 objective**. I’ve grouped them by theme so you can skim to what’s most relevant.\n\n---\n\n### 1. Frequency-aware / spectral losses — sharper details, quicker convergence\n\n| Paper                                                                      | Core idea (plain words)                                                                                                                                   | Why it matters                                                                                                                                                           |\n| -------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| **Guided Frequency Loss (GFL)**, ICCV 2023 ([arXiv][1])                    | Combines a Charbonnier (robust-L1), Laplacian-pyramid, and gradual-frequency term so the network balances low- and high-frequency energy while it learns. | Training stabilises sooner and PSNR jumps in image-restoration tasks; the same notion of “teach the model where in the spectrum it’s wrong” maps cleanly to FWI spectra. |\n| **Focal Frequency Loss (FFL)**, ICCV 2021 ([CVF Open Access][2])           | Lets the model *adaptively* up-weight frequency bins it currently struggles with, down-weighting the easy ones.                                           | Empirically speeds convergence and yields crisper outputs; conceptually similar to putting a moving spotlight on the hardest frequencies in seismic waveforms.           |\n| **ω-FWI: Fourier-metric Full Waveform Inversion**, arXiv 2022 ([arXiv][3]) | Replaces the sample-by-sample misfit with a Fourier-domain power-spectrum distance, giving the optimiser smoother gradients and less cycle-skipping.      | Demonstrates faster recovery of low-wavenumber structure from poor starting models—exactly the “go farther” property we want.                                            |\n\n---\n\n### 2. Structure-aware losses — reward “looks right”, not just “averages out”\n\n| Paper                                                               | Core idea                                                                                                            | Why it matters                                                                                                                                                                         |\n| ------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| **MS\\_ATpV-FWI**: Multi-scale SSIM + p-Variation, 2025 ([arXiv][4]) | Uses **multi-scale SSIM** (captures structural similarity) and **anisotropic total p-variation** as auxiliary terms. | Multiscale SSIM gives gradients that focus on phase + amplitude coherence, reducing cycle-skipping; p-Variation keeps faults sharp. Reported to converge where plain L2 or MAE stalls. |\n| **ML-misfit**, arXiv 2020 ([arXiv][5])                              | *Learns* the misfit itself (meta-learning) so it becomes convex around realistic waveform shifts.                    | Shows that a data-driven auxiliary loss can widen the basin of convergence in FWI—fewer restarts, deeper minima.                                                                       |\n\n---\n\n### 3. Optimal-transport & Wasserstein misfits — align whole waveforms, not points\n\n| Paper                                                                 | Core idea                                                                                                                      | Why it matters                                                                                                        |\n| --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------- |\n| **W2-FWI (Quadratic Wasserstein)**, Engquist et al. 2016 ([arXiv][6]) | Treats traces as mass distributions and measures the *cost to morph one into the other*; inherently accounts for phase shifts. | Strong convexity properties give larger, smoother descent steps—practically fewer iterations to reach a usable model. |\n\n---\n\n### 4. Expert-diversity / load-balancing losses (if you use MoE routers)\n\n| Paper                                                                       | Core idea                                                                                                                                         | Take-away                                                                                         |\n| --------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- |\n| **Auxiliary-Loss-Free Load Balancing**, 2024 ([arXiv][7])                   | Shows classic load-balance penalties can inject noisy gradients; proposes bias-based routing that *keeps balance* **without** an extra loss term. | Useful caution: auxiliary terms help, but ill-designed ones can hinder optimisation speed.        |\n| **Demons in the Detail: Revisiting Load-Balancing Loss**, 2025 ([arXiv][8]) | Demonstrates that *how* you compute the balance loss (micro- vs global-batch) dramatically affects expert specialisation and final quality.       | Highlights that even tiny auxiliary losses need thoughtful implementation to really “go farther.” |\n\n---\n\n### 5. Honorable mentions & cross-links\n\n* **Guided Frequency-aware SSIM**, **Amplitude-based misfit** models, and other **multi-scale FWI** variants extend the same philosophy: combine *where* it’s wrong (depth, frequency, edges) with *how much* it’s wrong (MAE/L2) for faster, deeper convergence.\n\n---\n\n#### TL;DR\n\nYes—there’s a growing body of work showing that *well-chosen* auxiliary losses (frequency, structural, transport, or balance-oriented) can **speed up training, escape shallow minima, and land on solutions MAE alone cannot reach**.  These papers offer concrete design patterns you can port straight into your OpenFWI experiments.\n\n[1]: https://arxiv.org/abs/2309.15563?utm_source=chatgpt.com \"Guided Frequency Loss for Image Restoration\"\n[2]: https://openaccess.thecvf.com/content/ICCV2021/papers/Jiang_Focal_Frequency_Loss_for_Image_Reconstruction_and_Synthesis_ICCV_2021_paper.pdf?utm_source=chatgpt.com \"[PDF] Focal Frequency Loss for Image Reconstruction and Synthesis\"\n[3]: https://arxiv.org/abs/2205.09234?utm_source=chatgpt.com \"$ω$-FWI: Robust full-waveform inversion with Fourier-based metric\"\n[4]: https://arxiv.org/html/2504.01695 \"MS_ATpV-FWI: Full Waveform Inversion based on Multi-scale Structural Similarity Index Measure and Anisotropic Total p-Variation Regularization\"\n[5]: https://arxiv.org/abs/2002.03163?utm_source=chatgpt.com \"ML-misfit: Learning a robust misfit function for full-waveform inversion using machine learning\"\n[6]: https://arxiv.org/abs/1612.05075?utm_source=chatgpt.com \"Application of Optimal Transport and the Quadratic Wasserstein Metric to Full-Waveform Inversion\"\n[7]: https://arxiv.org/abs/2408.15664?utm_source=chatgpt.com \"Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts\"\n[8]: https://arxiv.org/abs/2501.11873?utm_source=chatgpt.com \"Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models\"",
    "3202129": "There is a lot of text and figures both in this post and the notebook. However I did not find the most important thing- ablation study scores or something similar that proves:\n1. That auxiliary loss helps.  \n2. By how much compared to baseline...",
    "3202160": "Update:   ***The notebook goes over 100+ possible experimental combinations that I could not possibly do alone.***  This will take time and is better performed by the community as a whole.   Your comment misses the purpose of this discussion and my notebook.\n\nOn my own, so far, I have not had any luck with auxiliary loss functions on any of the models I've used... but I have not extensively tested them, nor do I have the means to.  ***I only have one 3060, and running any one experiment takes 2-5 days***, and that tends to only get me to epoch 30-40 with the whole dataset... It really isn't enough to come to any conclusion at all.  \n\nAuxiliary loss functions also seem to make the loss surface potentially more complicated, but ***don't overfit on this very small result of n=1***.  Results are super sensitive the hyper-parameters chosen, the weight values given to the loss functions, and the learning rate isn't the same when you use auxiliary loss functions etc.  ***This is precisely why I was hoping we'd have many people experimenting instead of just me.   I hope others play with these ideas and share their results too.***  The loss surface can become a bit more complicated with auxiliary loss functions, but the real problem is just low compute.\n\nThe real question is whether there is an area of the loss surface that is lower with auxiliary losses, and we just don't know that.  The research I cite suggests this may be the case.\n\nCurrent ideas are to realize that the depth based loss might need to be non-linear.  Spectral loss does correlate well with val loss, so that seems more promising.\n\n--Older message--\n\nMy current experiments are with different NN architectures, so results from them won't necessarily transfer to an ablation study of the NN architectures I wrote about above (which the community seems to be mostly using).  But this is besides the point.\n\nThe purpose of my posts is to clarify possible high-yield directions and experimental priorities to ***empower others (hundreds of Kagglers) to test more things than I could possibly test alone***.  \n\nThe notebook I've shared represents many hours of detailed research.  I wouldn't have done that unless I thought it was very valuable.  \n\nSo I disagree with the claim that the results alone are the \"most important thing.\"   ***Just like a data generating process is richer than the specific data points it outputs, I believe a series of research directions are far richer and deeper than any of the specific and temporary results they generate.***\n\nResults will come, but be patient.  It turns out that I do have a life outside of Kaggle ;).\n\nIf you're eager for quicker results from an ablation study, I warmly encourage you to perform one yourself and contribute your findings. That would be a meaningful addition to our collective effort, and clearly ***you*** would find it more meaningful than my \"text and figures\" 🤣"
  },
  "source": "meta"
}