{
  "id": 587422,
  "title": " Some Findings on Backpropagation with Forward Models",
  "url": "/competitions/waveform-inversion/discussion/587422",
  "author_name": "",
  "post_date": "2025-07-01T02:47:15.480778700Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congratulations to everyone who was able to make progress beyond Bartley’s strong model(s).</p>\n<p>I could not make further progress myself, so I am sharing some findings from my experiments on backpropagating through forward models.<br>\nThis technique is mainly used in generative image models [1, 2]; for example, obtaining input image which closely relate to the prompt text with CLIP embedding space.</p>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/2502.07753\" target=\"_blank\">https://arxiv.org/abs/2502.07753</a></li>\n<li>[2] <a href=\"https://arxiv.org/abs/2208.01618\" target=\"_blank\">https://arxiv.org/abs/2208.01618</a></li>\n</ul>\n<p>The experiment was as follows:</p>\n<ul>\n<li>I initialized the input image using Bartley’s model predictions (for example, an image from CurveFault-B with MAE=44.18).</li>\n<li>I optimized the input image to make the seismogram closer to the “ground truth” (see the 4th picture below).</li>\n</ul>\n<p><strong>Results:</strong></p>\n<ul>\n<li>Initial image (MAE=44.18)</li>\n<li>After 400 epochs (MAE=29.92)</li>\n<li>After 800 epochs (MAE=27.53)</li>\n</ul>\n<p>This approach can indeed improve the metric. However, I found it impractical due to computational efficiency: it took 150 seconds to optimize a single velocity map over 400 epochs.<br>\nThis means optimizing 65,000 images would require <strong>113 days</strong> on a single RTX 4090, which is clearly unrealistic.</p>\n<p>Perhaps a more efficient forward model could produce results in a reasonable timeframe.</p>\n<p><strong>Details of the experimental setup:</strong></p>\n<ul>\n<li>Learning rate: 20.0 (eta_min=0.05)</li>\n<li>Optimizer: Adam (betas=(0.9, 0.99))</li>\n<li>Schedule: constant LR</li>\n<li>Criterion: Huber Loss (delta=1.0) + 0.01 * Total Variation</li>\n<li>Using mu-law converted seismograms instead of the original ones</li>\n</ul>\n<p><strong>Tips on Efficiency</strong></p>\n<ul>\n<li>Converting rolling operations to 2D convolutions</li>\n<li>Converting 2D convolutions to FFT: we can perform FFT + element-wise multiplication + inverse FFT, which is equivalent to a Conv2D with cyclic padding (<strong>1.3~1.5 speedup</strong>)</li>\n</ul>\n<p><strong>Did not work:</strong></p>\n<ul>\n<li>Perceptual loss</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ffa6e55c3a97072379442a35f867a51ab%2FScreenshot%202025-07-01%20at%2011.20.48.png?generation=1751336509771951&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F21c7f0c6f65a52c834c22dacf83b809f%2FScreenshot%202025-07-01%20at%2011.20.58.png?generation=1751336521559770&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fac3399e4495e6487fe91069d114c3afa%2FScreenshot%202025-07-01%20at%2011.23.02.png?generation=1751336595264759&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F584e44a9f69905efd39ac58eea70bb80%2FScreenshot%202025-07-01%20at%2011.21.08.png?generation=1751336533273340&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3237291",
      "postDate": "07/01/2025 02:47:15",
      "content": "<p>Congratulations to everyone who was able to make progress beyond Bartley’s strong model(s).</p>\n<p>I could not make further progress myself, so I am sharing some findings from my experiments on backpropagating through forward models.<br>\nThis technique is mainly used in generative image models [1, 2]; for example, obtaining input image which closely relate to the prompt text with CLIP embedding space.</p>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/2502.07753\" target=\"_blank\">https://arxiv.org/abs/2502.07753</a></li>\n<li>[2] <a href=\"https://arxiv.org/abs/2208.01618\" target=\"_blank\">https://arxiv.org/abs/2208.01618</a></li>\n</ul>\n<p>The experiment was as follows:</p>\n<ul>\n<li>I initialized the input image using Bartley’s model predictions (for example, an image from CurveFault-B with MAE=44.18).</li>\n<li>I optimized the input image to make the seismogram closer to the “ground truth” (see the 4th picture below).</li>\n</ul>\n<p><strong>Results:</strong></p>\n<ul>\n<li>Initial image (MAE=44.18)</li>\n<li>After 400 epochs (MAE=29.92)</li>\n<li>After 800 epochs (MAE=27.53)</li>\n</ul>\n<p>This approach can indeed improve the metric. However, I found it impractical due to computational efficiency: it took 150 seconds to optimize a single velocity map over 400 epochs.<br>\nThis means optimizing 65,000 images would require <strong>113 days</strong> on a single RTX 4090, which is clearly unrealistic.</p>\n<p>Perhaps a more efficient forward model could produce results in a reasonable timeframe.</p>\n<p><strong>Details of the experimental setup:</strong></p>\n<ul>\n<li>Learning rate: 20.0 (eta_min=0.05)</li>\n<li>Optimizer: Adam (betas=(0.9, 0.99))</li>\n<li>Schedule: constant LR</li>\n<li>Criterion: Huber Loss (delta=1.0) + 0.01 * Total Variation</li>\n<li>Using mu-law converted seismograms instead of the original ones</li>\n</ul>\n<p><strong>Tips on Efficiency</strong></p>\n<ul>\n<li>Converting rolling operations to 2D convolutions</li>\n<li>Converting 2D convolutions to FFT: we can perform FFT + element-wise multiplication + inverse FFT, which is equivalent to a Conv2D with cyclic padding (<strong>1.3~1.5 speedup</strong>)</li>\n</ul>\n<p><strong>Did not work:</strong></p>\n<ul>\n<li>Perceptual loss</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ffa6e55c3a97072379442a35f867a51ab%2FScreenshot%202025-07-01%20at%2011.20.48.png?generation=1751336509771951&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F21c7f0c6f65a52c834c22dacf83b809f%2FScreenshot%202025-07-01%20at%2011.20.58.png?generation=1751336521559770&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fac3399e4495e6487fe91069d114c3afa%2FScreenshot%202025-07-01%20at%2011.23.02.png?generation=1751336595264759&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F584e44a9f69905efd39ac58eea70bb80%2FScreenshot%202025-07-01%20at%2011.21.08.png?generation=1751336533273340&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Congratulations to everyone who was able to make progress beyond Bartley’s strong model(s).\n\nI could not make further progress myself, so I am sharing some findings from my experiments on backpropagating through forward models.\nThis technique is mainly used in generative image models \\[1, 2]; for example, obtaining input image which closely relate to the prompt text with CLIP embedding space.\n\n* \\[1] [https://arxiv.org/abs/2502.07753](https://arxiv.org/abs/2502.07753)\n* \\[2] [https://arxiv.org/abs/2208.01618](https://arxiv.org/abs/2208.01618)\n\nThe experiment was as follows:\n\n* I initialized the input image using Bartley’s model predictions (for example, an image from CurveFault-B with MAE=44.18).\n* I optimized the input image to make the seismogram closer to the “ground truth” (see the 4th picture below).\n\n**Results:**\n\n* Initial image (MAE=44.18)\n* After 400 epochs (MAE=29.92)\n* After 800 epochs (MAE=27.53)\n\nThis approach can indeed improve the metric. However, I found it impractical due to computational efficiency: it took 150 seconds to optimize a single velocity map over 400 epochs.\nThis means optimizing 65,000 images would require **113 days** on a single RTX 4090, which is clearly unrealistic.\n\nPerhaps a more efficient forward model could produce results in a reasonable timeframe.\n\n**Details of the experimental setup:**\n\n* Learning rate: 20.0 (eta_min=0.05)\n* Optimizer: Adam (betas=(0.9, 0.99))\n* Schedule: constant LR\n* Criterion: Huber Loss (delta=1.0) + 0.01 * Total Variation\n* Using mu-law converted seismograms instead of the original ones\n\n**Tips on Efficiency**\n\n* Converting rolling operations to 2D convolutions\n* Converting 2D convolutions to FFT: we can perform FFT + element-wise multiplication + inverse FFT, which is equivalent to a Conv2D with cyclic padding (**1.3~1.5 speedup**)\n\n**Did not work:**\n\n* Perceptual loss\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ffa6e55c3a97072379442a35f867a51ab%2FScreenshot%202025-07-01%20at%2011.20.48.png?generation=1751336509771951&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F21c7f0c6f65a52c834c22dacf83b809f%2FScreenshot%202025-07-01%20at%2011.20.58.png?generation=1751336521559770&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fac3399e4495e6487fe91069d114c3afa%2FScreenshot%202025-07-01%20at%2011.23.02.png?generation=1751336595264759&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F584e44a9f69905efd39ac58eea70bb80%2FScreenshot%202025-07-01%20at%2011.21.08.png?generation=1751336533273340&alt=media)",
      "votes": null
    },
    {
      "id": "3237479",
      "postDate": "07/01/2025 06:00:55",
      "content": "<p>This is essentially what I did. Full details in the weekend, but some specific things responding to your suggestions:</p>\n<ul>\n<li>Using priors is critical (i.e. an approximation of the distribution from which the velocity map is taken). The pure inversion problem is far too ill-posed to be solved without one.</li>\n<li>Extremely efficient forward model and backpropagation is critical.</li>\n<li>Gradient descent with ADAM is fine for neural networks, but smaller scale problems like this need more dedicated solvers. In my case I used BFGS and Gauss-Newton.</li>\n<li>It is indeed computationally very expensive.</li>\n</ul>\n<p>To be continued!</p>",
      "rawMarkdown": "This is essentially what I did. Full details in the weekend, but some specific things responding to your suggestions:\n\n- Using priors is critical (i.e. an approximation of the distribution from which the velocity map is taken). The pure inversion problem is far too ill-posed to be solved without one.\n- Extremely efficient forward model and backpropagation is critical.\n- Gradient descent with ADAM is fine for neural networks, but smaller scale problems like this need more dedicated solvers. In my case I used BFGS and Gauss-Newton.\n- It is indeed computationally very expensive.\n\nTo be continued!",
      "votes": null
    },
    {
      "id": "3237495",
      "postDate": "07/01/2025 06:13:17",
      "content": "<blockquote>\n  <p>It is indeed computationally very expensive.</p>\n</blockquote>\n<p>I can't agree more. Back propagation for x1000 times 5x5 convolution is surely compute intensive.<br>\nI consulted Chat-GPT to further drastically reduced the computation amount, but I found it seems we can't reduce them under $O(n_t * n_z * n_x)$. </p>\n<blockquote>\n  <p>Using priors is critical</p>\n</blockquote>\n<p>Yes. I also tested for initializing with a gradation of 1500-4500 m/s vel maps, and it only resulted in ~MAE=300 with this prior after 400 epochs.<br>\n(Note that this implementation of total variation is wrong, so we can observe artifact of horizontal lines.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F226255ecea4cc582f39977be84c91749%2FScreenshot%202025-07-01%20at%2015.10.52.png?generation=1751350287562241&amp;alt=media\" alt=\"\"></p>\n<blockquote>\n  <p>Gradient descent with ADAM is fine for neural networks, but smaller scale problems like this need more dedicated solvers. In my case I used BFGS and Gauss-Newton.</p>\n</blockquote>\n<p>Interesting. Thank you for sharing!</p>",
      "rawMarkdown": "> It is indeed computationally very expensive.\n\nI can't agree more. Back propagation for x1000 times 5x5 convolution is surely compute intensive.\nI consulted Chat-GPT to further drastically reduced the computation amount, but I found it seems we can't reduce them under $O(n_t * n_z * n_x)$. \n\n> Using priors is critical\n\nYes. I also tested for initializing with a gradation of 1500-4500 m/s vel maps, and it only resulted in ~MAE=300 with this prior after 400 epochs.\n(Note that this implementation of total variation is wrong, so we can observe artifact of horizontal lines.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F226255ecea4cc582f39977be84c91749%2FScreenshot%202025-07-01%20at%2015.10.52.png?generation=1751350287562241&alt=media)\n\n> Gradient descent with ADAM is fine for neural networks, but smaller scale problems like this need more dedicated solvers. In my case I used BFGS and Gauss-Newton.\n\nInteresting. Thank you for sharing!",
      "votes": null
    },
    {
      "id": "3237640",
      "postDate": "07/01/2025 07:52:09",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> <br>\nFWI: </p>\n<ul>\n<li>the result of gradation velocity prior seems wrong; the same setup than the latest setup generates better result.</li>\n<li>I also tested LBFGS. It performs better per step, but worse result per actual computation time compared to Adam (more parameter tuning might improve the result)</li>\n</ul>\n<p><strong>Result</strong></p>\n<table>\n<thead>\n<tr>\n<th>Optimization</th>\n<th>Epochs</th>\n<th>Time [sec]</th>\n<th>MAE [m/s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Initial</td>\n<td>0</td>\n<td>0</td>\n<td>405.00</td>\n</tr>\n<tr>\n<td>Adam</td>\n<td>400</td>\n<td>161</td>\n<td>255.17</td>\n</tr>\n<tr>\n<td>Adam</td>\n<td>800</td>\n<td>323</td>\n<td>224.62</td>\n</tr>\n<tr>\n<td>LBFGS</td>\n<td>80</td>\n<td>160</td>\n<td>348.62</td>\n</tr>\n<tr>\n<td>LBFGS</td>\n<td>160</td>\n<td>312</td>\n<td>327.36</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Hyper Parameter</strong>:</p>\n<pre><code>opt = optim.LBFGS(\n    [vel_param], lr=, history_size=, max_iter=, line_search_fn=\n)\n</code></pre>\n<p><strong>Initial velocity map (gradation map over 1500-4000 [m/s])</strong>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fcb0b78e5b7cb234c4d3307a6d84da155%2FScreenshot%202025-07-01%20at%2016.44.36.png?generation=1751355894668105&amp;alt=media\" alt=\"\"></p>\n<p><strong>Adam (800 epoch)</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F08af2f683b884692cc573e3c553ff700%2FScreenshot%202025-07-01%20at%2016.44.23.png?generation=1751355910999978&amp;alt=media\" alt=\"\"></p>\n<p><strong>LBFGS (160epoch)</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F82bbd3067c784a4b607b090e5efd1f8f%2FScreenshot%202025-07-01%20at%2016.44.29.png?generation=1751355922783371&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "jeroencottaar \nFWI: \n- the result of gradation velocity prior seems wrong; the same setup than the latest setup generates better result.\n- I also tested LBFGS. It performs better per step, but worse result per actual computation time compared to Adam (more parameter tuning might improve the result)\n\n**Result**\n\n| Optimization | Epochs | Time [sec] | MAE [m/s] |\n|--------------|--------|------------|-----------|\n| Initial      | 0      | 0          | 405.00    |\n| Adam         | 400    | 161        | 255.17    |\n| Adam         | 800    | 323        | 224.62    |\n| LBFGS        | 80     | 160        | 348.62    |\n| LBFGS        | 160    | 312        | 327.36    |\n\n**Hyper Parameter**:\n\n```python\nopt = optim.LBFGS(\n    [vel_param], lr=5.0, history_size=10, max_iter=4, line_search_fn=\"strong_wolfe\"\n)\n```\n\n**Initial velocity map (gradation map over 1500-4000 [m/s])**:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fcb0b78e5b7cb234c4d3307a6d84da155%2FScreenshot%202025-07-01%20at%2016.44.36.png?generation=1751355894668105&alt=media)\n\n**Adam (800 epoch)**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F08af2f683b884692cc573e3c553ff700%2FScreenshot%202025-07-01%20at%2016.44.23.png?generation=1751355910999978&alt=media)\n\n**LBFGS (160epoch)**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F82bbd3067c784a4b607b090e5efd1f8f%2FScreenshot%202025-07-01%20at%2016.44.29.png?generation=1751355922783371&alt=media)",
      "votes": null
    },
    {
      "id": "3237641",
      "postDate": "07/01/2025 07:53:14",
      "content": "<p>Let's continue this when I've written up my full solution (after some work deadlines).</p>",
      "rawMarkdown": "Let's continue this when I've written up my full solution (after some work deadlines).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237479,
      "author_name": "jeroencottaar",
      "author_url": "",
      "post_date": "07/01/2025 06:00:55",
      "content": "<p>This is essentially what I did. Full details in the weekend, but some specific things responding to your suggestions:</p>\n<ul>\n<li>Using priors is critical (i.e. an approximation of the distribution from which the velocity map is taken). The pure inversion problem is far too ill-posed to be solved without one.</li>\n<li>Extremely efficient forward model and backpropagation is critical.</li>\n<li>Gradient descent with ADAM is fine for neural networks, but smaller scale problems like this need more dedicated solvers. In my case I used BFGS and Gauss-Newton.</li>\n<li>It is indeed computationally very expensive.</li>\n</ul>\n<p>To be continued!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237495,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "07/01/2025 06:13:17",
          "content": "<blockquote>\n  <p>It is indeed computationally very expensive.</p>\n</blockquote>\n<p>I can't agree more. Back propagation for x1000 times 5x5 convolution is surely compute intensive.<br>\nI consulted Chat-GPT to further drastically reduced the computation amount, but I found it seems we can't reduce them under $O(n_t * n_z * n_x)$. </p>\n<blockquote>\n  <p>Using priors is critical</p>\n</blockquote>\n<p>Yes. I also tested for initializing with a gradation of 1500-4500 m/s vel maps, and it only resulted in ~MAE=300 with this prior after 400 epochs.<br>\n(Note that this implementation of total variation is wrong, so we can observe artifact of horizontal lines.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F226255ecea4cc582f39977be84c91749%2FScreenshot%202025-07-01%20at%2015.10.52.png?generation=1751350287562241&amp;alt=media\" alt=\"\"></p>\n<blockquote>\n  <p>Gradient descent with ADAM is fine for neural networks, but smaller scale problems like this need more dedicated solvers. In my case I used BFGS and Gauss-Newton.</p>\n</blockquote>\n<p>Interesting. Thank you for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237640,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "07/01/2025 07:52:09",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> <br>\nFWI: </p>\n<ul>\n<li>the result of gradation velocity prior seems wrong; the same setup than the latest setup generates better result.</li>\n<li>I also tested LBFGS. It performs better per step, but worse result per actual computation time compared to Adam (more parameter tuning might improve the result)</li>\n</ul>\n<p><strong>Result</strong></p>\n<table>\n<thead>\n<tr>\n<th>Optimization</th>\n<th>Epochs</th>\n<th>Time [sec]</th>\n<th>MAE [m/s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Initial</td>\n<td>0</td>\n<td>0</td>\n<td>405.00</td>\n</tr>\n<tr>\n<td>Adam</td>\n<td>400</td>\n<td>161</td>\n<td>255.17</td>\n</tr>\n<tr>\n<td>Adam</td>\n<td>800</td>\n<td>323</td>\n<td>224.62</td>\n</tr>\n<tr>\n<td>LBFGS</td>\n<td>80</td>\n<td>160</td>\n<td>348.62</td>\n</tr>\n<tr>\n<td>LBFGS</td>\n<td>160</td>\n<td>312</td>\n<td>327.36</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Hyper Parameter</strong>:</p>\n<pre><code>opt = optim.LBFGS(\n    [vel_param], lr=, history_size=, max_iter=, line_search_fn=\n)\n</code></pre>\n<p><strong>Initial velocity map (gradation map over 1500-4000 [m/s])</strong>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fcb0b78e5b7cb234c4d3307a6d84da155%2FScreenshot%202025-07-01%20at%2016.44.36.png?generation=1751355894668105&amp;alt=media\" alt=\"\"></p>\n<p><strong>Adam (800 epoch)</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F08af2f683b884692cc573e3c553ff700%2FScreenshot%202025-07-01%20at%2016.44.23.png?generation=1751355910999978&amp;alt=media\" alt=\"\"></p>\n<p><strong>LBFGS (160epoch)</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F82bbd3067c784a4b607b090e5efd1f8f%2FScreenshot%202025-07-01%20at%2016.44.29.png?generation=1751355922783371&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3237641,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "07/01/2025 07:53:14",
          "content": "<p>Let's continue this when I've written up my full solution (after some work deadlines).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237291": "Congratulations to everyone who was able to make progress beyond Bartley’s strong model(s).\n\nI could not make further progress myself, so I am sharing some findings from my experiments on backpropagating through forward models.\nThis technique is mainly used in generative image models \\[1, 2]; for example, obtaining input image which closely relate to the prompt text with CLIP embedding space.\n\n* \\[1] [https://arxiv.org/abs/2502.07753](https://arxiv.org/abs/2502.07753)\n* \\[2] [https://arxiv.org/abs/2208.01618](https://arxiv.org/abs/2208.01618)\n\nThe experiment was as follows:\n\n* I initialized the input image using Bartley’s model predictions (for example, an image from CurveFault-B with MAE=44.18).\n* I optimized the input image to make the seismogram closer to the “ground truth” (see the 4th picture below).\n\n**Results:**\n\n* Initial image (MAE=44.18)\n* After 400 epochs (MAE=29.92)\n* After 800 epochs (MAE=27.53)\n\nThis approach can indeed improve the metric. However, I found it impractical due to computational efficiency: it took 150 seconds to optimize a single velocity map over 400 epochs.\nThis means optimizing 65,000 images would require **113 days** on a single RTX 4090, which is clearly unrealistic.\n\nPerhaps a more efficient forward model could produce results in a reasonable timeframe.\n\n**Details of the experimental setup:**\n\n* Learning rate: 20.0 (eta_min=0.05)\n* Optimizer: Adam (betas=(0.9, 0.99))\n* Schedule: constant LR\n* Criterion: Huber Loss (delta=1.0) + 0.01 * Total Variation\n* Using mu-law converted seismograms instead of the original ones\n\n**Tips on Efficiency**\n\n* Converting rolling operations to 2D convolutions\n* Converting 2D convolutions to FFT: we can perform FFT + element-wise multiplication + inverse FFT, which is equivalent to a Conv2D with cyclic padding (**1.3~1.5 speedup**)\n\n**Did not work:**\n\n* Perceptual loss\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Ffa6e55c3a97072379442a35f867a51ab%2FScreenshot%202025-07-01%20at%2011.20.48.png?generation=1751336509771951&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F21c7f0c6f65a52c834c22dacf83b809f%2FScreenshot%202025-07-01%20at%2011.20.58.png?generation=1751336521559770&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fac3399e4495e6487fe91069d114c3afa%2FScreenshot%202025-07-01%20at%2011.23.02.png?generation=1751336595264759&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F584e44a9f69905efd39ac58eea70bb80%2FScreenshot%202025-07-01%20at%2011.21.08.png?generation=1751336533273340&alt=media)",
    "3237479": "This is essentially what I did. Full details in the weekend, but some specific things responding to your suggestions:\n\n- Using priors is critical (i.e. an approximation of the distribution from which the velocity map is taken). The pure inversion problem is far too ill-posed to be solved without one.\n- Extremely efficient forward model and backpropagation is critical.\n- Gradient descent with ADAM is fine for neural networks, but smaller scale problems like this need more dedicated solvers. In my case I used BFGS and Gauss-Newton.\n- It is indeed computationally very expensive.\n\nTo be continued!",
    "3237495": "> It is indeed computationally very expensive.\n\nI can't agree more. Back propagation for x1000 times 5x5 convolution is surely compute intensive.\nI consulted Chat-GPT to further drastically reduced the computation amount, but I found it seems we can't reduce them under $O(n_t * n_z * n_x)$. \n\n> Using priors is critical\n\nYes. I also tested for initializing with a gradation of 1500-4500 m/s vel maps, and it only resulted in ~MAE=300 with this prior after 400 epochs.\n(Note that this implementation of total variation is wrong, so we can observe artifact of horizontal lines.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F226255ecea4cc582f39977be84c91749%2FScreenshot%202025-07-01%20at%2015.10.52.png?generation=1751350287562241&alt=media)\n\n> Gradient descent with ADAM is fine for neural networks, but smaller scale problems like this need more dedicated solvers. In my case I used BFGS and Gauss-Newton.\n\nInteresting. Thank you for sharing!",
    "3237640": "jeroencottaar \nFWI: \n- the result of gradation velocity prior seems wrong; the same setup than the latest setup generates better result.\n- I also tested LBFGS. It performs better per step, but worse result per actual computation time compared to Adam (more parameter tuning might improve the result)\n\n**Result**\n\n| Optimization | Epochs | Time [sec] | MAE [m/s] |\n|--------------|--------|------------|-----------|\n| Initial      | 0      | 0          | 405.00    |\n| Adam         | 400    | 161        | 255.17    |\n| Adam         | 800    | 323        | 224.62    |\n| LBFGS        | 80     | 160        | 348.62    |\n| LBFGS        | 160    | 312        | 327.36    |\n\n**Hyper Parameter**:\n\n```python\nopt = optim.LBFGS(\n    [vel_param], lr=5.0, history_size=10, max_iter=4, line_search_fn=\"strong_wolfe\"\n)\n```\n\n**Initial velocity map (gradation map over 1500-4000 [m/s])**:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fcb0b78e5b7cb234c4d3307a6d84da155%2FScreenshot%202025-07-01%20at%2016.44.36.png?generation=1751355894668105&alt=media)\n\n**Adam (800 epoch)**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F08af2f683b884692cc573e3c553ff700%2FScreenshot%202025-07-01%20at%2016.44.23.png?generation=1751355910999978&alt=media)\n\n**LBFGS (160epoch)**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F82bbd3067c784a4b607b090e5efd1f8f%2FScreenshot%202025-07-01%20at%2016.44.29.png?generation=1751355922783371&alt=media)",
    "3237641": "Let's continue this when I've written up my full solution (after some work deadlines)."
  },
  "source": "meta"
}