{
  "id": 587950,
  "title": "2nd place solution with code: refinement with FWI",
  "url": "/competitions/waveform-inversion/writeups/jeroen-cottaar-2nd-place-solution-with-code-refine",
  "author_name": "",
  "post_date": "2025-07-03T15:19:25.564118900Z",
  "votes": 49,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Link to full code: <a href=\"https://www.kaggle.com/code/jeroencottaar/geophysical-waveform-inversion-2nd-place\" target=\"_blank\">https://www.kaggle.com/code/jeroencottaar/geophysical-waveform-inversion-2nd-place</a></p>\n<p>Thanks to the organizers for a fun competition, to Kaggle for providing the infrastructure, and to <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for providing an excellent starting point - I probably wouldn't even have competed without it.</p>\n<h2>Introduction</h2>\n<p>How do you find out what's under your feet without digging? You give the ground a good thump, which sends seismic waves underground. These waves interact with and reflect off underground structures. By placing multiple detectors at the surface, you can pick up the reflected waves. This measurement is called a seismogram.</p>\n<p>It's fairly straightforward to model the impact on the seismogram of subsurface structures, described by a velocity profile. The challenge lies in inverting this function: how do we find the original velocity profile given a seismogram? This was the challenge posed to us in this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F862fb07e92298b675b70e3088b322f28%2FScreenshot%202025-07-03%20171603.png?generation=1751555797803926&amp;alt=media\" alt=\"\"></p>\n<h2>Outline of my solution</h2>\n<p>I start with a rough deep learning model that would score 28.8 on the competition metric (this is the model <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> shared a few weeks ago). This is then refined using full waveform inversion (FWI) to a score of 7.6. The challenge was making this feasible in terms of convergence and computation time. The key elements of my solution:</p>\n<ul>\n<li>Carefully designed priors for various subsets of the data.</li>\n<li>Choosing and tweaking of solvers (BFGS and Gauss-Newton) per subset.</li>\n<li>Optimizing computation speed, including custom CUDA kernels for the forward model and BFGS overhead.</li>\n</ul>\n<p>Even then, my solution cost about $700 in cloud compute costs to develop and run. This is in line with other top solutions, but I'm hoping this competition is an outlier in this regard or we'll have very few competitors going for top spots soon - not to mention the environmental impact.</p>\n<h2>Details</h2>\n<p>FWI is much easier when you start close to the solution. For this I use a deep learning model. I didn't develop this myself, but use a model <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> graciously shared earlier in the competition. I modified nothing about it (and also didn't train it myself), so if you want to know more about it, you might as well get it from him: <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved</a></p>\n<p>From there I perform regularized full waveform inversion. This means I'm trying to find <em>x</em> that minimizes:<br>\n$$c(x)=p(x) + ||s(x)-m||^2_2$$<br>\nHere, <em>x</em> is the velocity profile, <em>p(x)</em> is the cost function defined by the prior, <em>s(x)</em> is the seismogram that follows from <em>x</em>, and <em>m</em> is the actual measured seismogram. In words: we're trying to balance minimizing a cost function and minimizing the residual.</p>\n<p>Typical methods to achieve this are BFGS and Gauss-Newton; I won't discuss these in detail. Both of these methods require us differentiate the cost function above; the most expensive part of this is computing the gradient of the seismogram function <em>\\triangledown s(x)</em>. I heavily optimized this with custom CUDA kernels; the forward pass takes 30 ms and the backward pass 70 ms on a V100 at double precision, though faster is probably possible. I also reimplemented parts of BFGS with a CUDA kernel to reduce overhead.</p>\n<p>What remains is the choice of prior <em>p(x)</em> and solving strategy, which took up the bulk of development time. Different priors are needed for the different families. I won't discuss how I classify them here; as many of you have noticed, this is a fairly easy problem.</p>\n<h5>FlatVelA and FlatVelB</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F8fc5ef9592231a41259cda052c71bae3%2FScreenshot%202025-07-03%20165936.png?generation=1751554792560180&amp;alt=media\" alt=\"\"><br>\nThe solution is restricted to 1D profiles, with total variation as the cost function:<br>\n$$ p(x)=\\sum_{i=1}^{69} |x_{i+1}-x_i| $$<br>\n(Actually, it's not quite the absolute value, but smoothed around 0 to remain differentiable.)</p>\n<p>I use a single BFGS pass, terminated when the cost function stops decreasing significantly. My public LB score for this one is well under 0.1.</p>\n<h5>StyleA</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F4025fbd7ae9124c93460f6bf1c0a94bc%2FScreenshot%202025-07-03%20165958.png?generation=1751554813255044&amp;alt=media\" alt=\"\"><br>\nThe prior for this one is a Gaussian Process with a squared-exponential kernel, and additional terms for noise, offset, and slope. Hyperparameters were tuned on the training set using maximum likelihood estimation. There is actually no noise in StyleA, which makes the covariance matrix <em>K</em> very ill-conditioned. I resolve this by restricting the solution to the strongest 1073 modes of an SVD decomposition of <em>K</em>.</p>\n<p>The strategy is a BFGS pass, further refined by a single Gauss-Newton step. My public LB score for this one is around 3.4.</p>\n<h5>StyleB</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F872cb2fbc53d8ca3975539ae406ce082%2FScreenshot%202025-07-03%20165908.png?generation=1751554770111041&amp;alt=media\" alt=\"XXX picture\"><br>\nHere I also use a Gaussian Process, with hyperparameters different from above. This one is very noisy though - it's not very far from not being regularized at all.</p>\n<p>The strategy is a BFGS pass with 1500 iterations, followed by a second pass that is terminated based on a gradient criterion. My public LB score for StyleB is around 44.4 - I was unable to find a way to capture any of the high-frequency features.</p>\n<h5>Other datasets</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Fef50f49c655a76f1c190ac17160647d3%2FScreenshot%202025-07-03%20165806.png?generation=1751554725245359&amp;alt=media\" alt=\"\"><br>\nThese are again tackled with a total variation prior, but now in 2D, so the cost function penalizes absolute differences between a point and its neighbors.</p>\n<p>The strategy is 3 BFGS passes. For the last one, I scan for areas where the velocity profile is almost flat, and restrict these to be exactly flat. Some specific datasets with only 2 flat areas are handled differently; I won't discuss this here. My public LB score for this set, which is the bulk of the data, is 4.9.</p>\n<h2>Summary</h2>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Prior</th>\n<th>Strategy</th>\n<th>Public LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FlatVelA+B</td>\n<td>1D total variation</td>\n<td>BFGS</td>\n<td>0.0</td>\n</tr>\n<tr>\n<td>StyleA</td>\n<td>Gaussian Process, SVD restricted</td>\n<td>BFGS -&gt; GN</td>\n<td>3.4</td>\n</tr>\n<tr>\n<td>StyleB</td>\n<td>Gaussian Process</td>\n<td>2x BFGS</td>\n<td>44.4</td>\n</tr>\n<tr>\n<td>Other</td>\n<td>2D total variation</td>\n<td>3x BFGS</td>\n<td>4.9</td>\n</tr>\n</tbody>\n</table>\n<h2>Other notes and things that didn't work</h2>\n<ul>\n<li>I tried to describe the StyleB prior using a variational auto-encoder, but didn't get it to see anything but noise.</li>\n<li>I tried various forms of preconditioning for BFGS and GN, but they never brought any benefit.</li>\n<li>My forward modeling CUDA kernels are still pretty inefficient with memory. A lot more speedup might be possible by merging the timesteps (requiring a cooperative kernel), and managing caches cleverly.</li>\n<li>I didn't try to improve the original deep learning model; I was surprised this could achieve 28.8, and didn't think a significant improvement would be possible…</li>\n<li>The forward modeling involves taking the minimum of the velocity map - problematic, because this is not differentiable. I solved this by adding the minimum velocity as an additional degree of freedom (so decoupling it from the velocity profile entirely). It is never penalized in the prior.</li>\n</ul>",
  "messages": [
    {
      "id": "3240208",
      "postDate": "07/03/2025 15:19:25",
      "content": "<p>Link to full code: <a href=\"https://www.kaggle.com/code/jeroencottaar/geophysical-waveform-inversion-2nd-place\" target=\"_blank\">https://www.kaggle.com/code/jeroencottaar/geophysical-waveform-inversion-2nd-place</a></p>\n<p>Thanks to the organizers for a fun competition, to Kaggle for providing the infrastructure, and to <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for providing an excellent starting point - I probably wouldn't even have competed without it.</p>\n<h2>Introduction</h2>\n<p>How do you find out what's under your feet without digging? You give the ground a good thump, which sends seismic waves underground. These waves interact with and reflect off underground structures. By placing multiple detectors at the surface, you can pick up the reflected waves. This measurement is called a seismogram.</p>\n<p>It's fairly straightforward to model the impact on the seismogram of subsurface structures, described by a velocity profile. The challenge lies in inverting this function: how do we find the original velocity profile given a seismogram? This was the challenge posed to us in this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F862fb07e92298b675b70e3088b322f28%2FScreenshot%202025-07-03%20171603.png?generation=1751555797803926&amp;alt=media\" alt=\"\"></p>\n<h2>Outline of my solution</h2>\n<p>I start with a rough deep learning model that would score 28.8 on the competition metric (this is the model <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> shared a few weeks ago). This is then refined using full waveform inversion (FWI) to a score of 7.6. The challenge was making this feasible in terms of convergence and computation time. The key elements of my solution:</p>\n<ul>\n<li>Carefully designed priors for various subsets of the data.</li>\n<li>Choosing and tweaking of solvers (BFGS and Gauss-Newton) per subset.</li>\n<li>Optimizing computation speed, including custom CUDA kernels for the forward model and BFGS overhead.</li>\n</ul>\n<p>Even then, my solution cost about $700 in cloud compute costs to develop and run. This is in line with other top solutions, but I'm hoping this competition is an outlier in this regard or we'll have very few competitors going for top spots soon - not to mention the environmental impact.</p>\n<h2>Details</h2>\n<p>FWI is much easier when you start close to the solution. For this I use a deep learning model. I didn't develop this myself, but use a model <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> graciously shared earlier in the competition. I modified nothing about it (and also didn't train it myself), so if you want to know more about it, you might as well get it from him: <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved</a></p>\n<p>From there I perform regularized full waveform inversion. This means I'm trying to find <em>x</em> that minimizes:<br>\n$$c(x)=p(x) + ||s(x)-m||^2_2$$<br>\nHere, <em>x</em> is the velocity profile, <em>p(x)</em> is the cost function defined by the prior, <em>s(x)</em> is the seismogram that follows from <em>x</em>, and <em>m</em> is the actual measured seismogram. In words: we're trying to balance minimizing a cost function and minimizing the residual.</p>\n<p>Typical methods to achieve this are BFGS and Gauss-Newton; I won't discuss these in detail. Both of these methods require us differentiate the cost function above; the most expensive part of this is computing the gradient of the seismogram function <em>\\triangledown s(x)</em>. I heavily optimized this with custom CUDA kernels; the forward pass takes 30 ms and the backward pass 70 ms on a V100 at double precision, though faster is probably possible. I also reimplemented parts of BFGS with a CUDA kernel to reduce overhead.</p>\n<p>What remains is the choice of prior <em>p(x)</em> and solving strategy, which took up the bulk of development time. Different priors are needed for the different families. I won't discuss how I classify them here; as many of you have noticed, this is a fairly easy problem.</p>\n<h5>FlatVelA and FlatVelB</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F8fc5ef9592231a41259cda052c71bae3%2FScreenshot%202025-07-03%20165936.png?generation=1751554792560180&amp;alt=media\" alt=\"\"><br>\nThe solution is restricted to 1D profiles, with total variation as the cost function:<br>\n$$ p(x)=\\sum_{i=1}^{69} |x_{i+1}-x_i| $$<br>\n(Actually, it's not quite the absolute value, but smoothed around 0 to remain differentiable.)</p>\n<p>I use a single BFGS pass, terminated when the cost function stops decreasing significantly. My public LB score for this one is well under 0.1.</p>\n<h5>StyleA</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F4025fbd7ae9124c93460f6bf1c0a94bc%2FScreenshot%202025-07-03%20165958.png?generation=1751554813255044&amp;alt=media\" alt=\"\"><br>\nThe prior for this one is a Gaussian Process with a squared-exponential kernel, and additional terms for noise, offset, and slope. Hyperparameters were tuned on the training set using maximum likelihood estimation. There is actually no noise in StyleA, which makes the covariance matrix <em>K</em> very ill-conditioned. I resolve this by restricting the solution to the strongest 1073 modes of an SVD decomposition of <em>K</em>.</p>\n<p>The strategy is a BFGS pass, further refined by a single Gauss-Newton step. My public LB score for this one is around 3.4.</p>\n<h5>StyleB</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F872cb2fbc53d8ca3975539ae406ce082%2FScreenshot%202025-07-03%20165908.png?generation=1751554770111041&amp;alt=media\" alt=\"XXX picture\"><br>\nHere I also use a Gaussian Process, with hyperparameters different from above. This one is very noisy though - it's not very far from not being regularized at all.</p>\n<p>The strategy is a BFGS pass with 1500 iterations, followed by a second pass that is terminated based on a gradient criterion. My public LB score for StyleB is around 44.4 - I was unable to find a way to capture any of the high-frequency features.</p>\n<h5>Other datasets</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Fef50f49c655a76f1c190ac17160647d3%2FScreenshot%202025-07-03%20165806.png?generation=1751554725245359&amp;alt=media\" alt=\"\"><br>\nThese are again tackled with a total variation prior, but now in 2D, so the cost function penalizes absolute differences between a point and its neighbors.</p>\n<p>The strategy is 3 BFGS passes. For the last one, I scan for areas where the velocity profile is almost flat, and restrict these to be exactly flat. Some specific datasets with only 2 flat areas are handled differently; I won't discuss this here. My public LB score for this set, which is the bulk of the data, is 4.9.</p>\n<h2>Summary</h2>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Prior</th>\n<th>Strategy</th>\n<th>Public LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FlatVelA+B</td>\n<td>1D total variation</td>\n<td>BFGS</td>\n<td>0.0</td>\n</tr>\n<tr>\n<td>StyleA</td>\n<td>Gaussian Process, SVD restricted</td>\n<td>BFGS -&gt; GN</td>\n<td>3.4</td>\n</tr>\n<tr>\n<td>StyleB</td>\n<td>Gaussian Process</td>\n<td>2x BFGS</td>\n<td>44.4</td>\n</tr>\n<tr>\n<td>Other</td>\n<td>2D total variation</td>\n<td>3x BFGS</td>\n<td>4.9</td>\n</tr>\n</tbody>\n</table>\n<h2>Other notes and things that didn't work</h2>\n<ul>\n<li>I tried to describe the StyleB prior using a variational auto-encoder, but didn't get it to see anything but noise.</li>\n<li>I tried various forms of preconditioning for BFGS and GN, but they never brought any benefit.</li>\n<li>My forward modeling CUDA kernels are still pretty inefficient with memory. A lot more speedup might be possible by merging the timesteps (requiring a cooperative kernel), and managing caches cleverly.</li>\n<li>I didn't try to improve the original deep learning model; I was surprised this could achieve 28.8, and didn't think a significant improvement would be possible…</li>\n<li>The forward modeling involves taking the minimum of the velocity map - problematic, because this is not differentiable. I solved this by adding the minimum velocity as an additional degree of freedom (so decoupling it from the velocity profile entirely). It is never penalized in the prior.</li>\n</ul>",
      "rawMarkdown": "Link to full code: https://www.kaggle.com/code/jeroencottaar/geophysical-waveform-inversion-2nd-place\n\nThanks to the organizers for a fun competition, to Kaggle for providing the infrastructure, and to @brendanartley for providing an excellent starting point - I probably wouldn't even have competed without it.\n## Introduction\n\nHow do you find out what's under your feet without digging? You give the ground a good thump, which sends seismic waves underground. These waves interact with and reflect off underground structures. By placing multiple detectors at the surface, you can pick up the reflected waves. This measurement is called a seismogram.\n\nIt's fairly straightforward to model the impact on the seismogram of subsurface structures, described by a velocity profile. The challenge lies in inverting this function: how do we find the original velocity profile given a seismogram? This was the challenge posed to us in this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F862fb07e92298b675b70e3088b322f28%2FScreenshot%202025-07-03%20171603.png?generation=1751555797803926&alt=media)\n## Outline of my solution\n\nI start with a rough deep learning model that would score 28.8 on the competition metric (this is the model @brendanartley shared a few weeks ago). This is then refined using full waveform inversion (FWI) to a score of 7.6. The challenge was making this feasible in terms of convergence and computation time. The key elements of my solution:\n- Carefully designed priors for various subsets of the data.\n- Choosing and tweaking of solvers (BFGS and Gauss-Newton) per subset.\n- Optimizing computation speed, including custom CUDA kernels for the forward model and BFGS overhead.\n\nEven then, my solution cost about $700 in cloud compute costs to develop and run. This is in line with other top solutions, but I'm hoping this competition is an outlier in this regard or we'll have very few competitors going for top spots soon - not to mention the environmental impact.\n## Details\n\nFWI is much easier when you start close to the solution. For this I use a deep learning model. I didn't develop this myself, but use a model @brendanartley graciously shared earlier in the competition. I modified nothing about it (and also didn't train it myself), so if you want to know more about it, you might as well get it from him: https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\n\nFrom there I perform regularized full waveform inversion. This means I'm trying to find *x* that minimizes:\n$$c(x)=p(x) + ||s(x)-m||^2_2$$\nHere, *x* is the velocity profile, *p(x)* is the cost function defined by the prior, *s(x)* is the seismogram that follows from *x*, and *m* is the actual measured seismogram. In words: we're trying to balance minimizing a cost function and minimizing the residual.\n\nTypical methods to achieve this are BFGS and Gauss-Newton; I won't discuss these in detail. Both of these methods require us differentiate the cost function above; the most expensive part of this is computing the gradient of the seismogram function *\\triangledown s(x)*. I heavily optimized this with custom CUDA kernels; the forward pass takes 30 ms and the backward pass 70 ms on a V100 at double precision, though faster is probably possible. I also reimplemented parts of BFGS with a CUDA kernel to reduce overhead.\n\nWhat remains is the choice of prior *p(x)* and solving strategy, which took up the bulk of development time. Different priors are needed for the different families. I won't discuss how I classify them here; as many of you have noticed, this is a fairly easy problem.\n##### FlatVelA and FlatVelB\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F8fc5ef9592231a41259cda052c71bae3%2FScreenshot%202025-07-03%20165936.png?generation=1751554792560180&alt=media)\nThe solution is restricted to 1D profiles, with total variation as the cost function:\n$$ p(x)=\\sum_{i=1}^{69} |x_{i+1}-x_i| $$\n(Actually, it's not quite the absolute value, but smoothed around 0 to remain differentiable.)\n\nI use a single BFGS pass, terminated when the cost function stops decreasing significantly. My public LB score for this one is well under 0.1.\n##### StyleA\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F4025fbd7ae9124c93460f6bf1c0a94bc%2FScreenshot%202025-07-03%20165958.png?generation=1751554813255044&alt=media)\nThe prior for this one is a Gaussian Process with a squared-exponential kernel, and additional terms for noise, offset, and slope. Hyperparameters were tuned on the training set using maximum likelihood estimation. There is actually no noise in StyleA, which makes the covariance matrix *K* very ill-conditioned. I resolve this by restricting the solution to the strongest 1073 modes of an SVD decomposition of *K*.\n\nThe strategy is a BFGS pass, further refined by a single Gauss-Newton step. My public LB score for this one is around 3.4.\n##### StyleB\n![XXX picture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F872cb2fbc53d8ca3975539ae406ce082%2FScreenshot%202025-07-03%20165908.png?generation=1751554770111041&alt=media)\nHere I also use a Gaussian Process, with hyperparameters different from above. This one is very noisy though - it's not very far from not being regularized at all.\n\nThe strategy is a BFGS pass with 1500 iterations, followed by a second pass that is terminated based on a gradient criterion. My public LB score for StyleB is around 44.4 - I was unable to find a way to capture any of the high-frequency features.\n##### Other datasets\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Fef50f49c655a76f1c190ac17160647d3%2FScreenshot%202025-07-03%20165806.png?generation=1751554725245359&alt=media)\nThese are again tackled with a total variation prior, but now in 2D, so the cost function penalizes absolute differences between a point and its neighbors.\n\nThe strategy is 3 BFGS passes. For the last one, I scan for areas where the velocity profile is almost flat, and restrict these to be exactly flat. Some specific datasets with only 2 flat areas are handled differently; I won't discuss this here. My public LB score for this set, which is the bulk of the data, is 4.9.\n\n## Summary\n\n| Dataset    | Prior                            | Strategy   | Public LB score |\n| ---------- | -------------------------------- | ---------- | --------------- |\n| FlatVelA+B | 1D total variation               | BFGS       | 0.0             |\n| StyleA     | Gaussian Process, SVD restricted | BFGS -> GN | 3.4             |\n| StyleB     | Gaussian Process                 | 2x BFGS    | 44.4            |\n| Other      | 2D total variation               | 3x BFGS    | 4.9             |\n\n\n## Other notes and things that didn't work\n\n- I tried to describe the StyleB prior using a variational auto-encoder, but didn't get it to see anything but noise.\n- I tried various forms of preconditioning for BFGS and GN, but they never brought any benefit.\n- My forward modeling CUDA kernels are still pretty inefficient with memory. A lot more speedup might be possible by merging the timesteps (requiring a cooperative kernel), and managing caches cleverly.\n- I didn't try to improve the original deep learning model; I was surprised this could achieve 28.8, and didn't think a significant improvement would be possible...\n- The forward modeling involves taking the minimum of the velocity map - problematic, because this is not differentiable. I solved this by adding the minimum velocity as an additional degree of freedom (so decoupling it from the velocity profile entirely). It is never penalized in the prior.",
      "votes": null
    },
    {
      "id": "3240214",
      "postDate": "07/03/2025 15:25:26",
      "content": "<p>Thanks for sharing, I need to dig a bit more to fully grasp this, but appreciate the post as it helps me learn.</p>",
      "rawMarkdown": "Thanks for sharing, I need to dig a bit more to fully grasp this, but appreciate the post as it helps me learn.",
      "votes": null
    },
    {
      "id": "3240444",
      "postDate": "07/03/2025 19:52:39",
      "content": "<p>Thanks for the solution. I know you are professional Bayesian, but I didn't know that can write CUDA kernel!! Congratulations for the 2nd! Will you join the Ariel competition this year again?</p>",
      "rawMarkdown": "Thanks for the solution. I know you are professional Bayesian, but I didn't know that can write CUDA kernel!! Congratulations for the 2nd! Will you join the Ariel competition this year again?",
      "votes": null
    },
    {
      "id": "3240448",
      "postDate": "07/03/2025 20:11:42",
      "content": "<p>Hi, congrats on the solution and the result! Bayesian methods are undervalued on Kaggle indeed. Thanks for showing us the way.</p>\n<p>My only question is this: you said in the forum that you probed the distribution of private vs public test data. How did you do that?</p>",
      "rawMarkdown": "Hi, congrats on the solution and the result! Bayesian methods are undervalued on Kaggle indeed. Thanks for showing us the way.\n\nMy only question is this: you said in the forum that you probed the distribution of private vs public test data. How did you do that?",
      "votes": null
    },
    {
      "id": "3240480",
      "postDate": "07/03/2025 20:54:30",
      "content": "<p>Suppose we've run a classification model and identified a subset of the full test test (public and private) as StyleB. How do we now find our score on the public test set for this subset?</p>\n<p>We make three submissions:<br>\n(A) a submission as normal.<br>\n(B) a submission where we add 10,000 to all StyleB sets.<br>\n(C) a submission where we add 5,000 to all StyleB sets.</p>\n<p>The idea is that by adding or subtracting a large number, we make all our errors have the same sign. This means we effectively remove whatever error we were making on StyleB, and replace it by 5,000 or 10,000. Since we've removed the unkown error and put a known one back, we can reconstruct the original error.</p>\n<p>The reason we need 3 submissions is because we have 3 unkowns to solver for:</p>\n<ul>\n<li>Our score on StyleB</li>\n<li>Our score on the rest of the data</li>\n<li>The proportion of the public test data that is StyleB (note that we don't know this from the full test set, since we don't know the public/private split).</li>\n</ul>\n<p>There's an additional complication if the mean of our error on StyleB isn't zero, but it's typically not really an issue. But it can be solved for as well if we add a fourth submission where we subtract 10,000.</p>",
      "rawMarkdown": "Suppose we've run a classification model and identified a subset of the full test test (public and private) as StyleB. How do we now find our score on the public test set for this subset?\n\nWe make three submissions:\n(A) a submission as normal.\n(B) a submission where we add 10,000 to all StyleB sets.\n(C) a submission where we add 5,000 to all StyleB sets.\n\nThe idea is that by adding or subtracting a large number, we make all our errors have the same sign. This means we effectively remove whatever error we were making on StyleB, and replace it by 5,000 or 10,000. Since we've removed the unkown error and put a known one back, we can reconstruct the original error.\n\nThe reason we need 3 submissions is because we have 3 unkowns to solver for:\n- Our score on StyleB\n- Our score on the rest of the data\n- The proportion of the public test data that is StyleB (note that we don't know this from the full test set, since we don't know the public/private split).\n\nThere's an additional complication if the mean of our error on StyleB isn't zero, but it's typically not really an issue. But it can be solved for as well if we add a fourth submission where we subtract 10,000.",
      "votes": null
    },
    {
      "id": "3240483",
      "postDate": "07/03/2025 20:55:56",
      "content": "<p>I learned to write CUDA for this competition - or more specifically, ChatGPT taught me. I'm kind of relieved my input was still needed - it couldn't get it right all by itself. Not sure if it'll stay that way though…</p>\n<p>I do hope to join the Ariel competition again!</p>",
      "rawMarkdown": "I learned to write CUDA for this competition - or more specifically, ChatGPT taught me. I'm kind of relieved my input was still needed - it couldn't get it right all by itself. Not sure if it'll stay that way though...\n\nI do hope to join the Ariel competition again!",
      "votes": null
    },
    {
      "id": "3240516",
      "postDate": "07/03/2025 22:38:32",
      "content": "<p>How come you use BFGS/Gauss-Newton rather than gradient descent? </p>",
      "rawMarkdown": "How come you use BFGS/Gauss-Newton rather than gradient descent?",
      "votes": null
    },
    {
      "id": "3240523",
      "postDate": "07/03/2025 22:52:15",
      "content": "<p>Also how much did adding the prior to your cost function improve score compared to only optimizing reconstruction loss</p>",
      "rawMarkdown": "Also how much did adding the prior to your cost function improve score compared to only optimizing reconstruction loss",
      "votes": null
    },
    {
      "id": "3240526",
      "postDate": "07/03/2025 22:58:57",
      "content": "<p>Very cool! 🙂</p>\n<p>Just to confirm my understanding -- this would work only if we have a near perfect family classifier?</p>",
      "rawMarkdown": "Very cool! 🙂\n\nJust to confirm my understanding -- this would work only if we have a near perfect family classifier?",
      "votes": null
    },
    {
      "id": "3240636",
      "postDate": "07/04/2025 04:43:54",
      "content": "<p>If the subset you want to probe is a specific family, the method is indeed limited by the accuracy of your classifier. But I don't think it needs to be near-perfect - generally a few percent error in the probed score shouldn't really matter.</p>",
      "rawMarkdown": "If the subset you want to probe is a specific family, the method is indeed limited by the accuracy of your classifier. But I don't think it needs to be near-perfect - generally a few percent error in the probed score shouldn't really matter.",
      "votes": null
    },
    {
      "id": "3240637",
      "postDate": "07/04/2025 04:44:48",
      "content": "<p>Without having a prior, the method generally doesn't converge to begin with, so I'd be stuck with the initial deep learning solution. So I could say it brings the score from 28.8 to 7.6.</p>",
      "rawMarkdown": "Without having a prior, the method generally doesn't converge to begin with, so I'd be stuck with the initial deep learning solution. So I could say it brings the score from 28.8 to 7.6.",
      "votes": null
    },
    {
      "id": "3240642",
      "postDate": "07/04/2025 04:48:41",
      "content": "<p>Because we're now used to training neural networks, gradient descent has become the default choice for non-linear optimization problems for many people. But for lower-dimensional problems it's not usually the best choice.</p>\n<p>The essential thing that it misses is using the second derivative (Hessian). I have to say I didn't even try gradient descent, but intuitively I expect it to be 10x slower or worse. But I do think I should have tried it out actually.</p>",
      "rawMarkdown": "Because we're now used to training neural networks, gradient descent has become the default choice for non-linear optimization problems for many people. But for lower-dimensional problems it's not usually the best choice.\n\nThe essential thing that it misses is using the second derivative (Hessian). I have to say I didn't even try gradient descent, but intuitively I expect it to be 10x slower or worse. But I do think I should have tried it out actually.",
      "votes": null
    },
    {
      "id": "3240648",
      "postDate": "07/04/2025 05:02:22",
      "content": "<p>Thanks and congrats! </p>",
      "rawMarkdown": "Thanks and congrats!",
      "votes": null
    },
    {
      "id": "3240872",
      "postDate": "07/04/2025 09:59:03",
      "content": "<blockquote>\n  <p>this would work only if we have a near perfect family classifier?</p>\n</blockquote>\n<p>it was rather easy to get very good classifiers. I got 99.55% accuracy for 10 classes on my validation data for instance.</p>",
      "rawMarkdown": "> this would work only if we have a near perfect family classifier?\n\nit was rather easy to get very good classifiers. I got 99.55% accuracy for 10 classes on my validation data for instance.",
      "votes": null
    },
    {
      "id": "3240873",
      "postDate": "07/04/2025 10:00:56",
      "content": "<blockquote>\n  <p>if the mean of our error on StyleB isn't zero,</p>\n</blockquote>\n<p>I checked that the mean was very close to zero, as expected, on validation data.</p>",
      "rawMarkdown": "> if the mean of our error on StyleB isn't zero,\n\nI checked that the mean was very close to zero, as expected, on validation data.",
      "votes": null
    },
    {
      "id": "3240879",
      "postDate": "07/04/2025 10:07:55",
      "content": "<p>I tried without priors and could get a 0.1 improvement or so before it started diverging.</p>",
      "rawMarkdown": "I tried without priors and could get a 0.1 improvement or so before it started diverging.",
      "votes": null
    },
    {
      "id": "3241135",
      "postDate": "07/04/2025 15:20:08",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> , I tried to do something like this but gave up way too early because I had no idea how to, the implementation of priors seems to be the missing key, very smart.</p>\n<p>What happens if you get a better initial guess to begin with instead of 28.8, will it have significant affect on the score?</p>",
      "rawMarkdown": "Congratulations @jeroencottaar , I tried to do something like this but gave up way too early because I had no idea how to, the implementation of priors seems to be the missing key, very smart.\n\nWhat happens if you get a better initial guess to begin with instead of 28.8, will it have significant affect on the score?",
      "votes": null
    },
    {
      "id": "3241144",
      "postDate": "07/04/2025 15:28:02",
      "content": "<p>I'm not sure, but I expect the main gain will be in computation time. There's a floor to how well my approach can score - enforcing the prior also means that the correct solution is <em>not</em> a minimizer of the cost function; even if you start in the correct solution, you'll move away from it anyway.</p>\n<p>For StyleB specifically there is something else going on though: my method falls into a bad local minimum. Perhaps a better starting point would get it moving in the right direction instead. In that case the score might improve.</p>",
      "rawMarkdown": "I'm not sure, but I expect the main gain will be in computation time. There's a floor to how well my approach can score - enforcing the prior also means that the correct solution is *not* a minimizer of the cost function; even if you start in the correct solution, you'll move away from it anyway.\n\nFor StyleB specifically there is something else going on though: my method falls into a bad local minimum. Perhaps a better starting point would get it moving in the right direction instead. In that case the score might improve.",
      "votes": null
    },
    {
      "id": "3252916",
      "postDate": "07/23/2025 16:58:50",
      "content": "<p>Oh, I missed that you posted you solution!<br>\nIt's even more strange than I expected. You are the master of strange solutions that I don't understand at all 🤣<br>\nI have a Déjà vu from your solution to Ariel…<br>\nMaybe someday I will understand it.</p>",
      "rawMarkdown": "Oh, I missed that you posted you solution!\nIt's even more strange than I expected. You are the master of strange solutions that I don't understand at all 🤣\nI have a Déjà vu from your solution to Ariel...\nMaybe someday I will understand it.",
      "votes": null
    },
    {
      "id": "3407424",
      "postDate": "02/18/2026 10:25:09",
      "content": "<p>Hi there,\nthanks for the writeups which helps us to grow and learn in data science community.\nThank you sir.</p>",
      "rawMarkdown": "Hi there,\nthanks for the writeups which helps us to grow and learn in data science community.\nThank you sir.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3240214,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/03/2025 15:25:26",
      "content": "<p>Thanks for sharing, I need to dig a bit more to fully grasp this, but appreciate the post as it helps me learn.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3240444,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "07/03/2025 19:52:39",
      "content": "<p>Thanks for the solution. I know you are professional Bayesian, but I didn't know that can write CUDA kernel!! Congratulations for the 2nd! Will you join the Ariel competition this year again?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3240483,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "07/03/2025 20:55:56",
          "content": "<p>I learned to write CUDA for this competition - or more specifically, ChatGPT taught me. I'm kind of relieved my input was still needed - it couldn't get it right all by itself. Not sure if it'll stay that way though…</p>\n<p>I do hope to join the Ariel competition again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3240448,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/03/2025 20:11:42",
      "content": "<p>Hi, congrats on the solution and the result! Bayesian methods are undervalued on Kaggle indeed. Thanks for showing us the way.</p>\n<p>My only question is this: you said in the forum that you probed the distribution of private vs public test data. How did you do that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3240480,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "07/03/2025 20:54:30",
          "content": "<p>Suppose we've run a classification model and identified a subset of the full test test (public and private) as StyleB. How do we now find our score on the public test set for this subset?</p>\n<p>We make three submissions:<br>\n(A) a submission as normal.<br>\n(B) a submission where we add 10,000 to all StyleB sets.<br>\n(C) a submission where we add 5,000 to all StyleB sets.</p>\n<p>The idea is that by adding or subtracting a large number, we make all our errors have the same sign. This means we effectively remove whatever error we were making on StyleB, and replace it by 5,000 or 10,000. Since we've removed the unkown error and put a known one back, we can reconstruct the original error.</p>\n<p>The reason we need 3 submissions is because we have 3 unkowns to solver for:</p>\n<ul>\n<li>Our score on StyleB</li>\n<li>Our score on the rest of the data</li>\n<li>The proportion of the public test data that is StyleB (note that we don't know this from the full test set, since we don't know the public/private split).</li>\n</ul>\n<p>There's an additional complication if the mean of our error on StyleB isn't zero, but it's typically not really an issue. But it can be solved for as well if we add a fourth submission where we subtract 10,000.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3240526,
              "author_name": "radek1",
              "author_url": "",
              "post_date": "07/03/2025 22:58:57",
              "content": "<p>Very cool! 🙂</p>\n<p>Just to confirm my understanding -- this would work only if we have a near perfect family classifier?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3240636,
                  "author_name": "jeroencottaar",
                  "author_url": "",
                  "post_date": "07/04/2025 04:43:54",
                  "content": "<p>If the subset you want to probe is a specific family, the method is indeed limited by the accuracy of your classifier. But I don't think it needs to be near-perfect - generally a few percent error in the probed score shouldn't really matter.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3240872,
                      "author_name": "cpmpml",
                      "author_url": "",
                      "post_date": "07/04/2025 09:59:03",
                      "content": "<blockquote>\n  <p>this would work only if we have a near perfect family classifier?</p>\n</blockquote>\n<p>it was rather easy to get very good classifiers. I got 99.55% accuracy for 10 classes on my validation data for instance.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3240873,
                          "author_name": "cpmpml",
                          "author_url": "",
                          "post_date": "07/04/2025 10:00:56",
                          "content": "<blockquote>\n  <p>if the mean of our error on StyleB isn't zero,</p>\n</blockquote>\n<p>I checked that the mean was very close to zero, as expected, on validation data.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3240516,
      "author_name": "snehalverma10",
      "author_url": "",
      "post_date": "07/03/2025 22:38:32",
      "content": "<p>How come you use BFGS/Gauss-Newton rather than gradient descent? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3240642,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "07/04/2025 04:48:41",
          "content": "<p>Because we're now used to training neural networks, gradient descent has become the default choice for non-linear optimization problems for many people. But for lower-dimensional problems it's not usually the best choice.</p>\n<p>The essential thing that it misses is using the second derivative (Hessian). I have to say I didn't even try gradient descent, but intuitively I expect it to be 10x slower or worse. But I do think I should have tried it out actually.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3240648,
              "author_name": "snehalverma10",
              "author_url": "",
              "post_date": "07/04/2025 05:02:22",
              "content": "<p>Thanks and congrats! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3240523,
      "author_name": "snehalverma10",
      "author_url": "",
      "post_date": "07/03/2025 22:52:15",
      "content": "<p>Also how much did adding the prior to your cost function improve score compared to only optimizing reconstruction loss</p>",
      "votes": null,
      "replies": [
        {
          "id": 3240637,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "07/04/2025 04:44:48",
          "content": "<p>Without having a prior, the method generally doesn't converge to begin with, so I'd be stuck with the initial deep learning solution. So I could say it brings the score from 28.8 to 7.6.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3240879,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "07/04/2025 10:07:55",
              "content": "<p>I tried without priors and could get a 0.1 improvement or so before it started diverging.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3241135,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "07/04/2025 15:20:08",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> , I tried to do something like this but gave up way too early because I had no idea how to, the implementation of priors seems to be the missing key, very smart.</p>\n<p>What happens if you get a better initial guess to begin with instead of 28.8, will it have significant affect on the score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3241144,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "07/04/2025 15:28:02",
          "content": "<p>I'm not sure, but I expect the main gain will be in computation time. There's a floor to how well my approach can score - enforcing the prior also means that the correct solution is <em>not</em> a minimizer of the cost function; even if you start in the correct solution, you'll move away from it anyway.</p>\n<p>For StyleB specifically there is something else going on though: my method falls into a bad local minimum. Perhaps a better starting point would get it moving in the right direction instead. In that case the score might improve.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3252916,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "07/23/2025 16:58:50",
      "content": "<p>Oh, I missed that you posted you solution!<br>\nIt's even more strange than I expected. You are the master of strange solutions that I don't understand at all 🤣<br>\nI have a Déjà vu from your solution to Ariel…<br>\nMaybe someday I will understand it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3407424,
      "author_name": "navjyotarchitect",
      "author_url": "",
      "post_date": "02/18/2026 10:25:09",
      "content": "<p>Hi there,\nthanks for the writeups which helps us to grow and learn in data science community.\nThank you sir.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3240208": "Link to full code: https://www.kaggle.com/code/jeroencottaar/geophysical-waveform-inversion-2nd-place\n\nThanks to the organizers for a fun competition, to Kaggle for providing the infrastructure, and to @brendanartley for providing an excellent starting point - I probably wouldn't even have competed without it.\n## Introduction\n\nHow do you find out what's under your feet without digging? You give the ground a good thump, which sends seismic waves underground. These waves interact with and reflect off underground structures. By placing multiple detectors at the surface, you can pick up the reflected waves. This measurement is called a seismogram.\n\nIt's fairly straightforward to model the impact on the seismogram of subsurface structures, described by a velocity profile. The challenge lies in inverting this function: how do we find the original velocity profile given a seismogram? This was the challenge posed to us in this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F862fb07e92298b675b70e3088b322f28%2FScreenshot%202025-07-03%20171603.png?generation=1751555797803926&alt=media)\n## Outline of my solution\n\nI start with a rough deep learning model that would score 28.8 on the competition metric (this is the model @brendanartley shared a few weeks ago). This is then refined using full waveform inversion (FWI) to a score of 7.6. The challenge was making this feasible in terms of convergence and computation time. The key elements of my solution:\n- Carefully designed priors for various subsets of the data.\n- Choosing and tweaking of solvers (BFGS and Gauss-Newton) per subset.\n- Optimizing computation speed, including custom CUDA kernels for the forward model and BFGS overhead.\n\nEven then, my solution cost about $700 in cloud compute costs to develop and run. This is in line with other top solutions, but I'm hoping this competition is an outlier in this regard or we'll have very few competitors going for top spots soon - not to mention the environmental impact.\n## Details\n\nFWI is much easier when you start close to the solution. For this I use a deep learning model. I didn't develop this myself, but use a model @brendanartley graciously shared earlier in the competition. I modified nothing about it (and also didn't train it myself), so if you want to know more about it, you might as well get it from him: https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\n\nFrom there I perform regularized full waveform inversion. This means I'm trying to find *x* that minimizes:\n$$c(x)=p(x) + ||s(x)-m||^2_2$$\nHere, *x* is the velocity profile, *p(x)* is the cost function defined by the prior, *s(x)* is the seismogram that follows from *x*, and *m* is the actual measured seismogram. In words: we're trying to balance minimizing a cost function and minimizing the residual.\n\nTypical methods to achieve this are BFGS and Gauss-Newton; I won't discuss these in detail. Both of these methods require us differentiate the cost function above; the most expensive part of this is computing the gradient of the seismogram function *\\triangledown s(x)*. I heavily optimized this with custom CUDA kernels; the forward pass takes 30 ms and the backward pass 70 ms on a V100 at double precision, though faster is probably possible. I also reimplemented parts of BFGS with a CUDA kernel to reduce overhead.\n\nWhat remains is the choice of prior *p(x)* and solving strategy, which took up the bulk of development time. Different priors are needed for the different families. I won't discuss how I classify them here; as many of you have noticed, this is a fairly easy problem.\n##### FlatVelA and FlatVelB\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F8fc5ef9592231a41259cda052c71bae3%2FScreenshot%202025-07-03%20165936.png?generation=1751554792560180&alt=media)\nThe solution is restricted to 1D profiles, with total variation as the cost function:\n$$ p(x)=\\sum_{i=1}^{69} |x_{i+1}-x_i| $$\n(Actually, it's not quite the absolute value, but smoothed around 0 to remain differentiable.)\n\nI use a single BFGS pass, terminated when the cost function stops decreasing significantly. My public LB score for this one is well under 0.1.\n##### StyleA\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F4025fbd7ae9124c93460f6bf1c0a94bc%2FScreenshot%202025-07-03%20165958.png?generation=1751554813255044&alt=media)\nThe prior for this one is a Gaussian Process with a squared-exponential kernel, and additional terms for noise, offset, and slope. Hyperparameters were tuned on the training set using maximum likelihood estimation. There is actually no noise in StyleA, which makes the covariance matrix *K* very ill-conditioned. I resolve this by restricting the solution to the strongest 1073 modes of an SVD decomposition of *K*.\n\nThe strategy is a BFGS pass, further refined by a single Gauss-Newton step. My public LB score for this one is around 3.4.\n##### StyleB\n![XXX picture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F872cb2fbc53d8ca3975539ae406ce082%2FScreenshot%202025-07-03%20165908.png?generation=1751554770111041&alt=media)\nHere I also use a Gaussian Process, with hyperparameters different from above. This one is very noisy though - it's not very far from not being regularized at all.\n\nThe strategy is a BFGS pass with 1500 iterations, followed by a second pass that is terminated based on a gradient criterion. My public LB score for StyleB is around 44.4 - I was unable to find a way to capture any of the high-frequency features.\n##### Other datasets\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Fef50f49c655a76f1c190ac17160647d3%2FScreenshot%202025-07-03%20165806.png?generation=1751554725245359&alt=media)\nThese are again tackled with a total variation prior, but now in 2D, so the cost function penalizes absolute differences between a point and its neighbors.\n\nThe strategy is 3 BFGS passes. For the last one, I scan for areas where the velocity profile is almost flat, and restrict these to be exactly flat. Some specific datasets with only 2 flat areas are handled differently; I won't discuss this here. My public LB score for this set, which is the bulk of the data, is 4.9.\n\n## Summary\n\n| Dataset    | Prior                            | Strategy   | Public LB score |\n| ---------- | -------------------------------- | ---------- | --------------- |\n| FlatVelA+B | 1D total variation               | BFGS       | 0.0             |\n| StyleA     | Gaussian Process, SVD restricted | BFGS -> GN | 3.4             |\n| StyleB     | Gaussian Process                 | 2x BFGS    | 44.4            |\n| Other      | 2D total variation               | 3x BFGS    | 4.9             |\n\n\n## Other notes and things that didn't work\n\n- I tried to describe the StyleB prior using a variational auto-encoder, but didn't get it to see anything but noise.\n- I tried various forms of preconditioning for BFGS and GN, but they never brought any benefit.\n- My forward modeling CUDA kernels are still pretty inefficient with memory. A lot more speedup might be possible by merging the timesteps (requiring a cooperative kernel), and managing caches cleverly.\n- I didn't try to improve the original deep learning model; I was surprised this could achieve 28.8, and didn't think a significant improvement would be possible...\n- The forward modeling involves taking the minimum of the velocity map - problematic, because this is not differentiable. I solved this by adding the minimum velocity as an additional degree of freedom (so decoupling it from the velocity profile entirely). It is never penalized in the prior.",
    "3240214": "Thanks for sharing, I need to dig a bit more to fully grasp this, but appreciate the post as it helps me learn.",
    "3240444": "Thanks for the solution. I know you are professional Bayesian, but I didn't know that can write CUDA kernel!! Congratulations for the 2nd! Will you join the Ariel competition this year again?",
    "3240448": "Hi, congrats on the solution and the result! Bayesian methods are undervalued on Kaggle indeed. Thanks for showing us the way.\n\nMy only question is this: you said in the forum that you probed the distribution of private vs public test data. How did you do that?",
    "3240480": "Suppose we've run a classification model and identified a subset of the full test test (public and private) as StyleB. How do we now find our score on the public test set for this subset?\n\nWe make three submissions:\n(A) a submission as normal.\n(B) a submission where we add 10,000 to all StyleB sets.\n(C) a submission where we add 5,000 to all StyleB sets.\n\nThe idea is that by adding or subtracting a large number, we make all our errors have the same sign. This means we effectively remove whatever error we were making on StyleB, and replace it by 5,000 or 10,000. Since we've removed the unkown error and put a known one back, we can reconstruct the original error.\n\nThe reason we need 3 submissions is because we have 3 unkowns to solver for:\n- Our score on StyleB\n- Our score on the rest of the data\n- The proportion of the public test data that is StyleB (note that we don't know this from the full test set, since we don't know the public/private split).\n\nThere's an additional complication if the mean of our error on StyleB isn't zero, but it's typically not really an issue. But it can be solved for as well if we add a fourth submission where we subtract 10,000.",
    "3240483": "I learned to write CUDA for this competition - or more specifically, ChatGPT taught me. I'm kind of relieved my input was still needed - it couldn't get it right all by itself. Not sure if it'll stay that way though...\n\nI do hope to join the Ariel competition again!",
    "3240516": "How come you use BFGS/Gauss-Newton rather than gradient descent?",
    "3240523": "Also how much did adding the prior to your cost function improve score compared to only optimizing reconstruction loss",
    "3240526": "Very cool! 🙂\n\nJust to confirm my understanding -- this would work only if we have a near perfect family classifier?",
    "3240636": "If the subset you want to probe is a specific family, the method is indeed limited by the accuracy of your classifier. But I don't think it needs to be near-perfect - generally a few percent error in the probed score shouldn't really matter.",
    "3240637": "Without having a prior, the method generally doesn't converge to begin with, so I'd be stuck with the initial deep learning solution. So I could say it brings the score from 28.8 to 7.6.",
    "3240642": "Because we're now used to training neural networks, gradient descent has become the default choice for non-linear optimization problems for many people. But for lower-dimensional problems it's not usually the best choice.\n\nThe essential thing that it misses is using the second derivative (Hessian). I have to say I didn't even try gradient descent, but intuitively I expect it to be 10x slower or worse. But I do think I should have tried it out actually.",
    "3240648": "Thanks and congrats!",
    "3240872": "> this would work only if we have a near perfect family classifier?\n\nit was rather easy to get very good classifiers. I got 99.55% accuracy for 10 classes on my validation data for instance.",
    "3240873": "> if the mean of our error on StyleB isn't zero,\n\nI checked that the mean was very close to zero, as expected, on validation data.",
    "3240879": "I tried without priors and could get a 0.1 improvement or so before it started diverging.",
    "3241135": "Congratulations @jeroencottaar , I tried to do something like this but gave up way too early because I had no idea how to, the implementation of priors seems to be the missing key, very smart.\n\nWhat happens if you get a better initial guess to begin with instead of 28.8, will it have significant affect on the score?",
    "3241144": "I'm not sure, but I expect the main gain will be in computation time. There's a floor to how well my approach can score - enforcing the prior also means that the correct solution is *not* a minimizer of the cost function; even if you start in the correct solution, you'll move away from it anyway.\n\nFor StyleB specifically there is something else going on though: my method falls into a bad local minimum. Perhaps a better starting point would get it moving in the right direction instead. In that case the score might improve.",
    "3252916": "Oh, I missed that you posted you solution!\nIt's even more strange than I expected. You are the master of strange solutions that I don't understand at all 🤣\nI have a Déjà vu from your solution to Ariel...\nMaybe someday I will understand it.",
    "3407424": "Hi there,\nthanks for the writeups which helps us to grow and learn in data science community.\nThank you sir."
  },
  "source": "meta"
}