{
  "id": 587395,
  "title": "102nd Place Solution - Clustering And Refinement with Reconstruction Loss To Beat Public Codes",
  "url": "/competitions/waveform-inversion/discussion/587395",
  "author_name": "Harui-ig",
  "post_date": "2025-07-01T00:23:38.979000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to express my deepest gratitude to the competition organizers and all participants.</p>\n<h1>Our approach</h1>\n<h2>Solution Summary</h2>\n<p>In short, our solution is based on refining velocity maps using clustering and VelToSeisLoss, rather than purely training DNN models.</p>\n<ol>\n<li>Utilize public csvs</li>\n<li>Clustering velocity maps and replace cluster's value with median.</li>\n<li>For FlatVel data, refine velocity map using VelToSeisLoss.</li>\n</ol>\n<h2>Background</h2>\n<p>As many of you know, there were so many strong public codes. For example, Bartley's CAFormer achieved 28.8 MAE on the public LB(<a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved)\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved)</a>,  and Atom's ensemble even achieved an 25.6 MAE on the LB(<a href=\"https://www.kaggle.com/code/atom1231/ensemble-csv-for-more-files)\" target=\"_blank\">https://www.kaggle.com/code/atom1231/ensemble-csv-for-more-files)</a>.</p>\n<p>In short, we could not beat these public codes by training DNN models, a topic I will discuss this later.</p>\n<p>So, we adopted a refinement approach.</p>\n<h2>Refinement Procedure</h2>\n<p>Step1: First, we calculate or retrieve an initial velocity map, using models or public CSV files.</p>\n<p>Step2: Secondly, we apply a clustering algorithm to the velocity map, and replace each cluster's velocity value with its median of velocity value. <br>\nFor Style A/B datasets, this approach worsen the solution, so we did not apply the replacement when there were so many clusters; this is a sign of Style datasets.</p>\n<p>This median replacement improved the MAE from 25.6 to 25.4.</p>\n<p>As <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/585815\" target=\"_blank\">gguillard's discussion</a> pointed out, we can determine whether seismic data is in the FlatVel dataset by checking its symmetrisity among the different shots.</p>\n<p>Using this knowledge, we can improve the clustering better. You can then decide to put all the grids at the same z value into the same cluster.</p>\n<p>After clustering, the number of the components of typical FlatVel dataset ranges from 3 to 9. (Without clustering, there were 4900 components!)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F8d39f14d44f01e72f74feaf1f5ddf226%2FCEasy_0.63_to_0.37_clustered_vel_map.png?generation=1751329382973959&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F8299d4f7256127735312c8b7754b1991%2FCEasy_7.70_to_7.03_clustered_vel_map.png?generation=1751329408325463&amp;alt=media\" alt=\"\"></p>\n<p>Step3: Refinement using VelToSeisLoss(Physics Loss)</p>\n<p>This powerful constraints enable us to use VelToSeisLoss.(without the constraints, the process can easily become trapped in a local minimum) For example, I reduced MAE loss for a FlatVel data from 9.4 to 5.1 in 40 iteration steps(it took about a few minutes). However, backpropagation of forward modeling requires so much computational resources, we could only do 2 steps for each test data. (It alsore requires so high VRAM capacity that I could not do refinement without dividing shots and doing gradient accumulation.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F07d4e674a0f74f3685732c3f6e937d24%2FScreenshot%20from%202025-07-01%2009-33-44.png?generation=1751330050355324&amp;alt=media\" alt=\"\"></p>\n<h1>What didn't work</h1>\n<ul>\n<li>training UNet models with backbones(ConvNeXtV2Base/Large)(not strong enough to beat bartley's CAFormer)</li>\n<li>finetuning Bartley's CAFormer: finetuning even worsened validation loss, it may be because I didn't care about cv folds</li>\n</ul>\n<h1>Lessons what we learned</h1>\n<ul>\n<li>bigger model is better: given a fixed GPU/TPU hours, a bigger model achieved a better validation loss in my experiments ( 224x224 ConvNextV2Base &lt; 388x388 ConvNeXtV2Base &lt; 388x388ConvNeXtV2Large)</li>\n<li>Perceptual loss improved a speed of convergence at the beginning of training</li>\n<li>Pretrained Weights can improve loss about 10 % even when domains are totally different</li>\n<li>Google Cloud's storage costs too much for me</li>\n<li>TPU v4s are really fast and TPU Research Cloud is very helpful project</li>\n</ul>",
  "messages": [
    {
      "id": 3237155,
      "postDate": "2025-07-01T00:23:38.980Z",
      "content": "<p>First of all, I would like to express my deepest gratitude to the competition organizers and all participants.</p>\n<h1>Our approach</h1>\n<h2>Solution Summary</h2>\n<p>In short, our solution is based on refining velocity maps using clustering and VelToSeisLoss, rather than purely training DNN models.</p>\n<ol>\n<li>Utilize public csvs</li>\n<li>Clustering velocity maps and replace cluster's value with median.</li>\n<li>For FlatVel data, refine velocity map using VelToSeisLoss.</li>\n</ol>\n<h2>Background</h2>\n<p>As many of you know, there were so many strong public codes. For example, Bartley's CAFormer achieved 28.8 MAE on the public LB(<a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved)\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved)</a>,  and Atom's ensemble even achieved an 25.6 MAE on the LB(<a href=\"https://www.kaggle.com/code/atom1231/ensemble-csv-for-more-files)\" target=\"_blank\">https://www.kaggle.com/code/atom1231/ensemble-csv-for-more-files)</a>.</p>\n<p>In short, we could not beat these public codes by training DNN models, a topic I will discuss this later.</p>\n<p>So, we adopted a refinement approach.</p>\n<h2>Refinement Procedure</h2>\n<p>Step1: First, we calculate or retrieve an initial velocity map, using models or public CSV files.</p>\n<p>Step2: Secondly, we apply a clustering algorithm to the velocity map, and replace each cluster's velocity value with its median of velocity value. <br>\nFor Style A/B datasets, this approach worsen the solution, so we did not apply the replacement when there were so many clusters; this is a sign of Style datasets.</p>\n<p>This median replacement improved the MAE from 25.6 to 25.4.</p>\n<p>As <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/585815\" target=\"_blank\">gguillard's discussion</a> pointed out, we can determine whether seismic data is in the FlatVel dataset by checking its symmetrisity among the different shots.</p>\n<p>Using this knowledge, we can improve the clustering better. You can then decide to put all the grids at the same z value into the same cluster.</p>\n<p>After clustering, the number of the components of typical FlatVel dataset ranges from 3 to 9. (Without clustering, there were 4900 components!)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F8d39f14d44f01e72f74feaf1f5ddf226%2FCEasy_0.63_to_0.37_clustered_vel_map.png?generation=1751329382973959&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F8299d4f7256127735312c8b7754b1991%2FCEasy_7.70_to_7.03_clustered_vel_map.png?generation=1751329408325463&amp;alt=media\" alt=\"\"></p>\n<p>Step3: Refinement using VelToSeisLoss(Physics Loss)</p>\n<p>This powerful constraints enable us to use VelToSeisLoss.(without the constraints, the process can easily become trapped in a local minimum) For example, I reduced MAE loss for a FlatVel data from 9.4 to 5.1 in 40 iteration steps(it took about a few minutes). However, backpropagation of forward modeling requires so much computational resources, we could only do 2 steps for each test data. (It alsore requires so high VRAM capacity that I could not do refinement without dividing shots and doing gradient accumulation.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F07d4e674a0f74f3685732c3f6e937d24%2FScreenshot%20from%202025-07-01%2009-33-44.png?generation=1751330050355324&amp;alt=media\" alt=\"\"></p>\n<h1>What didn't work</h1>\n<ul>\n<li>training UNet models with backbones(ConvNeXtV2Base/Large)(not strong enough to beat bartley's CAFormer)</li>\n<li>finetuning Bartley's CAFormer: finetuning even worsened validation loss, it may be because I didn't care about cv folds</li>\n</ul>\n<h1>Lessons what we learned</h1>\n<ul>\n<li>bigger model is better: given a fixed GPU/TPU hours, a bigger model achieved a better validation loss in my experiments ( 224x224 ConvNextV2Base &lt; 388x388 ConvNeXtV2Base &lt; 388x388ConvNeXtV2Large)</li>\n<li>Perceptual loss improved a speed of convergence at the beginning of training</li>\n<li>Pretrained Weights can improve loss about 10 % even when domains are totally different</li>\n<li>Google Cloud's storage costs too much for me</li>\n<li>TPU v4s are really fast and TPU Research Cloud is very helpful project</li>\n</ul>",
      "rawMarkdown": "First of all, I would like to express my deepest gratitude to the competition organizers and all participants.\n\n\n# Our approach\n\n## Solution Summary\n\nIn short, our solution is based on refining velocity maps using clustering and VelToSeisLoss, rather than purely training DNN models.\n\n1. Utilize public csvs\n2. Clustering velocity maps and replace cluster's value with median.\n3. For FlatVel data, refine velocity map using VelToSeisLoss.\n\n## Background\n\nAs many of you know, there were so many strong public codes. For example, Bartley's CAFormer achieved 28.8 MAE on the public LB(https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved),  and Atom's ensemble even achieved an 25.6 MAE on the LB(https://www.kaggle.com/code/atom1231/ensemble-csv-for-more-files).\n\nIn short, we could not beat these public codes by training DNN models, a topic I will discuss this later.\n\nSo, we adopted a refinement approach.\n\n## Refinement Procedure\n\nStep1: First, we calculate or retrieve an initial velocity map, using models or public CSV files.\n\nStep2: Secondly, we apply a clustering algorithm to the velocity map, and replace each cluster's velocity value with its median of velocity value. \nFor Style A/B datasets, this approach worsen the solution, so we did not apply the replacement when there were so many clusters; this is a sign of Style datasets.\n\nThis median replacement improved the MAE from 25.6 to 25.4.\n\nAs [gguillard's discussion](https://www.kaggle.com/competitions/waveform-inversion/discussion/585815) pointed out, we can determine whether seismic data is in the FlatVel dataset by checking its symmetrisity among the different shots.\n\nUsing this knowledge, we can improve the clustering better. You can then decide to put all the grids at the same z value into the same cluster.\n\nAfter clustering, the number of the components of typical FlatVel dataset ranges from 3 to 9. (Without clustering, there were 4900 components!)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F8d39f14d44f01e72f74feaf1f5ddf226%2FCEasy_0.63_to_0.37_clustered_vel_map.png?generation=1751329382973959&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F8299d4f7256127735312c8b7754b1991%2FCEasy_7.70_to_7.03_clustered_vel_map.png?generation=1751329408325463&alt=media)\n\nStep3: Refinement using VelToSeisLoss(Physics Loss)\n\nThis powerful constraints enable us to use VelToSeisLoss.(without the constraints, the process can easily become trapped in a local minimum) For example, I reduced MAE loss for a FlatVel data from 9.4 to 5.1 in 40 iteration steps(it took about a few minutes). However, backpropagation of forward modeling requires so much computational resources, we could only do 2 steps for each test data. (It alsore requires so high VRAM capacity that I could not do refinement without dividing shots and doing gradient accumulation.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F07d4e674a0f74f3685732c3f6e937d24%2FScreenshot%20from%202025-07-01%2009-33-44.png?generation=1751330050355324&alt=media)\n\n# What didn't work\n\n- training UNet models with backbones(ConvNeXtV2Base/Large)(not strong enough to beat bartley's CAFormer)\n- finetuning Bartley's CAFormer: finetuning even worsened validation loss, it may be because I didn't care about cv folds\n\n\n# Lessons what we learned\n\n- bigger model is better: given a fixed GPU/TPU hours, a bigger model achieved a better validation loss in my experiments ( 224x224 ConvNextV2Base < 388x388 ConvNeXtV2Base < 388x388ConvNeXtV2Large)\n- Perceptual loss improved a speed of convergence at the beginning of training\n- Pretrained Weights can improve loss about 10 % even when domains are totally different\n- Google Cloud's storage costs too much for me\n- TPU v4s are really fast and TPU Research Cloud is very helpful project",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3237155": "First of all, I would like to express my deepest gratitude to the competition organizers and all participants.\n\n\n# Our approach\n\n## Solution Summary\n\nIn short, our solution is based on refining velocity maps using clustering and VelToSeisLoss, rather than purely training DNN models.\n\n1. Utilize public csvs\n2. Clustering velocity maps and replace cluster's value with median.\n3. For FlatVel data, refine velocity map using VelToSeisLoss.\n\n## Background\n\nAs many of you know, there were so many strong public codes. For example, Bartley's CAFormer achieved 28.8 MAE on the public LB(https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved),  and Atom's ensemble even achieved an 25.6 MAE on the LB(https://www.kaggle.com/code/atom1231/ensemble-csv-for-more-files).\n\nIn short, we could not beat these public codes by training DNN models, a topic I will discuss this later.\n\nSo, we adopted a refinement approach.\n\n## Refinement Procedure\n\nStep1: First, we calculate or retrieve an initial velocity map, using models or public CSV files.\n\nStep2: Secondly, we apply a clustering algorithm to the velocity map, and replace each cluster's velocity value with its median of velocity value. \nFor Style A/B datasets, this approach worsen the solution, so we did not apply the replacement when there were so many clusters; this is a sign of Style datasets.\n\nThis median replacement improved the MAE from 25.6 to 25.4.\n\nAs [gguillard's discussion](https://www.kaggle.com/competitions/waveform-inversion/discussion/585815) pointed out, we can determine whether seismic data is in the FlatVel dataset by checking its symmetrisity among the different shots.\n\nUsing this knowledge, we can improve the clustering better. You can then decide to put all the grids at the same z value into the same cluster.\n\nAfter clustering, the number of the components of typical FlatVel dataset ranges from 3 to 9. (Without clustering, there were 4900 components!)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F8d39f14d44f01e72f74feaf1f5ddf226%2FCEasy_0.63_to_0.37_clustered_vel_map.png?generation=1751329382973959&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F8299d4f7256127735312c8b7754b1991%2FCEasy_7.70_to_7.03_clustered_vel_map.png?generation=1751329408325463&alt=media)\n\nStep3: Refinement using VelToSeisLoss(Physics Loss)\n\nThis powerful constraints enable us to use VelToSeisLoss.(without the constraints, the process can easily become trapped in a local minimum) For example, I reduced MAE loss for a FlatVel data from 9.4 to 5.1 in 40 iteration steps(it took about a few minutes). However, backpropagation of forward modeling requires so much computational resources, we could only do 2 steps for each test data. (It alsore requires so high VRAM capacity that I could not do refinement without dividing shots and doing gradient accumulation.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F07d4e674a0f74f3685732c3f6e937d24%2FScreenshot%20from%202025-07-01%2009-33-44.png?generation=1751330050355324&alt=media)\n\n# What didn't work\n\n- training UNet models with backbones(ConvNeXtV2Base/Large)(not strong enough to beat bartley's CAFormer)\n- finetuning Bartley's CAFormer: finetuning even worsened validation loss, it may be because I didn't care about cv folds\n\n\n# Lessons what we learned\n\n- bigger model is better: given a fixed GPU/TPU hours, a bigger model achieved a better validation loss in my experiments ( 224x224 ConvNextV2Base < 388x388 ConvNeXtV2Base < 388x388ConvNeXtV2Large)\n- Perceptual loss improved a speed of convergence at the beginning of training\n- Pretrained Weights can improve loss about 10 % even when domains are totally different\n- Google Cloud's storage costs too much for me\n- TPU v4s are really fast and TPU Research Cloud is very helpful project"
  }
}