{
  "id": 587402,
  "title": "20th solution - Caformer + data generation",
  "url": "/competitions/waveform-inversion/writeups/cpmp-20th-solution-caformer-data-generation",
  "author_name": "",
  "post_date": "2025-07-01T12:13:08.133Z",
  "votes": 46,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I want to thank Kaggle and the host for a great competition. I joined when <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> shared his great notebooks,, esp the caformer one. Given I know nothing in computer vision I stuck to his model. The only variation i made was to also train it with the 5 input planes side by side, i.e. with a 1 channel, 1000x350 image. </p>\n<p>I won't describe my solution in detail as it is not a very good one, but here are few things I did that helped.</p>\n<p><strong>Data augmentation.</strong></p>\n<p>Besides the flip, I used a velocity scaling. If you multiply the velocity by alpha &gt; 1, then you can shrink the seismic image by 1/alpha along the time dimension, and pad with 0. If you multiply velocity by alpha &lt; 1, seismic data is expanded by 1/alpha, and we truncate it to the original size. This is basic physics. Well, I hope it is. I used this both at training time, and at test time. At training time, I randomly apply the transformation with alpha in [0.8, 1.2]. At test time, I computed the output with 10 values of alpha, and take the median of the outputs.</p>\n<p><strong>Forward modeling.</strong></p>\n<p>I found the code used by host to generate seismic data from velocity (<a href=\"https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py)\" target=\"_blank\">https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py)</a>, a week before it was shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> in the forum. This code can be used by batch, and I could run it on a single test file in one call. I always thought it would be great to leverage it. I tried various ways, including back propagation through it. It was too slow. What worked best was to generate additional data form test predictions. For each prediction, optionally apply some data augmentation to it (scaling + crop, velocity scaling, ect), then generate seismic data from it, then use the seismic data as input and the predicted velocity as target. There are a handful papers describing variants of this. The variants I used are the following ones. I started by generating 2 or 3 copies of test data that way, then tried an in batch method in the last few days. For that I add a batch from test data every 6 training batch. For that test batch I predict vel from the input. Then I generate seismic from vel, then predict on seismic to get a vel2 prediction. Then I back propagate the MAE between  vel2 and vel. This is quite effective. At the end I also ran this with only test data. This looked promising.</p>\n<p><strong>Other backbones</strong></p>\n<p>I let time fly and started looking at other backbones only a couple of days before end. Training them did not converge enough to help.</p>\n<p><strong>Deep supervision</strong></p>\n<p>I use the record (10 classes) as target with a linear head. The input to the linear head is the middle feature in unet, i.e. the one of size 9x9 in Bartley's model.  This helped quite a bit.</p>\n<p><strong>What did not work</strong></p>\n<p>A lot.</p>\n<p>Really a lot. Main ones are these:</p>\n<ul>\n<li><p>To start with, my CV did not correlate anymore with LB when I reached a CV of 12 or so. I am not sure why, but it is why I submitted a lot. I used the same 20 files form the public notebook, maybe I should have used a larger and random set.</p></li>\n<li><p>I tried to use the dataset prediction to route the decoding through one of 10 decoder, a bit like a mixture of experts. CV improved quite a bit, but LB didn't. I should have revisited it in the end maybe, given my dataset prediction accuracy was 99.55% on the validation data.</p></li>\n<li><p>I generated lots of data using frequency and location moves as in <a href=\"https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py\" target=\"_blank\">https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py</a> but training with it took ages, and led to a LBimprovement of 1. I now see that I should have reused it at the end most probably, but I didn't for lack of time.</p></li>\n</ul>\n<p><strong>My Takeaway</strong></p>\n<p>Entering a computer vision competition solo was a bit silly given how little I know there. But this was a great learning experience, and I am sure I'll learn a lot from the amazing solutions from top teams. </p>",
  "messages": [
    {
      "id": "3237192",
      "postDate": "07/01/2025 00:51:19",
      "content": "<p>I want to thank Kaggle and the host for a great competition. I joined when <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> shared his great notebooks,, esp the caformer one. Given I know nothing in computer vision I stuck to his model. The only variation i made was to also train it with the 5 input planes side by side, i.e. with a 1 channel, 1000x350 image. </p>\n<p>I won't describe my solution in detail as it is not a very good one, but here are few things I did that helped.</p>\n<p><strong>Data augmentation.</strong></p>\n<p>Besides the flip, I used a velocity scaling. If you multiply the velocity by alpha &gt; 1, then you can shrink the seismic image by 1/alpha along the time dimension, and pad with 0. If you multiply velocity by alpha &lt; 1, seismic data is expanded by 1/alpha, and we truncate it to the original size. This is basic physics. Well, I hope it is. I used this both at training time, and at test time. At training time, I randomly apply the transformation with alpha in [0.8, 1.2]. At test time, I computed the output with 10 values of alpha, and take the median of the outputs.</p>\n<p><strong>Forward modeling.</strong></p>\n<p>I found the code used by host to generate seismic data from velocity (<a href=\"https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py)\" target=\"_blank\">https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py)</a>, a week before it was shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> in the forum. This code can be used by batch, and I could run it on a single test file in one call. I always thought it would be great to leverage it. I tried various ways, including back propagation through it. It was too slow. What worked best was to generate additional data form test predictions. For each prediction, optionally apply some data augmentation to it (scaling + crop, velocity scaling, ect), then generate seismic data from it, then use the seismic data as input and the predicted velocity as target. There are a handful papers describing variants of this. The variants I used are the following ones. I started by generating 2 or 3 copies of test data that way, then tried an in batch method in the last few days. For that I add a batch from test data every 6 training batch. For that test batch I predict vel from the input. Then I generate seismic from vel, then predict on seismic to get a vel2 prediction. Then I back propagate the MAE between  vel2 and vel. This is quite effective. At the end I also ran this with only test data. This looked promising.</p>\n<p><strong>Other backbones</strong></p>\n<p>I let time fly and started looking at other backbones only a couple of days before end. Training them did not converge enough to help.</p>\n<p><strong>Deep supervision</strong></p>\n<p>I use the record (10 classes) as target with a linear head. The input to the linear head is the middle feature in unet, i.e. the one of size 9x9 in Bartley's model.  This helped quite a bit.</p>\n<p><strong>What did not work</strong></p>\n<p>A lot.</p>\n<p>Really a lot. Main ones are these:</p>\n<ul>\n<li><p>To start with, my CV did not correlate anymore with LB when I reached a CV of 12 or so. I am not sure why, but it is why I submitted a lot. I used the same 20 files form the public notebook, maybe I should have used a larger and random set.</p></li>\n<li><p>I tried to use the dataset prediction to route the decoding through one of 10 decoder, a bit like a mixture of experts. CV improved quite a bit, but LB didn't. I should have revisited it in the end maybe, given my dataset prediction accuracy was 99.55% on the validation data.</p></li>\n<li><p>I generated lots of data using frequency and location moves as in <a href=\"https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py\" target=\"_blank\">https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py</a> but training with it took ages, and led to a LBimprovement of 1. I now see that I should have reused it at the end most probably, but I didn't for lack of time.</p></li>\n</ul>\n<p><strong>My Takeaway</strong></p>\n<p>Entering a computer vision competition solo was a bit silly given how little I know there. But this was a great learning experience, and I am sure I'll learn a lot from the amazing solutions from top teams. </p>",
      "rawMarkdown": "I want to thank Kaggle and the host for a great competition. I joined when @brendanartley shared his great notebooks,, esp the caformer one. Given I know nothing in computer vision I stuck to his model. The only variation i made was to also train it with the 5 input planes side by side, i.e. with a 1 channel, 1000x350 image. \n\nI won't describe my solution in detail as it is not a very good one, but here are few things I did that helped.\n\n**Data augmentation.**\n\nBesides the flip, I used a velocity scaling. If you multiply the velocity by alpha > 1, then you can shrink the seismic image by 1/alpha along the time dimension, and pad with 0. If you multiply velocity by alpha < 1, seismic data is expanded by 1/alpha, and we truncate it to the original size. This is basic physics. Well, I hope it is. I used this both at training time, and at test time. At training time, I randomly apply the transformation with alpha in [0.8, 1.2]. At test time, I computed the output with 10 values of alpha, and take the median of the outputs.\n\n**Forward modeling.**\n\nI found the code used by host to generate seismic data from velocity (https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py), a week before it was shared by @hengck23 in the forum. This code can be used by batch, and I could run it on a single test file in one call. I always thought it would be great to leverage it. I tried various ways, including back propagation through it. It was too slow. What worked best was to generate additional data form test predictions. For each prediction, optionally apply some data augmentation to it (scaling + crop, velocity scaling, ect), then generate seismic data from it, then use the seismic data as input and the predicted velocity as target. There are a handful papers describing variants of this. The variants I used are the following ones. I started by generating 2 or 3 copies of test data that way, then tried an in batch method in the last few days. For that I add a batch from test data every 6 training batch. For that test batch I predict vel from the input. Then I generate seismic from vel, then predict on seismic to get a vel2 prediction. Then I back propagate the MAE between  vel2 and vel. This is quite effective. At the end I also ran this with only test data. This looked promising.\n\n**Other backbones**\n\nI let time fly and started looking at other backbones only a couple of days before end. Training them did not converge enough to help.\n\n**Deep supervision**\n\nI use the record (10 classes) as target with a linear head. The input to the linear head is the middle feature in unet, i.e. the one of size 9x9 in Bartley's model.  This helped quite a bit.\n\n**What did not work**\n\nA lot.\n\nReally a lot. Main ones are these:\n\n- To start with, my CV did not correlate anymore with LB when I reached a CV of 12 or so. I am not sure why, but it is why I submitted a lot. I used the same 20 files form the public notebook, maybe I should have used a larger and random set.\n\n- I tried to use the dataset prediction to route the decoding through one of 10 decoder, a bit like a mixture of experts. CV improved quite a bit, but LB didn't. I should have revisited it in the end maybe, given my dataset prediction accuracy was 99.55% on the validation data.\n\n- I generated lots of data using frequency and location moves as in https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py but training with it took ages, and led to a LBimprovement of 1. I now see that I should have reused it at the end most probably, but I didn't for lack of time.\n\n**My Takeaway**\n\nEntering a computer vision competition solo was a bit silly given how little I know there. But this was a great learning experience, and I am sure I'll learn a lot from the amazing solutions from top teams.",
      "votes": null
    },
    {
      "id": "3237712",
      "postDate": "07/01/2025 08:55:35",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nHi, thank you for sharing your approach — it was very insightful!</p>\n<p>I have a question regarding your Deep Supervision strategy.</p>\n<p>Could you please explain why you placed the linear head on the middle feature of the U-Net?<br>\nIn my implementation, I attached the head at the end of the decoder. It worked quite well for classification tasks, but didn’t help much for regression (velocity map prediction).</p>\n<p>Also, I’d like to confirm the structure of your classification head.<br>\nIs it like this? → 9×9 → flatten → linear head (81 → 10)</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "cpmpml \nHi, thank you for sharing your approach — it was very insightful!\n\nI have a question regarding your Deep Supervision strategy.\n\nCould you please explain why you placed the linear head on the middle feature of the U-Net?\nIn my implementation, I attached the head at the end of the decoder. It worked quite well for classification tasks, but didn’t help much for regression (velocity map prediction).\n\nAlso, I’d like to confirm the structure of your classification head.\nIs it like this? → 9×9 → flatten → linear head (81 → 10)\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "3237843",
      "postDate": "07/01/2025 10:39:13",
      "content": "<p>I start from a b x 768 x 9 x 9 input, average over space to get a b x 768 tensor, then a linear layer:</p>\n<p>In the model init:</p>\n<pre><code>        \n        .dataset_head = torch.nn.Linear(, num_datasets)\n</code></pre>\n<p>In the forward method:</p>\n<pre><code>     ():\n        d = x[-].mean(-).mean(-)\n        d = .dataset_head(d)\n         d\n</code></pre>",
      "rawMarkdown": "I start from a b x 768 x 9 x 9 input, average over space to get a b x 768 tensor, then a linear layer:\n\nIn the model init:\n\n```python\n        # deep supervision\n        self.dataset_head = torch.nn.Linear(768, num_datasets)\n\n```\n\nIn the forward method:\n\n```python\n    def dataset(self, x):\n        d = x[-1].mean(-1).mean(-1)\n        d = self.dataset_head(d)\n        return d\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237712,
      "author_name": "hidebu",
      "author_url": "",
      "post_date": "07/01/2025 08:55:35",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nHi, thank you for sharing your approach — it was very insightful!</p>\n<p>I have a question regarding your Deep Supervision strategy.</p>\n<p>Could you please explain why you placed the linear head on the middle feature of the U-Net?<br>\nIn my implementation, I attached the head at the end of the decoder. It worked quite well for classification tasks, but didn’t help much for regression (velocity map prediction).</p>\n<p>Also, I’d like to confirm the structure of your classification head.<br>\nIs it like this? → 9×9 → flatten → linear head (81 → 10)</p>\n<p>Thanks in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237843,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/01/2025 10:39:13",
          "content": "<p>I start from a b x 768 x 9 x 9 input, average over space to get a b x 768 tensor, then a linear layer:</p>\n<p>In the model init:</p>\n<pre><code>        \n        .dataset_head = torch.nn.Linear(, num_datasets)\n</code></pre>\n<p>In the forward method:</p>\n<pre><code>     ():\n        d = x[-].mean(-).mean(-)\n        d = .dataset_head(d)\n         d\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237192": "I want to thank Kaggle and the host for a great competition. I joined when @brendanartley shared his great notebooks,, esp the caformer one. Given I know nothing in computer vision I stuck to his model. The only variation i made was to also train it with the 5 input planes side by side, i.e. with a 1 channel, 1000x350 image. \n\nI won't describe my solution in detail as it is not a very good one, but here are few things I did that helped.\n\n**Data augmentation.**\n\nBesides the flip, I used a velocity scaling. If you multiply the velocity by alpha > 1, then you can shrink the seismic image by 1/alpha along the time dimension, and pad with 0. If you multiply velocity by alpha < 1, seismic data is expanded by 1/alpha, and we truncate it to the original size. This is basic physics. Well, I hope it is. I used this both at training time, and at test time. At training time, I randomly apply the transformation with alpha in [0.8, 1.2]. At test time, I computed the output with 10 values of alpha, and take the median of the outputs.\n\n**Forward modeling.**\n\nI found the code used by host to generate seismic data from velocity (https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py), a week before it was shared by @hengck23 in the forum. This code can be used by batch, and I could run it on a single test file in one call. I always thought it would be great to leverage it. I tried various ways, including back propagation through it. It was too slow. What worked best was to generate additional data form test predictions. For each prediction, optionally apply some data augmentation to it (scaling + crop, velocity scaling, ect), then generate seismic data from it, then use the seismic data as input and the predicted velocity as target. There are a handful papers describing variants of this. The variants I used are the following ones. I started by generating 2 or 3 copies of test data that way, then tried an in batch method in the last few days. For that I add a batch from test data every 6 training batch. For that test batch I predict vel from the input. Then I generate seismic from vel, then predict on seismic to get a vel2 prediction. Then I back propagate the MAE between  vel2 and vel. This is quite effective. At the end I also ran this with only test data. This looked promising.\n\n**Other backbones**\n\nI let time fly and started looking at other backbones only a couple of days before end. Training them did not converge enough to help.\n\n**Deep supervision**\n\nI use the record (10 classes) as target with a linear head. The input to the linear head is the middle feature in unet, i.e. the one of size 9x9 in Bartley's model.  This helped quite a bit.\n\n**What did not work**\n\nA lot.\n\nReally a lot. Main ones are these:\n\n- To start with, my CV did not correlate anymore with LB when I reached a CV of 12 or so. I am not sure why, but it is why I submitted a lot. I used the same 20 files form the public notebook, maybe I should have used a larger and random set.\n\n- I tried to use the dataset prediction to route the decoding through one of 10 decoder, a bit like a mixture of experts. CV improved quite a bit, but LB didn't. I should have revisited it in the end maybe, given my dataset prediction accuracy was 99.55% on the validation data.\n\n- I generated lots of data using frequency and location moves as in https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_loc_f.py but training with it took ages, and led to a LBimprovement of 1. I now see that I should have reused it at the end most probably, but I didn't for lack of time.\n\n**My Takeaway**\n\nEntering a computer vision competition solo was a bit silly given how little I know there. But this was a great learning experience, and I am sure I'll learn a lot from the amazing solutions from top teams.",
    "3237712": "cpmpml \nHi, thank you for sharing your approach — it was very insightful!\n\nI have a question regarding your Deep Supervision strategy.\n\nCould you please explain why you placed the linear head on the middle feature of the U-Net?\nIn my implementation, I attached the head at the end of the decoder. It worked quite well for classification tasks, but didn’t help much for regression (velocity map prediction).\n\nAlso, I’d like to confirm the structure of your classification head.\nIs it like this? → 9×9 → flatten → linear head (81 → 10)\n\nThanks in advance!",
    "3237843": "I start from a b x 768 x 9 x 9 input, average over space to get a b x 768 tensor, then a linear layer:\n\nIn the model init:\n\n```python\n        # deep supervision\n        self.dataset_head = torch.nn.Linear(768, num_datasets)\n\n```\n\nIn the forward method:\n\n```python\n    def dataset(self, x):\n        d = x[-1].mean(-1).mean(-1)\n        d = self.dataset_head(d)\n        return d\n```"
  },
  "source": "meta"
}