{
  "id": 587386,
  "title": "[Our findings] Did you read about how real seismic survey work when there is no ground truth of subsurface?",
  "url": "/competitions/waveform-inversion/discussion/587386",
  "author_name": "",
  "post_date": "2025-07-01T00:00:12.702083900Z",
  "votes": 11,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Most team may treat this competition as a regression task, which may not be very correct:</p>\n<pre><code>regression task:\n = model()  , here we want to predict \n\n = seismics\n = velocity\n</code></pre>\n<p>Actually, this is a solving task. We want to find the model parameters that can generate the obervation, it is a overfitting task!</p>\n<pre><code> solving task:\nx = physics() = physics(y), here \n\nyou \n</code></pre>\n<p>in real world real seismics survey,<br>\n<a href=\"https://www.youtube.com/watch?v=VWF--1OnitM\" target=\"_blank\">https://www.youtube.com/watch?v=VWF--1OnitM</a></p>\n<pre><code>observed seismics = physics(unkown parameter like velocity, ... , known parameters like source frequency, receiver locations, ...)\n\nsolution:\n) guess an initial parameter\n) simulate x\n) measure misfit: err(simulated x, observed x)\n)  misfit is small, parameter is found and . Else \n) update parameter and guess again\n</code></pre>\n<p>so to make submission csv, you train a neural net for each test sample, i.e. overfits the pretrained1.  net for that test sample</p>\n<pre><code>. net = pretain model  supervsied learning. This    inital guess.\n.  initial guess:  velocity0 = net(given seismics)\n. simulate: seismics = FWM(velocity0)\n. fmw loss = L1(simulated seismics , given seismics)\n.  loss  small, . Else \n. modify network by backpropgate\n.    guess\n\n you network  always changing  different samples \n</code></pre>\n<p>But there is a issue: fmw loss does not perfectly correlate  to kaggle mae loss  meaning.<br>\nwhen to stop or how to upate the  parameters, etc ….<br>\n(you can train a meta learning for that)</p>\n<pre><code>fmw =0 =&gt; mae =0\nbut fmw loss2&gt;fmw loss1 can mean  mae loss2 &gt; mae loss1 most of the time, but sometime otherwise\n</code></pre>\n<p>Also in real seismics, the physics model is a headache ( it cannot be smiply constant velocity FMW ). i.e. FMW is not given and you have to guess that as well</p>",
  "messages": [
    {
      "id": "3237122",
      "postDate": "07/01/2025 00:00:12",
      "content": "<p>Most team may treat this competition as a regression task, which may not be very correct:</p>\n<pre><code>regression task:\n = model()  , here we want to predict \n\n = seismics\n = velocity\n</code></pre>\n<p>Actually, this is a solving task. We want to find the model parameters that can generate the obervation, it is a overfitting task!</p>\n<pre><code> solving task:\nx = physics() = physics(y), here \n\nyou \n</code></pre>\n<p>in real world real seismics survey,<br>\n<a href=\"https://www.youtube.com/watch?v=VWF--1OnitM\" target=\"_blank\">https://www.youtube.com/watch?v=VWF--1OnitM</a></p>\n<pre><code>observed seismics = physics(unkown parameter like velocity, ... , known parameters like source frequency, receiver locations, ...)\n\nsolution:\n) guess an initial parameter\n) simulate x\n) measure misfit: err(simulated x, observed x)\n)  misfit is small, parameter is found and . Else \n) update parameter and guess again\n</code></pre>\n<p>so to make submission csv, you train a neural net for each test sample, i.e. overfits the pretrained1.  net for that test sample</p>\n<pre><code>. net = pretain model  supervsied learning. This    inital guess.\n.  initial guess:  velocity0 = net(given seismics)\n. simulate: seismics = FWM(velocity0)\n. fmw loss = L1(simulated seismics , given seismics)\n.  loss  small, . Else \n. modify network by backpropgate\n.    guess\n\n you network  always changing  different samples \n</code></pre>\n<p>But there is a issue: fmw loss does not perfectly correlate  to kaggle mae loss  meaning.<br>\nwhen to stop or how to upate the  parameters, etc ….<br>\n(you can train a meta learning for that)</p>\n<pre><code>fmw =0 =&gt; mae =0\nbut fmw loss2&gt;fmw loss1 can mean  mae loss2 &gt; mae loss1 most of the time, but sometime otherwise\n</code></pre>\n<p>Also in real seismics, the physics model is a headache ( it cannot be smiply constant velocity FMW ). i.e. FMW is not given and you have to guess that as well</p>",
      "rawMarkdown": "Most team may treat this competition as a regression task, which may not be very correct:\n```\nregression task:\ny = model(x)  , here we want to predict y\n\nx = seismics\ny = velocity\n\n```\n\nActually, this is a solving task. We want to find the model parameters that can generate the obervation, it is a overfitting task!\n\n```\nmodel solving task:\nx = physics(parameter) = physics(y), here we want to determine y\n\nyou can see this is an inversion task\n```\n\nin real world real seismics survey,\nhttps://www.youtube.com/watch?v=VWF--1OnitM\n\n```\nobserved seismics = physics(unkown parameter like velocity, ... , known parameters like source frequency, receiver locations, ...)\n\nsolution:\n1) guess an initial parameter\n2) simulate x\n3) measure misfit: err(simulated x, observed x)\n4) if misfit is small, parameter is found and exit. Else continue\n5) update parameter and guess again\n\n```\n\nso to make submission csv, you train a neural net for each test sample, i.e. overfits the pretrained1.  net for that test sample\n\n```\n1. net = pretain model on supervsied learning. This is to make inital guess.\n2. make initial guess:  velocity0 = net(given seismics)\n3. simulate: seismics = FWM(velocity0)\n4. fmw loss = L1(simulated seismics , given seismics)\n5. if loss is small, stop. Else continue\n6. modify network by backpropgate\n4. make a new guess\n\nso you network is always changing for different samples \n```\n\nBut there is a issue: fmw loss does not perfectly correlate  to kaggle mae loss  meaning.\nwhen to stop or how to upate the  parameters, etc ....\n(you can train a meta learning for that)\n\n```\nfmw loss=0 => mae loss=0\nbut fmw loss2>fmw loss1 can mean  mae loss2 > mae loss1 most of the time, but sometime otherwise\n\n```\n\nAlso in real seismics, the physics model is a headache ( it cannot be smiply constant velocity FMW ). i.e. FMW is not given and you have to guess that as well",
      "votes": null
    },
    {
      "id": "3237128",
      "postDate": "07/01/2025 00:05:20",
      "content": "<p>How much bump this fitting for each sample gave you compared to initial guess?<br>\nHow much compute it require?<br>\nI just trained a really big network with lot of data for very long time 🤣</p>",
      "rawMarkdown": "How much bump this fitting for each sample gave you compared to initial guess?\nHow much compute it require?\nI just trained a really big network with lot of data for very long time 🤣",
      "votes": null
    },
    {
      "id": "3237131",
      "postDate": "07/01/2025 00:07:17",
      "content": "<p>Also, your solution require simulating the seis from found vel in each step.<br>\nIt's lot of compute even after gpu optimization.<br>\nUnless you did really amazing optimization.</p>",
      "rawMarkdown": "Also, your solution require simulating the seis from found vel in each step.\nIt's lot of compute even after gpu optimization.\nUnless you did really amazing optimization.",
      "votes": null
    },
    {
      "id": "3237140",
      "postDate": "07/01/2025 00:12:02",
      "content": "<p>How much bump this fitting for each sample gave you compared to initial guess?<br>\nfrom 19.5 all the way to 12.3. My teammate will details the solution later.</p>\n<p>we cannot find reliable and fast FMW gradient compuatation method, so we did it without backprop from FMW</p>",
      "rawMarkdown": "How much bump this fitting for each sample gave you compared to initial guess?\nfrom 19.5 all the way to 12.3. My teammate will details the solution later.\n\nwe cannot find reliable and fast FMW gradient compuatation method, so we did it without backprop from FMW",
      "votes": null
    },
    {
      "id": "3237174",
      "postDate": "07/01/2025 00:36:24",
      "content": "<p>how it looks like when fitting network<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2444879ee6bacd9a20d759830ba26b40%2Fiterate.gif?generation=1751330170444490&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "how it looks like when fitting network\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2444879ee6bacd9a20d759830ba26b40%2Fiterate.gif?generation=1751330170444490&alt=media)",
      "votes": null
    },
    {
      "id": "3237176",
      "postDate": "07/01/2025 00:38:31",
      "content": "<p>I'm not refering to the backbrop step but the seis calculation. (you predict vel- now for overfitting you simulate the seis)<br>\nI guess it is manageable but still very demanding, no?<br>\nActually, this fitting on prediction is really the same as <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> ! Different implementation maybe but same idea at the core.  <br>\nThis is really the single most important thing probably to take from this competition if there is ever a similar one. </p>",
      "rawMarkdown": "I'm not refering to the backbrop step but the seis calculation. (you predict vel- now for overfitting you simulate the seis)\nI guess it is manageable but still very demanding, no?\nActually, this fitting on prediction is really the same as @harshitsheoran ! Different implementation maybe but same idea at the core.  \nThis is really the single most important thing probably to take from this competition if there is ever a similar one.",
      "votes": null
    },
    {
      "id": "3237178",
      "postDate": "07/01/2025 00:40:37",
      "content": "<p>Maybe you are looking for<br>\n<a href=\"https://github.com/liufeng2317/ADFWI\" target=\"_blank\">https://github.com/liufeng2317/ADFWI</a></p>",
      "rawMarkdown": "Maybe you are looking for\nhttps://github.com/liufeng2317/ADFWI",
      "votes": null
    },
    {
      "id": "3237182",
      "postDate": "07/01/2025 00:44:15",
      "content": "<p>our team leader <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> write very fast FWM (without gradient computation) code in cuda (faster then the one released in public code). it can do batch-size = 64 forward modelling in less than a second. fast enough for online augmentation and network fitting.</p>\n<p>Kudos to <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>!</p>",
      "rawMarkdown": "our team leader @lihaoweicvch write very fast FWM (without gradient computation) code in cuda (faster then the one released in public code). it can do batch-size = 64 forward modelling in less than a second. fast enough for online augmentation and network fitting.\n\nKudos to @lihaoweicvch!",
      "votes": null
    },
    {
      "id": "3237189",
      "postDate": "07/01/2025 00:48:37",
      "content": "<p>My code can also do batch size 64 in less than a second (&gt;5000 samples per minute), on a 5090… otherwise doing the calculation after every epoch would be too expensive</p>",
      "rawMarkdown": "My code can also do batch size 64 in less than a second (>5000 samples per minute), on a 5090... otherwise doing the calculation after every epoch would be too expensive",
      "votes": null
    },
    {
      "id": "3237205",
      "postDate": "07/01/2025 00:58:12",
      "content": "<p>Yes it's really fast.<br>\nYou intend to publish code?<br>\n<a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>  you did not publish the generation code, right? or I missed it.</p>",
      "rawMarkdown": "Yes it's really fast.\nYou intend to publish code?\n@harshitsheoran  you did not publish the generation code, right? or I missed it.",
      "votes": null
    },
    {
      "id": "3237209",
      "postDate": "07/01/2025 01:00:16",
      "content": "<p>I sued something quite similar.</p>",
      "rawMarkdown": "I sued something quite similar.",
      "votes": null
    },
    {
      "id": "3237219",
      "postDate": "07/01/2025 01:08:47",
      "content": "<p>I do plan to publish it after I wake up, will go to bed soon, been awake for a day now haha</p>",
      "rawMarkdown": "I do plan to publish it after I wake up, will go to bed soon, been awake for a day now haha",
      "votes": null
    },
    {
      "id": "3237231",
      "postDate": "07/01/2025 01:27:18",
      "content": "<p>I have tried the idea, but I can not find a method that can generate seis data fast enough. So I have to generate seis offline and regard it as pseudo labels. Then, finetune the model with these pseudo labels.</p>",
      "rawMarkdown": "I have tried the idea, but I can not find a method that can generate seis data fast enough. So I have to generate seis offline and regard it as pseudo labels. Then, finetune the model with these pseudo labels.",
      "votes": null
    },
    {
      "id": "3237319",
      "postDate": "07/01/2025 03:11:55",
      "content": "<p>I know you found this approch  in last last two days. </p>\n<p>Style A&amp;B benefit more than others in the family.</p>\n<p>For other families, BP with FWM are easy to find local optimum, but it seems that there is no way to give the gradient to those larger numbers in Velmap.</p>",
      "rawMarkdown": "I know you found this approch  in last last two days. \n\nStyle A&B benefit more than others in the family.\n\nFor other families, BP with FWM are easy to find local optimum, but it seems that there is no way to give the gradient to those larger numbers in Velmap.",
      "votes": null
    },
    {
      "id": "3237323",
      "postDate": "07/01/2025 03:18:09",
      "content": "<p>Since most Velmap in family except Style A/B is a discrete integer, if the model prediction's MAE is low enough(not mine), I think we can find the global optimum(MAE = 0) by some search strategy with FWM</p>",
      "rawMarkdown": "Since most Velmap in family except Style A/B is a discrete integer, if the model prediction's MAE is low enough(not mine), I think we can find the global optimum(MAE = 0) by some search strategy with FWM",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237128,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "07/01/2025 00:05:20",
      "content": "<p>How much bump this fitting for each sample gave you compared to initial guess?<br>\nHow much compute it require?<br>\nI just trained a really big network with lot of data for very long time 🤣</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237131,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "07/01/2025 00:07:17",
          "content": "<p>Also, your solution require simulating the seis from found vel in each step.<br>\nIt's lot of compute even after gpu optimization.<br>\nUnless you did really amazing optimization.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3237140,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/01/2025 00:12:02",
          "content": "<p>How much bump this fitting for each sample gave you compared to initial guess?<br>\nfrom 19.5 all the way to 12.3. My teammate will details the solution later.</p>\n<p>we cannot find reliable and fast FMW gradient compuatation method, so we did it without backprop from FMW</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237176,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "07/01/2025 00:38:31",
              "content": "<p>I'm not refering to the backbrop step but the seis calculation. (you predict vel- now for overfitting you simulate the seis)<br>\nI guess it is manageable but still very demanding, no?<br>\nActually, this fitting on prediction is really the same as <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> ! Different implementation maybe but same idea at the core.  <br>\nThis is really the single most important thing probably to take from this competition if there is ever a similar one. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3237182,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "07/01/2025 00:44:15",
                  "content": "<p>our team leader <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> write very fast FWM (without gradient computation) code in cuda (faster then the one released in public code). it can do batch-size = 64 forward modelling in less than a second. fast enough for online augmentation and network fitting.</p>\n<p>Kudos to <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3237189,
                      "author_name": "harshitsheoran",
                      "author_url": "",
                      "post_date": "07/01/2025 00:48:37",
                      "content": "<p>My code can also do batch size 64 in less than a second (&gt;5000 samples per minute), on a 5090… otherwise doing the calculation after every epoch would be too expensive</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3237205,
                          "author_name": "shlomoron",
                          "author_url": "",
                          "post_date": "07/01/2025 00:58:12",
                          "content": "<p>Yes it's really fast.<br>\nYou intend to publish code?<br>\n<a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>  you did not publish the generation code, right? or I missed it.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3237219,
                              "author_name": "harshitsheoran",
                              "author_url": "",
                              "post_date": "07/01/2025 01:08:47",
                              "content": "<p>I do plan to publish it after I wake up, will go to bed soon, been awake for a day now haha</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3237174,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/01/2025 00:36:24",
      "content": "<p>how it looks like when fitting network<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2444879ee6bacd9a20d759830ba26b40%2Fiterate.gif?generation=1751330170444490&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237178,
      "author_name": "uesugierii",
      "author_url": "",
      "post_date": "07/01/2025 00:40:37",
      "content": "<p>Maybe you are looking for<br>\n<a href=\"https://github.com/liufeng2317/ADFWI\" target=\"_blank\">https://github.com/liufeng2317/ADFWI</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237209,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/01/2025 01:00:16",
      "content": "<p>I sued something quite similar.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237231,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "07/01/2025 01:27:18",
      "content": "<p>I have tried the idea, but I can not find a method that can generate seis data fast enough. So I have to generate seis offline and regard it as pseudo labels. Then, finetune the model with these pseudo labels.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237319,
      "author_name": "sharksbeer",
      "author_url": "",
      "post_date": "07/01/2025 03:11:55",
      "content": "<p>I know you found this approch  in last last two days. </p>\n<p>Style A&amp;B benefit more than others in the family.</p>\n<p>For other families, BP with FWM are easy to find local optimum, but it seems that there is no way to give the gradient to those larger numbers in Velmap.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237323,
          "author_name": "sharksbeer",
          "author_url": "",
          "post_date": "07/01/2025 03:18:09",
          "content": "<p>Since most Velmap in family except Style A/B is a discrete integer, if the model prediction's MAE is low enough(not mine), I think we can find the global optimum(MAE = 0) by some search strategy with FWM</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237122": "Most team may treat this competition as a regression task, which may not be very correct:\n```\nregression task:\ny = model(x)  , here we want to predict y\n\nx = seismics\ny = velocity\n\n```\n\nActually, this is a solving task. We want to find the model parameters that can generate the obervation, it is a overfitting task!\n\n```\nmodel solving task:\nx = physics(parameter) = physics(y), here we want to determine y\n\nyou can see this is an inversion task\n```\n\nin real world real seismics survey,\nhttps://www.youtube.com/watch?v=VWF--1OnitM\n\n```\nobserved seismics = physics(unkown parameter like velocity, ... , known parameters like source frequency, receiver locations, ...)\n\nsolution:\n1) guess an initial parameter\n2) simulate x\n3) measure misfit: err(simulated x, observed x)\n4) if misfit is small, parameter is found and exit. Else continue\n5) update parameter and guess again\n\n```\n\nso to make submission csv, you train a neural net for each test sample, i.e. overfits the pretrained1.  net for that test sample\n\n```\n1. net = pretain model on supervsied learning. This is to make inital guess.\n2. make initial guess:  velocity0 = net(given seismics)\n3. simulate: seismics = FWM(velocity0)\n4. fmw loss = L1(simulated seismics , given seismics)\n5. if loss is small, stop. Else continue\n6. modify network by backpropgate\n4. make a new guess\n\nso you network is always changing for different samples \n```\n\nBut there is a issue: fmw loss does not perfectly correlate  to kaggle mae loss  meaning.\nwhen to stop or how to upate the  parameters, etc ....\n(you can train a meta learning for that)\n\n```\nfmw loss=0 => mae loss=0\nbut fmw loss2>fmw loss1 can mean  mae loss2 > mae loss1 most of the time, but sometime otherwise\n\n```\n\nAlso in real seismics, the physics model is a headache ( it cannot be smiply constant velocity FMW ). i.e. FMW is not given and you have to guess that as well",
    "3237128": "How much bump this fitting for each sample gave you compared to initial guess?\nHow much compute it require?\nI just trained a really big network with lot of data for very long time 🤣",
    "3237131": "Also, your solution require simulating the seis from found vel in each step.\nIt's lot of compute even after gpu optimization.\nUnless you did really amazing optimization.",
    "3237140": "How much bump this fitting for each sample gave you compared to initial guess?\nfrom 19.5 all the way to 12.3. My teammate will details the solution later.\n\nwe cannot find reliable and fast FMW gradient compuatation method, so we did it without backprop from FMW",
    "3237174": "how it looks like when fitting network\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2444879ee6bacd9a20d759830ba26b40%2Fiterate.gif?generation=1751330170444490&alt=media)",
    "3237176": "I'm not refering to the backbrop step but the seis calculation. (you predict vel- now for overfitting you simulate the seis)\nI guess it is manageable but still very demanding, no?\nActually, this fitting on prediction is really the same as @harshitsheoran ! Different implementation maybe but same idea at the core.  \nThis is really the single most important thing probably to take from this competition if there is ever a similar one.",
    "3237178": "Maybe you are looking for\nhttps://github.com/liufeng2317/ADFWI",
    "3237182": "our team leader @lihaoweicvch write very fast FWM (without gradient computation) code in cuda (faster then the one released in public code). it can do batch-size = 64 forward modelling in less than a second. fast enough for online augmentation and network fitting.\n\nKudos to @lihaoweicvch!",
    "3237189": "My code can also do batch size 64 in less than a second (>5000 samples per minute), on a 5090... otherwise doing the calculation after every epoch would be too expensive",
    "3237205": "Yes it's really fast.\nYou intend to publish code?\n@harshitsheoran  you did not publish the generation code, right? or I missed it.",
    "3237209": "I sued something quite similar.",
    "3237219": "I do plan to publish it after I wake up, will go to bed soon, been awake for a day now haha",
    "3237231": "I have tried the idea, but I can not find a method that can generate seis data fast enough. So I have to generate seis offline and regard it as pseudo labels. Then, finetune the model with these pseudo labels.",
    "3237319": "I know you found this approch  in last last two days. \n\nStyle A&B benefit more than others in the family.\n\nFor other families, BP with FWM are easy to find local optimum, but it seems that there is no way to give the gradient to those larger numbers in Velmap.",
    "3237323": "Since most Velmap in family except Style A/B is a discrete integer, if the model prediction's MAE is low enough(not mine), I think we can find the global optimum(MAE = 0) by some search strategy with FWM"
  },
  "source": "meta"
}