{
  "id": 587417,
  "title": "Share some findings",
  "url": "/competitions/waveform-inversion/writeups/uesugi-erii-share-some-findings",
  "author_name": "",
  "post_date": "2025-07-01T02:24:25.670Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to the organizers for hosting such an interesting and stable competition, and thanks to all the open source developers for their initial contributions.<br>\nConsidering that my score was not very good, I will not be publishing a solution. However, I would like to share some simple and interesting findings I discovered during the competition, and I hope they will be helpful to you.</p>\n<h2>Data generation</h2>\n<p>The formula for data generation in the OpenFWI paper is incorrect; the correct formula should be:</p>\n<p>$$c_{i}(x, y) = c_{i-1}(x, a_{i}sin(2πk_{i}(x+Φ)) + y) \\tag{2}$$<br>\n$$c_{i}(x, y) = c_0(x+s_i,y+s_i') \\tag{3}$$</p>\n<p>And the range of parameter selection is not completely random. Specifically, the initial flat velocity map consists of 3 to 8 regions, where the probability of having 3 or 8 regions is half that of having 4 to 7 regions. The curved wave is composed of a mixture of sine waves with periods of 70/1, 70/2, … 70/5, with the phase randomly sampled from -π to π. I was not able to successfully reverse-engineer the amplitude. </p>\n<p>The paper states that the weight of the background incremental velocity map in the style data generation process is 0.7–0.9, but in reality, it should be 0.1–0.3.</p>\n<h2>Large-scale pre-training</h2>\n<p>With the data generation method, it means that unlimited data can be generated at low cost. I tried pre-training on data that is 100 times the official data, which can increase the cv by about 4. We can expect that more data can achieve better performance. Unfortunately, considering the cost, I did not continue to try. 100 times the data pre-training costs about $64</p>\n<h2><a href=\"https://github.com/liufeng2317/ADFWI\" target=\"_blank\">ADFWI</a></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362030%2Fb4e9e061ec1b51e5e95b83874947ff03%2FFigure1-AISWIT-Workflow.jpg?generation=1751335562729776&amp;alt=media\" alt=\"\"></p>\n<p>You can refer to their paper for detailed methods. It is a bit like TTT (test time training) in the ARC competition, where each sample is trained during the test. I think this is a general post-processing solution that can reduce your neural network prediction to a local optimum to achieve a certain score improvement. For me, LB was optimized from 22.1 to 17.0. </p>\n<h2>Waveform loss function</h2>\n<p><a href=\"https://github.com/liufeng2317/ADFWI/blob/bv1.1/ADFWI/fwi/misfit/Envelope.py\" target=\"_blank\">https://github.com/liufeng2317/ADFWI/blob/bv1.1/ADFWI/fwi/misfit/Envelope.py</a></p>\n<p>For the loss functions Envelope loss and correlation loss that measure the difference between waveforms, the theoretically correct approach is to put the time step in the last dimension, but I found that putting the sensor in the last dimension works better.</p>",
  "messages": [
    {
      "id": "3237260",
      "postDate": "07/01/2025 02:19:50",
      "content": "<p>Thanks to the organizers for hosting such an interesting and stable competition, and thanks to all the open source developers for their initial contributions.<br>\nConsidering that my score was not very good, I will not be publishing a solution. However, I would like to share some simple and interesting findings I discovered during the competition, and I hope they will be helpful to you.</p>\n<h2>Data generation</h2>\n<p>The formula for data generation in the OpenFWI paper is incorrect; the correct formula should be:</p>\n<p>$$c_{i}(x, y) = c_{i-1}(x, a_{i}sin(2πk_{i}(x+Φ)) + y) \\tag{2}$$<br>\n$$c_{i}(x, y) = c_0(x+s_i,y+s_i') \\tag{3}$$</p>\n<p>And the range of parameter selection is not completely random. Specifically, the initial flat velocity map consists of 3 to 8 regions, where the probability of having 3 or 8 regions is half that of having 4 to 7 regions. The curved wave is composed of a mixture of sine waves with periods of 70/1, 70/2, … 70/5, with the phase randomly sampled from -π to π. I was not able to successfully reverse-engineer the amplitude. </p>\n<p>The paper states that the weight of the background incremental velocity map in the style data generation process is 0.7–0.9, but in reality, it should be 0.1–0.3.</p>\n<h2>Large-scale pre-training</h2>\n<p>With the data generation method, it means that unlimited data can be generated at low cost. I tried pre-training on data that is 100 times the official data, which can increase the cv by about 4. We can expect that more data can achieve better performance. Unfortunately, considering the cost, I did not continue to try. 100 times the data pre-training costs about $64</p>\n<h2><a href=\"https://github.com/liufeng2317/ADFWI\" target=\"_blank\">ADFWI</a></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362030%2Fb4e9e061ec1b51e5e95b83874947ff03%2FFigure1-AISWIT-Workflow.jpg?generation=1751335562729776&amp;alt=media\" alt=\"\"></p>\n<p>You can refer to their paper for detailed methods. It is a bit like TTT (test time training) in the ARC competition, where each sample is trained during the test. I think this is a general post-processing solution that can reduce your neural network prediction to a local optimum to achieve a certain score improvement. For me, LB was optimized from 22.1 to 17.0. </p>\n<h2>Waveform loss function</h2>\n<p><a href=\"https://github.com/liufeng2317/ADFWI/blob/bv1.1/ADFWI/fwi/misfit/Envelope.py\" target=\"_blank\">https://github.com/liufeng2317/ADFWI/blob/bv1.1/ADFWI/fwi/misfit/Envelope.py</a></p>\n<p>For the loss functions Envelope loss and correlation loss that measure the difference between waveforms, the theoretically correct approach is to put the time step in the last dimension, but I found that putting the sensor in the last dimension works better.</p>",
      "rawMarkdown": "Thanks to the organizers for hosting such an interesting and stable competition, and thanks to all the open source developers for their initial contributions.\nConsidering that my score was not very good, I will not be publishing a solution. However, I would like to share some simple and interesting findings I discovered during the competition, and I hope they will be helpful to you.\n\n## Data generation\n\nThe formula for data generation in the OpenFWI paper is incorrect; the correct formula should be:\n\n$$c_{i}(x, y) = c_{i-1}(x, a_{i}sin(2πk_{i}(x+Φ)) + y) \\tag{2}$$\n$$c_{i}(x, y) = c_0(x+s_i,y+s_i') \\tag{3}$$\n\nAnd the range of parameter selection is not completely random. Specifically, the initial flat velocity map consists of 3 to 8 regions, where the probability of having 3 or 8 regions is half that of having 4 to 7 regions. The curved wave is composed of a mixture of sine waves with periods of 70/1, 70/2, ... 70/5, with the phase randomly sampled from -π to π. I was not able to successfully reverse-engineer the amplitude. \n\nThe paper states that the weight of the background incremental velocity map in the style data generation process is 0.7–0.9, but in reality, it should be 0.1–0.3.\n\n## Large-scale pre-training\n\nWith the data generation method, it means that unlimited data can be generated at low cost. I tried pre-training on data that is 100 times the official data, which can increase the cv by about 4. We can expect that more data can achieve better performance. Unfortunately, considering the cost, I did not continue to try. 100 times the data pre-training costs about $64\n\n## [ADFWI](https://github.com/liufeng2317/ADFWI)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362030%2Fb4e9e061ec1b51e5e95b83874947ff03%2FFigure1-AISWIT-Workflow.jpg?generation=1751335562729776&alt=media)\n\nYou can refer to their paper for detailed methods. It is a bit like TTT (test time training) in the ARC competition, where each sample is trained during the test. I think this is a general post-processing solution that can reduce your neural network prediction to a local optimum to achieve a certain score improvement. For me, LB was optimized from 22.1 to 17.0. \n\n## Waveform loss function\n\nhttps://github.com/liufeng2317/ADFWI/blob/bv1.1/ADFWI/fwi/misfit/Envelope.py\n\nFor the loss functions Envelope loss and correlation loss that measure the difference between waveforms, the theoretically correct approach is to put the time step in the last dimension, but I found that putting the sensor in the last dimension works better.",
      "votes": null
    },
    {
      "id": "3237564",
      "postDate": "07/01/2025 07:06:18",
      "content": "<p>Thank you for sharing this, definitely an interesting read!</p>",
      "rawMarkdown": "Thank you for sharing this, definitely an interesting read!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237564,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/01/2025 07:06:18",
      "content": "<p>Thank you for sharing this, definitely an interesting read!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3237260": "Thanks to the organizers for hosting such an interesting and stable competition, and thanks to all the open source developers for their initial contributions.\nConsidering that my score was not very good, I will not be publishing a solution. However, I would like to share some simple and interesting findings I discovered during the competition, and I hope they will be helpful to you.\n\n## Data generation\n\nThe formula for data generation in the OpenFWI paper is incorrect; the correct formula should be:\n\n$$c_{i}(x, y) = c_{i-1}(x, a_{i}sin(2πk_{i}(x+Φ)) + y) \\tag{2}$$\n$$c_{i}(x, y) = c_0(x+s_i,y+s_i') \\tag{3}$$\n\nAnd the range of parameter selection is not completely random. Specifically, the initial flat velocity map consists of 3 to 8 regions, where the probability of having 3 or 8 regions is half that of having 4 to 7 regions. The curved wave is composed of a mixture of sine waves with periods of 70/1, 70/2, ... 70/5, with the phase randomly sampled from -π to π. I was not able to successfully reverse-engineer the amplitude. \n\nThe paper states that the weight of the background incremental velocity map in the style data generation process is 0.7–0.9, but in reality, it should be 0.1–0.3.\n\n## Large-scale pre-training\n\nWith the data generation method, it means that unlimited data can be generated at low cost. I tried pre-training on data that is 100 times the official data, which can increase the cv by about 4. We can expect that more data can achieve better performance. Unfortunately, considering the cost, I did not continue to try. 100 times the data pre-training costs about $64\n\n## [ADFWI](https://github.com/liufeng2317/ADFWI)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4362030%2Fb4e9e061ec1b51e5e95b83874947ff03%2FFigure1-AISWIT-Workflow.jpg?generation=1751335562729776&alt=media)\n\nYou can refer to their paper for detailed methods. It is a bit like TTT (test time training) in the ARC competition, where each sample is trained during the test. I think this is a general post-processing solution that can reduce your neural network prediction to a local optimum to achieve a certain score improvement. For me, LB was optimized from 22.1 to 17.0. \n\n## Waveform loss function\n\nhttps://github.com/liufeng2317/ADFWI/blob/bv1.1/ADFWI/fwi/misfit/Envelope.py\n\nFor the loss functions Envelope loss and correlation loss that measure the difference between waveforms, the theoretically correct approach is to put the time step in the last dimension, but I found that putting the sensor in the last dimension works better.",
    "3237564": "Thank you for sharing this, definitely an interesting read!"
  },
  "source": "meta"
}