{
  "id": 587429,
  "title": "10th Place Solution (yu4u's part)",
  "url": "/competitions/waveform-inversion/discussion/587429",
  "author_name": "",
  "post_date": "2025-07-01T03:17:51.545146200Z",
  "votes": 25,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I discovered this competition after the BYU competition had ended. I found it extremely interesting and regretted having spent so much time on the BYU competition instead.</p>\n<p>We would like to express our gratitude to the competition host and the Kaggle staff for organizing this outstanding competition. Below, we introduce the yu4u's part of the solution of Team monnu and yu4u. Please refer to awesome monnu’s solution <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/587412\" target=\"_blank\">here</a>.</p>\n<h1>Model</h1>\n<p>The input tensor with shape 5 × 1000 × 70 is resized to 224 × 224 and fed into the backbone; a forward pass produces a 7 × 7 output, which is then expanded to 70 × 70 using pixel shuffle. With this approach, no dedicated decoder is required.<br>\nDuring the forward pass in the backbone, the five channels are processed individually, and at the outputs of stages 1 and 2, the channels are aggregated from 5 to 3, and 3 to 1 using average pooling respectively.<br>\nThe implementation of this 2.5D model is based on <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561401\" target=\"_blank\">the version I employed in the CZII competition</a>.</p>\n<h1>Model Training</h1>\n<p>The initial model training was conducted on the entire OpenFWI dataset with the following settings: AdamW optimizer, 64 epochs, batch size = 8, learning rate = 2e-4, weight decay = 5e-2, and an exponential moving average (EMA) decay of 0.999. Data augmentation was limited to horizontal flipping only.</p>\n<h1>Finetuning</h1>\n<p>After the initial training phase, the inference outputs for the test and validation data were converted back to the input format (seismic data) via seismic forward modeling, and those pairs were used to finetune the pretrained model.<br>\nDuring finetuning, the learning rate was reduced to 1e-5 and learning-rate warm-up was also applied.</p>\n<h1>Iterative optimization</h1>\n<p>Using <a href=\"https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py\" target=\"_blank\">a back-propagation-capable forward-modeling implementation</a>, each test sample was processed individually: the pretrained model’s predictions were transformed back into seismic data, and the model was iteratively optimized to minimize the reconstruction error.<br>\nAlthough this procedure does not directly optimize the velocity map, we empirically found that the mean absolute error (MAE) improved as the number of iterations increased.</p>\n<p>A simplified implementation example is shown below:</p>\n<pre><code> = MyModel()\n x  test_xs:\n    .load_state_dict(torch.load())  \n    optimizer = AdamW(...)\n    .train()\n     i  range():\n        optimizer.zero_grad()\n        pred_y = (x)\n        reconst_x = FWM(pred_y)\n        loss = criterion(ori_x, reconst_x)\n        loss.backward()\n        optimizer.()\n\n</code></pre>\n<h1>Results</h1>\n<table>\n<thead>\n<tr>\n<th>Stage</th>\n<th>CV</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Initial Training</td>\n<td>19.50</td>\n<td>21.6</td>\n</tr>\n<tr>\n<td>Finetuning</td>\n<td>16.96</td>\n<td>17.5</td>\n</tr>\n<tr>\n<td>Iterative Optimization</td>\n<td>N/A</td>\n<td>13.7</td>\n</tr>\n<tr>\n<td>Finetuning</td>\n<td>13.36</td>\n<td>12.8</td>\n</tr>\n<tr>\n<td>Finetuning</td>\n<td>13.16</td>\n<td>12.5</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "3237322",
      "postDate": "07/01/2025 03:17:51",
      "content": "<p>I discovered this competition after the BYU competition had ended. I found it extremely interesting and regretted having spent so much time on the BYU competition instead.</p>\n<p>We would like to express our gratitude to the competition host and the Kaggle staff for organizing this outstanding competition. Below, we introduce the yu4u's part of the solution of Team monnu and yu4u. Please refer to awesome monnu’s solution <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/587412\" target=\"_blank\">here</a>.</p>\n<h1>Model</h1>\n<p>The input tensor with shape 5 × 1000 × 70 is resized to 224 × 224 and fed into the backbone; a forward pass produces a 7 × 7 output, which is then expanded to 70 × 70 using pixel shuffle. With this approach, no dedicated decoder is required.<br>\nDuring the forward pass in the backbone, the five channels are processed individually, and at the outputs of stages 1 and 2, the channels are aggregated from 5 to 3, and 3 to 1 using average pooling respectively.<br>\nThe implementation of this 2.5D model is based on <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561401\" target=\"_blank\">the version I employed in the CZII competition</a>.</p>\n<h1>Model Training</h1>\n<p>The initial model training was conducted on the entire OpenFWI dataset with the following settings: AdamW optimizer, 64 epochs, batch size = 8, learning rate = 2e-4, weight decay = 5e-2, and an exponential moving average (EMA) decay of 0.999. Data augmentation was limited to horizontal flipping only.</p>\n<h1>Finetuning</h1>\n<p>After the initial training phase, the inference outputs for the test and validation data were converted back to the input format (seismic data) via seismic forward modeling, and those pairs were used to finetune the pretrained model.<br>\nDuring finetuning, the learning rate was reduced to 1e-5 and learning-rate warm-up was also applied.</p>\n<h1>Iterative optimization</h1>\n<p>Using <a href=\"https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py\" target=\"_blank\">a back-propagation-capable forward-modeling implementation</a>, each test sample was processed individually: the pretrained model’s predictions were transformed back into seismic data, and the model was iteratively optimized to minimize the reconstruction error.<br>\nAlthough this procedure does not directly optimize the velocity map, we empirically found that the mean absolute error (MAE) improved as the number of iterations increased.</p>\n<p>A simplified implementation example is shown below:</p>\n<pre><code> = MyModel()\n x  test_xs:\n    .load_state_dict(torch.load())  \n    optimizer = AdamW(...)\n    .train()\n     i  range():\n        optimizer.zero_grad()\n        pred_y = (x)\n        reconst_x = FWM(pred_y)\n        loss = criterion(ori_x, reconst_x)\n        loss.backward()\n        optimizer.()\n\n</code></pre>\n<h1>Results</h1>\n<table>\n<thead>\n<tr>\n<th>Stage</th>\n<th>CV</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Initial Training</td>\n<td>19.50</td>\n<td>21.6</td>\n</tr>\n<tr>\n<td>Finetuning</td>\n<td>16.96</td>\n<td>17.5</td>\n</tr>\n<tr>\n<td>Iterative Optimization</td>\n<td>N/A</td>\n<td>13.7</td>\n</tr>\n<tr>\n<td>Finetuning</td>\n<td>13.36</td>\n<td>12.8</td>\n</tr>\n<tr>\n<td>Finetuning</td>\n<td>13.16</td>\n<td>12.5</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I discovered this competition after the BYU competition had ended. I found it extremely interesting and regretted having spent so much time on the BYU competition instead.\n\nWe would like to express our gratitude to the competition host and the Kaggle staff for organizing this outstanding competition. Below, we introduce the yu4u's part of the solution of Team monnu and yu4u. Please refer to awesome monnu’s solution [here](https://www.kaggle.com/competitions/waveform-inversion/discussion/587412).\n\n\n# Model\nThe input tensor with shape 5 × 1000 × 70 is resized to 224 × 224 and fed into the backbone; a forward pass produces a 7 × 7 output, which is then expanded to 70 × 70 using pixel shuffle. With this approach, no dedicated decoder is required.\nDuring the forward pass in the backbone, the five channels are processed individually, and at the outputs of stages 1 and 2, the channels are aggregated from 5 to 3, and 3 to 1 using average pooling respectively.\nThe implementation of this 2.5D model is based on [the version I employed in the CZII competition](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561401).\n\n\n# Model Training\nThe initial model training was conducted on the entire OpenFWI dataset with the following settings: AdamW optimizer, 64 epochs, batch size = 8, learning rate = 2e-4, weight decay = 5e-2, and an exponential moving average (EMA) decay of 0.999. Data augmentation was limited to horizontal flipping only.\n\n\n# Finetuning\nAfter the initial training phase, the inference outputs for the test and validation data were converted back to the input format (seismic data) via seismic forward modeling, and those pairs were used to finetune the pretrained model.\nDuring finetuning, the learning rate was reduced to 1e-5 and learning-rate warm-up was also applied.\n\n\n\n# Iterative optimization\nUsing [a back-propagation-capable forward-modeling implementation](https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py), each test sample was processed individually: the pretrained model’s predictions were transformed back into seismic data, and the model was iteratively optimized to minimize the reconstruction error.\nAlthough this procedure does not directly optimize the velocity map, we empirically found that the mean absolute error (MAE) improved as the number of iterations increased.\n\nA simplified implementation example is shown below:\n\n```\nmodel = MyModel()\nfor x in test_xs:\n    model.load_state_dict(torch.load(\"model.pth\"))  # reload for each test data\n    optimizer = AdamW(...)\n    model.train()\n    for i in range(50):\n        optimizer.zero_grad()\n        pred_y = model(x)\n        reconst_x = FWM(pred_y)\n        loss = criterion(ori_x, reconst_x)\n        loss.backward()\n        optimizer.step()\n# use final pred_y for submission or further finetuning\n```\n\n# Results\n\n| Stage                  | CV    | Private LB |\n| ---------------------- | ----- | ---------- |\n| Initial Training       | 19.50 | 21.6       |\n| Finetuning             | 16.96 | 17.5       |\n| Iterative Optimization | N/A   | 13.7       |\n| Finetuning             | 13.36 | 12.8       |\n| Finetuning             | 13.16 | 12.5       |",
      "votes": null
    },
    {
      "id": "3237349",
      "postDate": "07/01/2025 03:42:48",
      "content": "<p>Thanks for the writeup.<br>\n\"we empirically found that the mean absolute error (MAE) improved as the number of iterations increased.\"</p>\n<p>in my experiment, MAE will go down then it will go up if the iterations is run long enough. <br>\n\"loss = criterion(ori_x, reconst_x)\" always decreases. <br>\ndo you observe similarly?</p>",
      "rawMarkdown": "Thanks for the writeup.\n\"we empirically found that the mean absolute error (MAE) improved as the number of iterations increased.\"\n\nin my experiment, MAE will go down then it will go up if the iterations is run long enough. \n\"loss = criterion(ori_x, reconst_x)\" always decreases. \ndo you observe similarly?",
      "votes": null
    },
    {
      "id": "3237371",
      "postDate": "07/01/2025 03:55:47",
      "content": "<p>In my experiments, the optimization converged within around 100 steps, but I did not observe any degradation in performance even when the number of iterations was increased beyond that. I confirmed that the loss consistently decreased. However, slight oscillations were observed during the optimization process.</p>\n<p>My teammate <a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">@fuumin621</a> conducted an excellent experiment using the sampled validation data.<br>\nThis figure shows the gain in MAE as a function of the number of iterations performed and the percentage of data optimized, assuming optimization is applied in descending order of reconstruction error.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fe4a93277d74573f40280e27ecbfcabf5%2Fimage.png?generation=1751342036295658&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "In my experiments, the optimization converged within around 100 steps, but I did not observe any degradation in performance even when the number of iterations was increased beyond that. I confirmed that the loss consistently decreased. However, slight oscillations were observed during the optimization process.\n\nMy teammate @fuumin621 conducted an excellent experiment using the sampled validation data.\nThis figure shows the gain in MAE as a function of the number of iterations performed and the percentage of data optimized, assuming optimization is applied in descending order of reconstruction error.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fe4a93277d74573f40280e27ecbfcabf5%2Fimage.png?generation=1751342036295658&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237349,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/01/2025 03:42:48",
      "content": "<p>Thanks for the writeup.<br>\n\"we empirically found that the mean absolute error (MAE) improved as the number of iterations increased.\"</p>\n<p>in my experiment, MAE will go down then it will go up if the iterations is run long enough. <br>\n\"loss = criterion(ori_x, reconst_x)\" always decreases. <br>\ndo you observe similarly?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237371,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "07/01/2025 03:55:47",
          "content": "<p>In my experiments, the optimization converged within around 100 steps, but I did not observe any degradation in performance even when the number of iterations was increased beyond that. I confirmed that the loss consistently decreased. However, slight oscillations were observed during the optimization process.</p>\n<p>My teammate <a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">@fuumin621</a> conducted an excellent experiment using the sampled validation data.<br>\nThis figure shows the gain in MAE as a function of the number of iterations performed and the percentage of data optimized, assuming optimization is applied in descending order of reconstruction error.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fe4a93277d74573f40280e27ecbfcabf5%2Fimage.png?generation=1751342036295658&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237322": "I discovered this competition after the BYU competition had ended. I found it extremely interesting and regretted having spent so much time on the BYU competition instead.\n\nWe would like to express our gratitude to the competition host and the Kaggle staff for organizing this outstanding competition. Below, we introduce the yu4u's part of the solution of Team monnu and yu4u. Please refer to awesome monnu’s solution [here](https://www.kaggle.com/competitions/waveform-inversion/discussion/587412).\n\n\n# Model\nThe input tensor with shape 5 × 1000 × 70 is resized to 224 × 224 and fed into the backbone; a forward pass produces a 7 × 7 output, which is then expanded to 70 × 70 using pixel shuffle. With this approach, no dedicated decoder is required.\nDuring the forward pass in the backbone, the five channels are processed individually, and at the outputs of stages 1 and 2, the channels are aggregated from 5 to 3, and 3 to 1 using average pooling respectively.\nThe implementation of this 2.5D model is based on [the version I employed in the CZII competition](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561401).\n\n\n# Model Training\nThe initial model training was conducted on the entire OpenFWI dataset with the following settings: AdamW optimizer, 64 epochs, batch size = 8, learning rate = 2e-4, weight decay = 5e-2, and an exponential moving average (EMA) decay of 0.999. Data augmentation was limited to horizontal flipping only.\n\n\n# Finetuning\nAfter the initial training phase, the inference outputs for the test and validation data were converted back to the input format (seismic data) via seismic forward modeling, and those pairs were used to finetune the pretrained model.\nDuring finetuning, the learning rate was reduced to 1e-5 and learning-rate warm-up was also applied.\n\n\n\n# Iterative optimization\nUsing [a back-propagation-capable forward-modeling implementation](https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py), each test sample was processed individually: the pretrained model’s predictions were transformed back into seismic data, and the model was iteratively optimized to minimize the reconstruction error.\nAlthough this procedure does not directly optimize the velocity map, we empirically found that the mean absolute error (MAE) improved as the number of iterations increased.\n\nA simplified implementation example is shown below:\n\n```\nmodel = MyModel()\nfor x in test_xs:\n    model.load_state_dict(torch.load(\"model.pth\"))  # reload for each test data\n    optimizer = AdamW(...)\n    model.train()\n    for i in range(50):\n        optimizer.zero_grad()\n        pred_y = model(x)\n        reconst_x = FWM(pred_y)\n        loss = criterion(ori_x, reconst_x)\n        loss.backward()\n        optimizer.step()\n# use final pred_y for submission or further finetuning\n```\n\n# Results\n\n| Stage                  | CV    | Private LB |\n| ---------------------- | ----- | ---------- |\n| Initial Training       | 19.50 | 21.6       |\n| Finetuning             | 16.96 | 17.5       |\n| Iterative Optimization | N/A   | 13.7       |\n| Finetuning             | 13.36 | 12.8       |\n| Finetuning             | 13.16 | 12.5       |",
    "3237349": "Thanks for the writeup.\n\"we empirically found that the mean absolute error (MAE) improved as the number of iterations increased.\"\n\nin my experiment, MAE will go down then it will go up if the iterations is run long enough. \n\"loss = criterion(ori_x, reconst_x)\" always decreases. \ndo you observe similarly?",
    "3237371": "In my experiments, the optimization converged within around 100 steps, but I did not observe any degradation in performance even when the number of iterations was increased beyond that. I confirmed that the loss consistently decreased. However, slight oscillations were observed during the optimization process.\n\nMy teammate @fuumin621 conducted an excellent experiment using the sampled validation data.\nThis figure shows the gain in MAE as a function of the number of iterations performed and the percentage of data optimized, assuming optimization is applied in descending order of reconstruction error.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fe4a93277d74573f40280e27ecbfcabf5%2Fimage.png?generation=1751342036295658&alt=media)"
  },
  "source": "meta"
}