{
  "id": 587419,
  "title": "3rd place solution",
  "url": "/competitions/waveform-inversion/writeups/tascj-3rd-place-solution",
  "author_name": "",
  "post_date": "2025-07-02T10:49:33.640Z",
  "votes": 60,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thank you kaggle and the host team for hosting this competition. Congratulations to the winners!</p>\n<p>My solution leverages synthetic data generation and Vision Transformer for modeling. Requiring lots of storage and compute resources.<br>\nI used more than 13TB of storage and ~5.5 days of <code>A100 SXM 80G x4</code>.</p>\n<h2>Forward Modeling</h2>\n<p>I started with <a href=\"https://www.kaggle.com/code/manatoyo/improved-vel-to-seis\" target=\"_blank\">this implementation</a> by <a href=\"https://www.kaggle.com/manatoyo\" target=\"_blank\">@manatoyo</a> (developed with the help of Google's Gemini). And made the following changes:</p>\n<ol>\n<li>Change dtype from float64 to float32. This reduced errors a bit. Matlab uses float64 by default, and Gemini knows that. However, OpenFWI seems to be created using a float32 implementation.</li>\n<li>Generate 6 channels <code>isx=[120, 137, 154, 155, 172, 189]</code> to enable fast hflip augmentation.</li>\n<li>Port to PyTorch and enable batch execution on GPU.</li>\n</ol>\n<p>For some vel types, there are large errors in the last row of seismic data between OpenFWI data and generated data, so I zeroed the last row of all data during training.</p>\n<p>OpenFWI is not fully open. Their code may have some minor bugs, and if there are serious bugs in their forward modeling code, there may be little we can do about it. Fortunately, the issues don't seem significant at the moment.</p>\n<h2>Generating data</h2>\n<p>I generated the following data for training:</p>\n<ol>\n<li>6ch version of Full OpenFWI data.</li>\n<li>random blending (alpha in [0.3, 0.7]) of two vels from same fold, same vel-type. repeated 8 times with different seeds.</li>\n<li>random blending (alpha in [0.3, 0.7]) of two vels from same fold, not restricted to same vel-type. repeated 8 times with different seeds.</li>\n</ol>\n<p>In my short training runs, generating more data consistently leads to performance improvements.</p>\n<h2>Modeling</h2>\n<p>I started with ViT with RoPE (<code>eva02_base_patch16_clip_224</code>) for this task.</p>\n<ol>\n<li>I first re-aranged the input from (N, 5, 1000, 70) to (N, 5, 250, 280), then interpolated it to (N, 5, patch_size * 18, patch_size * 18).</li>\n<li>The input is then fed to the ViT model, predicted 16-dim outputs, then pixel-shuffled to (N, 1, 72, 72) and cropped to (N, 1, 70, 70).</li>\n<li>The logits are transformed to target by <code>logits.sigmoid() * 6000</code>.</li>\n</ol>\n<p>Then I split ViT layers:</p>\n<ol>\n<li>First 1/3 layers for intra-channel feature interaction.</li>\n<li>Average pool 5 channels to 1 channel.</li>\n<li>Last 2/3 layers for further feature interaction.</li>\n</ol>\n<h2>Optimizer</h2>\n<p>I started with a short <code>CurveFault_A</code> only training run. Switching from AdamW to <code>MuonWithAuxAdam</code> was a big improvement (9.x -&gt; 7.x). I adapted the implementation from <code>KellerJordan/Muon</code> and followed the practices from <code>MoonshotAI/Moonlight</code>. Muon was used for weights of all <code>Linear</code> layers except patch embedding, for other parameters AdamW was used.</p>\n<p>I haven't achieved significant performance improvements with optimizers other than SGD and Adam for many years, so this was a surprise for me, perhaps the biggest takeaway from this competition. The only drawback is that Muon's overhead becomes non-negligible when using small batch sizes.</p>\n<h2>Submission</h2>\n<p>The final model was <code>eva02_large_patch14_clip_224</code> trained in 3 stages.</p>\n<h3>Stage 1</h3>\n<ul>\n<li>HFlip augmentation.</li>\n<li>Input resized to (14 * 18, 14 * 18), 4x output upsampling.</li>\n<li>10 epochs on original data and generated data</li>\n<li>50% constant learning rate + 50% cosine decay learning rate</li>\n</ul>\n<p>MAE on 20% validation data: 8.1</p>\n<h3>Stage 2</h3>\n<ul>\n<li>Resume from stage 1</li>\n<li>HFlip augmentation.</li>\n<li>Input resized to (14 * 35, 14 * 35), 2x output upsampling.</li>\n<li>1 epochs on original data and generated data</li>\n<li>50% constant learning rate + 50% cosine decay learning rate</li>\n</ul>\n<p>MAE on 20% validation data: 7.6</p>\n<h3>Stage 3</h3>\n<ul>\n<li>Resume from stage 2</li>\n<li>HFlip augmentation.</li>\n<li>Input resized to (14 * 35, 14 * 35), 2x output upsampling.</li>\n<li>1 epochs on validation data</li>\n<li>cosine decay learning rate</li>\n</ul>\n<p>More synthetic data and more epochs could further improve the performance, this was the best I could finish with my GPU hours.</p>\n<p>Submission was a 20%+30%+50% blending of models from 3 stages. With <code>x.clip(1500, 4500).round()</code>.</p>\n<p><strong>Update</strong><br>\nI made a mistake in the inference script by not zeroing the last row, which is inconsistent with training. After fixing the bug, the score is (Public 7.96, Private 7.99).<br>\nFortunately, it doesn't affect the ranking.</p>\n<h2>Code</h2>\n<p><a href=\"https://www.kaggle.com/datasets/tascj0/fwi-submission\" target=\"_blank\">Submitted file</a><br>\n<a href=\"https://www.kaggle.com/code/tascj0/fwi-inference\" target=\"_blank\">Inference notebook</a><br>\n<a href=\"https://github.com/tascj/kaggle-waveform-inversion\" target=\"_blank\">Training code</a></p>",
  "messages": [
    {
      "id": "3237264",
      "postDate": "07/01/2025 02:24:27",
      "content": "<p>Thank you kaggle and the host team for hosting this competition. Congratulations to the winners!</p>\n<p>My solution leverages synthetic data generation and Vision Transformer for modeling. Requiring lots of storage and compute resources.<br>\nI used more than 13TB of storage and ~5.5 days of <code>A100 SXM 80G x4</code>.</p>\n<h2>Forward Modeling</h2>\n<p>I started with <a href=\"https://www.kaggle.com/code/manatoyo/improved-vel-to-seis\" target=\"_blank\">this implementation</a> by <a href=\"https://www.kaggle.com/manatoyo\" target=\"_blank\">@manatoyo</a> (developed with the help of Google's Gemini). And made the following changes:</p>\n<ol>\n<li>Change dtype from float64 to float32. This reduced errors a bit. Matlab uses float64 by default, and Gemini knows that. However, OpenFWI seems to be created using a float32 implementation.</li>\n<li>Generate 6 channels <code>isx=[120, 137, 154, 155, 172, 189]</code> to enable fast hflip augmentation.</li>\n<li>Port to PyTorch and enable batch execution on GPU.</li>\n</ol>\n<p>For some vel types, there are large errors in the last row of seismic data between OpenFWI data and generated data, so I zeroed the last row of all data during training.</p>\n<p>OpenFWI is not fully open. Their code may have some minor bugs, and if there are serious bugs in their forward modeling code, there may be little we can do about it. Fortunately, the issues don't seem significant at the moment.</p>\n<h2>Generating data</h2>\n<p>I generated the following data for training:</p>\n<ol>\n<li>6ch version of Full OpenFWI data.</li>\n<li>random blending (alpha in [0.3, 0.7]) of two vels from same fold, same vel-type. repeated 8 times with different seeds.</li>\n<li>random blending (alpha in [0.3, 0.7]) of two vels from same fold, not restricted to same vel-type. repeated 8 times with different seeds.</li>\n</ol>\n<p>In my short training runs, generating more data consistently leads to performance improvements.</p>\n<h2>Modeling</h2>\n<p>I started with ViT with RoPE (<code>eva02_base_patch16_clip_224</code>) for this task.</p>\n<ol>\n<li>I first re-aranged the input from (N, 5, 1000, 70) to (N, 5, 250, 280), then interpolated it to (N, 5, patch_size * 18, patch_size * 18).</li>\n<li>The input is then fed to the ViT model, predicted 16-dim outputs, then pixel-shuffled to (N, 1, 72, 72) and cropped to (N, 1, 70, 70).</li>\n<li>The logits are transformed to target by <code>logits.sigmoid() * 6000</code>.</li>\n</ol>\n<p>Then I split ViT layers:</p>\n<ol>\n<li>First 1/3 layers for intra-channel feature interaction.</li>\n<li>Average pool 5 channels to 1 channel.</li>\n<li>Last 2/3 layers for further feature interaction.</li>\n</ol>\n<h2>Optimizer</h2>\n<p>I started with a short <code>CurveFault_A</code> only training run. Switching from AdamW to <code>MuonWithAuxAdam</code> was a big improvement (9.x -&gt; 7.x). I adapted the implementation from <code>KellerJordan/Muon</code> and followed the practices from <code>MoonshotAI/Moonlight</code>. Muon was used for weights of all <code>Linear</code> layers except patch embedding, for other parameters AdamW was used.</p>\n<p>I haven't achieved significant performance improvements with optimizers other than SGD and Adam for many years, so this was a surprise for me, perhaps the biggest takeaway from this competition. The only drawback is that Muon's overhead becomes non-negligible when using small batch sizes.</p>\n<h2>Submission</h2>\n<p>The final model was <code>eva02_large_patch14_clip_224</code> trained in 3 stages.</p>\n<h3>Stage 1</h3>\n<ul>\n<li>HFlip augmentation.</li>\n<li>Input resized to (14 * 18, 14 * 18), 4x output upsampling.</li>\n<li>10 epochs on original data and generated data</li>\n<li>50% constant learning rate + 50% cosine decay learning rate</li>\n</ul>\n<p>MAE on 20% validation data: 8.1</p>\n<h3>Stage 2</h3>\n<ul>\n<li>Resume from stage 1</li>\n<li>HFlip augmentation.</li>\n<li>Input resized to (14 * 35, 14 * 35), 2x output upsampling.</li>\n<li>1 epochs on original data and generated data</li>\n<li>50% constant learning rate + 50% cosine decay learning rate</li>\n</ul>\n<p>MAE on 20% validation data: 7.6</p>\n<h3>Stage 3</h3>\n<ul>\n<li>Resume from stage 2</li>\n<li>HFlip augmentation.</li>\n<li>Input resized to (14 * 35, 14 * 35), 2x output upsampling.</li>\n<li>1 epochs on validation data</li>\n<li>cosine decay learning rate</li>\n</ul>\n<p>More synthetic data and more epochs could further improve the performance, this was the best I could finish with my GPU hours.</p>\n<p>Submission was a 20%+30%+50% blending of models from 3 stages. With <code>x.clip(1500, 4500).round()</code>.</p>\n<p><strong>Update</strong><br>\nI made a mistake in the inference script by not zeroing the last row, which is inconsistent with training. After fixing the bug, the score is (Public 7.96, Private 7.99).<br>\nFortunately, it doesn't affect the ranking.</p>\n<h2>Code</h2>\n<p><a href=\"https://www.kaggle.com/datasets/tascj0/fwi-submission\" target=\"_blank\">Submitted file</a><br>\n<a href=\"https://www.kaggle.com/code/tascj0/fwi-inference\" target=\"_blank\">Inference notebook</a><br>\n<a href=\"https://github.com/tascj/kaggle-waveform-inversion\" target=\"_blank\">Training code</a></p>",
      "rawMarkdown": "Thank you kaggle and the host team for hosting this competition. Congratulations to the winners!\n\n\nMy solution leverages synthetic data generation and Vision Transformer for modeling. Requiring lots of storage and compute resources.\nI used more than 13TB of storage and ~5.5 days of `A100 SXM 80G x4`.\n\n\n## Forward Modeling\n\nI started with [this implementation](https://www.kaggle.com/code/manatoyo/improved-vel-to-seis) by @manatoyo (developed with the help of Google's Gemini). And made the following changes:\n\n1. Change dtype from float64 to float32. This reduced errors a bit. Matlab uses float64 by default, and Gemini knows that. However, OpenFWI seems to be created using a float32 implementation.\n2. Generate 6 channels `isx=[120, 137, 154, 155, 172, 189]` to enable fast hflip augmentation.\n3. Port to PyTorch and enable batch execution on GPU.\n\nFor some vel types, there are large errors in the last row of seismic data between OpenFWI data and generated data, so I zeroed the last row of all data during training.\n\nOpenFWI is not fully open. Their code may have some minor bugs, and if there are serious bugs in their forward modeling code, there may be little we can do about it. Fortunately, the issues don't seem significant at the moment.\n\n## Generating data\n\nI generated the following data for training:\n1. 6ch version of Full OpenFWI data.\n2. random blending (alpha in [0.3, 0.7]) of two vels from same fold, same vel-type. repeated 8 times with different seeds.\n3. random blending (alpha in [0.3, 0.7]) of two vels from same fold, not restricted to same vel-type. repeated 8 times with different seeds.\n\nIn my short training runs, generating more data consistently leads to performance improvements.\n\n## Modeling\n\nI started with ViT with RoPE (`eva02_base_patch16_clip_224`) for this task.\n\n1. I first re-aranged the input from (N, 5, 1000, 70) to (N, 5, 250, 280), then interpolated it to (N, 5, patch_size * 18, patch_size * 18).\n2. The input is then fed to the ViT model, predicted 16-dim outputs, then pixel-shuffled to (N, 1, 72, 72) and cropped to (N, 1, 70, 70).\n3. The logits are transformed to target by `logits.sigmoid() * 6000`.\n\nThen I split ViT layers:\n1. First 1/3 layers for intra-channel feature interaction.\n2. Average pool 5 channels to 1 channel.\n3. Last 2/3 layers for further feature interaction.\n\n\n## Optimizer\n\nI started with a short `CurveFault_A` only training run. Switching from AdamW to `MuonWithAuxAdam` was a big improvement (9.x -> 7.x). I adapted the implementation from `KellerJordan/Muon` and followed the practices from `MoonshotAI/Moonlight`. Muon was used for weights of all `Linear` layers except patch embedding, for other parameters AdamW was used.\n\nI haven't achieved significant performance improvements with optimizers other than SGD and Adam for many years, so this was a surprise for me, perhaps the biggest takeaway from this competition. The only drawback is that Muon's overhead becomes non-negligible when using small batch sizes.\n\n\n## Submission\n\nThe final model was `eva02_large_patch14_clip_224` trained in 3 stages.\n\n### Stage 1\n\n* HFlip augmentation.\n* Input resized to (14 * 18, 14 * 18), 4x output upsampling.\n* 10 epochs on original data and generated data\n* 50% constant learning rate + 50% cosine decay learning rate\n\nMAE on 20% validation data: 8.1\n\n### Stage 2\n\n* Resume from stage 1\n* HFlip augmentation.\n* Input resized to (14 * 35, 14 * 35), 2x output upsampling.\n* 1 epochs on original data and generated data\n* 50% constant learning rate + 50% cosine decay learning rate\n\nMAE on 20% validation data: 7.6\n\n### Stage 3\n\n* Resume from stage 2\n* HFlip augmentation.\n* Input resized to (14 * 35, 14 * 35), 2x output upsampling.\n* 1 epochs on validation data\n* cosine decay learning rate\n\nMore synthetic data and more epochs could further improve the performance, this was the best I could finish with my GPU hours.\n\nSubmission was a 20%+30%+50% blending of models from 3 stages. With `x.clip(1500, 4500).round()`.\n\n**Update**\nI made a mistake in the inference script by not zeroing the last row, which is inconsistent with training. After fixing the bug, the score is (Public 7.96, Private 7.99).\nFortunately, it doesn't affect the ranking.\n\n## Code\n\n[Submitted file](https://www.kaggle.com/datasets/tascj0/fwi-submission)\n[Inference notebook](https://www.kaggle.com/code/tascj0/fwi-inference)\n[Training code](https://github.com/tascj/kaggle-waveform-inversion)",
      "votes": null
    },
    {
      "id": "3237299",
      "postDate": "07/01/2025 02:51:38",
      "content": "<p>Congratulations, looking forward to your <code>MuonWithAuxAdam</code>.</p>",
      "rawMarkdown": "Congratulations, looking forward to your `MuonWithAuxAdam`.",
      "votes": null
    },
    {
      "id": "3237307",
      "postDate": "07/01/2025 03:02:29",
      "content": "<p>I have tried MuonWithAuxAdam in early experiments, but it will output NaN in training loop and I have to give it up. Looking forward to code!</p>",
      "rawMarkdown": "I have tried MuonWithAuxAdam in early experiments, but it will output NaN in training loop and I have to give it up. Looking forward to code!",
      "votes": null
    },
    {
      "id": "3237312",
      "postDate": "07/01/2025 03:04:30",
      "content": "<p>Congratulations!</p>\n<p>I asked same question to another winner: What is your technique to prevent overfit in this one competition? </p>\n<p>Thank for reply.</p>",
      "rawMarkdown": "Congratulations!\n\nI asked same question to another winner: What is your technique to prevent overfit in this one competition? \n\nThank for reply.",
      "votes": null
    },
    {
      "id": "3237388",
      "postDate": "07/01/2025 04:26:47",
      "content": "<p>In this competition, you can generate a large amount of data. For me, train with more data is the key to prevent overfitting.</p>\n<p>However, don’t underestimate the impact of underfitting — you need a model with sufficient capacity and enough training to fully take advantage of data.</p>",
      "rawMarkdown": "In this competition, you can generate a large amount of data. For me, train with more data is the key to prevent overfitting.\n\nHowever, don’t underestimate the impact of underfitting — you need a model with sufficient capacity and enough training to fully take advantage of data.",
      "votes": null
    },
    {
      "id": "3237420",
      "postDate": "07/01/2025 04:59:13",
      "content": "<p>Yes, i was struggling with both overfit and underfit.<br>\nI am looking forward to your code solution (hopefully in  Kaggle notebook 😀)  </p>",
      "rawMarkdown": "Yes, i was struggling with both overfit and underfit.\nI am looking forward to your code solution (hopefully in  Kaggle notebook 😀)",
      "votes": null
    },
    {
      "id": "3237472",
      "postDate": "07/01/2025 05:51:51",
      "content": "<p>Congratulations on your third place.</p>\n<p>If I understand correctly, your method can efficiently infer new test data, right? In other words, your computationally intensive training step does not use the test data?</p>\n<p>This would make your solution quite different from <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>'s and mine. It would mean that yours is the highest solution that would do well if this were a code competition…</p>",
      "rawMarkdown": "Congratulations on your third place.\n\nIf I understand correctly, your method can efficiently infer new test data, right? In other words, your computationally intensive training step does not use the test data?\n\nThis would make your solution quite different from @harshitsheoran's and mine. It would mean that yours is the highest solution that would do well if this were a code competition...",
      "votes": null
    },
    {
      "id": "3237505",
      "postDate": "07/01/2025 06:17:44",
      "content": "<p>Yes. I did not use the test data for training.</p>\n<p>I started this competition downloading the full OpenFWI dataset and initially did not download the Kaggle data. I didn’t realize this is a CSV competition until I started preparing for submission (2 days before the deadline). By that time, my final training job was already scheduled, and I decided not to change it.</p>\n<p>By the way, I built pipeline to optimize predicted vel by forward modeling + backpropagation soon after I started working on this competition, but dropped that idea because it would be too slow for a code competition.</p>",
      "rawMarkdown": "Yes. I did not use the test data for training.\n\nI started this competition downloading the full OpenFWI dataset and initially did not download the Kaggle data. I didn’t realize this is a CSV competition until I started preparing for submission (2 days before the deadline). By that time, my final training job was already scheduled, and I decided not to change it.\n\nBy the way, I built pipeline to optimize predicted vel by forward modeling + backpropagation soon after I started working on this competition, but dropped that idea because it would be too slow for a code competition.",
      "votes": null
    },
    {
      "id": "3237514",
      "postDate": "07/01/2025 06:23:13",
      "content": "<p>Nice work, great improvement with the optimiser. I like the simplicity of the EVA head <code>then pixel-shuffled to (N, 1, 72, 72)</code>.</p>",
      "rawMarkdown": "Nice work, great improvement with the optimiser. I like the simplicity of the EVA head `then pixel-shuffled to (N, 1, 72, 72)`.",
      "votes": null
    },
    {
      "id": "3237519",
      "postDate": "07/01/2025 06:29:41",
      "content": "<p>Congratz. I followed the same path i.e. just training a large vit variation on lot of generated data without using the test data. You said 'By that time, my final training job was already scheduled, and I decided not to change it.'. So training your model is only 2 days? Or was only final fine-tuning? If so, how much time it take to train your full model and what is the required compute?</p>",
      "rawMarkdown": "Congratz. I followed the same path i.e. just training a large vit variation on lot of generated data without using the test data. You said 'By that time, my final training job was already scheduled, and I decided not to change it.'. So training your model is only 2 days? Or was only final fine-tuning? If so, how much time it take to train your full model and what is the required compute?",
      "votes": null
    },
    {
      "id": "3237577",
      "postDate": "07/01/2025 07:15:03",
      "content": "<p>I started preparing submission after stage 2 started.</p>\n<p>I used <code>A100 SXM 80G x4</code>.</p>\n<ol>\n<li>Data generation: ~25 hours</li>\n<li>Stage 1: ~74 hours</li>\n<li>Stage 2: ~26 hours</li>\n<li>Stage 3: ~6.5 hours</li>\n</ol>\n<p>So the full pipeline took ~5.5 days.</p>",
      "rawMarkdown": "I started preparing submission after stage 2 started.\n\nI used `A100 SXM 80G x4`.\n\n1. Data generation: ~25 hours\n2. Stage 1: ~74 hours\n3. Stage 2: ~26 hours\n4. Stage 3: ~6.5 hours\n\nSo the full pipeline took ~5.5 days.",
      "votes": null
    },
    {
      "id": "3238682",
      "postDate": "07/02/2025 05:29:45",
      "content": "<p>Inference code&amp;checkpoints, training code are available now.</p>",
      "rawMarkdown": "Inference code&checkpoints, training code are available now.",
      "votes": null
    },
    {
      "id": "3238685",
      "postDate": "07/02/2025 05:34:06",
      "content": "<p>Could you also share the notebook that runs the full test set and writes submission.csv (if you did this on Kaggle that is)?</p>",
      "rawMarkdown": "Could you also share the notebook that runs the full test set and writes submission.csv (if you did this on Kaggle that is)?",
      "votes": null
    },
    {
      "id": "3238730",
      "postDate": "07/02/2025 06:36:46",
      "content": "<p>I made predictions locally. Command shared on GitHub.<br>\nUploaded the submitted file <a href=\"https://www.kaggle.com/datasets/tascj0/fwi-submission\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I made predictions locally. Command shared on GitHub.\nUploaded the submitted file [here](https://www.kaggle.com/datasets/tascj0/fwi-submission).",
      "votes": null
    },
    {
      "id": "3239166",
      "postDate": "07/02/2025 14:54:50",
      "content": "<p>Thanks. What I was mainly curious about: how long does inferring the full test set take? (This being 1000s of GPU hours for most solutions, but not yours I think).</p>",
      "rawMarkdown": "Thanks. What I was mainly curious about: how long does inferring the full test set take? (This being 1000s of GPU hours for most solutions, but not yours I think).",
      "votes": null
    },
    {
      "id": "3240532",
      "postDate": "07/03/2025 23:14:06",
      "content": "<p>Inferring the full test set using stage2 model took 30+mins using single A100. Less than 90 mins to make the submission.</p>",
      "rawMarkdown": "Inferring the full test set using stage2 model took 30+mins using single A100. Less than 90 mins to make the submission.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237299,
      "author_name": "zhudong1949",
      "author_url": "",
      "post_date": "07/01/2025 02:51:38",
      "content": "<p>Congratulations, looking forward to your <code>MuonWithAuxAdam</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237307,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "07/01/2025 03:02:29",
      "content": "<p>I have tried MuonWithAuxAdam in early experiments, but it will output NaN in training loop and I have to give it up. Looking forward to code!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237312,
      "author_name": "overvalueawareness",
      "author_url": "",
      "post_date": "07/01/2025 03:04:30",
      "content": "<p>Congratulations!</p>\n<p>I asked same question to another winner: What is your technique to prevent overfit in this one competition? </p>\n<p>Thank for reply.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237388,
          "author_name": "tascj0",
          "author_url": "",
          "post_date": "07/01/2025 04:26:47",
          "content": "<p>In this competition, you can generate a large amount of data. For me, train with more data is the key to prevent overfitting.</p>\n<p>However, don’t underestimate the impact of underfitting — you need a model with sufficient capacity and enough training to fully take advantage of data.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237420,
              "author_name": "overvalueawareness",
              "author_url": "",
              "post_date": "07/01/2025 04:59:13",
              "content": "<p>Yes, i was struggling with both overfit and underfit.<br>\nI am looking forward to your code solution (hopefully in  Kaggle notebook 😀)  </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3238682,
                  "author_name": "tascj0",
                  "author_url": "",
                  "post_date": "07/02/2025 05:29:45",
                  "content": "<p>Inference code&amp;checkpoints, training code are available now.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3238685,
                      "author_name": "jeroencottaar",
                      "author_url": "",
                      "post_date": "07/02/2025 05:34:06",
                      "content": "<p>Could you also share the notebook that runs the full test set and writes submission.csv (if you did this on Kaggle that is)?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3238730,
                          "author_name": "tascj0",
                          "author_url": "",
                          "post_date": "07/02/2025 06:36:46",
                          "content": "<p>I made predictions locally. Command shared on GitHub.<br>\nUploaded the submitted file <a href=\"https://www.kaggle.com/datasets/tascj0/fwi-submission\" target=\"_blank\">here</a>.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3239166,
                              "author_name": "jeroencottaar",
                              "author_url": "",
                              "post_date": "07/02/2025 14:54:50",
                              "content": "<p>Thanks. What I was mainly curious about: how long does inferring the full test set take? (This being 1000s of GPU hours for most solutions, but not yours I think).</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3240532,
                                  "author_name": "tascj0",
                                  "author_url": "",
                                  "post_date": "07/03/2025 23:14:06",
                                  "content": "<p>Inferring the full test set using stage2 model took 30+mins using single A100. Less than 90 mins to make the submission.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3237472,
      "author_name": "jeroencottaar",
      "author_url": "",
      "post_date": "07/01/2025 05:51:51",
      "content": "<p>Congratulations on your third place.</p>\n<p>If I understand correctly, your method can efficiently infer new test data, right? In other words, your computationally intensive training step does not use the test data?</p>\n<p>This would make your solution quite different from <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>'s and mine. It would mean that yours is the highest solution that would do well if this were a code competition…</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237505,
          "author_name": "tascj0",
          "author_url": "",
          "post_date": "07/01/2025 06:17:44",
          "content": "<p>Yes. I did not use the test data for training.</p>\n<p>I started this competition downloading the full OpenFWI dataset and initially did not download the Kaggle data. I didn’t realize this is a CSV competition until I started preparing for submission (2 days before the deadline). By that time, my final training job was already scheduled, and I decided not to change it.</p>\n<p>By the way, I built pipeline to optimize predicted vel by forward modeling + backpropagation soon after I started working on this competition, but dropped that idea because it would be too slow for a code competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237514,
      "author_name": "darraghdog",
      "author_url": "",
      "post_date": "07/01/2025 06:23:13",
      "content": "<p>Nice work, great improvement with the optimiser. I like the simplicity of the EVA head <code>then pixel-shuffled to (N, 1, 72, 72)</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237519,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "07/01/2025 06:29:41",
      "content": "<p>Congratz. I followed the same path i.e. just training a large vit variation on lot of generated data without using the test data. You said 'By that time, my final training job was already scheduled, and I decided not to change it.'. So training your model is only 2 days? Or was only final fine-tuning? If so, how much time it take to train your full model and what is the required compute?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237577,
          "author_name": "tascj0",
          "author_url": "",
          "post_date": "07/01/2025 07:15:03",
          "content": "<p>I started preparing submission after stage 2 started.</p>\n<p>I used <code>A100 SXM 80G x4</code>.</p>\n<ol>\n<li>Data generation: ~25 hours</li>\n<li>Stage 1: ~74 hours</li>\n<li>Stage 2: ~26 hours</li>\n<li>Stage 3: ~6.5 hours</li>\n</ol>\n<p>So the full pipeline took ~5.5 days.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237264": "Thank you kaggle and the host team for hosting this competition. Congratulations to the winners!\n\n\nMy solution leverages synthetic data generation and Vision Transformer for modeling. Requiring lots of storage and compute resources.\nI used more than 13TB of storage and ~5.5 days of `A100 SXM 80G x4`.\n\n\n## Forward Modeling\n\nI started with [this implementation](https://www.kaggle.com/code/manatoyo/improved-vel-to-seis) by @manatoyo (developed with the help of Google's Gemini). And made the following changes:\n\n1. Change dtype from float64 to float32. This reduced errors a bit. Matlab uses float64 by default, and Gemini knows that. However, OpenFWI seems to be created using a float32 implementation.\n2. Generate 6 channels `isx=[120, 137, 154, 155, 172, 189]` to enable fast hflip augmentation.\n3. Port to PyTorch and enable batch execution on GPU.\n\nFor some vel types, there are large errors in the last row of seismic data between OpenFWI data and generated data, so I zeroed the last row of all data during training.\n\nOpenFWI is not fully open. Their code may have some minor bugs, and if there are serious bugs in their forward modeling code, there may be little we can do about it. Fortunately, the issues don't seem significant at the moment.\n\n## Generating data\n\nI generated the following data for training:\n1. 6ch version of Full OpenFWI data.\n2. random blending (alpha in [0.3, 0.7]) of two vels from same fold, same vel-type. repeated 8 times with different seeds.\n3. random blending (alpha in [0.3, 0.7]) of two vels from same fold, not restricted to same vel-type. repeated 8 times with different seeds.\n\nIn my short training runs, generating more data consistently leads to performance improvements.\n\n## Modeling\n\nI started with ViT with RoPE (`eva02_base_patch16_clip_224`) for this task.\n\n1. I first re-aranged the input from (N, 5, 1000, 70) to (N, 5, 250, 280), then interpolated it to (N, 5, patch_size * 18, patch_size * 18).\n2. The input is then fed to the ViT model, predicted 16-dim outputs, then pixel-shuffled to (N, 1, 72, 72) and cropped to (N, 1, 70, 70).\n3. The logits are transformed to target by `logits.sigmoid() * 6000`.\n\nThen I split ViT layers:\n1. First 1/3 layers for intra-channel feature interaction.\n2. Average pool 5 channels to 1 channel.\n3. Last 2/3 layers for further feature interaction.\n\n\n## Optimizer\n\nI started with a short `CurveFault_A` only training run. Switching from AdamW to `MuonWithAuxAdam` was a big improvement (9.x -> 7.x). I adapted the implementation from `KellerJordan/Muon` and followed the practices from `MoonshotAI/Moonlight`. Muon was used for weights of all `Linear` layers except patch embedding, for other parameters AdamW was used.\n\nI haven't achieved significant performance improvements with optimizers other than SGD and Adam for many years, so this was a surprise for me, perhaps the biggest takeaway from this competition. The only drawback is that Muon's overhead becomes non-negligible when using small batch sizes.\n\n\n## Submission\n\nThe final model was `eva02_large_patch14_clip_224` trained in 3 stages.\n\n### Stage 1\n\n* HFlip augmentation.\n* Input resized to (14 * 18, 14 * 18), 4x output upsampling.\n* 10 epochs on original data and generated data\n* 50% constant learning rate + 50% cosine decay learning rate\n\nMAE on 20% validation data: 8.1\n\n### Stage 2\n\n* Resume from stage 1\n* HFlip augmentation.\n* Input resized to (14 * 35, 14 * 35), 2x output upsampling.\n* 1 epochs on original data and generated data\n* 50% constant learning rate + 50% cosine decay learning rate\n\nMAE on 20% validation data: 7.6\n\n### Stage 3\n\n* Resume from stage 2\n* HFlip augmentation.\n* Input resized to (14 * 35, 14 * 35), 2x output upsampling.\n* 1 epochs on validation data\n* cosine decay learning rate\n\nMore synthetic data and more epochs could further improve the performance, this was the best I could finish with my GPU hours.\n\nSubmission was a 20%+30%+50% blending of models from 3 stages. With `x.clip(1500, 4500).round()`.\n\n**Update**\nI made a mistake in the inference script by not zeroing the last row, which is inconsistent with training. After fixing the bug, the score is (Public 7.96, Private 7.99).\nFortunately, it doesn't affect the ranking.\n\n## Code\n\n[Submitted file](https://www.kaggle.com/datasets/tascj0/fwi-submission)\n[Inference notebook](https://www.kaggle.com/code/tascj0/fwi-inference)\n[Training code](https://github.com/tascj/kaggle-waveform-inversion)",
    "3237299": "Congratulations, looking forward to your `MuonWithAuxAdam`.",
    "3237307": "I have tried MuonWithAuxAdam in early experiments, but it will output NaN in training loop and I have to give it up. Looking forward to code!",
    "3237312": "Congratulations!\n\nI asked same question to another winner: What is your technique to prevent overfit in this one competition? \n\nThank for reply.",
    "3237388": "In this competition, you can generate a large amount of data. For me, train with more data is the key to prevent overfitting.\n\nHowever, don’t underestimate the impact of underfitting — you need a model with sufficient capacity and enough training to fully take advantage of data.",
    "3237420": "Yes, i was struggling with both overfit and underfit.\nI am looking forward to your code solution (hopefully in  Kaggle notebook 😀)",
    "3237472": "Congratulations on your third place.\n\nIf I understand correctly, your method can efficiently infer new test data, right? In other words, your computationally intensive training step does not use the test data?\n\nThis would make your solution quite different from @harshitsheoran's and mine. It would mean that yours is the highest solution that would do well if this were a code competition...",
    "3237505": "Yes. I did not use the test data for training.\n\nI started this competition downloading the full OpenFWI dataset and initially did not download the Kaggle data. I didn’t realize this is a CSV competition until I started preparing for submission (2 days before the deadline). By that time, my final training job was already scheduled, and I decided not to change it.\n\nBy the way, I built pipeline to optimize predicted vel by forward modeling + backpropagation soon after I started working on this competition, but dropped that idea because it would be too slow for a code competition.",
    "3237514": "Nice work, great improvement with the optimiser. I like the simplicity of the EVA head `then pixel-shuffled to (N, 1, 72, 72)`.",
    "3237519": "Congratz. I followed the same path i.e. just training a large vit variation on lot of generated data without using the test data. You said 'By that time, my final training job was already scheduled, and I decided not to change it.'. So training your model is only 2 days? Or was only final fine-tuning? If so, how much time it take to train your full model and what is the required compute?",
    "3237577": "I started preparing submission after stage 2 started.\n\nI used `A100 SXM 80G x4`.\n\n1. Data generation: ~25 hours\n2. Stage 1: ~74 hours\n3. Stage 2: ~26 hours\n4. Stage 3: ~6.5 hours\n\nSo the full pipeline took ~5.5 days.",
    "3238682": "Inference code&checkpoints, training code are available now.",
    "3238685": "Could you also share the notebook that runs the full test set and writes submission.csv (if you did this on Kaggle that is)?",
    "3238730": "I made predictions locally. Command shared on GitHub.\nUploaded the submitted file [here](https://www.kaggle.com/datasets/tascj0/fwi-submission).",
    "3239166": "Thanks. What I was mainly curious about: how long does inferring the full test set take? (This being 1000s of GPU hours for most solutions, but not yours I think).",
    "3240532": "Inferring the full test set using stage2 model took 30+mins using single A100. Less than 90 mins to make the submission."
  },
  "source": "meta"
}