{
  "id": 587460,
  "title": "6th Place Solution Summary",
  "url": "/competitions/waveform-inversion/writeups/hyd-6th-place-solution-summary",
  "author_name": "",
  "post_date": "2025-07-01T06:54:31.790Z",
  "votes": 45,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First, I’d like to thank the competition organizers and the open-source community—your work and shared code were invaluable in helping me learn and improve.</p>\n<p><strong>1. Synthetic Training Data Generation</strong></p>\n<p>To augment the dataset, I generated approximately 10 million additional training samples. To optimize storage efficiency, inputs were downsampled to dimensions (5, 500, 70) and stored in np.float16 format. The synthetic data consisted of:</p>\n<p>40% Gaussian noise/random rotations/scaling/shifts</p>\n<p>40% Mixup-augmented samples</p>\n<p>20% Samples generated via Denoising Diffusion Implicit Models (DDIM)</p>\n<p>Code references:<br>\n<a href=\"https://www.kaggle.com/code/manatoyo/improved-vel-to-seis/notebook\" target=\"_blank\">Improved VEL to SEIS</a><br>\n<a href=\"https://www.kaggle.com/code/jaewook704/waveform-inversion-vel-to-seis\" target=\"_blank\">Waveform Inversion (VEL to SEIS)</a></p>\n<p><strong>2. Model Architecture</strong></p>\n<p>ConvNeXtV2-Base (convnextv2_base.fcmae_ft_in22k_in1k_384)</p>\n<p>CaFormer-B36 (caformer_b36.sail_in22k_ft_in1k_384)</p>\n<p>Modifications were adapted from:</p>\n<p><a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">CaFormer Full-Resolution Improved</a><br>\n<a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline\" target=\"_blank\">ConvNeXt Full-Resolution Baseline</a></p>\n<p><strong>3. Training Strategy</strong><br>\n<em>Stage 1 (Primary Training)</em></p>\n<p>Trained for 60 epochs on the synthetic dataset.</p>\n<p>Cross-validation (CV) scores:</p>\n<p>CaFormer: 9.2/ConvNeXtV2: 10.2/Ensemble: 8.1</p>\n<p>Leaderboard (LB) / Privateboard  (PB): 9.9 / 10.0</p>\n<p><em>Stage 2 (Test-Time Fine-Tuning)</em></p>\n<p>Further fine-tuned for 2 epochs on synthetic test-derived data.</p>\n<p>Generated additional training samples via:</p>\n<p>Forward modeling predictions on test data (y → x).</p>\n<p>Applied horizontal flipping (HFlip) for augmentation.</p>\n<p>Total synthetic data: 300k samples (from 75,818 test/validation samples × 2 models × 2 flips).</p>\n<p>Resulting CV scores:</p>\n<p>CaFormer: 8.6/ConvNeXtV2: 9.0/Ensemble: 7.6<br>\nLeaderboard (LB) / Privateboard  (PB): 8.8 / 8.9</p>\n<p><strong>4. Late-Stage Insights</strong></p>\n<p>Discovered Stage 2’s significant impact in the final hours (~1.1 LB improvement).</p>\n<p>Training was halted 30 minutes before the deadline—further data generation and training could likely have yielded additional gains.</p>",
  "messages": [
    {
      "id": "3237533",
      "postDate": "07/01/2025 06:41:22",
      "content": "<p>First, I’d like to thank the competition organizers and the open-source community—your work and shared code were invaluable in helping me learn and improve.</p>\n<p><strong>1. Synthetic Training Data Generation</strong></p>\n<p>To augment the dataset, I generated approximately 10 million additional training samples. To optimize storage efficiency, inputs were downsampled to dimensions (5, 500, 70) and stored in np.float16 format. The synthetic data consisted of:</p>\n<p>40% Gaussian noise/random rotations/scaling/shifts</p>\n<p>40% Mixup-augmented samples</p>\n<p>20% Samples generated via Denoising Diffusion Implicit Models (DDIM)</p>\n<p>Code references:<br>\n<a href=\"https://www.kaggle.com/code/manatoyo/improved-vel-to-seis/notebook\" target=\"_blank\">Improved VEL to SEIS</a><br>\n<a href=\"https://www.kaggle.com/code/jaewook704/waveform-inversion-vel-to-seis\" target=\"_blank\">Waveform Inversion (VEL to SEIS)</a></p>\n<p><strong>2. Model Architecture</strong></p>\n<p>ConvNeXtV2-Base (convnextv2_base.fcmae_ft_in22k_in1k_384)</p>\n<p>CaFormer-B36 (caformer_b36.sail_in22k_ft_in1k_384)</p>\n<p>Modifications were adapted from:</p>\n<p><a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">CaFormer Full-Resolution Improved</a><br>\n<a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline\" target=\"_blank\">ConvNeXt Full-Resolution Baseline</a></p>\n<p><strong>3. Training Strategy</strong><br>\n<em>Stage 1 (Primary Training)</em></p>\n<p>Trained for 60 epochs on the synthetic dataset.</p>\n<p>Cross-validation (CV) scores:</p>\n<p>CaFormer: 9.2/ConvNeXtV2: 10.2/Ensemble: 8.1</p>\n<p>Leaderboard (LB) / Privateboard  (PB): 9.9 / 10.0</p>\n<p><em>Stage 2 (Test-Time Fine-Tuning)</em></p>\n<p>Further fine-tuned for 2 epochs on synthetic test-derived data.</p>\n<p>Generated additional training samples via:</p>\n<p>Forward modeling predictions on test data (y → x).</p>\n<p>Applied horizontal flipping (HFlip) for augmentation.</p>\n<p>Total synthetic data: 300k samples (from 75,818 test/validation samples × 2 models × 2 flips).</p>\n<p>Resulting CV scores:</p>\n<p>CaFormer: 8.6/ConvNeXtV2: 9.0/Ensemble: 7.6<br>\nLeaderboard (LB) / Privateboard  (PB): 8.8 / 8.9</p>\n<p><strong>4. Late-Stage Insights</strong></p>\n<p>Discovered Stage 2’s significant impact in the final hours (~1.1 LB improvement).</p>\n<p>Training was halted 30 minutes before the deadline—further data generation and training could likely have yielded additional gains.</p>",
      "rawMarkdown": "First, I’d like to thank the competition organizers and the open-source community—your work and shared code were invaluable in helping me learn and improve.\n\n\n**1. Synthetic Training Data Generation**\n\nTo augment the dataset, I generated approximately 10 million additional training samples. To optimize storage efficiency, inputs were downsampled to dimensions (5, 500, 70) and stored in np.float16 format. The synthetic data consisted of:\n\n40% Gaussian noise/random rotations/scaling/shifts\n\n40% Mixup-augmented samples\n\n20% Samples generated via Denoising Diffusion Implicit Models (DDIM)\n\nCode references:\n[Improved VEL to SEIS]( https://www.kaggle.com/code/manatoyo/improved-vel-to-seis/notebook)\n[Waveform Inversion (VEL to SEIS)](https://www.kaggle.com/code/jaewook704/waveform-inversion-vel-to-seis)\n\n**2. Model Architecture**\n\nConvNeXtV2-Base (convnextv2_base.fcmae_ft_in22k_in1k_384)\n\nCaFormer-B36 (caformer_b36.sail_in22k_ft_in1k_384)\n\nModifications were adapted from:\n\n[CaFormer Full-Resolution Improved](https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved)\n[ConvNeXt Full-Resolution Baseline]( https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline)\n\n**3. Training Strategy**\n*Stage 1 (Primary Training)*\n\nTrained for 60 epochs on the synthetic dataset.\n\nCross-validation (CV) scores:\n\nCaFormer: 9.2/ConvNeXtV2: 10.2/Ensemble: 8.1\n\nLeaderboard (LB) / Privateboard  (PB): 9.9 / 10.0\n\n\n*Stage 2 (Test-Time Fine-Tuning)*\n\nFurther fine-tuned for 2 epochs on synthetic test-derived data.\n\nGenerated additional training samples via:\n\nForward modeling predictions on test data (y → x).\n\nApplied horizontal flipping (HFlip) for augmentation.\n\nTotal synthetic data: 300k samples (from 75,818 test/validation samples × 2 models × 2 flips).\n\nResulting CV scores:\n\nCaFormer: 8.6/ConvNeXtV2: 9.0/Ensemble: 7.6\nLeaderboard (LB) / Privateboard  (PB): 8.8 / 8.9\n\n**4. Late-Stage Insights**\n\nDiscovered Stage 2’s significant impact in the final hours (~1.1 LB improvement).\n\nTraining was halted 30 minutes before the deadline—further data generation and training could likely have yielded additional gains.",
      "votes": null
    },
    {
      "id": "3237544",
      "postDate": "07/01/2025 06:49:06",
      "content": "<p>I'm so lucky you didn't jump higher! 🤣  <br>\nTraining on test-predict-generation data was also a key for 1st place! Good job for finding it. </p>",
      "rawMarkdown": "I'm so lucky you didn't jump higher! 🤣  \nTraining on test-predict-generation data was also a key for 1st place! Good job for finding it.",
      "votes": null
    },
    {
      "id": "3237551",
      "postDate": "07/01/2025 06:56:48",
      "content": "<p>Great write up, appreciate you sharing this!</p>",
      "rawMarkdown": "Great write up, appreciate you sharing this!",
      "votes": null
    },
    {
      "id": "3237679",
      "postDate": "07/01/2025 08:24:35",
      "content": "<p>What an elegant solution and an elegant write-up! Huge congrats and thx for sharing 🙂 🙏</p>",
      "rawMarkdown": "What an elegant solution and an elegant write-up! Huge congrats and thx for sharing 🙂 🙏",
      "votes": null
    },
    {
      "id": "3237709",
      "postDate": "07/01/2025 08:52:17",
      "content": "<p>Fantastic job! thx for sharing! 😀</p>",
      "rawMarkdown": "Fantastic job! thx for sharing! 😀",
      "votes": null
    },
    {
      "id": "3237867",
      "postDate": "07/01/2025 10:58:20",
      "content": "<p>Hi! Great work.<br>\nHow have you got the Leaderboard (LB) / Privateboard (PB): 9.9 / 10.0 and metrics in other case, if you have only one submission on LB?</p>",
      "rawMarkdown": "Hi! Great work.\nHow have you got the Leaderboard (LB) / Privateboard (PB): 9.9 / 10.0 and metrics in other case, if you have only one submission on LB?",
      "votes": null
    },
    {
      "id": "3237964",
      "postDate": "07/01/2025 12:28:06",
      "content": "<p>Late Submission.</p>",
      "rawMarkdown": "Late Submission.",
      "votes": null
    },
    {
      "id": "3238399",
      "postDate": "07/01/2025 19:57:37",
      "content": "<p>I am curious what hardware did you use for training and about how long did this take? </p>",
      "rawMarkdown": "I am curious what hardware did you use for training and about how long did this take?",
      "votes": null
    },
    {
      "id": "3238603",
      "postDate": "07/02/2025 02:45:02",
      "content": "<p>2 8*4090 trained for 13-14 days, but maybe 8-9 days is enough, there was almost no improvement in the last few days. </p>",
      "rawMarkdown": "2 8*4090 trained for 13-14 days, but maybe 8-9 days is enough, there was almost no improvement in the last few days.",
      "votes": null
    },
    {
      "id": "3239757",
      "postDate": "07/03/2025 05:24:28",
      "content": "<p>Just one submission!😳</p>",
      "rawMarkdown": "Just one submission!😳",
      "votes": null
    },
    {
      "id": "3239813",
      "postDate": "07/03/2025 07:12:57",
      "content": "<p>I couldn't wrap my head around generating test-derived data. If I have an auto-encoder I could imagine training only the encoder on test input. But there is no VEL for VEL to SEIS for \"Forward modeling predictions on test data (y → x).\" Could you please share in more details what you did?</p>",
      "rawMarkdown": "I couldn't wrap my head around generating test-derived data. If I have an auto-encoder I could imagine training only the encoder on test input. But there is no VEL for VEL to SEIS for \"Forward modeling predictions on test data (y → x).\" Could you please share in more details what you did?",
      "votes": null
    },
    {
      "id": "3239817",
      "postDate": "07/03/2025 07:19:32",
      "content": "<p>VEL is your trained model predicitions on test data, then we can use Forward modeling to get the SEIS, and fine-tuning on the generating test-derived data.</p>",
      "rawMarkdown": "VEL is your trained model predicitions on test data, then we can use Forward modeling to get the SEIS, and fine-tuning on the generating test-derived data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237544,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "07/01/2025 06:49:06",
      "content": "<p>I'm so lucky you didn't jump higher! 🤣  <br>\nTraining on test-predict-generation data was also a key for 1st place! Good job for finding it. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237551,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/01/2025 06:56:48",
      "content": "<p>Great write up, appreciate you sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237679,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "07/01/2025 08:24:35",
      "content": "<p>What an elegant solution and an elegant write-up! Huge congrats and thx for sharing 🙂 🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237709,
      "author_name": "",
      "author_url": "",
      "post_date": "07/01/2025 08:52:17",
      "content": "<p>Fantastic job! thx for sharing! 😀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237867,
      "author_name": "ivanblch",
      "author_url": "",
      "post_date": "07/01/2025 10:58:20",
      "content": "<p>Hi! Great work.<br>\nHow have you got the Leaderboard (LB) / Privateboard (PB): 9.9 / 10.0 and metrics in other case, if you have only one submission on LB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237964,
          "author_name": "hydantess",
          "author_url": "",
          "post_date": "07/01/2025 12:28:06",
          "content": "<p>Late Submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3238399,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "07/01/2025 19:57:37",
      "content": "<p>I am curious what hardware did you use for training and about how long did this take? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3238603,
          "author_name": "hydantess",
          "author_url": "",
          "post_date": "07/02/2025 02:45:02",
          "content": "<p>2 8*4090 trained for 13-14 days, but maybe 8-9 days is enough, there was almost no improvement in the last few days. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3239757,
      "author_name": "graceqianshun",
      "author_url": "",
      "post_date": "07/03/2025 05:24:28",
      "content": "<p>Just one submission!😳</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3239813,
      "author_name": "tinkei",
      "author_url": "",
      "post_date": "07/03/2025 07:12:57",
      "content": "<p>I couldn't wrap my head around generating test-derived data. If I have an auto-encoder I could imagine training only the encoder on test input. But there is no VEL for VEL to SEIS for \"Forward modeling predictions on test data (y → x).\" Could you please share in more details what you did?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3239817,
          "author_name": "hydantess",
          "author_url": "",
          "post_date": "07/03/2025 07:19:32",
          "content": "<p>VEL is your trained model predicitions on test data, then we can use Forward modeling to get the SEIS, and fine-tuning on the generating test-derived data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237533": "First, I’d like to thank the competition organizers and the open-source community—your work and shared code were invaluable in helping me learn and improve.\n\n\n**1. Synthetic Training Data Generation**\n\nTo augment the dataset, I generated approximately 10 million additional training samples. To optimize storage efficiency, inputs were downsampled to dimensions (5, 500, 70) and stored in np.float16 format. The synthetic data consisted of:\n\n40% Gaussian noise/random rotations/scaling/shifts\n\n40% Mixup-augmented samples\n\n20% Samples generated via Denoising Diffusion Implicit Models (DDIM)\n\nCode references:\n[Improved VEL to SEIS]( https://www.kaggle.com/code/manatoyo/improved-vel-to-seis/notebook)\n[Waveform Inversion (VEL to SEIS)](https://www.kaggle.com/code/jaewook704/waveform-inversion-vel-to-seis)\n\n**2. Model Architecture**\n\nConvNeXtV2-Base (convnextv2_base.fcmae_ft_in22k_in1k_384)\n\nCaFormer-B36 (caformer_b36.sail_in22k_ft_in1k_384)\n\nModifications were adapted from:\n\n[CaFormer Full-Resolution Improved](https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved)\n[ConvNeXt Full-Resolution Baseline]( https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline)\n\n**3. Training Strategy**\n*Stage 1 (Primary Training)*\n\nTrained for 60 epochs on the synthetic dataset.\n\nCross-validation (CV) scores:\n\nCaFormer: 9.2/ConvNeXtV2: 10.2/Ensemble: 8.1\n\nLeaderboard (LB) / Privateboard  (PB): 9.9 / 10.0\n\n\n*Stage 2 (Test-Time Fine-Tuning)*\n\nFurther fine-tuned for 2 epochs on synthetic test-derived data.\n\nGenerated additional training samples via:\n\nForward modeling predictions on test data (y → x).\n\nApplied horizontal flipping (HFlip) for augmentation.\n\nTotal synthetic data: 300k samples (from 75,818 test/validation samples × 2 models × 2 flips).\n\nResulting CV scores:\n\nCaFormer: 8.6/ConvNeXtV2: 9.0/Ensemble: 7.6\nLeaderboard (LB) / Privateboard  (PB): 8.8 / 8.9\n\n**4. Late-Stage Insights**\n\nDiscovered Stage 2’s significant impact in the final hours (~1.1 LB improvement).\n\nTraining was halted 30 minutes before the deadline—further data generation and training could likely have yielded additional gains.",
    "3237544": "I'm so lucky you didn't jump higher! 🤣  \nTraining on test-predict-generation data was also a key for 1st place! Good job for finding it.",
    "3237551": "Great write up, appreciate you sharing this!",
    "3237679": "What an elegant solution and an elegant write-up! Huge congrats and thx for sharing 🙂 🙏",
    "3237709": "Fantastic job! thx for sharing! 😀",
    "3237867": "Hi! Great work.\nHow have you got the Leaderboard (LB) / Privateboard (PB): 9.9 / 10.0 and metrics in other case, if you have only one submission on LB?",
    "3237964": "Late Submission.",
    "3238399": "I am curious what hardware did you use for training and about how long did this take?",
    "3238603": "2 8*4090 trained for 13-14 days, but maybe 8-9 days is enough, there was almost no improvement in the last few days.",
    "3239757": "Just one submission!😳",
    "3239813": "I couldn't wrap my head around generating test-derived data. If I have an auto-encoder I could imagine training only the encoder on test input. But there is no VEL for VEL to SEIS for \"Forward modeling predictions on test data (y → x).\" Could you please share in more details what you did?",
    "3239817": "VEL is your trained model predicitions on test data, then we can use Forward modeling to get the SEIS, and fine-tuning on the generating test-derived data."
  },
  "source": "meta"
}