{
  "id": 587500,
  "title": "4th place solution for the GWI competition  ",
  "url": "/competitions/waveform-inversion/writeups/greysnow-4th-place-solution-for-the-gwi-competitio",
  "author_name": "",
  "post_date": "2025-07-08T20:30:49.347Z",
  "votes": 42,
  "comment_count": 9,
  "views": 0,
  "content": "<p>It was a great competition. Many thanks to the Kaggle staff, the hosts, and all the active discussion and code section participants.</p>\n<h2>Context section</h2>\n<p><a href=\"https://www.kaggle.com/competitions/waveform-inversion\" target=\"_blank\">Business context</a>.  <br>\n<a href=\"https://www.kaggle.com/competitions/waveform-inversion/data\" target=\"_blank\">Data context</a>.  </p>\n<h2>1. Overview of the Approach</h2>\n<p>I generated a lot of additional data (probably more than 10M samples- at some point, I stopped counting). Then I trained a ~100M parameters custom ViT+2D conv model for ~3 weeks. Then I added a bit of postprocessing for the final 0.01 bump.  </p>\n<h3>1.1. Data generation</h3>\n<p>I implemented Python vel-to-seis on GPU and could generate ~13k samples per hour on T4. Then I generated for many hours on Kaggle GPU (60 hours in a week is sweet!) and colab.<br>\nIt's really slow compared to other top solutions, tho.<br>\nAt first, I generated very heavily augmented data (cut-mix between all families+scaling+shifting, etc.). I started to train my model with this, then during the training, I refined my data generation by implementing the host method from the paper for generating CurveVel and CurveFault. I could not get the same CurveFault, but I was close. For Style, I just used cutmix between samples from the same style family. I downcast the seis to 5*288*70 and saved as bfloat16.  </p>\n<h3>1.2. model</h3>\n<p>If you are familiar with my LEAP solution, it's very similar- interlaced layers of transformers and 2D conv block (it helps a lot with convergence due to bias!). The input was casted 288*70*5 -&gt; 18*18*384 patches, then passed through 28 blocks of 2Dconv+transformer, then casted back 18*18*384 -&gt; 70*70*1. Optimizar was AdamW with a half-cosine lr scheduler, starting from a max lr of 1e-3.  </p>\n<h3>1.3. Loss</h3>\n<p>MAE+confidence Loss (confidence Loss : MAE(confidence, MAE(targets, preds))).</p>\n<h3>1.4 Post-processing</h3>\n<p>It was a small bump in the last day (only 0.01!), but the difference between 5th and 4th place. But it can be ignored in the big picture.  </p>\n<h4>1.4.1</h4>\n<p>Clipped predictions between 1500 and 4500- very very minor.  </p>\n<h4>1.4.2.</h4>\n<p>I trained a classifier on the prediction. It was a very good classifier. Then, for all families except Style A+B, I rounded to int. This, together with clipping, was 0.03 in val.  </p>\n<h4>1.4.3</h4>\n<p>Again, for all families except style, for each sample I found grouped values and changed the values of each group to the most prominent one within the group (basically binned the predictions since those families had only a small number of values in vel maps- I call it 'restoring bias'). This was 0.1 in val.  </p>\n<h2>2. Details of the submission</h2>\n<h3>2.1. Ensembling.</h3>\n<p>At first, I started training a large model with horizontal flip augmentation + TTA. Then I realised, I can cheaply generate exact flipped samples by only recalculating the central seis out of 5- the rest are flip-identical. But for the middle, the source has one pixel shift after flipping. Training on exact flipping was better than flip-augmenting, so I continued my training with two versions- one with no flipping and the second with everything flipped (to use the flip-augmentation data the model already learned). I ensembled these two models for 8.7 LB.  <br>\nIn the last week, I noticed the models were a bit weaker on Style B- during training, I focused more on Curve families. So I decided to train an additional, third model in the last week, focused on Style- it was a success and achieved a bump of 2 for StyleB. Then I ensembled it with the other two, using the classifier model for calculating different weights for each family based on my validation split (10K samples in total, 1K for each family). This gave me another 0.1 bump to 0.86 (and the final 0.1 was from post processing, as I mentioned).  <br>\nAlso, I ensembled several checkpoints for each model (about 10 checkpints give or take).  </p>\n<h3>2.2 Hardware</h3>\n<p>I used Kaggle and Colab T4 for data generation and TPU for training. Spent about 400$ in total on colab credits.  </p>\n<h2>3. Sources</h2>\n<p><a href=\"https://github.com/tensorflow/models/tree/f007603b50b4db38907594a156994a4e983d2d31/official/projects/maxvit\" target=\"_blank\">MaxViT implementation in tensorflow</a>- I based my 2D conv block on their implementation with minor modifications.  </p>\n<h2>Code</h2>\n<p><a href=\"https://github.com/shlomoron/GWI-solution\" target=\"_blank\">Here</a>.</p>",
  "messages": [
    {
      "id": "3237749",
      "postDate": "07/01/2025 09:25:25",
      "content": "<p>It was a great competition. Many thanks to the Kaggle staff, the hosts, and all the active discussion and code section participants.</p>\n<h2>Context section</h2>\n<p><a href=\"https://www.kaggle.com/competitions/waveform-inversion\" target=\"_blank\">Business context</a>.  <br>\n<a href=\"https://www.kaggle.com/competitions/waveform-inversion/data\" target=\"_blank\">Data context</a>.  </p>\n<h2>1. Overview of the Approach</h2>\n<p>I generated a lot of additional data (probably more than 10M samples- at some point, I stopped counting). Then I trained a ~100M parameters custom ViT+2D conv model for ~3 weeks. Then I added a bit of postprocessing for the final 0.01 bump.  </p>\n<h3>1.1. Data generation</h3>\n<p>I implemented Python vel-to-seis on GPU and could generate ~13k samples per hour on T4. Then I generated for many hours on Kaggle GPU (60 hours in a week is sweet!) and colab.<br>\nIt's really slow compared to other top solutions, tho.<br>\nAt first, I generated very heavily augmented data (cut-mix between all families+scaling+shifting, etc.). I started to train my model with this, then during the training, I refined my data generation by implementing the host method from the paper for generating CurveVel and CurveFault. I could not get the same CurveFault, but I was close. For Style, I just used cutmix between samples from the same style family. I downcast the seis to 5*288*70 and saved as bfloat16.  </p>\n<h3>1.2. model</h3>\n<p>If you are familiar with my LEAP solution, it's very similar- interlaced layers of transformers and 2D conv block (it helps a lot with convergence due to bias!). The input was casted 288*70*5 -&gt; 18*18*384 patches, then passed through 28 blocks of 2Dconv+transformer, then casted back 18*18*384 -&gt; 70*70*1. Optimizar was AdamW with a half-cosine lr scheduler, starting from a max lr of 1e-3.  </p>\n<h3>1.3. Loss</h3>\n<p>MAE+confidence Loss (confidence Loss : MAE(confidence, MAE(targets, preds))).</p>\n<h3>1.4 Post-processing</h3>\n<p>It was a small bump in the last day (only 0.01!), but the difference between 5th and 4th place. But it can be ignored in the big picture.  </p>\n<h4>1.4.1</h4>\n<p>Clipped predictions between 1500 and 4500- very very minor.  </p>\n<h4>1.4.2.</h4>\n<p>I trained a classifier on the prediction. It was a very good classifier. Then, for all families except Style A+B, I rounded to int. This, together with clipping, was 0.03 in val.  </p>\n<h4>1.4.3</h4>\n<p>Again, for all families except style, for each sample I found grouped values and changed the values of each group to the most prominent one within the group (basically binned the predictions since those families had only a small number of values in vel maps- I call it 'restoring bias'). This was 0.1 in val.  </p>\n<h2>2. Details of the submission</h2>\n<h3>2.1. Ensembling.</h3>\n<p>At first, I started training a large model with horizontal flip augmentation + TTA. Then I realised, I can cheaply generate exact flipped samples by only recalculating the central seis out of 5- the rest are flip-identical. But for the middle, the source has one pixel shift after flipping. Training on exact flipping was better than flip-augmenting, so I continued my training with two versions- one with no flipping and the second with everything flipped (to use the flip-augmentation data the model already learned). I ensembled these two models for 8.7 LB.  <br>\nIn the last week, I noticed the models were a bit weaker on Style B- during training, I focused more on Curve families. So I decided to train an additional, third model in the last week, focused on Style- it was a success and achieved a bump of 2 for StyleB. Then I ensembled it with the other two, using the classifier model for calculating different weights for each family based on my validation split (10K samples in total, 1K for each family). This gave me another 0.1 bump to 0.86 (and the final 0.1 was from post processing, as I mentioned).  <br>\nAlso, I ensembled several checkpoints for each model (about 10 checkpints give or take).  </p>\n<h3>2.2 Hardware</h3>\n<p>I used Kaggle and Colab T4 for data generation and TPU for training. Spent about 400$ in total on colab credits.  </p>\n<h2>3. Sources</h2>\n<p><a href=\"https://github.com/tensorflow/models/tree/f007603b50b4db38907594a156994a4e983d2d31/official/projects/maxvit\" target=\"_blank\">MaxViT implementation in tensorflow</a>- I based my 2D conv block on their implementation with minor modifications.  </p>\n<h2>Code</h2>\n<p><a href=\"https://github.com/shlomoron/GWI-solution\" target=\"_blank\">Here</a>.</p>",
      "rawMarkdown": "It was a great competition. Many thanks to the Kaggle staff, the hosts, and all the active discussion and code section participants.\n\n## Context section  \n[Business context](https://www.kaggle.com/competitions/waveform-inversion).  \n[Data context](https://www.kaggle.com/competitions/waveform-inversion/data).  \n\n## 1. Overview of the Approach  \nI generated a lot of additional data (probably more than 10M samples- at some point, I stopped counting). Then I trained a ~100M parameters custom ViT+2D conv model for ~3 weeks. Then I added a bit of postprocessing for the final 0.01 bump.  \n\n### 1.1. Data generation  \nI implemented Python vel-to-seis on GPU and could generate ~13k samples per hour on T4. Then I generated for many hours on Kaggle GPU (60 hours in a week is sweet!) and colab.\nIt's really slow compared to other top solutions, tho.\nAt first, I generated very heavily augmented data (cut-mix between all families+scaling+shifting, etc.). I started to train my model with this, then during the training, I refined my data generation by implementing the host method from the paper for generating CurveVel and CurveFault. I could not get the same CurveFault, but I was close. For Style, I just used cutmix between samples from the same style family. I downcast the seis to 5\\*288\\*70 and saved as bfloat16.  \n\n### 1.2. model\nIf you are familiar with my LEAP solution, it's very similar- interlaced layers of transformers and 2D conv block (it helps a lot with convergence due to bias!). The input was casted 288\\*70\\*5 -> 18\\*18\\*384 patches, then passed through 28 blocks of 2Dconv+transformer, then casted back 18\\*18\\*384 -> 70\\*70\\*1. Optimizar was AdamW with a half-cosine lr scheduler, starting from a max lr of 1e-3.  \n\n### 1.3. Loss  \nMAE+confidence Loss (confidence Loss : MAE(confidence, MAE(targets, preds))).\n\n### 1.4 Post-processing  \nIt was a small bump in the last day (only 0.01!), but the difference between 5th and 4th place. But it can be ignored in the big picture.  \n#### 1.4.1  \nClipped predictions between 1500 and 4500- very very minor.  \n\n#### 1.4.2.  \nI trained a classifier on the prediction. It was a very good classifier. Then, for all families except Style A+B, I rounded to int. This, together with clipping, was 0.03 in val.  \n\n#### 1.4.3   \nAgain, for all families except style, for each sample I found grouped values and changed the values of each group to the most prominent one within the group (basically binned the predictions since those families had only a small number of values in vel maps- I call it 'restoring bias'). This was 0.1 in val.  \n\n## 2. Details of the submission  \n### 2.1. Ensembling.  \nAt first, I started training a large model with horizontal flip augmentation + TTA. Then I realised, I can cheaply generate exact flipped samples by only recalculating the central seis out of 5- the rest are flip-identical. But for the middle, the source has one pixel shift after flipping. Training on exact flipping was better than flip-augmenting, so I continued my training with two versions- one with no flipping and the second with everything flipped (to use the flip-augmentation data the model already learned). I ensembled these two models for 8.7 LB.  \nIn the last week, I noticed the models were a bit weaker on Style B- during training, I focused more on Curve families. So I decided to train an additional, third model in the last week, focused on Style- it was a success and achieved a bump of 2 for StyleB. Then I ensembled it with the other two, using the classifier model for calculating different weights for each family based on my validation split (10K samples in total, 1K for each family). This gave me another 0.1 bump to 0.86 (and the final 0.1 was from post processing, as I mentioned).  \nAlso, I ensembled several checkpoints for each model (about 10 checkpints give or take).  \n\n### 2.2 Hardware\nI used Kaggle and Colab T4 for data generation and TPU for training. Spent about 400$ in total on colab credits.  \n\n## 3. Sources\n[MaxViT implementation in tensorflow](https://github.com/tensorflow/models/tree/f007603b50b4db38907594a156994a4e983d2d31/official/projects/maxvit)- I based my 2D conv block on their implementation with minor modifications.  \n\n## Code\n[Here](https://github.com/shlomoron/GWI-solution).",
      "votes": null
    },
    {
      "id": "3238285",
      "postDate": "07/01/2025 17:45:27",
      "content": "<p>Congrats and thanks for sharing your solution! I thought colab has limits on runtime, how did you train for 3 weeks? Does the pro+ version have unlimited runtime as long as you have compute units?</p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution! I thought colab has limits on runtime, how did you train for 3 weeks? Does the pro+ version have unlimited runtime as long as you have compute units?",
      "votes": null
    },
    {
      "id": "3238300",
      "postDate": "07/01/2025 18:08:58",
      "content": "<p>I train for 15 hours or so, save Checkpoint, then load it in notebook that continue the train</p>",
      "rawMarkdown": "I train for 15 hours or so, save Checkpoint, then load it in notebook that continue the train",
      "votes": null
    },
    {
      "id": "3238324",
      "postDate": "07/01/2025 18:28:25",
      "content": "<p>It only cost $400—such a great deal! A few years ago, I also rented cards on Chinese platforms to compete in Kaggle Back then, I had to rely on sampling and accelerated experiments to save money. For you to make it to the top 4 in this comp with just $400 is truly impressive!</p>",
      "rawMarkdown": "It only cost $400—such a great deal! A few years ago, I also rented cards on Chinese platforms to compete in Kaggle Back then, I had to rely on sampling and accelerated experiments to save money. For you to make it to the top 4 in this comp with just $400 is truly impressive!",
      "votes": null
    },
    {
      "id": "3238335",
      "postDate": "07/01/2025 18:42:33",
      "content": "<p>Still, the competition I paid most for compute by a very large margin! <br>\nAlso, I was not so cheap compared to the other top 5, maybe a little bit cheaper, but not by order of magnitude.</p>",
      "rawMarkdown": "Still, the competition I paid most for compute by a very large margin! \nAlso, I was not so cheap compared to the other top 5, maybe a little bit cheaper, but not by order of magnitude.",
      "votes": null
    },
    {
      "id": "3238348",
      "postDate": "07/01/2025 18:56:05",
      "content": "<p>In the end, I trained two models on 8*4090 GPUs for 13-14 days, plus earlier experiments—probably cost me almost $1,000 in total!</p>",
      "rawMarkdown": "In the end, I trained two models on 8*4090 GPUs for 13-14 days, plus earlier experiments—probably cost me almost $1,000 in total!",
      "votes": null
    },
    {
      "id": "3238358",
      "postDate": "07/01/2025 19:08:18",
      "content": "<p>Damn you are rich! 😱</p>",
      "rawMarkdown": "Damn you are rich! 😱",
      "votes": null
    },
    {
      "id": "3239626",
      "postDate": "07/03/2025 02:29:03",
      "content": "<p>Thanks for sharing. Can't wait to learn from your code:)</p>",
      "rawMarkdown": "Thanks for sharing. Can't wait to learn from your code:)",
      "votes": null
    },
    {
      "id": "3239941",
      "postDate": "07/03/2025 09:55:10",
      "content": "<p>It’s nice to see confidence loss again. How much it improves your score?</p>",
      "rawMarkdown": "It’s nice to see confidence loss again. How much it improves your score?",
      "votes": null
    },
    {
      "id": "3239955",
      "postDate": "07/03/2025 10:16:23",
      "content": "<p>It improved my val on full dataset (without additional generated data) 28.37 -&gt; 27.72</p>",
      "rawMarkdown": "It improved my val on full dataset (without additional generated data) 28.37 -> 27.72",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3238285,
      "author_name": "snehalverma10",
      "author_url": "",
      "post_date": "07/01/2025 17:45:27",
      "content": "<p>Congrats and thanks for sharing your solution! I thought colab has limits on runtime, how did you train for 3 weeks? Does the pro+ version have unlimited runtime as long as you have compute units?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3238300,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "07/01/2025 18:08:58",
          "content": "<p>I train for 15 hours or so, save Checkpoint, then load it in notebook that continue the train</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3238324,
      "author_name": "hydantess",
      "author_url": "",
      "post_date": "07/01/2025 18:28:25",
      "content": "<p>It only cost $400—such a great deal! A few years ago, I also rented cards on Chinese platforms to compete in Kaggle Back then, I had to rely on sampling and accelerated experiments to save money. For you to make it to the top 4 in this comp with just $400 is truly impressive!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3238335,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "07/01/2025 18:42:33",
          "content": "<p>Still, the competition I paid most for compute by a very large margin! <br>\nAlso, I was not so cheap compared to the other top 5, maybe a little bit cheaper, but not by order of magnitude.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3238348,
              "author_name": "hydantess",
              "author_url": "",
              "post_date": "07/01/2025 18:56:05",
              "content": "<p>In the end, I trained two models on 8*4090 GPUs for 13-14 days, plus earlier experiments—probably cost me almost $1,000 in total!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3238358,
                  "author_name": "shlomoron",
                  "author_url": "",
                  "post_date": "07/01/2025 19:08:18",
                  "content": "<p>Damn you are rich! 😱</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3239626,
      "author_name": "cherylwang",
      "author_url": "",
      "post_date": "07/03/2025 02:29:03",
      "content": "<p>Thanks for sharing. Can't wait to learn from your code:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3239941,
      "author_name": "zzy990106",
      "author_url": "",
      "post_date": "07/03/2025 09:55:10",
      "content": "<p>It’s nice to see confidence loss again. How much it improves your score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3239955,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "07/03/2025 10:16:23",
          "content": "<p>It improved my val on full dataset (without additional generated data) 28.37 -&gt; 27.72</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237749": "It was a great competition. Many thanks to the Kaggle staff, the hosts, and all the active discussion and code section participants.\n\n## Context section  \n[Business context](https://www.kaggle.com/competitions/waveform-inversion).  \n[Data context](https://www.kaggle.com/competitions/waveform-inversion/data).  \n\n## 1. Overview of the Approach  \nI generated a lot of additional data (probably more than 10M samples- at some point, I stopped counting). Then I trained a ~100M parameters custom ViT+2D conv model for ~3 weeks. Then I added a bit of postprocessing for the final 0.01 bump.  \n\n### 1.1. Data generation  \nI implemented Python vel-to-seis on GPU and could generate ~13k samples per hour on T4. Then I generated for many hours on Kaggle GPU (60 hours in a week is sweet!) and colab.\nIt's really slow compared to other top solutions, tho.\nAt first, I generated very heavily augmented data (cut-mix between all families+scaling+shifting, etc.). I started to train my model with this, then during the training, I refined my data generation by implementing the host method from the paper for generating CurveVel and CurveFault. I could not get the same CurveFault, but I was close. For Style, I just used cutmix between samples from the same style family. I downcast the seis to 5\\*288\\*70 and saved as bfloat16.  \n\n### 1.2. model\nIf you are familiar with my LEAP solution, it's very similar- interlaced layers of transformers and 2D conv block (it helps a lot with convergence due to bias!). The input was casted 288\\*70\\*5 -> 18\\*18\\*384 patches, then passed through 28 blocks of 2Dconv+transformer, then casted back 18\\*18\\*384 -> 70\\*70\\*1. Optimizar was AdamW with a half-cosine lr scheduler, starting from a max lr of 1e-3.  \n\n### 1.3. Loss  \nMAE+confidence Loss (confidence Loss : MAE(confidence, MAE(targets, preds))).\n\n### 1.4 Post-processing  \nIt was a small bump in the last day (only 0.01!), but the difference between 5th and 4th place. But it can be ignored in the big picture.  \n#### 1.4.1  \nClipped predictions between 1500 and 4500- very very minor.  \n\n#### 1.4.2.  \nI trained a classifier on the prediction. It was a very good classifier. Then, for all families except Style A+B, I rounded to int. This, together with clipping, was 0.03 in val.  \n\n#### 1.4.3   \nAgain, for all families except style, for each sample I found grouped values and changed the values of each group to the most prominent one within the group (basically binned the predictions since those families had only a small number of values in vel maps- I call it 'restoring bias'). This was 0.1 in val.  \n\n## 2. Details of the submission  \n### 2.1. Ensembling.  \nAt first, I started training a large model with horizontal flip augmentation + TTA. Then I realised, I can cheaply generate exact flipped samples by only recalculating the central seis out of 5- the rest are flip-identical. But for the middle, the source has one pixel shift after flipping. Training on exact flipping was better than flip-augmenting, so I continued my training with two versions- one with no flipping and the second with everything flipped (to use the flip-augmentation data the model already learned). I ensembled these two models for 8.7 LB.  \nIn the last week, I noticed the models were a bit weaker on Style B- during training, I focused more on Curve families. So I decided to train an additional, third model in the last week, focused on Style- it was a success and achieved a bump of 2 for StyleB. Then I ensembled it with the other two, using the classifier model for calculating different weights for each family based on my validation split (10K samples in total, 1K for each family). This gave me another 0.1 bump to 0.86 (and the final 0.1 was from post processing, as I mentioned).  \nAlso, I ensembled several checkpoints for each model (about 10 checkpints give or take).  \n\n### 2.2 Hardware\nI used Kaggle and Colab T4 for data generation and TPU for training. Spent about 400$ in total on colab credits.  \n\n## 3. Sources\n[MaxViT implementation in tensorflow](https://github.com/tensorflow/models/tree/f007603b50b4db38907594a156994a4e983d2d31/official/projects/maxvit)- I based my 2D conv block on their implementation with minor modifications.  \n\n## Code\n[Here](https://github.com/shlomoron/GWI-solution).",
    "3238285": "Congrats and thanks for sharing your solution! I thought colab has limits on runtime, how did you train for 3 weeks? Does the pro+ version have unlimited runtime as long as you have compute units?",
    "3238300": "I train for 15 hours or so, save Checkpoint, then load it in notebook that continue the train",
    "3238324": "It only cost $400—such a great deal! A few years ago, I also rented cards on Chinese platforms to compete in Kaggle Back then, I had to rely on sampling and accelerated experiments to save money. For you to make it to the top 4 in this comp with just $400 is truly impressive!",
    "3238335": "Still, the competition I paid most for compute by a very large margin! \nAlso, I was not so cheap compared to the other top 5, maybe a little bit cheaper, but not by order of magnitude.",
    "3238348": "In the end, I trained two models on 8*4090 GPUs for 13-14 days, plus earlier experiments—probably cost me almost $1,000 in total!",
    "3238358": "Damn you are rich! 😱",
    "3239626": "Thanks for sharing. Can't wait to learn from your code:)",
    "3239941": "It’s nice to see confidence loss again. How much it improves your score?",
    "3239955": "It improved my val on full dataset (without additional generated data) 28.37 -> 27.72"
  },
  "source": "meta"
}