{
  "id": 587443,
  "title": "Team FAMAS. 5th place solution.",
  "url": "/competitions/waveform-inversion/discussion/587443",
  "author_name": "Yurnero",
  "post_date": "2025-07-01T05:01:53.264000",
  "votes": 43,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thanks to kaggle and everyone involved for hosting such an intense competition. It was a great 3-week run for our team FAMAS — <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>, <a href=\"https://www.kaggle.com/alehandreus\" target=\"_blank\">@alehandreus</a>, <a href=\"https://www.kaggle.com/kevin1742064161\" target=\"_blank\">@kevin1742064161</a>, <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a>, and <a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a>. The solution we achieved would not have been nearly as strong without the contributions of every team member.</p>\n<h1>Brief summary</h1>\n<p>Our solution is a hill-climb ensemble of 6 models trained on additionaly augmented data and post-processed family-wise using a differentiable version of a data simulator.</p>\n<h1>Detailed summary</h1>\n<h2>1. Resources.</h2>\n<p>In total, we used 3x8 A100 + 2x4 A100 + 8xL20 to run our experiments, train models and apply various optimization techniques. Moreover, we rented cloud 32x4090 and 4x5070, which enabled us to further optimize our results.</p>\n<h2>2. Data augmentations / additional data generation.</h2>\n<p>The 'seis-images' are the solutions of the wave equation <code>p(r,t)</code> with the medium given by the 'vel-images' <code>v(r)</code>. Therefore, we can extend the training data by solving a forward problem: generating seis-images based on vel-images. <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a> successfully managed to restore host's simulator which, which made this data generation process possible. That allowed us to use complex vel-image transformations, e.g. CutMix:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Fcd02851cb8a5763db6a8e0778d616fd1%2Fsss.jpg?generation=1751315922683308&amp;alt=media\" alt=\"\"></p>\n<p>We have also created MixUp, both vertical and horizontal stretches, wave deformations, smooth multiplications and artificial faults. Basically, each unique combination could be counted as a new family. Moreover, we believe that some families, e.g. CFB are already CutMix'ed versions of some 'easier' families.</p>\n<p>There is a problem though. Generating one additional sample takes approximately 0.4s (even with CUDA acceleration). That's why the proposed approach is only useful for creating additional train data.</p>\n<p>Of course, there are ways to apply some augmentations 'online' also. The most trivial augmentation is the flip augmentation that was used in most of the public notebooks. Additionally, if <code>p(r,t)</code> is described by <code>v(r)</code>, then <code>p(r,a⋅t)</code> is described by <code>a⋅v(r)</code>. Physically, this means that if we multiply the vel-image by <code>a</code>, the time axis of the corresponding seis-image will be compressed by <code>a</code>, or, in other words, if the velocity is higher, the propagation time is lower. Strictly speaking, this reasoning is only valid for a delta-impulse signal source. However, even when exposed to a long-duration source (as in this competition), we were able to slightly improve the model. This is effectively a way to implement runtime augmentations. We haven't used it in our final pipeline, but we still consider it as an interesting finding. </p>\n<p>Also, 5 days before the competition end deadline we found out, that SA looks like perlin noise. So we created blob-like structures using Perlin noise, applying Gaussian smoothing, then adding linear gradients and normalizing the values to the required velocity range (1500-4500 m/s):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Faef297c365cf7f91d20c91511209fda9%2Fphoto_2025-06-26_05-14-30.jpg?generation=1751320014272979&amp;alt=media\" alt=\"\"></p>\n<p>Basically, we could have generated as much additional SA data as we wanted to improve the validation score on that specific family. Additionally, we found how to improve SA even further (see the post-processing section).</p>\n<h2>3. Model architecture and training.</h2>\n<p>1. <strong>Model architecture</strong>: All the models we used can be considered as <code>caformer/convformer + unet_{pubic/modified}</code> in structure. Before teaming up and implementing data augmentation, we all observed that increasing the resolution of feature maps could significantly improve CV/LB performance. We found that Convformer showed better performance than Caformer in both CV and LB (-1.0). Our final 6 models include: 3 Convformers with 144 resolution, Caformers with 144, 160, and 256 resolutions. Among these, the  144-resolution performed the best.</p>\n<p>The following is the structure diagram of the model:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2F0a3bdd091836b3d5eb69f004fe36bea7%2Fgwi_model.png?generation=1751343067950054&amp;alt=media\" alt=\"\"></p>\n<p>2. <strong>Speed up training</strong>: Directly training with a larger resolution is a option—for example, caformer256 follows this approach. However, a more effective strategy is to first train with a smaller resolution and then scale up. For instance, we initially trained convformer72 and later fine-tuned convformer144 by resuming the weights from convformer72 before we figure out how to make data augs.<br>\n3. <strong>Data Generation</strong>: We generated a dataset that is 15 times the size of the official dataset, which can be roughly divided into 8 versions, utilizing all the data augmentation mentioned above. Among these, the more challenging classes have higher weights, with CFB, CVB, FFB, SA, and SB having significantly higher weights than the remaining classes. In the last stages of training, CFB even accounted for more than half of the training set. It is important to note that no matter how is synthetic dataset, it must be combined with the competition dataset for training; otherwise, the performance on the simpler classes will significantly deteriorate.<br>\n4. <strong>Resume training</strong>: We did not attempt to train the model with the entire 15x-sized dataset at once. The first model was trained sequentially on the first 7 versions of the synthetic dataset (e.g., augv1 → augv2 → … → augv7), ultimately achieving a CV-score of 10.26.The remaining models were trained simultaneously on multiple synthetic datasets in stages, (such as:augv1v2 → augv3v4v5 → augv6v7 → augv8).<br>\n5. <strong>Lower the LR gradually</strong>. The LR strategy for all of the models are close. Before the models haven't break a 13-CV score, we use a LR of 1e-4. After that, if we observed the CV score decreases slowly, we switch to a new version of dataset (e.g. augv6-&gt;v7) and lower the LR (1e-4 -&gt; 5e-5 -&gt; 2e-5 -&gt; 1e-5 -&gt; 5e-6 -&gt; 1e-6).</p>\n<p>The following table contains information about our models in the final ensemble.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>backbone</th>\n<th>decoder</th>\n<th>CV-Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>augv7_conv144</td>\n<td>Convformer</td>\n<td>modified Unet</td>\n<td>10.26</td>\n</tr>\n<tr>\n<td>augv8_conv144</td>\n<td>Convformer</td>\n<td>modified Unet</td>\n<td>10.28</td>\n</tr>\n<tr>\n<td>augv567_conv144</td>\n<td>Convformer</td>\n<td>modified Unet</td>\n<td>10.53</td>\n</tr>\n<tr>\n<td>augv8_ca144</td>\n<td>Caformer</td>\n<td>modified Unet</td>\n<td>11.3</td>\n</tr>\n<tr>\n<td>augv8_ca256</td>\n<td>Caformer</td>\n<td>public Unet</td>\n<td>11.9</td>\n</tr>\n<tr>\n<td>augv8_ca160</td>\n<td>Caformer</td>\n<td>public Unet</td>\n<td>12.0</td>\n</tr>\n</tbody>\n</table>\n<h2>4. Post-processing.</h2>\n<p>4.1. <strong>Ensemble</strong>: We used the hill‐climb method to ensemble our models. Hill‐climb method is an optimization method that, at each step, selects the local change yielding the largest immediate performance gain and repeats until no significant improvement remains. Compared with mean and median methods, hill-climb improves cv by about 0.3-0.4. However, the median method of CFB is 0.4 higher than hill-climb.</p>\n<p>4.2. <a href=\"https://www.kaggle.com/alehandreus\" target=\"_blank\">@alehandreus</a> found a way to modify ours vel-to-seis forward simulator in PyTorch as a differentiable function. We used it to refine the predictions of our ensemble:</p>\n<ol>\n<li>Given <code>seis_true</code>, the ensemble outputs <code>vel_pred</code>;</li>\n<li>Differentiable forward simulation on <code>vel_pred</code> transforms it into <code>seis_pred</code>;</li>\n<li><code>vel_pred</code> pixels are optimized with gradient descent to minimize the difference between <code>seis_pred</code> and <code>seis_true</code>.</li>\n</ol>\n<h4>Style A/B families</h4>\n<p>On these families the method works without any modifications. Still, there were a couple of modifications to improve the result:</p>\n<ol>\n<li>Set individual LR for each pixel proportional to <code>vel_pred</code> values. Larger values usually have more MAE and thus receive larger LR.</li>\n<li><code>seis_pred</code> and <code>seis_true</code> have 1000 time steps. To prevent vanishing gradients at time steps close to 1000, we multipled seis by <code>linspace</code> from 0.1 to 10.</li>\n</ol>\n<h4>\"Discrete\" families</h4>\n<p>Firstly, most of the mistakes in our model predictions come from the boundaries/edges between layers:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Fe4aefabecf90a39c9f32859a6ec92cd9%2Fphoto_2025-06-22_01-05-45%20(2).jpg?generation=1751345175204562&amp;alt=media\" alt=\"\"></p>\n<p>Because of the strict layer structure, these samples required more sophisticated modifications:</p>\n<ol>\n<li>Apply Total Variation loss on <code>vel_pred</code>. A crucial thing we discovered yesterday is to use p &lt; 1 for the p-norm to encourage sharp edges.</li>\n<li>We replaced the value of the LR with the maximum value of it and its neighbours. Thus pixels one unit away from the edges still recieved larger LR.</li>\n</ol>\n<h4>FlatVel_B</h4>\n<p>Here we used a completely different approach. Predicted stripes are nearly perfect and require only slight color tuning. We replaced each stripe with one color and optimized only this one color.</p>\n<h4>Results</h4>\n<table>\n<thead>\n<tr>\n<th>Family</th>\n<th>Estimated improvement</th>\n<th>Required iterations</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CurveFault_A</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>CurveFault_B</td>\n<td>1.5</td>\n<td>250</td>\n</tr>\n<tr>\n<td>CurveVel_A</td>\n<td>0.4</td>\n<td>300</td>\n</tr>\n<tr>\n<td>CurveVel_B</td>\n<td>2.0</td>\n<td>600</td>\n</tr>\n<tr>\n<td>FlatFault_A</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>FlatFault_B</td>\n<td>1.0</td>\n<td>300</td>\n</tr>\n<tr>\n<td>FlatVel_A</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>FlatVel_B</td>\n<td>0.95</td>\n<td>80</td>\n</tr>\n<tr>\n<td>Style_A</td>\n<td>14</td>\n<td>5000</td>\n</tr>\n<tr>\n<td>Style_B</td>\n<td>6</td>\n<td>5000</td>\n</tr>\n</tbody>\n</table>\n<h4>Classifier</h4>\n<p>To split the test set into families we trained a model with the classifier head. The following are the results on the validation data subset:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2F8ce3cd8644bdd98f1986b9ec51ff2a26%2Fphoto_2025-06-15_15-31-19.jpg?generation=1751344089857990&amp;alt=media\" alt=\"\"></p>\n<p>We used that classifier to differ Style families from \"Discrete\" families in the test data, since these family types requires different optimization strategies. Moreover, we used it to perform a family-wise hill-climb ensembling.</p>\n<p>Thanks for reading. Questions welcome.</p>\n<p>Code: (in progress)</p>",
  "messages": [
    {
      "id": 3237427,
      "postDate": "2025-07-01T05:01:53.263Z",
      "content": "<p>Thanks to kaggle and everyone involved for hosting such an intense competition. It was a great 3-week run for our team FAMAS — <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>, <a href=\"https://www.kaggle.com/alehandreus\" target=\"_blank\">@alehandreus</a>, <a href=\"https://www.kaggle.com/kevin1742064161\" target=\"_blank\">@kevin1742064161</a>, <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a>, and <a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a>. The solution we achieved would not have been nearly as strong without the contributions of every team member.</p>\n<h1>Brief summary</h1>\n<p>Our solution is a hill-climb ensemble of 6 models trained on additionaly augmented data and post-processed family-wise using a differentiable version of a data simulator.</p>\n<h1>Detailed summary</h1>\n<h2>1. Resources.</h2>\n<p>In total, we used 3x8 A100 + 2x4 A100 + 8xL20 to run our experiments, train models and apply various optimization techniques. Moreover, we rented cloud 32x4090 and 4x5070, which enabled us to further optimize our results.</p>\n<h2>2. Data augmentations / additional data generation.</h2>\n<p>The 'seis-images' are the solutions of the wave equation <code>p(r,t)</code> with the medium given by the 'vel-images' <code>v(r)</code>. Therefore, we can extend the training data by solving a forward problem: generating seis-images based on vel-images. <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a> successfully managed to restore host's simulator which, which made this data generation process possible. That allowed us to use complex vel-image transformations, e.g. CutMix:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Fcd02851cb8a5763db6a8e0778d616fd1%2Fsss.jpg?generation=1751315922683308&amp;alt=media\" alt=\"\"></p>\n<p>We have also created MixUp, both vertical and horizontal stretches, wave deformations, smooth multiplications and artificial faults. Basically, each unique combination could be counted as a new family. Moreover, we believe that some families, e.g. CFB are already CutMix'ed versions of some 'easier' families.</p>\n<p>There is a problem though. Generating one additional sample takes approximately 0.4s (even with CUDA acceleration). That's why the proposed approach is only useful for creating additional train data.</p>\n<p>Of course, there are ways to apply some augmentations 'online' also. The most trivial augmentation is the flip augmentation that was used in most of the public notebooks. Additionally, if <code>p(r,t)</code> is described by <code>v(r)</code>, then <code>p(r,a⋅t)</code> is described by <code>a⋅v(r)</code>. Physically, this means that if we multiply the vel-image by <code>a</code>, the time axis of the corresponding seis-image will be compressed by <code>a</code>, or, in other words, if the velocity is higher, the propagation time is lower. Strictly speaking, this reasoning is only valid for a delta-impulse signal source. However, even when exposed to a long-duration source (as in this competition), we were able to slightly improve the model. This is effectively a way to implement runtime augmentations. We haven't used it in our final pipeline, but we still consider it as an interesting finding. </p>\n<p>Also, 5 days before the competition end deadline we found out, that SA looks like perlin noise. So we created blob-like structures using Perlin noise, applying Gaussian smoothing, then adding linear gradients and normalizing the values to the required velocity range (1500-4500 m/s):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Faef297c365cf7f91d20c91511209fda9%2Fphoto_2025-06-26_05-14-30.jpg?generation=1751320014272979&amp;alt=media\" alt=\"\"></p>\n<p>Basically, we could have generated as much additional SA data as we wanted to improve the validation score on that specific family. Additionally, we found how to improve SA even further (see the post-processing section).</p>\n<h2>3. Model architecture and training.</h2>\n<p>1. <strong>Model architecture</strong>: All the models we used can be considered as <code>caformer/convformer + unet_{pubic/modified}</code> in structure. Before teaming up and implementing data augmentation, we all observed that increasing the resolution of feature maps could significantly improve CV/LB performance. We found that Convformer showed better performance than Caformer in both CV and LB (-1.0). Our final 6 models include: 3 Convformers with 144 resolution, Caformers with 144, 160, and 256 resolutions. Among these, the  144-resolution performed the best.</p>\n<p>The following is the structure diagram of the model:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2F0a3bdd091836b3d5eb69f004fe36bea7%2Fgwi_model.png?generation=1751343067950054&amp;alt=media\" alt=\"\"></p>\n<p>2. <strong>Speed up training</strong>: Directly training with a larger resolution is a option—for example, caformer256 follows this approach. However, a more effective strategy is to first train with a smaller resolution and then scale up. For instance, we initially trained convformer72 and later fine-tuned convformer144 by resuming the weights from convformer72 before we figure out how to make data augs.<br>\n3. <strong>Data Generation</strong>: We generated a dataset that is 15 times the size of the official dataset, which can be roughly divided into 8 versions, utilizing all the data augmentation mentioned above. Among these, the more challenging classes have higher weights, with CFB, CVB, FFB, SA, and SB having significantly higher weights than the remaining classes. In the last stages of training, CFB even accounted for more than half of the training set. It is important to note that no matter how is synthetic dataset, it must be combined with the competition dataset for training; otherwise, the performance on the simpler classes will significantly deteriorate.<br>\n4. <strong>Resume training</strong>: We did not attempt to train the model with the entire 15x-sized dataset at once. The first model was trained sequentially on the first 7 versions of the synthetic dataset (e.g., augv1 → augv2 → … → augv7), ultimately achieving a CV-score of 10.26.The remaining models were trained simultaneously on multiple synthetic datasets in stages, (such as:augv1v2 → augv3v4v5 → augv6v7 → augv8).<br>\n5. <strong>Lower the LR gradually</strong>. The LR strategy for all of the models are close. Before the models haven't break a 13-CV score, we use a LR of 1e-4. After that, if we observed the CV score decreases slowly, we switch to a new version of dataset (e.g. augv6-&gt;v7) and lower the LR (1e-4 -&gt; 5e-5 -&gt; 2e-5 -&gt; 1e-5 -&gt; 5e-6 -&gt; 1e-6).</p>\n<p>The following table contains information about our models in the final ensemble.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>backbone</th>\n<th>decoder</th>\n<th>CV-Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>augv7_conv144</td>\n<td>Convformer</td>\n<td>modified Unet</td>\n<td>10.26</td>\n</tr>\n<tr>\n<td>augv8_conv144</td>\n<td>Convformer</td>\n<td>modified Unet</td>\n<td>10.28</td>\n</tr>\n<tr>\n<td>augv567_conv144</td>\n<td>Convformer</td>\n<td>modified Unet</td>\n<td>10.53</td>\n</tr>\n<tr>\n<td>augv8_ca144</td>\n<td>Caformer</td>\n<td>modified Unet</td>\n<td>11.3</td>\n</tr>\n<tr>\n<td>augv8_ca256</td>\n<td>Caformer</td>\n<td>public Unet</td>\n<td>11.9</td>\n</tr>\n<tr>\n<td>augv8_ca160</td>\n<td>Caformer</td>\n<td>public Unet</td>\n<td>12.0</td>\n</tr>\n</tbody>\n</table>\n<h2>4. Post-processing.</h2>\n<p>4.1. <strong>Ensemble</strong>: We used the hill‐climb method to ensemble our models. Hill‐climb method is an optimization method that, at each step, selects the local change yielding the largest immediate performance gain and repeats until no significant improvement remains. Compared with mean and median methods, hill-climb improves cv by about 0.3-0.4. However, the median method of CFB is 0.4 higher than hill-climb.</p>\n<p>4.2. <a href=\"https://www.kaggle.com/alehandreus\" target=\"_blank\">@alehandreus</a> found a way to modify ours vel-to-seis forward simulator in PyTorch as a differentiable function. We used it to refine the predictions of our ensemble:</p>\n<ol>\n<li>Given <code>seis_true</code>, the ensemble outputs <code>vel_pred</code>;</li>\n<li>Differentiable forward simulation on <code>vel_pred</code> transforms it into <code>seis_pred</code>;</li>\n<li><code>vel_pred</code> pixels are optimized with gradient descent to minimize the difference between <code>seis_pred</code> and <code>seis_true</code>.</li>\n</ol>\n<h4>Style A/B families</h4>\n<p>On these families the method works without any modifications. Still, there were a couple of modifications to improve the result:</p>\n<ol>\n<li>Set individual LR for each pixel proportional to <code>vel_pred</code> values. Larger values usually have more MAE and thus receive larger LR.</li>\n<li><code>seis_pred</code> and <code>seis_true</code> have 1000 time steps. To prevent vanishing gradients at time steps close to 1000, we multipled seis by <code>linspace</code> from 0.1 to 10.</li>\n</ol>\n<h4>\"Discrete\" families</h4>\n<p>Firstly, most of the mistakes in our model predictions come from the boundaries/edges between layers:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Fe4aefabecf90a39c9f32859a6ec92cd9%2Fphoto_2025-06-22_01-05-45%20(2).jpg?generation=1751345175204562&amp;alt=media\" alt=\"\"></p>\n<p>Because of the strict layer structure, these samples required more sophisticated modifications:</p>\n<ol>\n<li>Apply Total Variation loss on <code>vel_pred</code>. A crucial thing we discovered yesterday is to use p &lt; 1 for the p-norm to encourage sharp edges.</li>\n<li>We replaced the value of the LR with the maximum value of it and its neighbours. Thus pixels one unit away from the edges still recieved larger LR.</li>\n</ol>\n<h4>FlatVel_B</h4>\n<p>Here we used a completely different approach. Predicted stripes are nearly perfect and require only slight color tuning. We replaced each stripe with one color and optimized only this one color.</p>\n<h4>Results</h4>\n<table>\n<thead>\n<tr>\n<th>Family</th>\n<th>Estimated improvement</th>\n<th>Required iterations</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CurveFault_A</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>CurveFault_B</td>\n<td>1.5</td>\n<td>250</td>\n</tr>\n<tr>\n<td>CurveVel_A</td>\n<td>0.4</td>\n<td>300</td>\n</tr>\n<tr>\n<td>CurveVel_B</td>\n<td>2.0</td>\n<td>600</td>\n</tr>\n<tr>\n<td>FlatFault_A</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>FlatFault_B</td>\n<td>1.0</td>\n<td>300</td>\n</tr>\n<tr>\n<td>FlatVel_A</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>FlatVel_B</td>\n<td>0.95</td>\n<td>80</td>\n</tr>\n<tr>\n<td>Style_A</td>\n<td>14</td>\n<td>5000</td>\n</tr>\n<tr>\n<td>Style_B</td>\n<td>6</td>\n<td>5000</td>\n</tr>\n</tbody>\n</table>\n<h4>Classifier</h4>\n<p>To split the test set into families we trained a model with the classifier head. The following are the results on the validation data subset:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2F8ce3cd8644bdd98f1986b9ec51ff2a26%2Fphoto_2025-06-15_15-31-19.jpg?generation=1751344089857990&amp;alt=media\" alt=\"\"></p>\n<p>We used that classifier to differ Style families from \"Discrete\" families in the test data, since these family types requires different optimization strategies. Moreover, we used it to perform a family-wise hill-climb ensembling.</p>\n<p>Thanks for reading. Questions welcome.</p>\n<p>Code: (in progress)</p>",
      "rawMarkdown": "Thanks to kaggle and everyone involved for hosting such an intense competition. It was a great 3-week run for our team FAMAS — @forcewithme, @alehandreus, @kevin1742064161, @arsenypoyda, and @samson8. The solution we achieved would not have been nearly as strong without the contributions of every team member.\n\n# Brief summary\n\nOur solution is a hill-climb ensemble of 6 models trained on additionaly augmented data and post-processed family-wise using a differentiable version of a data simulator.\n\n# Detailed summary\n\n## 1. Resources.\n\nIn total, we used 3x8 A100 + 2x4 A100 + 8xL20 to run our experiments, train models and apply various optimization techniques. Moreover, we rented cloud 32x4090 and 4x5070, which enabled us to further optimize our results.\n\n## 2. Data augmentations / additional data generation.\n\nThe 'seis-images' are the solutions of the wave equation `p(r,t)` with the medium given by the 'vel-images' `v(r)`. Therefore, we can extend the training data by solving a forward problem: generating seis-images based on vel-images. @arsenypoyda successfully managed to restore host's simulator which, which made this data generation process possible. That allowed us to use complex vel-image transformations, e.g. CutMix:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Fcd02851cb8a5763db6a8e0778d616fd1%2Fsss.jpg?generation=1751315922683308&alt=media)\n\nWe have also created MixUp, both vertical and horizontal stretches, wave deformations, smooth multiplications and artificial faults. Basically, each unique combination could be counted as a new family. Moreover, we believe that some families, e.g. CFB are already CutMix'ed versions of some 'easier' families.\n\nThere is a problem though. Generating one additional sample takes approximately 0.4s (even with CUDA acceleration). That's why the proposed approach is only useful for creating additional train data.\n\nOf course, there are ways to apply some augmentations 'online' also. The most trivial augmentation is the flip augmentation that was used in most of the public notebooks. Additionally, if `p(r,t)` is described by `v(r)`, then `p(r,a⋅t)` is described by `a⋅v(r)`. Physically, this means that if we multiply the vel-image by `a`, the time axis of the corresponding seis-image will be compressed by `a`, or, in other words, if the velocity is higher, the propagation time is lower. Strictly speaking, this reasoning is only valid for a delta-impulse signal source. However, even when exposed to a long-duration source (as in this competition), we were able to slightly improve the model. This is effectively a way to implement runtime augmentations. We haven't used it in our final pipeline, but we still consider it as an interesting finding. \n\nAlso, 5 days before the competition end deadline we found out, that SA looks like perlin noise. So we created blob-like structures using Perlin noise, applying Gaussian smoothing, then adding linear gradients and normalizing the values to the required velocity range (1500-4500 m/s):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Faef297c365cf7f91d20c91511209fda9%2Fphoto_2025-06-26_05-14-30.jpg?generation=1751320014272979&alt=media)\n\nBasically, we could have generated as much additional SA data as we wanted to improve the validation score on that specific family. Additionally, we found how to improve SA even further (see the post-processing section).\n\n## 3. Model architecture and training.\n\n1\\. **Model architecture**: All the models we used can be considered as `caformer/convformer + unet_{pubic/modified}` in structure. Before teaming up and implementing data augmentation, we all observed that increasing the resolution of feature maps could significantly improve CV/LB performance. We found that Convformer showed better performance than Caformer in both CV and LB (-1.0). Our final 6 models include: 3 Convformers with 144 resolution, Caformers with 144, 160, and 256 resolutions. Among these, the  144-resolution performed the best.\n\nThe following is the structure diagram of the model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2F0a3bdd091836b3d5eb69f004fe36bea7%2Fgwi_model.png?generation=1751343067950054&alt=media)\n\n\n2\\. **Speed up training**: Directly training with a larger resolution is a option—for example, caformer256 follows this approach. However, a more effective strategy is to first train with a smaller resolution and then scale up. For instance, we initially trained convformer72 and later fine-tuned convformer144 by resuming the weights from convformer72 before we figure out how to make data augs.\n3\\. **Data Generation**: We generated a dataset that is 15 times the size of the official dataset, which can be roughly divided into 8 versions, utilizing all the data augmentation mentioned above. Among these, the more challenging classes have higher weights, with CFB, CVB, FFB, SA, and SB having significantly higher weights than the remaining classes. In the last stages of training, CFB even accounted for more than half of the training set. It is important to note that no matter how is synthetic dataset, it must be combined with the competition dataset for training; otherwise, the performance on the simpler classes will significantly deteriorate.\n4\\. **Resume training**: We did not attempt to train the model with the entire 15x-sized dataset at once. The first model was trained sequentially on the first 7 versions of the synthetic dataset (e.g., augv1 → augv2 → ... → augv7), ultimately achieving a CV-score of 10.26.The remaining models were trained simultaneously on multiple synthetic datasets in stages, (such as:augv1v2 → augv3v4v5 → augv6v7 → augv8).\n5\\. **Lower the LR gradually**. The LR strategy for all of the models are close. Before the models haven't break a 13-CV score, we use a LR of 1e-4. After that, if we observed the CV score decreases slowly, we switch to a new version of dataset (e.g. augv6->v7) and lower the LR (1e-4 -> 5e-5 -> 2e-5 -> 1e-5 -> 5e-6 -> 1e-6).\n\nThe following table contains information about our models in the final ensemble.\n\n| model | backbone | decoder | CV-Score|\n| --- | --- |--- |--- |\n| augv7_conv144 | Convformer | modified Unet|10.26 |\n| augv8_conv144 | Convformer | modified Unet|10.28 |\n| augv567_conv144 | Convformer |modified Unet|10.53 |\n| augv8_ca144 | Caformer |modified Unet|11.3 |\n| augv8_ca256 | Caformer |public Unet|11.9 |\n| augv8_ca160 | Caformer |public Unet|12.0 |\n\n## 4. Post-processing.\n\n4.1. **Ensemble**: We used the hill‐climb method to ensemble our models. Hill‐climb method is an optimization method that, at each step, selects the local change yielding the largest immediate performance gain and repeats until no significant improvement remains. Compared with mean and median methods, hill-climb improves cv by about 0.3-0.4. However, the median method of CFB is 0.4 higher than hill-climb.\n\n4.2. @alehandreus found a way to modify ours vel-to-seis forward simulator in PyTorch as a differentiable function. We used it to refine the predictions of our ensemble:\n\n1. Given `seis_true`, the ensemble outputs `vel_pred`;\n2. Differentiable forward simulation on `vel_pred` transforms it into `seis_pred`;\n3. `vel_pred` pixels are optimized with gradient descent to minimize the difference between `seis_pred` and `seis_true`.\n\n#### Style A/B families\nOn these families the method works without any modifications. Still, there were a couple of modifications to improve the result:\n1. Set individual LR for each pixel proportional to `vel_pred` values. Larger values usually have more MAE and thus receive larger LR.\n2. `seis_pred` and `seis_true` have 1000 time steps. To prevent vanishing gradients at time steps close to 1000, we multipled seis by `linspace` from 0.1 to 10.\n\n#### \"Discrete\" families\n Firstly, most of the mistakes in our model predictions come from the boundaries/edges between layers:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Fe4aefabecf90a39c9f32859a6ec92cd9%2Fphoto_2025-06-22_01-05-45%20(2).jpg?generation=1751345175204562&alt=media)\n\nBecause of the strict layer structure, these samples required more sophisticated modifications:\n\n1. Apply Total Variation loss on `vel_pred`. A crucial thing we discovered yesterday is to use p < 1 for the p-norm to encourage sharp edges.\n2. We replaced the value of the LR with the maximum value of it and its neighbours. Thus pixels one unit away from the edges still recieved larger LR.\n\n#### FlatVel_B\nHere we used a completely different approach. Predicted stripes are nearly perfect and require only slight color tuning. We replaced each stripe with one color and optimized only this one color.\n\n#### Results\n\n| Family    | Estimated improvement | Required iterations |\n| -------- | ------- | ------- |\n| CurveFault_A  | -    | -    |\n| CurveFault_B | 1.5     | 250  |\n| CurveVel_A    | 0.4    | 300 |\n| CurveVel_B    | 2.0    | 600 |\n| FlatFault_A    | -   | - |\n| FlatFault_B    | 1.0   | 300 |\n| FlatVel_A    |  - | - |\n| FlatVel_B    | 0.95 | 80 |\n| Style_A    | 14 | 5000 |\n| Style_B    | 6 | 5000 |\n\n#### Classifier\n\nTo split the test set into families we trained a model with the classifier head. The following are the results on the validation data subset:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2F8ce3cd8644bdd98f1986b9ec51ff2a26%2Fphoto_2025-06-15_15-31-19.jpg?generation=1751344089857990&alt=media)\n\nWe used that classifier to differ Style families from \"Discrete\" families in the test data, since these family types requires different optimization strategies. Moreover, we used it to perform a family-wise hill-climb ensembling.\n\nThanks for reading. Questions welcome.\n\nCode: (in progress)",
      "votes": 43
    },
    {
      "id": 3237431,
      "postDate": "2025-07-01T05:06:26.690Z",
      "content": "<p><strong>Addendum</strong>. Huge congratulations to <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a> on achieving Kaggle Competitions GM rank and to <a href=\"https://www.kaggle.com/alehandreus\" target=\"_blank\">@alehandreus</a> on achieving Kaggle Competitions Master rank! </p>",
      "rawMarkdown": "**Addendum**. Huge congratulations to @arsenypoyda on achieving Kaggle Competitions GM rank and to @alehandreus on achieving Kaggle Competitions Master rank! ",
      "votes": 8,
      "replies": [
        {
          "id": 3237478,
          "postDate": "2025-07-01T06:00:21.827Z",
          "content": "<p><a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a> Your coding skills have impressed me, and the data aug tools you provided are the main reason we've consistently stayed in the gold medal zone. You absolutely deserve the title of Kaggle Grandmaster. Congrats. 💯 </p>",
          "rawMarkdown": "@arsenypoyda Your coding skills have impressed me, and the data aug tools you provided are the main reason we've consistently stayed in the gold medal zone. You absolutely deserve the title of Kaggle Grandmaster. Congrats. 💯 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 3241701,
      "postDate": "2025-07-05T04:48:32.750Z",
      "content": "<p>Big congrats, and thank you for sharing. I'm eager to learn and your write up helped!</p>",
      "rawMarkdown": "Big congrats, and thank you for sharing. I'm eager to learn and your write up helped!",
      "votes": 6
    },
    {
      "id": 3237489,
      "postDate": "2025-07-01T06:10:44.803Z",
      "content": "<p>Good job. It was really close, I think if the competition was a bit longer you would have left me behind. The compute you used is crazy! Can I ask how much $ you estimate your total compute cost? Like, 32*4090 is just crasy. 3*8A100 also.</p>",
      "rawMarkdown": "Good job. It was really close, I think if the competition was a bit longer you would have left me behind. The compute you used is crazy! Can I ask how much $ you estimate your total compute cost? Like, 32\\*4090 is just crasy. 3\\*8A100 also.",
      "votes": 3,
      "replies": [
        {
          "id": 3237706,
          "postDate": "2025-07-01T08:49:49.510Z",
          "content": "<p>That is a good question.</p>\n<p>I have an access to 8xA100 instance + 4xA100 instance (thanks to my company). All the other instances were by provided by <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> (mostly we used his cluster to train all of our models).</p>\n<p>The interesting thing is why we rented 32x4090 and 4x5070. It's all about our post-processing. The bottlneck we encoutered was the core frequency, which is much higher for the ADA family GPUs (e.g. 4090) vs the Ampere family (e.g. A100). 1 iteration on a batch take ~3.2-3.7s on A100 vs 1.2s on 4090. And we needed a lot of iterations 1-2 days before the competition deadline…</p>\n<p>In total we spent ~$500-600</p>",
          "rawMarkdown": "That is a good question.\n\nI have an access to 8xA100 instance + 4xA100 instance (thanks to my company). All the other instances were by provided by @forcewithme (mostly we used his cluster to train all of our models).\n\nThe interesting thing is why we rented 32x4090 and 4x5070. It's all about our post-processing. The bottlneck we encoutered was the core frequency, which is much higher for the ADA family GPUs (e.g. 4090) vs the Ampere family (e.g. A100). 1 iteration on a batch take ~3.2-3.7s on A100 vs 1.2s on 4090. And we needed a lot of iterations 1-2 days before the competition deadline...\n\nIn total we spent ~$500-600",
          "votes": 1,
          "replies": [
            {
              "id": 3238027,
              "postDate": "2025-07-01T13:22:36.130Z",
              "content": "<p>Damn you have so powerfull compute for free?<br>\nI want to work in such conpany too! 🤣<br>\nThis conpetition is not fair 🤣</p>",
              "rawMarkdown": "Damn you have so powerfull compute for free?\nI want to work in such conpany too! 🤣\nThis conpetition is not fair 🤣",
              "votes": 1
            },
            {
              "id": 3238645,
              "postDate": "2025-07-02T04:24:27.940Z",
              "content": "<p>\"I want to work in such conpany too!\"</p>\n<p>why don't top kaggler start a gpu server company? During competition, founding members have exclusive access to it (or others needs to pay membership fee, etc). When they are not competition, the gpu are rent to cloud provider to generate income.</p>\n<p>i would like to have  system that i can block reseved gpu (e.g. over 3 week period in competition time)</p>\n<p>actually google cloud should comes up with special package/plan for kaggler. There should be a better buiness model. The gpu usage pattern/requirement for competition is quire different from that of other business</p>",
              "rawMarkdown": "\"I want to work in such conpany too!\"\n\nwhy don't top kaggler start a gpu server company? During competition, founding members have exclusive access to it (or others needs to pay membership fee, etc). When they are not competition, the gpu are rent to cloud provider to generate income.\n\ni would like to have  system that i can block reseved gpu (e.g. over 3 week period in competition time)\n\nactually google cloud should comes up with special package/plan for kaggler. There should be a better buiness model. The gpu usage pattern/requirement for competition is quire different from that of other business"
            }
          ]
        }
      ]
    },
    {
      "id": 3237483,
      "postDate": "2025-07-01T06:04:16.293Z",
      "content": "<p>Congratulations! Your solution is very similar to mine, it's quite possible the difference in our scores is mainly down to how much compute cost and environmental impact we were willing to accept…</p>",
      "rawMarkdown": "Congratulations! Your solution is very similar to mine, it's quite possible the difference in our scores is mainly down to how much compute cost and environmental impact we were willing to accept...",
      "votes": 3,
      "replies": [
        {
          "id": 3237492,
          "postDate": "2025-07-01T06:12:25.040Z",
          "content": "<p>You used <em>more</em> compute?! How do you have the resources, damn I jealous.  </p>",
          "rawMarkdown": "You used *more* compute?! How do you have the resources, damn I jealous.  ",
          "votes": 1,
          "replies": [
            {
              "id": 3237496,
              "postDate": "2025-07-01T06:13:52.243Z",
              "content": "<p>Ah I missed where they explained how much compute they actually used. No, I don't think I used more, though I don't know how long they ran that massive GPU cluster.</p>\n<p>My main calculations ran on 8xV100 (relatively cheap low-power GPU) over 3 weeks.</p>",
              "rawMarkdown": "Ah I missed where they explained how much compute they actually used. No, I don't think I used more, though I don't know how long they ran that massive GPU cluster.\n\nMy main calculations ran on 8xV100 (relatively cheap low-power GPU) over 3 weeks.",
              "votes": 1
            },
            {
              "id": 3237512,
              "postDate": "2025-07-01T06:22:26.977Z",
              "content": "<p>You call that cheap?! If we take 0.7$ per 1 A100 than 0.7*8*24*21=2800 $ approx, am I correct?</p>",
              "rawMarkdown": "You call that cheap?! If we take 0.7$ per 1 A100 than 0.7\\*8\\*24\\*21=2800 $ approx, am I correct?"
            },
            {
              "id": 3237517,
              "postDate": "2025-07-01T06:26:42.567Z",
              "content": "<p>V100, not A100. They cost me about $0.12 per hour. I call them <strong>relatively</strong> cheap 😄</p>",
              "rawMarkdown": "V100, not A100. They cost me about $0.12 per hour. I call them **relatively** cheap 😄",
              "votes": 2
            },
            {
              "id": 3237525,
              "postDate": "2025-07-01T06:35:25.587Z",
              "content": "<p>Ah, this is much better, yes. So we spent about the same amount. Good job!</p>",
              "rawMarkdown": "Ah, this is much better, yes. So we spent about the same amount. Good job!"
            },
            {
              "id": 3241689,
              "postDate": "2025-07-05T04:14:15.100Z",
              "content": "<p>I also missed… </p>",
              "rawMarkdown": "I also missed... "
            }
          ]
        }
      ]
    },
    {
      "id": 3241987,
      "postDate": "2025-07-05T12:18:10.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3237431,
      "author_name": "Yurnero",
      "author_url": "",
      "post_date": "2025-07-01T05:06:26.690000",
      "content": "<p><strong>Addendum</strong>. Huge congratulations to <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a> on achieving Kaggle Competitions GM rank and to <a href=\"https://www.kaggle.com/alehandreus\" target=\"_blank\">@alehandreus</a> on achieving Kaggle Competitions Master rank! </p>",
      "votes": 8,
      "replies": [
        {
          "id": 3237478,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2025-07-01T06:00:21.827000",
          "content": "<p><a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a> Your coding skills have impressed me, and the data aug tools you provided are the main reason we've consistently stayed in the gold medal zone. You absolutely deserve the title of Kaggle Grandmaster. Congrats. 💯 </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3241701,
      "author_name": "Taylor S. Amarel",
      "author_url": "",
      "post_date": "2025-07-05T04:48:32.750000",
      "content": "<p>Big congrats, and thank you for sharing. I'm eager to learn and your write up helped!</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 3237489,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2025-07-01T06:10:44.803000",
      "content": "<p>Good job. It was really close, I think if the competition was a bit longer you would have left me behind. The compute you used is crazy! Can I ask how much $ you estimate your total compute cost? Like, 32*4090 is just crasy. 3*8A100 also.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3237706,
          "author_name": "Yurnero",
          "author_url": "",
          "post_date": "2025-07-01T08:49:49.510000",
          "content": "<p>That is a good question.</p>\n<p>I have an access to 8xA100 instance + 4xA100 instance (thanks to my company). All the other instances were by provided by <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> (mostly we used his cluster to train all of our models).</p>\n<p>The interesting thing is why we rented 32x4090 and 4x5070. It's all about our post-processing. The bottlneck we encoutered was the core frequency, which is much higher for the ADA family GPUs (e.g. 4090) vs the Ampere family (e.g. A100). 1 iteration on a batch take ~3.2-3.7s on A100 vs 1.2s on 4090. And we needed a lot of iterations 1-2 days before the competition deadline…</p>\n<p>In total we spent ~$500-600</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3238027,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-07-01T13:22:36.130000",
              "content": "<p>Damn you have so powerfull compute for free?<br>\nI want to work in such conpany too! 🤣<br>\nThis conpetition is not fair 🤣</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3238645,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-07-02T04:24:27.940000",
              "content": "<p>\"I want to work in such conpany too!\"</p>\n<p>why don't top kaggler start a gpu server company? During competition, founding members have exclusive access to it (or others needs to pay membership fee, etc). When they are not competition, the gpu are rent to cloud provider to generate income.</p>\n<p>i would like to have  system that i can block reseved gpu (e.g. over 3 week period in competition time)</p>\n<p>actually google cloud should comes up with special package/plan for kaggler. There should be a better buiness model. The gpu usage pattern/requirement for competition is quire different from that of other business</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3237483,
      "author_name": "Jeroen Cottaar",
      "author_url": "",
      "post_date": "2025-07-01T06:04:16.293000",
      "content": "<p>Congratulations! Your solution is very similar to mine, it's quite possible the difference in our scores is mainly down to how much compute cost and environmental impact we were willing to accept…</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3237492,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2025-07-01T06:12:25.040000",
          "content": "<p>You used <em>more</em> compute?! How do you have the resources, damn I jealous.  </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3237496,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-07-01T06:13:52.243000",
              "content": "<p>Ah I missed where they explained how much compute they actually used. No, I don't think I used more, though I don't know how long they ran that massive GPU cluster.</p>\n<p>My main calculations ran on 8xV100 (relatively cheap low-power GPU) over 3 weeks.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3237512,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-07-01T06:22:26.977000",
              "content": "<p>You call that cheap?! If we take 0.7$ per 1 A100 than 0.7*8*24*21=2800 $ approx, am I correct?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3237517,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-07-01T06:26:42.567000",
              "content": "<p>V100, not A100. They cost me about $0.12 per hour. I call them <strong>relatively</strong> cheap 😄</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3237525,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-07-01T06:35:25.587000",
              "content": "<p>Ah, this is much better, yes. So we spent about the same amount. Good job!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3241689,
              "author_name": "Mehwish",
              "author_url": "",
              "post_date": "2025-07-05T04:14:15.100000",
              "content": "<p>I also missed… </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3241987,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-05T12:18:10.170000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3237427": "Thanks to kaggle and everyone involved for hosting such an intense competition. It was a great 3-week run for our team FAMAS — @forcewithme, @alehandreus, @kevin1742064161, @arsenypoyda, and @samson8. The solution we achieved would not have been nearly as strong without the contributions of every team member.\n\n# Brief summary\n\nOur solution is a hill-climb ensemble of 6 models trained on additionaly augmented data and post-processed family-wise using a differentiable version of a data simulator.\n\n# Detailed summary\n\n## 1. Resources.\n\nIn total, we used 3x8 A100 + 2x4 A100 + 8xL20 to run our experiments, train models and apply various optimization techniques. Moreover, we rented cloud 32x4090 and 4x5070, which enabled us to further optimize our results.\n\n## 2. Data augmentations / additional data generation.\n\nThe 'seis-images' are the solutions of the wave equation `p(r,t)` with the medium given by the 'vel-images' `v(r)`. Therefore, we can extend the training data by solving a forward problem: generating seis-images based on vel-images. @arsenypoyda successfully managed to restore host's simulator which, which made this data generation process possible. That allowed us to use complex vel-image transformations, e.g. CutMix:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Fcd02851cb8a5763db6a8e0778d616fd1%2Fsss.jpg?generation=1751315922683308&alt=media)\n\nWe have also created MixUp, both vertical and horizontal stretches, wave deformations, smooth multiplications and artificial faults. Basically, each unique combination could be counted as a new family. Moreover, we believe that some families, e.g. CFB are already CutMix'ed versions of some 'easier' families.\n\nThere is a problem though. Generating one additional sample takes approximately 0.4s (even with CUDA acceleration). That's why the proposed approach is only useful for creating additional train data.\n\nOf course, there are ways to apply some augmentations 'online' also. The most trivial augmentation is the flip augmentation that was used in most of the public notebooks. Additionally, if `p(r,t)` is described by `v(r)`, then `p(r,a⋅t)` is described by `a⋅v(r)`. Physically, this means that if we multiply the vel-image by `a`, the time axis of the corresponding seis-image will be compressed by `a`, or, in other words, if the velocity is higher, the propagation time is lower. Strictly speaking, this reasoning is only valid for a delta-impulse signal source. However, even when exposed to a long-duration source (as in this competition), we were able to slightly improve the model. This is effectively a way to implement runtime augmentations. We haven't used it in our final pipeline, but we still consider it as an interesting finding. \n\nAlso, 5 days before the competition end deadline we found out, that SA looks like perlin noise. So we created blob-like structures using Perlin noise, applying Gaussian smoothing, then adding linear gradients and normalizing the values to the required velocity range (1500-4500 m/s):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Faef297c365cf7f91d20c91511209fda9%2Fphoto_2025-06-26_05-14-30.jpg?generation=1751320014272979&alt=media)\n\nBasically, we could have generated as much additional SA data as we wanted to improve the validation score on that specific family. Additionally, we found how to improve SA even further (see the post-processing section).\n\n## 3. Model architecture and training.\n\n1\\. **Model architecture**: All the models we used can be considered as `caformer/convformer + unet_{pubic/modified}` in structure. Before teaming up and implementing data augmentation, we all observed that increasing the resolution of feature maps could significantly improve CV/LB performance. We found that Convformer showed better performance than Caformer in both CV and LB (-1.0). Our final 6 models include: 3 Convformers with 144 resolution, Caformers with 144, 160, and 256 resolutions. Among these, the  144-resolution performed the best.\n\nThe following is the structure diagram of the model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2F0a3bdd091836b3d5eb69f004fe36bea7%2Fgwi_model.png?generation=1751343067950054&alt=media)\n\n\n2\\. **Speed up training**: Directly training with a larger resolution is a option—for example, caformer256 follows this approach. However, a more effective strategy is to first train with a smaller resolution and then scale up. For instance, we initially trained convformer72 and later fine-tuned convformer144 by resuming the weights from convformer72 before we figure out how to make data augs.\n3\\. **Data Generation**: We generated a dataset that is 15 times the size of the official dataset, which can be roughly divided into 8 versions, utilizing all the data augmentation mentioned above. Among these, the more challenging classes have higher weights, with CFB, CVB, FFB, SA, and SB having significantly higher weights than the remaining classes. In the last stages of training, CFB even accounted for more than half of the training set. It is important to note that no matter how is synthetic dataset, it must be combined with the competition dataset for training; otherwise, the performance on the simpler classes will significantly deteriorate.\n4\\. **Resume training**: We did not attempt to train the model with the entire 15x-sized dataset at once. The first model was trained sequentially on the first 7 versions of the synthetic dataset (e.g., augv1 → augv2 → ... → augv7), ultimately achieving a CV-score of 10.26.The remaining models were trained simultaneously on multiple synthetic datasets in stages, (such as:augv1v2 → augv3v4v5 → augv6v7 → augv8).\n5\\. **Lower the LR gradually**. The LR strategy for all of the models are close. Before the models haven't break a 13-CV score, we use a LR of 1e-4. After that, if we observed the CV score decreases slowly, we switch to a new version of dataset (e.g. augv6->v7) and lower the LR (1e-4 -> 5e-5 -> 2e-5 -> 1e-5 -> 5e-6 -> 1e-6).\n\nThe following table contains information about our models in the final ensemble.\n\n| model | backbone | decoder | CV-Score|\n| --- | --- |--- |--- |\n| augv7_conv144 | Convformer | modified Unet|10.26 |\n| augv8_conv144 | Convformer | modified Unet|10.28 |\n| augv567_conv144 | Convformer |modified Unet|10.53 |\n| augv8_ca144 | Caformer |modified Unet|11.3 |\n| augv8_ca256 | Caformer |public Unet|11.9 |\n| augv8_ca160 | Caformer |public Unet|12.0 |\n\n## 4. Post-processing.\n\n4.1. **Ensemble**: We used the hill‐climb method to ensemble our models. Hill‐climb method is an optimization method that, at each step, selects the local change yielding the largest immediate performance gain and repeats until no significant improvement remains. Compared with mean and median methods, hill-climb improves cv by about 0.3-0.4. However, the median method of CFB is 0.4 higher than hill-climb.\n\n4.2. @alehandreus found a way to modify ours vel-to-seis forward simulator in PyTorch as a differentiable function. We used it to refine the predictions of our ensemble:\n\n1. Given `seis_true`, the ensemble outputs `vel_pred`;\n2. Differentiable forward simulation on `vel_pred` transforms it into `seis_pred`;\n3. `vel_pred` pixels are optimized with gradient descent to minimize the difference between `seis_pred` and `seis_true`.\n\n#### Style A/B families\nOn these families the method works without any modifications. Still, there were a couple of modifications to improve the result:\n1. Set individual LR for each pixel proportional to `vel_pred` values. Larger values usually have more MAE and thus receive larger LR.\n2. `seis_pred` and `seis_true` have 1000 time steps. To prevent vanishing gradients at time steps close to 1000, we multipled seis by `linspace` from 0.1 to 10.\n\n#### \"Discrete\" families\n Firstly, most of the mistakes in our model predictions come from the boundaries/edges between layers:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2Fe4aefabecf90a39c9f32859a6ec92cd9%2Fphoto_2025-06-22_01-05-45%20(2).jpg?generation=1751345175204562&alt=media)\n\nBecause of the strict layer structure, these samples required more sophisticated modifications:\n\n1. Apply Total Variation loss on `vel_pred`. A crucial thing we discovered yesterday is to use p < 1 for the p-norm to encourage sharp edges.\n2. We replaced the value of the LR with the maximum value of it and its neighbours. Thus pixels one unit away from the edges still recieved larger LR.\n\n#### FlatVel_B\nHere we used a completely different approach. Predicted stripes are nearly perfect and require only slight color tuning. We replaced each stripe with one color and optimized only this one color.\n\n#### Results\n\n| Family    | Estimated improvement | Required iterations |\n| -------- | ------- | ------- |\n| CurveFault_A  | -    | -    |\n| CurveFault_B | 1.5     | 250  |\n| CurveVel_A    | 0.4    | 300 |\n| CurveVel_B    | 2.0    | 600 |\n| FlatFault_A    | -   | - |\n| FlatFault_B    | 1.0   | 300 |\n| FlatVel_A    |  - | - |\n| FlatVel_B    | 0.95 | 80 |\n| Style_A    | 14 | 5000 |\n| Style_B    | 6 | 5000 |\n\n#### Classifier\n\nTo split the test set into families we trained a model with the classifier head. The following are the results on the validation data subset:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13977249%2F8ce3cd8644bdd98f1986b9ec51ff2a26%2Fphoto_2025-06-15_15-31-19.jpg?generation=1751344089857990&alt=media)\n\nWe used that classifier to differ Style families from \"Discrete\" families in the test data, since these family types requires different optimization strategies. Moreover, we used it to perform a family-wise hill-climb ensembling.\n\nThanks for reading. Questions welcome.\n\nCode: (in progress)",
    "3237431": "**Addendum**. Huge congratulations to @arsenypoyda on achieving Kaggle Competitions GM rank and to @alehandreus on achieving Kaggle Competitions Master rank! ",
    "3241701": "Big congrats, and thank you for sharing. I'm eager to learn and your write up helped!",
    "3237489": "Good job. It was really close, I think if the competition was a bit longer you would have left me behind. The compute you used is crazy! Can I ask how much $ you estimate your total compute cost? Like, 32\\*4090 is just crasy. 3\\*8A100 also.",
    "3237483": "Congratulations! Your solution is very similar to mine, it's quite possible the difference in our scores is mainly down to how much compute cost and environmental impact we were willing to accept...",
    "3241987": ""
  }
}