{
  "id": 587388,
  "title": "1st Place Solution",
  "url": "/competitions/waveform-inversion/writeups/harshit-sheoran-1st-place-solution",
  "author_name": "",
  "post_date": "2025-07-01T00:22:09.413Z",
  "votes": 194,
  "comment_count": 66,
  "views": 0,
  "content": "<p>First and foremost, a sincere thank you to the competition organizers for this challenge, and congratulations to all my fellow competitors on their impressive work.</p>\n<p>Crazy how nobody moved a single place up or down in top 28.</p>\n<p>Goals List:<br>\n✔️Single digit submissions solo gold<br>\n✔️Solo win a competition #1</p>\n<p>Anyways, here's what I have done:</p>\n<h1><strong>Preprocessing</strong></h1>\n<p>This is not a segmentation task. The velocity model pixels don't line up spatially with the seismic data.<br>\nAn early experiment tested it. I had a baseline that resized the input from the original (5, 1000, 70) to (5, 350, 70) and got a CV of 46.9. , I reshape the input to (1, 350, 350), basically forcing the channels into a spatial layout. CV jumped to 32.5</p>\n<p>This is how the non-resized 1x1000x350 input looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F2c0bdcf39b2d105431911c182de5ec99%2Finput.png?generation=1751317173814874&amp;alt=media\"></p>\n<h1><strong>Architecture</strong></h1>\n<p>Unet didn’t make sense to me, my input was 350x350, the pixel values of the velocity model did not spatially align with the input, so it did not make much sense to me to go back to previous layers to get intermediate layer’s embeddings like a Unet does</p>\n<p>Anyways, it turns out, vision transformers don’t necessarily need a Unet to do their bidding 🙂</p>\n<p>Vision transformers are secretly segmentation-like regression models:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F51344cf7cb6fdeef3087d0bcbe72acd5%2Farch1.png?generation=1751317505425731&amp;alt=media\" alt=\"\"></p>\n<p>This image size was a nice square so it was easy to work with and scale.<br>\nI started using this without a decoder with the best vision transformer I found, EVA02-small model. The performance was really cool, even at this stage it got me a close to 30 score on CV with just 40 epochs of training.</p>\n<p>To scale up the model, I experimented with base and large variants which were scoring better, but what was scoring even more was repeating the architecture as an encoder+decoder setup:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F84e7795e777d616c22ffc264e10dec37%2Farch2.png?generation=1751317556231754&amp;alt=media\" alt=\"\"></p>\n<p>Training this scored CV 28, repeating this (encoder+decoder) with base or large variant of the model did not give improvement and were overfitting</p>\n<p>Training EVA, it was still 20-30% slower than training the original ViT, I found a better backbone after experimenting to find out what was making EVA score much better than ViT, Depth? Channels? Number of attention heads? MLP layer scaling? Gated activations? Turns out, RoPE (Rotary Positional Embeddings) was contributing to the majority of the bottleneck, then I chose better pretrained weights and I got vit_small_patch14_reg4_dinov2.lvd142m as my ideal choice of backbone. On the same training, ViT+RoPE got me under 26 MAE.</p>\n<p>Here is the final architecture:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F8415956230f0715e1c7a910602e211dc%2Farch3.png?generation=1751317568008933&amp;alt=media\" alt=\"\"></p>\n<p>Code: <a href=\"https://www.kaggle.com/code/harshitsheoran/yale-fwi-vit-architecture?scriptVersionId=248201360\" target=\"_blank\">https://www.kaggle.com/code/harshitsheoran/yale-fwi-vit-architecture?scriptVersionId=248201360</a></p>\n<p>My loss function is MAE loss, I apply sigmoid on my logits and scale them in range of (1500, 4500)</p>\n<h1><strong>Training</strong></h1>\n<p>Pretty much every training run used the same recipe:</p>\n<ul>\n<li>LR: 1e-4 with cosine annealing down to ~1e-5</li>\n<li>Optimizer: AdamW</li>\n<li>Loss: MAE</li>\n<li>HFlip and TTA</li>\n<li>EMA</li>\n</ul>\n<p>I used 440k out of 470k for training, 30k for validation, the core of my training at MAE of 26 with 50 epochs is that I am severely underfitting, increasing image size and restarting training can lead to much better scores, for a total of 300 epochs:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2Fe47301b125f7736663b1cc57e63e095b%2Ftrain1.png?generation=1751317762528020&amp;alt=media\" alt=\"\"></p>\n<p>This model was done training on May 28 (more than a month before the deadline), the last month was done exploring the next part of the solution.</p>\n<h1><strong>Generating Data</strong></h1>\n<p>Data Augmentation is a key part of my solution, before an almost perfect replication of the forward modeling function was available, I used forward-modeling + denoiser model setup where denoiser model was the architecture above designed to predict the noise still left in the forward-modeled data</p>\n<p>After failing experiments with generating new velocity models with heuristics or diffusion, I went ahead with using FiveCrop and 5xRandomAffine augmentations {RandomAffine(p=1.0, degrees=15, translate=(0.2, 0), shear=(-15, 15), padding_mode='reflection', resample='nearest')} on existing velocity models to generate more velocity models, this gave me 10x more data of 4.7 million seismic samples, I use this data for pretraining on every step, the training graph looks like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F5fd975fcb6ba850fa56b16a706bf6373%2Ftrain2.png?generation=1751317778801585&amp;alt=media\" alt=\"\"></p>\n<p>Later in the competition, a perfect method for forward modeling was publicly available, at the time, the original notebook was running at a speed of about 10 images per minute which was unusable, using pytorch compile and a 5090, I sped it up to &gt;5000 images per minute per gpu which is a lot more bearable for the next time consuming step.</p>\n<p>I call it “Iterative Pseudo” where after every epoch, I predict on my validation and competition’s test set, about ~95k samples in total, this prediction is a pseudo prediction (for the test set) as we do not know how good or bad it is, I take the predicted velocity models and forward-model them to generate what should be the input for them and use this input in training the next epoch. This step is repeated for every epoch, this does add extra time to each epoch but the payout in terms of score is so worth it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F60f4f0f2f6421676e6d86a7d478b199e%2Ftrain3.png?generation=1751317800532720&amp;alt=media\" alt=\"\"></p>\n<p>As pretraining epochs are 10x larger, all epochs weighted, this was trained for a total of 570 epochs.</p>\n<p>Another observation was made that most of the new score that the model is improving for a while, has been coming from mainly 4 out of 10 methods. So, model was further trained with removing the data that doesn’t come from [‘CurveFault_B’, ‘CurveVel_B’, ‘Style_A’, ‘Style_B’], this removes more than half of the data, speeds up training, frees up the model’s parameters to focus more on this part.</p>\n<p>Then for inference, I would replace the predictions that were from these 4 methods with the predictions from the model below, which yielded a score of 7.5 on CV, 7.0/7.0 on Public/Private LB</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F04ec4c683c1f8d4607cd3dcefb7d1df1%2Ftrain_top4.png?generation=1751317941360967&amp;alt=media\"></p>\n<p>Training this with larger image size at the end, with 896x896 and ensembling that Top4 model to the existing Top4 model yields my best submission at 7.28 CV, 6.9/6.9 on Public/Private LB</p>\n<p>The whole process from start to finish takes about 15 days on 4x 5090.</p>\n<p>My solution does not feature much for ensembling, and can be cleanly trained for longer to generate a hopeful sub 5 leaderboard.</p>\n<h1><strong>That didn’t work / Didn’t get to try</strong></h1>\n<p>MAE/MIM self-supervised pretraining<br>\nStandard Augmentations<br>\nUnet decoder, UperNet, Mask2Former, ViT-Adapter, Fusion heads, more complicated heads than a linear layer</p>\n<p>Due to being short on time, I did not go back and recreate the pretraining data with perfect simulation, that could still improve the score more.</p>\n<p>Idea that I didn’t get to try: Prediction from the model forward modeled to generate simulated seismic input - original seismic input to compute gradient, the original seismic input, this gradient and the predicted velocity model as input to a new model that will refine the prediction according to the gradient</p>\n<p>Thank you for reading!</p>",
  "messages": [
    {
      "id": "3237125",
      "postDate": "07/01/2025 00:02:26",
      "content": "<p>First and foremost, a sincere thank you to the competition organizers for this challenge, and congratulations to all my fellow competitors on their impressive work.</p>\n<p>Crazy how nobody moved a single place up or down in top 28.</p>\n<p>Goals List:<br>\n✔️Single digit submissions solo gold<br>\n✔️Solo win a competition #1</p>\n<p>Anyways, here's what I have done:</p>\n<h1><strong>Preprocessing</strong></h1>\n<p>This is not a segmentation task. The velocity model pixels don't line up spatially with the seismic data.<br>\nAn early experiment tested it. I had a baseline that resized the input from the original (5, 1000, 70) to (5, 350, 70) and got a CV of 46.9. , I reshape the input to (1, 350, 350), basically forcing the channels into a spatial layout. CV jumped to 32.5</p>\n<p>This is how the non-resized 1x1000x350 input looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F2c0bdcf39b2d105431911c182de5ec99%2Finput.png?generation=1751317173814874&amp;alt=media\"></p>\n<h1><strong>Architecture</strong></h1>\n<p>Unet didn’t make sense to me, my input was 350x350, the pixel values of the velocity model did not spatially align with the input, so it did not make much sense to me to go back to previous layers to get intermediate layer’s embeddings like a Unet does</p>\n<p>Anyways, it turns out, vision transformers don’t necessarily need a Unet to do their bidding 🙂</p>\n<p>Vision transformers are secretly segmentation-like regression models:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F51344cf7cb6fdeef3087d0bcbe72acd5%2Farch1.png?generation=1751317505425731&amp;alt=media\" alt=\"\"></p>\n<p>This image size was a nice square so it was easy to work with and scale.<br>\nI started using this without a decoder with the best vision transformer I found, EVA02-small model. The performance was really cool, even at this stage it got me a close to 30 score on CV with just 40 epochs of training.</p>\n<p>To scale up the model, I experimented with base and large variants which were scoring better, but what was scoring even more was repeating the architecture as an encoder+decoder setup:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F84e7795e777d616c22ffc264e10dec37%2Farch2.png?generation=1751317556231754&amp;alt=media\" alt=\"\"></p>\n<p>Training this scored CV 28, repeating this (encoder+decoder) with base or large variant of the model did not give improvement and were overfitting</p>\n<p>Training EVA, it was still 20-30% slower than training the original ViT, I found a better backbone after experimenting to find out what was making EVA score much better than ViT, Depth? Channels? Number of attention heads? MLP layer scaling? Gated activations? Turns out, RoPE (Rotary Positional Embeddings) was contributing to the majority of the bottleneck, then I chose better pretrained weights and I got vit_small_patch14_reg4_dinov2.lvd142m as my ideal choice of backbone. On the same training, ViT+RoPE got me under 26 MAE.</p>\n<p>Here is the final architecture:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F8415956230f0715e1c7a910602e211dc%2Farch3.png?generation=1751317568008933&amp;alt=media\" alt=\"\"></p>\n<p>Code: <a href=\"https://www.kaggle.com/code/harshitsheoran/yale-fwi-vit-architecture?scriptVersionId=248201360\" target=\"_blank\">https://www.kaggle.com/code/harshitsheoran/yale-fwi-vit-architecture?scriptVersionId=248201360</a></p>\n<p>My loss function is MAE loss, I apply sigmoid on my logits and scale them in range of (1500, 4500)</p>\n<h1><strong>Training</strong></h1>\n<p>Pretty much every training run used the same recipe:</p>\n<ul>\n<li>LR: 1e-4 with cosine annealing down to ~1e-5</li>\n<li>Optimizer: AdamW</li>\n<li>Loss: MAE</li>\n<li>HFlip and TTA</li>\n<li>EMA</li>\n</ul>\n<p>I used 440k out of 470k for training, 30k for validation, the core of my training at MAE of 26 with 50 epochs is that I am severely underfitting, increasing image size and restarting training can lead to much better scores, for a total of 300 epochs:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2Fe47301b125f7736663b1cc57e63e095b%2Ftrain1.png?generation=1751317762528020&amp;alt=media\" alt=\"\"></p>\n<p>This model was done training on May 28 (more than a month before the deadline), the last month was done exploring the next part of the solution.</p>\n<h1><strong>Generating Data</strong></h1>\n<p>Data Augmentation is a key part of my solution, before an almost perfect replication of the forward modeling function was available, I used forward-modeling + denoiser model setup where denoiser model was the architecture above designed to predict the noise still left in the forward-modeled data</p>\n<p>After failing experiments with generating new velocity models with heuristics or diffusion, I went ahead with using FiveCrop and 5xRandomAffine augmentations {RandomAffine(p=1.0, degrees=15, translate=(0.2, 0), shear=(-15, 15), padding_mode='reflection', resample='nearest')} on existing velocity models to generate more velocity models, this gave me 10x more data of 4.7 million seismic samples, I use this data for pretraining on every step, the training graph looks like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F5fd975fcb6ba850fa56b16a706bf6373%2Ftrain2.png?generation=1751317778801585&amp;alt=media\" alt=\"\"></p>\n<p>Later in the competition, a perfect method for forward modeling was publicly available, at the time, the original notebook was running at a speed of about 10 images per minute which was unusable, using pytorch compile and a 5090, I sped it up to &gt;5000 images per minute per gpu which is a lot more bearable for the next time consuming step.</p>\n<p>I call it “Iterative Pseudo” where after every epoch, I predict on my validation and competition’s test set, about ~95k samples in total, this prediction is a pseudo prediction (for the test set) as we do not know how good or bad it is, I take the predicted velocity models and forward-model them to generate what should be the input for them and use this input in training the next epoch. This step is repeated for every epoch, this does add extra time to each epoch but the payout in terms of score is so worth it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F60f4f0f2f6421676e6d86a7d478b199e%2Ftrain3.png?generation=1751317800532720&amp;alt=media\" alt=\"\"></p>\n<p>As pretraining epochs are 10x larger, all epochs weighted, this was trained for a total of 570 epochs.</p>\n<p>Another observation was made that most of the new score that the model is improving for a while, has been coming from mainly 4 out of 10 methods. So, model was further trained with removing the data that doesn’t come from [‘CurveFault_B’, ‘CurveVel_B’, ‘Style_A’, ‘Style_B’], this removes more than half of the data, speeds up training, frees up the model’s parameters to focus more on this part.</p>\n<p>Then for inference, I would replace the predictions that were from these 4 methods with the predictions from the model below, which yielded a score of 7.5 on CV, 7.0/7.0 on Public/Private LB</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F04ec4c683c1f8d4607cd3dcefb7d1df1%2Ftrain_top4.png?generation=1751317941360967&amp;alt=media\"></p>\n<p>Training this with larger image size at the end, with 896x896 and ensembling that Top4 model to the existing Top4 model yields my best submission at 7.28 CV, 6.9/6.9 on Public/Private LB</p>\n<p>The whole process from start to finish takes about 15 days on 4x 5090.</p>\n<p>My solution does not feature much for ensembling, and can be cleanly trained for longer to generate a hopeful sub 5 leaderboard.</p>\n<h1><strong>That didn’t work / Didn’t get to try</strong></h1>\n<p>MAE/MIM self-supervised pretraining<br>\nStandard Augmentations<br>\nUnet decoder, UperNet, Mask2Former, ViT-Adapter, Fusion heads, more complicated heads than a linear layer</p>\n<p>Due to being short on time, I did not go back and recreate the pretraining data with perfect simulation, that could still improve the score more.</p>\n<p>Idea that I didn’t get to try: Prediction from the model forward modeled to generate simulated seismic input - original seismic input to compute gradient, the original seismic input, this gradient and the predicted velocity model as input to a new model that will refine the prediction according to the gradient</p>\n<p>Thank you for reading!</p>",
      "rawMarkdown": "First and foremost, a sincere thank you to the competition organizers for this challenge, and congratulations to all my fellow competitors on their impressive work.\n\nCrazy how nobody moved a single place up or down in top 28.\n\n\nGoals List:\n✔️Single digit submissions solo gold\n✔️Solo win a competition #1\n\nAnyways, here's what I have done:\n\n# **Preprocessing**\n\nThis is not a segmentation task. The velocity model pixels don't line up spatially with the seismic data.\nAn early experiment tested it. I had a baseline that resized the input from the original (5, 1000, 70) to (5, 350, 70) and got a CV of 46.9. , I reshape the input to (1, 350, 350), basically forcing the channels into a spatial layout. CV jumped to 32.5\n\nThis is how the non-resized 1x1000x350 input looks like:\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F2c0bdcf39b2d105431911c182de5ec99%2Finput.png?generation=1751317173814874&alt=media\" width=\"100%\">\n\n# **Architecture**\n\nUnet didn’t make sense to me, my input was 350x350, the pixel values of the velocity model did not spatially align with the input, so it did not make much sense to me to go back to previous layers to get intermediate layer’s embeddings like a Unet does\n\nAnyways, it turns out, vision transformers don’t necessarily need a Unet to do their bidding 🙂\n\nVision transformers are secretly segmentation-like regression models:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F51344cf7cb6fdeef3087d0bcbe72acd5%2Farch1.png?generation=1751317505425731&alt=media)\n\nThis image size was a nice square so it was easy to work with and scale.\nI started using this without a decoder with the best vision transformer I found, EVA02-small model. The performance was really cool, even at this stage it got me a close to 30 score on CV with just 40 epochs of training.\n\nTo scale up the model, I experimented with base and large variants which were scoring better, but what was scoring even more was repeating the architecture as an encoder+decoder setup:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F84e7795e777d616c22ffc264e10dec37%2Farch2.png?generation=1751317556231754&alt=media)\n\nTraining this scored CV 28, repeating this (encoder+decoder) with base or large variant of the model did not give improvement and were overfitting\n\nTraining EVA, it was still 20-30% slower than training the original ViT, I found a better backbone after experimenting to find out what was making EVA score much better than ViT, Depth? Channels? Number of attention heads? MLP layer scaling? Gated activations? Turns out, RoPE (Rotary Positional Embeddings) was contributing to the majority of the bottleneck, then I chose better pretrained weights and I got vit_small_patch14_reg4_dinov2.lvd142m as my ideal choice of backbone. On the same training, ViT+RoPE got me under 26 MAE.\n\nHere is the final architecture:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F8415956230f0715e1c7a910602e211dc%2Farch3.png?generation=1751317568008933&alt=media)\n\nCode: https://www.kaggle.com/code/harshitsheoran/yale-fwi-vit-architecture?scriptVersionId=248201360\n\nMy loss function is MAE loss, I apply sigmoid on my logits and scale them in range of (1500, 4500)\n\n\n# **Training**\n\nPretty much every training run used the same recipe:\n- LR: 1e-4 with cosine annealing down to ~1e-5\n- Optimizer: AdamW\n- Loss: MAE\n- HFlip and TTA\n- EMA\n\nI used 440k out of 470k for training, 30k for validation, the core of my training at MAE of 26 with 50 epochs is that I am severely underfitting, increasing image size and restarting training can lead to much better scores, for a total of 300 epochs:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2Fe47301b125f7736663b1cc57e63e095b%2Ftrain1.png?generation=1751317762528020&alt=media)\n\nThis model was done training on May 28 (more than a month before the deadline), the last month was done exploring the next part of the solution.\n\n# **Generating Data**\n\nData Augmentation is a key part of my solution, before an almost perfect replication of the forward modeling function was available, I used forward-modeling + denoiser model setup where denoiser model was the architecture above designed to predict the noise still left in the forward-modeled data\n\nAfter failing experiments with generating new velocity models with heuristics or diffusion, I went ahead with using FiveCrop and 5xRandomAffine augmentations {RandomAffine(p=1.0, degrees=15, translate=(0.2, 0), shear=(-15, 15), padding_mode='reflection', resample='nearest')} on existing velocity models to generate more velocity models, this gave me 10x more data of 4.7 million seismic samples, I use this data for pretraining on every step, the training graph looks like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F5fd975fcb6ba850fa56b16a706bf6373%2Ftrain2.png?generation=1751317778801585&alt=media)\n\nLater in the competition, a perfect method for forward modeling was publicly available, at the time, the original notebook was running at a speed of about 10 images per minute which was unusable, using pytorch compile and a 5090, I sped it up to >5000 images per minute per gpu which is a lot more bearable for the next time consuming step.\n\nI call it “Iterative Pseudo” where after every epoch, I predict on my validation and competition’s test set, about ~95k samples in total, this prediction is a pseudo prediction (for the test set) as we do not know how good or bad it is, I take the predicted velocity models and forward-model them to generate what should be the input for them and use this input in training the next epoch. This step is repeated for every epoch, this does add extra time to each epoch but the payout in terms of score is so worth it.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F60f4f0f2f6421676e6d86a7d478b199e%2Ftrain3.png?generation=1751317800532720&alt=media)\n\nAs pretraining epochs are 10x larger, all epochs weighted, this was trained for a total of 570 epochs.\n\nAnother observation was made that most of the new score that the model is improving for a while, has been coming from mainly 4 out of 10 methods. So, model was further trained with removing the data that doesn’t come from [‘CurveFault_B’, ‘CurveVel_B’, ‘Style_A’, ‘Style_B’], this removes more than half of the data, speeds up training, frees up the model’s parameters to focus more on this part.\n\nThen for inference, I would replace the predictions that were from these 4 methods with the predictions from the model below, which yielded a score of 7.5 on CV, 7.0/7.0 on Public/Private LB\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F04ec4c683c1f8d4607cd3dcefb7d1df1%2Ftrain_top4.png?generation=1751317941360967&alt=media\" width=\"50%\">\n\nTraining this with larger image size at the end, with 896x896 and ensembling that Top4 model to the existing Top4 model yields my best submission at 7.28 CV, 6.9/6.9 on Public/Private LB\n\nThe whole process from start to finish takes about 15 days on 4x 5090.\n\nMy solution does not feature much for ensembling, and can be cleanly trained for longer to generate a hopeful sub 5 leaderboard.\n\n# **That didn’t work / Didn’t get to try**\n\nMAE/MIM self-supervised pretraining\nStandard Augmentations\nUnet decoder, UperNet, Mask2Former, ViT-Adapter, Fusion heads, more complicated heads than a linear layer\n\nDue to being short on time, I did not go back and recreate the pretraining data with perfect simulation, that could still improve the score more.\n\nIdea that I didn’t get to try: Prediction from the model forward modeled to generate simulated seismic input - original seismic input to compute gradient, the original seismic input, this gradient and the predicted velocity model as input to a new model that will refine the prediction according to the gradient\n\n\n\nThank you for reading!",
      "votes": null
    },
    {
      "id": "3237146",
      "postDate": "07/01/2025 00:18:04",
      "content": "<p>Huge congrats, <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>! 🙂 Stellar performance and an awesome write-up! Really neat how you explain your reasoning at various steps and the experiments you ran based on your observations 🙏</p>\n<p>I also didn't think unets were the necessary arch here, but in my explorations I tried various archs (perceiver, custom arch, etc) from scratch, but the one take away from this competition for me was just how unbelievably powerful pretrained vision models are (both in terms of the weights and the priors the models themselves embody).</p>\n<p>Your exploration and the insights around vision transformers are on another level! 🙂 What an awesome 1st place and so well deserved!</p>",
      "rawMarkdown": "Huge congrats, @harshitsheoran! 🙂 Stellar performance and an awesome write-up! Really neat how you explain your reasoning at various steps and the experiments you ran based on your observations 🙏\n\nI also didn't think unets were the necessary arch here, but in my explorations I tried various archs (perceiver, custom arch, etc) from scratch, but the one take away from this competition for me was just how unbelievably powerful pretrained vision models are (both in terms of the weights and the priors the models themselves embody).\n\nYour exploration and the insights around vision transformers are on another level! 🙂 What an awesome 1st place and so well deserved!",
      "votes": null
    },
    {
      "id": "3237163",
      "postDate": "07/01/2025 00:27:39",
      "content": "<p>Someone once said that all top solution are realy simple. This time, it's not! Really amazing job. So many interesting ideas. Now I really want to do with you a competition sometime 🤣</p>",
      "rawMarkdown": "Someone once said that all top solution are realy simple. This time, it's not! Really amazing job. So many interesting ideas. Now I really want to do with you a competition sometime 🤣",
      "votes": null
    },
    {
      "id": "3237165",
      "postDate": "07/01/2025 00:28:50",
      "content": "<p>Thank you, I would be looking forward to it 😊</p>",
      "rawMarkdown": "Thank you, I would be looking forward to it 😊",
      "votes": null
    },
    {
      "id": "3237181",
      "postDate": "07/01/2025 00:42:45",
      "content": "<p>As a beginner and self-taught student of data science, I really love the culture of sharing their knowledge and solutions in Kaggle. Thank you for your sharing! Even tho, I did not participate this competition, but I will study this, following your notebook. Thank you😙</p>",
      "rawMarkdown": "As a beginner and self-taught student of data science, I really love the culture of sharing their knowledge and solutions in Kaggle. Thank you for your sharing! Even tho, I did not participate this competition, but I will study this, following your notebook. Thank you😙",
      "votes": null
    },
    {
      "id": "3237184",
      "postDate": "07/01/2025 00:44:21",
      "content": "<p>Congratulations!  <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> Excellent work. You have a very good baseline model, and then you keep iterating on it. <br>\nI think the essence is to constantly overfit the data, because they all come from the same simulation model.<br>\nAnyway， congratulations again！</p>",
      "rawMarkdown": "Congratulations!  @harshitsheoran Excellent work. You have a very good baseline model, and then you keep iterating on it. \nI think the essence is to constantly overfit the data, because they all come from the same simulation model.\nAnyway， congratulations again！",
      "votes": null
    },
    {
      "id": "3237190",
      "postDate": "07/01/2025 00:48:58",
      "content": "<p>Congratulations! If optimized to the limit, is there a chance to break 6?</p>",
      "rawMarkdown": "Congratulations! If optimized to the limit, is there a chance to break 6?",
      "votes": null
    },
    {
      "id": "3237194",
      "postDate": "07/01/2025 00:51:49",
      "content": "<p>If optimized to the limit, there is a chance to break 5…? Maybe… 6, most likely</p>",
      "rawMarkdown": "If optimized to the limit, there is a chance to break 5...? Maybe... 6, most likely",
      "votes": null
    },
    {
      "id": "3237198",
      "postDate": "07/01/2025 00:54:14",
      "content": "<p>Amazing jobs!Thank you for sharing!</p>",
      "rawMarkdown": "Amazing jobs!Thank you for sharing!",
      "votes": null
    },
    {
      "id": "3237207",
      "postDate": "07/01/2025 00:59:15",
      "content": "<p>Amazing result😧</p>",
      "rawMarkdown": "Amazing result😧",
      "votes": null
    },
    {
      "id": "3237238",
      "postDate": "07/01/2025 01:40:21",
      "content": "<p>Hm can you check something? The StyleB samples are very hard to predict - there is a solution where you don't predict local high-frequent features but still match the given seismograms almost perfectly. This is worse than just numerical noise - it's actually a local minimum. Without solving this StyleB alone contributes about 4.5 to the LB score for me.</p>\n<p>I would expect your performance on StyleB has stagnated long before the others. (I can share a list of which test samples are StyleB later if needed.)</p>",
      "rawMarkdown": "Hm can you check something? The StyleB samples are very hard to predict - there is a solution where you don't predict local high-frequent features but still match the given seismograms almost perfectly. This is worse than just numerical noise - it's actually a local minimum. Without solving this StyleB alone contributes about 4.5 to the LB score for me.\n\nI would expect your performance on StyleB has stagnated long before the others. (I can share a list of which test samples are StyleB later if needed.)",
      "votes": null
    },
    {
      "id": "3237240",
      "postDate": "07/01/2025 01:45:39",
      "content": "<p>Congratulations! I'll be sharing my solution around the weekend, but in short I used classical full waveform inversion, using <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>'s solution as a starting point. Key point to get this working was to have strong priors.</p>\n<p>Essentially I think your solution is not so very different; in your case the priors are basically learned by the network. Both our solutions require significant compute per test sample (unlike some of the other solutions that are expensive to train but cheap to infer), and I think our compute costs are surprisingly similar.</p>\n<p>More details later this week!</p>",
      "rawMarkdown": "Congratulations! I'll be sharing my solution around the weekend, but in short I used classical full waveform inversion, using @brendanartley's solution as a starting point. Key point to get this working was to have strong priors.\n\nEssentially I think your solution is not so very different; in your case the priors are basically learned by the network. Both our solutions require significant compute per test sample (unlike some of the other solutions that are expensive to train but cheap to infer), and I think our compute costs are surprisingly similar.\n\nMore details later this week!",
      "votes": null
    },
    {
      "id": "3237257",
      "postDate": "07/01/2025 02:17:28",
      "content": "<p>Congrats! Well deserved. </p>",
      "rawMarkdown": "Congrats! Well deserved.",
      "votes": null
    },
    {
      "id": "3237269",
      "postDate": "07/01/2025 02:29:34",
      "content": "<p>Congratulations!</p>\n<p>I just wonder how do you sense a model will overfit or not? Did you sense a model will overfit by seeing loss graph for ten epochs ?</p>\n<p>Thanks for reply.</p>",
      "rawMarkdown": "Congratulations!\n\nI just wonder how do you sense a model will overfit or not? Did you sense a model will overfit by seeing loss graph for ten epochs ?\n\nThanks for reply.",
      "votes": null
    },
    {
      "id": "3237409",
      "postDate": "07/01/2025 04:49:56",
      "content": "<p>Congratulations Harshit, great work and great result !</p>",
      "rawMarkdown": "Congratulations Harshit, great work and great result !",
      "votes": null
    },
    {
      "id": "3237432",
      "postDate": "07/01/2025 05:09:23",
      "content": "<p>can someone in simple language explain how much time in hrs did it take for training and on what resources. Is kaggle weekly 30 hr gpu sufficient. This was my first competition.</p>",
      "rawMarkdown": "can someone in simple language explain how much time in hrs did it take for training and on what resources. Is kaggle weekly 30 hr gpu sufficient. This was my first competition.",
      "votes": null
    },
    {
      "id": "3237459",
      "postDate": "07/01/2025 05:37:55",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> what was your score only for StyleB?</p>",
      "rawMarkdown": "jeroencottaar what was your score only for StyleB?",
      "votes": null
    },
    {
      "id": "3237470",
      "postDate": "07/01/2025 05:49:44",
      "content": "<p>My average score for the StyleB maps in the public test set is ~44, contributing ~4.3 to my score. CV score is similar.</p>",
      "rawMarkdown": "My average score for the StyleB maps in the public test set is ~44, contributing ~4.3 to my score. CV score is similar.",
      "votes": null
    },
    {
      "id": "3237486",
      "postDate": "07/01/2025 06:07:32",
      "content": "<p>Congrats!!! Thanks so much for your detailed explanation, but I'm wonder do you have more detail code snippet? Like how you do the reshape how you do the pseudo iteration. It's really helpful and I gained a lot from your post. </p>",
      "rawMarkdown": "Congrats!!! Thanks so much for your detailed explanation, but I'm wonder do you have more detail code snippet? Like how you do the reshape how you do the pseudo iteration. It's really helpful and I gained a lot from your post.",
      "votes": null
    },
    {
      "id": "3237557",
      "postDate": "07/01/2025 07:01:12",
      "content": "<p>The second last section mentioned:</p>\n<blockquote>\n  <p>The whole process from start to finish takes about 15 days on 4x 5090.</p>\n</blockquote>",
      "rawMarkdown": "The second last section mentioned:\n>The whole process from start to finish takes about 15 days on 4x 5090.",
      "votes": null
    },
    {
      "id": "3237569",
      "postDate": "07/01/2025 07:07:34",
      "content": "<p>Hah. So only for this family, I have MUCH better score.<br>\nValidation is 22 and I can get it lower with better training.  <br>\nWhich mean that for other family your method is so much more powerful than mine.</p>",
      "rawMarkdown": "Hah. So only for this family, I have MUCH better score.\nValidation is 22 and I can get it lower with better training.  \nWhich mean that for other family your method is so much more powerful than mine.",
      "votes": null
    },
    {
      "id": "3237572",
      "postDate": "07/01/2025 07:09:50",
      "content": "<p>Looking forward to your writeup. I started out by caching Bartley's ensemble models output, and ran scalar wave FWI to refine only CurveFault_B. (In my CV, FWI mostly helps with <code>_B</code> family but not <code>_A</code>. I don't have enough compute so I focused on the most difficult family.) It ended up performing worse. So I added a loss term to not let the physical inversion deviate too much from the ensemble output (L1 loss between the two velocity maps in Deepwave library). Still, ended up being 0.1 worse than the public model's starting point. Curious to learn from you experience!</p>",
      "rawMarkdown": "Looking forward to your writeup. I started out by caching Bartley's ensemble models output, and ran scalar wave FWI to refine only CurveFault_B. (In my CV, FWI mostly helps with `_B` family but not `_A`. I don't have enough compute so I focused on the most difficult family.) It ended up performing worse. So I added a loss term to not let the physical inversion deviate too much from the ensemble output (L1 loss between the two velocity maps in Deepwave library). Still, ended up being 0.1 worse than the public model's starting point. Curious to learn from you experience!",
      "votes": null
    },
    {
      "id": "3237583",
      "postDate": "07/01/2025 07:19:26",
      "content": "<p>So 15 continuous days of training, how does one do it on kaggle resources, or is it done on private resources. Does it mean from competitive pov, one has little chance of breaking into top ranking with just depending upon kaggle resources?</p>",
      "rawMarkdown": "So 15 continuous days of training, how does one do it on kaggle resources, or is it done on private resources. Does it mean from competitive pov, one has little chance of breaking into top ranking with just depending upon kaggle resources?",
      "votes": null
    },
    {
      "id": "3237604",
      "postDate": "07/01/2025 07:29:32",
      "content": "<p>You have to do it in the cloud or using your own GPUs. For this competition in particular I don't think a top ranking using only Kaggle resources was possible. Luckily this is not the case for all competitions.</p>",
      "rawMarkdown": "You have to do it in the cloud or using your own GPUs. For this competition in particular I don't think a top ranking using only Kaggle resources was possible. Luckily this is not the case for all competitions.",
      "votes": null
    },
    {
      "id": "3237606",
      "postDate": "07/01/2025 07:31:07",
      "content": "<p>I think this is one where the deep learning approaches may be able to learn a better prior than my classical ones. In particular, they may know where to expect the high-frequent features that I have no hope of finding. </p>\n<p>It also means that sub-5 may be very well possible, and may not need more than selecting the right solution per family.</p>",
      "rawMarkdown": "I think this is one where the deep learning approaches may be able to learn a better prior than my classical ones. In particular, they may know where to expect the high-frequent features that I have no hope of finding. \n\nIt also means that sub-5 may be very well possible, and may not need more than selecting the right solution per family.",
      "votes": null
    },
    {
      "id": "3237624",
      "postDate": "07/01/2025 07:43:00",
      "content": "<blockquote>\n  <p>It also means that sub-5 may be very well possible, and may not need more than selecting the right solution per family.  </p>\n</blockquote>\n<p>Which is very easy, very simple for classifier to have 100% accuracy for classify between style and the rest- I did it.  </p>",
      "rawMarkdown": ">It also means that sub-5 may be very well possible, and may not need more than selecting the right solution per family.  \n\nWhich is very easy, very simple for classifier to have 100% accuracy for classify between style and the rest- I did it.",
      "votes": null
    },
    {
      "id": "3237636",
      "postDate": "07/01/2025 07:50:44",
      "content": "<p>Here's everything I know about my public test set performance. This is not CV, but based on probing, though it mostly tracks CV. See any more opportunities? Also <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> and other top competitors.</p>\n<table>\n<thead>\n<tr>\n<th>Set</th>\n<th>Percentage of public test set</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FlatVelA+B</td>\n<td>17.764</td>\n<td>0.0</td>\n</tr>\n<tr>\n<td>Fault maps with only 2 velocities</td>\n<td>6.92</td>\n<td>8.7</td>\n</tr>\n<tr>\n<td>StyleA</td>\n<td>8.892</td>\n<td>3.4</td>\n</tr>\n<tr>\n<td>StyleB</td>\n<td>9.57</td>\n<td>44.4</td>\n</tr>\n<tr>\n<td>All others</td>\n<td>56.85</td>\n<td>4.5</td>\n</tr>\n</tbody>\n</table>\n<p>Pretty sure I messed something up in my model for the 2-velocity maps, they should be easy…</p>",
      "rawMarkdown": "Here's everything I know about my public test set performance. This is not CV, but based on probing, though it mostly tracks CV. See any more opportunities? Also @harshitsheoran and other top competitors.\n\n\n| Set                                   | Percentage of public test set | Score    |\n| ------------------------------------- | ------------------------------ | -------- |\n| FlatVelA+B                            | 17.764                        | 0.0        |\n| Fault maps with only 2 velocities | 6.92                         | 8.7  |\n| StyleA                                | 8.892                        | 3.4 |\n| StyleB                                | 9.57                         | 44.4 |\n| All others                            | 56.85 | 4.5 |\n\n\nPretty sure I messed something up in my model for the 2-velocity maps, they should be easy...",
      "votes": null
    },
    {
      "id": "3237646",
      "postDate": "07/01/2025 07:59:46",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> let's try it!  <br>\n<a href=\"https://www.kaggle.com/code/shlomoron/gwi-final-sub-with-labels/\" target=\"_blank\">Here is</a> a notebook with my final ensemble and labels.  <br>\nTry to replace the predictions in test_ensemble with yours, except for the indices where test_kind_labels=9.  </p>",
      "rawMarkdown": "jeroencottaar let's try it!  \n[Here is](https://www.kaggle.com/code/shlomoron/gwi-final-sub-with-labels/) a notebook with my final ensemble and labels.  \nTry to replace the predictions in test_ensemble with yours, except for the indices where test_kind_labels=9.",
      "votes": null
    },
    {
      "id": "3237721",
      "postDate": "07/01/2025 09:05:20",
      "content": "<p>congratulation well done</p>",
      "rawMarkdown": "congratulation well done",
      "votes": null
    },
    {
      "id": "3237724",
      "postDate": "07/01/2025 09:10:25",
      "content": "<p>Cool, will pick this up on Thursday (work deadlines first…)</p>",
      "rawMarkdown": "Cool, will pick this up on Thursday (work deadlines first...)",
      "votes": null
    },
    {
      "id": "3237754",
      "postDate": "07/01/2025 09:30:27",
      "content": "<p>Congrats and thanks for sharing the results here. I've totally learned a lot. Its one of my freshest experiences while trying to dive deep in the Data science course.</p>",
      "rawMarkdown": "Congrats and thanks for sharing the results here. I've totally learned a lot. Its one of my freshest experiences while trying to dive deep in the Data science course.",
      "votes": null
    },
    {
      "id": "3237904",
      "postDate": "07/01/2025 11:18:41",
      "content": "<p>Congrats on the win and the lead throughout the competition!</p>\n<blockquote>\n  <p>I apply sigmoid on my logits and scale them in range of (1500, 4500)</p>\n</blockquote>\n<p>I had this on my todo list. Did you compare without sigmoid?</p>",
      "rawMarkdown": "Congrats on the win and the lead throughout the competition!\n\n>  I apply sigmoid on my logits and scale them in range of (1500, 4500)\n\nI had this on my todo list. Did you compare without sigmoid?",
      "votes": null
    },
    {
      "id": "3237914",
      "postDate": "07/01/2025 11:29:08",
      "content": "<p>Thank you, I have been doing this from day 1 actually, without sigmoid was worse but my testing was primilinary at the time.</p>",
      "rawMarkdown": "Thank you, I have been doing this from day 1 actually, without sigmoid was worse but my testing was primilinary at the time.",
      "votes": null
    },
    {
      "id": "3237925",
      "postDate": "07/01/2025 11:53:41",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "3237929",
      "postDate": "07/01/2025 12:03:34",
      "content": "<p>Not sure I got it right : (700,700) is more pixels than (5,1000,70), I guess you interpolated to increase the size ?</p>\n<p>So did I understand correctly that you did (5,1000,70) -&gt; (350,350) -&gt; (476,476) -&gt; (588,588) -&gt; (700,700) ?</p>\n<p>If so, why not (5,1000,70) -&gt; (588,588) -&gt; (700,700), or even (5,1000,70) -&gt; (700,700) directly ?</p>",
      "rawMarkdown": "Not sure I got it right : (700,700) is more pixels than (5,1000,70), I guess you interpolated to increase the size ?\n\nSo did I understand correctly that you did (5,1000,70) -> (350,350) -> (476,476) -> (588,588) -> (700,700) ?\n\nIf so, why not (5,1000,70) -> (588,588) -> (700,700), or even (5,1000,70) -> (700,700) directly ?",
      "votes": null
    },
    {
      "id": "3237934",
      "postDate": "07/01/2025 12:09:34",
      "content": "<p>It is much more time efficient and stable to gradually increase the image size this way</p>",
      "rawMarkdown": "It is much more time efficient and stable to gradually increase the image size this way",
      "votes": null
    },
    {
      "id": "3237941",
      "postDate": "07/01/2025 12:14:31",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>! Very interesting ideas overall, you deserved it</p>",
      "rawMarkdown": "Congrats @harshitsheoran! Very interesting ideas overall, you deserved it",
      "votes": null
    },
    {
      "id": "3237945",
      "postDate": "07/01/2025 12:16:45",
      "content": "<p>Congrats!!  </p>",
      "rawMarkdown": "Congrats!!",
      "votes": null
    },
    {
      "id": "3238292",
      "postDate": "07/01/2025 17:54:57",
      "content": "<p>Yes. This is really cool. The iterative psuedo step is pretty wild, also the 4/10 types is a really nice find I wish I'd seen! awesome work.</p>",
      "rawMarkdown": "Yes. This is really cool. The iterative psuedo step is pretty wild, also the 4/10 types is a really nice find I wish I'd seen! awesome work.",
      "votes": null
    },
    {
      "id": "3238383",
      "postDate": "07/01/2025 19:41:01",
      "content": "<p>there is some mistakes, but good</p>",
      "rawMarkdown": "there is some mistakes, but good",
      "votes": null
    },
    {
      "id": "3238402",
      "postDate": "07/01/2025 19:59:41",
      "content": "<p>Yes, your CV do look off for some classes, FlatVel is perfect, StyleA is too good, \"All others\" have curves and faults, so 4.5 average for them is good</p>\n<p>CV For single model:<br>\nMAE 7.940401554107666<br>\nCurveFault_A: 1.8529449701309204<br>\nCurveFault_B: 12.928738594055176<br>\nCurveVel_A: 1.51606023311615<br>\nCurveVel_B: 5.029183864593506<br>\nFlatFault_A: 0.9475402235984802<br>\nFlatFault_B: 2.63285231590271<br>\nFlatVel_A: 0.3224563002586365<br>\nFlatVel_B: 0.5822372436523438<br>\nStyle_A: 10.023091316223145<br>\nStyle_B: 25.63881492614746</p>\n<p>For Top4 methods are slightly better in my best submission file as focused model<br>\nYou can find the submission here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/yale-fwi-best-submission-csv-file\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/yale-fwi-best-submission-csv-file</a></p>",
      "rawMarkdown": "Yes, your CV do look off for some classes, FlatVel is perfect, StyleA is too good, \"All others\" have curves and faults, so 4.5 average for them is good\n\nCV For single model:\nMAE 7.940401554107666\nCurveFault_A: 1.8529449701309204\nCurveFault_B: 12.928738594055176\nCurveVel_A: 1.51606023311615\nCurveVel_B: 5.029183864593506\nFlatFault_A: 0.9475402235984802\nFlatFault_B: 2.63285231590271\nFlatVel_A: 0.3224563002586365\nFlatVel_B: 0.5822372436523438\nStyle_A: 10.023091316223145\nStyle_B: 25.63881492614746\n\n\nFor Top4 methods are slightly better in my best submission file as focused model\nYou can find the submission here:\n\nhttps://www.kaggle.com/datasets/harshitsheoran/yale-fwi-best-submission-csv-file",
      "votes": null
    },
    {
      "id": "3238435",
      "postDate": "07/01/2025 20:24:17",
      "content": "<p>Yeah I wans't able to make a good prior for StyleB, so that one hurt me bad. I think something weird also happened to some of the CurveFault_A and FaltFault_A datasets; either I messed up my probing or I messed up the model there.</p>\n<p>Anyway, pretty sure there's a sub-5 in ensembling our combined submissions. I'll only be picking this up in a few days, but here's my best in case anyone else wants to have a go: <a href=\"https://www.kaggle.com/datasets/jeroencottaar/best-yale-submission\" target=\"_blank\">https://www.kaggle.com/datasets/jeroencottaar/best-yale-submission</a></p>",
      "rawMarkdown": "Yeah I wans't able to make a good prior for StyleB, so that one hurt me bad. I think something weird also happened to some of the CurveFault_A and FaltFault_A datasets; either I messed up my probing or I messed up the model there.\n\nAnyway, pretty sure there's a sub-5 in ensembling our combined submissions. I'll only be picking this up in a few days, but here's my best in case anyone else wants to have a go: https://www.kaggle.com/datasets/jeroencottaar/best-yale-submission",
      "votes": null
    },
    {
      "id": "3238465",
      "postDate": "07/01/2025 20:55:47",
      "content": "<p>The average best solution (from csv file) is 6.7<br>\n<a href=\"https://www.kaggle.com/code/overvalueawareness/ensemble-best-solution-seismic25?scriptVersionId=248381868\" target=\"_blank\">https://www.kaggle.com/code/overvalueawareness/ensemble-best-solution-seismic25?scriptVersionId=248381868</a></p>",
      "rawMarkdown": "The average best solution (from csv file) is 6.7\nhttps://www.kaggle.com/code/overvalueawareness/ensemble-best-solution-seismic25?scriptVersionId=248381868",
      "votes": null
    },
    {
      "id": "3238477",
      "postDate": "07/01/2025 21:25:57",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> I tried replacing your preds with mine for StyleB, did not work. Maybe I have a bug, or my classifier is much worse than I thought. Or your probing maybe wrong.  </p>",
      "rawMarkdown": "jeroencottaar I tried replacing your preds with mine for StyleB, did not work. Maybe I have a bug, or my classifier is much worse than I thought. Or your probing maybe wrong.",
      "votes": null
    },
    {
      "id": "3238484",
      "postDate": "07/01/2025 21:41:43",
      "content": "<p>Same, replacing StyleB with mine gives 7.47 private, which is only a bit better</p>",
      "rawMarkdown": "Same, replacing StyleB with mine gives 7.47 private, which is only a bit better",
      "votes": null
    },
    {
      "id": "3238504",
      "postDate": "07/01/2025 22:22:21",
      "content": "<p>Almost 6.5 can be achieved with replacing your Style_B, and all of the Faults families with mine</p>",
      "rawMarkdown": "Almost 6.5 can be achieved with replacing your Style_B, and all of the Faults families with mine",
      "votes": null
    },
    {
      "id": "3238535",
      "postDate": "07/01/2025 23:53:49",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    },
    {
      "id": "3238546",
      "postDate": "07/02/2025 00:23:33",
      "content": "<p>Congratulations and thanks for sharing your solution!</p>\n<p>Do you have any validation plots of the predicted and actual velocity/geology. Since I'm a geologist, I'm just curious what the results look like.Thanks!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your solution!\n\nDo you have any validation plots of the predicted and actual velocity/geology. Since I'm a geologist, I'm just curious what the results look like.Thanks!",
      "votes": null
    },
    {
      "id": "3238678",
      "postDate": "07/02/2025 05:27:05",
      "content": "<p>Aww no easy sub-5 just yet.</p>\n<p>Another possibility is that both of your StyleB performances are not as good on LB as CV. I've seen some clues that StyleB in LB may be a bit different from training.</p>\n<p>Anyway, I'll share my probing method later so that you can audit it, and so that we can check it on your submissions.</p>",
      "rawMarkdown": "Aww no easy sub-5 just yet.\n\nAnother possibility is that both of your StyleB performances are not as good on LB as CV. I've seen some clues that StyleB in LB may be a bit different from training.\n\nAnyway, I'll share my probing method later so that you can audit it, and so that we can check it on your submissions.",
      "votes": null
    },
    {
      "id": "3239020",
      "postDate": "07/02/2025 11:53:45",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "3239494",
      "postDate": "07/02/2025 21:27:25",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> Hey would it be possible for you to tell what GPUs you used, as a student i was only able to use kaggle's T4 [recently got intern, now have little money]. In upcoming competition that required high vram and time for training.<br>\nI would go for rented compute, if you can tell some info about what GPUs you used, rented or you own them and from where you think it's better to rent(you think they provide a friendly env for training).<br>\nThanks.</p>",
      "rawMarkdown": "jeroencottaar Hey would it be possible for you to tell what GPUs you used, as a student i was only able to use kaggle's T4 [recently got intern, now have little money]. In upcoming competition that required high vram and time for training.\nI would go for rented compute, if you can tell some info about what GPUs you used, rented or you own them and from where you think it's better to rent(you think they provide a friendly env for training).\nThanks.",
      "votes": null
    },
    {
      "id": "3239663",
      "postDate": "07/03/2025 03:14:35",
      "content": "<p>Congratulations! Impressive job, learnt a lot!</p>",
      "rawMarkdown": "Congratulations! Impressive job, learnt a lot!",
      "votes": null
    },
    {
      "id": "3239682",
      "postDate": "07/03/2025 03:45:25",
      "content": "<p>This is absolutely amazing, and I loved your thought process to arrive at such insights. As a newcomer to this field, however, I feel like this is where I have to stop. The computational resources needed in most competitions are truly discouraging and insane for me. Wishing you all the best going forward!</p>",
      "rawMarkdown": "This is absolutely amazing, and I loved your thought process to arrive at such insights. As a newcomer to this field, however, I feel like this is where I have to stop. The computational resources needed in most competitions are truly discouraging and insane for me. Wishing you all the best going forward!",
      "votes": null
    },
    {
      "id": "3239721",
      "postDate": "07/03/2025 04:22:25",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "votes": null
    },
    {
      "id": "3239732",
      "postDate": "07/03/2025 04:37:19",
      "content": "<p>I used vast.ai. I'm not sure if they're the cheapest, but you get full control over your instance (i.e. you just get a Unix environment and can do what you want, including running notebooks for any duration).</p>",
      "rawMarkdown": "I used vast.ai. I'm not sure if they're the cheapest, but you get full control over your instance (i.e. you just get a Unix environment and can do what you want, including running notebooks for any duration).",
      "votes": null
    },
    {
      "id": "3239733",
      "postDate": "07/03/2025 04:38:51",
      "content": "<p>Please don't be discouraged. This competition really is an outlier in terms of necessary compute. For example, I took second place in this one with 0 GPU hours: <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/overview\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/overview</a></p>",
      "rawMarkdown": "Please don't be discouraged. This competition really is an outlier in terms of necessary compute. For example, I took second place in this one with 0 GPU hours: https://www.kaggle.com/competitions/ariel-data-challenge-2024/overview",
      "votes": null
    },
    {
      "id": "3239822",
      "postDate": "07/03/2025 07:26:55",
      "content": "<p>I went through your solution, and it's incredible how you deconstructed and simplified the problem—that's really something extraordinary. Figuring out how you even got to that solution is a whole other challenge?<br>\nAnyway, thank you for the encouragement. I'll try to keep going, though I'm not sure for how long.<br>\nIt's just very exciting for me to see brilliant minds like yours at work here.</p>",
      "rawMarkdown": "I went through your solution, and it's incredible how you deconstructed and simplified the problem—that's really something extraordinary. Figuring out how you even got to that solution is a whole other challenge?\nAnyway, thank you for the encouragement. I'll try to keep going, though I'm not sure for how long.\nIt's just very exciting for me to see brilliant minds like yours at work here.",
      "votes": null
    },
    {
      "id": "3239912",
      "postDate": "07/03/2025 09:18:20",
      "content": "<p>Congrats man! interesting….</p>",
      "rawMarkdown": "Congrats man! interesting....",
      "votes": null
    },
    {
      "id": "3239998",
      "postDate": "07/03/2025 11:19:16",
      "content": "<p>May I know how does iteration improve the MAE.</p>\n<p>PS: I'm just a beginner</p>",
      "rawMarkdown": "May I know how does iteration improve the MAE.\n\nPS: I'm just a beginner",
      "votes": null
    },
    {
      "id": "3240031",
      "postDate": "07/03/2025 11:51:03",
      "content": "<p>Congratulations🎉🎉</p>",
      "rawMarkdown": "Congratulations🎉🎉",
      "votes": null
    },
    {
      "id": "3240177",
      "postDate": "07/03/2025 14:26:30",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>  Just read you solution in ariel competition, Truly smart, hats off to you.</p>",
      "rawMarkdown": "jeroencottaar  Just read you solution in ariel competition, Truly smart, hats off to you.",
      "votes": null
    },
    {
      "id": "3241459",
      "postDate": "07/04/2025 20:11:04",
      "content": "<p>Smart preprocessing insight</p>",
      "rawMarkdown": "Smart preprocessing insight",
      "votes": null
    },
    {
      "id": "3242154",
      "postDate": "07/05/2025 16:27:38",
      "content": "<p>woah super cool! Physics-guided ML models are up and coming.</p>",
      "rawMarkdown": "woah super cool! Physics-guided ML models are up and coming.",
      "votes": null
    },
    {
      "id": "3242462",
      "postDate": "07/05/2025 22:55:15",
      "content": "<p>Congrats! That's amazing!</p>",
      "rawMarkdown": "Congrats! That's amazing!",
      "votes": null
    },
    {
      "id": "3243409",
      "postDate": "07/07/2025 05:34:32",
      "content": "<p>Quite detail ths</p>",
      "rawMarkdown": "Quite detail ths",
      "votes": null
    },
    {
      "id": "3246618",
      "postDate": "07/11/2025 08:31:46",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": null
    },
    {
      "id": "3247375",
      "postDate": "07/12/2025 15:55:38",
      "content": "<p>Congratulations….. </p>",
      "rawMarkdown": "Congratulations.....",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237146,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "07/01/2025 00:18:04",
      "content": "<p>Huge congrats, <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>! 🙂 Stellar performance and an awesome write-up! Really neat how you explain your reasoning at various steps and the experiments you ran based on your observations 🙏</p>\n<p>I also didn't think unets were the necessary arch here, but in my explorations I tried various archs (perceiver, custom arch, etc) from scratch, but the one take away from this competition for me was just how unbelievably powerful pretrained vision models are (both in terms of the weights and the priors the models themselves embody).</p>\n<p>Your exploration and the insights around vision transformers are on another level! 🙂 What an awesome 1st place and so well deserved!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237163,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "07/01/2025 00:27:39",
      "content": "<p>Someone once said that all top solution are realy simple. This time, it's not! Really amazing job. So many interesting ideas. Now I really want to do with you a competition sometime 🤣</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237165,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "07/01/2025 00:28:50",
          "content": "<p>Thank you, I would be looking forward to it 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237181,
      "author_name": "pablolee789",
      "author_url": "",
      "post_date": "07/01/2025 00:42:45",
      "content": "<p>As a beginner and self-taught student of data science, I really love the culture of sharing their knowledge and solutions in Kaggle. Thank you for your sharing! Even tho, I did not participate this competition, but I will study this, following your notebook. Thank you😙</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237184,
      "author_name": "wisley1024",
      "author_url": "",
      "post_date": "07/01/2025 00:44:21",
      "content": "<p>Congratulations!  <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> Excellent work. You have a very good baseline model, and then you keep iterating on it. <br>\nI think the essence is to constantly overfit the data, because they all come from the same simulation model.<br>\nAnyway， congratulations again！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237190,
      "author_name": "jezzlin",
      "author_url": "",
      "post_date": "07/01/2025 00:48:58",
      "content": "<p>Congratulations! If optimized to the limit, is there a chance to break 6?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237194,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "07/01/2025 00:51:49",
          "content": "<p>If optimized to the limit, there is a chance to break 5…? Maybe… 6, most likely</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237207,
              "author_name": "jezzlin",
              "author_url": "",
              "post_date": "07/01/2025 00:59:15",
              "content": "<p>Amazing result😧</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3237238,
              "author_name": "jeroencottaar",
              "author_url": "",
              "post_date": "07/01/2025 01:40:21",
              "content": "<p>Hm can you check something? The StyleB samples are very hard to predict - there is a solution where you don't predict local high-frequent features but still match the given seismograms almost perfectly. This is worse than just numerical noise - it's actually a local minimum. Without solving this StyleB alone contributes about 4.5 to the LB score for me.</p>\n<p>I would expect your performance on StyleB has stagnated long before the others. (I can share a list of which test samples are StyleB later if needed.)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3237459,
                  "author_name": "shlomoron",
                  "author_url": "",
                  "post_date": "07/01/2025 05:37:55",
                  "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> what was your score only for StyleB?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3237470,
                      "author_name": "jeroencottaar",
                      "author_url": "",
                      "post_date": "07/01/2025 05:49:44",
                      "content": "<p>My average score for the StyleB maps in the public test set is ~44, contributing ~4.3 to my score. CV score is similar.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3237569,
                          "author_name": "shlomoron",
                          "author_url": "",
                          "post_date": "07/01/2025 07:07:34",
                          "content": "<p>Hah. So only for this family, I have MUCH better score.<br>\nValidation is 22 and I can get it lower with better training.  <br>\nWhich mean that for other family your method is so much more powerful than mine.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3237606,
                              "author_name": "jeroencottaar",
                              "author_url": "",
                              "post_date": "07/01/2025 07:31:07",
                              "content": "<p>I think this is one where the deep learning approaches may be able to learn a better prior than my classical ones. In particular, they may know where to expect the high-frequent features that I have no hope of finding. </p>\n<p>It also means that sub-5 may be very well possible, and may not need more than selecting the right solution per family.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3237624,
                                  "author_name": "shlomoron",
                                  "author_url": "",
                                  "post_date": "07/01/2025 07:43:00",
                                  "content": "<blockquote>\n  <p>It also means that sub-5 may be very well possible, and may not need more than selecting the right solution per family.  </p>\n</blockquote>\n<p>Which is very easy, very simple for classifier to have 100% accuracy for classify between style and the rest- I did it.  </p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3237636,
                                      "author_name": "jeroencottaar",
                                      "author_url": "",
                                      "post_date": "07/01/2025 07:50:44",
                                      "content": "<p>Here's everything I know about my public test set performance. This is not CV, but based on probing, though it mostly tracks CV. See any more opportunities? Also <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> and other top competitors.</p>\n<table>\n<thead>\n<tr>\n<th>Set</th>\n<th>Percentage of public test set</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FlatVelA+B</td>\n<td>17.764</td>\n<td>0.0</td>\n</tr>\n<tr>\n<td>Fault maps with only 2 velocities</td>\n<td>6.92</td>\n<td>8.7</td>\n</tr>\n<tr>\n<td>StyleA</td>\n<td>8.892</td>\n<td>3.4</td>\n</tr>\n<tr>\n<td>StyleB</td>\n<td>9.57</td>\n<td>44.4</td>\n</tr>\n<tr>\n<td>All others</td>\n<td>56.85</td>\n<td>4.5</td>\n</tr>\n</tbody>\n</table>\n<p>Pretty sure I messed something up in my model for the 2-velocity maps, they should be easy…</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 3237646,
                                          "author_name": "shlomoron",
                                          "author_url": "",
                                          "post_date": "07/01/2025 07:59:46",
                                          "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> let's try it!  <br>\n<a href=\"https://www.kaggle.com/code/shlomoron/gwi-final-sub-with-labels/\" target=\"_blank\">Here is</a> a notebook with my final ensemble and labels.  <br>\nTry to replace the predictions in test_ensemble with yours, except for the indices where test_kind_labels=9.  </p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 3237724,
                                              "author_name": "jeroencottaar",
                                              "author_url": "",
                                              "post_date": "07/01/2025 09:10:25",
                                              "content": "<p>Cool, will pick this up on Thursday (work deadlines first…)</p>",
                                              "votes": null,
                                              "replies": []
                                            }
                                          ]
                                        },
                                        {
                                          "id": 3238402,
                                          "author_name": "harshitsheoran",
                                          "author_url": "",
                                          "post_date": "07/01/2025 19:59:41",
                                          "content": "<p>Yes, your CV do look off for some classes, FlatVel is perfect, StyleA is too good, \"All others\" have curves and faults, so 4.5 average for them is good</p>\n<p>CV For single model:<br>\nMAE 7.940401554107666<br>\nCurveFault_A: 1.8529449701309204<br>\nCurveFault_B: 12.928738594055176<br>\nCurveVel_A: 1.51606023311615<br>\nCurveVel_B: 5.029183864593506<br>\nFlatFault_A: 0.9475402235984802<br>\nFlatFault_B: 2.63285231590271<br>\nFlatVel_A: 0.3224563002586365<br>\nFlatVel_B: 0.5822372436523438<br>\nStyle_A: 10.023091316223145<br>\nStyle_B: 25.63881492614746</p>\n<p>For Top4 methods are slightly better in my best submission file as focused model<br>\nYou can find the submission here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/yale-fwi-best-submission-csv-file\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/yale-fwi-best-submission-csv-file</a></p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 3238435,
                                              "author_name": "jeroencottaar",
                                              "author_url": "",
                                              "post_date": "07/01/2025 20:24:17",
                                              "content": "<p>Yeah I wans't able to make a good prior for StyleB, so that one hurt me bad. I think something weird also happened to some of the CurveFault_A and FaltFault_A datasets; either I messed up my probing or I messed up the model there.</p>\n<p>Anyway, pretty sure there's a sub-5 in ensembling our combined submissions. I'll only be picking this up in a few days, but here's my best in case anyone else wants to have a go: <a href=\"https://www.kaggle.com/datasets/jeroencottaar/best-yale-submission\" target=\"_blank\">https://www.kaggle.com/datasets/jeroencottaar/best-yale-submission</a></p>",
                                              "votes": null,
                                              "replies": [
                                                {
                                                  "id": 3238465,
                                                  "author_name": "overvalueawareness",
                                                  "author_url": "",
                                                  "post_date": "07/01/2025 20:55:47",
                                                  "content": "<p>The average best solution (from csv file) is 6.7<br>\n<a href=\"https://www.kaggle.com/code/overvalueawareness/ensemble-best-solution-seismic25?scriptVersionId=248381868\" target=\"_blank\">https://www.kaggle.com/code/overvalueawareness/ensemble-best-solution-seismic25?scriptVersionId=248381868</a></p>",
                                                  "votes": null,
                                                  "replies": []
                                                },
                                                {
                                                  "id": 3238477,
                                                  "author_name": "shlomoron",
                                                  "author_url": "",
                                                  "post_date": "07/01/2025 21:25:57",
                                                  "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> I tried replacing your preds with mine for StyleB, did not work. Maybe I have a bug, or my classifier is much worse than I thought. Or your probing maybe wrong.  </p>",
                                                  "votes": null,
                                                  "replies": [
                                                    {
                                                      "id": 3238484,
                                                      "author_name": "harshitsheoran",
                                                      "author_url": "",
                                                      "post_date": "07/01/2025 21:41:43",
                                                      "content": "<p>Same, replacing StyleB with mine gives 7.47 private, which is only a bit better</p>",
                                                      "votes": null,
                                                      "replies": [
                                                        {
                                                          "id": 3238504,
                                                          "author_name": "harshitsheoran",
                                                          "author_url": "",
                                                          "post_date": "07/01/2025 22:22:21",
                                                          "content": "<p>Almost 6.5 can be achieved with replacing your Style_B, and all of the Faults families with mine</p>",
                                                          "votes": null,
                                                          "replies": [
                                                            {
                                                              "id": 3238678,
                                                              "author_name": "jeroencottaar",
                                                              "author_url": "",
                                                              "post_date": "07/02/2025 05:27:05",
                                                              "content": "<p>Aww no easy sub-5 just yet.</p>\n<p>Another possibility is that both of your StyleB performances are not as good on LB as CV. I've seen some clues that StyleB in LB may be a bit different from training.</p>\n<p>Anyway, I'll share my probing method later so that you can audit it, and so that we can check it on your submissions.</p>",
                                                              "votes": null,
                                                              "replies": []
                                                            }
                                                          ]
                                                        }
                                                      ]
                                                    }
                                                  ]
                                                }
                                              ]
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3237198,
      "author_name": "lytcnaorevision2018",
      "author_url": "",
      "post_date": "07/01/2025 00:54:14",
      "content": "<p>Amazing jobs!Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237240,
      "author_name": "jeroencottaar",
      "author_url": "",
      "post_date": "07/01/2025 01:45:39",
      "content": "<p>Congratulations! I'll be sharing my solution around the weekend, but in short I used classical full waveform inversion, using <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>'s solution as a starting point. Key point to get this working was to have strong priors.</p>\n<p>Essentially I think your solution is not so very different; in your case the priors are basically learned by the network. Both our solutions require significant compute per test sample (unlike some of the other solutions that are expensive to train but cheap to infer), and I think our compute costs are surprisingly similar.</p>\n<p>More details later this week!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237572,
          "author_name": "tinkei",
          "author_url": "",
          "post_date": "07/01/2025 07:09:50",
          "content": "<p>Looking forward to your writeup. I started out by caching Bartley's ensemble models output, and ran scalar wave FWI to refine only CurveFault_B. (In my CV, FWI mostly helps with <code>_B</code> family but not <code>_A</code>. I don't have enough compute so I focused on the most difficult family.) It ended up performing worse. So I added a loss term to not let the physical inversion deviate too much from the ensemble output (L1 loss between the two velocity maps in Deepwave library). Still, ended up being 0.1 worse than the public model's starting point. Curious to learn from you experience!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3239494,
          "author_name": "arpit1bansal",
          "author_url": "",
          "post_date": "07/02/2025 21:27:25",
          "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> Hey would it be possible for you to tell what GPUs you used, as a student i was only able to use kaggle's T4 [recently got intern, now have little money]. In upcoming competition that required high vram and time for training.<br>\nI would go for rented compute, if you can tell some info about what GPUs you used, rented or you own them and from where you think it's better to rent(you think they provide a friendly env for training).<br>\nThanks.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3239732,
              "author_name": "jeroencottaar",
              "author_url": "",
              "post_date": "07/03/2025 04:37:19",
              "content": "<p>I used vast.ai. I'm not sure if they're the cheapest, but you get full control over your instance (i.e. you just get a Unix environment and can do what you want, including running notebooks for any duration).</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3237257,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "07/01/2025 02:17:28",
      "content": "<p>Congrats! Well deserved. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3247375,
          "author_name": "mehwish004",
          "author_url": "",
          "post_date": "07/12/2025 15:55:38",
          "content": "<p>Congratulations….. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237269,
      "author_name": "overvalueawareness",
      "author_url": "",
      "post_date": "07/01/2025 02:29:34",
      "content": "<p>Congratulations!</p>\n<p>I just wonder how do you sense a model will overfit or not? Did you sense a model will overfit by seeing loss graph for ten epochs ?</p>\n<p>Thanks for reply.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237409,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "07/01/2025 04:49:56",
      "content": "<p>Congratulations Harshit, great work and great result !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237432,
      "author_name": "golden9east",
      "author_url": "",
      "post_date": "07/01/2025 05:09:23",
      "content": "<p>can someone in simple language explain how much time in hrs did it take for training and on what resources. Is kaggle weekly 30 hr gpu sufficient. This was my first competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237557,
          "author_name": "tinkei",
          "author_url": "",
          "post_date": "07/01/2025 07:01:12",
          "content": "<p>The second last section mentioned:</p>\n<blockquote>\n  <p>The whole process from start to finish takes about 15 days on 4x 5090.</p>\n</blockquote>",
          "votes": null,
          "replies": [
            {
              "id": 3237583,
              "author_name": "golden9east",
              "author_url": "",
              "post_date": "07/01/2025 07:19:26",
              "content": "<p>So 15 continuous days of training, how does one do it on kaggle resources, or is it done on private resources. Does it mean from competitive pov, one has little chance of breaking into top ranking with just depending upon kaggle resources?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3237604,
                  "author_name": "jeroencottaar",
                  "author_url": "",
                  "post_date": "07/01/2025 07:29:32",
                  "content": "<p>You have to do it in the cloud or using your own GPUs. For this competition in particular I don't think a top ranking using only Kaggle resources was possible. Luckily this is not the case for all competitions.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3237486,
      "author_name": "dan5280",
      "author_url": "",
      "post_date": "07/01/2025 06:07:32",
      "content": "<p>Congrats!!! Thanks so much for your detailed explanation, but I'm wonder do you have more detail code snippet? Like how you do the reshape how you do the pseudo iteration. It's really helpful and I gained a lot from your post. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237721,
      "author_name": "anglolodorf",
      "author_url": "",
      "post_date": "07/01/2025 09:05:20",
      "content": "<p>congratulation well done</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237754,
      "author_name": "kelvinmwaisaka",
      "author_url": "",
      "post_date": "07/01/2025 09:30:27",
      "content": "<p>Congrats and thanks for sharing the results here. I've totally learned a lot. Its one of my freshest experiences while trying to dive deep in the Data science course.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237904,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/01/2025 11:18:41",
      "content": "<p>Congrats on the win and the lead throughout the competition!</p>\n<blockquote>\n  <p>I apply sigmoid on my logits and scale them in range of (1500, 4500)</p>\n</blockquote>\n<p>I had this on my todo list. Did you compare without sigmoid?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237914,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "07/01/2025 11:29:08",
          "content": "<p>Thank you, I have been doing this from day 1 actually, without sigmoid was worse but my testing was primilinary at the time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237925,
      "author_name": "mjarosz",
      "author_url": "",
      "post_date": "07/01/2025 11:53:41",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237929,
      "author_name": "gguillard",
      "author_url": "",
      "post_date": "07/01/2025 12:03:34",
      "content": "<p>Not sure I got it right : (700,700) is more pixels than (5,1000,70), I guess you interpolated to increase the size ?</p>\n<p>So did I understand correctly that you did (5,1000,70) -&gt; (350,350) -&gt; (476,476) -&gt; (588,588) -&gt; (700,700) ?</p>\n<p>If so, why not (5,1000,70) -&gt; (588,588) -&gt; (700,700), or even (5,1000,70) -&gt; (700,700) directly ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237934,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "07/01/2025 12:09:34",
          "content": "<p>It is much more time efficient and stable to gradually increase the image size this way</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237941,
      "author_name": "lucascandianisouza",
      "author_url": "",
      "post_date": "07/01/2025 12:14:31",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>! Very interesting ideas overall, you deserved it</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237945,
      "author_name": "mahiradev",
      "author_url": "",
      "post_date": "07/01/2025 12:16:45",
      "content": "<p>Congrats!!  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238292,
      "author_name": "johnnyhyland",
      "author_url": "",
      "post_date": "07/01/2025 17:54:57",
      "content": "<p>Yes. This is really cool. The iterative psuedo step is pretty wild, also the 4/10 types is a really nice find I wish I'd seen! awesome work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238383,
      "author_name": "algre231",
      "author_url": "",
      "post_date": "07/01/2025 19:41:01",
      "content": "<p>there is some mistakes, but good</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238535,
      "author_name": "justforfun44",
      "author_url": "",
      "post_date": "07/01/2025 23:53:49",
      "content": "<p>Thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238546,
      "author_name": "lauracatalano",
      "author_url": "",
      "post_date": "07/02/2025 00:23:33",
      "content": "<p>Congratulations and thanks for sharing your solution!</p>\n<p>Do you have any validation plots of the predicted and actual velocity/geology. Since I'm a geologist, I'm just curious what the results look like.Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3239020,
      "author_name": "towhid121",
      "author_url": "",
      "post_date": "07/02/2025 11:53:45",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3239663,
      "author_name": "",
      "author_url": "",
      "post_date": "07/03/2025 03:14:35",
      "content": "<p>Congratulations! Impressive job, learnt a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3239682,
      "author_name": "arashkaroudi",
      "author_url": "",
      "post_date": "07/03/2025 03:45:25",
      "content": "<p>This is absolutely amazing, and I loved your thought process to arrive at such insights. As a newcomer to this field, however, I feel like this is where I have to stop. The computational resources needed in most competitions are truly discouraging and insane for me. Wishing you all the best going forward!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3239733,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "07/03/2025 04:38:51",
          "content": "<p>Please don't be discouraged. This competition really is an outlier in terms of necessary compute. For example, I took second place in this one with 0 GPU hours: <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/overview\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/overview</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3239822,
              "author_name": "arashkaroudi",
              "author_url": "",
              "post_date": "07/03/2025 07:26:55",
              "content": "<p>I went through your solution, and it's incredible how you deconstructed and simplified the problem—that's really something extraordinary. Figuring out how you even got to that solution is a whole other challenge?<br>\nAnyway, thank you for the encouragement. I'll try to keep going, though I'm not sure for how long.<br>\nIt's just very exciting for me to see brilliant minds like yours at work here.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3240177,
              "author_name": "arpit1bansal",
              "author_url": "",
              "post_date": "07/03/2025 14:26:30",
              "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>  Just read you solution in ariel competition, Truly smart, hats off to you.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3239721,
      "author_name": "renukaoladhri",
      "author_url": "",
      "post_date": "07/03/2025 04:22:25",
      "content": "<p>Congratulations!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3239912,
      "author_name": "saraharshadbcs",
      "author_url": "",
      "post_date": "07/03/2025 09:18:20",
      "content": "<p>Congrats man! interesting….</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3239998,
      "author_name": "benedictpangilinan",
      "author_url": "",
      "post_date": "07/03/2025 11:19:16",
      "content": "<p>May I know how does iteration improve the MAE.</p>\n<p>PS: I'm just a beginner</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3240031,
      "author_name": "prabhatyadav369",
      "author_url": "",
      "post_date": "07/03/2025 11:51:03",
      "content": "<p>Congratulations🎉🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3241459,
      "author_name": "bakhtawarshabbir",
      "author_url": "",
      "post_date": "07/04/2025 20:11:04",
      "content": "<p>Smart preprocessing insight</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3242154,
      "author_name": "abelabraham77",
      "author_url": "",
      "post_date": "07/05/2025 16:27:38",
      "content": "<p>woah super cool! Physics-guided ML models are up and coming.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3242462,
      "author_name": "kyleprior",
      "author_url": "",
      "post_date": "07/05/2025 22:55:15",
      "content": "<p>Congrats! That's amazing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3243409,
      "author_name": "nickyjliang",
      "author_url": "",
      "post_date": "07/07/2025 05:34:32",
      "content": "<p>Quite detail ths</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3246618,
      "author_name": "kritikapawar511",
      "author_url": "",
      "post_date": "07/11/2025 08:31:46",
      "content": "<p>congratulations</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3237125": "First and foremost, a sincere thank you to the competition organizers for this challenge, and congratulations to all my fellow competitors on their impressive work.\n\nCrazy how nobody moved a single place up or down in top 28.\n\n\nGoals List:\n✔️Single digit submissions solo gold\n✔️Solo win a competition #1\n\nAnyways, here's what I have done:\n\n# **Preprocessing**\n\nThis is not a segmentation task. The velocity model pixels don't line up spatially with the seismic data.\nAn early experiment tested it. I had a baseline that resized the input from the original (5, 1000, 70) to (5, 350, 70) and got a CV of 46.9. , I reshape the input to (1, 350, 350), basically forcing the channels into a spatial layout. CV jumped to 32.5\n\nThis is how the non-resized 1x1000x350 input looks like:\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F2c0bdcf39b2d105431911c182de5ec99%2Finput.png?generation=1751317173814874&alt=media\" width=\"100%\">\n\n# **Architecture**\n\nUnet didn’t make sense to me, my input was 350x350, the pixel values of the velocity model did not spatially align with the input, so it did not make much sense to me to go back to previous layers to get intermediate layer’s embeddings like a Unet does\n\nAnyways, it turns out, vision transformers don’t necessarily need a Unet to do their bidding 🙂\n\nVision transformers are secretly segmentation-like regression models:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F51344cf7cb6fdeef3087d0bcbe72acd5%2Farch1.png?generation=1751317505425731&alt=media)\n\nThis image size was a nice square so it was easy to work with and scale.\nI started using this without a decoder with the best vision transformer I found, EVA02-small model. The performance was really cool, even at this stage it got me a close to 30 score on CV with just 40 epochs of training.\n\nTo scale up the model, I experimented with base and large variants which were scoring better, but what was scoring even more was repeating the architecture as an encoder+decoder setup:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F84e7795e777d616c22ffc264e10dec37%2Farch2.png?generation=1751317556231754&alt=media)\n\nTraining this scored CV 28, repeating this (encoder+decoder) with base or large variant of the model did not give improvement and were overfitting\n\nTraining EVA, it was still 20-30% slower than training the original ViT, I found a better backbone after experimenting to find out what was making EVA score much better than ViT, Depth? Channels? Number of attention heads? MLP layer scaling? Gated activations? Turns out, RoPE (Rotary Positional Embeddings) was contributing to the majority of the bottleneck, then I chose better pretrained weights and I got vit_small_patch14_reg4_dinov2.lvd142m as my ideal choice of backbone. On the same training, ViT+RoPE got me under 26 MAE.\n\nHere is the final architecture:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F8415956230f0715e1c7a910602e211dc%2Farch3.png?generation=1751317568008933&alt=media)\n\nCode: https://www.kaggle.com/code/harshitsheoran/yale-fwi-vit-architecture?scriptVersionId=248201360\n\nMy loss function is MAE loss, I apply sigmoid on my logits and scale them in range of (1500, 4500)\n\n\n# **Training**\n\nPretty much every training run used the same recipe:\n- LR: 1e-4 with cosine annealing down to ~1e-5\n- Optimizer: AdamW\n- Loss: MAE\n- HFlip and TTA\n- EMA\n\nI used 440k out of 470k for training, 30k for validation, the core of my training at MAE of 26 with 50 epochs is that I am severely underfitting, increasing image size and restarting training can lead to much better scores, for a total of 300 epochs:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2Fe47301b125f7736663b1cc57e63e095b%2Ftrain1.png?generation=1751317762528020&alt=media)\n\nThis model was done training on May 28 (more than a month before the deadline), the last month was done exploring the next part of the solution.\n\n# **Generating Data**\n\nData Augmentation is a key part of my solution, before an almost perfect replication of the forward modeling function was available, I used forward-modeling + denoiser model setup where denoiser model was the architecture above designed to predict the noise still left in the forward-modeled data\n\nAfter failing experiments with generating new velocity models with heuristics or diffusion, I went ahead with using FiveCrop and 5xRandomAffine augmentations {RandomAffine(p=1.0, degrees=15, translate=(0.2, 0), shear=(-15, 15), padding_mode='reflection', resample='nearest')} on existing velocity models to generate more velocity models, this gave me 10x more data of 4.7 million seismic samples, I use this data for pretraining on every step, the training graph looks like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F5fd975fcb6ba850fa56b16a706bf6373%2Ftrain2.png?generation=1751317778801585&alt=media)\n\nLater in the competition, a perfect method for forward modeling was publicly available, at the time, the original notebook was running at a speed of about 10 images per minute which was unusable, using pytorch compile and a 5090, I sped it up to >5000 images per minute per gpu which is a lot more bearable for the next time consuming step.\n\nI call it “Iterative Pseudo” where after every epoch, I predict on my validation and competition’s test set, about ~95k samples in total, this prediction is a pseudo prediction (for the test set) as we do not know how good or bad it is, I take the predicted velocity models and forward-model them to generate what should be the input for them and use this input in training the next epoch. This step is repeated for every epoch, this does add extra time to each epoch but the payout in terms of score is so worth it.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F60f4f0f2f6421676e6d86a7d478b199e%2Ftrain3.png?generation=1751317800532720&alt=media)\n\nAs pretraining epochs are 10x larger, all epochs weighted, this was trained for a total of 570 epochs.\n\nAnother observation was made that most of the new score that the model is improving for a while, has been coming from mainly 4 out of 10 methods. So, model was further trained with removing the data that doesn’t come from [‘CurveFault_B’, ‘CurveVel_B’, ‘Style_A’, ‘Style_B’], this removes more than half of the data, speeds up training, frees up the model’s parameters to focus more on this part.\n\nThen for inference, I would replace the predictions that were from these 4 methods with the predictions from the model below, which yielded a score of 7.5 on CV, 7.0/7.0 on Public/Private LB\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1794509%2F04ec4c683c1f8d4607cd3dcefb7d1df1%2Ftrain_top4.png?generation=1751317941360967&alt=media\" width=\"50%\">\n\nTraining this with larger image size at the end, with 896x896 and ensembling that Top4 model to the existing Top4 model yields my best submission at 7.28 CV, 6.9/6.9 on Public/Private LB\n\nThe whole process from start to finish takes about 15 days on 4x 5090.\n\nMy solution does not feature much for ensembling, and can be cleanly trained for longer to generate a hopeful sub 5 leaderboard.\n\n# **That didn’t work / Didn’t get to try**\n\nMAE/MIM self-supervised pretraining\nStandard Augmentations\nUnet decoder, UperNet, Mask2Former, ViT-Adapter, Fusion heads, more complicated heads than a linear layer\n\nDue to being short on time, I did not go back and recreate the pretraining data with perfect simulation, that could still improve the score more.\n\nIdea that I didn’t get to try: Prediction from the model forward modeled to generate simulated seismic input - original seismic input to compute gradient, the original seismic input, this gradient and the predicted velocity model as input to a new model that will refine the prediction according to the gradient\n\n\n\nThank you for reading!",
    "3237146": "Huge congrats, @harshitsheoran! 🙂 Stellar performance and an awesome write-up! Really neat how you explain your reasoning at various steps and the experiments you ran based on your observations 🙏\n\nI also didn't think unets were the necessary arch here, but in my explorations I tried various archs (perceiver, custom arch, etc) from scratch, but the one take away from this competition for me was just how unbelievably powerful pretrained vision models are (both in terms of the weights and the priors the models themselves embody).\n\nYour exploration and the insights around vision transformers are on another level! 🙂 What an awesome 1st place and so well deserved!",
    "3237163": "Someone once said that all top solution are realy simple. This time, it's not! Really amazing job. So many interesting ideas. Now I really want to do with you a competition sometime 🤣",
    "3237165": "Thank you, I would be looking forward to it 😊",
    "3237181": "As a beginner and self-taught student of data science, I really love the culture of sharing their knowledge and solutions in Kaggle. Thank you for your sharing! Even tho, I did not participate this competition, but I will study this, following your notebook. Thank you😙",
    "3237184": "Congratulations!  @harshitsheoran Excellent work. You have a very good baseline model, and then you keep iterating on it. \nI think the essence is to constantly overfit the data, because they all come from the same simulation model.\nAnyway， congratulations again！",
    "3237190": "Congratulations! If optimized to the limit, is there a chance to break 6?",
    "3237194": "If optimized to the limit, there is a chance to break 5...? Maybe... 6, most likely",
    "3237198": "Amazing jobs!Thank you for sharing!",
    "3237207": "Amazing result😧",
    "3237238": "Hm can you check something? The StyleB samples are very hard to predict - there is a solution where you don't predict local high-frequent features but still match the given seismograms almost perfectly. This is worse than just numerical noise - it's actually a local minimum. Without solving this StyleB alone contributes about 4.5 to the LB score for me.\n\nI would expect your performance on StyleB has stagnated long before the others. (I can share a list of which test samples are StyleB later if needed.)",
    "3237240": "Congratulations! I'll be sharing my solution around the weekend, but in short I used classical full waveform inversion, using @brendanartley's solution as a starting point. Key point to get this working was to have strong priors.\n\nEssentially I think your solution is not so very different; in your case the priors are basically learned by the network. Both our solutions require significant compute per test sample (unlike some of the other solutions that are expensive to train but cheap to infer), and I think our compute costs are surprisingly similar.\n\nMore details later this week!",
    "3237257": "Congrats! Well deserved.",
    "3237269": "Congratulations!\n\nI just wonder how do you sense a model will overfit or not? Did you sense a model will overfit by seeing loss graph for ten epochs ?\n\nThanks for reply.",
    "3237409": "Congratulations Harshit, great work and great result !",
    "3237432": "can someone in simple language explain how much time in hrs did it take for training and on what resources. Is kaggle weekly 30 hr gpu sufficient. This was my first competition.",
    "3237459": "jeroencottaar what was your score only for StyleB?",
    "3237470": "My average score for the StyleB maps in the public test set is ~44, contributing ~4.3 to my score. CV score is similar.",
    "3237486": "Congrats!!! Thanks so much for your detailed explanation, but I'm wonder do you have more detail code snippet? Like how you do the reshape how you do the pseudo iteration. It's really helpful and I gained a lot from your post.",
    "3237557": "The second last section mentioned:\n>The whole process from start to finish takes about 15 days on 4x 5090.",
    "3237569": "Hah. So only for this family, I have MUCH better score.\nValidation is 22 and I can get it lower with better training.  \nWhich mean that for other family your method is so much more powerful than mine.",
    "3237572": "Looking forward to your writeup. I started out by caching Bartley's ensemble models output, and ran scalar wave FWI to refine only CurveFault_B. (In my CV, FWI mostly helps with `_B` family but not `_A`. I don't have enough compute so I focused on the most difficult family.) It ended up performing worse. So I added a loss term to not let the physical inversion deviate too much from the ensemble output (L1 loss between the two velocity maps in Deepwave library). Still, ended up being 0.1 worse than the public model's starting point. Curious to learn from you experience!",
    "3237583": "So 15 continuous days of training, how does one do it on kaggle resources, or is it done on private resources. Does it mean from competitive pov, one has little chance of breaking into top ranking with just depending upon kaggle resources?",
    "3237604": "You have to do it in the cloud or using your own GPUs. For this competition in particular I don't think a top ranking using only Kaggle resources was possible. Luckily this is not the case for all competitions.",
    "3237606": "I think this is one where the deep learning approaches may be able to learn a better prior than my classical ones. In particular, they may know where to expect the high-frequent features that I have no hope of finding. \n\nIt also means that sub-5 may be very well possible, and may not need more than selecting the right solution per family.",
    "3237624": ">It also means that sub-5 may be very well possible, and may not need more than selecting the right solution per family.  \n\nWhich is very easy, very simple for classifier to have 100% accuracy for classify between style and the rest- I did it.",
    "3237636": "Here's everything I know about my public test set performance. This is not CV, but based on probing, though it mostly tracks CV. See any more opportunities? Also @harshitsheoran and other top competitors.\n\n\n| Set                                   | Percentage of public test set | Score    |\n| ------------------------------------- | ------------------------------ | -------- |\n| FlatVelA+B                            | 17.764                        | 0.0        |\n| Fault maps with only 2 velocities | 6.92                         | 8.7  |\n| StyleA                                | 8.892                        | 3.4 |\n| StyleB                                | 9.57                         | 44.4 |\n| All others                            | 56.85 | 4.5 |\n\n\nPretty sure I messed something up in my model for the 2-velocity maps, they should be easy...",
    "3237646": "jeroencottaar let's try it!  \n[Here is](https://www.kaggle.com/code/shlomoron/gwi-final-sub-with-labels/) a notebook with my final ensemble and labels.  \nTry to replace the predictions in test_ensemble with yours, except for the indices where test_kind_labels=9.",
    "3237721": "congratulation well done",
    "3237724": "Cool, will pick this up on Thursday (work deadlines first...)",
    "3237754": "Congrats and thanks for sharing the results here. I've totally learned a lot. Its one of my freshest experiences while trying to dive deep in the Data science course.",
    "3237904": "Congrats on the win and the lead throughout the competition!\n\n>  I apply sigmoid on my logits and scale them in range of (1500, 4500)\n\nI had this on my todo list. Did you compare without sigmoid?",
    "3237914": "Thank you, I have been doing this from day 1 actually, without sigmoid was worse but my testing was primilinary at the time.",
    "3237925": "Congratulations!",
    "3237929": "Not sure I got it right : (700,700) is more pixels than (5,1000,70), I guess you interpolated to increase the size ?\n\nSo did I understand correctly that you did (5,1000,70) -> (350,350) -> (476,476) -> (588,588) -> (700,700) ?\n\nIf so, why not (5,1000,70) -> (588,588) -> (700,700), or even (5,1000,70) -> (700,700) directly ?",
    "3237934": "It is much more time efficient and stable to gradually increase the image size this way",
    "3237941": "Congrats @harshitsheoran! Very interesting ideas overall, you deserved it",
    "3237945": "Congrats!!",
    "3238292": "Yes. This is really cool. The iterative psuedo step is pretty wild, also the 4/10 types is a really nice find I wish I'd seen! awesome work.",
    "3238383": "there is some mistakes, but good",
    "3238402": "Yes, your CV do look off for some classes, FlatVel is perfect, StyleA is too good, \"All others\" have curves and faults, so 4.5 average for them is good\n\nCV For single model:\nMAE 7.940401554107666\nCurveFault_A: 1.8529449701309204\nCurveFault_B: 12.928738594055176\nCurveVel_A: 1.51606023311615\nCurveVel_B: 5.029183864593506\nFlatFault_A: 0.9475402235984802\nFlatFault_B: 2.63285231590271\nFlatVel_A: 0.3224563002586365\nFlatVel_B: 0.5822372436523438\nStyle_A: 10.023091316223145\nStyle_B: 25.63881492614746\n\n\nFor Top4 methods are slightly better in my best submission file as focused model\nYou can find the submission here:\n\nhttps://www.kaggle.com/datasets/harshitsheoran/yale-fwi-best-submission-csv-file",
    "3238435": "Yeah I wans't able to make a good prior for StyleB, so that one hurt me bad. I think something weird also happened to some of the CurveFault_A and FaltFault_A datasets; either I messed up my probing or I messed up the model there.\n\nAnyway, pretty sure there's a sub-5 in ensembling our combined submissions. I'll only be picking this up in a few days, but here's my best in case anyone else wants to have a go: https://www.kaggle.com/datasets/jeroencottaar/best-yale-submission",
    "3238465": "The average best solution (from csv file) is 6.7\nhttps://www.kaggle.com/code/overvalueawareness/ensemble-best-solution-seismic25?scriptVersionId=248381868",
    "3238477": "jeroencottaar I tried replacing your preds with mine for StyleB, did not work. Maybe I have a bug, or my classifier is much worse than I thought. Or your probing maybe wrong.",
    "3238484": "Same, replacing StyleB with mine gives 7.47 private, which is only a bit better",
    "3238504": "Almost 6.5 can be achieved with replacing your Style_B, and all of the Faults families with mine",
    "3238535": "Thank you for sharing",
    "3238546": "Congratulations and thanks for sharing your solution!\n\nDo you have any validation plots of the predicted and actual velocity/geology. Since I'm a geologist, I'm just curious what the results look like.Thanks!",
    "3238678": "Aww no easy sub-5 just yet.\n\nAnother possibility is that both of your StyleB performances are not as good on LB as CV. I've seen some clues that StyleB in LB may be a bit different from training.\n\nAnyway, I'll share my probing method later so that you can audit it, and so that we can check it on your submissions.",
    "3239020": "Congratulations!",
    "3239494": "jeroencottaar Hey would it be possible for you to tell what GPUs you used, as a student i was only able to use kaggle's T4 [recently got intern, now have little money]. In upcoming competition that required high vram and time for training.\nI would go for rented compute, if you can tell some info about what GPUs you used, rented or you own them and from where you think it's better to rent(you think they provide a friendly env for training).\nThanks.",
    "3239663": "Congratulations! Impressive job, learnt a lot!",
    "3239682": "This is absolutely amazing, and I loved your thought process to arrive at such insights. As a newcomer to this field, however, I feel like this is where I have to stop. The computational resources needed in most competitions are truly discouraging and insane for me. Wishing you all the best going forward!",
    "3239721": "Congratulations!!",
    "3239732": "I used vast.ai. I'm not sure if they're the cheapest, but you get full control over your instance (i.e. you just get a Unix environment and can do what you want, including running notebooks for any duration).",
    "3239733": "Please don't be discouraged. This competition really is an outlier in terms of necessary compute. For example, I took second place in this one with 0 GPU hours: https://www.kaggle.com/competitions/ariel-data-challenge-2024/overview",
    "3239822": "I went through your solution, and it's incredible how you deconstructed and simplified the problem—that's really something extraordinary. Figuring out how you even got to that solution is a whole other challenge?\nAnyway, thank you for the encouragement. I'll try to keep going, though I'm not sure for how long.\nIt's just very exciting for me to see brilliant minds like yours at work here.",
    "3239912": "Congrats man! interesting....",
    "3239998": "May I know how does iteration improve the MAE.\n\nPS: I'm just a beginner",
    "3240031": "Congratulations🎉🎉",
    "3240177": "jeroencottaar  Just read you solution in ariel competition, Truly smart, hats off to you.",
    "3241459": "Smart preprocessing insight",
    "3242154": "woah super cool! Physics-guided ML models are up and coming.",
    "3242462": "Congrats! That's amazing!",
    "3243409": "Quite detail ths",
    "3246618": "congratulations",
    "3247375": "Congratulations....."
  },
  "source": "meta"
}