{
  "id": 587389,
  "title": "How far did you get by continuing to train Bartley's type models?",
  "url": "/competitions/waveform-inversion/discussion/587389",
  "author_name": "",
  "post_date": "2025-07-01T00:04:28.863895200Z",
  "votes": 9,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>First of all, thank you so much to the organizers for such a fun and truly deep learning competition! Lots of high quality data, no shenanigans with the train - test split, was such a joy to participate! 😁</p>\n<p>And thank you so much to <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for the amazing shares! Made for such a great time learning and trying out novel approaches 🙏</p>\n<p>The question I would like to ask is this: did you manage to train a model like in Bartley's notebooks (could be with modifications) beyond the 22 - 24 CV range?</p>\n<p>The reason I am asking is that I made some modifications to the training to optimize for speed / theoretically help with convergence, but in the end I didn't manage to break the ~22 - 24 CV barrier.</p>\n<p>The changes I made:</p>\n<ul>\n<li>training on normalized data (mean and std normalization)</li>\n<li>creating custom train - val split with possibly better shuffling</li>\n<li>training in bfloat16</li>\n<li>switching to <code>convnext_large.fb_in22k_ft_in1k</code> for the resnet version</li>\n<li>increasing the number of channels in the decoder</li>\n<li>removing the EMA model due to how long the weights update takes</li>\n</ul>\n<p>If you could please share your experience, that would be very helpful 🙏 I would like to \"debug\" my approach in the hopes of learning how I could improve.</p>\n<p>I plan to keep my single GPU at home busy over the next couple of weeks, attempting to get to the root of the issue, but if you could please share your experience, that would be greatly appreciated!</p>\n<p>Seems this competition has definitely raised my appetite for more after a couple of years' hiatus 😅 Quite looking forward to the future proper deep learning ones!</p>",
  "messages": [
    {
      "id": "3237127",
      "postDate": "07/01/2025 00:04:28",
      "content": "<p>Hi everyone,</p>\n<p>First of all, thank you so much to the organizers for such a fun and truly deep learning competition! Lots of high quality data, no shenanigans with the train - test split, was such a joy to participate! 😁</p>\n<p>And thank you so much to <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for the amazing shares! Made for such a great time learning and trying out novel approaches 🙏</p>\n<p>The question I would like to ask is this: did you manage to train a model like in Bartley's notebooks (could be with modifications) beyond the 22 - 24 CV range?</p>\n<p>The reason I am asking is that I made some modifications to the training to optimize for speed / theoretically help with convergence, but in the end I didn't manage to break the ~22 - 24 CV barrier.</p>\n<p>The changes I made:</p>\n<ul>\n<li>training on normalized data (mean and std normalization)</li>\n<li>creating custom train - val split with possibly better shuffling</li>\n<li>training in bfloat16</li>\n<li>switching to <code>convnext_large.fb_in22k_ft_in1k</code> for the resnet version</li>\n<li>increasing the number of channels in the decoder</li>\n<li>removing the EMA model due to how long the weights update takes</li>\n</ul>\n<p>If you could please share your experience, that would be very helpful 🙏 I would like to \"debug\" my approach in the hopes of learning how I could improve.</p>\n<p>I plan to keep my single GPU at home busy over the next couple of weeks, attempting to get to the root of the issue, but if you could please share your experience, that would be greatly appreciated!</p>\n<p>Seems this competition has definitely raised my appetite for more after a couple of years' hiatus 😅 Quite looking forward to the future proper deep learning ones!</p>",
      "rawMarkdown": "Hi everyone,\n\nFirst of all, thank you so much to the organizers for such a fun and truly deep learning competition! Lots of high quality data, no shenanigans with the train - test split, was such a joy to participate! 😁\n\nAnd thank you so much to @brendanartley for the amazing shares! Made for such a great time learning and trying out novel approaches 🙏\n\nThe question I would like to ask is this: did you manage to train a model like in Bartley's notebooks (could be with modifications) beyond the 22 - 24 CV range?\n\nThe reason I am asking is that I made some modifications to the training to optimize for speed / theoretically help with convergence, but in the end I didn't manage to break the ~22 - 24 CV barrier.\n\nThe changes I made:\n- training on normalized data (mean and std normalization)\n- creating custom train - val split with possibly better shuffling\n- training in bfloat16\n- switching to `convnext_large.fb_in22k_ft_in1k` for the resnet version\n- increasing the number of channels in the decoder\n- removing the EMA model due to how long the weights update takes\n\nIf you could please share your experience, that would be very helpful 🙏 I would like to \"debug\" my approach in the hopes of learning how I could improve.\n\nI plan to keep my single GPU at home busy over the next couple of weeks, attempting to get to the root of the issue, but if you could please share your experience, that would be greatly appreciated!\n\nSeems this competition has definitely raised my appetite for more after a couple of years' hiatus 😅 Quite looking forward to the future proper deep learning ones!",
      "votes": null
    },
    {
      "id": "3237132",
      "postDate": "07/01/2025 00:07:42",
      "content": "<p>add augmentation and you can get lb20.</p>\n<pre><code>augmentation:\n) mix   of same family : aug  = (-mask)*v1 + mask*v2\n)aug seismic = FMW( aug  )\n\nmask  half mask, etc\n</code></pre>\n<p>alternatively, use gan to generate aug velocity</p>",
      "rawMarkdown": "add augmentation and you can get lb20.\n```\naugmentation:\n1) mix 2 velocity of same family : aug velocity = (1-mask)*v1 + mask*v2\n2)aug seismic = FMW( aug velocity )\n\nmask random half mask, etc\n```\n\nalternatively, use gan to generate aug velocity",
      "votes": null
    },
    {
      "id": "3237149",
      "postDate": "07/01/2025 00:19:55",
      "content": "<p>my score (13.4) is only using Bartley caformer. I don't know anything in computer vision, hence focused on how to best train the best public model.</p>\n<p>The only change I made (very proud it was also used by <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> ) was to also train it with the 5 seismic planes side by side, i.e. a single channel 1000 x 350 image.</p>",
      "rawMarkdown": "my score (13.4) is only using Bartley caformer. I don't know anything in computer vision, hence focused on how to best train the best public model.\n\nThe only change I made (very proud it was also used by @harshitsheoran ) was to also train it with the 5 seismic planes side by side, i.e. a single channel 1000 x 350 image.",
      "votes": null
    },
    {
      "id": "3237151",
      "postDate": "07/01/2025 00:20:29",
      "content": "<p>Thank you so much for your reply <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>!</p>\n<p>So essentially, it seems there might have not been anything wrong with my setup, there might have been a hard ceiling for these models at around ~22lb - ~24lb without other tricks (like data aug)?</p>\n<p>That would be a great relief 😅</p>",
      "rawMarkdown": "Thank you so much for your reply @hengck23!\n\nSo essentially, it seems there might have not been anything wrong with my setup, there might have been a hard ceiling for these models at around ~22lb - ~24lb without other tricks (like data aug)?\n\nThat would be a great relief 😅",
      "votes": null
    },
    {
      "id": "3237157",
      "postDate": "07/01/2025 00:25:13",
      "content": "<p>example of augmention</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa1044b1c3294f19fde8c19bb5af5d27f%2FSelection_133.png?generation=1751329401848736&amp;alt=media\" alt=\"\"></p>\n<p>using mix mask<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F539f32dc55b130cf92fac2b47d93c5f7%2FSelection_197.png?generation=1751329511441305&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "example of augmention\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa1044b1c3294f19fde8c19bb5af5d27f%2FSelection_133.png?generation=1751329401848736&alt=media)\n\n\nusing mix mask\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F539f32dc55b130cf92fac2b47d93c5f7%2FSelection_197.png?generation=1751329511441305&alt=media)",
      "votes": null
    },
    {
      "id": "3237166",
      "postDate": "07/01/2025 00:28:55",
      "content": "<p>here is the breakdown:<br>\n28.8 : bartely's caformer<br>\n25.+ : ensemble of (bartely's caformers)<br>\n20.+ : augmentation<br>\n15.+  : add test samples into the learning set (predicted velocity, simulated seismic)<br>\nsome more gain if you fine-tune for each family</p>",
      "rawMarkdown": "here is the breakdown:\n28.8 : bartely's caformer\n25.+ : ensemble of (bartely's caformers)\n20.+ : augmentation\n15.+  : add test samples into the learning set (predicted velocity, simulated seismic)\nsome more gain if you fine-tune for each family",
      "votes": null
    },
    {
      "id": "3237172",
      "postDate": "07/01/2025 00:33:35",
      "content": "<p>by adding 2m generated samples (from our own velocity simulation code + forward modeling), long training (150-300 epochs) and larger convnext backbones (large, xxlarge or hugev2), we can achieve 10.x - 11.x scores with the exactly same Bartley type's model.</p>\n<p>We didn't experiment with anything related to architecture (collapsing the time dimension, resizing images to a square or larger resolution), which turned out to be a mistake. We got stuck after reaching 11 for quite a while :D. </p>",
      "rawMarkdown": "by adding 2m generated samples (from our own velocity simulation code + forward modeling), long training (150-300 epochs) and larger convnext backbones (large, xxlarge or hugev2), we can achieve 10.x - 11.x scores with the exactly same Bartley type's model.\n\nWe didn't experiment with anything related to architecture (collapsing the time dimension, resizing images to a square or larger resolution), which turned out to be a mistake. We got stuck after reaching 11 for quite a while :D.",
      "votes": null
    },
    {
      "id": "3237177",
      "postDate": "07/01/2025 00:40:25",
      "content": "<p>so it proves that with 10x data can approximate the physics inverse model</p>",
      "rawMarkdown": "so it proves that with 10x data can approximate the physics inverse model",
      "votes": null
    },
    {
      "id": "3237183",
      "postDate": "07/01/2025 00:44:20",
      "content": "<p>Thank you so much for your reply, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>! 🙏🏻</p>\n<p>Could I please ask you if you trained in float32? Technically, bfloat16 shouldn’t make a difference, as far as I understand, but that is one thing about my approach I didn’t verify…</p>",
      "rawMarkdown": "Thank you so much for your reply, @cpmpml! 🙏🏻\n\nCould I please ask you if you trained in float32? Technically, bfloat16 shouldn’t make a difference, as far as I understand, but that is one thing about my approach I didn’t verify…",
      "votes": null
    },
    {
      "id": "3237187",
      "postDate": "07/01/2025 00:46:18",
      "content": "<p>Thank you for sharing, <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a>, super helpful!</p>\n<p>Could I please ask you if you trained in float32?</p>",
      "rawMarkdown": "Thank you for sharing, @andy2709, super helpful!\n\nCould I please ask you if you trained in float32?",
      "votes": null
    },
    {
      "id": "3237188",
      "postDate": "07/01/2025 00:47:46",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, extremely helpful 🙏🏻🙂</p>",
      "rawMarkdown": "Thank you, @hengck23, extremely helpful 🙏🏻🙂",
      "votes": null
    },
    {
      "id": "3237193",
      "postDate": "07/01/2025 00:51:43",
      "content": "<p>no <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, we trained all models in bfloat16. </p>",
      "rawMarkdown": "no @radek1, we trained all models in bfloat16.",
      "votes": null
    },
    {
      "id": "3237195",
      "postDate": "07/01/2025 00:51:50",
      "content": "<p>I trained with bf16.</p>",
      "rawMarkdown": "I trained with bf16.",
      "votes": null
    },
    {
      "id": "3237196",
      "postDate": "07/01/2025 00:51:53",
      "content": "<p>And <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, if I could please ask you, did you train in float32?</p>\n<p>I’m curious if that could be one of the things I missed…</p>",
      "rawMarkdown": "And @hengck23, if I could please ask you, did you train in float32?\n\nI’m curious if that could be one of the things I missed…",
      "votes": null
    },
    {
      "id": "3237223",
      "postDate": "07/01/2025 01:13:33",
      "content": "<p>I only use bartley's caformer and finetune it with pseudo labels, and then I can get 24.3.</p>",
      "rawMarkdown": "I only use bartley's caformer and finetune it with pseudo labels, and then I can get 24.3.",
      "votes": null
    },
    {
      "id": "3237228",
      "postDate": "07/01/2025 01:22:20",
      "content": "<p>Bfloat16 is good enough</p>",
      "rawMarkdown": "Bfloat16 is good enough",
      "votes": null
    },
    {
      "id": "3237256",
      "postDate": "07/01/2025 02:15:05",
      "content": "<p>Glad to hear you are coming back. Your and Chris D’s efforts in LLM science exam really changed my entire career trajectory! :)</p>",
      "rawMarkdown": "Glad to hear you are coming back. Your and Chris D’s efforts in LLM science exam really changed my entire career trajectory! :)",
      "votes": null
    },
    {
      "id": "3237270",
      "postDate": "07/01/2025 02:29:48",
      "content": "<p>Wow, that is terrific to hear! 😁 🙏</p>",
      "rawMarkdown": "Wow, that is terrific to hear! 😁 🙏",
      "votes": null
    },
    {
      "id": "3237503",
      "postDate": "07/01/2025 06:17:00",
      "content": "<p>May i ask how do you finetune with pseudo labels?</p>",
      "rawMarkdown": "May i ask how do you finetune with pseudo labels?",
      "votes": null
    },
    {
      "id": "3237652",
      "postDate": "07/01/2025 08:03:47",
      "content": "<p>Just want to add that this competition should be put in ML competition textbook.<br>\nIt is actually is very good for tutorial.<br>\nit has all aspects and tricks : novel network design, synthetic data generation, physics modeling, multi-gpu training , fast cuda coding …</p>",
      "rawMarkdown": "Just want to add that this competition should be put in ML competition textbook.\nIt is actually is very good for tutorial.\nit has all aspects and tricks : novel network design, synthetic data generation, physics modeling, multi-gpu training , fast cuda coding ...",
      "votes": null
    },
    {
      "id": "3237657",
      "postDate": "07/01/2025 08:05:57",
      "content": "<p>1) Predict given test seis data to get vel data <strong>y_hat</strong><br>\n2) Use Forward Modeling to generate new seis data <strong>x_hat</strong> with predicted vel data <strong>y_hat</strong><br>\n3) Finetune model with (<strong>x_hat</strong>, <strong>y_hat</strong>)</p>",
      "rawMarkdown": "1) Predict given test seis data to get vel data **y_hat**\n2) Use Forward Modeling to generate new seis data **x_hat** with predicted vel data **y_hat**\n3) Finetune model with (**x_hat**, **y_hat**)",
      "votes": null
    },
    {
      "id": "3237689",
      "postDate": "07/01/2025 08:34:56",
      "content": "<p>The following is forward modeling. <br>\nAm i correct?</p>\n<p><a href=\"https://www.kaggle.com/code/manatoyo/improved-vel-to-seis\" target=\"_blank\">https://www.kaggle.com/code/manatoyo/improved-vel-to-seis</a></p>",
      "rawMarkdown": "The following is forward modeling. \nAm i correct?\n\nhttps://www.kaggle.com/code/manatoyo/improved-vel-to-seis",
      "votes": null
    },
    {
      "id": "3237717",
      "postDate": "07/01/2025 09:02:27",
      "content": "<p>Yes, but I use this code: <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/585236\" target=\"_blank\">https://www.kaggle.com/competitions/waveform-inversion/discussion/585236</a></p>",
      "rawMarkdown": "Yes, but I use this code: https://www.kaggle.com/competitions/waveform-inversion/discussion/585236",
      "votes": null
    },
    {
      "id": "3238557",
      "postDate": "07/02/2025 01:04:30",
      "content": "<p>Yes, 100%! 🙂 Also the data is great and there is a lot of data to go by, what an awesome competition this was!!!!</p>",
      "rawMarkdown": "Yes, 100%! 🙂 Also the data is great and there is a lot of data to go by, what an awesome competition this was!!!!",
      "votes": null
    },
    {
      "id": "3238635",
      "postDate": "07/02/2025 04:01:40",
      "content": "<p>i have a suggestino for you</p>\n<p>1) start a community competition, entitled:<br>\n\"post competition study of  Yale/UNC-CH - Geophysicale …\"<br>\nobjective:<br>\nafter reading the top solution, you may think \"hey, the top solution is smiliar to mine. why his results is much better when i did what he did?\" This study aim to algorithm debug. Let's find out what is done wrong and help to do better in future competition.</p>\n<p>dataset:<br>\nscaled version of full data (e.g. just choose curvefault^ ,….</p>\n<p>plan:<br>\nlayout what you want to do:</p>\n<ul>\n<li>design new network (that is better than what you have used in the current competitioon)</li>\n<li>understand the physics of wave modeling and  implement cuda  code</li>\n<li>… repeat winning solution …</li>\n</ul>\n<p>duration:<br>\ni would say 4 weeks is a good timeline.</p>\n<ul>\n<li>consistently published your code notebook for the plan above</li>\n<li>start a discord group, people can join and debug each other code,algorithm</li>\n</ul>\n<p>note that is a \"group-study competition\".everyone can share code and discuss. the leaderboard is just a check to see if everything was done correctly: you should get same score as others have</p>\n<hr>\n<p>why this effort?</p>\n<ul>\n<li>my experience is that a (large) group of people trying to debug and improve your code/agorithm is more efficient then doing it yourself.<br>\nother may give you unexpected improvement </li>\n</ul>",
      "rawMarkdown": "i have a suggestino for you\n\n1) start a community competition, entitled:\n\"post competition study of  Yale/UNC-CH - Geophysicale ...\"\nobjective:\nafter reading the top solution, you may think \"hey, the top solution is smiliar to mine. why his results is much better when i did what he did?\" This study aim to algorithm debug. Let's find out what is done wrong and help to do better in future competition.\n\ndataset:\nscaled version of full data (e.g. just choose curvefault^ ,....\n\nplan:\nlayout what you want to do:\n- design new network (that is better than what you have used in the current competitioon)\n- understand the physics of wave modeling and  implement cuda  code\n- ... repeat winning solution ...\n\nduration:\ni would say 4 weeks is a good timeline.\n- consistently published your code notebook for the plan above\n- start a discord group, people can join and debug each other code,algorithm\n\nnote that is a \"group-study competition\".everyone can share code and discuss. the leaderboard is just a check to see if everything was done correctly: you should get same score as others have\n\n---\n\nwhy this effort?\n- my experience is that a (large) group of people trying to debug and improve your code/agorithm is more efficient then doing it yourself.\nother may give you unexpected improvement",
      "votes": null
    },
    {
      "id": "3238831",
      "postDate": "07/02/2025 08:31:02",
      "content": "<p>That is a great idea! I have been actually thinking of doing something along these lines, but using the late submission to this competition 🙂</p>\n<p>Let me work on this a bit more and will share 🙂 But thank you so much for all your thoughts, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, really appreciated 🙏</p>",
      "rawMarkdown": "That is a great idea! I have been actually thinking of doing something along these lines, but using the late submission to this competition 🙂\n\nLet me work on this a bit more and will share 🙂 But thank you so much for all your thoughts, @hengck23, really appreciated 🙏",
      "votes": null
    },
    {
      "id": "3238912",
      "postDate": "07/02/2025 10:08:06",
      "content": "<blockquote>\n  <p>1) Predict given test seis data to get vel data y_hat<br>\n  2) Use Forward Modeling to generate new seis data x_hat with predicted vel data y_hat<br>\n  3) Finetune model with (x_hat, y_hat)</p>\n</blockquote>\n<p>I got caformer to 13.4 with this method, see my writeup.</p>",
      "rawMarkdown": "> 1) Predict given test seis data to get vel data y_hat\n> 2) Use Forward Modeling to generate new seis data x_hat with predicted vel data y_hat\n> 3) Finetune model with (x_hat, y_hat)\n\nI got caformer to 13.4 with this method, see my writeup.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237132,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/01/2025 00:07:42",
      "content": "<p>add augmentation and you can get lb20.</p>\n<pre><code>augmentation:\n) mix   of same family : aug  = (-mask)*v1 + mask*v2\n)aug seismic = FMW( aug  )\n\nmask  half mask, etc\n</code></pre>\n<p>alternatively, use gan to generate aug velocity</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237151,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/01/2025 00:20:29",
          "content": "<p>Thank you so much for your reply <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>!</p>\n<p>So essentially, it seems there might have not been anything wrong with my setup, there might have been a hard ceiling for these models at around ~22lb - ~24lb without other tricks (like data aug)?</p>\n<p>That would be a great relief 😅</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237166,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "07/01/2025 00:28:55",
              "content": "<p>here is the breakdown:<br>\n28.8 : bartely's caformer<br>\n25.+ : ensemble of (bartely's caformers)<br>\n20.+ : augmentation<br>\n15.+  : add test samples into the learning set (predicted velocity, simulated seismic)<br>\nsome more gain if you fine-tune for each family</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3237157,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/01/2025 00:25:13",
          "content": "<p>example of augmention</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa1044b1c3294f19fde8c19bb5af5d27f%2FSelection_133.png?generation=1751329401848736&amp;alt=media\" alt=\"\"></p>\n<p>using mix mask<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F539f32dc55b130cf92fac2b47d93c5f7%2FSelection_197.png?generation=1751329511441305&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 3237188,
              "author_name": "radek1",
              "author_url": "",
              "post_date": "07/01/2025 00:47:46",
              "content": "<p>Thank you, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, extremely helpful 🙏🏻🙂</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3237196,
                  "author_name": "radek1",
                  "author_url": "",
                  "post_date": "07/01/2025 00:51:53",
                  "content": "<p>And <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, if I could please ask you, did you train in float32?</p>\n<p>I’m curious if that could be one of the things I missed…</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3237228,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "07/01/2025 01:22:20",
                      "content": "<p>Bfloat16 is good enough</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3237149,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/01/2025 00:19:55",
      "content": "<p>my score (13.4) is only using Bartley caformer. I don't know anything in computer vision, hence focused on how to best train the best public model.</p>\n<p>The only change I made (very proud it was also used by <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> ) was to also train it with the 5 seismic planes side by side, i.e. a single channel 1000 x 350 image.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237183,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/01/2025 00:44:20",
          "content": "<p>Thank you so much for your reply, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>! 🙏🏻</p>\n<p>Could I please ask you if you trained in float32? Technically, bfloat16 shouldn’t make a difference, as far as I understand, but that is one thing about my approach I didn’t verify…</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237195,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "07/01/2025 00:51:50",
              "content": "<p>I trained with bf16.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3237172,
      "author_name": "andy2709",
      "author_url": "",
      "post_date": "07/01/2025 00:33:35",
      "content": "<p>by adding 2m generated samples (from our own velocity simulation code + forward modeling), long training (150-300 epochs) and larger convnext backbones (large, xxlarge or hugev2), we can achieve 10.x - 11.x scores with the exactly same Bartley type's model.</p>\n<p>We didn't experiment with anything related to architecture (collapsing the time dimension, resizing images to a square or larger resolution), which turned out to be a mistake. We got stuck after reaching 11 for quite a while :D. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3237177,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/01/2025 00:40:25",
          "content": "<p>so it proves that with 10x data can approximate the physics inverse model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3237187,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/01/2025 00:46:18",
          "content": "<p>Thank you for sharing, <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a>, super helpful!</p>\n<p>Could I please ask you if you trained in float32?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237193,
              "author_name": "andy2709",
              "author_url": "",
              "post_date": "07/01/2025 00:51:43",
              "content": "<p>no <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, we trained all models in bfloat16. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3237223,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "07/01/2025 01:13:33",
      "content": "<p>I only use bartley's caformer and finetune it with pseudo labels, and then I can get 24.3.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237503,
          "author_name": "overvalueawareness",
          "author_url": "",
          "post_date": "07/01/2025 06:17:00",
          "content": "<p>May i ask how do you finetune with pseudo labels?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237657,
              "author_name": "i2nfinit3y",
              "author_url": "",
              "post_date": "07/01/2025 08:05:57",
              "content": "<p>1) Predict given test seis data to get vel data <strong>y_hat</strong><br>\n2) Use Forward Modeling to generate new seis data <strong>x_hat</strong> with predicted vel data <strong>y_hat</strong><br>\n3) Finetune model with (<strong>x_hat</strong>, <strong>y_hat</strong>)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3237689,
                  "author_name": "overvalueawareness",
                  "author_url": "",
                  "post_date": "07/01/2025 08:34:56",
                  "content": "<p>The following is forward modeling. <br>\nAm i correct?</p>\n<p><a href=\"https://www.kaggle.com/code/manatoyo/improved-vel-to-seis\" target=\"_blank\">https://www.kaggle.com/code/manatoyo/improved-vel-to-seis</a></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3237717,
                      "author_name": "i2nfinit3y",
                      "author_url": "",
                      "post_date": "07/01/2025 09:02:27",
                      "content": "<p>Yes, but I use this code: <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/585236\" target=\"_blank\">https://www.kaggle.com/competitions/waveform-inversion/discussion/585236</a></p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3238912,
                          "author_name": "cpmpml",
                          "author_url": "",
                          "post_date": "07/02/2025 10:08:06",
                          "content": "<blockquote>\n  <p>1) Predict given test seis data to get vel data y_hat<br>\n  2) Use Forward Modeling to generate new seis data x_hat with predicted vel data y_hat<br>\n  3) Finetune model with (x_hat, y_hat)</p>\n</blockquote>\n<p>I got caformer to 13.4 with this method, see my writeup.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3237256,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "07/01/2025 02:15:05",
      "content": "<p>Glad to hear you are coming back. Your and Chris D’s efforts in LLM science exam really changed my entire career trajectory! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237270,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/01/2025 02:29:48",
          "content": "<p>Wow, that is terrific to hear! 😁 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237652,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/01/2025 08:03:47",
      "content": "<p>Just want to add that this competition should be put in ML competition textbook.<br>\nIt is actually is very good for tutorial.<br>\nit has all aspects and tricks : novel network design, synthetic data generation, physics modeling, multi-gpu training , fast cuda coding …</p>",
      "votes": null,
      "replies": [
        {
          "id": 3238557,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/02/2025 01:04:30",
          "content": "<p>Yes, 100%! 🙂 Also the data is great and there is a lot of data to go by, what an awesome competition this was!!!!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3238635,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "07/02/2025 04:01:40",
              "content": "<p>i have a suggestino for you</p>\n<p>1) start a community competition, entitled:<br>\n\"post competition study of  Yale/UNC-CH - Geophysicale …\"<br>\nobjective:<br>\nafter reading the top solution, you may think \"hey, the top solution is smiliar to mine. why his results is much better when i did what he did?\" This study aim to algorithm debug. Let's find out what is done wrong and help to do better in future competition.</p>\n<p>dataset:<br>\nscaled version of full data (e.g. just choose curvefault^ ,….</p>\n<p>plan:<br>\nlayout what you want to do:</p>\n<ul>\n<li>design new network (that is better than what you have used in the current competitioon)</li>\n<li>understand the physics of wave modeling and  implement cuda  code</li>\n<li>… repeat winning solution …</li>\n</ul>\n<p>duration:<br>\ni would say 4 weeks is a good timeline.</p>\n<ul>\n<li>consistently published your code notebook for the plan above</li>\n<li>start a discord group, people can join and debug each other code,algorithm</li>\n</ul>\n<p>note that is a \"group-study competition\".everyone can share code and discuss. the leaderboard is just a check to see if everything was done correctly: you should get same score as others have</p>\n<hr>\n<p>why this effort?</p>\n<ul>\n<li>my experience is that a (large) group of people trying to debug and improve your code/agorithm is more efficient then doing it yourself.<br>\nother may give you unexpected improvement </li>\n</ul>",
              "votes": null,
              "replies": [
                {
                  "id": 3238831,
                  "author_name": "radek1",
                  "author_url": "",
                  "post_date": "07/02/2025 08:31:02",
                  "content": "<p>That is a great idea! I have been actually thinking of doing something along these lines, but using the late submission to this competition 🙂</p>\n<p>Let me work on this a bit more and will share 🙂 But thank you so much for all your thoughts, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, really appreciated 🙏</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237127": "Hi everyone,\n\nFirst of all, thank you so much to the organizers for such a fun and truly deep learning competition! Lots of high quality data, no shenanigans with the train - test split, was such a joy to participate! 😁\n\nAnd thank you so much to @brendanartley for the amazing shares! Made for such a great time learning and trying out novel approaches 🙏\n\nThe question I would like to ask is this: did you manage to train a model like in Bartley's notebooks (could be with modifications) beyond the 22 - 24 CV range?\n\nThe reason I am asking is that I made some modifications to the training to optimize for speed / theoretically help with convergence, but in the end I didn't manage to break the ~22 - 24 CV barrier.\n\nThe changes I made:\n- training on normalized data (mean and std normalization)\n- creating custom train - val split with possibly better shuffling\n- training in bfloat16\n- switching to `convnext_large.fb_in22k_ft_in1k` for the resnet version\n- increasing the number of channels in the decoder\n- removing the EMA model due to how long the weights update takes\n\nIf you could please share your experience, that would be very helpful 🙏 I would like to \"debug\" my approach in the hopes of learning how I could improve.\n\nI plan to keep my single GPU at home busy over the next couple of weeks, attempting to get to the root of the issue, but if you could please share your experience, that would be greatly appreciated!\n\nSeems this competition has definitely raised my appetite for more after a couple of years' hiatus 😅 Quite looking forward to the future proper deep learning ones!",
    "3237132": "add augmentation and you can get lb20.\n```\naugmentation:\n1) mix 2 velocity of same family : aug velocity = (1-mask)*v1 + mask*v2\n2)aug seismic = FMW( aug velocity )\n\nmask random half mask, etc\n```\n\nalternatively, use gan to generate aug velocity",
    "3237149": "my score (13.4) is only using Bartley caformer. I don't know anything in computer vision, hence focused on how to best train the best public model.\n\nThe only change I made (very proud it was also used by @harshitsheoran ) was to also train it with the 5 seismic planes side by side, i.e. a single channel 1000 x 350 image.",
    "3237151": "Thank you so much for your reply @hengck23!\n\nSo essentially, it seems there might have not been anything wrong with my setup, there might have been a hard ceiling for these models at around ~22lb - ~24lb without other tricks (like data aug)?\n\nThat would be a great relief 😅",
    "3237157": "example of augmention\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa1044b1c3294f19fde8c19bb5af5d27f%2FSelection_133.png?generation=1751329401848736&alt=media)\n\n\nusing mix mask\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F539f32dc55b130cf92fac2b47d93c5f7%2FSelection_197.png?generation=1751329511441305&alt=media)",
    "3237166": "here is the breakdown:\n28.8 : bartely's caformer\n25.+ : ensemble of (bartely's caformers)\n20.+ : augmentation\n15.+  : add test samples into the learning set (predicted velocity, simulated seismic)\nsome more gain if you fine-tune for each family",
    "3237172": "by adding 2m generated samples (from our own velocity simulation code + forward modeling), long training (150-300 epochs) and larger convnext backbones (large, xxlarge or hugev2), we can achieve 10.x - 11.x scores with the exactly same Bartley type's model.\n\nWe didn't experiment with anything related to architecture (collapsing the time dimension, resizing images to a square or larger resolution), which turned out to be a mistake. We got stuck after reaching 11 for quite a while :D.",
    "3237177": "so it proves that with 10x data can approximate the physics inverse model",
    "3237183": "Thank you so much for your reply, @cpmpml! 🙏🏻\n\nCould I please ask you if you trained in float32? Technically, bfloat16 shouldn’t make a difference, as far as I understand, but that is one thing about my approach I didn’t verify…",
    "3237187": "Thank you for sharing, @andy2709, super helpful!\n\nCould I please ask you if you trained in float32?",
    "3237188": "Thank you, @hengck23, extremely helpful 🙏🏻🙂",
    "3237193": "no @radek1, we trained all models in bfloat16.",
    "3237195": "I trained with bf16.",
    "3237196": "And @hengck23, if I could please ask you, did you train in float32?\n\nI’m curious if that could be one of the things I missed…",
    "3237223": "I only use bartley's caformer and finetune it with pseudo labels, and then I can get 24.3.",
    "3237228": "Bfloat16 is good enough",
    "3237256": "Glad to hear you are coming back. Your and Chris D’s efforts in LLM science exam really changed my entire career trajectory! :)",
    "3237270": "Wow, that is terrific to hear! 😁 🙏",
    "3237503": "May i ask how do you finetune with pseudo labels?",
    "3237652": "Just want to add that this competition should be put in ML competition textbook.\nIt is actually is very good for tutorial.\nit has all aspects and tricks : novel network design, synthetic data generation, physics modeling, multi-gpu training , fast cuda coding ...",
    "3237657": "1) Predict given test seis data to get vel data **y_hat**\n2) Use Forward Modeling to generate new seis data **x_hat** with predicted vel data **y_hat**\n3) Finetune model with (**x_hat**, **y_hat**)",
    "3237689": "The following is forward modeling. \nAm i correct?\n\nhttps://www.kaggle.com/code/manatoyo/improved-vel-to-seis",
    "3237717": "Yes, but I use this code: https://www.kaggle.com/competitions/waveform-inversion/discussion/585236",
    "3238557": "Yes, 100%! 🙂 Also the data is great and there is a lot of data to go by, what an awesome competition this was!!!!",
    "3238635": "i have a suggestino for you\n\n1) start a community competition, entitled:\n\"post competition study of  Yale/UNC-CH - Geophysicale ...\"\nobjective:\nafter reading the top solution, you may think \"hey, the top solution is smiliar to mine. why his results is much better when i did what he did?\" This study aim to algorithm debug. Let's find out what is done wrong and help to do better in future competition.\n\ndataset:\nscaled version of full data (e.g. just choose curvefault^ ,....\n\nplan:\nlayout what you want to do:\n- design new network (that is better than what you have used in the current competitioon)\n- understand the physics of wave modeling and  implement cuda  code\n- ... repeat winning solution ...\n\nduration:\ni would say 4 weeks is a good timeline.\n- consistently published your code notebook for the plan above\n- start a discord group, people can join and debug each other code,algorithm\n\nnote that is a \"group-study competition\".everyone can share code and discuss. the leaderboard is just a check to see if everything was done correctly: you should get same score as others have\n\n---\n\nwhy this effort?\n- my experience is that a (large) group of people trying to debug and improve your code/agorithm is more efficient then doing it yourself.\nother may give you unexpected improvement",
    "3238831": "That is a great idea! I have been actually thinking of doing something along these lines, but using the late submission to this competition 🙂\n\nLet me work on this a bit more and will share 🙂 But thank you so much for all your thoughts, @hengck23, really appreciated 🙏",
    "3238912": "> 1) Predict given test seis data to get vel data y_hat\n> 2) Use Forward Modeling to generate new seis data x_hat with predicted vel data y_hat\n> 3) Finetune model with (x_hat, y_hat)\n\nI got caformer to 13.4 with this method, see my writeup."
  },
  "source": "meta"
}