{
  "id": 579841,
  "title": "ConvNeXt Approach - [CV 31.9 LB 36.4]",
  "url": "/competitions/waveform-inversion/discussion/579841",
  "author_name": "Bartley",
  "post_date": "2025-05-20T18:43:51.935000",
  "votes": 84,
  "comment_count": 53,
  "views": 0,
  "content": "<p>Here is another baseline trained on the full-resolution data. Like my previous notebook, I provide a few trained checkpoints and leave lots of room for customization.</p>\n<p>Notebook <a href=\"https://www.kaggle.com/code/brendanartley/convnext-baseline\" target=\"_blank\">here</a><br>\nWeights <a href=\"https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72\" target=\"_blank\">here</a></p>\n<p>Here are the CV scores for this approach. Cheers!</p>\n<pre><code>+--------------+--------+\n| Dataset      | Score  |\n+--------------+--------+\n| CurveFault_A |    |\n| CurveFault_B |   |\n| CurveVel_A   |   |\n| CurveVel_B   |   |\n| FlatFault_A  |    |\n| FlatFault_B  |   |\n| FlatVel_A    |    |\n| FlatVel_B    |   |\n| Style_A      |   |\n| Style_B      |   |\n+--------------+--------+\n</code></pre>",
  "messages": [
    {
      "id": 3206031,
      "postDate": "2025-05-20T18:43:51.937Z",
      "content": "<p>Here is another baseline trained on the full-resolution data. Like my previous notebook, I provide a few trained checkpoints and leave lots of room for customization.</p>\n<p>Notebook <a href=\"https://www.kaggle.com/code/brendanartley/convnext-baseline\" target=\"_blank\">here</a><br>\nWeights <a href=\"https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72\" target=\"_blank\">here</a></p>\n<p>Here are the CV scores for this approach. Cheers!</p>\n<pre><code>+--------------+--------+\n| Dataset      | Score  |\n+--------------+--------+\n| CurveFault_A |    |\n| CurveFault_B |   |\n| CurveVel_A   |   |\n| CurveVel_B   |   |\n| FlatFault_A  |    |\n| FlatFault_B  |   |\n| FlatVel_A    |    |\n| FlatVel_B    |   |\n| Style_A      |   |\n| Style_B      |   |\n+--------------+--------+\n</code></pre>",
      "rawMarkdown": "Here is another baseline trained on the full-resolution data. Like my previous notebook, I provide a few trained checkpoints and leave lots of room for customization.\n\nNotebook [here](https://www.kaggle.com/code/brendanartley/convnext-baseline)\nWeights [here](https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72)\n\nHere are the CV scores for this approach. Cheers!\n\n```python\n+--------------+--------+\n| Dataset      | Score  |\n+--------------+--------+\n| CurveFault_A |   6.07 |\n| CurveFault_B |  92.54 |\n| CurveVel_A   |  15.07 |\n| CurveVel_B   |  53.67 |\n| FlatFault_A  |   4.32 |\n| FlatFault_B  |  37.60 |\n| FlatVel_A    |   2.62 |\n| FlatVel_B    |  12.89 |\n| Style_A      |  37.38 |\n| Style_B      |  57.70 |\n+--------------+--------+\n```",
      "votes": 84
    },
    {
      "id": 3206525,
      "postDate": "2025-05-21T12:47:47.330Z",
      "content": "<p>A simple regression model shows that the LB score of the next \"baseline\" Bartley model, due out next week, is less than 20. Tremble, leaders!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291846%2Fbbda54814df51697ff7445fdb8482044%2FBartley-LB.png?generation=1747831623968838&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "A simple regression model shows that the LB score of the next \"baseline\" Bartley model, due out next week, is less than 20. Tremble, leaders!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291846%2Fbbda54814df51697ff7445fdb8482044%2FBartley-LB.png?generation=1747831623968838&alt=media)",
      "votes": 34,
      "replies": [
        {
          "id": 3206569,
          "postDate": "2025-05-21T14:15:24.590Z",
          "content": "<p>Huge Respect!!</p>",
          "rawMarkdown": "Huge Respect!!",
          "votes": 2
        }
      ]
    },
    {
      "id": 3206037,
      "postDate": "2025-05-20T18:54:33.627Z",
      "content": "<p>Loss curves below. Each model takes ~10hrs for 50 epochs on 1×4090 with optimizations (e.g., torch.compile(), fused optimizer, increased num_workers, mixed precision, etc.).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F893d5096069da688ce1249ab739ef7de%2Flogs_wandb.JPG?generation=1747767254477184&amp;alt=media\" alt=\"WandbLogs\"></p>",
      "rawMarkdown": "Loss curves below. Each model takes ~10hrs for 50 epochs on 1×4090 with optimizations (e.g., torch.compile(), fused optimizer, increased num_workers, mixed precision, etc.).\n\n![WandbLogs](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F893d5096069da688ce1249ab739ef7de%2Flogs_wandb.JPG?generation=1747767254477184&alt=media)",
      "votes": 18,
      "replies": [
        {
          "id": 3206054,
          "postDate": "2025-05-20T19:14:28.547Z",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> are you evaluated every 20% of epoch?</p>\n<blockquote>\n  <p>These charts helps to cross check the stability of the local codes, Thanks you.</p>\n</blockquote>",
          "rawMarkdown": "@brendanartley are you evaluated every 20% of epoch?\n> These charts helps to cross check the stability of the local codes, Thanks you.",
          "votes": 1,
          "replies": [
            {
              "id": 3206132,
              "postDate": "2025-05-20T22:04:42.027Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a>, evaluation happens once each epoch.</p>",
              "rawMarkdown": "Hi @seshurajup, evaluation happens once each epoch.",
              "votes": 2
            }
          ]
        },
        {
          "id": 3206948,
          "postDate": "2025-05-22T02:34:35.700Z",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Thank you for sharing your approach in this competition—it’s impressive. On my end, even with multiple GPUs（float16 full resolution), a single epoch still takes around 30 minutes.<br>\nyou’re able to run 50 epochs on a single RTX 4090 in roughly the same time. What optimizations or tricks are you using to achieve that speed?</p>",
          "rawMarkdown": "@brendanartley Thank you for sharing your approach in this competition—it’s impressive. On my end, even with multiple GPUs（float16 full resolution), a single epoch still takes around 30 minutes.\nyou’re able to run 50 epochs on a single RTX 4090 in roughly the same time. What optimizations or tricks are you using to achieve that speed?",
          "votes": 1,
          "replies": [
            {
              "id": 3207207,
              "postDate": "2025-05-22T12:04:56.330Z",
              "content": "<p>When using multiple gpu's, are you properly sharding the dataset? Otherwise, 30m per epoch is not too bad either.</p>",
              "rawMarkdown": "When using multiple gpu's, are you properly sharding the dataset? Otherwise, 30m per epoch is not too bad either.",
              "votes": 3
            },
            {
              "id": 3207213,
              "postDate": "2025-05-22T12:11:35.550Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/koukinn\" target=\"_blank\">@koukinn</a>, its hard to give tips as each machine could have different bottlenecks. </p>\n<p>For me, the biggest speedups came from using <code>AdamW(..., fused=True)</code> and <code>torch.compile(..., mode=\"max-autotune\")</code>. These two changes made my pipeline ~2x faster.</p>",
              "rawMarkdown": "Hi @koukinn, its hard to give tips as each machine could have different bottlenecks. \n\nFor me, the biggest speedups came from using `AdamW(..., fused=True)` and `torch.compile(..., mode=\"max-autotune\")`. These two changes made my pipeline ~2x faster.",
              "votes": 6
            },
            {
              "id": 3207286,
              "postDate": "2025-05-22T14:27:12.573Z",
              "content": "<p>Since you said RTX 4090 I assume you ran it outside of Kaggle? <br>\nIf yes, how did you do so please? (overall as I am struggling to do so; training on kaggle takes2 hours per episode)</p>\n<p>Thank you in advance</p>",
              "rawMarkdown": "Since you said RTX 4090 I assume you ran it outside of Kaggle? \nIf yes, how did you do so please? (overall as I am struggling to do so; training on kaggle takes2 hours per episode)\n\nThank you in advance",
              "votes": 1
            },
            {
              "id": 3207705,
              "postDate": "2025-05-23T06:42:28.770Z",
              "content": "<p>Thank you so much</p>",
              "rawMarkdown": "Thank you so much",
              "votes": 1
            },
            {
              "id": 3209113,
              "postDate": "2025-05-25T08:49:29.897Z",
              "content": "<p>For my model, the difference between max-autotune and reduce-overhead is minor, but max-autotune takes a lot (!) longer for initial optimization. Like 2m vs 30m+. Might be worth checking out as well.</p>",
              "rawMarkdown": "For my model, the difference between max-autotune and reduce-overhead is minor, but max-autotune takes a lot (!) longer for initial optimization. Like 2m vs 30m+. Might be worth checking out as well.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3214227,
          "postDate": "2025-05-31T06:07:54.490Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3207307,
      "postDate": "2025-05-22T15:04:10.927Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> </p>\n<p>Awesome Open!, I am sharing a little bit more finetuned models of yours</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/simple-further-finetuned-bartley-open-models\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/simple-further-finetuned-bartley-open-models</a></p>\n<p>You can add them to your notebook if you want ( don't want the credit, the hard work was yours ), I have shared you the notebook that would use these weights, the link is at<br>\n<a href=\"https://www.kaggle.com/code/harshitsheoran/convnext-baseline/\" target=\"_blank\">https://www.kaggle.com/code/harshitsheoran/convnext-baseline/</a></p>\n<p>As long as there is no leak that is happening in my finetuning, it's CV 0.289 :)</p>\n<p>If you do choose to use these models, ofc, I will say what the finetuning was :D</p>\n<p>Update:<br>\nHere is my exact code that I used to finetune the models:</p>\n<p><a href=\"https://www.kaggle.com/code/harshitsheoran/recipe-simple-further-finetuned-bartley-models\" target=\"_blank\">https://www.kaggle.com/code/harshitsheoran/recipe-simple-further-finetuned-bartley-models</a></p>\n<p>Have a good day!</p>",
      "rawMarkdown": "Hi @brendanartley \n\nAwesome Open!, I am sharing a little bit more finetuned models of yours\n\nhttps://www.kaggle.com/datasets/harshitsheoran/simple-further-finetuned-bartley-open-models\n\nYou can add them to your notebook if you want ( don't want the credit, the hard work was yours ), I have shared you the notebook that would use these weights, the link is at\nhttps://www.kaggle.com/code/harshitsheoran/convnext-baseline/\n\nAs long as there is no leak that is happening in my finetuning, it's CV 0.289 :)\n\nIf you do choose to use these models, ofc, I will say what the finetuning was :D\n\nUpdate:\nHere is my exact code that I used to finetune the models:\n\nhttps://www.kaggle.com/code/harshitsheoran/recipe-simple-further-finetuned-bartley-models\n\nHave a good day!",
      "votes": 13,
      "replies": [
        {
          "id": 3207338,
          "postDate": "2025-05-22T15:43:37.623Z",
          "content": "<p>Amazing <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>, you deserve the credit on this one! Just updated the public notebook with your weights 😄</p>\n<p>CV: 31.9 -&gt; 28.9<br>\nLB: 36.4 -&gt; 33.2</p>",
          "rawMarkdown": "Amazing @harshitsheoran, you deserve the credit on this one! Just updated the public notebook with your weights 😄\n\nCV: 31.9 -> 28.9\nLB: 36.4 -> 33.2",
          "votes": 5,
          "replies": [
            {
              "id": 3207548,
              "postDate": "2025-05-22T23:39:41.587Z",
              "content": "<p>As you have seen in the code, Finetuning had no change in model or data, the only difference was that I copied your model code and placed it my notebook, which had changes in:</p>\n<ul>\n<li>Learning Rate and Global Batch Size</li>\n<li>The way EMA is applied</li>\n<li>The finetuning was short, like 10-20 epochs, then the score was converging slowly so I stopped the training</li>\n</ul>",
              "rawMarkdown": "As you have seen in the code, Finetuning had no change in model or data, the only difference was that I copied your model code and placed it my notebook, which had changes in:\n- Learning Rate and Global Batch Size\n- The way EMA is applied\n- The finetuning was short, like 10-20 epochs, then the score was converging slowly so I stopped the training",
              "votes": 2
            },
            {
              "id": 3207724,
              "postDate": "2025-05-23T07:02:51.893Z",
              "content": "<p>Not to be a party spoiler , but <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> if you have shared the code with only <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> , then would this not be Private sharing ? since you both are not part of the same team .  Thanks for sharing the weights and giving hints . Be careful man ..</p>",
              "rawMarkdown": "Not to be a party spoiler , but @harshitsheoran if you have shared the code with only @brendanartley , then would this not be Private sharing ? since you both are not part of the same team .  Thanks for sharing the weights and giving hints . Be careful man ..",
              "votes": 7
            },
            {
              "id": 3207755,
              "postDate": "2025-05-23T07:35:02.947Z",
              "content": "<p><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> I will open the code to everyone, I had the intention that I am opening the weights to everyone and the method to use those weights is very simple, as it's just plug and play, but I wanted a secure line where if I say that weights can just be used as is, someone else might copy his work before him, ofc, now it's not needed and I have opened the notebook so we can see that there is nothing I have shared that isn't the same code as his original notebook haha :)</p>\n<p>Thanks for advice, I should have thought this through</p>",
              "rawMarkdown": "@phoenix9032 I will open the code to everyone, I had the intention that I am opening the weights to everyone and the method to use those weights is very simple, as it's just plug and play, but I wanted a secure line where if I say that weights can just be used as is, someone else might copy his work before him, ofc, now it's not needed and I have opened the notebook so we can see that there is nothing I have shared that isn't the same code as his original notebook haha :)\n\nThanks for advice, I should have thought this through",
              "votes": 4
            },
            {
              "id": 3207801,
              "postDate": "2025-05-23T08:36:01.603Z",
              "content": "<p>So what  is The way EMA is applied  mean?</p>",
              "rawMarkdown": "So what  is The way EMA is applied  mean?",
              "votes": 2
            },
            {
              "id": 3207803,
              "postDate": "2025-05-23T08:41:49.667Z",
              "content": "<p>I don't think EMA is any different, I am just using a different library to apply EMA which I have been using for an year, impact should be negligible regardless, my guess is that score comes from just having a larger global batch size and a slight decay in learning rate</p>",
              "rawMarkdown": "I don't think EMA is any different, I am just using a different library to apply EMA which I have been using for an year, impact should be negligible regardless, my guess is that score comes from just having a larger global batch size and a slight decay in learning rate",
              "votes": 4
            },
            {
              "id": 3207807,
              "postDate": "2025-05-23T08:52:22.077Z",
              "content": "<p>No worries … both of you have been doing great work in this competition and putting a lot of effort , so just wanted to caution … </p>",
              "rawMarkdown": "No worries ... both of you have been doing great work in this competition and putting a lot of effort , so just wanted to caution ... ",
              "votes": 3
            },
            {
              "id": 3214068,
              "postDate": "2025-05-30T22:03:38.980Z",
              "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> You have not shared publicly the code you commented as follows:</p>\n<blockquote>\n  <p>As you have seen in the code, Finetuning had no change in model or data, the only difference was that I copied your model code and placed it my notebook, which had changes in:</p>\n  <p>Learning Rate and Global Batch Size<br>\n     The way EMA is applied<br>\n    The finetuning was short, like 10-20 epochs, then the score was converging slowly so I stopped the training<br>\n  and<br>\n   I am just using a different library to apply EMA which I have been using for an year,</p>\n</blockquote>\n<p>You shared a copy of <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> notebook but it is not the one you used to finetune the models.</p>\n<p></p>\n<p>I react because it is not fine to claim you shared your code publicly. Please don't lie to us.</p>\n<p>I don't want to shadow that you shared your improved weights. This is great, thanks for that.</p>\n<p>Edit: I update this because there was no private sharing.</p>",
              "rawMarkdown": "@harshitsheoran You have not shared publicly the code you commented as follows:\n\n> As you have seen in the code, Finetuning had no change in model or data, the only difference was that I copied your model code and placed it my notebook, which had changes in:\n>\n >   Learning Rate and Global Batch Size\n>    The way EMA is applied\n >   The finetuning was short, like 10-20 epochs, then the score was converging slowly so I stopped the training\nand\n>  I am just using a different library to apply EMA which I have been using for an year,\n\nYou shared a copy of @brendanartley notebook but it is not the one you used to finetune the models.\n\n~~It is not fine to privately share code but that's Kaggle business to deal with it, not mine.~~\n\nI react because it is not fine to claim you shared your code publicly. Please don't lie to us.\n\nI don't want to shadow that you shared your improved weights. This is great, thanks for that.\n\nEdit: I update this because there was no private sharing.\n\n",
              "votes": -1
            },
            {
              "id": 3214089,
              "postDate": "2025-05-30T23:18:06.680Z",
              "content": "<p>Totally agree <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>. The right way is to make the notebook public before sharing - though I cant control what is shared with me. The only code I have seen is <a href=\"https://www.kaggle.com/code/harshitsheoran/convnext-baseline/\" target=\"_blank\">here</a></p>\n<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> Would be great if you could give more details, but understand if you would rather not 🙂</p>",
              "rawMarkdown": "Totally agree @cpmpml. The right way is to make the notebook public before sharing - though I cant control what is shared with me. The only code I have seen is [here](https://www.kaggle.com/code/harshitsheoran/convnext-baseline/)\n\n@harshitsheoran Would be great if you could give more details, but understand if you would rather not 🙂",
              "votes": 2
            },
            {
              "id": 3214112,
              "postDate": "2025-05-31T00:58:07.477Z",
              "content": "<p>Great, then there is no private sharing. I'll edit my previous reply then.</p>",
              "rawMarkdown": "Great, then there is no private sharing. I'll edit my previous reply then.",
              "votes": 2
            },
            {
              "id": 3214140,
              "postDate": "2025-05-31T02:26:57.960Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> !</p>\n<p>I am sorry but I think I am misunderstanding the quote</p>\n<blockquote>\n  <p>It is not fine to claim that you shared your code publicly</p>\n</blockquote>\n<p>Are we talking about the code that was shared first only to bartley? Yes, that was an unintentional mistake I made as there was no new code in that notebook so I failed to think it through before sharing it and I opened that code to everyone as soon as that mistake was pointed out to show that there was no new code in that notebook.</p>\n<p>Sorry, I re-read your message from the start:<br>\nTalking about the code that was used to finetune (it was finetuned on local machine), where I said \"I will say what the finetuning was\", although I meant it literally, I have no issues making the code public because there is nothing good in that, I believe that this great work can be pushed much further than my dumb finetune.</p>\n<p>By \"As you have seen in the code\" I meant that as everyone has noticed in the inference code that I did not change any data or model so inference remains the same</p>\n<p>Edit: Updated the original comment with the link to the code that was used for finetuning</p>",
              "rawMarkdown": "Hi @cpmpml !\n\nI am sorry but I think I am misunderstanding the quote\n>It is not fine to claim that you shared your code publicly\n\nAre we talking about the code that was shared first only to bartley? Yes, that was an unintentional mistake I made as there was no new code in that notebook so I failed to think it through before sharing it and I opened that code to everyone as soon as that mistake was pointed out to show that there was no new code in that notebook.\n\nSorry, I re-read your message from the start:\nTalking about the code that was used to finetune (it was finetuned on local machine), where I said \"I will say what the finetuning was\", although I meant it literally, I have no issues making the code public because there is nothing good in that, I believe that this great work can be pushed much further than my dumb finetune.\n\nBy \"As you have seen in the code\" I meant that as everyone has noticed in the inference code that I did not change any data or model so inference remains the same\n\nEdit: Updated the original comment with the link to the code that was used for finetuning",
              "votes": 4
            },
            {
              "id": 3214297,
              "postDate": "2025-05-31T08:25:39.283Z",
              "content": "<p>Awesome, thanks.</p>",
              "rawMarkdown": "Awesome, thanks."
            }
          ]
        }
      ]
    },
    {
      "id": 3214069,
      "postDate": "2025-05-30T22:06:38.087Z",
      "content": "<p>There is a lot to learn from your notebook, thanks for sharing.</p>\n<p>I never used mmap in numpy, this is really neat. Not sure who first share dit here, but kudos to whom did it.</p>\n<p>I still have to dig into how you modified convnext!</p>",
      "rawMarkdown": "There is a lot to learn from your notebook, thanks for sharing.\n\nI never used mmap in numpy, this is really neat. Not sure who first share dit here, but kudos to whom did it.\n\nI still have to dig into how you modified convnext!",
      "votes": 5
    },
    {
      "id": 3212157,
      "postDate": "2025-05-29T11:54:01.607Z",
      "content": "<p>Based on past few days experimentation, I realized how strong the baseline <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> has provided.  There are lot of tricks that are implemented in the network, vanilla network would not perform this well even with out of the box encoder and decoder ! Great work!</p>",
      "rawMarkdown": "Based on past few days experimentation, I realized how strong the baseline @brendanartley has provided.  There are lot of tricks that are implemented in the network, vanilla network would not perform this well even with out of the box encoder and decoder ! Great work!",
      "votes": 5
    },
    {
      "id": 3206983,
      "postDate": "2025-05-22T04:10:19.450Z",
      "content": "<p>Huge respect! Although it is honestly kinda depressing to see my ranking drop from 14 -&gt; 150+ in two days.😂</p>",
      "rawMarkdown": "Huge respect! Although it is honestly kinda depressing to see my ranking drop from 14 -> 150+ in two days.😂",
      "votes": 5
    },
    {
      "id": 3206224,
      "postDate": "2025-05-21T03:58:34.537Z",
      "content": "<p>Show biggest respect to you!</p>",
      "rawMarkdown": "Show biggest respect to you!",
      "votes": 6
    },
    {
      "id": 3214010,
      "postDate": "2025-05-30T18:36:35.933Z",
      "content": "<p>Excellent notebook, thanks for sharing!</p>\n<p>Just 1 question, how did you train on \"1x4090\" GPU?</p>\n<p>The code seems to require \"Requires &gt;= 2 GPUs with CUDA enabled.\"</p>",
      "rawMarkdown": "Excellent notebook, thanks for sharing!\n\nJust 1 question, how did you train on \"1x4090\" GPU?\n\nThe code seems to require \"Requires >= 2 GPUs with CUDA enabled.\"",
      "votes": 1,
      "replies": [
        {
          "id": 3214086,
          "postDate": "2025-05-30T23:03:48.723Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/williamwu88\" target=\"_blank\">@williamwu88</a>, good point. You will have to modify the training loop for the script to work with 1 GPU.</p>",
          "rawMarkdown": "Hi @williamwu88, good point. You will have to modify the training loop for the script to work with 1 GPU.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3212174,
      "postDate": "2025-05-29T12:31:38.427Z",
      "content": "<p>Can someone please explain why was the stride custom changed in this case? </p>",
      "rawMarkdown": "Can someone please explain why was the stride custom changed in this case? ",
      "votes": 1,
      "replies": [
        {
          "id": 3212189,
          "postDate": "2025-05-29T13:00:33.073Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/meddhifallah\" target=\"_blank\">@meddhifallah</a>, does <a href=\"https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200035\" target=\"_blank\">this comment</a> help?</p>",
          "rawMarkdown": "Hi @meddhifallah, does [this comment](https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200035) help?",
          "votes": 2
        }
      ]
    },
    {
      "id": 3210285,
      "postDate": "2025-05-27T02:09:52.153Z",
      "content": "<p>That's truly amazing. I'm learning a lot.</p>",
      "rawMarkdown": "That's truly amazing. I'm learning a lot.",
      "votes": 1
    },
    {
      "id": 3210078,
      "postDate": "2025-05-26T17:31:44.220Z",
      "content": "<p>Looks truly awesome </p>",
      "rawMarkdown": "Looks truly awesome ",
      "votes": 1
    },
    {
      "id": 3208519,
      "postDate": "2025-05-24T09:22:36.323Z",
      "content": "<p>That's really great!!</p>",
      "rawMarkdown": "That's really great!!",
      "votes": 1
    },
    {
      "id": 3207893,
      "postDate": "2025-05-23T11:42:16.180Z",
      "content": "<p>it will be very helpful in nova days!</p>",
      "rawMarkdown": "it will be very helpful in nova days!",
      "votes": 1
    },
    {
      "id": 3207856,
      "postDate": "2025-05-23T10:32:20.167Z",
      "content": "<p>hi <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> couple of quick thoughts/questions, and I'd love to hear your perspective if you have time:</p>\n<p>Resolution scaling: Did you observe any significant performance changes or convergence issues when training on the full resolution compared to down sampled versions? Given the computational cost, I just wanted to understanding the trade-off<br>\nLoss function Insights: What loss function did you primarily use for this ConvNeXt model? Did you explore any alternatives, perhaps something more tailored to the characteristics of geophysical waveforms or a hybrid loss?<br>\nFuture Directions: Based on these results, where do you see the most immediate opportunities for improvement? Perhaps exploring different attention mechanisms, a deeper understanding of the \"B\" datasets, or specialized data augmentation techniques?</p>",
      "rawMarkdown": "hi @brendanartley couple of quick thoughts/questions, and I'd love to hear your perspective if you have time:\n\nResolution scaling: Did you observe any significant performance changes or convergence issues when training on the full resolution compared to down sampled versions? Given the computational cost, I just wanted to understanding the trade-off\nLoss function Insights: What loss function did you primarily use for this ConvNeXt model? Did you explore any alternatives, perhaps something more tailored to the characteristics of geophysical waveforms or a hybrid loss?\nFuture Directions: Based on these results, where do you see the most immediate opportunities for improvement? Perhaps exploring different attention mechanisms, a deeper understanding of the \"B\" datasets, or specialized data augmentation techniques?",
      "votes": 1,
      "replies": [
        {
          "id": 3207918,
          "postDate": "2025-05-23T12:32:27.870Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/bruhjjwal\" target=\"_blank\">@bruhjjwal</a>, thanks for the comment.</p>\n<ol>\n<li><p>Using the full resolution data helped significantly.</p></li>\n<li><p>L1 loss (have not tried other loss functions yet)</p></li>\n<li><p>Its hard to tell where to focus efforts at the moment. Currently I am playing with batch size / learning rate combinations. In addition, I would like to have a fully converged model before analyzing out of fold predictions. </p></li>\n</ol>\n<p>Would recommend reading <a href=\"https://www.kaggle.com/code/tpmeli/beyond-mae-depth-curves-spectra-and-residuals/comments#3203025\" target=\"_blank\">this</a> thread. <a href=\"https://www.kaggle.com/hoffmanns\" target=\"_blank\">@hoffmanns</a> makes some good points!</p>",
          "rawMarkdown": "Hi @bruhjjwal, thanks for the comment.\n\n1. Using the full resolution data helped significantly.\n\n2. L1 loss (have not tried other loss functions yet)\n\n3. Its hard to tell where to focus efforts at the moment. Currently I am playing with batch size / learning rate combinations. In addition, I would like to have a fully converged model before analyzing out of fold predictions. \n\nWould recommend reading [this](https://www.kaggle.com/code/tpmeli/beyond-mae-depth-curves-spectra-and-residuals/comments#3203025) thread. @hoffmanns makes some good points!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3207848,
      "postDate": "2025-05-23T10:16:42.233Z",
      "content": "<p>What batch size do you train the model with? 256 or 16?</p>",
      "rawMarkdown": "What batch size do you train the model with? 256 or 16?\n",
      "votes": 1,
      "replies": [
        {
          "id": 3207913,
          "postDate": "2025-05-23T12:23:37.177Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/vonmainstein\" target=\"_blank\">@vonmainstein</a>, I used a batch size of 16 for this model.</p>",
          "rawMarkdown": "Hi @vonmainstein, I used a batch size of 16 for this model.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3207633,
      "postDate": "2025-05-23T04:36:22.677Z",
      "content": "<p>Huge Respect!</p>",
      "rawMarkdown": "Huge Respect!",
      "votes": 1
    },
    {
      "id": 3206499,
      "postDate": "2025-05-21T11:58:55Z",
      "content": "<p>HUGE RESPECT!</p>",
      "rawMarkdown": "HUGE RESPECT!",
      "votes": 1
    },
    {
      "id": 3206193,
      "postDate": "2025-05-21T02:42:39.023Z",
      "content": "<p>Is the main improvement due to the use of full-resolution data?🥰</p>",
      "rawMarkdown": "Is the main improvement due to the use of full-resolution data?🥰",
      "votes": 1,
      "replies": [
        {
          "id": 3206195,
          "postDate": "2025-05-21T02:51:56.513Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/guoooooooss\" target=\"_blank\">@guoooooooss</a>, I think so. Though this was my first experiment using the full-resolution data.</p>",
          "rawMarkdown": "Hi @guoooooooss, I think so. Though this was my first experiment using the full-resolution data.",
          "votes": 3,
          "replies": [
            {
              "id": 3206200,
              "postDate": "2025-05-21T02:57:29.917Z",
              "content": "<p>thank you very much!🥰</p>",
              "rawMarkdown": "thank you very much!🥰",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3207695,
      "postDate": "2025-05-23T06:22:24.473Z",
      "content": "<p>Nice work! Appreciate so much</p>",
      "rawMarkdown": "Nice work! Appreciate so much\n",
      "votes": 2
    },
    {
      "id": 3206070,
      "postDate": "2025-05-20T19:28:02.483Z",
      "content": "<p>Bro casually messing up the LB with \"Baseline\" and \"Starter\" notebooks :)</p>\n<p>Jokes aside, I love seeing your notebooks popping up.</p>",
      "rawMarkdown": "Bro casually messing up the LB with \"Baseline\" and \"Starter\" notebooks :)\n\nJokes aside, I love seeing your notebooks popping up.",
      "votes": 2
    },
    {
      "id": 3206580,
      "postDate": "2025-05-21T14:26:12.983Z",
      "content": "<p>nice sharing 👍<br>\nDoes pretrained encoder help?</p>",
      "rawMarkdown": "nice sharing 👍\nDoes pretrained encoder help?"
    },
    {
      "id": 3207057,
      "postDate": "2025-05-22T06:54:42.817Z",
      "content": "<p>Beautiful LB.<br>\nIn this match, we achieved communism.<br>\nHuge respect to Bartley. Love you bro!</p>",
      "rawMarkdown": "Beautiful LB.\nIn this match, we achieved communism.\nHuge respect to Bartley. Love you bro!",
      "votes": -1
    },
    {
      "id": 3207055,
      "postDate": "2025-05-22T06:53:18.650Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3206865,
      "postDate": "2025-05-21T21:17:51.613Z",
      "content": "<p>Great work :) :)</p>",
      "rawMarkdown": "Great work :) :)",
      "votes": 1
    },
    {
      "id": 3206771,
      "postDate": "2025-05-21T18:43:15.097Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3206525,
      "author_name": "ln",
      "author_url": "",
      "post_date": "2025-05-21T12:47:47.330000",
      "content": "<p>A simple regression model shows that the LB score of the next \"baseline\" Bartley model, due out next week, is less than 20. Tremble, leaders!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291846%2Fbbda54814df51697ff7445fdb8482044%2FBartley-LB.png?generation=1747831623968838&amp;alt=media\" alt=\"\"></p>",
      "votes": 34,
      "replies": [
        {
          "id": 3206569,
          "author_name": "Marpini Pavan Srikar",
          "author_url": "",
          "post_date": "2025-05-21T14:15:24.590000",
          "content": "<p>Huge Respect!!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3206037,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2025-05-20T18:54:33.627000",
      "content": "<p>Loss curves below. Each model takes ~10hrs for 50 epochs on 1×4090 with optimizations (e.g., torch.compile(), fused optimizer, increased num_workers, mixed precision, etc.).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F893d5096069da688ce1249ab739ef7de%2Flogs_wandb.JPG?generation=1747767254477184&amp;alt=media\" alt=\"WandbLogs\"></p>",
      "votes": 18,
      "replies": [
        {
          "id": 3206054,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2025-05-20T19:14:28.547000",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> are you evaluated every 20% of epoch?</p>\n<blockquote>\n  <p>These charts helps to cross check the stability of the local codes, Thanks you.</p>\n</blockquote>",
          "votes": 1,
          "replies": [
            {
              "id": 3206132,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2025-05-20T22:04:42.027000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a>, evaluation happens once each epoch.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 3206948,
          "author_name": "Gao Xin",
          "author_url": "",
          "post_date": "2025-05-22T02:34:35.700000",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Thank you for sharing your approach in this competition—it’s impressive. On my end, even with multiple GPUs（float16 full resolution), a single epoch still takes around 30 minutes.<br>\nyou’re able to run 50 epochs on a single RTX 4090 in roughly the same time. What optimizations or tricks are you using to achieve that speed?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3207207,
              "author_name": "Sebastian Hoffmann",
              "author_url": "",
              "post_date": "2025-05-22T12:04:56.330000",
              "content": "<p>When using multiple gpu's, are you properly sharding the dataset? Otherwise, 30m per epoch is not too bad either.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3207213,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2025-05-22T12:11:35.550000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/koukinn\" target=\"_blank\">@koukinn</a>, its hard to give tips as each machine could have different bottlenecks. </p>\n<p>For me, the biggest speedups came from using <code>AdamW(..., fused=True)</code> and <code>torch.compile(..., mode=\"max-autotune\")</code>. These two changes made my pipeline ~2x faster.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 3207286,
              "author_name": "Med Dhifallah",
              "author_url": "",
              "post_date": "2025-05-22T14:27:12.573000",
              "content": "<p>Since you said RTX 4090 I assume you ran it outside of Kaggle? <br>\nIf yes, how did you do so please? (overall as I am struggling to do so; training on kaggle takes2 hours per episode)</p>\n<p>Thank you in advance</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3207705,
              "author_name": "Gao Xin",
              "author_url": "",
              "post_date": "2025-05-23T06:42:28.770000",
              "content": "<p>Thank you so much</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3209113,
              "author_name": "Sebastian Hoffmann",
              "author_url": "",
              "post_date": "2025-05-25T08:49:29.897000",
              "content": "<p>For my model, the difference between max-autotune and reduce-overhead is minor, but max-autotune takes a lot (!) longer for initial optimization. Like 2m vs 30m+. Might be worth checking out as well.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3214227,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-31T06:07:54.490000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3207307,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2025-05-22T15:04:10.927000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> </p>\n<p>Awesome Open!, I am sharing a little bit more finetuned models of yours</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/simple-further-finetuned-bartley-open-models\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/simple-further-finetuned-bartley-open-models</a></p>\n<p>You can add them to your notebook if you want ( don't want the credit, the hard work was yours ), I have shared you the notebook that would use these weights, the link is at<br>\n<a href=\"https://www.kaggle.com/code/harshitsheoran/convnext-baseline/\" target=\"_blank\">https://www.kaggle.com/code/harshitsheoran/convnext-baseline/</a></p>\n<p>As long as there is no leak that is happening in my finetuning, it's CV 0.289 :)</p>\n<p>If you do choose to use these models, ofc, I will say what the finetuning was :D</p>\n<p>Update:<br>\nHere is my exact code that I used to finetune the models:</p>\n<p><a href=\"https://www.kaggle.com/code/harshitsheoran/recipe-simple-further-finetuned-bartley-models\" target=\"_blank\">https://www.kaggle.com/code/harshitsheoran/recipe-simple-further-finetuned-bartley-models</a></p>\n<p>Have a good day!</p>",
      "votes": 13,
      "replies": [
        {
          "id": 3207338,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-05-22T15:43:37.623000",
          "content": "<p>Amazing <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>, you deserve the credit on this one! Just updated the public notebook with your weights 😄</p>\n<p>CV: 31.9 -&gt; 28.9<br>\nLB: 36.4 -&gt; 33.2</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3207548,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2025-05-22T23:39:41.587000",
              "content": "<p>As you have seen in the code, Finetuning had no change in model or data, the only difference was that I copied your model code and placed it my notebook, which had changes in:</p>\n<ul>\n<li>Learning Rate and Global Batch Size</li>\n<li>The way EMA is applied</li>\n<li>The finetuning was short, like 10-20 epochs, then the score was converging slowly so I stopped the training</li>\n</ul>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3207724,
              "author_name": "Nirjhar Roy",
              "author_url": "",
              "post_date": "2025-05-23T07:02:51.893000",
              "content": "<p>Not to be a party spoiler , but <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> if you have shared the code with only <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> , then would this not be Private sharing ? since you both are not part of the same team .  Thanks for sharing the weights and giving hints . Be careful man ..</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 3207755,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2025-05-23T07:35:02.947000",
              "content": "<p><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> I will open the code to everyone, I had the intention that I am opening the weights to everyone and the method to use those weights is very simple, as it's just plug and play, but I wanted a secure line where if I say that weights can just be used as is, someone else might copy his work before him, ofc, now it's not needed and I have opened the notebook so we can see that there is nothing I have shared that isn't the same code as his original notebook haha :)</p>\n<p>Thanks for advice, I should have thought this through</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3207801,
              "author_name": "guo dashuai",
              "author_url": "",
              "post_date": "2025-05-23T08:36:01.603000",
              "content": "<p>So what  is The way EMA is applied  mean?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3207803,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2025-05-23T08:41:49.667000",
              "content": "<p>I don't think EMA is any different, I am just using a different library to apply EMA which I have been using for an year, impact should be negligible regardless, my guess is that score comes from just having a larger global batch size and a slight decay in learning rate</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3207807,
              "author_name": "Nirjhar Roy",
              "author_url": "",
              "post_date": "2025-05-23T08:52:22.077000",
              "content": "<p>No worries … both of you have been doing great work in this competition and putting a lot of effort , so just wanted to caution … </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3214068,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-05-30T22:03:38.980000",
              "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> You have not shared publicly the code you commented as follows:</p>\n<blockquote>\n  <p>As you have seen in the code, Finetuning had no change in model or data, the only difference was that I copied your model code and placed it my notebook, which had changes in:</p>\n  <p>Learning Rate and Global Batch Size<br>\n     The way EMA is applied<br>\n    The finetuning was short, like 10-20 epochs, then the score was converging slowly so I stopped the training<br>\n  and<br>\n   I am just using a different library to apply EMA which I have been using for an year,</p>\n</blockquote>\n<p>You shared a copy of <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> notebook but it is not the one you used to finetune the models.</p>\n<p></p>\n<p>I react because it is not fine to claim you shared your code publicly. Please don't lie to us.</p>\n<p>I don't want to shadow that you shared your improved weights. This is great, thanks for that.</p>\n<p>Edit: I update this because there was no private sharing.</p>",
              "votes": -1,
              "replies": []
            },
            {
              "id": 3214089,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2025-05-30T23:18:06.680000",
              "content": "<p>Totally agree <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>. The right way is to make the notebook public before sharing - though I cant control what is shared with me. The only code I have seen is <a href=\"https://www.kaggle.com/code/harshitsheoran/convnext-baseline/\" target=\"_blank\">here</a></p>\n<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> Would be great if you could give more details, but understand if you would rather not 🙂</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3214112,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-05-31T00:58:07.477000",
              "content": "<p>Great, then there is no private sharing. I'll edit my previous reply then.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3214140,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2025-05-31T02:26:57.960000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> !</p>\n<p>I am sorry but I think I am misunderstanding the quote</p>\n<blockquote>\n  <p>It is not fine to claim that you shared your code publicly</p>\n</blockquote>\n<p>Are we talking about the code that was shared first only to bartley? Yes, that was an unintentional mistake I made as there was no new code in that notebook so I failed to think it through before sharing it and I opened that code to everyone as soon as that mistake was pointed out to show that there was no new code in that notebook.</p>\n<p>Sorry, I re-read your message from the start:<br>\nTalking about the code that was used to finetune (it was finetuned on local machine), where I said \"I will say what the finetuning was\", although I meant it literally, I have no issues making the code public because there is nothing good in that, I believe that this great work can be pushed much further than my dumb finetune.</p>\n<p>By \"As you have seen in the code\" I meant that as everyone has noticed in the inference code that I did not change any data or model so inference remains the same</p>\n<p>Edit: Updated the original comment with the link to the code that was used for finetuning</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3214297,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-05-31T08:25:39.283000",
              "content": "<p>Awesome, thanks.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3214069,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2025-05-30T22:06:38.087000",
      "content": "<p>There is a lot to learn from your notebook, thanks for sharing.</p>\n<p>I never used mmap in numpy, this is really neat. Not sure who first share dit here, but kudos to whom did it.</p>\n<p>I still have to dig into how you modified convnext!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3212157,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2025-05-29T11:54:01.607000",
      "content": "<p>Based on past few days experimentation, I realized how strong the baseline <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> has provided.  There are lot of tricks that are implemented in the network, vanilla network would not perform this well even with out of the box encoder and decoder ! Great work!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3206983,
      "author_name": "Siming Shan",
      "author_url": "",
      "post_date": "2025-05-22T04:10:19.450000",
      "content": "<p>Huge respect! Although it is honestly kinda depressing to see my ranking drop from 14 -&gt; 150+ in two days.😂</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3206224,
      "author_name": "hyd",
      "author_url": "",
      "post_date": "2025-05-21T03:58:34.537000",
      "content": "<p>Show biggest respect to you!</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 3214010,
      "author_name": "William Wu Chengyuan",
      "author_url": "",
      "post_date": "2025-05-30T18:36:35.933000",
      "content": "<p>Excellent notebook, thanks for sharing!</p>\n<p>Just 1 question, how did you train on \"1x4090\" GPU?</p>\n<p>The code seems to require \"Requires &gt;= 2 GPUs with CUDA enabled.\"</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3214086,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-05-30T23:03:48.723000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/williamwu88\" target=\"_blank\">@williamwu88</a>, good point. You will have to modify the training loop for the script to work with 1 GPU.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3212174,
      "author_name": "Med Dhifallah",
      "author_url": "",
      "post_date": "2025-05-29T12:31:38.427000",
      "content": "<p>Can someone please explain why was the stride custom changed in this case? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3212189,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-05-29T13:00:33.073000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/meddhifallah\" target=\"_blank\">@meddhifallah</a>, does <a href=\"https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200035\" target=\"_blank\">this comment</a> help?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3210285,
      "author_name": "habibi9901",
      "author_url": "",
      "post_date": "2025-05-27T02:09:52.153000",
      "content": "<p>That's truly amazing. I'm learning a lot.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3210078,
      "author_name": "Stanislav Sobolev",
      "author_url": "",
      "post_date": "2025-05-26T17:31:44.220000",
      "content": "<p>Looks truly awesome </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3208519,
      "author_name": "Nihant At",
      "author_url": "",
      "post_date": "2025-05-24T09:22:36.323000",
      "content": "<p>That's really great!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3207893,
      "author_name": "kasey644",
      "author_url": "",
      "post_date": "2025-05-23T11:42:16.180000",
      "content": "<p>it will be very helpful in nova days!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3207856,
      "author_name": "Prajjwal Pratap Singh",
      "author_url": "",
      "post_date": "2025-05-23T10:32:20.167000",
      "content": "<p>hi <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> couple of quick thoughts/questions, and I'd love to hear your perspective if you have time:</p>\n<p>Resolution scaling: Did you observe any significant performance changes or convergence issues when training on the full resolution compared to down sampled versions? Given the computational cost, I just wanted to understanding the trade-off<br>\nLoss function Insights: What loss function did you primarily use for this ConvNeXt model? Did you explore any alternatives, perhaps something more tailored to the characteristics of geophysical waveforms or a hybrid loss?<br>\nFuture Directions: Based on these results, where do you see the most immediate opportunities for improvement? Perhaps exploring different attention mechanisms, a deeper understanding of the \"B\" datasets, or specialized data augmentation techniques?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3207918,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-05-23T12:32:27.870000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/bruhjjwal\" target=\"_blank\">@bruhjjwal</a>, thanks for the comment.</p>\n<ol>\n<li><p>Using the full resolution data helped significantly.</p></li>\n<li><p>L1 loss (have not tried other loss functions yet)</p></li>\n<li><p>Its hard to tell where to focus efforts at the moment. Currently I am playing with batch size / learning rate combinations. In addition, I would like to have a fully converged model before analyzing out of fold predictions. </p></li>\n</ol>\n<p>Would recommend reading <a href=\"https://www.kaggle.com/code/tpmeli/beyond-mae-depth-curves-spectra-and-residuals/comments#3203025\" target=\"_blank\">this</a> thread. <a href=\"https://www.kaggle.com/hoffmanns\" target=\"_blank\">@hoffmanns</a> makes some good points!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3207848,
      "author_name": "Erich von Mainstein",
      "author_url": "",
      "post_date": "2025-05-23T10:16:42.233000",
      "content": "<p>What batch size do you train the model with? 256 or 16?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3207913,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-05-23T12:23:37.177000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/vonmainstein\" target=\"_blank\">@vonmainstein</a>, I used a batch size of 16 for this model.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3207633,
      "author_name": "Hima Bindu Marpini",
      "author_url": "",
      "post_date": "2025-05-23T04:36:22.677000",
      "content": "<p>Huge Respect!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3206499,
      "author_name": "Med Dhifallah",
      "author_url": "",
      "post_date": "2025-05-21T11:58:55",
      "content": "<p>HUGE RESPECT!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3206193,
      "author_name": "guo dashuai",
      "author_url": "",
      "post_date": "2025-05-21T02:42:39.023000",
      "content": "<p>Is the main improvement due to the use of full-resolution data?🥰</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3206195,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-05-21T02:51:56.513000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/guoooooooss\" target=\"_blank\">@guoooooooss</a>, I think so. Though this was my first experiment using the full-resolution data.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3206200,
              "author_name": "guo dashuai",
              "author_url": "",
              "post_date": "2025-05-21T02:57:29.917000",
              "content": "<p>thank you very much!🥰</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3207695,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-23T06:22:24.473000",
      "content": "<p>Nice work! Appreciate so much</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3206070,
      "author_name": "MGöksu",
      "author_url": "",
      "post_date": "2025-05-20T19:28:02.483000",
      "content": "<p>Bro casually messing up the LB with \"Baseline\" and \"Starter\" notebooks :)</p>\n<p>Jokes aside, I love seeing your notebooks popping up.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3206580,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2025-05-21T14:26:12.983000",
      "content": "<p>nice sharing 👍<br>\nDoes pretrained encoder help?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3207057,
      "author_name": "Shane Wang",
      "author_url": "",
      "post_date": "2025-05-22T06:54:42.817000",
      "content": "<p>Beautiful LB.<br>\nIn this match, we achieved communism.<br>\nHuge respect to Bartley. Love you bro!</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3207055,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-22T06:53:18.650000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3206865,
      "author_name": "Sudo",
      "author_url": "",
      "post_date": "2025-05-21T21:17:51.613000",
      "content": "<p>Great work :) :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3206771,
      "author_name": "Taha Saleem",
      "author_url": "",
      "post_date": "2025-05-21T18:43:15.097000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3206031": "Here is another baseline trained on the full-resolution data. Like my previous notebook, I provide a few trained checkpoints and leave lots of room for customization.\n\nNotebook [here](https://www.kaggle.com/code/brendanartley/convnext-baseline)\nWeights [here](https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72)\n\nHere are the CV scores for this approach. Cheers!\n\n```python\n+--------------+--------+\n| Dataset      | Score  |\n+--------------+--------+\n| CurveFault_A |   6.07 |\n| CurveFault_B |  92.54 |\n| CurveVel_A   |  15.07 |\n| CurveVel_B   |  53.67 |\n| FlatFault_A  |   4.32 |\n| FlatFault_B  |  37.60 |\n| FlatVel_A    |   2.62 |\n| FlatVel_B    |  12.89 |\n| Style_A      |  37.38 |\n| Style_B      |  57.70 |\n+--------------+--------+\n```",
    "3206525": "A simple regression model shows that the LB score of the next \"baseline\" Bartley model, due out next week, is less than 20. Tremble, leaders!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291846%2Fbbda54814df51697ff7445fdb8482044%2FBartley-LB.png?generation=1747831623968838&alt=media)",
    "3206037": "Loss curves below. Each model takes ~10hrs for 50 epochs on 1×4090 with optimizations (e.g., torch.compile(), fused optimizer, increased num_workers, mixed precision, etc.).\n\n![WandbLogs](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F893d5096069da688ce1249ab739ef7de%2Flogs_wandb.JPG?generation=1747767254477184&alt=media)",
    "3207307": "Hi @brendanartley \n\nAwesome Open!, I am sharing a little bit more finetuned models of yours\n\nhttps://www.kaggle.com/datasets/harshitsheoran/simple-further-finetuned-bartley-open-models\n\nYou can add them to your notebook if you want ( don't want the credit, the hard work was yours ), I have shared you the notebook that would use these weights, the link is at\nhttps://www.kaggle.com/code/harshitsheoran/convnext-baseline/\n\nAs long as there is no leak that is happening in my finetuning, it's CV 0.289 :)\n\nIf you do choose to use these models, ofc, I will say what the finetuning was :D\n\nUpdate:\nHere is my exact code that I used to finetune the models:\n\nhttps://www.kaggle.com/code/harshitsheoran/recipe-simple-further-finetuned-bartley-models\n\nHave a good day!",
    "3214069": "There is a lot to learn from your notebook, thanks for sharing.\n\nI never used mmap in numpy, this is really neat. Not sure who first share dit here, but kudos to whom did it.\n\nI still have to dig into how you modified convnext!",
    "3212157": "Based on past few days experimentation, I realized how strong the baseline @brendanartley has provided.  There are lot of tricks that are implemented in the network, vanilla network would not perform this well even with out of the box encoder and decoder ! Great work!",
    "3206983": "Huge respect! Although it is honestly kinda depressing to see my ranking drop from 14 -> 150+ in two days.😂",
    "3206224": "Show biggest respect to you!",
    "3214010": "Excellent notebook, thanks for sharing!\n\nJust 1 question, how did you train on \"1x4090\" GPU?\n\nThe code seems to require \"Requires >= 2 GPUs with CUDA enabled.\"",
    "3212174": "Can someone please explain why was the stride custom changed in this case? ",
    "3210285": "That's truly amazing. I'm learning a lot.",
    "3210078": "Looks truly awesome ",
    "3208519": "That's really great!!",
    "3207893": "it will be very helpful in nova days!",
    "3207856": "hi @brendanartley couple of quick thoughts/questions, and I'd love to hear your perspective if you have time:\n\nResolution scaling: Did you observe any significant performance changes or convergence issues when training on the full resolution compared to down sampled versions? Given the computational cost, I just wanted to understanding the trade-off\nLoss function Insights: What loss function did you primarily use for this ConvNeXt model? Did you explore any alternatives, perhaps something more tailored to the characteristics of geophysical waveforms or a hybrid loss?\nFuture Directions: Based on these results, where do you see the most immediate opportunities for improvement? Perhaps exploring different attention mechanisms, a deeper understanding of the \"B\" datasets, or specialized data augmentation techniques?",
    "3207848": "What batch size do you train the model with? 256 or 16?\n",
    "3207633": "Huge Respect!",
    "3206499": "HUGE RESPECT!",
    "3206193": "Is the main improvement due to the use of full-resolution data?🥰",
    "3207695": "Nice work! Appreciate so much\n",
    "3206070": "Bro casually messing up the LB with \"Baseline\" and \"Starter\" notebooks :)\n\nJokes aside, I love seeing your notebooks popping up.",
    "3206580": "nice sharing 👍\nDoes pretrained encoder help?",
    "3207057": "Beautiful LB.\nIn this match, we achieved communism.\nHuge respect to Bartley. Love you bro!",
    "3207055": "",
    "3206865": "Great work :) :)",
    "3206771": "Thanks for sharing !\n"
  }
}