{
  "id": 582785,
  "title": "CAFormer Improved - [CV 24.2 LB 28.8]",
  "url": "/competitions/waveform-inversion/discussion/582785",
  "author_name": "Bartley",
  "post_date": "2025-06-02T18:38:52.268000",
  "votes": 78,
  "comment_count": 70,
  "views": 0,
  "content": "<p>Happy to share an improved full resolution model. This model uses a CAFormer encoder with an improved decoder (pixel shuffle, SCSE, intermediate convolutions, etc.).</p>\n<p>Notebook <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">here</a><br>\nDataset <a href=\"https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72\" target=\"_blank\">here</a></p>\n<pre><code>+--------------+----------+\n|||\n+--------------+----------+\n|||\n|||\n|||\n|||\n|||\n|||\n|||\n|||\n|||\n|||\n+--------------+----------+\n|||\n+--------------+----------+\n</code></pre>\n<p>This will be my last notebook in this competition, but I am hopeful that we can push this architecture further as a community!</p>\n<p>Happy Kaggling 😄</p>",
  "messages": [
    {
      "id": 3215873,
      "postDate": "2025-06-02T18:38:52.270Z",
      "content": "<p>Happy to share an improved full resolution model. This model uses a CAFormer encoder with an improved decoder (pixel shuffle, SCSE, intermediate convolutions, etc.).</p>\n<p>Notebook <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">here</a><br>\nDataset <a href=\"https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72\" target=\"_blank\">here</a></p>\n<pre><code>+--------------+----------+\n|||\n+--------------+----------+\n|||\n|||\n|||\n|||\n|||\n|||\n|||\n|||\n|||\n|||\n+--------------+----------+\n|||\n+--------------+----------+\n</code></pre>\n<p>This will be my last notebook in this competition, but I am hopeful that we can push this architecture further as a community!</p>\n<p>Happy Kaggling 😄</p>",
      "rawMarkdown": "Happy to share an improved full resolution model. This model uses a CAFormer encoder with an improved decoder (pixel shuffle, SCSE, intermediate convolutions, etc.).\n\nNotebook [here](https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved)\nDataset [here](https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72)\n\n```\n+--------------+----------+\n| Dataset      | Caformer |\n+--------------+----------+\n| CurveFault_A |     4.00 |\n| CurveFault_B |    71.06 |\n| CurveVel_A   |     9.19 |\n| CurveVel_B   |    38.60 |\n| FlatFault_A  |     2.58 |\n| FlatFault_B  |    24.29 |\n| FlatVel_A    |     1.31 |\n| FlatVel_B    |     7.20 |\n| Style_A      |    34.34 |\n| Style_B      |    48.90 |\n+--------------+----------+\n| Overall      |    24.15 |\n+--------------+----------+\n```\n\nThis will be my last notebook in this competition, but I am hopeful that we can push this architecture further as a community!\n\nHappy Kaggling 😄",
      "votes": 77
    },
    {
      "id": 3217147,
      "postDate": "2025-06-04T15:15:24.837Z",
      "content": "<p>\"this pipeline takes far too long to converge.\"</p>\n<ul>\n<li>you can classify the input type ( 3 families—Vel, Fault, and Style) for the test with a classifier (on the predicted velocity)</li>\n<li>then you can train seperate models for each type (this happens for most paper)</li>\n<li>if you have many machines, you can train in parallel</li>\n</ul>",
      "rawMarkdown": "\"this pipeline takes far too long to converge.\"\n- you can classify the input type ( 3 families—Vel, Fault, and Style) for the test with a classifier (on the predicted velocity)\n- then you can train seperate models for each type (this happens for most paper)\n- if you have many machines, you can train in parallel",
      "votes": 8,
      "replies": [
        {
          "id": 3217172,
          "postDate": "2025-06-04T15:49:40.440Z",
          "content": "<p><a href=\"https://arxiv.org/pdf/2412.19510v1\" target=\"_blank\">https://arxiv.org/pdf/2412.19510v1</a><br>\nParameter Efficient Fine-Tuning for Deep Learning-Based Full-Waveform Inversion</p>\n<p>try : Full fine-tuning(FFT-PFM) and LoRA-PFM ?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F44a7cfd8c762cf2da6b321f4a1c1686f%2FSelection_999(8225).png?generation=1749052179015479&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "https://arxiv.org/pdf/2412.19510v1\nParameter Efficient Fine-Tuning for Deep Learning-Based Full-Waveform Inversion\n\ntry : Full fine-tuning(FFT-PFM) and LoRA-PFM ?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F44a7cfd8c762cf2da6b321f4a1c1686f%2FSelection_999(8225).png?generation=1749052179015479&alt=media)",
          "votes": 4,
          "replies": [
            {
              "id": 3219822,
              "postDate": "2025-06-08T10:57:39.090Z",
              "content": "<p>From a quick glance, the results of that paper look subpar to any of the current LB models. Seems convoluted and unintuitive to me. At least do a proper MOE model… Happy to be disproved</p>",
              "rawMarkdown": "From a quick glance, the results of that paper look subpar to any of the current LB models. Seems convoluted and unintuitive to me. At least do a proper MOE model... Happy to be disproved",
              "votes": 2
            },
            {
              "id": 3219834,
              "postDate": "2025-06-08T11:13:00.057Z",
              "content": "<blockquote>\n  <p>From a quick glance, the results of that paper look subpar to any of the current LB models.  </p>\n</blockquote>\n<p>Well, if we can't get better results than non-Kagglers, what are we even doing here 😆</p>",
              "rawMarkdown": "> From a quick glance, the results of that paper look subpar to any of the current LB models.  \n\nWell, if we can't get better results than non-Kagglers, what are we even doing here 😆",
              "votes": 4
            },
            {
              "id": 3220620,
              "postDate": "2025-06-09T15:50:30.773Z",
              "content": "<p>But it's pretty rare to see a Kaggle competition where methods from papers in the past few years can't even beat some brute-force training of a public timm model. Kinda makes you wonder if it's even worth continuing to research this…</p>",
              "rawMarkdown": "But it's pretty rare to see a Kaggle competition where methods from papers in the past few years can't even beat some brute-force training of a public timm model. Kinda makes you wonder if it's even worth continuing to research this..."
            },
            {
              "id": 3220658,
              "postDate": "2025-06-09T17:14:03.743Z",
              "content": "<p>I always wonder - is this frowned upon in research to use pretrained model ? Or is it because the goal of research is to build networks from scratch ?  (I kaggle as hobby and not from ML/Data Science background) so it is a genuine question …</p>",
              "rawMarkdown": "I always wonder - is this frowned upon in research to use pretrained model ? Or is it because the goal of research is to build networks from scratch ?  (I kaggle as hobby and not from ML/Data Science background) so it is a genuine question ..."
            },
            {
              "id": 3220677,
              "postDate": "2025-06-09T17:35:34.087Z",
              "content": "<blockquote>\n  <p>Kinda makes you wonder if it's even worth continuing to research this…</p>\n</blockquote>\n<p>Why wouldn't it? The novelty is in defining the problem, creating data, applying, verifying etc. I dare to say, few researchers can get competitive results against top Kagglers, on almost any dataset that is not trivial (i.e. small/simple). But we still need researchers 😄</p>\n<blockquote>\n  <p>But it's pretty rare to see a Kaggle competition where methods from papers in the past few years can't even beat some brute-force training of a public timm model  </p>\n</blockquote>\n<p>It's far less trivial than what you imply. Without a 'certain' Kaggler, most people here would probably still be stuck somewhere at ~LB50+</p>",
              "rawMarkdown": "> Kinda makes you wonder if it's even worth continuing to research this…\n\nWhy wouldn't it? The novelty is in defining the problem, creating data, applying, verifying etc. I dare to say, few researchers can get competitive results against top Kagglers, on almost any dataset that is not trivial (i.e. small/simple). But we still need researchers 😄\n\n>But it's pretty rare to see a Kaggle competition where methods from papers in the past few years can't even beat some brute-force training of a public timm model  \n\nIt's far less trivial than what you imply. Without a 'certain' Kaggler, most people here would probably still be stuck somewhere at ~LB50+",
              "votes": 1
            },
            {
              "id": 3220712,
              "postDate": "2025-06-09T18:42:49.967Z",
              "content": "<blockquote>\n  <p>some brute-force training of a public timm model.</p>\n</blockquote>\n<p>If this is what you think <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> notebooks are about then you miss a lot. I suggest reading his code. You'll see that he did a bit more than using a timm model.</p>",
              "rawMarkdown": "> some brute-force training of a public timm model.\n\nIf this is what you think @brendanartley notebooks are about then you miss a lot. I suggest reading his code. You'll see that he did a bit more than using a timm model.",
              "votes": 4
            },
            {
              "id": 3220721,
              "postDate": "2025-06-09T19:15:26.503Z",
              "content": "<blockquote>\n  <p>wonder if it's even worth continuing to research</p>\n</blockquote>\n<p>Alex Krizhevsky is probably more a \"researcher\" than a \"Kaggler\".  But he basically kicked everyone's ass when he did a competition in 2012.</p>\n<p>It's not an exaggeration to say that because of that event, Alex brought neural network back to life after decades of it being shunned away, and the rest was history.</p>\n<p>Or let's take the researcher Tianqi Chen, who developed practical gradient boosting, which I don't think any Kaggler today can live without.</p>\n<p>I think it's not only worth it but in fact <em>essential</em> to keep doing research.</p>",
              "rawMarkdown": ">wonder if it's even worth continuing to research\n\nAlex Krizhevsky is probably more a \"researcher\" than a \"Kaggler\".  But he basically kicked everyone's ass when he did a competition in 2012.\n\nIt's not an exaggeration to say that because of that event, Alex brought neural network back to life after decades of it being shunned away, and the rest was history.\n\nOr let's take the researcher Tianqi Chen, who developed practical gradient boosting, which I don't think any Kaggler today can live without.\n\nI think it's not only worth it but in fact *essential* to keep doing research.",
              "votes": 3
            },
            {
              "id": 3220778,
              "postDate": "2025-06-10T00:35:19.477Z",
              "content": "<p>This topic can easily spiral into a much broader discussion—no one would dare deny the importance of research in general. That’s why I originally limited the scope of my comment to “papers from recent years related to this specific competition.”</p>\n<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> did a great job with the task, but I wouldn’t consider what he did to be traditional research in this specific field. It’s more like applying a general-purpose method to this domain.</p>\n<p>But if I were a researcher in this field and saw these results, I’d probably cough up blood and die.</p>",
              "rawMarkdown": "This topic can easily spiral into a much broader discussion—no one would dare deny the importance of research in general. That’s why I originally limited the scope of my comment to “papers from recent years related to this specific competition.”\n\n@brendanartley did a great job with the task, but I wouldn’t consider what he did to be traditional research in this specific field. It’s more like applying a general-purpose method to this domain.\n\nBut if I were a researcher in this field and saw these results, I’d probably cough up blood and die.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3217142,
      "postDate": "2025-06-04T15:12:18.253Z",
      "content": "<p>a trick is use fp32  in input and model for inference. you will get something like +0.0025 in lb (?)</p>",
      "rawMarkdown": "a trick is use fp32  in input and model for inference. you will get something like +0.0025 in lb (?)",
      "votes": 8,
      "replies": [
        {
          "id": 3217254,
          "postDate": "2025-06-04T18:02:46.680Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nhow you do that? I tried use <code>model = model.to(torch.float)</code> and test input with fp32 in <code>/kaggle/input/waveform-inversion/test</code> but seem nothing improved.</p>",
          "rawMarkdown": "@hengck23 \nhow you do that? I tried use `model = model.to(torch.float)` and test input with fp32 in `/kaggle/input/waveform-inversion/test` but seem nothing improved.",
          "votes": 1,
          "replies": [
            {
              "id": 3217267,
              "postDate": "2025-06-04T18:42:03.560Z",
              "content": "<p>Disable auto cast. Lb score is the same (due to decimal improvement only)but better ranking</p>",
              "rawMarkdown": "Disable auto cast. Lb score is the same (due to decimal improvement only)but better ranking",
              "votes": 5
            },
            {
              "id": 3217273,
              "postDate": "2025-06-04T18:50:48.220Z",
              "content": "<p>I will try it. Thanks.</p>",
              "rawMarkdown": "I will try it. Thanks.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3220406,
          "postDate": "2025-06-09T08:55:08.300Z",
          "content": "<p>Did the score improve during your local testing?</p>",
          "rawMarkdown": "Did the score improve during your local testing?"
        }
      ]
    },
    {
      "id": 3216127,
      "postDate": "2025-06-03T06:36:51.947Z",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Have you tried re-training convnext large with the same settings as caformer i.e. 150 epochs + improved decoder + cosine scheduler ? <br>\nCaformer training time was 2x longer than convnext large so I'm not sure it's worth the extra cost.</p>",
      "rawMarkdown": "@brendanartley Have you tried re-training convnext large with the same settings as caformer i.e. 150 epochs + improved decoder + cosine scheduler ? \nCaformer training time was 2x longer than convnext large so I'm not sure it's worth the extra cost.",
      "votes": 5,
      "replies": [
        {
          "id": 3216334,
          "postDate": "2025-06-03T12:30:26.557Z",
          "content": "<p>Good question <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a>. </p>\n<p>I have not tried <code>convnext large</code> due to compute limits, but 2x faster training sounds nice!</p>",
          "rawMarkdown": "Good question @andy2709. \n\nI have not tried `convnext large` due to compute limits, but 2x faster training sounds nice!"
        }
      ]
    },
    {
      "id": 3219000,
      "postDate": "2025-06-07T03:36:08.287Z",
      "content": "<p>this is segmentation head in your code:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .conv = nn.Conv2d(\n            in_channels, out_channels, kernel_size=kernel_size,\n            padding=kernel_size // \n        )\n        .upsample = UpSample(\n            spatial_dims=,\n            in_channels=out_channels,\n            out_channels=out_channels,\n            scale_factor=scale_factor,\n            mode=mode,\n        )\n\n     ():\n        x = .conv(x)\n        x = .upsample(x)\n         x\n</code></pre>\n<p>i haven't done in-depth experiment yet, but below are from my experiences:</p>\n<ul>\n<li>this is regression problem, so value of the prediction is important</li>\n<li>do not use upsample as last layer. you will be bounded by up scaling error.</li>\n<li>use original input data resolution  (or 72x72 sampled version) in final decoder layers</li>\n<li>the last prediction module should be 1x1 conv  (not even 3x3)</li>\n</ul>\n<p>```<br>\ne.g. something like</p>\n<p>last_feature = lastdecoder(input = from previous decoder, skip = original input)<br>\npredicted_velocity = prediction(last_feature)</p>\n<p>where  prediction= nn.Seqential(<br>\n….<br>\nnn.Conv2d( 1x1 kernel)<br>\n)<br>\n``</p>",
      "rawMarkdown": "this is segmentation head in your code:\n\n```\n\n\nclass SegmentationHead2d(nn.Module):\n    def __init__(\n            self,\n            in_channels,\n            out_channels,\n            scale_factor: tuple[int] = (2, 2),\n            kernel_size: int = 3,\n            mode: str = \"nontrainable\",\n    ):\n        super().__init__()\n        self.conv = nn.Conv2d(\n            in_channels, out_channels, kernel_size=kernel_size,\n            padding=kernel_size // 2\n        )\n        self.upsample = UpSample(\n            spatial_dims=2,\n            in_channels=out_channels,\n            out_channels=out_channels,\n            scale_factor=scale_factor,\n            mode=mode,\n        )\n\n    def forward(self, x):\n        x = self.conv(x)\n        x = self.upsample(x)\n        return x\n\n```\n\ni haven't done in-depth experiment yet, but below are from my experiences:\n- this is regression problem, so value of the prediction is important\n- do not use upsample as last layer. you will be bounded by up scaling error.\n- use original input data resolution  (or 72x72 sampled version) in final decoder layers\n- the last prediction module should be 1x1 conv  (not even 3x3)\n\n```\ne.g. something like\n\nlast_feature = lastdecoder(input = from previous decoder, skip = original input)\npredicted_velocity = prediction(last_feature)\n\nwhere  prediction= nn.Seqential(\n....\nnn.Conv2d( 1x1 kernel)\n)\n``",
      "votes": 3,
      "replies": [
        {
          "id": 3219011,
          "postDate": "2025-06-07T04:06:05.330Z",
          "content": "<p>You are bound by upscaling error only if you upsample with something like bilinear. If you upsample with ConvTranspose2d there is no bound (although it might be harder to learn than upscaling to full resolution before final layer, but then it's trade-off against compute and time, no bound).<br>\nTheoretically we can compress all the seis into a single vector and then predict from this vector each pixel - there is nothing imposing any bound.</p>",
          "rawMarkdown": "You are bound by upscaling error only if you upsample with something like bilinear. If you upsample with ConvTranspose2d there is no bound (although it might be harder to learn than upscaling to full resolution before final layer, but then it's trade-off against compute and time, no bound).\nTheoretically we can compress all the seis into a single vector and then predict from this vector each pixel - there is nothing imposing any bound.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3216409,
      "postDate": "2025-06-03T14:38:38.930Z",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <br>\nWhen turning a Train, the IO seems to be a problem(?). What is your countermeasure for this?</p>",
      "rawMarkdown": "@brendanartley \nWhen turning a Train, the IO seems to be a problem(?). What is your countermeasure for this?",
      "votes": 3,
      "replies": [
        {
          "id": 3216437,
          "postDate": "2025-06-03T15:22:55.247Z",
          "content": "<p>To be precise, the speed decreases around the 800-step mark.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3790528%2F06453f06b67c9a8dede330ce4e0d66b2%2F.png?generation=1748964167653260&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "To be precise, the speed decreases around the 800-step mark.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3790528%2F06453f06b67c9a8dede330ce4e0d66b2%2F.png?generation=1748964167653260&alt=media)",
          "votes": 1,
          "replies": [
            {
              "id": 3216726,
              "postDate": "2025-06-04T03:37:23.660Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3216093,
      "postDate": "2025-06-03T05:27:23.170Z",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Thanks for sharing one more milestone very hard to break <strong>struggle to reach &lt;100</strong> --&gt; <strong>69.7</strong> --&gt; <strong>33.2</strong> --&gt; <strong>28.8</strong>, could you share <a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500\" target=\"_blank\">CV scores</a> of training 20-25 epochs of CurveFault-B ( your cv and training logs if possible )</p>",
      "rawMarkdown": "@brendanartley Thanks for sharing one more milestone very hard to break **struggle to reach <100** --> **69.7** --> **33.2** --> **28.8**, could you share [CV scores](https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500) of training 20-25 epochs of CurveFault-B ( your cv and training logs if possible )",
      "votes": 4
    },
    {
      "id": 3224682,
      "postDate": "2025-06-15T10:12:01.490Z",
      "content": "<p>This is very interesting</p>",
      "rawMarkdown": "This is very interesting",
      "votes": 1
    },
    {
      "id": 3223522,
      "postDate": "2025-06-13T12:44:36.783Z",
      "content": "<p>whats the upsample mode , pixel shuffle or deconv ??</p>",
      "rawMarkdown": "whats the upsample mode , pixel shuffle or deconv ??",
      "votes": 1,
      "replies": [
        {
          "id": 3224846,
          "postDate": "2025-06-15T15:02:57.960Z",
          "content": "<p>Read the code :D</p>",
          "rawMarkdown": "Read the code :D",
          "votes": 2
        }
      ]
    },
    {
      "id": 3217724,
      "postDate": "2025-06-05T10:59:35.717Z",
      "content": "<p>When you're testing model performance on a small order of magnitude, do you use CurveFault_B, and if so, what the results are, and then you use them for training on all datasets?</p>",
      "rawMarkdown": "When you're testing model performance on a small order of magnitude, do you use CurveFault_B, and if so, what the results are, and then you use them for training on all datasets?",
      "votes": 1,
      "replies": [
        {
          "id": 3219155,
          "postDate": "2025-06-07T08:13:20.073Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500</a></p>",
          "rawMarkdown": "https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500",
          "votes": 2
        }
      ]
    },
    {
      "id": 3217157,
      "postDate": "2025-06-04T15:33:31.020Z",
      "content": "<p>Thanks for sharing best model with full training….</p>",
      "rawMarkdown": "Thanks for sharing best model with full training....",
      "votes": 1
    },
    {
      "id": 3216520,
      "postDate": "2025-06-03T17:32:56.413Z",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">Bartley</a></p>\n<p>Ok - maybe its a lack of coffee or the fact that I am an old fart - but I think I read all the other comments several times and failed to find the string that leads us to the code for training this latest nice bit of work.  I tried re-using the previous train but that failed.</p>\n<p>Is it your intent to not share the training code for your last model?</p>\n<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/billfan88\" target=\"_blank\">@billfan88</a>, I appreciate your questions but please read other comments first. Cheers 🙂</p>\n</blockquote>",
      "rawMarkdown": "[Bartley](https://www.kaggle.com/brendanartley)\n\nOk - maybe its a lack of coffee or the fact that I am an old fart - but I think I read all the other comments several times and failed to find the string that leads us to the code for training this latest nice bit of work.  I tried re-using the previous train but that failed.\n\nIs it your intent to not share the training code for your last model?\n\n>Hi @billfan88, I appreciate your questions but please read other comments first. Cheers 🙂",
      "votes": 1,
      "replies": [
        {
          "id": 3216559,
          "postDate": "2025-06-03T18:31:55.850Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, <a href=\"https://www.kaggle.com/ttyn4519\" target=\"_blank\">@ttyn4519</a>, thanks for your comments.</p>\n<p>I did not include training code in this notebook because the model is trained outside of the Kaggle environment. I do not see issues when running locally.</p>",
          "rawMarkdown": "Hi @pcjimmmy, @ttyn4519, thanks for your comments.\n\nI did not include training code in this notebook because the model is trained outside of the Kaggle environment. I do not see issues when running locally.",
          "votes": 1,
          "replies": [
            {
              "id": 3216567,
              "postDate": "2025-06-03T18:37:43.663Z",
              "content": "<p>Yes - and I fork it outside of kaggle :)   That's the only place that multiple days of running are possible.  Your previous ran 10 days on one of my local machines.</p>\n<p>Are you telling me that you can re-use the previous train.py, util.py and the full cfg and it runs OK?  </p>",
              "rawMarkdown": "Yes - and I fork it outside of kaggle :)   That's the only place that multiple days of running are possible.  Your previous ran 10 days on one of my local machines.\n\nAre you telling me that you can re-use the previous train.py, util.py and the full cfg and it runs OK?  ",
              "votes": 1
            },
            {
              "id": 3216578,
              "postDate": "2025-06-03T18:45:13.180Z",
              "content": "<p>I use a slightly different training script, but the cfg is the same. It's hard to predict how the pipeline will act across different GPUs, CUDA versions, and environments, but it works for me 🤷‍♂️</p>",
              "rawMarkdown": "I use a slightly different training script, but the cfg is the same. It's hard to predict how the pipeline will act across different GPUs, CUDA versions, and environments, but it works for me 🤷‍♂️",
              "votes": 2
            },
            {
              "id": 3216633,
              "postDate": "2025-06-03T21:37:16.347Z",
              "content": "<p>So- in summary - your not sharing the full code for this last model?</p>",
              "rawMarkdown": "So- in summary - your not sharing the full code for this last model?",
              "votes": -1
            },
            {
              "id": 3220810,
              "postDate": "2025-06-10T02:24:29.347Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3216040,
      "postDate": "2025-06-03T03:24:04.390Z",
      "content": "<p>Thanks for sharing best model with full training - one more break through notebook and solution.<br>\nCould you share your training logs of these models after 50 epochs?</p>",
      "rawMarkdown": "Thanks for sharing best model with full training - one more break through notebook and solution.\nCould you share your training logs of these models after 50 epochs?\n",
      "votes": 1,
      "replies": [
        {
          "id": 3216066,
          "postDate": "2025-06-03T04:19:17.360Z",
          "content": "<p><a href=\"https://www.kaggle.com/waterjoe\" target=\"_blank\">@waterjoe</a> <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> here are the loss values. After 50 epochs training loss is ~21.5 and validation loss is 29.3.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F1900a9c815836e8b891baeed2ed69461%2FScreen%20Shot%202025-06-02%20at%209.16.55%20PM.png?generation=1748924280482744&amp;alt=media\" alt=\"Loss curves\"></p>",
          "rawMarkdown": "@waterjoe @seshurajup here are the loss values. After 50 epochs training loss is ~21.5 and validation loss is 29.3.\n\n![Loss curves](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F1900a9c815836e8b891baeed2ed69461%2FScreen%20Shot%202025-06-02%20at%209.16.55%20PM.png?generation=1748924280482744&alt=media)",
          "votes": 3
        }
      ]
    },
    {
      "id": 3215914,
      "postDate": "2025-06-02T20:38:25.347Z",
      "content": "<p>Was the model trained as in previous baselines?</p>",
      "rawMarkdown": "Was the model trained as in previous baselines?",
      "votes": 1,
      "replies": [
        {
          "id": 3215926,
          "postDate": "2025-06-02T21:28:33.530Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/xbar19\" target=\"_blank\">@xbar19</a> and <a href=\"https://www.kaggle.com/billfan88\" target=\"_blank\">@billfan88</a>, thanks for the comments.</p>\n<p>The training is very similar. The differences are that this model is trained for 150 epochs and uses a custom learning rate scheduler shown <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved?scriptVersionId=243291347&amp;cellId=4\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Hi @xbar19 and @billfan88, thanks for the comments.\n \nThe training is very similar. The differences are that this model is trained for 150 epochs and uses a custom learning rate scheduler shown [here](https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved?scriptVersionId=243291347&cellId=4)",
          "votes": 2,
          "replies": [
            {
              "id": 3215938,
              "postDate": "2025-06-02T21:48:06.603Z",
              "content": "<p>Thank you for your reply and your work!</p>",
              "rawMarkdown": "Thank you for your reply and your work!",
              "votes": 1
            },
            {
              "id": 3215946,
              "postDate": "2025-06-02T22:21:46.487Z",
              "content": "<p>How much time it took to train on 150 epochs? Damn this is one long training</p>",
              "rawMarkdown": "How much time it took to train on 150 epochs? Damn this is one long training",
              "votes": 4
            },
            {
              "id": 3215958,
              "postDate": "2025-06-02T23:25:38.563Z",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a>, it took 8 days (~192 hours) on a single 4090.</p>\n<p>You make a good point — this pipeline takes far too long to converge. I am hoping people can share ideas on how we can decrease this time!</p>",
              "rawMarkdown": "@shlomoron, it took 8 days (~192 hours) on a single 4090.\n\nYou make a good point — this pipeline takes far too long to converge. I am hoping people can share ideas on how we can decrease this time!",
              "votes": 5
            },
            {
              "id": 3215959,
              "postDate": "2025-06-02T23:29:23.150Z",
              "content": "<p>Holy <strong><em>*</em></strong> shit! 8 days…<br>\nFull 8*24 days? Damn</p>",
              "rawMarkdown": "Holy ******* shit! 8 days...\nFull 8\\*24 days? Damn",
              "votes": 4
            },
            {
              "id": 3215962,
              "postDate": "2025-06-02T23:43:49.617Z",
              "content": "<blockquote>\n  <p>this pipeline takes far too long to converge. </p>\n</blockquote>\n<p>MAE is always slow to converge. Alas.</p>",
              "rawMarkdown": "> this pipeline takes far too long to converge. \n\nMAE is always slow to converge. Alas.",
              "votes": 1
            },
            {
              "id": 3217497,
              "postDate": "2025-06-05T04:30:26.740Z",
              "content": "<p>It would take me ~8 days on a single 4090 to hit my current LB (17.1), the trainings take a long time to converge </p>",
              "rawMarkdown": "It would take me ~8 days on a single 4090 to hit my current LB (17.1), the trainings take a long time to converge ",
              "votes": 4
            },
            {
              "id": 3217741,
              "postDate": "2025-06-05T11:18:10.863Z",
              "content": "<p><strong>Magic of L1 --&gt; Run Long</strong> - Thanks for sharing <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>  <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>\n<ul>\n<li>it took 8 days (~192 hours) on a single 4090 -- <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> </li>\n<li>It would take me ~8 days on a single 4090 to hit my current LB (17.1), the trainings take a long time to converge <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </li>\n</ul>\n<hr>\n<blockquote>\n  <p>-- How to decided when to stop, most of epochs getting negligible loss improvement (small improvement - my experiments based on only CurveFault_B)<br>\n   -- some batches are too high even using clip the grads<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fc49c05fe040a84a4cb8bd2e1b3e1644e%2FScreenshot%202025-06-05%20at%204.46.22PM.png?generation=1749122195332342&amp;alt=media\" alt=\"\"></p>\n</blockquote>",
              "rawMarkdown": "**Magic of L1 --> Run Long** - Thanks for sharing @brendanartley  @harshitsheoran \n- it took 8 days (~192 hours) on a single 4090 -- @brendanartley \n- It would take me ~8 days on a single 4090 to hit my current LB (17.1), the trainings take a long time to converge @harshitsheoran \n\n--- \n\n> -- How to decided when to stop, most of epochs getting negligible loss improvement (small improvement - my experiments based on only CurveFault_B)\n -- some batches are too high even using clip the grads\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fc49c05fe040a84a4cb8bd2e1b3e1644e%2FScreenshot%202025-06-05%20at%204.46.22PM.png?generation=1749122195332342&alt=media)",
              "votes": 1
            },
            {
              "id": 3217866,
              "postDate": "2025-06-05T13:15:42.077Z",
              "content": "<p>Since there is a lot of information on 4090, I will share the results of V100x8 DDP.</p>\n<p>To reproduce this old notebook's convnext[CV 31.9 LB 36.4] locally, it took 2 days for 150 epochs of training, and I was able to obtain the result of [CV30.07, LB35.1].</p>\n<pre><code>=========================\nCurveFault_A    6.15\nCurveFault_B    86.72\nCurveVel_A      13.69\nCurveVel_B      48.81\nFlatFault_A     4.34\nFlatFault_B     33.73\nFlatVel_A       2.92\nFlatVel_B       12.11\nStyle_A         37.07\nStyle_B         55.18\n=========================\n\n</code></pre>",
              "rawMarkdown": "Since there is a lot of information on 4090, I will share the results of V100x8 DDP.\n\nTo reproduce this old notebook's convnext[CV 31.9 LB 36.4] locally, it took 2 days for 150 epochs of training, and I was able to obtain the result of [CV30.07, LB35.1].\n\n```\n=========================\nCurveFault_A    6.15\nCurveFault_B    86.72\nCurveVel_A      13.69\nCurveVel_B      48.81\nFlatFault_A     4.34\nFlatFault_B     33.73\nFlatVel_A       2.92\nFlatVel_B       12.11\nStyle_A         37.07\nStyle_B         55.18\n=========================\nVal MAE: 30.07\n=========================\n```",
              "votes": 7
            },
            {
              "id": 3220281,
              "postDate": "2025-06-09T05:17:50.173Z",
              "content": "<p>Could you please tell me whether the data you used was the original FP32 data? When I tried to reproduce Caformer’s training results using Kaggle’s FP16 open dataset, the cross-validation converged around 34, and I’m not sure if this is a dataset issue.</p>",
              "rawMarkdown": "Could you please tell me whether the data you used was the original FP32 data? When I tried to reproduce Caformer’s training results using Kaggle’s FP16 open dataset, the cross-validation converged around 34, and I’m not sure if this is a dataset issue.",
              "votes": 1
            },
            {
              "id": 3220288,
              "postDate": "2025-06-09T05:39:21.670Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/oigckko\" target=\"_blank\">@oigckko</a> <br>\nIf you have any questions regarding my above post \"Val MAE: 30.07\", I  hope this helps.</p>\n<p><strong>The training dataset remains unchanged from the public notebook.</strong></p>\n<pre><code>df= pd()\n...\np1 = os(, row)\n            p2 = os(, row(), , row())\n            p3 = os(, row)\n            p4 = os(, row(), , row())\n</code></pre>",
              "rawMarkdown": "Hi @oigckko \nIf you have any questions regarding my above post \"Val MAE: 30.07\", I  hope this helps.\n\n**The training dataset remains unchanged from the public notebook.**\n```\ndf= pd.read_csv(\"/kaggle/input/openfwi-preprocessed-72x72/folds.csv\")\n...\np1 = os.path.join(\"/kaggle/input/open-wfi-1/openfwi_float16_1/\", row[\"data_fpath\"])\n            p2 = os.path.join(\"/kaggle/input/open-wfi-1/openfwi_float16_1/\", row[\"data_fpath\"].split(\"/\")[0], \"*\", row[\"data_fpath\"].split(\"/\")[-1])\n            p3 = os.path.join(\"/kaggle/input/open-wfi-2/openfwi_float16_2/\", row[\"data_fpath\"])\n            p4 = os.path.join(\"/kaggle/input/open-wfi-2/openfwi_float16_2/\", row[\"data_fpath\"].split(\"/\")[0], \"*\", row[\"data_fpath\"].split(\"/\")[-1])\n```\n"
            }
          ]
        }
      ]
    },
    {
      "id": 3215969,
      "postDate": "2025-06-03T00:22:00.277Z",
      "content": "<p>Train section  Same as starter notebook?</p>",
      "rawMarkdown": "Train section  Same as starter notebook?",
      "votes": 2,
      "replies": [
        {
          "id": 3216322,
          "postDate": "2025-06-03T12:18:06.103Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/billfan88\" target=\"_blank\">@billfan88</a>, I appreciate your questions but please read other comments first. Cheers 🙂</p>",
          "rawMarkdown": "Hi @billfan88, I appreciate your questions but please read other comments first. Cheers 🙂"
        }
      ]
    },
    {
      "id": 3215917,
      "postDate": "2025-06-02T20:57:34.400Z",
      "content": "<p>Awesome!</p>\n<p>I wonder how many backbones you tried. You obviously have a good methodology to pick the ones you test.</p>",
      "rawMarkdown": "Awesome!\n\nI wonder how many backbones you tried. You obviously have a good methodology to pick the ones you test.",
      "votes": 2,
      "replies": [
        {
          "id": 3215933,
          "postDate": "2025-06-02T21:42:40.240Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, I have tried Convnext, Caformer, Hgnet, NextVIT and Hornet so far.</p>\n<p>My methodology is quite simple. First I try and modify the backbone for 1000 x 70 in as few steps as possible. Then, I follow <a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500\" target=\"_blank\">this</a> validation scheme. It may not be the most accurate, but it seems to work well given limited compute!</p>",
          "rawMarkdown": "Hi @cpmpml, I have tried Convnext, Caformer, Hgnet, NextVIT and Hornet so far.\n\nMy methodology is quite simple. First I try and modify the backbone for 1000 x 70 in as few steps as possible. Then, I follow [this](https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500) validation scheme. It may not be the most accurate, but it seems to work well given limited compute!\n",
          "votes": 7,
          "replies": [
            {
              "id": 3215939,
              "postDate": "2025-06-02T21:49:12.983Z",
              "content": "<p>Yes, I noticed you use a subset to validate your choice. This is a good idea.</p>",
              "rawMarkdown": "Yes, I noticed you use a subset to validate your choice. This is a good idea.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 3221540,
      "postDate": "2025-06-11T06:45:49.307Z",
      "content": "<p>Why is my train loss always higher than vaild loss during training, is this normal?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25710556%2F53e5d71b04b7954200dfa184996221ba%2F2025-06-11%20144329.png?generation=1749624341620051&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Why is my train loss always higher than vaild loss during training, is this normal?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25710556%2F53e5d71b04b7954200dfa184996221ba%2F2025-06-11%20144329.png?generation=1749624341620051&alt=media)",
      "replies": [
        {
          "id": 3221552,
          "postDate": "2025-06-11T07:03:30.613Z",
          "content": "<p><a href=\"https://www.kaggle.com/oliver34\" target=\"_blank\">@oliver34</a> its same for me too, but it got switch after few epochs when i tested on <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/580892#3221434\" target=\"_blank\">CurveFault_B</a> dataset experiments.</p>\n<p>-- How many epochs you are running of CAFormer? Did you run from backbone or from Bartley weights?</p>",
          "rawMarkdown": "@oliver34 its same for me too, but it got switch after few epochs when i tested on [CurveFault_B](https://www.kaggle.com/competitions/waveform-inversion/discussion/580892#3221434) dataset experiments.\n\n-- How many epochs you are running of CAFormer? Did you run from backbone or from Bartley weights?",
          "votes": 1,
          "replies": [
            {
              "id": 3221557,
              "postDate": "2025-06-11T07:11:46.630Z",
              "content": "<p>44epoch,I only used the weights of the backbone</p>",
              "rawMarkdown": "44epoch,I only used the weights of the backbone",
              "votes": 2
            },
            {
              "id": 3221558,
              "postDate": "2025-06-11T07:15:54.113Z",
              "content": "<p><a href=\"https://www.kaggle.com/oliver34\" target=\"_blank\">@oliver34</a> will you share your train/val loss for epochs 5, 10, 20, 30 ? -- which gpu - training finishing in 1h 25mins for epoch?</p>",
              "rawMarkdown": "@oliver34 will you share your train/val loss for epochs 5, 10, 20, 30 ? -- which gpu - training finishing in 1h 25mins for epoch?",
              "votes": 1
            },
            {
              "id": 3221562,
              "postDate": "2025-06-11T07:23:45.047Z",
              "content": "<p>They are 47.70, 40.59, 31.86, 28.70. I also encountered the situation you mentioned at CurveFault_B. --I'm using two 4090s and it seems like it takes 9 days to complete a 150epoch😥</p>",
              "rawMarkdown": "They are 47.70, 40.59, 31.86, 28.70. I also encountered the situation you mentioned at CurveFault_B. --I'm using two 4090s and it seems like it takes 9 days to complete a 150epoch😥",
              "votes": 1
            },
            {
              "id": 3221639,
              "postDate": "2025-06-11T09:46:46.287Z",
              "content": "<p>Hello, may I ask if you are using the fp32 data or the fp16 dataset processed by Kaggle?</p>",
              "rawMarkdown": "Hello, may I ask if you are using the fp32 data or the fp16 dataset processed by Kaggle?"
            },
            {
              "id": 3221666,
              "postDate": "2025-06-11T10:30:15.037Z",
              "content": "<p>the fp16 dataset processed by Kaggle</p>",
              "rawMarkdown": "the fp16 dataset processed by Kaggle",
              "votes": 1
            },
            {
              "id": 3224699,
              "postDate": "2025-06-15T10:45:50.537Z",
              "content": "<p><a href=\"https://www.kaggle.com/oliver34\" target=\"_blank\">@oliver34</a>  could you share cv scores of 40, 60, 80, 100? </p>",
              "rawMarkdown": "@oliver34  could you share cv scores of 40, 60, 80, 100? ",
              "votes": 1
            },
            {
              "id": 3224761,
              "postDate": "2025-06-15T12:45:32.147Z",
              "content": "<p>I would love to share them, 26.88, 24.58, 23.08, 22.11.</p>",
              "rawMarkdown": "I would love to share them, 26.88, 24.58, 23.08, 22.11.",
              "votes": 1
            },
            {
              "id": 3225362,
              "postDate": "2025-06-16T09:50:55.937Z",
              "content": "<p>With the \"<a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/583896\" target=\"_blank\">Tips to speed up training</a>\" optimizations, training can be completed within 2 days using 2×4090 </p>",
              "rawMarkdown": "With the \"[Tips to speed up training](https://www.kaggle.com/competitions/waveform-inversion/discussion/583896)\" optimizations, training can be completed within 2 days using 2×4090 ",
              "votes": 1
            },
            {
              "id": 3225364,
              "postDate": "2025-06-16T09:53:50.547Z",
              "content": "<p>I don’t think 4090 can train caformer for 150 epoch for 2 days</p>",
              "rawMarkdown": "I don’t think 4090 can train caformer for 150 epoch for 2 days",
              "votes": 1
            }
          ]
        },
        {
          "id": 3222445,
          "postDate": "2025-06-12T07:46:55.827Z",
          "content": "<p>batch-wise average vs overall average</p>",
          "rawMarkdown": "batch-wise average vs overall average",
          "votes": 1
        },
        {
          "id": 3223545,
          "postDate": "2025-06-13T13:28:12.613Z",
          "content": "<blockquote>\n  <p>my train loss always higher than vaild loss during training</p>\n</blockquote>\n<p>I also observe this.  I haven't observed a single epoch in which my val loss is worse than my train loss.  I'm still relatively early in my training epochs, though.  Glad to know I'm not the only one. 😀</p>",
          "rawMarkdown": ">my train loss always higher than vaild loss during training\n\nI also observe this.  I haven't observed a single epoch in which my val loss is worse than my train loss.  I'm still relatively early in my training epochs, though.  Glad to know I'm not the only one. 😀"
        }
      ]
    },
    {
      "id": 3221539,
      "postDate": "2025-06-11T06:45:11.713Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3225006,
      "postDate": "2025-06-15T19:37:03.183Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": 1
    },
    {
      "id": 3223364,
      "postDate": "2025-06-13T07:42:11.890Z",
      "content": "<p>Thanks for the information</p>",
      "rawMarkdown": "Thanks for the information"
    }
  ],
  "comments": [
    {
      "id": 3217147,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-04T15:15:24.837000",
      "content": "<p>\"this pipeline takes far too long to converge.\"</p>\n<ul>\n<li>you can classify the input type ( 3 families—Vel, Fault, and Style) for the test with a classifier (on the predicted velocity)</li>\n<li>then you can train seperate models for each type (this happens for most paper)</li>\n<li>if you have many machines, you can train in parallel</li>\n</ul>",
      "votes": 8,
      "replies": [
        {
          "id": 3217172,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-06-04T15:49:40.440000",
          "content": "<p><a href=\"https://arxiv.org/pdf/2412.19510v1\" target=\"_blank\">https://arxiv.org/pdf/2412.19510v1</a><br>\nParameter Efficient Fine-Tuning for Deep Learning-Based Full-Waveform Inversion</p>\n<p>try : Full fine-tuning(FFT-PFM) and LoRA-PFM ?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F44a7cfd8c762cf2da6b321f4a1c1686f%2FSelection_999(8225).png?generation=1749052179015479&amp;alt=media\" alt=\"\"></p>",
          "votes": 4,
          "replies": [
            {
              "id": 3219822,
              "author_name": "Sebastian Hoffmann",
              "author_url": "",
              "post_date": "2025-06-08T10:57:39.090000",
              "content": "<p>From a quick glance, the results of that paper look subpar to any of the current LB models. Seems convoluted and unintuitive to me. At least do a proper MOE model… Happy to be disproved</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3219834,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-06-08T11:13:00.057000",
              "content": "<blockquote>\n  <p>From a quick glance, the results of that paper look subpar to any of the current LB models.  </p>\n</blockquote>\n<p>Well, if we can't get better results than non-Kagglers, what are we even doing here 😆</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3220620,
              "author_name": "atom1231",
              "author_url": "",
              "post_date": "2025-06-09T15:50:30.773000",
              "content": "<p>But it's pretty rare to see a Kaggle competition where methods from papers in the past few years can't even beat some brute-force training of a public timm model. Kinda makes you wonder if it's even worth continuing to research this…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220658,
              "author_name": "Nirjhar Roy",
              "author_url": "",
              "post_date": "2025-06-09T17:14:03.743000",
              "content": "<p>I always wonder - is this frowned upon in research to use pretrained model ? Or is it because the goal of research is to build networks from scratch ?  (I kaggle as hobby and not from ML/Data Science background) so it is a genuine question …</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220677,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-06-09T17:35:34.087000",
              "content": "<blockquote>\n  <p>Kinda makes you wonder if it's even worth continuing to research this…</p>\n</blockquote>\n<p>Why wouldn't it? The novelty is in defining the problem, creating data, applying, verifying etc. I dare to say, few researchers can get competitive results against top Kagglers, on almost any dataset that is not trivial (i.e. small/simple). But we still need researchers 😄</p>\n<blockquote>\n  <p>But it's pretty rare to see a Kaggle competition where methods from papers in the past few years can't even beat some brute-force training of a public timm model  </p>\n</blockquote>\n<p>It's far less trivial than what you imply. Without a 'certain' Kaggler, most people here would probably still be stuck somewhere at ~LB50+</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3220712,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-06-09T18:42:49.967000",
              "content": "<blockquote>\n  <p>some brute-force training of a public timm model.</p>\n</blockquote>\n<p>If this is what you think <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> notebooks are about then you miss a lot. I suggest reading his code. You'll see that he did a bit more than using a timm model.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3220721,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2025-06-09T19:15:26.503000",
              "content": "<blockquote>\n  <p>wonder if it's even worth continuing to research</p>\n</blockquote>\n<p>Alex Krizhevsky is probably more a \"researcher\" than a \"Kaggler\".  But he basically kicked everyone's ass when he did a competition in 2012.</p>\n<p>It's not an exaggeration to say that because of that event, Alex brought neural network back to life after decades of it being shunned away, and the rest was history.</p>\n<p>Or let's take the researcher Tianqi Chen, who developed practical gradient boosting, which I don't think any Kaggler today can live without.</p>\n<p>I think it's not only worth it but in fact <em>essential</em> to keep doing research.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3220778,
              "author_name": "atom1231",
              "author_url": "",
              "post_date": "2025-06-10T00:35:19.477000",
              "content": "<p>This topic can easily spiral into a much broader discussion—no one would dare deny the importance of research in general. That’s why I originally limited the scope of my comment to “papers from recent years related to this specific competition.”</p>\n<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> did a great job with the task, but I wouldn’t consider what he did to be traditional research in this specific field. It’s more like applying a general-purpose method to this domain.</p>\n<p>But if I were a researcher in this field and saw these results, I’d probably cough up blood and die.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3217142,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-04T15:12:18.253000",
      "content": "<p>a trick is use fp32  in input and model for inference. you will get something like +0.0025 in lb (?)</p>",
      "votes": 8,
      "replies": [
        {
          "id": 3217254,
          "author_name": "Kawa",
          "author_url": "",
          "post_date": "2025-06-04T18:02:46.680000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nhow you do that? I tried use <code>model = model.to(torch.float)</code> and test input with fp32 in <code>/kaggle/input/waveform-inversion/test</code> but seem nothing improved.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3217267,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-04T18:42:03.560000",
              "content": "<p>Disable auto cast. Lb score is the same (due to decimal improvement only)but better ranking</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3217273,
              "author_name": "Kawa",
              "author_url": "",
              "post_date": "2025-06-04T18:50:48.220000",
              "content": "<p>I will try it. Thanks.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3220406,
          "author_name": "Siyuan Yin",
          "author_url": "",
          "post_date": "2025-06-09T08:55:08.300000",
          "content": "<p>Did the score improve during your local testing?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3216127,
      "author_name": "NguyenThanhNhan",
      "author_url": "",
      "post_date": "2025-06-03T06:36:51.947000",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Have you tried re-training convnext large with the same settings as caformer i.e. 150 epochs + improved decoder + cosine scheduler ? <br>\nCaformer training time was 2x longer than convnext large so I'm not sure it's worth the extra cost.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3216334,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-03T12:30:26.557000",
          "content": "<p>Good question <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a>. </p>\n<p>I have not tried <code>convnext large</code> due to compute limits, but 2x faster training sounds nice!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3219000,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-07T03:36:08.287000",
      "content": "<p>this is segmentation head in your code:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .conv = nn.Conv2d(\n            in_channels, out_channels, kernel_size=kernel_size,\n            padding=kernel_size // \n        )\n        .upsample = UpSample(\n            spatial_dims=,\n            in_channels=out_channels,\n            out_channels=out_channels,\n            scale_factor=scale_factor,\n            mode=mode,\n        )\n\n     ():\n        x = .conv(x)\n        x = .upsample(x)\n         x\n</code></pre>\n<p>i haven't done in-depth experiment yet, but below are from my experiences:</p>\n<ul>\n<li>this is regression problem, so value of the prediction is important</li>\n<li>do not use upsample as last layer. you will be bounded by up scaling error.</li>\n<li>use original input data resolution  (or 72x72 sampled version) in final decoder layers</li>\n<li>the last prediction module should be 1x1 conv  (not even 3x3)</li>\n</ul>\n<p>```<br>\ne.g. something like</p>\n<p>last_feature = lastdecoder(input = from previous decoder, skip = original input)<br>\npredicted_velocity = prediction(last_feature)</p>\n<p>where  prediction= nn.Seqential(<br>\n….<br>\nnn.Conv2d( 1x1 kernel)<br>\n)<br>\n``</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3219011,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2025-06-07T04:06:05.330000",
          "content": "<p>You are bound by upscaling error only if you upsample with something like bilinear. If you upsample with ConvTranspose2d there is no bound (although it might be harder to learn than upscaling to full resolution before final layer, but then it's trade-off against compute and time, no bound).<br>\nTheoretically we can compress all the seis into a single vector and then predict from this vector each pixel - there is nothing imposing any bound.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3216409,
      "author_name": "t fuku",
      "author_url": "",
      "post_date": "2025-06-03T14:38:38.930000",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <br>\nWhen turning a Train, the IO seems to be a problem(?). What is your countermeasure for this?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3216437,
          "author_name": "t fuku",
          "author_url": "",
          "post_date": "2025-06-03T15:22:55.247000",
          "content": "<p>To be precise, the speed decreases around the 800-step mark.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3790528%2F06453f06b67c9a8dede330ce4e0d66b2%2F.png?generation=1748964167653260&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 3216726,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-06-04T03:37:23.660000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3216093,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2025-06-03T05:27:23.170000",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Thanks for sharing one more milestone very hard to break <strong>struggle to reach &lt;100</strong> --&gt; <strong>69.7</strong> --&gt; <strong>33.2</strong> --&gt; <strong>28.8</strong>, could you share <a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500\" target=\"_blank\">CV scores</a> of training 20-25 epochs of CurveFault-B ( your cv and training logs if possible )</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3224682,
      "author_name": "Tony Yang",
      "author_url": "",
      "post_date": "2025-06-15T10:12:01.490000",
      "content": "<p>This is very interesting</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3223522,
      "author_name": "Surya_trainer",
      "author_url": "",
      "post_date": "2025-06-13T12:44:36.783000",
      "content": "<p>whats the upsample mode , pixel shuffle or deconv ??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3224846,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-06-15T15:02:57.960000",
          "content": "<p>Read the code :D</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3217724,
      "author_name": "Flowers",
      "author_url": "",
      "post_date": "2025-06-05T10:59:35.717000",
      "content": "<p>When you're testing model performance on a small order of magnitude, do you use CurveFault_B, and if so, what the results are, and then you use them for training on all datasets?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3219155,
          "author_name": "water joe",
          "author_url": "",
          "post_date": "2025-06-07T08:13:20.073000",
          "content": "<p><a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3217157,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-06-04T15:33:31.020000",
      "content": "<p>Thanks for sharing best model with full training….</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3216520,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2025-06-03T17:32:56.413000",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">Bartley</a></p>\n<p>Ok - maybe its a lack of coffee or the fact that I am an old fart - but I think I read all the other comments several times and failed to find the string that leads us to the code for training this latest nice bit of work.  I tried re-using the previous train but that failed.</p>\n<p>Is it your intent to not share the training code for your last model?</p>\n<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/billfan88\" target=\"_blank\">@billfan88</a>, I appreciate your questions but please read other comments first. Cheers 🙂</p>\n</blockquote>",
      "votes": 1,
      "replies": [
        {
          "id": 3216559,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-03T18:31:55.850000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, <a href=\"https://www.kaggle.com/ttyn4519\" target=\"_blank\">@ttyn4519</a>, thanks for your comments.</p>\n<p>I did not include training code in this notebook because the model is trained outside of the Kaggle environment. I do not see issues when running locally.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3216567,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2025-06-03T18:37:43.663000",
              "content": "<p>Yes - and I fork it outside of kaggle :)   That's the only place that multiple days of running are possible.  Your previous ran 10 days on one of my local machines.</p>\n<p>Are you telling me that you can re-use the previous train.py, util.py and the full cfg and it runs OK?  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3216578,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2025-06-03T18:45:13.180000",
              "content": "<p>I use a slightly different training script, but the cfg is the same. It's hard to predict how the pipeline will act across different GPUs, CUDA versions, and environments, but it works for me 🤷‍♂️</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3216633,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2025-06-03T21:37:16.347000",
              "content": "<p>So- in summary - your not sharing the full code for this last model?</p>",
              "votes": -1,
              "replies": []
            },
            {
              "id": 3220810,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-06-10T02:24:29.347000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3216040,
      "author_name": "water joe",
      "author_url": "",
      "post_date": "2025-06-03T03:24:04.390000",
      "content": "<p>Thanks for sharing best model with full training - one more break through notebook and solution.<br>\nCould you share your training logs of these models after 50 epochs?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3216066,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-03T04:19:17.360000",
          "content": "<p><a href=\"https://www.kaggle.com/waterjoe\" target=\"_blank\">@waterjoe</a> <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> here are the loss values. After 50 epochs training loss is ~21.5 and validation loss is 29.3.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F1900a9c815836e8b891baeed2ed69461%2FScreen%20Shot%202025-06-02%20at%209.16.55%20PM.png?generation=1748924280482744&amp;alt=media\" alt=\"Loss curves\"></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3215914,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-02T20:38:25.347000",
      "content": "<p>Was the model trained as in previous baselines?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3215926,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-02T21:28:33.530000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/xbar19\" target=\"_blank\">@xbar19</a> and <a href=\"https://www.kaggle.com/billfan88\" target=\"_blank\">@billfan88</a>, thanks for the comments.</p>\n<p>The training is very similar. The differences are that this model is trained for 150 epochs and uses a custom learning rate scheduler shown <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved?scriptVersionId=243291347&amp;cellId=4\" target=\"_blank\">here</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 3215938,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-06-02T21:48:06.603000",
              "content": "<p>Thank you for your reply and your work!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3215946,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-06-02T22:21:46.487000",
              "content": "<p>How much time it took to train on 150 epochs? Damn this is one long training</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3215958,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2025-06-02T23:25:38.563000",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a>, it took 8 days (~192 hours) on a single 4090.</p>\n<p>You make a good point — this pipeline takes far too long to converge. I am hoping people can share ideas on how we can decrease this time!</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3215959,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-06-02T23:29:23.150000",
              "content": "<p>Holy <strong><em>*</em></strong> shit! 8 days…<br>\nFull 8*24 days? Damn</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3215962,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-06-02T23:43:49.617000",
              "content": "<blockquote>\n  <p>this pipeline takes far too long to converge. </p>\n</blockquote>\n<p>MAE is always slow to converge. Alas.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3217497,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2025-06-05T04:30:26.740000",
              "content": "<p>It would take me ~8 days on a single 4090 to hit my current LB (17.1), the trainings take a long time to converge </p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3217741,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2025-06-05T11:18:10.863000",
              "content": "<p><strong>Magic of L1 --&gt; Run Long</strong> - Thanks for sharing <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>  <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>\n<ul>\n<li>it took 8 days (~192 hours) on a single 4090 -- <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> </li>\n<li>It would take me ~8 days on a single 4090 to hit my current LB (17.1), the trainings take a long time to converge <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </li>\n</ul>\n<hr>\n<blockquote>\n  <p>-- How to decided when to stop, most of epochs getting negligible loss improvement (small improvement - my experiments based on only CurveFault_B)<br>\n   -- some batches are too high even using clip the grads<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fc49c05fe040a84a4cb8bd2e1b3e1644e%2FScreenshot%202025-06-05%20at%204.46.22PM.png?generation=1749122195332342&amp;alt=media\" alt=\"\"></p>\n</blockquote>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3217866,
              "author_name": "yukiZ",
              "author_url": "",
              "post_date": "2025-06-05T13:15:42.077000",
              "content": "<p>Since there is a lot of information on 4090, I will share the results of V100x8 DDP.</p>\n<p>To reproduce this old notebook's convnext[CV 31.9 LB 36.4] locally, it took 2 days for 150 epochs of training, and I was able to obtain the result of [CV30.07, LB35.1].</p>\n<pre><code>=========================\nCurveFault_A    6.15\nCurveFault_B    86.72\nCurveVel_A      13.69\nCurveVel_B      48.81\nFlatFault_A     4.34\nFlatFault_B     33.73\nFlatVel_A       2.92\nFlatVel_B       12.11\nStyle_A         37.07\nStyle_B         55.18\n=========================\n\n</code></pre>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 3220281,
              "author_name": "lin-kukuoreoa",
              "author_url": "",
              "post_date": "2025-06-09T05:17:50.173000",
              "content": "<p>Could you please tell me whether the data you used was the original FP32 data? When I tried to reproduce Caformer’s training results using Kaggle’s FP16 open dataset, the cross-validation converged around 34, and I’m not sure if this is a dataset issue.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3220288,
              "author_name": "yukiZ",
              "author_url": "",
              "post_date": "2025-06-09T05:39:21.670000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/oigckko\" target=\"_blank\">@oigckko</a> <br>\nIf you have any questions regarding my above post \"Val MAE: 30.07\", I  hope this helps.</p>\n<p><strong>The training dataset remains unchanged from the public notebook.</strong></p>\n<pre><code>df= pd()\n...\np1 = os(, row)\n            p2 = os(, row(), , row())\n            p3 = os(, row)\n            p4 = os(, row(), , row())\n</code></pre>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3215969,
      "author_name": "Bill Fan",
      "author_url": "",
      "post_date": "2025-06-03T00:22:00.277000",
      "content": "<p>Train section  Same as starter notebook?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3216322,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-03T12:18:06.103000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/billfan88\" target=\"_blank\">@billfan88</a>, I appreciate your questions but please read other comments first. Cheers 🙂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3215917,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2025-06-02T20:57:34.400000",
      "content": "<p>Awesome!</p>\n<p>I wonder how many backbones you tried. You obviously have a good methodology to pick the ones you test.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3215933,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-02T21:42:40.240000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, I have tried Convnext, Caformer, Hgnet, NextVIT and Hornet so far.</p>\n<p>My methodology is quite simple. First I try and modify the backbone for 1000 x 70 in as few steps as possible. Then, I follow <a href=\"https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500\" target=\"_blank\">this</a> validation scheme. It may not be the most accurate, but it seems to work well given limited compute!</p>",
          "votes": 7,
          "replies": [
            {
              "id": 3215939,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-06-02T21:49:12.983000",
              "content": "<p>Yes, I noticed you use a subset to validate your choice. This is a good idea.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3221540,
      "author_name": "Flowers",
      "author_url": "",
      "post_date": "2025-06-11T06:45:49.307000",
      "content": "<p>Why is my train loss always higher than vaild loss during training, is this normal?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25710556%2F53e5d71b04b7954200dfa184996221ba%2F2025-06-11%20144329.png?generation=1749624341620051&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3221552,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2025-06-11T07:03:30.613000",
          "content": "<p><a href=\"https://www.kaggle.com/oliver34\" target=\"_blank\">@oliver34</a> its same for me too, but it got switch after few epochs when i tested on <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/580892#3221434\" target=\"_blank\">CurveFault_B</a> dataset experiments.</p>\n<p>-- How many epochs you are running of CAFormer? Did you run from backbone or from Bartley weights?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3221557,
              "author_name": "Flowers",
              "author_url": "",
              "post_date": "2025-06-11T07:11:46.630000",
              "content": "<p>44epoch,I only used the weights of the backbone</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3221558,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2025-06-11T07:15:54.113000",
              "content": "<p><a href=\"https://www.kaggle.com/oliver34\" target=\"_blank\">@oliver34</a> will you share your train/val loss for epochs 5, 10, 20, 30 ? -- which gpu - training finishing in 1h 25mins for epoch?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3221562,
              "author_name": "Flowers",
              "author_url": "",
              "post_date": "2025-06-11T07:23:45.047000",
              "content": "<p>They are 47.70, 40.59, 31.86, 28.70. I also encountered the situation you mentioned at CurveFault_B. --I'm using two 4090s and it seems like it takes 9 days to complete a 150epoch😥</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3221639,
              "author_name": "lin-kukuoreoa",
              "author_url": "",
              "post_date": "2025-06-11T09:46:46.287000",
              "content": "<p>Hello, may I ask if you are using the fp32 data or the fp16 dataset processed by Kaggle?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3221666,
              "author_name": "Flowers",
              "author_url": "",
              "post_date": "2025-06-11T10:30:15.037000",
              "content": "<p>the fp16 dataset processed by Kaggle</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3224699,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2025-06-15T10:45:50.537000",
              "content": "<p><a href=\"https://www.kaggle.com/oliver34\" target=\"_blank\">@oliver34</a>  could you share cv scores of 40, 60, 80, 100? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3224761,
              "author_name": "Flowers",
              "author_url": "",
              "post_date": "2025-06-15T12:45:32.147000",
              "content": "<p>I would love to share them, 26.88, 24.58, 23.08, 22.11.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3225362,
              "author_name": "DJ_Xia",
              "author_url": "",
              "post_date": "2025-06-16T09:50:55.937000",
              "content": "<p>With the \"<a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/583896\" target=\"_blank\">Tips to speed up training</a>\" optimizations, training can be completed within 2 days using 2×4090 </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3225364,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-16T09:53:50.547000",
              "content": "<p>I don’t think 4090 can train caformer for 150 epoch for 2 days</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3222445,
          "author_name": "M Sato",
          "author_url": "",
          "post_date": "2025-06-12T07:46:55.827000",
          "content": "<p>batch-wise average vs overall average</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3223545,
          "author_name": "Truth Seeker",
          "author_url": "",
          "post_date": "2025-06-13T13:28:12.613000",
          "content": "<blockquote>\n  <p>my train loss always higher than vaild loss during training</p>\n</blockquote>\n<p>I also observe this.  I haven't observed a single epoch in which my val loss is worse than my train loss.  I'm still relatively early in my training epochs, though.  Glad to know I'm not the only one. 😀</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3221539,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-11T06:45:11.713000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3225006,
      "author_name": "Iqra Rafique",
      "author_url": "",
      "post_date": "2025-06-15T19:37:03.183000",
      "content": "<p>Great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3223364,
      "author_name": "Du Qinglin",
      "author_url": "",
      "post_date": "2025-06-13T07:42:11.890000",
      "content": "<p>Thanks for the information</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3215873": "Happy to share an improved full resolution model. This model uses a CAFormer encoder with an improved decoder (pixel shuffle, SCSE, intermediate convolutions, etc.).\n\nNotebook [here](https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved)\nDataset [here](https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72)\n\n```\n+--------------+----------+\n| Dataset      | Caformer |\n+--------------+----------+\n| CurveFault_A |     4.00 |\n| CurveFault_B |    71.06 |\n| CurveVel_A   |     9.19 |\n| CurveVel_B   |    38.60 |\n| FlatFault_A  |     2.58 |\n| FlatFault_B  |    24.29 |\n| FlatVel_A    |     1.31 |\n| FlatVel_B    |     7.20 |\n| Style_A      |    34.34 |\n| Style_B      |    48.90 |\n+--------------+----------+\n| Overall      |    24.15 |\n+--------------+----------+\n```\n\nThis will be my last notebook in this competition, but I am hopeful that we can push this architecture further as a community!\n\nHappy Kaggling 😄",
    "3217147": "\"this pipeline takes far too long to converge.\"\n- you can classify the input type ( 3 families—Vel, Fault, and Style) for the test with a classifier (on the predicted velocity)\n- then you can train seperate models for each type (this happens for most paper)\n- if you have many machines, you can train in parallel",
    "3217142": "a trick is use fp32  in input and model for inference. you will get something like +0.0025 in lb (?)",
    "3216127": "@brendanartley Have you tried re-training convnext large with the same settings as caformer i.e. 150 epochs + improved decoder + cosine scheduler ? \nCaformer training time was 2x longer than convnext large so I'm not sure it's worth the extra cost.",
    "3219000": "this is segmentation head in your code:\n\n```\n\n\nclass SegmentationHead2d(nn.Module):\n    def __init__(\n            self,\n            in_channels,\n            out_channels,\n            scale_factor: tuple[int] = (2, 2),\n            kernel_size: int = 3,\n            mode: str = \"nontrainable\",\n    ):\n        super().__init__()\n        self.conv = nn.Conv2d(\n            in_channels, out_channels, kernel_size=kernel_size,\n            padding=kernel_size // 2\n        )\n        self.upsample = UpSample(\n            spatial_dims=2,\n            in_channels=out_channels,\n            out_channels=out_channels,\n            scale_factor=scale_factor,\n            mode=mode,\n        )\n\n    def forward(self, x):\n        x = self.conv(x)\n        x = self.upsample(x)\n        return x\n\n```\n\ni haven't done in-depth experiment yet, but below are from my experiences:\n- this is regression problem, so value of the prediction is important\n- do not use upsample as last layer. you will be bounded by up scaling error.\n- use original input data resolution  (or 72x72 sampled version) in final decoder layers\n- the last prediction module should be 1x1 conv  (not even 3x3)\n\n```\ne.g. something like\n\nlast_feature = lastdecoder(input = from previous decoder, skip = original input)\npredicted_velocity = prediction(last_feature)\n\nwhere  prediction= nn.Seqential(\n....\nnn.Conv2d( 1x1 kernel)\n)\n``",
    "3216409": "@brendanartley \nWhen turning a Train, the IO seems to be a problem(?). What is your countermeasure for this?",
    "3216093": "@brendanartley Thanks for sharing one more milestone very hard to break **struggle to reach <100** --> **69.7** --> **33.2** --> **28.8**, could you share [CV scores](https://www.kaggle.com/code/brendanartley/convnext-full-resolution-baseline/comments#3206500) of training 20-25 epochs of CurveFault-B ( your cv and training logs if possible )",
    "3224682": "This is very interesting",
    "3223522": "whats the upsample mode , pixel shuffle or deconv ??",
    "3217724": "When you're testing model performance on a small order of magnitude, do you use CurveFault_B, and if so, what the results are, and then you use them for training on all datasets?",
    "3217157": "Thanks for sharing best model with full training....",
    "3216520": "[Bartley](https://www.kaggle.com/brendanartley)\n\nOk - maybe its a lack of coffee or the fact that I am an old fart - but I think I read all the other comments several times and failed to find the string that leads us to the code for training this latest nice bit of work.  I tried re-using the previous train but that failed.\n\nIs it your intent to not share the training code for your last model?\n\n>Hi @billfan88, I appreciate your questions but please read other comments first. Cheers 🙂",
    "3216040": "Thanks for sharing best model with full training - one more break through notebook and solution.\nCould you share your training logs of these models after 50 epochs?\n",
    "3215914": "Was the model trained as in previous baselines?",
    "3215969": "Train section  Same as starter notebook?",
    "3215917": "Awesome!\n\nI wonder how many backbones you tried. You obviously have a good methodology to pick the ones you test.",
    "3221540": "Why is my train loss always higher than vaild loss during training, is this normal?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25710556%2F53e5d71b04b7954200dfa184996221ba%2F2025-06-11%20144329.png?generation=1749624341620051&alt=media)",
    "3221539": "",
    "3225006": "Great work!",
    "3223364": "Thanks for the information"
  }
}