{
  "id": 584772,
  "title": "my maybe efficient trainer?",
  "url": "/competitions/waveform-inversion/discussion/584772",
  "author_name": "hengck23",
  "post_date": "2025-06-16T01:30:25.182000",
  "votes": 16,
  "comment_count": 39,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8a05eebaa2abfdbbc4903c0972ff0ba7%2FSelection_252.png?generation=1750036637016034&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ec93d83b8121bf5e188b3b901555fc2%2FSelection_261.png?generation=1750068630074827&amp;alt=media\" alt=\"\"><br>\ncode: <a href=\"https://www.kaggle.com/datasets/hengck23/kaggle-competition-waveform-inversion-discussion-0\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/kaggle-competition-waveform-inversion-discussion-0</a></p>\n<hr>\n<p>gpu = 2x RTX 6000 (close to 4090)<br>\nvalid samples = 20 np arrays = 2 npy from each family (1,10 npy, following Bartley's fold0) <br>\ntrain samples = rest of arrays<br>\n(no data subsampling used)</p>\n<h2>training speed:</h2>\n<ul>\n<li>convnext base at 30min per epoch</li>\n<li>convnext small at 24min</li>\n<li>covnnext large at 40min</li>\n</ul>\n<p>at most 3 days for 150 epoch(?)</p>\n<h2>implementation:</h2>\n<ul>\n<li><p>basically Bartley'smodel with my modifications</p>\n<ul>\n<li>my own stem</li>\n<li>batch norm in unet</li>\n<li>instance-norm2d #1 in convnext block  but remove contiguous() in forward</li></ul></li>\n<li><p>train hyper-parameters:</p>\n<ul>\n<li>I wrote my own trainer, the train MAE shown is moving average over the last 100 batches. validation MAE exclude flip TTA</li>\n<li>batch size = 64 for each GPU (faster if I use batch size =128, which I intend to use at later epoch of fine tunning with smaller lr). i did not use sync batch norm for multi-gpu training, so batch size has to be large</li>\n<li>bfloat16, hence autocast scaler is not required</li>\n<li>Bartley's EMA #1</li>\n<li>I cannot use optimizer fuse and torch compile. i have errors below.</li>\n<li>I cannot use FP8 training (torchAO and FSDP2) for now</li></ul></li>\n</ul>\n<p>for other configurations, please refer to code and log in the public dataset link</p>\n<pre><code>\n</code></pre>\n<p>ow</p>",
  "messages": [
    {
      "id": 3225091,
      "postDate": "2025-06-16T01:30:25.183Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8a05eebaa2abfdbbc4903c0972ff0ba7%2FSelection_252.png?generation=1750036637016034&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ec93d83b8121bf5e188b3b901555fc2%2FSelection_261.png?generation=1750068630074827&amp;alt=media\" alt=\"\"><br>\ncode: <a href=\"https://www.kaggle.com/datasets/hengck23/kaggle-competition-waveform-inversion-discussion-0\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/kaggle-competition-waveform-inversion-discussion-0</a></p>\n<hr>\n<p>gpu = 2x RTX 6000 (close to 4090)<br>\nvalid samples = 20 np arrays = 2 npy from each family (1,10 npy, following Bartley's fold0) <br>\ntrain samples = rest of arrays<br>\n(no data subsampling used)</p>\n<h2>training speed:</h2>\n<ul>\n<li>convnext base at 30min per epoch</li>\n<li>convnext small at 24min</li>\n<li>covnnext large at 40min</li>\n</ul>\n<p>at most 3 days for 150 epoch(?)</p>\n<h2>implementation:</h2>\n<ul>\n<li><p>basically Bartley'smodel with my modifications</p>\n<ul>\n<li>my own stem</li>\n<li>batch norm in unet</li>\n<li>instance-norm2d #1 in convnext block  but remove contiguous() in forward</li></ul></li>\n<li><p>train hyper-parameters:</p>\n<ul>\n<li>I wrote my own trainer, the train MAE shown is moving average over the last 100 batches. validation MAE exclude flip TTA</li>\n<li>batch size = 64 for each GPU (faster if I use batch size =128, which I intend to use at later epoch of fine tunning with smaller lr). i did not use sync batch norm for multi-gpu training, so batch size has to be large</li>\n<li>bfloat16, hence autocast scaler is not required</li>\n<li>Bartley's EMA #1</li>\n<li>I cannot use optimizer fuse and torch compile. i have errors below.</li>\n<li>I cannot use FP8 training (torchAO and FSDP2) for now</li></ul></li>\n</ul>\n<p>for other configurations, please refer to code and log in the public dataset link</p>\n<pre><code>\n</code></pre>\n<p>ow</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8a05eebaa2abfdbbc4903c0972ff0ba7%2FSelection_252.png?generation=1750036637016034&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ec93d83b8121bf5e188b3b901555fc2%2FSelection_261.png?generation=1750068630074827&alt=media)\ncode: https://www.kaggle.com/datasets/hengck23/kaggle-competition-waveform-inversion-discussion-0\n\n---\n\ngpu = 2x RTX 6000 (close to 4090)\nvalid samples = 20 np arrays = 2 npy from each family (1,10 npy, following Bartley's fold0) \ntrain samples = rest of arrays\n(no data subsampling used)\n\n##training speed:\n- convnext base at 30min per epoch\n- convnext small at 24min\n- covnnext large at 40min\n\nat most 3 days for 150 epoch(?)\n\n##implementation:\n- basically Bartley'smodel with my modifications\n  - my own stem\n  - batch norm in unet\n  - instance-norm2d #1 in convnext block  but remove contiguous() in forward\n\n- train hyper-parameters:\n  - I wrote my own trainer, the train MAE shown is moving average over the last 100 batches. validation MAE exclude flip TTA\n  - batch size = 64 for each GPU (faster if I use batch size =128, which I intend to use at later epoch of fine tunning with smaller lr). i did not use sync batch norm for multi-gpu training, so batch size has to be large\n  - bfloat16, hence autocast scaler is not required\n  - Bartley's EMA #1\n  - I cannot use optimizer fuse and torch compile. i have errors below.\n  - I cannot use FP8 training (torchAO and FSDP2) for now\n\n\nfor other configurations, please refer to code and log in the public dataset link\n\n```\n#1 : I find this to have great effect in accuracy\n```ow",
      "votes": 16
    },
    {
      "id": 3227610,
      "postDate": "2025-06-19T05:54:55.447Z",
      "content": "<p>how to do experiment fast. Bartley's caformer use more complicated decoder with attention and intermediate layers. i want to see if decoder is important in this competition. hence:</p>\n<ol>\n<li>I change the decoder to Bartley's</li>\n<li>load weights of previously trained encoder and freeze them</li>\n<li>I train with new decoder only<br>\nnow I can compare if Bartley's is better:</li>\n</ol>\n<p>I can further unfreeze encoder and train and train all after initial warmup.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F044a2cbc3848434184e879595c187149%2FSelection_093.png?generation=1750312490647785&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "how to do experiment fast. Bartley's caformer use more complicated decoder with attention and intermediate layers. i want to see if decoder is important in this competition. hence:\n1. I change the decoder to Bartley's\n2. load weights of previously trained encoder and freeze them\n3. I train with new decoder only\nnow I can compare if Bartley's is better:\n\nI can further unfreeze encoder and train and train all after initial warmup.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F044a2cbc3848434184e879595c187149%2FSelection_093.png?generation=1750312490647785&alt=media)",
      "votes": 1
    },
    {
      "id": 3225771,
      "postDate": "2025-06-16T20:27:43.420Z",
      "content": "<p>other worth trying:</p>\n<ul>\n<li><p>densenetersion of covnext:<br>\n<a href=\"https://huggingface.co/naver-ai/rdnet_base.nv_in1k\" target=\"_blank\">https://huggingface.co/naver-ai/rdnet_base.nv_in1k</a></p></li>\n<li><p><a href=\"https://github.com/AILab-CVC/UniRepLKNet\" target=\"_blank\">https://github.com/AILab-CVC/UniRepLKNet</a></p></li>\n<li><p><a href=\"https://github.com/LMMMEng/OverLoCK\" target=\"_blank\">https://github.com/LMMMEng/OverLoCK</a></p></li>\n</ul>\n<p>if strong encoder doesn't help, try strong decoder (e.g. attention layer, skip layer as in Bartley's caformer; unet++ or multi connect unet)</p>",
      "rawMarkdown": "other worth trying:\n- densenetersion of covnext:\nhttps://huggingface.co/naver-ai/rdnet_base.nv_in1k\n\n- https://github.com/AILab-CVC/UniRepLKNet\n- https://github.com/LMMMEng/OverLoCK\n\nif strong encoder doesn't help, try strong decoder (e.g. attention layer, skip layer as in Bartley's caformer; unet++ or multi connect unet)",
      "votes": 1
    },
    {
      "id": 3225630,
      "postDate": "2025-06-16T17:01:53.410Z",
      "content": "<p>a stronger and fast backbone (? or/and !)<br>\n<a href=\"https://huggingface.co/timm/inception_next_base.sail_in1k\" target=\"_blank\">https://huggingface.co/timm/inception_next_base.sail_in1k</a></p>\n<p>This is another metaformer like caformer. Both are from the same lab SAIL.<br>\nNOTE: inception-next is batchnorm based, so I used \"net = torch.nn.SyncBatchNorm.convert_sync_batchnorm(net) \"</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc7232f69b59601495251b2062df593d4%2FSelection_264.png?generation=1750104201525874&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "a stronger and fast backbone (? or/and !)\nhttps://huggingface.co/timm/inception_next_base.sail_in1k\n\nThis is another metaformer like caformer. Both are from the same lab SAIL.\nNOTE: inception-next is batchnorm based, so I used \"net = torch.nn.SyncBatchNorm.convert_sync_batchnorm(net) \"\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc7232f69b59601495251b2062df593d4%2FSelection_264.png?generation=1750104201525874&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 3225656,
          "postDate": "2025-06-16T17:30:39.687Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, when do you decided to stop training and try next experiment?</p>\n<p>As i tested multiple experiments on CurveFault_B ( with your setup - 1e-3 and with 1e-4 ) --&gt; 1e-4 got better improvement --&gt; not tested on all dataset. as we train more - gap between training loss improving 3x and validation improving with 0.75x.</p>\n<p>will you test to scale the velocity / 100 and see how your training converging ? Nice to see you in this compatition <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "rawMarkdown": "@hengck23, when do you decided to stop training and try next experiment?\n\nAs i tested multiple experiments on CurveFault_B ( with your setup - 1e-3 and with 1e-4 ) --> 1e-4 got better improvement --> not tested on all dataset. as we train more - gap between training loss improving 3x and validation improving with 0.75x.\n\nwill you test to scale the velocity / 100 and see how your training converging ? Nice to see you in this compatition @hengck23 ",
          "votes": 1,
          "replies": [
            {
              "id": 3225671,
              "postDate": "2025-06-16T17:44:16.307Z",
              "content": "<p>i first train a fast model:<br>\n1e-3 for 35 epoch, 1e-4 for next 15 epoch, 1e-5 for next 5 epoch.</p>\n<p>that confirms:<br>\n1e-4 at at most improve mae by 3 to4 , 1e-5 can almost improve 0.5 to 1</p>\n<p>also, from your validation, you can estimate if current mae loss = xxx,<br>\nto reduce 1 mae loss, you need e.g.  1 epoch<br>\nreduce next 1 mae, you need 2<br>\nreduce next 1 mae, you need 4<br>\nreduce next 1 mae, you need 8<br>\nreduce next 1 mae, you need 12 ….</p>\n<p>so you can estimate the final performance at the middle of training if you only have 150 or 200 epoch</p>\n<p>you should confirm this with public models (eg, small convnext) where other already train for 50, 100 and 150 epoches.</p>\n<hr>\n<p>lastly assume  model A has validation =40 at epoch 30,<br>\nfor model B to achieve +5 better at 150, B must be better than A by eg +6 at epoch 30.</p>\n<p>so you can try different backbone till epoch 30.</p>\n<p>(30,5,6 are all exmpale numbers, you need to do experiments to find out)</p>",
              "rawMarkdown": "i first train a fast model:\n1e-3 for 35 epoch, 1e-4 for next 15 epoch, 1e-5 for next 5 epoch.\n\nthat confirms:\n1e-4 at at most improve mae by 3 to4 , 1e-5 can almost improve 0.5 to 1\n\nalso, from your validation, you can estimate if current mae loss = xxx,\nto reduce 1 mae loss, you need e.g.  1 epoch\nreduce next 1 mae, you need 2\nreduce next 1 mae, you need 4\nreduce next 1 mae, you need 8\nreduce next 1 mae, you need 12 ....\n\nso you can estimate the final performance at the middle of training if you only have 150 or 200 epoch\n\nyou should confirm this with public models (eg, small convnext) where other already train for 50, 100 and 150 epoches.\n\n\n---\n\nlastly assume  model A has validation =40 at epoch 30,\nfor model B to achieve +5 better at 150, B must be better than A by eg +6 at epoch 30.\n\nso you can try different backbone till epoch 30.\n\n(30,5,6 are all exmpale numbers, you need to do experiments to find out)",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3225615,
      "postDate": "2025-06-16T16:40:15.917Z",
      "content": "<p>torch.compile(… mode='max) is very helpful.</p>",
      "rawMarkdown": "torch.compile(... mode='max) is very helpful.",
      "votes": 1
    },
    {
      "id": 3227013,
      "postDate": "2025-06-18T11:26:27.640Z",
      "content": "<p>covnext-large with data sampling.<br>\nCV after training for 24hr:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F09632d6ed0ddcf89643fc6e3a1b6db9a%2FSelection_089.png?generation=1750269732895607&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "covnext-large with data sampling.\nCV after training for 24hr:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F09632d6ed0ddcf89643fc6e3a1b6db9a%2FSelection_089.png?generation=1750269732895607&alt=media)",
      "votes": 2,
      "replies": [
        {
          "id": 3227095,
          "postDate": "2025-06-18T13:50:11.077Z",
          "content": "<p>Pardon me , but why does epoch has values after decimal?? <br>\nLike are you considering batches done also??</p>",
          "rawMarkdown": "Pardon me , but why does epoch has values after decimal?? \nLike are you considering batches done also??"
        }
      ]
    },
    {
      "id": 3226100,
      "postDate": "2025-06-17T08:44:44.513Z",
      "content": "<p>update!!!</p>\n<p>i change machine.<br>\nOLD: 2xRTX A6000<br>\nNEW: 2xRTX 6000 ada, shift all data to SSD</p>\n<p>previous: convnext base at 30min per epoch<br>\nnew: convnext base at 21 min per epoch !!!</p>\n<p>i think a lot is due to ssd</p>",
      "rawMarkdown": "update!!!\n\ni change machine.\nOLD: 2xRTX A6000\nNEW: 2xRTX 6000 ada, shift all data to SSD\n\nprevious: convnext base at 30min per epoch\nnew: convnext base at 21 min per epoch !!!\n\ni think a lot is due to ssd",
      "votes": 2,
      "replies": [
        {
          "id": 3226103,
          "postDate": "2025-06-17T08:51:39.017Z",
          "content": "<p>did you add weighted random sampler during train ? that would probably be a big gain. </p>",
          "rawMarkdown": "did you add weighted random sampler during train ? that would probably be a big gain. ",
          "votes": 2,
          "replies": [
            {
              "id": 3226110,
              "postDate": "2025-06-17T08:58:25.767Z",
              "content": "<p>yes,i note that for different epoch the mae loss for simple family doesn't change. the imporvment comes for difficult family.<br>\nnow i am testing different backbone(upt o 30 epoch) and ifferent train config (eg fp8), so i use all train samples first.</p>",
              "rawMarkdown": "yes,i note that for different epoch the mae loss for simple family doesn't change. the imporvment comes for difficult family.\nnow i am testing different backbone(upt o 30 epoch) and ifferent train config (eg fp8), so i use all train samples first.",
              "votes": 2
            },
            {
              "id": 3226504,
              "postDate": "2025-06-17T18:51:56.767Z",
              "content": "<p>thanks <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> for the hint. effects of sampling difficult train data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5873887d0fa2dd34ba9f1becfa1c5d74%2FSelection_084.png?generation=1750186244010579&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5eab51a72bb9196a19538ef380402df0%2FSelection_085.png?generation=1750186258173534&amp;alt=media\" alt=\"\"></p>\n<pre><code>class MyFWIDataset(Dataset):\n    def __init__(\n            self,\n            df,\n            cfg,\n            mode='train', \n    ):\n   ...\n\n    def __getitem__(self, ):\n\n         .mode == :\n            family = .family[np..choice(\n                np.arange(len(.family)),\n                p=.family_p\n            )]\n            df = ..get_group(family)\n            i = np..choice(df.)\n            j = np..choice()\n        :\n            i =  // \n            j =  % \n\n        seismic = .seismic[i][j]\n        velocity = .velocity[i][j]\n</code></pre>\n<p>but there is a catch: you are using error of validation IN training, i.e. you are overfitting validation data.<br>\nbut is this a problem and do we need or need not solve it? This is homework for kagglers to think about.</p>\n<p>hint: it is all about if public test is in-distribution or out-distribution</p>",
              "rawMarkdown": "thanks @darraghdog for the hint. effects of sampling difficult train data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5873887d0fa2dd34ba9f1becfa1c5d74%2FSelection_084.png?generation=1750186244010579&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5eab51a72bb9196a19538ef380402df0%2FSelection_085.png?generation=1750186258173534&alt=media)\n\n\n```\nclass MyFWIDataset(Dataset):\n    def __init__(\n            self,\n            df,\n            cfg,\n            mode='train', \n    ):\n   ...\n\n    def __getitem__(self, index):\n\n        if self.mode == 'train':\n            family = self.family[np.random.choice(\n                np.arange(len(self.family)),\n                p=self.family_p\n            )]\n            df = self.group.get_group(family)\n            i = np.random.choice(df.index)\n            j = np.random.choice(LENGTH)\n        else:\n            i = index // LENGTH\n            j = index % LENGTH\n\n        seismic = self.seismic[i][j]\n        velocity = self.velocity[i][j]\n\n```\n\nbut there is a catch: you are using error of validation IN training, i.e. you are overfitting validation data.\nbut is this a problem and do we need or need not solve it? This is homework for kagglers to think about.\n\nhint: it is all about if public test is in-distribution or out-distribution"
            },
            {
              "id": 3226567,
              "postDate": "2025-06-17T20:22:18.710Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  is the family_p approach as similar to torch.utils.data.WeightedRandomSampler?</p>",
              "rawMarkdown": "@hengck23  is the family_p approach as similar to torch.utils.data.WeightedRandomSampler?"
            },
            {
              "id": 3226570,
              "postDate": "2025-06-17T20:26:49.227Z",
              "content": "<p>probabilty the same, but i need to use distributed parallel sampler  for DDP multi gpu training</p>",
              "rawMarkdown": "probabilty the same, but i need to use distributed parallel sampler  for DDP multi gpu training",
              "votes": 1
            },
            {
              "id": 3226717,
              "postDate": "2025-06-18T03:30:09.400Z",
              "content": "<p>in my experiment, it overfitting the training dataset well but not well for validation <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> -&gt; </p>\n<table>\n<thead>\n<tr>\n<th>epoch</th>\n<th>CV</th>\n<th>Train</th>\n<th>Gap</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>1</td>\n<td>84.8</td>\n<td>37.7</td>\n<td>47.1</td>\n</tr>\n<tr>\n<td>2</td>\n<td>73.6</td>\n<td>29.1</td>\n<td>44.5</td>\n</tr>\n<tr>\n<td>3</td>\n<td>67.7</td>\n<td>25.6</td>\n<td>42.1</td>\n</tr>\n<tr>\n<td>4</td>\n<td>63.7</td>\n<td>23.8</td>\n<td>38.1</td>\n</tr>\n<tr>\n<td>5</td>\n<td>57.8</td>\n<td>19.8</td>\n<td>38</td>\n</tr>\n<tr>\n<td>6</td>\n<td>55.7</td>\n<td>18.4</td>\n<td>37.3</td>\n</tr>\n<tr>\n<td>7</td>\n<td>55.6</td>\n<td>18.8</td>\n<td>36.8</td>\n</tr>\n<tr>\n<td>8</td>\n<td>55.0</td>\n<td>18.6</td>\n<td>36.4</td>\n</tr>\n<tr>\n<td>9</td>\n<td>51.8</td>\n<td>15.9</td>\n<td>35. 9</td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "in my experiment, it overfitting the training dataset well but not well for validation @hengck23 -> \n| epoch | CV | Train | Gap |\n| --- | --- | ---- | ---- |\n|  |  | | |\n| 1 | 84.8 | 37.7 | 47.1 |\n| 2 | 73.6 | 29.1 | 44.5 |\n| 3| 67.7 | 25.6 | 42.1 |\n| 4 | 63.7 | 23.8 | 38.1 |\n| 5 | 57.8 | 19.8 | 38 |\n| 6 | 55.7 | 18.4 | 37.3 |\n| 7 | 55.6 | 18.8 | 36.8 |\n| 8 | 55.0 | 18.6 | 36.4 |\n| 9 | 51.8 | 15.9 | 35. 9 |\n\n",
              "votes": 1
            },
            {
              "id": 3226800,
              "postDate": "2025-06-18T05:24:20.713Z",
              "content": "<p>something is very wrong, without code or train parameters/train log, i cannot diagnose the error.<br>\nplease post them if you can.</p>\n<hr>\n<p>i can only say that:<br>\n1) assume that you are using my fold.csv<br>\n2) then,  training at constant 1e-3, validation loss is always better than train loss(moving average) until they converge.</p>\n<p>they converge at:<br>\nmae loss = 35 for covnext small,<br>\n33 for base<br>\n30 for large</p>\n<p>(this is also the approximate LB score)</p>\n<hr>\n<p>i think there is a bug in your code. it is easy to check</p>\n<ul>\n<li>since you are sample more difficult sample, train loss is higher than the unsampled version (btw: high loss means greater gradient backprop, so faster converge)</li>\n</ul>\n<p>did you subsample more difficult samples or mistaken sample the easier ones?</p>\n<ul>\n<li>getting train loss at 37.7 is impossible  (assume you start from imagenet pretrain). you can check train and validation error of imagenet pretrain initialisation. this is initial gap. </li>\n</ul>",
              "rawMarkdown": "something is very wrong, without code or train parameters/train log, i cannot diagnose the error.\nplease post them if you can.\n\n---\n\ni can only say that:\n1) assume that you are using my fold.csv\n2) then,  training at constant 1e-3, validation loss is always better than train loss(moving average) until they converge.\n\nthey converge at:\nmae loss = 35 for covnext small,\n33 for base\n30 for large\n\n(this is also the approximate LB score)\n\n---\n\ni think there is a bug in your code. it is easy to check\n- since you are sample more difficult sample, train loss is higher than the unsampled version (btw: high loss means greater gradient backprop, so faster converge)\n\ndid you subsample more difficult samples or mistaken sample the easier ones?\n\n- getting train loss at 37.7 is impossible  (assume you start from imagenet pretrain). you can check train and validation error of imagenet pretrain initialisation. this is initial gap. \n\n \n\n",
              "votes": 1
            },
            {
              "id": 3226929,
              "postDate": "2025-06-18T09:08:19.213Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - its because <strong>since you are sample more difficult sample, train loss is higher than the unsampled version (btw: high loss means greater gradient backprop, so faster converge)</strong></p>",
              "rawMarkdown": "@hengck23 - its because **since you are sample more difficult sample, train loss is higher than the unsampled version (btw: high loss means greater gradient backprop, so faster converge)**"
            },
            {
              "id": 3227061,
              "postDate": "2025-06-18T12:39:24.733Z",
              "content": "<p>it is either you oversample (too many difficult and too few simple) or wrong/under sampling  (too many simple and too few difficult).</p>\n<p>assuming there is no other bug …</p>",
              "rawMarkdown": "it is either you oversample (too many difficult and too few simple) or wrong/under sampling  (too many simple and too few difficult).\n \nassuming there is no other bug ...",
              "votes": 1
            },
            {
              "id": 3227430,
              "postDate": "2025-06-19T02:35:03.770Z",
              "content": "<p>Thanks for your implementation <code>WeightedSampler</code> for DDP training. Besides, could you please share your experiences of setting <code>self.family_p</code>? Are they all equal to 0.1 ?</p>",
              "rawMarkdown": "Thanks for your implementation `WeightedSampler` for DDP training. Besides, could you please share your experiences of setting `self.family_p`? Are they all equal to 0.1 ?"
            },
            {
              "id": 3227606,
              "postDate": "2025-06-19T05:50:20.080Z",
              "content": "<p>my solution may not be the best:</p>\n<pre><code> ():\n     ():\n\n        sampling_alpha=\n        .cfg  = cfg\n        .mode = mode\n        .df = df\n        .seismic, .velocity = load_data(df)\n        .length = ((.df)*LENGTH*subsample_ratio)\n        .subsample_ratio=subsample_ratio\n\n        \n        .group = df.groupby()\n        .mapping={\n            :  ,\n            :  ,\n                :  ,\n                : ,\n              :  ,\n              : ,\n               :  ,\n               : ,\n                   : ,\n                   : ,\n        }\n        \n\n        .family = (.mapping.keys())\n        .family_p = np.array((.mapping.values()))\n        .family_p = (.family_p/.family_p.()) ** sampling_alpha \n        .family_p = .family_p/.family_p.()\n</code></pre>",
              "rawMarkdown": "my solution may not be the best:\n```\nclass MyFWIDataset(Dataset):\n    def __init__(\n            self,\n            df,\n            cfg,\n            mode='train',\n            subsample_ratio=1.0,\n    ):\n\n        sampling_alpha=1\n        self.cfg  = cfg\n        self.mode = mode\n        self.df = df\n        self.seismic, self.velocity = load_data(df)\n        self.length = int(len(self.df)*LENGTH*subsample_ratio)\n        self.subsample_ratio=subsample_ratio\n\n        #----\n        self.group = df.groupby('dataset')\n        self.mapping={\n            'FlatVel_A':  1.31,\n            'FlatVel_B':  7.20,\n            'CurveVel_A'    :  9.19,\n            'CurveVel_B'    : 38.60,\n            'CurveFault_A'  :  4.00,\n            'CurveFault_B'  : 71.06,\n            'FlatFault_A'   :  2.58,\n            'FlatFault_B'   : 24.29,\n            'Style_A'       : 34.34,\n            'Style_B'       : 48.90,\n        }\n        #df = self.group.get_group('FlatVel_A')\n\n        self.family = list(self.mapping.keys())\n        self.family_p = np.array(list(self.mapping.values()))\n        self.family_p = (self.family_p/self.family_p.sum()) ** sampling_alpha #control the sharpness aka temperature\n        self.family_p = self.family_p/self.family_p.sum()\n\n```",
              "votes": 1
            }
          ]
        },
        {
          "id": 3227967,
          "postDate": "2025-06-19T13:57:41.163Z",
          "content": "<p>Does your college give these resources, or do you source them yourself?</p>",
          "rawMarkdown": "Does your college give these resources, or do you source them yourself?",
          "votes": 1
        }
      ]
    },
    {
      "id": 3225094,
      "postDate": "2025-06-16T01:33:48.533Z",
      "content": "<p>i wonder did anyone have similiar error:</p>\n<p>torch compile:</p>\n<pre><code> \n</code></pre>\n<p>fuse = True in Adam optimizer</p>\n<pre><code>[]:   \n[]:     \n[]: \n</code></pre>",
      "rawMarkdown": "i wonder did anyone have similiar error:\n\ntorch compile:\n```\n... torch/fx/experimental/symbolic_shapes.py:4449] [0/1] xindex is not in var_ranges, defaulting to unknown range.\n```\n\nfuse = True in Adam optimizer\n```\n[rank0]:   File \"/home/hp/app/anaconda3.11-vision/lib/python3.11/site-packages/torch/optim/adam.py\", line 676, in _fused_adam\n[rank0]:     torch._fused_adam_(\n[rank0]: RuntimeError: params, grads, exp_avgs, and exp_avg_sqs must have same dtype, device, and layout\n```\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 3225119,
          "postDate": "2025-06-16T02:53:19.747Z",
          "content": "<p>Same here. It seems that torch.compile can not be applied with fuse optimizer in distributing training</p>",
          "rawMarkdown": "Same here. It seems that torch.compile can not be applied with fuse optimizer in distributing training",
          "votes": 1,
          "replies": [
            {
              "id": 3225123,
              "postDate": "2025-06-16T02:58:03.673Z",
              "content": "<p>No. It can be done for if i use encoder only.</p>",
              "rawMarkdown": "No. It can be done for if i use encoder only.",
              "votes": 1
            },
            {
              "id": 3228015,
              "postDate": "2025-06-19T15:11:12.990Z",
              "content": "<p>UPDATE:<br>\n    The reason maybe I delete the \".contiguous()\" in the convnext_block forward function. In my experiment, when I add \".contiguous()\" into forward function, model can work with fuse optimizer.</p>",
              "rawMarkdown": "UPDATE:\n    The reason maybe I delete the \".contiguous()\" in the convnext_block forward function. In my experiment, when I add \".contiguous()\" into forward function, model can work with fuse optimizer.\n"
            },
            {
              "id": 3228440,
              "postDate": "2025-06-20T06:12:14.273Z",
              "content": "<p>thanks, it tried that. but fused optimizer doesn't seems to improve speed for my case.<br>\ncontiguous() seems to be slow too. maybe it is only for my case.</p>",
              "rawMarkdown": "thanks, it tried that. but fused optimizer doesn't seems to improve speed for my case.\ncontiguous() seems to be slow too. maybe it is only for my case.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3229148,
      "postDate": "2025-06-21T04:37:42.347Z",
      "content": "<p>i hope chatgpt is not giving me hallucinations …</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff70e64cd4ee9ba207b21c0ddfe712f81%2FSelection_999(8255).png?generation=1750480646134888&amp;alt=media\" alt=\"\"></p>\n<p>speedup trick<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8f97fcddb20b3029863fb3e0605b3493%2FSelection_999(8256).png?generation=1750480660497901&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "i hope chatgpt is not giving me hallucinations ...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff70e64cd4ee9ba207b21c0ddfe712f81%2FSelection_999(8255).png?generation=1750480646134888&alt=media)\n\nspeedup trick\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8f97fcddb20b3029863fb3e0605b3493%2FSelection_999(8256).png?generation=1750480660497901&alt=media)"
    },
    {
      "id": 3228844,
      "postDate": "2025-06-20T15:21:49.137Z",
      "content": "<h2>update!!!</h2>\n<p>for those using my code please change upsampling in decoder layer from nearest neighbour to  deconv (I,e, transpose conv)</p>\n<p>I make a mistake. we are not doing segmentation or pixel based context.for each pixel location in the target velocity map, the context can be anywhere in the semsic data (since wave can be reflected etc). more importantly the axis in semsic data is time.</p>\n<p>so we need feature to be propagated in scaling up. (i.e. we need the scaling-up weights that is dependent on feature value)<br>\ninterploation uses constant weight, which is not suitable for our task</p>\n<hr>\n<p>on a side note large conv/deconv kernel in decoder may work better. but I haven't try yet. e.g. convnext decoder</p>\n<hr>\n<p>segmentation is image to image translation problem.</p>\n<p>here are are doing wave (time aixs) to velocity. it is not \"same coordinate\" translation problem. rather it is \"converting from one modality to another, sharing one common axis(x) and yet differ in (z and time)\".</p>\n<p>it is more like encoder to semsic latent, then semsic latent to velocity latent, then decoder to upscale .</p>",
      "rawMarkdown": "##update!!!\n\nfor those using my code please change upsampling in decoder layer from nearest neighbour to  deconv (I,e, transpose conv)\n\nI make a mistake. we are not doing segmentation or pixel based context.for each pixel location in the target velocity map, the context can be anywhere in the semsic data (since wave can be reflected etc). more importantly the axis in semsic data is time.\n\nso we need feature to be propagated in scaling up. (i.e. we need the scaling-up weights that is dependent on feature value)\ninterploation uses constant weight, which is not suitable for our task\n\n---\n\non a side note large conv/deconv kernel in decoder may work better. but I haven't try yet. e.g. convnext decoder\n\n---\n\nsegmentation is image to image translation problem.\n\nhere are are doing wave (time aixs) to velocity. it is not \"same coordinate\" translation problem. rather it is \"converting from one modality to another, sharing one common axis(x) and yet differ in (z and time)\".\n\nit is more like encoder to semsic latent, then semsic latent to velocity latent, then decoder to upscale .",
      "replies": [
        {
          "id": 3229317,
          "postDate": "2025-06-21T09:58:50.523Z",
          "content": "<p>deconv is worse than nearest neighbour in my experiment. So weird.</p>",
          "rawMarkdown": "deconv is worse than nearest neighbour in my experiment. So weird."
        }
      ]
    },
    {
      "id": 3228826,
      "postDate": "2025-06-20T14:59:24.087Z",
      "content": "<p>i think my idea work (?) <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> here are some initial experiments</p>\n<ol>\n<li>given a model m</li>\n<li>do validation on set v (mae 30). obtain p = model(v)</li>\n<li>do forward wave modeling r = fwi(p)</li>\n<li>create a set (p,r) = (p, fwi(p. this set has mae loss = 18.</li>\n<li>create new train set = old train set + (p,r)</li>\n</ol>\n<p>here are the results</p>\n<hr>\n<h3>fine-tune: model initialisation=m, traindset = old train set</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc873cd4fb1906a15b3bbcae57c41b204%2FSelection_284.png?generation=1750431510814355&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h3>fine-tune: model initialisation=m, traindset = old train set + (p,r)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3fc9a283130634a1258827c80c6a161e%2FSelection_287.png?generation=1750434856419984&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h3>fine-tune: model initialisation=m, traindset = old train set + (p,r) + second round (p,r)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1167c056e31da787c2833aa1b63b8b52%2FSelection_289.png?generation=1750446143887763&amp;alt=media\" alt=\"\"></p>\n<p>training in progress … update later …</p>",
      "rawMarkdown": "i think my idea work (?) @shlomoron here are some initial experiments\n\n1. given a model m\n2. do validation on set v (mae 30). obtain p = model(v)\n3. do forward wave modeling r = fwi(p)\n4. create a set (p,r) = (p, fwi(p. this set has mae loss = 18.\n4. create new train set = old train set + (p,r)\n\nhere are the results\n\n---\n\n###fine-tune: model initialisation=m, traindset = old train set\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc873cd4fb1906a15b3bbcae57c41b204%2FSelection_284.png?generation=1750431510814355&alt=media)\n\n---\n###fine-tune: model initialisation=m, traindset = old train set + (p,r)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3fc9a283130634a1258827c80c6a161e%2FSelection_287.png?generation=1750434856419984&alt=media)\n\n---\n###fine-tune: model initialisation=m, traindset = old train set + (p,r) + second round (p,r)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1167c056e31da787c2833aa1b63b8b52%2FSelection_289.png?generation=1750446143887763&alt=media)\n\ntraining in progress ... update later ...",
      "replies": [
        {
          "id": 3228834,
          "postDate": "2025-06-20T15:07:18.347Z",
          "content": "<p>related paper (?)</p>\n<h3>[1] Physics-Consistent Data-driven Waveform Inversion with Adaptive Data Augmentation - Rojas-Gómez et al. (2020, arXiv)</h3>\n<p>chatgpt is really helpful to find paper<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffe0a723baac66d3be6682d240fdf42fb%2FSelection_286.png?generation=1750432036254103&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "\nrelated paper (?)\n\n###[1] Physics-Consistent Data-driven Waveform Inversion with Adaptive Data Augmentation - Rojas-Gómez et al. (2020, arXiv)\n\nchatgpt is really helpful to find paper\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffe0a723baac66d3be6682d240fdf42fb%2FSelection_286.png?generation=1750432036254103&alt=media)\n\n",
          "replies": [
            {
              "id": 3228878,
              "postDate": "2025-06-20T16:05:42.310Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7b81e6295e829140d29f062ad89ce39d%2FSelection_288.png?generation=1750435433963643&amp;alt=media\" alt=\"\"></p>\n<p>error of different models are show above</p>\n<p>left : m<br>\nright :fine-tune: model initialisation=m, traindset = old train set + (p,r)</p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7b81e6295e829140d29f062ad89ce39d%2FSelection_288.png?generation=1750435433963643&alt=media)\n\nerror of different models are show above\n\nleft : m\nright :fine-tune: model initialisation=m, traindset = old train set + (p,r)"
            }
          ]
        }
      ]
    },
    {
      "id": 3228586,
      "postDate": "2025-06-20T10:28:17.770Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc9235a28c9a1912684b7bd6c40ceff97%2FSelection_094.png?generation=1750415260721070&amp;alt=media\" alt=\"\"></p>\n<p>with the torch version of forward code, I wonder if direct route 1 is possible?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc9235a28c9a1912684b7bd6c40ceff97%2FSelection_094.png?generation=1750415260721070&alt=media)\n\nwith the torch version of forward code, I wonder if direct route 1 is possible?",
      "replies": [
        {
          "id": 3228595,
          "postDate": "2025-06-20T10:40:15.550Z",
          "content": "<p>navie method (just refinement)</p>\n<pre><code>\n():\n      .velocity = nn.parameter( from unet ....)\n\n():\n      just run forward wave modeling code ....\n /</code></pre>",
          "rawMarkdown": "navie method (just refinement)\n\n```\nclass myModel(nn.Module):\ndef __init__():\n      self.velocity = nn.parameter( from unet ....)\n\ndef forward(self):\n      just run forward wave modeling code ....\nsee: https://www.kaggle.com/competitions/waveform-inversion/discussion/585236#3228229\n\n```"
        }
      ]
    },
    {
      "id": 3227180,
      "postDate": "2025-06-18T16:19:16.640Z",
      "content": "<p>Didn't you continue to try using the forward modeling loss function?</p>",
      "rawMarkdown": "Didn't you continue to try using the forward modeling loss function?",
      "replies": [
        {
          "id": 3227181,
          "postDate": "2025-06-18T16:20:54.037Z",
          "content": "<p>No. I am trying few things now, including physics solution. U cannot win by just unet. This thread is more for efficient implementation </p>",
          "rawMarkdown": "No. I am trying few things now, including physics solution. U cannot win by just unet. This thread is more for efficient implementation ",
          "votes": 1,
          "replies": [
            {
              "id": 3227186,
              "postDate": "2025-06-18T16:23:38.973Z",
              "content": "<p>Yes, I think so too now.</p>",
              "rawMarkdown": "Yes, I think so too now."
            },
            {
              "id": 3227456,
              "postDate": "2025-06-19T02:58:17.093Z",
              "content": "<p>I would love to see you post about the physics loss , i had trouble doing forward modeling….</p>",
              "rawMarkdown": "I would love to see you post about the physics loss , i had trouble doing forward modeling...."
            }
          ]
        }
      ]
    },
    {
      "id": 3227458,
      "postDate": "2025-06-19T03:00:14.493Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3227610,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-19T05:54:55.447000",
      "content": "<p>how to do experiment fast. Bartley's caformer use more complicated decoder with attention and intermediate layers. i want to see if decoder is important in this competition. hence:</p>\n<ol>\n<li>I change the decoder to Bartley's</li>\n<li>load weights of previously trained encoder and freeze them</li>\n<li>I train with new decoder only<br>\nnow I can compare if Bartley's is better:</li>\n</ol>\n<p>I can further unfreeze encoder and train and train all after initial warmup.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F044a2cbc3848434184e879595c187149%2FSelection_093.png?generation=1750312490647785&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3225771,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-16T20:27:43.420000",
      "content": "<p>other worth trying:</p>\n<ul>\n<li><p>densenetersion of covnext:<br>\n<a href=\"https://huggingface.co/naver-ai/rdnet_base.nv_in1k\" target=\"_blank\">https://huggingface.co/naver-ai/rdnet_base.nv_in1k</a></p></li>\n<li><p><a href=\"https://github.com/AILab-CVC/UniRepLKNet\" target=\"_blank\">https://github.com/AILab-CVC/UniRepLKNet</a></p></li>\n<li><p><a href=\"https://github.com/LMMMEng/OverLoCK\" target=\"_blank\">https://github.com/LMMMEng/OverLoCK</a></p></li>\n</ul>\n<p>if strong encoder doesn't help, try strong decoder (e.g. attention layer, skip layer as in Bartley's caformer; unet++ or multi connect unet)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3225630,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-16T17:01:53.410000",
      "content": "<p>a stronger and fast backbone (? or/and !)<br>\n<a href=\"https://huggingface.co/timm/inception_next_base.sail_in1k\" target=\"_blank\">https://huggingface.co/timm/inception_next_base.sail_in1k</a></p>\n<p>This is another metaformer like caformer. Both are from the same lab SAIL.<br>\nNOTE: inception-next is batchnorm based, so I used \"net = torch.nn.SyncBatchNorm.convert_sync_batchnorm(net) \"</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc7232f69b59601495251b2062df593d4%2FSelection_264.png?generation=1750104201525874&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 3225656,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2025-06-16T17:30:39.687000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, when do you decided to stop training and try next experiment?</p>\n<p>As i tested multiple experiments on CurveFault_B ( with your setup - 1e-3 and with 1e-4 ) --&gt; 1e-4 got better improvement --&gt; not tested on all dataset. as we train more - gap between training loss improving 3x and validation improving with 0.75x.</p>\n<p>will you test to scale the velocity / 100 and see how your training converging ? Nice to see you in this compatition <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3225671,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-16T17:44:16.307000",
              "content": "<p>i first train a fast model:<br>\n1e-3 for 35 epoch, 1e-4 for next 15 epoch, 1e-5 for next 5 epoch.</p>\n<p>that confirms:<br>\n1e-4 at at most improve mae by 3 to4 , 1e-5 can almost improve 0.5 to 1</p>\n<p>also, from your validation, you can estimate if current mae loss = xxx,<br>\nto reduce 1 mae loss, you need e.g.  1 epoch<br>\nreduce next 1 mae, you need 2<br>\nreduce next 1 mae, you need 4<br>\nreduce next 1 mae, you need 8<br>\nreduce next 1 mae, you need 12 ….</p>\n<p>so you can estimate the final performance at the middle of training if you only have 150 or 200 epoch</p>\n<p>you should confirm this with public models (eg, small convnext) where other already train for 50, 100 and 150 epoches.</p>\n<hr>\n<p>lastly assume  model A has validation =40 at epoch 30,<br>\nfor model B to achieve +5 better at 150, B must be better than A by eg +6 at epoch 30.</p>\n<p>so you can try different backbone till epoch 30.</p>\n<p>(30,5,6 are all exmpale numbers, you need to do experiments to find out)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3225615,
      "author_name": "Wu Qiuyi",
      "author_url": "",
      "post_date": "2025-06-16T16:40:15.917000",
      "content": "<p>torch.compile(… mode='max) is very helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3227013,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-18T11:26:27.640000",
      "content": "<p>covnext-large with data sampling.<br>\nCV after training for 24hr:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F09632d6ed0ddcf89643fc6e3a1b6db9a%2FSelection_089.png?generation=1750269732895607&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 3227095,
          "author_name": "Suryansh Mishra ",
          "author_url": "",
          "post_date": "2025-06-18T13:50:11.077000",
          "content": "<p>Pardon me , but why does epoch has values after decimal?? <br>\nLike are you considering batches done also??</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3226100,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-17T08:44:44.513000",
      "content": "<p>update!!!</p>\n<p>i change machine.<br>\nOLD: 2xRTX A6000<br>\nNEW: 2xRTX 6000 ada, shift all data to SSD</p>\n<p>previous: convnext base at 30min per epoch<br>\nnew: convnext base at 21 min per epoch !!!</p>\n<p>i think a lot is due to ssd</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3226103,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2025-06-17T08:51:39.017000",
          "content": "<p>did you add weighted random sampler during train ? that would probably be a big gain. </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3226110,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-17T08:58:25.767000",
              "content": "<p>yes,i note that for different epoch the mae loss for simple family doesn't change. the imporvment comes for difficult family.<br>\nnow i am testing different backbone(upt o 30 epoch) and ifferent train config (eg fp8), so i use all train samples first.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3226504,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-17T18:51:56.767000",
              "content": "<p>thanks <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> for the hint. effects of sampling difficult train data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5873887d0fa2dd34ba9f1becfa1c5d74%2FSelection_084.png?generation=1750186244010579&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5eab51a72bb9196a19538ef380402df0%2FSelection_085.png?generation=1750186258173534&amp;alt=media\" alt=\"\"></p>\n<pre><code>class MyFWIDataset(Dataset):\n    def __init__(\n            self,\n            df,\n            cfg,\n            mode='train', \n    ):\n   ...\n\n    def __getitem__(self, ):\n\n         .mode == :\n            family = .family[np..choice(\n                np.arange(len(.family)),\n                p=.family_p\n            )]\n            df = ..get_group(family)\n            i = np..choice(df.)\n            j = np..choice()\n        :\n            i =  // \n            j =  % \n\n        seismic = .seismic[i][j]\n        velocity = .velocity[i][j]\n</code></pre>\n<p>but there is a catch: you are using error of validation IN training, i.e. you are overfitting validation data.<br>\nbut is this a problem and do we need or need not solve it? This is homework for kagglers to think about.</p>\n<p>hint: it is all about if public test is in-distribution or out-distribution</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3226567,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2025-06-17T20:22:18.710000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  is the family_p approach as similar to torch.utils.data.WeightedRandomSampler?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3226570,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-17T20:26:49.227000",
              "content": "<p>probabilty the same, but i need to use distributed parallel sampler  for DDP multi gpu training</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3226717,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2025-06-18T03:30:09.400000",
              "content": "<p>in my experiment, it overfitting the training dataset well but not well for validation <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> -&gt; </p>\n<table>\n<thead>\n<tr>\n<th>epoch</th>\n<th>CV</th>\n<th>Train</th>\n<th>Gap</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>1</td>\n<td>84.8</td>\n<td>37.7</td>\n<td>47.1</td>\n</tr>\n<tr>\n<td>2</td>\n<td>73.6</td>\n<td>29.1</td>\n<td>44.5</td>\n</tr>\n<tr>\n<td>3</td>\n<td>67.7</td>\n<td>25.6</td>\n<td>42.1</td>\n</tr>\n<tr>\n<td>4</td>\n<td>63.7</td>\n<td>23.8</td>\n<td>38.1</td>\n</tr>\n<tr>\n<td>5</td>\n<td>57.8</td>\n<td>19.8</td>\n<td>38</td>\n</tr>\n<tr>\n<td>6</td>\n<td>55.7</td>\n<td>18.4</td>\n<td>37.3</td>\n</tr>\n<tr>\n<td>7</td>\n<td>55.6</td>\n<td>18.8</td>\n<td>36.8</td>\n</tr>\n<tr>\n<td>8</td>\n<td>55.0</td>\n<td>18.6</td>\n<td>36.4</td>\n</tr>\n<tr>\n<td>9</td>\n<td>51.8</td>\n<td>15.9</td>\n<td>35. 9</td>\n</tr>\n</tbody>\n</table>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3226800,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-18T05:24:20.713000",
              "content": "<p>something is very wrong, without code or train parameters/train log, i cannot diagnose the error.<br>\nplease post them if you can.</p>\n<hr>\n<p>i can only say that:<br>\n1) assume that you are using my fold.csv<br>\n2) then,  training at constant 1e-3, validation loss is always better than train loss(moving average) until they converge.</p>\n<p>they converge at:<br>\nmae loss = 35 for covnext small,<br>\n33 for base<br>\n30 for large</p>\n<p>(this is also the approximate LB score)</p>\n<hr>\n<p>i think there is a bug in your code. it is easy to check</p>\n<ul>\n<li>since you are sample more difficult sample, train loss is higher than the unsampled version (btw: high loss means greater gradient backprop, so faster converge)</li>\n</ul>\n<p>did you subsample more difficult samples or mistaken sample the easier ones?</p>\n<ul>\n<li>getting train loss at 37.7 is impossible  (assume you start from imagenet pretrain). you can check train and validation error of imagenet pretrain initialisation. this is initial gap. </li>\n</ul>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3226929,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2025-06-18T09:08:19.213000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - its because <strong>since you are sample more difficult sample, train loss is higher than the unsampled version (btw: high loss means greater gradient backprop, so faster converge)</strong></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3227061,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-18T12:39:24.733000",
              "content": "<p>it is either you oversample (too many difficult and too few simple) or wrong/under sampling  (too many simple and too few difficult).</p>\n<p>assuming there is no other bug …</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3227430,
              "author_name": "AnnieGo",
              "author_url": "",
              "post_date": "2025-06-19T02:35:03.770000",
              "content": "<p>Thanks for your implementation <code>WeightedSampler</code> for DDP training. Besides, could you please share your experiences of setting <code>self.family_p</code>? Are they all equal to 0.1 ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3227606,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-19T05:50:20.080000",
              "content": "<p>my solution may not be the best:</p>\n<pre><code> ():\n     ():\n\n        sampling_alpha=\n        .cfg  = cfg\n        .mode = mode\n        .df = df\n        .seismic, .velocity = load_data(df)\n        .length = ((.df)*LENGTH*subsample_ratio)\n        .subsample_ratio=subsample_ratio\n\n        \n        .group = df.groupby()\n        .mapping={\n            :  ,\n            :  ,\n                :  ,\n                : ,\n              :  ,\n              : ,\n               :  ,\n               : ,\n                   : ,\n                   : ,\n        }\n        \n\n        .family = (.mapping.keys())\n        .family_p = np.array((.mapping.values()))\n        .family_p = (.family_p/.family_p.()) ** sampling_alpha \n        .family_p = .family_p/.family_p.()\n</code></pre>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3227967,
          "author_name": "Justin Wallace",
          "author_url": "",
          "post_date": "2025-06-19T13:57:41.163000",
          "content": "<p>Does your college give these resources, or do you source them yourself?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3225094,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-16T01:33:48.533000",
      "content": "<p>i wonder did anyone have similiar error:</p>\n<p>torch compile:</p>\n<pre><code> \n</code></pre>\n<p>fuse = True in Adam optimizer</p>\n<pre><code>[]:   \n[]:     \n[]: \n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 3225119,
          "author_name": "I2nfinit3y",
          "author_url": "",
          "post_date": "2025-06-16T02:53:19.747000",
          "content": "<p>Same here. It seems that torch.compile can not be applied with fuse optimizer in distributing training</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3225123,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-16T02:58:03.673000",
              "content": "<p>No. It can be done for if i use encoder only.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3228015,
              "author_name": "I2nfinit3y",
              "author_url": "",
              "post_date": "2025-06-19T15:11:12.990000",
              "content": "<p>UPDATE:<br>\n    The reason maybe I delete the \".contiguous()\" in the convnext_block forward function. In my experiment, when I add \".contiguous()\" into forward function, model can work with fuse optimizer.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3228440,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-20T06:12:14.273000",
              "content": "<p>thanks, it tried that. but fused optimizer doesn't seems to improve speed for my case.<br>\ncontiguous() seems to be slow too. maybe it is only for my case.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3229148,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-21T04:37:42.347000",
      "content": "<p>i hope chatgpt is not giving me hallucinations …</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff70e64cd4ee9ba207b21c0ddfe712f81%2FSelection_999(8255).png?generation=1750480646134888&amp;alt=media\" alt=\"\"></p>\n<p>speedup trick<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8f97fcddb20b3029863fb3e0605b3493%2FSelection_999(8256).png?generation=1750480660497901&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3228844,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-20T15:21:49.137000",
      "content": "<h2>update!!!</h2>\n<p>for those using my code please change upsampling in decoder layer from nearest neighbour to  deconv (I,e, transpose conv)</p>\n<p>I make a mistake. we are not doing segmentation or pixel based context.for each pixel location in the target velocity map, the context can be anywhere in the semsic data (since wave can be reflected etc). more importantly the axis in semsic data is time.</p>\n<p>so we need feature to be propagated in scaling up. (i.e. we need the scaling-up weights that is dependent on feature value)<br>\ninterploation uses constant weight, which is not suitable for our task</p>\n<hr>\n<p>on a side note large conv/deconv kernel in decoder may work better. but I haven't try yet. e.g. convnext decoder</p>\n<hr>\n<p>segmentation is image to image translation problem.</p>\n<p>here are are doing wave (time aixs) to velocity. it is not \"same coordinate\" translation problem. rather it is \"converting from one modality to another, sharing one common axis(x) and yet differ in (z and time)\".</p>\n<p>it is more like encoder to semsic latent, then semsic latent to velocity latent, then decoder to upscale .</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3229317,
          "author_name": "I2nfinit3y",
          "author_url": "",
          "post_date": "2025-06-21T09:58:50.523000",
          "content": "<p>deconv is worse than nearest neighbour in my experiment. So weird.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3228826,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-20T14:59:24.087000",
      "content": "<p>i think my idea work (?) <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> here are some initial experiments</p>\n<ol>\n<li>given a model m</li>\n<li>do validation on set v (mae 30). obtain p = model(v)</li>\n<li>do forward wave modeling r = fwi(p)</li>\n<li>create a set (p,r) = (p, fwi(p. this set has mae loss = 18.</li>\n<li>create new train set = old train set + (p,r)</li>\n</ol>\n<p>here are the results</p>\n<hr>\n<h3>fine-tune: model initialisation=m, traindset = old train set</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc873cd4fb1906a15b3bbcae57c41b204%2FSelection_284.png?generation=1750431510814355&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h3>fine-tune: model initialisation=m, traindset = old train set + (p,r)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3fc9a283130634a1258827c80c6a161e%2FSelection_287.png?generation=1750434856419984&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h3>fine-tune: model initialisation=m, traindset = old train set + (p,r) + second round (p,r)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1167c056e31da787c2833aa1b63b8b52%2FSelection_289.png?generation=1750446143887763&amp;alt=media\" alt=\"\"></p>\n<p>training in progress … update later …</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3228834,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-06-20T15:07:18.347000",
          "content": "<p>related paper (?)</p>\n<h3>[1] Physics-Consistent Data-driven Waveform Inversion with Adaptive Data Augmentation - Rojas-Gómez et al. (2020, arXiv)</h3>\n<p>chatgpt is really helpful to find paper<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ffe0a723baac66d3be6682d240fdf42fb%2FSelection_286.png?generation=1750432036254103&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 3228878,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-20T16:05:42.310000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7b81e6295e829140d29f062ad89ce39d%2FSelection_288.png?generation=1750435433963643&amp;alt=media\" alt=\"\"></p>\n<p>error of different models are show above</p>\n<p>left : m<br>\nright :fine-tune: model initialisation=m, traindset = old train set + (p,r)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3228586,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-20T10:28:17.770000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc9235a28c9a1912684b7bd6c40ceff97%2FSelection_094.png?generation=1750415260721070&amp;alt=media\" alt=\"\"></p>\n<p>with the torch version of forward code, I wonder if direct route 1 is possible?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3228595,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-06-20T10:40:15.550000",
          "content": "<p>navie method (just refinement)</p>\n<pre><code>\n():\n      .velocity = nn.parameter( from unet ....)\n\n():\n      just run forward wave modeling code ....\n /</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3227180,
      "author_name": "guo dashuai",
      "author_url": "",
      "post_date": "2025-06-18T16:19:16.640000",
      "content": "<p>Didn't you continue to try using the forward modeling loss function?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3227181,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-06-18T16:20:54.037000",
          "content": "<p>No. I am trying few things now, including physics solution. U cannot win by just unet. This thread is more for efficient implementation </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3227186,
              "author_name": "guo dashuai",
              "author_url": "",
              "post_date": "2025-06-18T16:23:38.973000",
              "content": "<p>Yes, I think so too now.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3227456,
              "author_name": "Suryansh Mishra ",
              "author_url": "",
              "post_date": "2025-06-19T02:58:17.093000",
              "content": "<p>I would love to see you post about the physics loss , i had trouble doing forward modeling….</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3227458,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-19T03:00:14.493000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3225091": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8a05eebaa2abfdbbc4903c0972ff0ba7%2FSelection_252.png?generation=1750036637016034&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ec93d83b8121bf5e188b3b901555fc2%2FSelection_261.png?generation=1750068630074827&alt=media)\ncode: https://www.kaggle.com/datasets/hengck23/kaggle-competition-waveform-inversion-discussion-0\n\n---\n\ngpu = 2x RTX 6000 (close to 4090)\nvalid samples = 20 np arrays = 2 npy from each family (1,10 npy, following Bartley's fold0) \ntrain samples = rest of arrays\n(no data subsampling used)\n\n##training speed:\n- convnext base at 30min per epoch\n- convnext small at 24min\n- covnnext large at 40min\n\nat most 3 days for 150 epoch(?)\n\n##implementation:\n- basically Bartley'smodel with my modifications\n  - my own stem\n  - batch norm in unet\n  - instance-norm2d #1 in convnext block  but remove contiguous() in forward\n\n- train hyper-parameters:\n  - I wrote my own trainer, the train MAE shown is moving average over the last 100 batches. validation MAE exclude flip TTA\n  - batch size = 64 for each GPU (faster if I use batch size =128, which I intend to use at later epoch of fine tunning with smaller lr). i did not use sync batch norm for multi-gpu training, so batch size has to be large\n  - bfloat16, hence autocast scaler is not required\n  - Bartley's EMA #1\n  - I cannot use optimizer fuse and torch compile. i have errors below.\n  - I cannot use FP8 training (torchAO and FSDP2) for now\n\n\nfor other configurations, please refer to code and log in the public dataset link\n\n```\n#1 : I find this to have great effect in accuracy\n```ow",
    "3227610": "how to do experiment fast. Bartley's caformer use more complicated decoder with attention and intermediate layers. i want to see if decoder is important in this competition. hence:\n1. I change the decoder to Bartley's\n2. load weights of previously trained encoder and freeze them\n3. I train with new decoder only\nnow I can compare if Bartley's is better:\n\nI can further unfreeze encoder and train and train all after initial warmup.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F044a2cbc3848434184e879595c187149%2FSelection_093.png?generation=1750312490647785&alt=media)",
    "3225771": "other worth trying:\n- densenetersion of covnext:\nhttps://huggingface.co/naver-ai/rdnet_base.nv_in1k\n\n- https://github.com/AILab-CVC/UniRepLKNet\n- https://github.com/LMMMEng/OverLoCK\n\nif strong encoder doesn't help, try strong decoder (e.g. attention layer, skip layer as in Bartley's caformer; unet++ or multi connect unet)",
    "3225630": "a stronger and fast backbone (? or/and !)\nhttps://huggingface.co/timm/inception_next_base.sail_in1k\n\nThis is another metaformer like caformer. Both are from the same lab SAIL.\nNOTE: inception-next is batchnorm based, so I used \"net = torch.nn.SyncBatchNorm.convert_sync_batchnorm(net) \"\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc7232f69b59601495251b2062df593d4%2FSelection_264.png?generation=1750104201525874&alt=media)",
    "3225615": "torch.compile(... mode='max) is very helpful.",
    "3227013": "covnext-large with data sampling.\nCV after training for 24hr:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F09632d6ed0ddcf89643fc6e3a1b6db9a%2FSelection_089.png?generation=1750269732895607&alt=media)",
    "3226100": "update!!!\n\ni change machine.\nOLD: 2xRTX A6000\nNEW: 2xRTX 6000 ada, shift all data to SSD\n\nprevious: convnext base at 30min per epoch\nnew: convnext base at 21 min per epoch !!!\n\ni think a lot is due to ssd",
    "3225094": "i wonder did anyone have similiar error:\n\ntorch compile:\n```\n... torch/fx/experimental/symbolic_shapes.py:4449] [0/1] xindex is not in var_ranges, defaulting to unknown range.\n```\n\nfuse = True in Adam optimizer\n```\n[rank0]:   File \"/home/hp/app/anaconda3.11-vision/lib/python3.11/site-packages/torch/optim/adam.py\", line 676, in _fused_adam\n[rank0]:     torch._fused_adam_(\n[rank0]: RuntimeError: params, grads, exp_avgs, and exp_avg_sqs must have same dtype, device, and layout\n```\n\n",
    "3229148": "i hope chatgpt is not giving me hallucinations ...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff70e64cd4ee9ba207b21c0ddfe712f81%2FSelection_999(8255).png?generation=1750480646134888&alt=media)\n\nspeedup trick\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8f97fcddb20b3029863fb3e0605b3493%2FSelection_999(8256).png?generation=1750480660497901&alt=media)",
    "3228844": "##update!!!\n\nfor those using my code please change upsampling in decoder layer from nearest neighbour to  deconv (I,e, transpose conv)\n\nI make a mistake. we are not doing segmentation or pixel based context.for each pixel location in the target velocity map, the context can be anywhere in the semsic data (since wave can be reflected etc). more importantly the axis in semsic data is time.\n\nso we need feature to be propagated in scaling up. (i.e. we need the scaling-up weights that is dependent on feature value)\ninterploation uses constant weight, which is not suitable for our task\n\n---\n\non a side note large conv/deconv kernel in decoder may work better. but I haven't try yet. e.g. convnext decoder\n\n---\n\nsegmentation is image to image translation problem.\n\nhere are are doing wave (time aixs) to velocity. it is not \"same coordinate\" translation problem. rather it is \"converting from one modality to another, sharing one common axis(x) and yet differ in (z and time)\".\n\nit is more like encoder to semsic latent, then semsic latent to velocity latent, then decoder to upscale .",
    "3228826": "i think my idea work (?) @shlomoron here are some initial experiments\n\n1. given a model m\n2. do validation on set v (mae 30). obtain p = model(v)\n3. do forward wave modeling r = fwi(p)\n4. create a set (p,r) = (p, fwi(p. this set has mae loss = 18.\n4. create new train set = old train set + (p,r)\n\nhere are the results\n\n---\n\n###fine-tune: model initialisation=m, traindset = old train set\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc873cd4fb1906a15b3bbcae57c41b204%2FSelection_284.png?generation=1750431510814355&alt=media)\n\n---\n###fine-tune: model initialisation=m, traindset = old train set + (p,r)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3fc9a283130634a1258827c80c6a161e%2FSelection_287.png?generation=1750434856419984&alt=media)\n\n---\n###fine-tune: model initialisation=m, traindset = old train set + (p,r) + second round (p,r)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1167c056e31da787c2833aa1b63b8b52%2FSelection_289.png?generation=1750446143887763&alt=media)\n\ntraining in progress ... update later ...",
    "3228586": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc9235a28c9a1912684b7bd6c40ceff97%2FSelection_094.png?generation=1750415260721070&alt=media)\n\nwith the torch version of forward code, I wonder if direct route 1 is possible?",
    "3227180": "Didn't you continue to try using the forward modeling loss function?",
    "3227458": ""
  }
}