{
  "id": 583896,
  "title": "Tips to speed up training",
  "url": "/competitions/waveform-inversion/discussion/583896",
  "author_name": "Darragh",
  "post_date": "2025-06-10T09:07:29.520000",
  "votes": 66,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Here is what I use to speed up training, a combination of ideas in the forum, some from Anthropic's Claude Opus based on the training scripts….</p>\n<ul>\n<li>Memory map in dataloading <code>np.load(path, mmap_mode='r')</code> like in many public scripts. </li>\n<li>Mixed precision training, <a href=\"https://docs.pytorch.org/docs/stable/notes/amp_examples.html\" target=\"_blank\">examples</a></li>\n<li>Compile model in torch, <code>model = torch.compile(model, **dict(fullgraph=False, dynamic=False, mode=\"default\"))</code>. </li>\n<li>Subsample easy families, eg., in <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/582785\" target=\"_blank\">Bartley's CV table</a> , <code>FlatVel_A</code> scores 1.31 - subsample it with torch <a href=\"https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler\" target=\"_blank\">WeightedRandomSampler</a>. </li>\n<li>Downsize along time axis, - a bit risky. </li>\n<li>Increase prefetch <code>prefetch_factor</code> in DataLoader, not sure if this helps, but I use it.</li>\n<li><code>channels_last</code> format in torch (small difference, if any)</li>\n<li>FP8 training, <a href=\"https://pytorch.org/blog/training-using-float8-fsdp2/\" target=\"_blank\">linky</a>, gain seems to be limited for smaller models. Only works for newer GPU's - 4090, 5090, H100, etc. </li>\n</ul>\n<p>I did not try the last one. Any other ideas to make the most of the last three weeks ?</p>",
  "messages": [
    {
      "id": 3221005,
      "postDate": "2025-06-10T09:07:29.520Z",
      "content": "<p>Here is what I use to speed up training, a combination of ideas in the forum, some from Anthropic's Claude Opus based on the training scripts….</p>\n<ul>\n<li>Memory map in dataloading <code>np.load(path, mmap_mode='r')</code> like in many public scripts. </li>\n<li>Mixed precision training, <a href=\"https://docs.pytorch.org/docs/stable/notes/amp_examples.html\" target=\"_blank\">examples</a></li>\n<li>Compile model in torch, <code>model = torch.compile(model, **dict(fullgraph=False, dynamic=False, mode=\"default\"))</code>. </li>\n<li>Subsample easy families, eg., in <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/582785\" target=\"_blank\">Bartley's CV table</a> , <code>FlatVel_A</code> scores 1.31 - subsample it with torch <a href=\"https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler\" target=\"_blank\">WeightedRandomSampler</a>. </li>\n<li>Downsize along time axis, - a bit risky. </li>\n<li>Increase prefetch <code>prefetch_factor</code> in DataLoader, not sure if this helps, but I use it.</li>\n<li><code>channels_last</code> format in torch (small difference, if any)</li>\n<li>FP8 training, <a href=\"https://pytorch.org/blog/training-using-float8-fsdp2/\" target=\"_blank\">linky</a>, gain seems to be limited for smaller models. Only works for newer GPU's - 4090, 5090, H100, etc. </li>\n</ul>\n<p>I did not try the last one. Any other ideas to make the most of the last three weeks ?</p>",
      "rawMarkdown": "Here is what I use to speed up training, a combination of ideas in the forum, some from Anthropic's Claude Opus based on the training scripts....\n\n- Memory map in dataloading `np.load(path, mmap_mode='r')` like in many public scripts. \n- Mixed precision training, [examples](https://docs.pytorch.org/docs/stable/notes/amp_examples.html)\n- Compile model in torch, `model = torch.compile(model, **dict(fullgraph=False, dynamic=False, mode=\"default\"))`. \n- Subsample easy families, eg., in [Bartley's CV table](https://www.kaggle.com/competitions/waveform-inversion/discussion/582785) , `FlatVel_A` scores 1.31 - subsample it with torch [WeightedRandomSampler](https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler). \n- Downsize along time axis, - a bit risky. \n- Increase prefetch `prefetch_factor` in DataLoader, not sure if this helps, but I use it.\n- `channels_last` format in torch (small difference, if any)\n- FP8 training, [linky](https://pytorch.org/blog/training-using-float8-fsdp2/), gain seems to be limited for smaller models. Only works for newer GPU's - 4090, 5090, H100, etc. \n\n\nI did not try the last one. Any other ideas to make the most of the last three weeks ?",
      "votes": 65
    },
    {
      "id": 3221016,
      "postDate": "2025-06-10T09:25:55.813Z",
      "content": "<p>Fused optimizer</p>",
      "rawMarkdown": "Fused optimizer",
      "votes": 7,
      "replies": [
        {
          "id": 3221224,
          "postDate": "2025-06-10T15:33:36.080Z",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> would you mind to share what speed up Fused optimiser gave; I remember it not being huge but might try again. </p>",
          "rawMarkdown": "@harshitsheoran would you mind to share what speed up Fused optimiser gave; I remember it not being huge but might try again. ",
          "replies": [
            {
              "id": 3221242,
              "postDate": "2025-06-10T16:06:40.017Z",
              "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> Just going from memory, Compile+Fused gave me 30-35%, I think fused alone was 10%+</p>\n<p><code>optimizer = optim.AdamW(model.parameters(), lr=CFG.lr, weight_decay=CFG.wd, fused=True)</code></p>",
              "rawMarkdown": "@darraghdog Just going from memory, Compile+Fused gave me 30-35%, I think fused alone was 10%+\n\n`optimizer = optim.AdamW(model.parameters(), lr=CFG.lr, weight_decay=CFG.wd, fused=True)`",
              "votes": 4
            },
            {
              "id": 3221247,
              "postDate": "2025-06-10T16:15:13.290Z",
              "content": "<p>Thanks, will try fused optimiser again.</p>\n<p>Maybe I had it wrong how much torch.compile sped up; or else on Hopper the difference is not so big as I got on a100. <a href=\"https://lightning.ai/blog/training-compiled-pytorch-2.0-with-pytorch-lightning/\" target=\"_blank\">This</a> shows a100 speed ups. </p>",
              "rawMarkdown": "Thanks, will try fused optimiser again.\n\nMaybe I had it wrong how much torch.compile sped up; or else on Hopper the difference is not so big as I got on a100. [This](https://lightning.ai/blog/training-compiled-pytorch-2.0-with-pytorch-lightning/) shows a100 speed ups. "
            },
            {
              "id": 3225677,
              "postDate": "2025-06-16T17:55:53.353Z",
              "content": "<p>The Compile+Fused gave really good boost.</p>",
              "rawMarkdown": "The Compile+Fused gave really good boost.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3224767,
      "postDate": "2025-06-15T12:55:11.847Z",
      "content": "<p>Thanks for sharing, it's friendly to newcomers!</p>",
      "rawMarkdown": "Thanks for sharing, it's friendly to newcomers!",
      "votes": 3,
      "replies": [
        {
          "id": 3225546,
          "postDate": "2025-06-16T14:38:19.823Z",
          "content": "<p>yes, so good for me too. thanks for author.</p>",
          "rawMarkdown": "yes, so good for me too. thanks for author.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3223457,
      "postDate": "2025-06-13T10:26:12.620Z",
      "content": "<p>Thanks for your kind shareing! Would you tell me the reason why not use<br>\n<code>torch.compile(mode = \"reduce overhead\")</code>?</p>",
      "rawMarkdown": "Thanks for your kind shareing! Would you tell me the reason why not use\n`torch.compile(mode = \"reduce overhead\")`?",
      "votes": 1,
      "replies": [
        {
          "id": 3225685,
          "postDate": "2025-06-16T18:01:20.110Z",
          "content": "<p>I played around with a few settings and got some issues with some settings and didn’t see a huge difference in the others which worked, so settled on this. I see most people use reduce overhead, I might try again. Thanks 🙏 </p>",
          "rawMarkdown": "I played around with a few settings and got some issues with some settings and didn’t see a huge difference in the others which worked, so settled on this. I see most people use reduce overhead, I might try again. Thanks 🙏 "
        }
      ]
    },
    {
      "id": 3221494,
      "postDate": "2025-06-11T04:47:40.400Z",
      "content": "<p>Thanks for your kind sharing! But when I combine torch.compile and fused optimizer, codes always raise a error like \"RuntimeError: params, grads, exp_avgs, and exp_avg_sqs must have same dtype, device, and layout\" on RTX A40. Have you met this?</p>",
      "rawMarkdown": "Thanks for your kind sharing! But when I combine torch.compile and fused optimizer, codes always raise a error like \"RuntimeError: params, grads, exp_avgs, and exp_avg_sqs must have same dtype, device, and layout\" on RTX A40. Have you met this?",
      "votes": 2,
      "replies": [
        {
          "id": 3221606,
          "postDate": "2025-06-11T08:37:08.770Z",
          "content": "<p>Sorry, I have not seen this. I'd recommend use your choice of free [deepseek.com, chatgpt.com, etc] or paid [clause-sonnet,openai-o3, etc] LLM apis. This is my workflow for the competition 😀 ps. I test and my gains on a100 for fused optimiser were not much. </p>\n<pre><code>\n\nAnd write your question\n</code></pre>",
          "rawMarkdown": "Sorry, I have not seen this. I'd recommend use your choice of free [deepseek.com, chatgpt.com, etc] or paid [clause-sonnet,openai-o3, etc] LLM apis. This is my workflow for the competition 😀 ps. I test and my gains on a100 for fused optimiser were not much. \n```\n'''RELEVANT_CODEBLOCK'''\n'''ERROR_TRACE'''\nAnd write your question\n```\n ",
          "votes": 1
        },
        {
          "id": 3227150,
          "postDate": "2025-06-18T15:26:12.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/i2nfinit3y\" target=\"_blank\">@i2nfinit3y</a> I got this error as well. Did you fix it?</p>",
          "rawMarkdown": "@i2nfinit3y I got this error as well. Did you fix it?"
        }
      ]
    },
    {
      "id": 3226959,
      "postDate": "2025-06-18T10:00:27.457Z",
      "content": "<p>Thanks for sharing, great opportunity.🌟</p>",
      "rawMarkdown": "Thanks for sharing, great opportunity.🌟"
    },
    {
      "id": 3224857,
      "postDate": "2025-06-15T15:19:32.253Z",
      "content": "<p>I get this message, since I started using the caformer code:</p>\n<pre><code>/usr//lib/python3/dist-packages/torch/autograd/graph.py:: UserWarning: Grad strides do  match bucket view strides. This may indicate grad was  created according   gradient layout contract,    param's strides changed  DDP was constructed.  This   an ,  may impair performance.\n\ns      grad.sizes() = [, , , ], strides() = [, , , ]\n\ns      bucket_view.sizes() = [, , , ], strides() = [, , , ] (Triggered internally  /pytorch/torch/csrc/distributed/c10d/reducer.cpp:)\n\ns         Variable._execution_engine.run_backward(  \n</code></pre>\n<p>Does anyone know how to fix it or what's causing it?<br>\nI tried checking for non-contiguous tensors, but didnt find any.<br>\nI feel like my training slows down alot</p>",
      "rawMarkdown": "I get this message, since I started using the caformer code:\n\n```\n/usr/local/lib/python3.11/dist-packages/torch/autograd/graph.py:823: UserWarning: Grad strides do not match bucket view strides. This may indicate grad was not created according to the gradient layout contract, or that the param's strides changed since DDP was constructed.  This is not an error, but may impair performance.\n\n485.7s\t55\tgrad.sizes() = [64, 128, 1, 1], strides() = [128, 1, 128, 128]\n\n485.7s\t56\tbucket_view.sizes() = [64, 128, 1, 1], strides() = [128, 1, 1, 1] (Triggered internally at /pytorch/torch/csrc/distributed/c10d/reducer.cpp:327.)\n\n485.7s\t57\t  return Variable._execution_engine.run_backward(  # Calls into the C++ engine to run the backward pass\n```\n\nDoes anyone know how to fix it or what's causing it?\nI tried checking for non-contiguous tensors, but didnt find any.\nI feel like my training slows down alot",
      "replies": [
        {
          "id": 3225010,
          "postDate": "2025-06-15T19:47:04.440Z",
          "content": "<p>I got the same. I did not dig into it.</p>",
          "rawMarkdown": "I got the same. I did not dig into it.",
          "votes": 1
        },
        {
          "id": 3225182,
          "postDate": "2025-06-16T04:34:30.133Z",
          "content": "<p>Hey Marius,</p>\n<p>That warning seems to be related to how <code>DistributedDataParallel (DDP)</code> handles gradient bucketing — especially when the tensor strides mismatch between the parameter and the bucket view. Since you're using Caformer, it could be that one of the layers or reshaping operations is altering tensor memory layout in a way that DDP doesn't expect.</p>\n<p>A few ideas to try:</p>\n<ul>\n<li>Check if any tensor operations like <code>.permute()</code>, <code>.transpose()</code>, or <code>.view()</code> are used without following them with <code>.contiguous()</code>. DDP can be sensitive to non-contiguous gradients.</li>\n<li>You can try wrapping your model parameters or outputs with <code>.contiguous()</code> where it makes sense, especially just before loss computation or backpropagation.</li>\n<li>Make sure the model is initialized <strong>before</strong> wrapping with <code>DDP</code>, and that no parameter reshaping is done afterward.</li>\n<li>Also worth trying <code>model = model.to(memory_format=torch.channels_last)</code> early in your pipeline — Caformer might be using channels_last format internally, which can create stride inconsistencies.</li>\n</ul>\n<p>This warning doesn't usually cause incorrect results, but as you noted, it <strong>can</strong> hurt performance if grad bucketing isn’t efficient. If you’re seeing slower training, it might be worth running a profiler (<code>torch.profiler</code>) to verify where the slowdown is coming from.</p>\n<p>Let us know if you find a clean workaround — I'm sure others using Caformer in this comp are hitting similar issues! 🔍</p>",
          "rawMarkdown": "Hey Marius,\n\nThat warning seems to be related to how `DistributedDataParallel (DDP)` handles gradient bucketing — especially when the tensor strides mismatch between the parameter and the bucket view. Since you're using Caformer, it could be that one of the layers or reshaping operations is altering tensor memory layout in a way that DDP doesn't expect.\n\nA few ideas to try:\n\n* Check if any tensor operations like `.permute()`, `.transpose()`, or `.view()` are used without following them with `.contiguous()`. DDP can be sensitive to non-contiguous gradients.\n* You can try wrapping your model parameters or outputs with `.contiguous()` where it makes sense, especially just before loss computation or backpropagation.\n* Make sure the model is initialized **before** wrapping with `DDP`, and that no parameter reshaping is done afterward.\n* Also worth trying `model = model.to(memory_format=torch.channels_last)` early in your pipeline — Caformer might be using channels\\_last format internally, which can create stride inconsistencies.\n\nThis warning doesn't usually cause incorrect results, but as you noted, it **can** hurt performance if grad bucketing isn’t efficient. If you’re seeing slower training, it might be worth running a profiler (`torch.profiler`) to verify where the slowdown is coming from.\n\nLet us know if you find a clean workaround — I'm sure others using Caformer in this comp are hitting similar issues! 🔍"
        }
      ]
    },
    {
      "id": 3223478,
      "postDate": "2025-06-13T11:04:37.523Z",
      "content": "<p>Use torch.compile, mixed precision (torch.cuda.amp), and gradient checkpointing to boost speed and reduce memory. Optimize DataLoader with more workers, pin_memory, and persistent_workers for faster data throughput.</p>",
      "rawMarkdown": "Use torch.compile, mixed precision (torch.cuda.amp), and gradient checkpointing to boost speed and reduce memory. Optimize DataLoader with more workers, pin_memory, and persistent_workers for faster data throughput.",
      "replies": [
        {
          "id": 3224848,
          "postDate": "2025-06-15T15:04:07.663Z",
          "content": "<p>Gradient checkpointing reduces memory but it doesn't speed up training.</p>",
          "rawMarkdown": "Gradient checkpointing reduces memory but it doesn't speed up training."
        },
        {
          "id": 3225053,
          "postDate": "2025-06-15T23:17:57.933Z",
          "content": "<p>It could if your vram is really low and your original batch size is small.</p>",
          "rawMarkdown": "It could if your vram is really low and your original batch size is small."
        }
      ]
    },
    {
      "id": 3223318,
      "postDate": "2025-06-13T06:47:37.250Z",
      "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> thank you, except FP8 training - training speed improved 18%-20% with fuse=True on rtx 5090</p>",
      "rawMarkdown": "@darraghdog thank you, except FP8 training - training speed improved 18%-20% with fuse=True on rtx 5090"
    },
    {
      "id": 3221209,
      "postDate": "2025-06-10T15:14:18.280Z",
      "content": "<p>How do you avoid NaN issues with fp16 training ?</p>",
      "rawMarkdown": "How do you avoid NaN issues with fp16 training ?",
      "replies": [
        {
          "id": 3221217,
          "postDate": "2025-06-10T15:26:51.867Z",
          "content": "<p>I get it sometimes, but not too often; feed in fp32 inputs &amp; labels and use <code>from torch.amp import autocast</code> like below.</p>\n<pre><code>                with autocast(=, =torch.bfloat16):\n                    output = model(batch)\n</code></pre>",
          "rawMarkdown": "I get it sometimes, but not too often; feed in fp32 inputs & labels and use `from torch.amp import autocast` like below.\n```\n                with autocast(device_type='cuda', dtype=torch.bfloat16):\n                    output = model(batch)\n```",
          "votes": 4
        },
        {
          "id": 3226859,
          "postDate": "2025-06-18T07:09:25.310Z",
          "content": "<p><code>with autocast(device_type='cuda', dtype=torch.float32):</code> </p>",
          "rawMarkdown": "`with autocast(device_type='cuda', dtype=torch.float32):` \n"
        }
      ]
    },
    {
      "id": 3221204,
      "postDate": "2025-06-10T15:01:56.650Z",
      "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> Hi, what are the pros and cons of loading all the .npy files with <code>mmap_mode='r'</code> in <code>__init__</code> versus lazy loading with <code>mmap_mode='r'</code> in <code>__getitem__</code>?</p>",
      "rawMarkdown": "@darraghdog Hi, what are the pros and cons of loading all the .npy files with `mmap_mode='r'` in `__init__` versus lazy loading with `mmap_mode='r'` in `__getitem__`?",
      "replies": [
        {
          "id": 3221222,
          "postDate": "2025-06-10T15:32:37.730Z",
          "content": "<p>I just tried in <code>__getitem__</code> .. .would need some research to see the difference. </p>",
          "rawMarkdown": "I just tried in `__getitem__` .. .would need some research to see the difference. "
        }
      ]
    },
    {
      "id": 3221189,
      "postDate": "2025-06-10T14:30:51.920Z",
      "content": "<p>Thanks, I get about 13% speedup with torch.compile().</p>",
      "rawMarkdown": "Thanks, I get about 13% speedup with torch.compile().",
      "replies": [
        {
          "id": 3221195,
          "postDate": "2025-06-10T14:42:00.633Z",
          "content": "<p>Same model 12hrs to 7hrs on a100 for <code>torch.compile()</code> like above - will double check I got that correct. Does that sound right for others ?</p>",
          "rawMarkdown": "Same model 12hrs to 7hrs on a100 for `torch.compile()` like above - will double check I got that correct. Does that sound right for others ?",
          "votes": 2
        }
      ]
    },
    {
      "id": 3221116,
      "postDate": "2025-06-10T12:27:31.543Z",
      "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> - How much training speed improved? </p>",
      "rawMarkdown": "@darraghdog - How much training speed improved? ",
      "replies": [
        {
          "id": 3221196,
          "postDate": "2025-06-10T14:44:52.290Z",
          "content": "<p>Depends on your baseline, amp / compile / subsampling … all were the big hitters for me - in the 1.5x to 2x range. Subsampling and downsizing (a type of subsampling) depends on your choices on how aggressively to do it. </p>",
          "rawMarkdown": "Depends on your baseline, amp / compile / subsampling ... all were the big hitters for me - in the 1.5x to 2x range. Subsampling and downsizing (a type of subsampling) depends on your choices on how aggressively to do it. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3227462,
      "postDate": "2025-06-19T03:02:18.767Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 3225181,
      "postDate": "2025-06-16T04:33:16.623Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3225683,
      "postDate": "2025-06-16T18:00:08.483Z",
      "content": "<p>Thanks for it.</p>",
      "rawMarkdown": "Thanks for it."
    },
    {
      "id": 3222491,
      "postDate": "2025-06-12T08:20:14.520Z",
      "content": "<p>Thanks for your kind sharing…!</p>",
      "rawMarkdown": "Thanks for your kind sharing...!"
    }
  ],
  "comments": [
    {
      "id": 3221016,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2025-06-10T09:25:55.813000",
      "content": "<p>Fused optimizer</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3221224,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2025-06-10T15:33:36.080000",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> would you mind to share what speed up Fused optimiser gave; I remember it not being huge but might try again. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3221242,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2025-06-10T16:06:40.017000",
              "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> Just going from memory, Compile+Fused gave me 30-35%, I think fused alone was 10%+</p>\n<p><code>optimizer = optim.AdamW(model.parameters(), lr=CFG.lr, weight_decay=CFG.wd, fused=True)</code></p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3221247,
              "author_name": "Darragh",
              "author_url": "",
              "post_date": "2025-06-10T16:15:13.290000",
              "content": "<p>Thanks, will try fused optimiser again.</p>\n<p>Maybe I had it wrong how much torch.compile sped up; or else on Hopper the difference is not so big as I got on a100. <a href=\"https://lightning.ai/blog/training-compiled-pytorch-2.0-with-pytorch-lightning/\" target=\"_blank\">This</a> shows a100 speed ups. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3225677,
              "author_name": "Gowri Shankar Penugonda",
              "author_url": "",
              "post_date": "2025-06-16T17:55:53.353000",
              "content": "<p>The Compile+Fused gave really good boost.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3224767,
      "author_name": "Jiaming Tao",
      "author_url": "",
      "post_date": "2025-06-15T12:55:11.847000",
      "content": "<p>Thanks for sharing, it's friendly to newcomers!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3225546,
          "author_name": "Chaiho Wang",
          "author_url": "",
          "post_date": "2025-06-16T14:38:19.823000",
          "content": "<p>yes, so good for me too. thanks for author.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3223457,
      "author_name": "Nakapen",
      "author_url": "",
      "post_date": "2025-06-13T10:26:12.620000",
      "content": "<p>Thanks for your kind shareing! Would you tell me the reason why not use<br>\n<code>torch.compile(mode = \"reduce overhead\")</code>?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3225685,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2025-06-16T18:01:20.110000",
          "content": "<p>I played around with a few settings and got some issues with some settings and didn’t see a huge difference in the others which worked, so settled on this. I see most people use reduce overhead, I might try again. Thanks 🙏 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3221494,
      "author_name": "I2nfinit3y",
      "author_url": "",
      "post_date": "2025-06-11T04:47:40.400000",
      "content": "<p>Thanks for your kind sharing! But when I combine torch.compile and fused optimizer, codes always raise a error like \"RuntimeError: params, grads, exp_avgs, and exp_avg_sqs must have same dtype, device, and layout\" on RTX A40. Have you met this?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3221606,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2025-06-11T08:37:08.770000",
          "content": "<p>Sorry, I have not seen this. I'd recommend use your choice of free [deepseek.com, chatgpt.com, etc] or paid [clause-sonnet,openai-o3, etc] LLM apis. This is my workflow for the competition 😀 ps. I test and my gains on a100 for fused optimiser were not much. </p>\n<pre><code>\n\nAnd write your question\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3227150,
          "author_name": "Waylon Wu",
          "author_url": "",
          "post_date": "2025-06-18T15:26:12.163000",
          "content": "<p><a href=\"https://www.kaggle.com/i2nfinit3y\" target=\"_blank\">@i2nfinit3y</a> I got this error as well. Did you fix it?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3226959,
      "author_name": "Khushi Yadav",
      "author_url": "",
      "post_date": "2025-06-18T10:00:27.457000",
      "content": "<p>Thanks for sharing, great opportunity.🌟</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3224857,
      "author_name": "Marius Heuser",
      "author_url": "",
      "post_date": "2025-06-15T15:19:32.253000",
      "content": "<p>I get this message, since I started using the caformer code:</p>\n<pre><code>/usr//lib/python3/dist-packages/torch/autograd/graph.py:: UserWarning: Grad strides do  match bucket view strides. This may indicate grad was  created according   gradient layout contract,    param's strides changed  DDP was constructed.  This   an ,  may impair performance.\n\ns      grad.sizes() = [, , , ], strides() = [, , , ]\n\ns      bucket_view.sizes() = [, , , ], strides() = [, , , ] (Triggered internally  /pytorch/torch/csrc/distributed/c10d/reducer.cpp:)\n\ns         Variable._execution_engine.run_backward(  \n</code></pre>\n<p>Does anyone know how to fix it or what's causing it?<br>\nI tried checking for non-contiguous tensors, but didnt find any.<br>\nI feel like my training slows down alot</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3225010,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-06-15T19:47:04.440000",
          "content": "<p>I got the same. I did not dig into it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3225182,
          "author_name": "Bertnardo Mario Uskono",
          "author_url": "",
          "post_date": "2025-06-16T04:34:30.133000",
          "content": "<p>Hey Marius,</p>\n<p>That warning seems to be related to how <code>DistributedDataParallel (DDP)</code> handles gradient bucketing — especially when the tensor strides mismatch between the parameter and the bucket view. Since you're using Caformer, it could be that one of the layers or reshaping operations is altering tensor memory layout in a way that DDP doesn't expect.</p>\n<p>A few ideas to try:</p>\n<ul>\n<li>Check if any tensor operations like <code>.permute()</code>, <code>.transpose()</code>, or <code>.view()</code> are used without following them with <code>.contiguous()</code>. DDP can be sensitive to non-contiguous gradients.</li>\n<li>You can try wrapping your model parameters or outputs with <code>.contiguous()</code> where it makes sense, especially just before loss computation or backpropagation.</li>\n<li>Make sure the model is initialized <strong>before</strong> wrapping with <code>DDP</code>, and that no parameter reshaping is done afterward.</li>\n<li>Also worth trying <code>model = model.to(memory_format=torch.channels_last)</code> early in your pipeline — Caformer might be using channels_last format internally, which can create stride inconsistencies.</li>\n</ul>\n<p>This warning doesn't usually cause incorrect results, but as you noted, it <strong>can</strong> hurt performance if grad bucketing isn’t efficient. If you’re seeing slower training, it might be worth running a profiler (<code>torch.profiler</code>) to verify where the slowdown is coming from.</p>\n<p>Let us know if you find a clean workaround — I'm sure others using Caformer in this comp are hitting similar issues! 🔍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3223478,
      "author_name": "Abhishek Dave",
      "author_url": "",
      "post_date": "2025-06-13T11:04:37.523000",
      "content": "<p>Use torch.compile, mixed precision (torch.cuda.amp), and gradient checkpointing to boost speed and reduce memory. Optimize DataLoader with more workers, pin_memory, and persistent_workers for faster data throughput.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3224848,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-06-15T15:04:07.663000",
          "content": "<p>Gradient checkpointing reduces memory but it doesn't speed up training.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3225053,
          "author_name": "c-number",
          "author_url": "",
          "post_date": "2025-06-15T23:17:57.933000",
          "content": "<p>It could if your vram is really low and your original batch size is small.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3223318,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2025-06-13T06:47:37.250000",
      "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> thank you, except FP8 training - training speed improved 18%-20% with fuse=True on rtx 5090</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3221209,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2025-06-10T15:14:18.280000",
      "content": "<p>How do you avoid NaN issues with fp16 training ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3221217,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2025-06-10T15:26:51.867000",
          "content": "<p>I get it sometimes, but not too often; feed in fp32 inputs &amp; labels and use <code>from torch.amp import autocast</code> like below.</p>\n<pre><code>                with autocast(=, =torch.bfloat16):\n                    output = model(batch)\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3226859,
          "author_name": "QIXUAN HE",
          "author_url": "",
          "post_date": "2025-06-18T07:09:25.310000",
          "content": "<p><code>with autocast(device_type='cuda', dtype=torch.float32):</code> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3221204,
      "author_name": "Victor",
      "author_url": "",
      "post_date": "2025-06-10T15:01:56.650000",
      "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> Hi, what are the pros and cons of loading all the .npy files with <code>mmap_mode='r'</code> in <code>__init__</code> versus lazy loading with <code>mmap_mode='r'</code> in <code>__getitem__</code>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3221222,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2025-06-10T15:32:37.730000",
          "content": "<p>I just tried in <code>__getitem__</code> .. .would need some research to see the difference. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3221189,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2025-06-10T14:30:51.920000",
      "content": "<p>Thanks, I get about 13% speedup with torch.compile().</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3221195,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2025-06-10T14:42:00.633000",
          "content": "<p>Same model 12hrs to 7hrs on a100 for <code>torch.compile()</code> like above - will double check I got that correct. Does that sound right for others ?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3221116,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2025-06-10T12:27:31.543000",
      "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> - How much training speed improved? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3221196,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2025-06-10T14:44:52.290000",
          "content": "<p>Depends on your baseline, amp / compile / subsampling … all were the big hitters for me - in the 1.5x to 2x range. Subsampling and downsizing (a type of subsampling) depends on your choices on how aggressively to do it. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3227462,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-19T03:02:18.767000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3225181,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-16T04:33:16.623000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3225683,
      "author_name": "Mandip Sapkota",
      "author_url": "",
      "post_date": "2025-06-16T18:00:08.483000",
      "content": "<p>Thanks for it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3222491,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-06-12T08:20:14.520000",
      "content": "<p>Thanks for your kind sharing…!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3221005": "Here is what I use to speed up training, a combination of ideas in the forum, some from Anthropic's Claude Opus based on the training scripts....\n\n- Memory map in dataloading `np.load(path, mmap_mode='r')` like in many public scripts. \n- Mixed precision training, [examples](https://docs.pytorch.org/docs/stable/notes/amp_examples.html)\n- Compile model in torch, `model = torch.compile(model, **dict(fullgraph=False, dynamic=False, mode=\"default\"))`. \n- Subsample easy families, eg., in [Bartley's CV table](https://www.kaggle.com/competitions/waveform-inversion/discussion/582785) , `FlatVel_A` scores 1.31 - subsample it with torch [WeightedRandomSampler](https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler). \n- Downsize along time axis, - a bit risky. \n- Increase prefetch `prefetch_factor` in DataLoader, not sure if this helps, but I use it.\n- `channels_last` format in torch (small difference, if any)\n- FP8 training, [linky](https://pytorch.org/blog/training-using-float8-fsdp2/), gain seems to be limited for smaller models. Only works for newer GPU's - 4090, 5090, H100, etc. \n\n\nI did not try the last one. Any other ideas to make the most of the last three weeks ?",
    "3221016": "Fused optimizer",
    "3224767": "Thanks for sharing, it's friendly to newcomers!",
    "3223457": "Thanks for your kind shareing! Would you tell me the reason why not use\n`torch.compile(mode = \"reduce overhead\")`?",
    "3221494": "Thanks for your kind sharing! But when I combine torch.compile and fused optimizer, codes always raise a error like \"RuntimeError: params, grads, exp_avgs, and exp_avg_sqs must have same dtype, device, and layout\" on RTX A40. Have you met this?",
    "3226959": "Thanks for sharing, great opportunity.🌟",
    "3224857": "I get this message, since I started using the caformer code:\n\n```\n/usr/local/lib/python3.11/dist-packages/torch/autograd/graph.py:823: UserWarning: Grad strides do not match bucket view strides. This may indicate grad was not created according to the gradient layout contract, or that the param's strides changed since DDP was constructed.  This is not an error, but may impair performance.\n\n485.7s\t55\tgrad.sizes() = [64, 128, 1, 1], strides() = [128, 1, 128, 128]\n\n485.7s\t56\tbucket_view.sizes() = [64, 128, 1, 1], strides() = [128, 1, 1, 1] (Triggered internally at /pytorch/torch/csrc/distributed/c10d/reducer.cpp:327.)\n\n485.7s\t57\t  return Variable._execution_engine.run_backward(  # Calls into the C++ engine to run the backward pass\n```\n\nDoes anyone know how to fix it or what's causing it?\nI tried checking for non-contiguous tensors, but didnt find any.\nI feel like my training slows down alot",
    "3223478": "Use torch.compile, mixed precision (torch.cuda.amp), and gradient checkpointing to boost speed and reduce memory. Optimize DataLoader with more workers, pin_memory, and persistent_workers for faster data throughput.",
    "3223318": "@darraghdog thank you, except FP8 training - training speed improved 18%-20% with fuse=True on rtx 5090",
    "3221209": "How do you avoid NaN issues with fp16 training ?",
    "3221204": "@darraghdog Hi, what are the pros and cons of loading all the .npy files with `mmap_mode='r'` in `__init__` versus lazy loading with `mmap_mode='r'` in `__getitem__`?",
    "3221189": "Thanks, I get about 13% speedup with torch.compile().",
    "3221116": "@darraghdog - How much training speed improved? ",
    "3227462": "",
    "3225181": "",
    "3225683": "Thanks for it.",
    "3222491": "Thanks for your kind sharing...!"
  }
}