{
  "id": 580980,
  "title": "Failed to load the model",
  "url": "/competitions/waveform-inversion/discussion/580980",
  "author_name": "",
  "post_date": "2025-05-27T15:49:25.395572Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>May I ask everyone why there was an error when I loaded the model like this?😭😭😭<br>\n`   cfg.resume_path = \"/kaggle/input/openfwi-preprocessed-72x72/models/unet2d_best1.pt\" <br>\n    model = Net(backbone=cfg.backbone)<br>\n    model = model.to(cfg.local_rank)</p>\n<pre><code>\n cfg.resume_path is  None  os.path.exists(cfg.resume_path):\n     cfg.local_rank == 0:\n        (f)\n    map_location = { % 0:  % cfg.local_rank}\n    state_dict = torch.load(cfg.resume_path, =map_location)\n    model.load_state_dict(state_dict)\n\nmodel = DistributedDataParallel(model, device_ids=[cfg.local_rank])\n\n\n cfg.ema:\n    ema_model = ModelEMA(model, =cfg.ema_decay, =cfg.local_rank)\n     cfg.resume_path is  None  os.path.exists(cfg.resume_path):\n        ema_model.module.load_state_dict(state_dict)\n:\n    ema_model = None\n\nmodel= DistributedDataParallel(\n    model, \n    device_ids=[cfg.local_rank], \n    )\n</code></pre>\n<p>`<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16297364%2Fc25933880c4b4cef2d74a9a14beb8537%2Fb80f1b58-268f-4b47-b80d-ead55bf0689e.png?generation=1748360871110119&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3210698",
      "postDate": "05/27/2025 15:49:25",
      "content": "<p>May I ask everyone why there was an error when I loaded the model like this?😭😭😭<br>\n`   cfg.resume_path = \"/kaggle/input/openfwi-preprocessed-72x72/models/unet2d_best1.pt\" <br>\n    model = Net(backbone=cfg.backbone)<br>\n    model = model.to(cfg.local_rank)</p>\n<pre><code>\n cfg.resume_path is  None  os.path.exists(cfg.resume_path):\n     cfg.local_rank == 0:\n        (f)\n    map_location = { % 0:  % cfg.local_rank}\n    state_dict = torch.load(cfg.resume_path, =map_location)\n    model.load_state_dict(state_dict)\n\nmodel = DistributedDataParallel(model, device_ids=[cfg.local_rank])\n\n\n cfg.ema:\n    ema_model = ModelEMA(model, =cfg.ema_decay, =cfg.local_rank)\n     cfg.resume_path is  None  os.path.exists(cfg.resume_path):\n        ema_model.module.load_state_dict(state_dict)\n:\n    ema_model = None\n\nmodel= DistributedDataParallel(\n    model, \n    device_ids=[cfg.local_rank], \n    )\n</code></pre>\n<p>`<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16297364%2Fc25933880c4b4cef2d74a9a14beb8537%2Fb80f1b58-268f-4b47-b80d-ead55bf0689e.png?generation=1748360871110119&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "May I ask everyone why there was an error when I loaded the model like this?😭😭😭\n`   cfg.resume_path = \"/kaggle/input/openfwi-preprocessed-72x72/models/unet2d_best1.pt\" \n    model = Net(backbone=cfg.backbone)\n    model = model.to(cfg.local_rank)\n    \n    # Resume training\n    if cfg.resume_path is not None and os.path.exists(cfg.resume_path):\n        if cfg.local_rank == 0:\n            print(f\"Resuming training from {cfg.resume_path}\")\n        map_location = {\"cuda:%d\" % 0: \"cuda:%d\" % cfg.local_rank}\n        state_dict = torch.load(cfg.resume_path, map_location=map_location)\n        model.load_state_dict(state_dict)\n    \n    model = DistributedDataParallel(model, device_ids=[cfg.local_rank])\n    \n    # Initialize EMA\n    if cfg.ema:\n        ema_model = ModelEMA(model, decay=cfg.ema_decay, device=cfg.local_rank)\n        if cfg.resume_path is not None and os.path.exists(cfg.resume_path):\n            ema_model.module.load_state_dict(state_dict)\n    else:\n        ema_model = None\n\n    model= DistributedDataParallel(\n        model, \n        device_ids=[cfg.local_rank], \n        )\n    \n`![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16297364%2Fc25933880c4b4cef2d74a9a14beb8537%2Fb80f1b58-268f-4b47-b80d-ead55bf0689e.png?generation=1748360871110119&alt=media)",
      "votes": null
    },
    {
      "id": "3210850",
      "postDate": "05/27/2025 18:36:21",
      "content": "<p>what's your cfg.backbone is it? set it to \"hgnetv2_b4.ssld_stage2_ft_in1k\"</p>",
      "rawMarkdown": "what's your cfg.backbone is it? set it to \"hgnetv2_b4.ssld_stage2_ft_in1k\"",
      "votes": null
    },
    {
      "id": "3211029",
      "postDate": "05/28/2025 02:45:45",
      "content": "<p>model = Net(backbone=cfg.backbone) \"hgnetv2_b2.ssld_stage2_ft_in1k\"Yesz, I set it up, It's still the same mistake</p>",
      "rawMarkdown": "model = Net(backbone=cfg.backbone) \"hgnetv2_b2.ssld_stage2_ft_in1k\"Yesz, I set it up, It's still the same mistake",
      "votes": null
    },
    {
      "id": "3211062",
      "postDate": "05/28/2025 03:31:34",
      "content": "<p>b4 not b2. They are different</p>\n<p>\"<br>\nPretrained Models<br>\nNext, we load in 3x pretrained models. These models were trained with with an effective batch_size of 512 (256 per GPU) and use the B4 variant of the HgnetV2 backbone.<br>\n\"</p>",
      "rawMarkdown": "b4 not b2. They are different\n\n\"\nPretrained Models\nNext, we load in 3x pretrained models. These models were trained with with an effective batch_size of 512 (256 per GPU) and use the B4 variant of the HgnetV2 backbone.\n\"",
      "votes": null
    },
    {
      "id": "3211063",
      "postDate": "05/28/2025 03:32:02",
      "content": "<p>in the cfg it write b2 variant but actually the model is trained on b4 variant</p>",
      "rawMarkdown": "in the cfg it write b2 variant but actually the model is trained on b4 variant",
      "votes": null
    },
    {
      "id": "3211498",
      "postDate": "05/28/2025 13:49:33",
      "content": "<p>cfg.backbone = \"hgnetv2_b4.ssld_stage2_ft_in1k\", now I have modified it and the incorrect parameters have changed，Missing key(s) in state_dict: \"module.backbone.stem.stem1.conv.weight\", \"module.backbone.stem.stem1.bn.weight\", \"module.backbone.stem.stem1.bn.bias\", \"module.backbone.stem.stem1.bn.running_mean\", \"module.backbone.stem.stem1.bn.running_var\", and so on</p>",
      "rawMarkdown": "cfg.backbone = \"hgnetv2_b4.ssld_stage2_ft_in1k\", now I have modified it and the incorrect parameters have changed，Missing key(s) in state_dict: \"module.backbone.stem.stem1.conv.weight\", \"module.backbone.stem.stem1.bn.weight\", \"module.backbone.stem.stem1.bn.bias\", \"module.backbone.stem.stem1.bn.running_mean\", \"module.backbone.stem.stem1.bn.running_var\", and so on",
      "votes": null
    },
    {
      "id": "3211698",
      "postDate": "05/28/2025 18:13:07",
      "content": "<p>When you wrap a model with DistributedDataParallel or DataParallel, PyTorch adds a \"module\" attribute that contains your original model.<br>\nWhen you call model.state_dict() on a wrapped model, all parameter keys are prefixed with \"module.\" because they're actually stored under the .module attribute of the wrapper.<br>\nLater, when you try to load these weights into a non-wrapped model, the keys don't match because your new model doesn't have the \"module.\" prefix in its parameter names.</p>\n<p>Just delete the module. prefix</p>",
      "rawMarkdown": "When you wrap a model with DistributedDataParallel or DataParallel, PyTorch adds a \"module\" attribute that contains your original model.\nWhen you call model.state_dict() on a wrapped model, all parameter keys are prefixed with \"module.\" because they're actually stored under the .module attribute of the wrapper.\nLater, when you try to load these weights into a non-wrapped model, the keys don't match because your new model doesn't have the \"module.\" prefix in its parameter names.\n\nJust delete the module. prefix",
      "votes": null
    },
    {
      "id": "3213589",
      "postDate": "05/30/2025 08:02:54",
      "content": "<p>thanks bro，now i use state_dict = {k.replace('module.', ''): v for k, v in state_dict.items()}</p>\n<pre><code>model = Net(backbone=cfg.backbone)\n    model = model.to(cfg.local_rank)\n\n    \n     cfg.resume_path     os.path.exists(cfg.resume_path):\n         cfg.local_rank == :\n            ()\n        map_location = { % :  % cfg.local_rank}\n        state_dict = torch.load(cfg.resume_path, map_location=map_location)\n        state_dict = {k.replace(, ): v  k, v  state_dict.items()} \n        model.load_state_dict(state_dict)\n\n    model = DistributedDataParallel(model, device_ids=[cfg.local_rank])\n\n    \n     cfg.ema:\n        ema_model = ModelEMA(model, decay=cfg.ema_decay, device=cfg.local_rank)\n         cfg.resume_path     os.path.exists(cfg.resume_path):\n            ema_model.module.load_state_dict(state_dict)\n    :\n        ema_model = \n\n    model= DistributedDataParallel(\n        model, \n        device_ids=[cfg.local_rank], \n        )\n</code></pre>\n<p>It seems to be of no use. The error remains the same. Missing key(s) in state_dict: \"module.backbone.stem.stem1.conv.weight\", \"module.backbone.stem.stem1.bn.weight\", \"module.backbone.stem.stem1.bn.bias\",  \"module.backbone.stem.stem1.bn.running_mean\", \"module.backbone.stem.stem1.bn.running_var\", and so on。How did you continue your training?</p>",
      "rawMarkdown": "thanks bro，now i use state_dict = {k.replace('module.', ''): v for k, v in state_dict.items()}\n\n```python\nmodel = Net(backbone=cfg.backbone)\n    model = model.to(cfg.local_rank)\n    \n    # Resume training\n    if cfg.resume_path is not None and os.path.exists(cfg.resume_path):\n        if cfg.local_rank == 0:\n            print(f\"Resuming training from {cfg.resume_path}\")\n        map_location = {\"cuda:%d\" % 0: \"cuda:%d\" % cfg.local_rank}\n        state_dict = torch.load(cfg.resume_path, map_location=map_location)\n        state_dict = {k.replace('module.', ''): v for k, v in state_dict.items()} \n        model.load_state_dict(state_dict)\n    \n    model = DistributedDataParallel(model, device_ids=[cfg.local_rank])\n    \n    # Initialize EMA\n    if cfg.ema:\n        ema_model = ModelEMA(model, decay=cfg.ema_decay, device=cfg.local_rank)\n        if cfg.resume_path is not None and os.path.exists(cfg.resume_path):\n            ema_model.module.load_state_dict(state_dict)\n    else:\n        ema_model = None\n\n    model= DistributedDataParallel(\n        model, \n        device_ids=[cfg.local_rank], \n        )\n```\nIt seems to be of no use. The error remains the same. Missing key(s) in state_dict: \"module.backbone.stem.stem1.conv.weight\", \"module.backbone.stem.stem1.bn.weight\", \"module.backbone.stem.stem1.bn.bias\",  \"module.backbone.stem.stem1.bn.running_mean\", \"module.backbone.stem.stem1.bn.running_var\", and so on。How did you continue your training?",
      "votes": null
    },
    {
      "id": "3213701",
      "postDate": "05/30/2025 10:48:05",
      "content": "<p>I still have a question. If the author trains for 150 epochs, will it still be three early stops or longer</p>",
      "rawMarkdown": "I still have a question. If the author trains for 150 epochs, will it still be three early stops or longer",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3210850,
      "author_name": "dan5280",
      "author_url": "",
      "post_date": "05/27/2025 18:36:21",
      "content": "<p>what's your cfg.backbone is it? set it to \"hgnetv2_b4.ssld_stage2_ft_in1k\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 3211029,
          "author_name": "waterjoe",
          "author_url": "",
          "post_date": "05/28/2025 02:45:45",
          "content": "<p>model = Net(backbone=cfg.backbone) \"hgnetv2_b2.ssld_stage2_ft_in1k\"Yesz, I set it up, It's still the same mistake</p>",
          "votes": null,
          "replies": [
            {
              "id": 3211062,
              "author_name": "dan5280",
              "author_url": "",
              "post_date": "05/28/2025 03:31:34",
              "content": "<p>b4 not b2. They are different</p>\n<p>\"<br>\nPretrained Models<br>\nNext, we load in 3x pretrained models. These models were trained with with an effective batch_size of 512 (256 per GPU) and use the B4 variant of the HgnetV2 backbone.<br>\n\"</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3211063,
              "author_name": "dan5280",
              "author_url": "",
              "post_date": "05/28/2025 03:32:02",
              "content": "<p>in the cfg it write b2 variant but actually the model is trained on b4 variant</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3211498,
                  "author_name": "waterjoe",
                  "author_url": "",
                  "post_date": "05/28/2025 13:49:33",
                  "content": "<p>cfg.backbone = \"hgnetv2_b4.ssld_stage2_ft_in1k\", now I have modified it and the incorrect parameters have changed，Missing key(s) in state_dict: \"module.backbone.stem.stem1.conv.weight\", \"module.backbone.stem.stem1.bn.weight\", \"module.backbone.stem.stem1.bn.bias\", \"module.backbone.stem.stem1.bn.running_mean\", \"module.backbone.stem.stem1.bn.running_var\", and so on</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3211698,
                      "author_name": "dan5280",
                      "author_url": "",
                      "post_date": "05/28/2025 18:13:07",
                      "content": "<p>When you wrap a model with DistributedDataParallel or DataParallel, PyTorch adds a \"module\" attribute that contains your original model.<br>\nWhen you call model.state_dict() on a wrapped model, all parameter keys are prefixed with \"module.\" because they're actually stored under the .module attribute of the wrapper.<br>\nLater, when you try to load these weights into a non-wrapped model, the keys don't match because your new model doesn't have the \"module.\" prefix in its parameter names.</p>\n<p>Just delete the module. prefix</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3213589,
                          "author_name": "waterjoe",
                          "author_url": "",
                          "post_date": "05/30/2025 08:02:54",
                          "content": "<p>thanks bro，now i use state_dict = {k.replace('module.', ''): v for k, v in state_dict.items()}</p>\n<pre><code>model = Net(backbone=cfg.backbone)\n    model = model.to(cfg.local_rank)\n\n    \n     cfg.resume_path     os.path.exists(cfg.resume_path):\n         cfg.local_rank == :\n            ()\n        map_location = { % :  % cfg.local_rank}\n        state_dict = torch.load(cfg.resume_path, map_location=map_location)\n        state_dict = {k.replace(, ): v  k, v  state_dict.items()} \n        model.load_state_dict(state_dict)\n\n    model = DistributedDataParallel(model, device_ids=[cfg.local_rank])\n\n    \n     cfg.ema:\n        ema_model = ModelEMA(model, decay=cfg.ema_decay, device=cfg.local_rank)\n         cfg.resume_path     os.path.exists(cfg.resume_path):\n            ema_model.module.load_state_dict(state_dict)\n    :\n        ema_model = \n\n    model= DistributedDataParallel(\n        model, \n        device_ids=[cfg.local_rank], \n        )\n</code></pre>\n<p>It seems to be of no use. The error remains the same. Missing key(s) in state_dict: \"module.backbone.stem.stem1.conv.weight\", \"module.backbone.stem.stem1.bn.weight\", \"module.backbone.stem.stem1.bn.bias\",  \"module.backbone.stem.stem1.bn.running_mean\", \"module.backbone.stem.stem1.bn.running_var\", and so on。How did you continue your training?</p>",
                          "votes": null,
                          "replies": []
                        },
                        {
                          "id": 3213701,
                          "author_name": "waterjoe",
                          "author_url": "",
                          "post_date": "05/30/2025 10:48:05",
                          "content": "<p>I still have a question. If the author trains for 150 epochs, will it still be three early stops or longer</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3210698": "May I ask everyone why there was an error when I loaded the model like this?😭😭😭\n`   cfg.resume_path = \"/kaggle/input/openfwi-preprocessed-72x72/models/unet2d_best1.pt\" \n    model = Net(backbone=cfg.backbone)\n    model = model.to(cfg.local_rank)\n    \n    # Resume training\n    if cfg.resume_path is not None and os.path.exists(cfg.resume_path):\n        if cfg.local_rank == 0:\n            print(f\"Resuming training from {cfg.resume_path}\")\n        map_location = {\"cuda:%d\" % 0: \"cuda:%d\" % cfg.local_rank}\n        state_dict = torch.load(cfg.resume_path, map_location=map_location)\n        model.load_state_dict(state_dict)\n    \n    model = DistributedDataParallel(model, device_ids=[cfg.local_rank])\n    \n    # Initialize EMA\n    if cfg.ema:\n        ema_model = ModelEMA(model, decay=cfg.ema_decay, device=cfg.local_rank)\n        if cfg.resume_path is not None and os.path.exists(cfg.resume_path):\n            ema_model.module.load_state_dict(state_dict)\n    else:\n        ema_model = None\n\n    model= DistributedDataParallel(\n        model, \n        device_ids=[cfg.local_rank], \n        )\n    \n`![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16297364%2Fc25933880c4b4cef2d74a9a14beb8537%2Fb80f1b58-268f-4b47-b80d-ead55bf0689e.png?generation=1748360871110119&alt=media)",
    "3210850": "what's your cfg.backbone is it? set it to \"hgnetv2_b4.ssld_stage2_ft_in1k\"",
    "3211029": "model = Net(backbone=cfg.backbone) \"hgnetv2_b2.ssld_stage2_ft_in1k\"Yesz, I set it up, It's still the same mistake",
    "3211062": "b4 not b2. They are different\n\n\"\nPretrained Models\nNext, we load in 3x pretrained models. These models were trained with with an effective batch_size of 512 (256 per GPU) and use the B4 variant of the HgnetV2 backbone.\n\"",
    "3211063": "in the cfg it write b2 variant but actually the model is trained on b4 variant",
    "3211498": "cfg.backbone = \"hgnetv2_b4.ssld_stage2_ft_in1k\", now I have modified it and the incorrect parameters have changed，Missing key(s) in state_dict: \"module.backbone.stem.stem1.conv.weight\", \"module.backbone.stem.stem1.bn.weight\", \"module.backbone.stem.stem1.bn.bias\", \"module.backbone.stem.stem1.bn.running_mean\", \"module.backbone.stem.stem1.bn.running_var\", and so on",
    "3211698": "When you wrap a model with DistributedDataParallel or DataParallel, PyTorch adds a \"module\" attribute that contains your original model.\nWhen you call model.state_dict() on a wrapped model, all parameter keys are prefixed with \"module.\" because they're actually stored under the .module attribute of the wrapper.\nLater, when you try to load these weights into a non-wrapped model, the keys don't match because your new model doesn't have the \"module.\" prefix in its parameter names.\n\nJust delete the module. prefix",
    "3213589": "thanks bro，now i use state_dict = {k.replace('module.', ''): v for k, v in state_dict.items()}\n\n```python\nmodel = Net(backbone=cfg.backbone)\n    model = model.to(cfg.local_rank)\n    \n    # Resume training\n    if cfg.resume_path is not None and os.path.exists(cfg.resume_path):\n        if cfg.local_rank == 0:\n            print(f\"Resuming training from {cfg.resume_path}\")\n        map_location = {\"cuda:%d\" % 0: \"cuda:%d\" % cfg.local_rank}\n        state_dict = torch.load(cfg.resume_path, map_location=map_location)\n        state_dict = {k.replace('module.', ''): v for k, v in state_dict.items()} \n        model.load_state_dict(state_dict)\n    \n    model = DistributedDataParallel(model, device_ids=[cfg.local_rank])\n    \n    # Initialize EMA\n    if cfg.ema:\n        ema_model = ModelEMA(model, decay=cfg.ema_decay, device=cfg.local_rank)\n        if cfg.resume_path is not None and os.path.exists(cfg.resume_path):\n            ema_model.module.load_state_dict(state_dict)\n    else:\n        ema_model = None\n\n    model= DistributedDataParallel(\n        model, \n        device_ids=[cfg.local_rank], \n        )\n```\nIt seems to be of no use. The error remains the same. Missing key(s) in state_dict: \"module.backbone.stem.stem1.conv.weight\", \"module.backbone.stem.stem1.bn.weight\", \"module.backbone.stem.stem1.bn.bias\",  \"module.backbone.stem.stem1.bn.running_mean\", \"module.backbone.stem.stem1.bn.running_var\", and so on。How did you continue your training?",
    "3213701": "I still have a question. If the author trains for 150 epochs, will it still be three early stops or longer"
  },
  "source": "meta"
}