{
  "id": 265738,
  "title": "Training with pre-loaded weights using pytorch-pfn-extras ?",
  "url": "/competitions/seti-breakthrough-listen/discussion/265738",
  "author_name": "Ace",
  "post_date": "2021-08-16T19:40:04.217000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p>A number of public notebooks, derived from excellent <a href=\"https://www.kaggle.com/ttahara/seti-e-t-resnet18d-baseline\" target=\"_blank\">ttahara</a> are using pytorch-pfn-extras. Due to GPU running time limitations the most max_epoch can be set up. Does anybody know how to re-start training from max_epoch+1, pre-loading the best model weights so far? I suppose that some of extension parametres: {type: snapshot, target: \"@/model\", filename: \"xx_{.epoch}.pth\"} should be changed.</p>",
  "messages": [
    {
      "id": 1476199,
      "postDate": "2021-08-17T03:15:41.470Z",
      "content": "<p>I had experimented with one of <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>'s kernels in a previous competition. Great kernel and work, really. But I quickly fell out of interest with pfn-extras. Trying to maneuver around framework only usurped valuable time from me. Sorry, I don't have an answer to your question, but it is almost for this very reason when I persist my models, i don't just save the model's state_dict, but also the optimizer's as well. Makes it painless to load the state and immediately pick off without having to do any hackery.</p>",
      "rawMarkdown": "I had experimented with one of @ttahara's kernels in a previous competition. Great kernel and work, really. But I quickly fell out of interest with pfn-extras. Trying to maneuver around framework only usurped valuable time from me. Sorry, I don't have an answer to your question, but it is almost for this very reason when I persist my models, i don't just save the model's state_dict, but also the optimizer's as well. Makes it painless to load the state and immediately pick off without having to do any hackery.",
      "votes": 3
    },
    {
      "id": 1477885,
      "postDate": "2021-08-17T17:26:08.523Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a>. <br>\nYour suggestion for straightforward model weights loading works:<br>\n          model.load_state_dict(torch.load(RESUME_WEIGHTS, map_location=device))<br>\nMy mistake was, that i was trying to modify args.snapshots to RESUME_WEIGHTS, and for the initial iteration only i did:<br>\n          manager.load_state_dict(torch.load(args.snapshot))</p>",
      "rawMarkdown": "Thank you @imeintanis. \nYour suggestion for straightforward model weights loading works:\n          model.load_state_dict(torch.load(RESUME_WEIGHTS, map_location=device))\nMy mistake was, that i was trying to modify args.snapshots to RESUME_WEIGHTS, and for the initial iteration only i did:\n          manager.load_state_dict(torch.load(args.snapshot))",
      "votes": 1
    },
    {
      "id": 1477175,
      "postDate": "2021-08-17T11:36:54.387Z",
      "content": "<p>bTW, </p>\n<blockquote>\n  <p>I suppose that some of extension parametres: {type: snapshot, target: \"@/model\", filename: \"xx_{.epoch}.pth\"} should be changed</p>\n</blockquote>\n<p>this ext has to do with saving the \"snapshot\" best checkpoint at every epoch where the monitored metric is improved. i.e. shouldn't conflict with pre-loading model weights</p>",
      "rawMarkdown": "bTW, \n\n> I suppose that some of extension parametres: {type: snapshot, target: \"@/model\", filename: \"xx_{.epoch}.pth\"} should be changed\n\nthis ext has to do with saving the \"snapshot\" best checkpoint at every epoch where the monitored metric is improved. i.e. shouldn't conflict with pre-loading model weights",
      "votes": 1
    },
    {
      "id": 1477168,
      "postDate": "2021-08-17T11:32:23.603Z",
      "content": "<p>I think if you modify a bit <code>train_one_fold()</code> should make it work, e.g.</p>\n<pre><code>   resume_weights = 'path/to/model'\n\n   model = cfg[\"/model\"]\n   if resume_weights is not None:\n        model.load_state_dict(torch.load(resume_weights, map_location=device))\n        print('---&gt; Resume from previous checkpoint')\n\n   model = nn.DataParallel(model, device_ids=[0])  # check GPUs\n   model.to(device)\n</code></pre>",
      "rawMarkdown": "I think if you modify a bit `train_one_fold()` should make it work, e.g.\n\n```\n   resume_weights = 'path/to/model'\n\n   model = cfg[\"/model\"]\n   if resume_weights is not None:\n        model.load_state_dict(torch.load(resume_weights, map_location=device))\n        print('---> Resume from previous checkpoint')\n\n   model = nn.DataParallel(model, device_ids=[0])  # check GPUs\n   model.to(device)\n```",
      "votes": 1
    },
    {
      "id": 1479542,
      "postDate": "2021-08-18T14:13:31.073Z",
      "content": "<p>Unfortunately the problem is still open, but there is no more time left for this contest :( . In suggested solution below, the model is loaded with resume_weights:<br>\n     model.load_state_dict(torch.load(resume_weights, map_location=device))<br>\nbut apparently the <br>\n   manager.load_state_dict(…) overrides the weights automatically at later steps.<br>\nSeveral training sessions, each one loading the previous model weights,  produce the same val/loss, val/metric statistics.</p>",
      "rawMarkdown": "Unfortunately the problem is still open, but there is no more time left for this contest :( . In suggested solution below, the model is loaded with resume_weights:\n     model.load_state_dict(torch.load(resume_weights, map_location=device))\nbut apparently the \n   manager.load_state_dict(...) overrides the weights automatically at later steps.\nSeveral training sessions, each one loading the previous model weights,  produce the same val/loss, val/metric statistics.",
      "replies": [
        {
          "id": 1480139,
          "postDate": "2021-08-18T20:21:42.550Z",
          "content": "<p>In my case is working fine I have to say - I used that for only few exp so I'm not very familiar either, <br>\nhowever, for e.g. I trained for 10 epochs (AUC 0.85x) - stage 1 and resume for other 10 (stage 2) starting from AUC ~0.82x to reach eventually 0.87x</p>\n<p>are you sure you haven't messed anything ? anyway since the comp is over I'd suggest to move to a pure Torch pipeline, or move to Lightning/Fast-ai if you like working in a \"higher-level\"</p>",
          "rawMarkdown": "In my case is working fine I have to say - I used that for only few exp so I'm not very familiar either, \nhowever, for e.g. I trained for 10 epochs (AUC 0.85x) - stage 1 and resume for other 10 (stage 2) starting from AUC ~0.82x to reach eventually 0.87x\n\nare you sure you haven't messed anything ? anyway since the comp is over I'd suggest to move to a pure Torch pipeline, or move to Lightning/Fast-ai if you like working in a \"higher-level\""
        }
      ]
    },
    {
      "id": 1475897,
      "postDate": "2021-08-16T22:04:12.363Z",
      "content": "<p>I don't know pure pytorch but this is straighforward with FastAI. The key parts you need to do is to save your checkpoint model for every n epoch.<br>\nwhen restarting there are two key points you need to pay attention:<br>\n1- Load the model from the last checkpoint<br>\n2- reinitialize the learning rate (and any other hyper parameter) scheduler you have, so you start with the same learning rate you ended</p>",
      "rawMarkdown": "I don't know pure pytorch but this is straighforward with FastAI. The key parts you need to do is to save your checkpoint model for every n epoch.\nwhen restarting there are two key points you need to pay attention:\n1- Load the model from the last checkpoint\n2- reinitialize the learning rate (and any other hyper parameter) scheduler you have, so you start with the same learning rate you ended",
      "replies": [
        {
          "id": 1476296,
          "postDate": "2021-08-17T04:30:32.193Z",
          "content": "<p>The problem with these rapid development tools is that they rely on their own framework and parameters and often are not well documented. In pfn-extras case, the attempt to load the model from the last checkpoint: <br>\nstate = torch.load(LOADED_MODEL, map_location=device)<br>\nconflicts with:<br>\n manager.load_state_dict(state) =&gt; self._start_iteration = to_load['_start_iteration']</p>",
          "rawMarkdown": "The problem with these rapid development tools is that they rely on their own framework and parameters and often are not well documented. In pfn-extras case, the attempt to load the model from the last checkpoint: \nstate = torch.load(LOADED_MODEL, map_location=device)\nconflicts with:\n manager.load_state_dict(state) => self._start_iteration = to_load['_start_iteration']"
        }
      ]
    },
    {
      "id": 1475664,
      "postDate": "2021-08-16T19:40:04.217Z",
      "content": "<p>A number of public notebooks, derived from excellent <a href=\"https://www.kaggle.com/ttahara/seti-e-t-resnet18d-baseline\" target=\"_blank\">ttahara</a> are using pytorch-pfn-extras. Due to GPU running time limitations the most max_epoch can be set up. Does anybody know how to re-start training from max_epoch+1, pre-loading the best model weights so far? I suppose that some of extension parametres: {type: snapshot, target: \"@/model\", filename: \"xx_{.epoch}.pth\"} should be changed.</p>",
      "rawMarkdown": "A number of public notebooks, derived from excellent [ttahara](https://www.kaggle.com/ttahara/seti-e-t-resnet18d-baseline) are using pytorch-pfn-extras. Due to GPU running time limitations the most max_epoch can be set up. Does anybody know how to re-start training from max_epoch+1, pre-loading the best model weights so far? I suppose that some of extension parametres: {type: snapshot, target: \"@/model\", filename: \"xx_{.epoch}.pth\"} should be changed."
    }
  ],
  "comments": [
    {
      "id": 1476199,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2021-08-17T03:15:41.470000",
      "content": "<p>I had experimented with one of <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>'s kernels in a previous competition. Great kernel and work, really. But I quickly fell out of interest with pfn-extras. Trying to maneuver around framework only usurped valuable time from me. Sorry, I don't have an answer to your question, but it is almost for this very reason when I persist my models, i don't just save the model's state_dict, but also the optimizer's as well. Makes it painless to load the state and immediately pick off without having to do any hackery.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1477885,
      "author_name": "Ace",
      "author_url": "",
      "post_date": "2021-08-17T17:26:08.523000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a>. <br>\nYour suggestion for straightforward model weights loading works:<br>\n          model.load_state_dict(torch.load(RESUME_WEIGHTS, map_location=device))<br>\nMy mistake was, that i was trying to modify args.snapshots to RESUME_WEIGHTS, and for the initial iteration only i did:<br>\n          manager.load_state_dict(torch.load(args.snapshot))</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1477175,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-08-17T11:36:54.387000",
      "content": "<p>bTW, </p>\n<blockquote>\n  <p>I suppose that some of extension parametres: {type: snapshot, target: \"@/model\", filename: \"xx_{.epoch}.pth\"} should be changed</p>\n</blockquote>\n<p>this ext has to do with saving the \"snapshot\" best checkpoint at every epoch where the monitored metric is improved. i.e. shouldn't conflict with pre-loading model weights</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1477168,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-08-17T11:32:23.603000",
      "content": "<p>I think if you modify a bit <code>train_one_fold()</code> should make it work, e.g.</p>\n<pre><code>   resume_weights = 'path/to/model'\n\n   model = cfg[\"/model\"]\n   if resume_weights is not None:\n        model.load_state_dict(torch.load(resume_weights, map_location=device))\n        print('---&gt; Resume from previous checkpoint')\n\n   model = nn.DataParallel(model, device_ids=[0])  # check GPUs\n   model.to(device)\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1479542,
      "author_name": "Ace",
      "author_url": "",
      "post_date": "2021-08-18T14:13:31.073000",
      "content": "<p>Unfortunately the problem is still open, but there is no more time left for this contest :( . In suggested solution below, the model is loaded with resume_weights:<br>\n     model.load_state_dict(torch.load(resume_weights, map_location=device))<br>\nbut apparently the <br>\n   manager.load_state_dict(…) overrides the weights automatically at later steps.<br>\nSeveral training sessions, each one loading the previous model weights,  produce the same val/loss, val/metric statistics.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480139,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2021-08-18T20:21:42.550000",
          "content": "<p>In my case is working fine I have to say - I used that for only few exp so I'm not very familiar either, <br>\nhowever, for e.g. I trained for 10 epochs (AUC 0.85x) - stage 1 and resume for other 10 (stage 2) starting from AUC ~0.82x to reach eventually 0.87x</p>\n<p>are you sure you haven't messed anything ? anyway since the comp is over I'd suggest to move to a pure Torch pipeline, or move to Lightning/Fast-ai if you like working in a \"higher-level\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1475897,
      "author_name": "Adriano Passos",
      "author_url": "",
      "post_date": "2021-08-16T22:04:12.363000",
      "content": "<p>I don't know pure pytorch but this is straighforward with FastAI. The key parts you need to do is to save your checkpoint model for every n epoch.<br>\nwhen restarting there are two key points you need to pay attention:<br>\n1- Load the model from the last checkpoint<br>\n2- reinitialize the learning rate (and any other hyper parameter) scheduler you have, so you start with the same learning rate you ended</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1476296,
          "author_name": "Ace",
          "author_url": "",
          "post_date": "2021-08-17T04:30:32.193000",
          "content": "<p>The problem with these rapid development tools is that they rely on their own framework and parameters and often are not well documented. In pfn-extras case, the attempt to load the model from the last checkpoint: <br>\nstate = torch.load(LOADED_MODEL, map_location=device)<br>\nconflicts with:<br>\n manager.load_state_dict(state) =&gt; self._start_iteration = to_load['_start_iteration']</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1476199": "I had experimented with one of @ttahara's kernels in a previous competition. Great kernel and work, really. But I quickly fell out of interest with pfn-extras. Trying to maneuver around framework only usurped valuable time from me. Sorry, I don't have an answer to your question, but it is almost for this very reason when I persist my models, i don't just save the model's state_dict, but also the optimizer's as well. Makes it painless to load the state and immediately pick off without having to do any hackery.",
    "1477885": "Thank you @imeintanis. \nYour suggestion for straightforward model weights loading works:\n          model.load_state_dict(torch.load(RESUME_WEIGHTS, map_location=device))\nMy mistake was, that i was trying to modify args.snapshots to RESUME_WEIGHTS, and for the initial iteration only i did:\n          manager.load_state_dict(torch.load(args.snapshot))",
    "1477175": "bTW, \n\n> I suppose that some of extension parametres: {type: snapshot, target: \"@/model\", filename: \"xx_{.epoch}.pth\"} should be changed\n\nthis ext has to do with saving the \"snapshot\" best checkpoint at every epoch where the monitored metric is improved. i.e. shouldn't conflict with pre-loading model weights",
    "1477168": "I think if you modify a bit `train_one_fold()` should make it work, e.g.\n\n```\n   resume_weights = 'path/to/model'\n\n   model = cfg[\"/model\"]\n   if resume_weights is not None:\n        model.load_state_dict(torch.load(resume_weights, map_location=device))\n        print('---> Resume from previous checkpoint')\n\n   model = nn.DataParallel(model, device_ids=[0])  # check GPUs\n   model.to(device)\n```",
    "1479542": "Unfortunately the problem is still open, but there is no more time left for this contest :( . In suggested solution below, the model is loaded with resume_weights:\n     model.load_state_dict(torch.load(resume_weights, map_location=device))\nbut apparently the \n   manager.load_state_dict(...) overrides the weights automatically at later steps.\nSeveral training sessions, each one loading the previous model weights,  produce the same val/loss, val/metric statistics.",
    "1475897": "I don't know pure pytorch but this is straighforward with FastAI. The key parts you need to do is to save your checkpoint model for every n epoch.\nwhen restarting there are two key points you need to pay attention:\n1- Load the model from the last checkpoint\n2- reinitialize the learning rate (and any other hyper parameter) scheduler you have, so you start with the same learning rate you ended",
    "1475664": "A number of public notebooks, derived from excellent [ttahara](https://www.kaggle.com/ttahara/seti-e-t-resnet18d-baseline) are using pytorch-pfn-extras. Due to GPU running time limitations the most max_epoch can be set up. Does anybody know how to re-start training from max_epoch+1, pre-loading the best model weights so far? I suppose that some of extension parametres: {type: snapshot, target: \"@/model\", filename: \"xx_{.epoch}.pth\"} should be changed."
  }
}