{
  "id": 165669,
  "title": "TPU - how to save progress?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165669",
  "author_name": "",
  "post_date": "2020-07-10T14:29:21.169984100Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am now trying to implement mutlicore support for TPU training and I found following problem: how to save progress to continue training later?</p>\n\n<p>in following kernel by @abhishek only model state is saved:\n<a href=\"https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel\">https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel</a>\nwith code:\n<code>xm.save(model.state_dict(), f\"model_{fold}.bin\")</code></p>\n\n<p>however, in my usual PyTorch code I do following:</p>\n\n<pre><code>torch.save(model.state_dict(), path + \"model_{}.pth\".format(i))\ntorch.save(optimizer.state_dict(), path + \"optimizer_{}.pth\".format(i))\n</code></pre>\n\n<p>The problem is - which optimizer should I save? In multicore environment there is one model but many optimizers, right? Should I save then reload all 8? </p>\n\n<p>Anyone was trying to continue TPU training or always start from beginning?</p>",
  "messages": [
    {
      "id": "923104",
      "postDate": "07/10/2020 14:29:21",
      "content": "<p>I am now trying to implement mutlicore support for TPU training and I found following problem: how to save progress to continue training later?</p>\n\n<p>in following kernel by @abhishek only model state is saved:\n<a href=\"https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel\">https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel</a>\nwith code:\n<code>xm.save(model.state_dict(), f\"model_{fold}.bin\")</code></p>\n\n<p>however, in my usual PyTorch code I do following:</p>\n\n<pre><code>torch.save(model.state_dict(), path + \"model_{}.pth\".format(i))\ntorch.save(optimizer.state_dict(), path + \"optimizer_{}.pth\".format(i))\n</code></pre>\n\n<p>The problem is - which optimizer should I save? In multicore environment there is one model but many optimizers, right? Should I save then reload all 8? </p>\n\n<p>Anyone was trying to continue TPU training or always start from beginning?</p>",
      "rawMarkdown": "I am now trying to implement mutlicore support for TPU training and I found following problem: how to save progress to continue training later?\n\nin following kernel by @abhishek only model state is saved:\nhttps://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel\nwith code:\n`xm.save(model.state_dict(), f\"model_{fold}.bin\")`\n\nhowever, in my usual PyTorch code I do following:\n\n    torch.save(model.state_dict(), path + \"model_{}.pth\".format(i))\n    torch.save(optimizer.state_dict(), path + \"optimizer_{}.pth\".format(i))\n\nThe problem is - which optimizer should I save? In multicore environment there is one model but many optimizers, right? Should I save then reload all 8? \n\nAnyone was trying to continue TPU training or always start from beginning?",
      "votes": null
    },
    {
      "id": "923211",
      "postDate": "07/10/2020 16:08:06",
      "content": "<p>Hey <a href=\"/jacekpoplawski\">@jacekpoplawski</a> \nYou have 8 of everything, 8 copies of the model, 8 optimisers, 8 datasets, etc\nHowever, models and optimisers are synchronised on every batch, so they are always the same\nTherefore, you only need to save once :)\nHere is a list the algorithms used in multi-GPU training but its about the same in XLA/TF: <a href=\"https://pytorch.org/docs/stable/distributed.html\">https://pytorch.org/docs/stable/distributed.html</a></p>",
      "rawMarkdown": "Hey @jacekpoplawski \nYou have 8 of everything, 8 copies of the model, 8 optimisers, 8 datasets, etc\nHowever, models and optimisers are synchronised on every batch, so they are always the same\nTherefore, you only need to save once :)\nHere is a list the algorithms used in multi-GPU training but its about the same in XLA/TF: https://pytorch.org/docs/stable/distributed.html",
      "votes": null
    },
    {
      "id": "923384",
      "postDate": "07/10/2020 19:34:26",
      "content": "<p>\"However, models and optimisers are synchronised on every batch\" - that's a news for me, are you sure?\nEven if linked doc there is \"Each process maintains its own optimizer and performs a complete optimization step with each iteration\"</p>",
      "rawMarkdown": "\"However, models and optimisers are synchronised on every batch\" - that's a news for me, are you sure?\nEven if linked doc there is \"Each process maintains its own optimizer and performs a complete optimization step with each iteration\"",
      "votes": null
    },
    {
      "id": "923417",
      "postDate": "07/10/2020 20:22:04",
      "content": "<p>Yes <a href=\"/jacekpoplawski\">@jacekpoplawski</a> that's the whole point of the all reduce algo and sync training</p>\n\n<p>if you look at XLA optimizer_step <a href=\"https://pytorch.org/xla/release/1.5/_modules/torch_xla/core/xla_model.html#optimizer_step\">https://pytorch.org/xla/release/1.5/_modules/torch_xla/core/xla_model.html#optimizer_step</a> it call <code>reduce_gradients</code> already before updating the model.\nIf you want to check, you can predict (e.g. the submission) with all 8 cores, and all 8 outputs will be the same.</p>\n\n<p>The same is valid for TF: <a href=\"https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy\">https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy</a></p>",
      "rawMarkdown": "Yes @jacekpoplawski that's the whole point of the all reduce algo and sync training\n\nif you look at XLA optimizer_step https://pytorch.org/xla/release/1.5/_modules/torch_xla/core/xla_model.html#optimizer_step it call `reduce_gradients` already before updating the model.\nIf you want to check, you can predict (e.g. the submission) with all 8 cores, and all 8 outputs will be the same.\n\nThe same is valid for TF: https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy",
      "votes": null
    },
    {
      "id": "923430",
      "postDate": "07/10/2020 20:46:25",
      "content": "<p>\"If you want to check, you can predict (e.g. the submission) with all 8 cores, and all 8 outputs will be the same.\"</p>\n\n<p><a href=\"/hmendonca\">@hmendonca</a> correct me if I am wrong, but that's only means that model is shared, not the optimizer</p>\n\n<p>also when I look at code:\n<code>\n  reduce_gradients(optimizer)\n  loss = optimizer.step(**optimizer_args)\n</code>\nit means the loss is calculated from multiple optimizers and it is the same, but optimizers stay different</p>\n\n<p>However I am not sure what \"reduce_gradients\" does:</p>\n\n<p><code>\ngradients = _fetch_gradients(optimizer)\nall_reduce('sum', gradients, scale=1.0 / count)\n</code></p>",
      "rawMarkdown": "\"If you want to check, you can predict (e.g. the submission) with all 8 cores, and all 8 outputs will be the same.\"\n\n@hmendonca correct me if I am wrong, but that's only means that model is shared, not the optimizer\n\nalso when I look at code:\n```\n  reduce_gradients(optimizer)\n  loss = optimizer.step(**optimizer_args)\n```\nit means the loss is calculated from multiple optimizers and it is the same, but optimizers stay different\n\nHowever I am not sure what \"reduce_gradients\" does:\n\n```\ngradients = _fetch_gradients(optimizer)\nall_reduce('sum', gradients, scale=1.0 / count)\n```",
      "votes": null
    },
    {
      "id": "923455",
      "postDate": "07/10/2020 21:41:44",
      "content": "<p>reduce should sum/average all the gradients across the 8 optimizers\nas the state of the opt only depends on the gradients, if the gradients are the same so are the optimizers</p>",
      "rawMarkdown": "reduce should sum/average all the gradients across the 8 optimizers\nas the state of the opt only depends on the gradients, if the gradients are the same so are the optimizers",
      "votes": null
    },
    {
      "id": "923468",
      "postDate": "07/10/2020 22:12:40",
      "content": "<p>I see - thanks for clarification! :)</p>",
      "rawMarkdown": "I see - thanks for clarification! :)",
      "votes": null
    },
    {
      "id": "924857",
      "postDate": "07/11/2020 16:58:50",
      "content": "<p>Do you try to use ModelCheckpoint. It save the best model</p>",
      "rawMarkdown": "Do you try to use ModelCheckpoint. It save the best model",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 923211,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "07/10/2020 16:08:06",
      "content": "<p>Hey <a href=\"/jacekpoplawski\">@jacekpoplawski</a> \nYou have 8 of everything, 8 copies of the model, 8 optimisers, 8 datasets, etc\nHowever, models and optimisers are synchronised on every batch, so they are always the same\nTherefore, you only need to save once :)\nHere is a list the algorithms used in multi-GPU training but its about the same in XLA/TF: <a href=\"https://pytorch.org/docs/stable/distributed.html\">https://pytorch.org/docs/stable/distributed.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 923384,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/10/2020 19:34:26",
          "content": "<p>\"However, models and optimisers are synchronised on every batch\" - that's a news for me, are you sure?\nEven if linked doc there is \"Each process maintains its own optimizer and performs a complete optimization step with each iteration\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 923417,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/10/2020 20:22:04",
          "content": "<p>Yes <a href=\"/jacekpoplawski\">@jacekpoplawski</a> that's the whole point of the all reduce algo and sync training</p>\n\n<p>if you look at XLA optimizer_step <a href=\"https://pytorch.org/xla/release/1.5/_modules/torch_xla/core/xla_model.html#optimizer_step\">https://pytorch.org/xla/release/1.5/_modules/torch_xla/core/xla_model.html#optimizer_step</a> it call <code>reduce_gradients</code> already before updating the model.\nIf you want to check, you can predict (e.g. the submission) with all 8 cores, and all 8 outputs will be the same.</p>\n\n<p>The same is valid for TF: <a href=\"https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy\">https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 923430,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/10/2020 20:46:25",
          "content": "<p>\"If you want to check, you can predict (e.g. the submission) with all 8 cores, and all 8 outputs will be the same.\"</p>\n\n<p><a href=\"/hmendonca\">@hmendonca</a> correct me if I am wrong, but that's only means that model is shared, not the optimizer</p>\n\n<p>also when I look at code:\n<code>\n  reduce_gradients(optimizer)\n  loss = optimizer.step(**optimizer_args)\n</code>\nit means the loss is calculated from multiple optimizers and it is the same, but optimizers stay different</p>\n\n<p>However I am not sure what \"reduce_gradients\" does:</p>\n\n<p><code>\ngradients = _fetch_gradients(optimizer)\nall_reduce('sum', gradients, scale=1.0 / count)\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 923455,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/10/2020 21:41:44",
          "content": "<p>reduce should sum/average all the gradients across the 8 optimizers\nas the state of the opt only depends on the gradients, if the gradients are the same so are the optimizers</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 923468,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/10/2020 22:12:40",
          "content": "<p>I see - thanks for clarification! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 924857,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "07/11/2020 16:58:50",
      "content": "<p>Do you try to use ModelCheckpoint. It save the best model</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "923104": "I am now trying to implement mutlicore support for TPU training and I found following problem: how to save progress to continue training later?\n\nin following kernel by @abhishek only model state is saved:\nhttps://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel\nwith code:\n`xm.save(model.state_dict(), f\"model_{fold}.bin\")`\n\nhowever, in my usual PyTorch code I do following:\n\n    torch.save(model.state_dict(), path + \"model_{}.pth\".format(i))\n    torch.save(optimizer.state_dict(), path + \"optimizer_{}.pth\".format(i))\n\nThe problem is - which optimizer should I save? In multicore environment there is one model but many optimizers, right? Should I save then reload all 8? \n\nAnyone was trying to continue TPU training or always start from beginning?",
    "923211": "Hey @jacekpoplawski \nYou have 8 of everything, 8 copies of the model, 8 optimisers, 8 datasets, etc\nHowever, models and optimisers are synchronised on every batch, so they are always the same\nTherefore, you only need to save once :)\nHere is a list the algorithms used in multi-GPU training but its about the same in XLA/TF: https://pytorch.org/docs/stable/distributed.html",
    "923384": "\"However, models and optimisers are synchronised on every batch\" - that's a news for me, are you sure?\nEven if linked doc there is \"Each process maintains its own optimizer and performs a complete optimization step with each iteration\"",
    "923417": "Yes @jacekpoplawski that's the whole point of the all reduce algo and sync training\n\nif you look at XLA optimizer_step https://pytorch.org/xla/release/1.5/_modules/torch_xla/core/xla_model.html#optimizer_step it call `reduce_gradients` already before updating the model.\nIf you want to check, you can predict (e.g. the submission) with all 8 cores, and all 8 outputs will be the same.\n\nThe same is valid for TF: https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy",
    "923430": "\"If you want to check, you can predict (e.g. the submission) with all 8 cores, and all 8 outputs will be the same.\"\n\n@hmendonca correct me if I am wrong, but that's only means that model is shared, not the optimizer\n\nalso when I look at code:\n```\n  reduce_gradients(optimizer)\n  loss = optimizer.step(**optimizer_args)\n```\nit means the loss is calculated from multiple optimizers and it is the same, but optimizers stay different\n\nHowever I am not sure what \"reduce_gradients\" does:\n\n```\ngradients = _fetch_gradients(optimizer)\nall_reduce('sum', gradients, scale=1.0 / count)\n```",
    "923455": "reduce should sum/average all the gradients across the 8 optimizers\nas the state of the opt only depends on the gradients, if the gradients are the same so are the optimizers",
    "923468": "I see - thanks for clarification! :)",
    "924857": "Do you try to use ModelCheckpoint. It save the best model"
  },
  "source": "meta"
}