{
  "id": 130657,
  "title": "How to save checkpoint?",
  "url": "/competitions/flower-classification-with-tpus/discussion/130657",
  "author_name": "",
  "post_date": "2020-02-15T15:25:33.423115200Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I normally use the following code to save checkpoints</p>\n\n<pre><code>checkpoint_path = GCS_DS_PATH_CKPT\nckpt = tf.train.Checkpoint(model=flower_classifier, optimizer=optimizer)\nckpt_manager = tf.train.CheckpointManager(ckpt, checkpoint_path, max_to_keep=10)\n\n# if a checkpoint exists, restore the latest checkpoint.\nif ckpt_manager.latest_checkpoint:\n    ckpt.restore(ckpt_manager.latest_checkpoint)\n    last_epoch = int(ckpt_manager.latest_checkpoint.split(\"-\")[-1])\n    print (f'Latest checkpoint restored -- Model trained for {last_epoch} epochs')\nelse:\n    print('Checkpoint not found. Train from scratch')\n    last_epoch = 0\n\n--- some training ---\n\nckpt_save_path = ckpt_manager.save()\nprint ('\\nSaving checkpoint for epoch {} at {}'.format(epoch + 1, ckpt_save_path))\n</code></pre>\n\n<p>With TPU, the checkpoints are saved automatically GCP buckets. However, with Kaggle TPU, this is not working. I got error like</p>\n\n<pre><code>\"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.create access to kds-1c8dede3674a34644c1d44089dc67551957904190b386b496dc45e21/ckpt-1_temp_08325598b2824c369bfb6'\n\"when initiating an upload to gs://kds-1c8dede3674a34644c1d44089dc67551957904190b386b496dc45e21/ckpt-1_temp_08325598b2824c369bfb69fe562844b5/\"    \n</code></pre>\n\n<p>Saving checkpoints is probably not important for this competition, since the training is quite fast. However, for future competitions, the need of saving and accessing to saved checkpoints are necessary, I think.</p>\n\n<p>So the question is -- How to save checkpoints and access to them, when TPU is used in a Kaggle competition?</p>",
  "messages": [
    {
      "id": "746812",
      "postDate": "02/15/2020 15:25:33",
      "content": "<p>I normally use the following code to save checkpoints</p>\n\n<pre><code>checkpoint_path = GCS_DS_PATH_CKPT\nckpt = tf.train.Checkpoint(model=flower_classifier, optimizer=optimizer)\nckpt_manager = tf.train.CheckpointManager(ckpt, checkpoint_path, max_to_keep=10)\n\n# if a checkpoint exists, restore the latest checkpoint.\nif ckpt_manager.latest_checkpoint:\n    ckpt.restore(ckpt_manager.latest_checkpoint)\n    last_epoch = int(ckpt_manager.latest_checkpoint.split(\"-\")[-1])\n    print (f'Latest checkpoint restored -- Model trained for {last_epoch} epochs')\nelse:\n    print('Checkpoint not found. Train from scratch')\n    last_epoch = 0\n\n--- some training ---\n\nckpt_save_path = ckpt_manager.save()\nprint ('\\nSaving checkpoint for epoch {} at {}'.format(epoch + 1, ckpt_save_path))\n</code></pre>\n\n<p>With TPU, the checkpoints are saved automatically GCP buckets. However, with Kaggle TPU, this is not working. I got error like</p>\n\n<pre><code>\"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.create access to kds-1c8dede3674a34644c1d44089dc67551957904190b386b496dc45e21/ckpt-1_temp_08325598b2824c369bfb6'\n\"when initiating an upload to gs://kds-1c8dede3674a34644c1d44089dc67551957904190b386b496dc45e21/ckpt-1_temp_08325598b2824c369bfb69fe562844b5/\"    \n</code></pre>\n\n<p>Saving checkpoints is probably not important for this competition, since the training is quite fast. However, for future competitions, the need of saving and accessing to saved checkpoints are necessary, I think.</p>\n\n<p>So the question is -- How to save checkpoints and access to them, when TPU is used in a Kaggle competition?</p>",
      "rawMarkdown": "I normally use the following code to save checkpoints\n\n    checkpoint_path = GCS_DS_PATH_CKPT\n    ckpt = tf.train.Checkpoint(model=flower_classifier, optimizer=optimizer)\n    ckpt_manager = tf.train.CheckpointManager(ckpt, checkpoint_path, max_to_keep=10)\n\n    # if a checkpoint exists, restore the latest checkpoint.\n    if ckpt_manager.latest_checkpoint:\n        ckpt.restore(ckpt_manager.latest_checkpoint)\n        last_epoch = int(ckpt_manager.latest_checkpoint.split(\"-\")[-1])\n        print (f'Latest checkpoint restored -- Model trained for {last_epoch} epochs')\n    else:\n        print('Checkpoint not found. Train from scratch')\n        last_epoch = 0\n\n    --- some training ---\n\n    ckpt_save_path = ckpt_manager.save()\n    print ('\\nSaving checkpoint for epoch {} at {}'.format(epoch + 1, ckpt_save_path))\n\nWith TPU, the checkpoints are saved automatically GCP buckets. However, with Kaggle TPU, this is not working. I got error like\n\n    \"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.create access to kds-1c8dede3674a34644c1d44089dc67551957904190b386b496dc45e21/ckpt-1_temp_08325598b2824c369bfb6'\n    \"when initiating an upload to gs://kds-1c8dede3674a34644c1d44089dc67551957904190b386b496dc45e21/ckpt-1_temp_08325598b2824c369bfb69fe562844b5/\"    \n \nSaving checkpoints is probably not important for this competition, since the training is quite fast. However, for future competitions, the need of saving and accessing to saved checkpoints are necessary, I think.\n\nSo the question is -- How to save checkpoints and access to them, when TPU is used in a Kaggle competition?",
      "votes": null
    },
    {
      "id": "749557",
      "postDate": "02/18/2020 19:22:33",
      "content": "<p>Indeed, on Kaggle you will not be able to save a checkpoint from a TPU as we do not offer writable GCS buckets for TPUs to write to.\nIf you want to save a model at the end of a training run and restore from there, there is an example of that in the TPU documentation sample: <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">Five flowers with Keras and Xception on TPU</a></p>",
      "rawMarkdown": "Indeed, on Kaggle you will not be able to save a checkpoint from a TPU as we do not offer writable GCS buckets for TPUs to write to.\nIf you want to save a model at the end of a training run and restore from there, there is an example of that in the TPU documentation sample: [Five flowers with Keras and Xception on TPU](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu)",
      "votes": null
    },
    {
      "id": "749558",
      "postDate": "02/18/2020 19:23:22",
      "content": "<p>And yes, saving the model requires some juggling: copy the TPU model to CPU, save from there to localhost.</p>",
      "rawMarkdown": "And yes, saving the model requires some juggling: copy the TPU model to CPU, save from there to localhost.",
      "votes": null
    },
    {
      "id": "757476",
      "postDate": "02/26/2020 19:53:27",
      "content": "<p>It looks like <a href=\"/dimitreoliveira\">@dimitreoliveira</a> has found a way to use <code>ModelCheckpoint</code> with a TPU here:\n<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455</a></p>",
      "rawMarkdown": "It looks like @dimitreoliveira has found a way to use `ModelCheckpoint` with a TPU here:\nhttps://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455",
      "votes": null
    },
    {
      "id": "814773",
      "postDate": "04/21/2020 01:03:04",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> But what if I want to restore optimizer weights ? Is it possible ?</p>",
      "rawMarkdown": "mgornergoogle But what if I want to restore optimizer weights ? Is it possible ?",
      "votes": null
    },
    {
      "id": "816953",
      "postDate": "04/22/2020 18:22:00",
      "content": "<p>Another option is to train on colab, although the tpus there I believe are the v2 versions instead of the v3 versions in kaggle so have less memory. To get the checkpoint in to kaggle you can create a dataset with the checkpoint in it and then port that over to gcs in kaggle the same way you would for the competition dataset with something like:</p>\n\n<p><code>\nDATASET_DIR = '/kaggle/input'/DATASET_DIR\nGCS_DATASET_DIR = KaggleDatasets().get_gcs_path(DATASET_DIR.parts[-1])\n</code></p>\n\n<p>That puts the checkpoint on gcs in the same zone as the tpu, so you should then be able to load it.</p>",
      "rawMarkdown": "Another option is to train on colab, although the tpus there I believe are the v2 versions instead of the v3 versions in kaggle so have less memory. To get the checkpoint in to kaggle you can create a dataset with the checkpoint in it and then port that over to gcs in kaggle the same way you would for the competition dataset with something like:\n\n```\nDATASET_DIR = '/kaggle/input'/DATASET_DIR\nGCS_DATASET_DIR = KaggleDatasets().get_gcs_path(DATASET_DIR.parts[-1])\n```\n\nThat puts the checkpoint on gcs in the same zone as the tpu, so you should then be able to load it.",
      "votes": null
    },
    {
      "id": "817319",
      "postDate": "04/23/2020 04:08:13",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1353016",
      "postDate": "06/16/2021 19:17:40",
      "content": "<p>You can save to the local drive by using \"options=tf.train.CheckpointOptions(experimental_io_device='/job:localhost')\" when defining the checkpoint object, like this:</p>\n<p><code>checkpoint = tf.keras.callbacks.ModelCheckpoint(SAVE_CHECKPOINT_FILE, monitor='loss',                                  \n                                                  save_weights_only=True, save_best_only=True, save_freq='epoch',                                  \n                                                  options=tf.train.CheckpointOptions(experimental_io_device='/job:localhost'))\n</code></p>",
      "rawMarkdown": "You can save to the local drive by using \"options=tf.train.CheckpointOptions(experimental_io_device='/job:localhost')\" when defining the checkpoint object, like this:\n\n`        checkpoint = tf.keras.callbacks.ModelCheckpoint(SAVE_CHECKPOINT_FILE, monitor='loss',                                  \n                                                  save_weights_only=True, save_best_only=True, save_freq='epoch',                                  \n                                                  options=tf.train.CheckpointOptions(experimental_io_device='/job:localhost'))\n`",
      "votes": null
    },
    {
      "id": "1593761",
      "postDate": "11/24/2021 08:36:16",
      "content": "<p>Thanks, I was looking for this answer for a while :) </p>",
      "rawMarkdown": "Thanks, I was looking for this answer for a while :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1353016,
      "author_name": "mvenou",
      "author_url": "",
      "post_date": "06/16/2021 19:17:40",
      "content": "<p>You can save to the local drive by using \"options=tf.train.CheckpointOptions(experimental_io_device='/job:localhost')\" when defining the checkpoint object, like this:</p>\n<p><code>checkpoint = tf.keras.callbacks.ModelCheckpoint(SAVE_CHECKPOINT_FILE, monitor='loss',                                  \n                                                  save_weights_only=True, save_best_only=True, save_freq='epoch',                                  \n                                                  options=tf.train.CheckpointOptions(experimental_io_device='/job:localhost'))\n</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1593761,
          "author_name": "maciejdzieyc",
          "author_url": "",
          "post_date": "11/24/2021 08:36:16",
          "content": "<p>Thanks, I was looking for this answer for a while :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749557,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/18/2020 19:22:33",
      "content": "<p>Indeed, on Kaggle you will not be able to save a checkpoint from a TPU as we do not offer writable GCS buckets for TPUs to write to.\nIf you want to save a model at the end of a training run and restore from there, there is an example of that in the TPU documentation sample: <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">Five flowers with Keras and Xception on TPU</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 749558,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 19:23:22",
          "content": "<p>And yes, saving the model requires some juggling: copy the TPU model to CPU, save from there to localhost.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 814773,
          "author_name": "goldenlock",
          "author_url": "",
          "post_date": "04/21/2020 01:03:04",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> But what if I want to restore optimizer weights ? Is it possible ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 757476,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/26/2020 19:53:27",
      "content": "<p>It looks like <a href=\"/dimitreoliveira\">@dimitreoliveira</a> has found a way to use <code>ModelCheckpoint</code> with a TPU here:\n<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 816953,
      "author_name": "calebeverett",
      "author_url": "",
      "post_date": "04/22/2020 18:22:00",
      "content": "<p>Another option is to train on colab, although the tpus there I believe are the v2 versions instead of the v3 versions in kaggle so have less memory. To get the checkpoint in to kaggle you can create a dataset with the checkpoint in it and then port that over to gcs in kaggle the same way you would for the competition dataset with something like:</p>\n\n<p><code>\nDATASET_DIR = '/kaggle/input'/DATASET_DIR\nGCS_DATASET_DIR = KaggleDatasets().get_gcs_path(DATASET_DIR.parts[-1])\n</code></p>\n\n<p>That puts the checkpoint on gcs in the same zone as the tpu, so you should then be able to load it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 817319,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "04/23/2020 04:08:13",
          "content": "<p>Thanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "746812": "I normally use the following code to save checkpoints\n\n    checkpoint_path = GCS_DS_PATH_CKPT\n    ckpt = tf.train.Checkpoint(model=flower_classifier, optimizer=optimizer)\n    ckpt_manager = tf.train.CheckpointManager(ckpt, checkpoint_path, max_to_keep=10)\n\n    # if a checkpoint exists, restore the latest checkpoint.\n    if ckpt_manager.latest_checkpoint:\n        ckpt.restore(ckpt_manager.latest_checkpoint)\n        last_epoch = int(ckpt_manager.latest_checkpoint.split(\"-\")[-1])\n        print (f'Latest checkpoint restored -- Model trained for {last_epoch} epochs')\n    else:\n        print('Checkpoint not found. Train from scratch')\n        last_epoch = 0\n\n    --- some training ---\n\n    ckpt_save_path = ckpt_manager.save()\n    print ('\\nSaving checkpoint for epoch {} at {}'.format(epoch + 1, ckpt_save_path))\n\nWith TPU, the checkpoints are saved automatically GCP buckets. However, with Kaggle TPU, this is not working. I got error like\n\n    \"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.create access to kds-1c8dede3674a34644c1d44089dc67551957904190b386b496dc45e21/ckpt-1_temp_08325598b2824c369bfb6'\n    \"when initiating an upload to gs://kds-1c8dede3674a34644c1d44089dc67551957904190b386b496dc45e21/ckpt-1_temp_08325598b2824c369bfb69fe562844b5/\"    \n \nSaving checkpoints is probably not important for this competition, since the training is quite fast. However, for future competitions, the need of saving and accessing to saved checkpoints are necessary, I think.\n\nSo the question is -- How to save checkpoints and access to them, when TPU is used in a Kaggle competition?",
    "749557": "Indeed, on Kaggle you will not be able to save a checkpoint from a TPU as we do not offer writable GCS buckets for TPUs to write to.\nIf you want to save a model at the end of a training run and restore from there, there is an example of that in the TPU documentation sample: [Five flowers with Keras and Xception on TPU](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu)",
    "749558": "And yes, saving the model requires some juggling: copy the TPU model to CPU, save from there to localhost.",
    "757476": "It looks like @dimitreoliveira has found a way to use `ModelCheckpoint` with a TPU here:\nhttps://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455",
    "814773": "mgornergoogle But what if I want to restore optimizer weights ? Is it possible ?",
    "816953": "Another option is to train on colab, although the tpus there I believe are the v2 versions instead of the v3 versions in kaggle so have less memory. To get the checkpoint in to kaggle you can create a dataset with the checkpoint in it and then port that over to gcs in kaggle the same way you would for the competition dataset with something like:\n\n```\nDATASET_DIR = '/kaggle/input'/DATASET_DIR\nGCS_DATASET_DIR = KaggleDatasets().get_gcs_path(DATASET_DIR.parts[-1])\n```\n\nThat puts the checkpoint on gcs in the same zone as the tpu, so you should then be able to load it.",
    "817319": "Thanks for sharing",
    "1353016": "You can save to the local drive by using \"options=tf.train.CheckpointOptions(experimental_io_device='/job:localhost')\" when defining the checkpoint object, like this:\n\n`        checkpoint = tf.keras.callbacks.ModelCheckpoint(SAVE_CHECKPOINT_FILE, monitor='loss',                                  \n                                                  save_weights_only=True, save_best_only=True, save_freq='epoch',                                  \n                                                  options=tf.train.CheckpointOptions(experimental_io_device='/job:localhost'))\n`",
    "1593761": "Thanks, I was looking for this answer for a while :)"
  },
  "source": "meta"
}