{
  "id": 142248,
  "title": "Error when training a loaded model[RESOLVED]",
  "url": "/competitions/flower-classification-with-tpus/discussion/142248",
  "author_name": "Victor Paslay",
  "post_date": "2020-04-09T16:06:48.564000",
  "votes": 5,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I have faced with the error when training a model using TPU:</p>\n\n<p><code>InvalidArgumentError: Unable to find a context_id matching the specified one (9692482182556506708). Perhaps the worker was restarted, or the context was GC'd?</code>\nThe epoch starts and on the first step it throws this exception.</p>\n\n<p>Before that I have trained EfficientNetB7, did a checkpoint, created a dataset and used this dataset in another kernel.\nThe code is the following:\n<code>\ncheckpoint = tf.keras.callbacks.ModelCheckpoint(filepath, monitor='val_loss', verbose=1, save_best_only=True)\nmodel = tf.keras.models.load_model(\"/kaggle/input/tpu-flowers-efficientnetb7/tpu-flowers-weights16-0.18.hdf5\")\nhistory = model.fit(\n        get_training_dataset(train_dataset, do_aug=True, transform=transform), \n        steps_per_epoch = STEPS_PER_EPOCH,\n        epochs = EPOCHS,\n        callbacks = [lr_callback, checkpoint],\n        validation_data = get_validation_dataset(val_dataset),\n        initial_epoch = 16,\n        verbose=1\n    )\n</code></p>\n\n<p>Did anybody face this issue? Were you able to use checkpoints in another kernels with TPUs?\nThanks.</p>\n\n<h2>Solution</h2>\n\n<p>I forgot to load the model under the scope. Instead of the loading code about I should have done this:\n<code>\nwith strategy.scope():\n    model = tf.keras.models.load_model(\"/kaggle/input/tpu-flowers-efficientnetb7/tpu-flowers-weights16-0.18.hdf5\")\n</code></p>\n\n<h2>Update</h2>\n\n<p>I have faced with numerous issues and my current approach is to save the model in tensorflow format(<code>save_format=\"tf\"</code>) and upload to Google Cloud Storage. Then I read the model from GCS bucket during the further training and inference.</p>",
  "messages": [
    {
      "id": 802552,
      "postDate": "2020-04-09T16:06:48.563Z",
      "content": "<p>I have faced with the error when training a model using TPU:</p>\n\n<p><code>InvalidArgumentError: Unable to find a context_id matching the specified one (9692482182556506708). Perhaps the worker was restarted, or the context was GC'd?</code>\nThe epoch starts and on the first step it throws this exception.</p>\n\n<p>Before that I have trained EfficientNetB7, did a checkpoint, created a dataset and used this dataset in another kernel.\nThe code is the following:\n<code>\ncheckpoint = tf.keras.callbacks.ModelCheckpoint(filepath, monitor='val_loss', verbose=1, save_best_only=True)\nmodel = tf.keras.models.load_model(\"/kaggle/input/tpu-flowers-efficientnetb7/tpu-flowers-weights16-0.18.hdf5\")\nhistory = model.fit(\n        get_training_dataset(train_dataset, do_aug=True, transform=transform), \n        steps_per_epoch = STEPS_PER_EPOCH,\n        epochs = EPOCHS,\n        callbacks = [lr_callback, checkpoint],\n        validation_data = get_validation_dataset(val_dataset),\n        initial_epoch = 16,\n        verbose=1\n    )\n</code></p>\n\n<p>Did anybody face this issue? Were you able to use checkpoints in another kernels with TPUs?\nThanks.</p>\n\n<h2>Solution</h2>\n\n<p>I forgot to load the model under the scope. Instead of the loading code about I should have done this:\n<code>\nwith strategy.scope():\n    model = tf.keras.models.load_model(\"/kaggle/input/tpu-flowers-efficientnetb7/tpu-flowers-weights16-0.18.hdf5\")\n</code></p>\n\n<h2>Update</h2>\n\n<p>I have faced with numerous issues and my current approach is to save the model in tensorflow format(<code>save_format=\"tf\"</code>) and upload to Google Cloud Storage. Then I read the model from GCS bucket during the further training and inference.</p>",
      "rawMarkdown": "I have faced with the error when training a model using TPU:\n\n```InvalidArgumentError: Unable to find a context_id matching the specified one (9692482182556506708). Perhaps the worker was restarted, or the context was GC'd?```\nThe epoch starts and on the first step it throws this exception.\n\nBefore that I have trained EfficientNetB7, did a checkpoint, created a dataset and used this dataset in another kernel.\nThe code is the following:\n```\ncheckpoint = tf.keras.callbacks.ModelCheckpoint(filepath, monitor='val_loss', verbose=1, save_best_only=True)\nmodel = tf.keras.models.load_model(\"/kaggle/input/tpu-flowers-efficientnetb7/tpu-flowers-weights16-0.18.hdf5\")\nhistory = model.fit(\n        get_training_dataset(train_dataset, do_aug=True, transform=transform), \n        steps_per_epoch = STEPS_PER_EPOCH,\n        epochs = EPOCHS,\n        callbacks = [lr_callback, checkpoint],\n        validation_data = get_validation_dataset(val_dataset),\n        initial_epoch = 16,\n        verbose=1\n    )\n```\n\nDid anybody face this issue? Were you able to use checkpoints in another kernels with TPUs?\nThanks.\n\n## Solution\nI forgot to load the model under the scope. Instead of the loading code about I should have done this:\n```\nwith strategy.scope():\n    model = tf.keras.models.load_model(\"/kaggle/input/tpu-flowers-efficientnetb7/tpu-flowers-weights16-0.18.hdf5\")\n```\n\n## Update\nI have faced with numerous issues and my current approach is to save the model in tensorflow format(`save_format=\"tf\"`) and upload to Google Cloud Storage. Then I read the model from GCS bucket during the further training and inference.",
      "votes": 5
    },
    {
      "id": 834812,
      "postDate": "2020-05-05T19:44:13.870Z",
      "content": "<p>This is a known bug indeed. Here is a workaround. If you copy the model from TPU to CPU first and then save it, you will be able to reload it in the strategy scope:</p>\n\n<p>```\nmodel_copy = create_model() # whatever your model creation code is\nmodel_copy.set_weights(trained_model.get_weights()) # copy your trained weights\nmodel_copy.save(\"model.h5\")</p>\n\n<p>with strategy.scope():\n  reload_model = tf.keras.models.load_model('model.h5')\n```</p>\n\n<p>I made a notebook with this code:\n<a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-train-save-and-reload-on-tpu\">https://www.kaggle.com/mgornergoogle/five-flowers-train-save-and-reload-on-tpu</a></p>",
      "rawMarkdown": "This is a known bug indeed. Here is a workaround. If you copy the model from TPU to CPU first and then save it, you will be able to reload it in the strategy scope:\n\n```\nmodel_copy = create_model() # whatever your model creation code is\nmodel_copy.set_weights(trained_model.get_weights()) # copy your trained weights\nmodel_copy.save(\"model.h5\")\n\nwith strategy.scope():\n  reload_model = tf.keras.models.load_model('model.h5')\n```\n\nI made a notebook with this code:\nhttps://www.kaggle.com/mgornergoogle/five-flowers-train-save-and-reload-on-tpu",
      "votes": 3,
      "replies": [
        {
          "id": 898739,
          "postDate": "2020-06-23T17:45:56.827Z",
          "content": "<p>Thanks Martin, I had a tough time saving the weights. This is an easy workaround to the problem.</p>",
          "rawMarkdown": "Thanks Martin, I had a tough time saving the weights. This is an easy workaround to the problem."
        }
      ]
    },
    {
      "id": 820357,
      "postDate": "2020-04-25T11:28:04.743Z",
      "content": "<p>I am facing the same issue. But I dont load the model, this error throws during training and it happens after some epochs but, there are no defined epochs .. It just happens in the middle of the training at any epoch.  </p>",
      "rawMarkdown": "I am facing the same issue. But I dont load the model, this error throws during training and it happens after some epochs but, there are no defined epochs .. It just happens in the middle of the training at any epoch.  ",
      "votes": 2
    },
    {
      "id": 852484,
      "postDate": "2020-05-18T13:29:24.027Z",
      "content": "<p>Thank astzls, I have changed the matrix images to np.float32 and it work.🙏 </p>",
      "rawMarkdown": "Thank astzls, I have changed the matrix images to np.float32 and it work.🙏 "
    },
    {
      "id": 830236,
      "postDate": "2020-05-02T12:28:03.823Z",
      "content": "<p>Guys, I have faced with numerous issues and my current approach is to save the model in tensorflow format(save_format=\"tf\") and upload to Google Cloud Storage.  Then I read the model from GCS bucket during the further training and inference.</p>",
      "rawMarkdown": "Guys, I have faced with numerous issues and my current approach is to save the model in tensorflow format(save_format=\"tf\") and upload to Google Cloud Storage.  Then I read the model from GCS bucket during the further training and inference."
    },
    {
      "id": 830160,
      "postDate": "2020-05-02T11:20:15.910Z",
      "content": "<p>When using efn b7, occurs same issue...And I write code in 'with' statement, your solution did not work T_T</p>",
      "rawMarkdown": "When using efn b7, occurs same issue...And I write code in 'with' statement, your solution did not work T_T",
      "replies": [
        {
          "id": 830195,
          "postDate": "2020-05-02T11:54:04.070Z",
          "content": "<p>I change my image matrix type from float64 to float32, and it works well!!!!</p>",
          "rawMarkdown": "I change my image matrix type from float64 to float32, and it works well!!!!"
        },
        {
          "id": 830232,
          "postDate": "2020-05-02T12:21:40.997Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 831152,
          "postDate": "2020-05-03T07:31:59.857Z",
          "content": "<p>Sure, tf.cast(example['image'],tf.float32) in read_labeled_dataset func </p>",
          "rawMarkdown": "Sure, tf.cast(example['image'],tf.float32) in read_labeled_dataset func ",
          "votes": 1
        },
        {
          "id": 833815,
          "postDate": "2020-05-05T04:34:27.170Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 802846,
      "postDate": "2020-04-09T21:30:34.353Z",
      "content": "<p>I'm not sure, but you may need to move the checkpoint to gcs before loading into the model on the TPU. I've been training off kaggle and pushing notebooks in with a checkpoint in a dataset and then creating a copy of the dataset on gcs with something like this:</p>\n\n<p><code>\nfrom kaggle_datasets import KaggleDatasets\nDATASET_DIR = Path('/kaggle/input/your-dataset')\nGCS_DATASET_DIR = KaggleDatasets().get_gcs_path(DATASET_DIR.parts[-1])\n</code>\nMy understanding is that this creates a copy of the dataset on gcs in the same zone/region as the TPU so the TPU can read from it. Then you should be able to load the checkpoint from wherever your checkpoint is in <code>GCS_DATASET_DIR</code>.</p>",
      "rawMarkdown": "I'm not sure, but you may need to move the checkpoint to gcs before loading into the model on the TPU. I've been training off kaggle and pushing notebooks in with a checkpoint in a dataset and then creating a copy of the dataset on gcs with something like this:\n\n```\nfrom kaggle_datasets import KaggleDatasets\nDATASET_DIR = Path('/kaggle/input/your-dataset')\nGCS_DATASET_DIR = KaggleDatasets().get_gcs_path(DATASET_DIR.parts[-1])\n```\nMy understanding is that this creates a copy of the dataset on gcs in the same zone/region as the TPU so the TPU can read from it. Then you should be able to load the checkpoint from wherever your checkpoint is in `GCS_DATASET_DIR`.",
      "replies": [
        {
          "id": 804480,
          "postDate": "2020-04-11T16:04:04.900Z",
          "content": "<p>Did not help, the same problem but thank you.</p>",
          "rawMarkdown": "Did not help, the same problem but thank you."
        }
      ]
    },
    {
      "id": 852482,
      "postDate": "2020-05-18T13:28:17.483Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 829852,
      "postDate": "2020-05-02T05:57:03.363Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 829864,
          "postDate": "2020-05-02T06:09:26.080Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 833725,
          "postDate": "2020-05-05T03:09:19.713Z",
          "content": "<p>Same problem hah, tf developer suggested me saving model as tf format, but I have not tried yet..</p>",
          "rawMarkdown": "Same problem hah, tf developer suggested me saving model as tf format, but I have not tried yet.."
        },
        {
          "id": 833814,
          "postDate": "2020-05-05T04:34:09.763Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 834812,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-05-05T19:44:13.870000",
      "content": "<p>This is a known bug indeed. Here is a workaround. If you copy the model from TPU to CPU first and then save it, you will be able to reload it in the strategy scope:</p>\n\n<p>```\nmodel_copy = create_model() # whatever your model creation code is\nmodel_copy.set_weights(trained_model.get_weights()) # copy your trained weights\nmodel_copy.save(\"model.h5\")</p>\n\n<p>with strategy.scope():\n  reload_model = tf.keras.models.load_model('model.h5')\n```</p>\n\n<p>I made a notebook with this code:\n<a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-train-save-and-reload-on-tpu\">https://www.kaggle.com/mgornergoogle/five-flowers-train-save-and-reload-on-tpu</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 898739,
          "author_name": "Jesudas DSouza",
          "author_url": "",
          "post_date": "2020-06-23T17:45:56.827000",
          "content": "<p>Thanks Martin, I had a tough time saving the weights. This is an easy workaround to the problem.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 820357,
      "author_name": "torch",
      "author_url": "",
      "post_date": "2020-04-25T11:28:04.743000",
      "content": "<p>I am facing the same issue. But I dont load the model, this error throws during training and it happens after some epochs but, there are no defined epochs .. It just happens in the middle of the training at any epoch.  </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 852484,
      "author_name": "Lan Dinh",
      "author_url": "",
      "post_date": "2020-05-18T13:29:24.027000",
      "content": "<p>Thank astzls, I have changed the matrix images to np.float32 and it work.🙏 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 830236,
      "author_name": "Victor Paslay",
      "author_url": "",
      "post_date": "2020-05-02T12:28:03.823000",
      "content": "<p>Guys, I have faced with numerous issues and my current approach is to save the model in tensorflow format(save_format=\"tf\") and upload to Google Cloud Storage.  Then I read the model from GCS bucket during the further training and inference.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 830160,
      "author_name": "astzls",
      "author_url": "",
      "post_date": "2020-05-02T11:20:15.910000",
      "content": "<p>When using efn b7, occurs same issue...And I write code in 'with' statement, your solution did not work T_T</p>",
      "votes": 0,
      "replies": [
        {
          "id": 830195,
          "author_name": "astzls",
          "author_url": "",
          "post_date": "2020-05-02T11:54:04.070000",
          "content": "<p>I change my image matrix type from float64 to float32, and it works well!!!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 830232,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-02T12:21:40.997000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 831152,
          "author_name": "astzls",
          "author_url": "",
          "post_date": "2020-05-03T07:31:59.857000",
          "content": "<p>Sure, tf.cast(example['image'],tf.float32) in read_labeled_dataset func </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 833815,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-05T04:34:27.170000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 802846,
      "author_name": "Caleb",
      "author_url": "",
      "post_date": "2020-04-09T21:30:34.353000",
      "content": "<p>I'm not sure, but you may need to move the checkpoint to gcs before loading into the model on the TPU. I've been training off kaggle and pushing notebooks in with a checkpoint in a dataset and then creating a copy of the dataset on gcs with something like this:</p>\n\n<p><code>\nfrom kaggle_datasets import KaggleDatasets\nDATASET_DIR = Path('/kaggle/input/your-dataset')\nGCS_DATASET_DIR = KaggleDatasets().get_gcs_path(DATASET_DIR.parts[-1])\n</code>\nMy understanding is that this creates a copy of the dataset on gcs in the same zone/region as the TPU so the TPU can read from it. Then you should be able to load the checkpoint from wherever your checkpoint is in <code>GCS_DATASET_DIR</code>.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 804480,
          "author_name": "Victor Paslay",
          "author_url": "",
          "post_date": "2020-04-11T16:04:04.900000",
          "content": "<p>Did not help, the same problem but thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 852482,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-18T13:28:17.483000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 829852,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-02T05:57:03.363000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 829864,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-02T06:09:26.080000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 833725,
          "author_name": "astzls",
          "author_url": "",
          "post_date": "2020-05-05T03:09:19.713000",
          "content": "<p>Same problem hah, tf developer suggested me saving model as tf format, but I have not tried yet..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 833814,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-05T04:34:09.763000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "802552": "I have faced with the error when training a model using TPU:\n\n```InvalidArgumentError: Unable to find a context_id matching the specified one (9692482182556506708). Perhaps the worker was restarted, or the context was GC'd?```\nThe epoch starts and on the first step it throws this exception.\n\nBefore that I have trained EfficientNetB7, did a checkpoint, created a dataset and used this dataset in another kernel.\nThe code is the following:\n```\ncheckpoint = tf.keras.callbacks.ModelCheckpoint(filepath, monitor='val_loss', verbose=1, save_best_only=True)\nmodel = tf.keras.models.load_model(\"/kaggle/input/tpu-flowers-efficientnetb7/tpu-flowers-weights16-0.18.hdf5\")\nhistory = model.fit(\n        get_training_dataset(train_dataset, do_aug=True, transform=transform), \n        steps_per_epoch = STEPS_PER_EPOCH,\n        epochs = EPOCHS,\n        callbacks = [lr_callback, checkpoint],\n        validation_data = get_validation_dataset(val_dataset),\n        initial_epoch = 16,\n        verbose=1\n    )\n```\n\nDid anybody face this issue? Were you able to use checkpoints in another kernels with TPUs?\nThanks.\n\n## Solution\nI forgot to load the model under the scope. Instead of the loading code about I should have done this:\n```\nwith strategy.scope():\n    model = tf.keras.models.load_model(\"/kaggle/input/tpu-flowers-efficientnetb7/tpu-flowers-weights16-0.18.hdf5\")\n```\n\n## Update\nI have faced with numerous issues and my current approach is to save the model in tensorflow format(`save_format=\"tf\"`) and upload to Google Cloud Storage. Then I read the model from GCS bucket during the further training and inference.",
    "834812": "This is a known bug indeed. Here is a workaround. If you copy the model from TPU to CPU first and then save it, you will be able to reload it in the strategy scope:\n\n```\nmodel_copy = create_model() # whatever your model creation code is\nmodel_copy.set_weights(trained_model.get_weights()) # copy your trained weights\nmodel_copy.save(\"model.h5\")\n\nwith strategy.scope():\n  reload_model = tf.keras.models.load_model('model.h5')\n```\n\nI made a notebook with this code:\nhttps://www.kaggle.com/mgornergoogle/five-flowers-train-save-and-reload-on-tpu",
    "820357": "I am facing the same issue. But I dont load the model, this error throws during training and it happens after some epochs but, there are no defined epochs .. It just happens in the middle of the training at any epoch.  ",
    "852484": "Thank astzls, I have changed the matrix images to np.float32 and it work.🙏 ",
    "830236": "Guys, I have faced with numerous issues and my current approach is to save the model in tensorflow format(save_format=\"tf\") and upload to Google Cloud Storage.  Then I read the model from GCS bucket during the further training and inference.",
    "830160": "When using efn b7, occurs same issue...And I write code in 'with' statement, your solution did not work T_T",
    "802846": "I'm not sure, but you may need to move the checkpoint to gcs before loading into the model on the TPU. I've been training off kaggle and pushing notebooks in with a checkpoint in a dataset and then creating a copy of the dataset on gcs with something like this:\n\n```\nfrom kaggle_datasets import KaggleDatasets\nDATASET_DIR = Path('/kaggle/input/your-dataset')\nGCS_DATASET_DIR = KaggleDatasets().get_gcs_path(DATASET_DIR.parts[-1])\n```\nMy understanding is that this creates a copy of the dataset on gcs in the same zone/region as the TPU so the TPU can read from it. Then you should be able to load the checkpoint from wherever your checkpoint is in `GCS_DATASET_DIR`.",
    "852482": "",
    "829852": ""
  }
}