{
  "id": 222190,
  "title": "Keras callbacks not saving model on TPU",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/222190",
  "author_name": "",
  "post_date": "2021-02-25T18:46:44.774392500Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Apologies if this is a beginner problem, but I can't get my Tensorflow models to save using the same callbacks that I've encountered in several different notebooks when using TPU. I've gone so far as copying line-by-line the exact callbacks that were used in:</p>\n<p><a href=\"https://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline\" target=\"_blank\">https://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline</a><br>\nand <a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training\" target=\"_blank\">https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training</a></p>\n<p>Still, nothing has worked. The kaggle instructions for saving checkpoints with TPU models explain how to save them using this code snippet:</p>\n<pre><code>save_locally = tf.saved_model.SaveOptions(experimental_io_device='/job:localhost')\ncheckpoints_cb = tf.keras.callbacks.ModelCheckpoint('./checkpoints', options=save_locally)\nmodel.fit(…, callbacks=[checkpoints_cb])\n</code></pre>\n<p>But it doesn't explain how to retrieve it. I can save the model to the working directory after training has completed, but this will only save the model in its final form and this is not necessarily the best performing one (<a href=\"https://github.com/keras-team/keras/issues/2768)\" target=\"_blank\">https://github.com/keras-team/keras/issues/2768)</a>. I've also made sure that the models I've tried saving with callbacks all have generic names that end in '.h5'.</p>\n<p>I've been tearing my hair out over this since I can't identify anything in my code that would cause the callbacks not to save their output unlike the other notebooks I linked. If anyone has experienced a similar issue I would appreciate your input, as I've already lost 7 hours of TPU time. </p>",
  "messages": [
    {
      "id": "1218356",
      "postDate": "02/25/2021 18:46:44",
      "content": "<p>Apologies if this is a beginner problem, but I can't get my Tensorflow models to save using the same callbacks that I've encountered in several different notebooks when using TPU. I've gone so far as copying line-by-line the exact callbacks that were used in:</p>\n<p><a href=\"https://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline\" target=\"_blank\">https://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline</a><br>\nand <a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training\" target=\"_blank\">https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training</a></p>\n<p>Still, nothing has worked. The kaggle instructions for saving checkpoints with TPU models explain how to save them using this code snippet:</p>\n<pre><code>save_locally = tf.saved_model.SaveOptions(experimental_io_device='/job:localhost')\ncheckpoints_cb = tf.keras.callbacks.ModelCheckpoint('./checkpoints', options=save_locally)\nmodel.fit(…, callbacks=[checkpoints_cb])\n</code></pre>\n<p>But it doesn't explain how to retrieve it. I can save the model to the working directory after training has completed, but this will only save the model in its final form and this is not necessarily the best performing one (<a href=\"https://github.com/keras-team/keras/issues/2768)\" target=\"_blank\">https://github.com/keras-team/keras/issues/2768)</a>. I've also made sure that the models I've tried saving with callbacks all have generic names that end in '.h5'.</p>\n<p>I've been tearing my hair out over this since I can't identify anything in my code that would cause the callbacks not to save their output unlike the other notebooks I linked. If anyone has experienced a similar issue I would appreciate your input, as I've already lost 7 hours of TPU time. </p>",
      "rawMarkdown": "Apologies if this is a beginner problem, but I can't get my Tensorflow models to save using the same callbacks that I've encountered in several different notebooks when using TPU. I've gone so far as copying line-by-line the exact callbacks that were used in:\n\nhttps://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline\nand https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training\n\nStill, nothing has worked. The kaggle instructions for saving checkpoints with TPU models explain how to save them using this code snippet:\n\n```\nsave_locally = tf.saved_model.SaveOptions(experimental_io_device='/job:localhost')\ncheckpoints_cb = tf.keras.callbacks.ModelCheckpoint('./checkpoints', options=save_locally)\nmodel.fit(…, callbacks=[checkpoints_cb])\n```\n\nBut it doesn't explain how to retrieve it. I can save the model to the working directory after training has completed, but this will only save the model in its final form and this is not necessarily the best performing one (https://github.com/keras-team/keras/issues/2768). I've also made sure that the models I've tried saving with callbacks all have generic names that end in '.h5'.\n\nI've been tearing my hair out over this since I can't identify anything in my code that would cause the callbacks not to save their output unlike the other notebooks I linked. If anyone has experienced a similar issue I would appreciate your input, as I've already lost 7 hours of TPU time.",
      "votes": null
    },
    {
      "id": "1218445",
      "postDate": "02/25/2021 22:04:21",
      "content": "<p>Shoudn't you use <code>save_best_only=True</code> parameter in <code>ModelCheckpoint</code>?</p>",
      "rawMarkdown": "Shoudn't you use `save_best_only=True` parameter in `ModelCheckpoint`?",
      "votes": null
    },
    {
      "id": "1218459",
      "postDate": "02/25/2021 22:19:18",
      "content": "<p>Yes, I have been using that parameter. The main code I've used is:</p>\n<pre><code>checkpoint = tf.keras.callbacks.ModelCheckpoint(model_name, save_best_only=True, save_weights_only=True, monitor='val_auc', mode='max')\nlr_reducer = tf.keras.callbacks.ReduceLROnPlateau(\n    monitor=\"val_auc\", patience=1, min_lr=1e-6, mode='max')   \n\nhistory = model.fit(\n    train_dataset, \n    epochs=25,\n    verbose=1,\n    callbacks=[checkpoint, lr_reducer],\n    steps_per_epoch=steps_per_epoch,\n    validation_data=valid_dataset)\n</code></pre>\n<p>I've tried removing <code>save_weights_only=True</code>, results are the same - nothing is saved to the working directory.</p>",
      "rawMarkdown": "Yes, I have been using that parameter. The main code I've used is:\n\n```\ncheckpoint = tf.keras.callbacks.ModelCheckpoint(model_name, save_best_only=True, save_weights_only=True, monitor='val_auc', mode='max')\nlr_reducer = tf.keras.callbacks.ReduceLROnPlateau(\n    monitor=\"val_auc\", patience=1, min_lr=1e-6, mode='max')   \n\nhistory = model.fit(\n    train_dataset, \n    epochs=25,\n    verbose=1,\n    callbacks=[checkpoint, lr_reducer],\n    steps_per_epoch=steps_per_epoch,\n    validation_data=valid_dataset)\n```\nI've tried removing `save_weights_only=True`, results are the same - nothing is saved to the working directory.",
      "votes": null
    },
    {
      "id": "1219070",
      "postDate": "02/26/2021 12:21:58",
      "content": "<p>Can you show how your <code>model.compile(...)</code> looks?</p>",
      "rawMarkdown": "Can you show how your `model.compile(...)` looks?",
      "votes": null
    },
    {
      "id": "1219215",
      "postDate": "02/26/2021 14:59:52",
      "content": "<pre><code> model.compile(            \n            optimizer=tf.keras.optimizers.Adam(lr=1e-4),\n            loss='binary_crossentropy',\n            metrics=[tf.keras.metrics.AUC(multi_label=True)])\n</code></pre>",
      "rawMarkdown": "```\n model.compile(            \n            optimizer=tf.keras.optimizers.Adam(lr=1e-4),\n            loss='binary_crossentropy',\n            metrics=[tf.keras.metrics.AUC(multi_label=True)])\n```",
      "votes": null
    },
    {
      "id": "1219299",
      "postDate": "02/26/2021 16:48:11",
      "content": "<p>Looks OK. Can you try to save model based on <code>val_loss</code>, not on <code>val_auc</code>(and don't forget to change <code>mode</code> to <code>min</code>)?</p>",
      "rawMarkdown": "Looks OK. Can you try to save model based on `val_loss`, not on `val_auc `(and don't forget to change `mode ` to `min`)?",
      "votes": null
    },
    {
      "id": "1220196",
      "postDate": "02/27/2021 17:52:39",
      "content": "<p>I faced the same issue, looks like some issue with the recent upgrade of tensorflow version and h5py. The sequential api is working but issue with the functional one. I tried running exactly same code which is using functional api  and was working before the tf upgrade but now it does not seem to work. </p>",
      "rawMarkdown": "I faced the same issue, looks like some issue with the recent upgrade of tensorflow version and h5py. The sequential api is working but issue with the functional one. I tried running exactly same code which is using functional api  and was working before the tf upgrade but now it does not seem to work.",
      "votes": null
    },
    {
      "id": "1220206",
      "postDate": "02/27/2021 17:58:58",
      "content": "<p>I tried that last night, again no positive results.</p>\n<p>I've been very careful with my code and can't see anything that would obstruct the callback saving its output. I'm considering contacting the Kaggle technical team over this since I can't share my code privately in this competition.</p>",
      "rawMarkdown": "I tried that last night, again no positive results.\n\nI've been very careful with my code and can't see anything that would obstruct the callback saving its output. I'm considering contacting the Kaggle technical team over this since I can't share my code privately in this competition.",
      "votes": null
    },
    {
      "id": "1220233",
      "postDate": "02/27/2021 18:47:47",
      "content": "<p>You should try previous TF versions with <code>!pip install tensorflow==&lt;version&gt;</code></p>",
      "rawMarkdown": "You should try previous TF versions with `!pip install tensorflow==<version>`",
      "votes": null
    },
    {
      "id": "1221100",
      "postDate": "02/28/2021 17:08:15",
      "content": "<p>model_save = ModelCheckpoint('./model.h5', <br>\n                             save_best_only = True, <br>\n                             save_weights_only = True,<br>\n                             monitor = 'val_loss', <br>\n                             mode = 'min', verbose = 1)</p>\n<p>See here :<br>\n<a href=\"https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline\" target=\"_blank\">https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline</a></p>",
      "rawMarkdown": "model_save = ModelCheckpoint('./model.h5', \n                             save_best_only = True, \n                             save_weights_only = True,\n                             monitor = 'val_loss', \n                             mode = 'min', verbose = 1)\n\n\nSee here :\nhttps://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1218445,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "02/25/2021 22:04:21",
      "content": "<p>Shoudn't you use <code>save_best_only=True</code> parameter in <code>ModelCheckpoint</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1218459,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "02/25/2021 22:19:18",
          "content": "<p>Yes, I have been using that parameter. The main code I've used is:</p>\n<pre><code>checkpoint = tf.keras.callbacks.ModelCheckpoint(model_name, save_best_only=True, save_weights_only=True, monitor='val_auc', mode='max')\nlr_reducer = tf.keras.callbacks.ReduceLROnPlateau(\n    monitor=\"val_auc\", patience=1, min_lr=1e-6, mode='max')   \n\nhistory = model.fit(\n    train_dataset, \n    epochs=25,\n    verbose=1,\n    callbacks=[checkpoint, lr_reducer],\n    steps_per_epoch=steps_per_epoch,\n    validation_data=valid_dataset)\n</code></pre>\n<p>I've tried removing <code>save_weights_only=True</code>, results are the same - nothing is saved to the working directory.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1219070,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/26/2021 12:21:58",
          "content": "<p>Can you show how your <code>model.compile(...)</code> looks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1219215,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "02/26/2021 14:59:52",
          "content": "<pre><code> model.compile(            \n            optimizer=tf.keras.optimizers.Adam(lr=1e-4),\n            loss='binary_crossentropy',\n            metrics=[tf.keras.metrics.AUC(multi_label=True)])\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1219299,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/26/2021 16:48:11",
          "content": "<p>Looks OK. Can you try to save model based on <code>val_loss</code>, not on <code>val_auc</code>(and don't forget to change <code>mode</code> to <code>min</code>)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220206,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "02/27/2021 17:58:58",
          "content": "<p>I tried that last night, again no positive results.</p>\n<p>I've been very careful with my code and can't see anything that would obstruct the callback saving its output. I'm considering contacting the Kaggle technical team over this since I can't share my code privately in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220233,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "02/27/2021 18:47:47",
          "content": "<p>You should try previous TF versions with <code>!pip install tensorflow==&lt;version&gt;</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1220196,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/27/2021 17:52:39",
      "content": "<p>I faced the same issue, looks like some issue with the recent upgrade of tensorflow version and h5py. The sequential api is working but issue with the functional one. I tried running exactly same code which is using functional api  and was working before the tf upgrade but now it does not seem to work. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1221100,
      "author_name": "faisalalsrheed",
      "author_url": "",
      "post_date": "02/28/2021 17:08:15",
      "content": "<p>model_save = ModelCheckpoint('./model.h5', <br>\n                             save_best_only = True, <br>\n                             save_weights_only = True,<br>\n                             monitor = 'val_loss', <br>\n                             mode = 'min', verbose = 1)</p>\n<p>See here :<br>\n<a href=\"https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline\" target=\"_blank\">https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1218356": "Apologies if this is a beginner problem, but I can't get my Tensorflow models to save using the same callbacks that I've encountered in several different notebooks when using TPU. I've gone so far as copying line-by-line the exact callbacks that were used in:\n\nhttps://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline\nand https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training\n\nStill, nothing has worked. The kaggle instructions for saving checkpoints with TPU models explain how to save them using this code snippet:\n\n```\nsave_locally = tf.saved_model.SaveOptions(experimental_io_device='/job:localhost')\ncheckpoints_cb = tf.keras.callbacks.ModelCheckpoint('./checkpoints', options=save_locally)\nmodel.fit(…, callbacks=[checkpoints_cb])\n```\n\nBut it doesn't explain how to retrieve it. I can save the model to the working directory after training has completed, but this will only save the model in its final form and this is not necessarily the best performing one (https://github.com/keras-team/keras/issues/2768). I've also made sure that the models I've tried saving with callbacks all have generic names that end in '.h5'.\n\nI've been tearing my hair out over this since I can't identify anything in my code that would cause the callbacks not to save their output unlike the other notebooks I linked. If anyone has experienced a similar issue I would appreciate your input, as I've already lost 7 hours of TPU time.",
    "1218445": "Shoudn't you use `save_best_only=True` parameter in `ModelCheckpoint`?",
    "1218459": "Yes, I have been using that parameter. The main code I've used is:\n\n```\ncheckpoint = tf.keras.callbacks.ModelCheckpoint(model_name, save_best_only=True, save_weights_only=True, monitor='val_auc', mode='max')\nlr_reducer = tf.keras.callbacks.ReduceLROnPlateau(\n    monitor=\"val_auc\", patience=1, min_lr=1e-6, mode='max')   \n\nhistory = model.fit(\n    train_dataset, \n    epochs=25,\n    verbose=1,\n    callbacks=[checkpoint, lr_reducer],\n    steps_per_epoch=steps_per_epoch,\n    validation_data=valid_dataset)\n```\nI've tried removing `save_weights_only=True`, results are the same - nothing is saved to the working directory.",
    "1219070": "Can you show how your `model.compile(...)` looks?",
    "1219215": "```\n model.compile(            \n            optimizer=tf.keras.optimizers.Adam(lr=1e-4),\n            loss='binary_crossentropy',\n            metrics=[tf.keras.metrics.AUC(multi_label=True)])\n```",
    "1219299": "Looks OK. Can you try to save model based on `val_loss`, not on `val_auc `(and don't forget to change `mode ` to `min`)?",
    "1220196": "I faced the same issue, looks like some issue with the recent upgrade of tensorflow version and h5py. The sequential api is working but issue with the functional one. I tried running exactly same code which is using functional api  and was working before the tf upgrade but now it does not seem to work.",
    "1220206": "I tried that last night, again no positive results.\n\nI've been very careful with my code and can't see anything that would obstruct the callback saving its output. I'm considering contacting the Kaggle technical team over this since I can't share my code privately in this competition.",
    "1220233": "You should try previous TF versions with `!pip install tensorflow==<version>`",
    "1221100": "model_save = ModelCheckpoint('./model.h5', \n                             save_best_only = True, \n                             save_weights_only = True,\n                             monitor = 'val_loss', \n                             mode = 'min', verbose = 1)\n\n\nSee here :\nhttps://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline"
  },
  "source": "meta"
}