{
  "id": 131045,
  "title": "How to release  the TPU memory?",
  "url": "/competitions/flower-classification-with-tpus/discussion/131045",
  "author_name": "",
  "post_date": "2020-02-18T00:52:51.537019200Z",
  "votes": 9,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I tried to train 5 fold Effientnet B6 in a kernel. But after two folds only , then I got the  TPU ResourceExhaustedError.  I tried to use 'K.clear_session()' to release the TPU memory, but it did not work.  : ( </p>",
  "messages": [
    {
      "id": "748767",
      "postDate": "02/18/2020 00:52:51",
      "content": "<p>Hi,</p>\n\n<p>I tried to train 5 fold Effientnet B6 in a kernel. But after two folds only , then I got the  TPU ResourceExhaustedError.  I tried to use 'K.clear_session()' to release the TPU memory, but it did not work.  : ( </p>",
      "rawMarkdown": "Hi,\n\nI tried to train 5 fold Effientnet B6 in a kernel. But after two folds only , then I got the  TPU ResourceExhaustedError.  I tried to use 'K.clear_session()' to release the TPU memory, but it did not work.  : (",
      "votes": null
    },
    {
      "id": "749774",
      "postDate": "02/18/2020 22:29:30",
      "content": "<p>Not sure what you are doing here. Could you explain further what you mean by \"5 fold\" or maybe share some code ?</p>",
      "rawMarkdown": "Not sure what you are doing here. Could you explain further what you mean by \"5 fold\" or maybe share some code ?",
      "votes": null
    },
    {
      "id": "749877",
      "postDate": "02/19/2020 00:48:06",
      "content": "<p>'5 fold' mean that we use kfold api(from sklearn ) to split the all train data into 5 parts.  Everytime we choose one of them for validation, four of them for train. Do the same thing 5 times to train 5 different models. You can see this <a href=\"https://www.kaggle.com/ragnar123/5-kfold-densenet201\">kernel</a> . If you replace the densenet201 with efficentnet b6, you will get the error infomation. </p>",
      "rawMarkdown": "'5 fold' mean that we use kfold api(from sklearn ) to split the all train data into 5 parts.  Everytime we choose one of them for validation, four of them for train. Do the same thing 5 times to train 5 different models. You can see this [kernel](https://www.kaggle.com/ragnar123/5-kfold-densenet201) . If you replace the densenet201 with efficentnet b6, you will get the error infomation.",
      "votes": null
    },
    {
      "id": "749903",
      "postDate": "02/19/2020 01:44:21",
      "content": "<p>Got it. Let me check.</p>",
      "rawMarkdown": "Got it. Let me check.",
      "votes": null
    },
    {
      "id": "749923",
      "postDate": "02/19/2020 02:20:00",
      "content": "<p>I had the same issue and as a workaround, I changed the Image size from 512 to 331.</p>\n\n<p>but with 512 image size I was also getting ResourceExhaustedError after 2 folds and tf.keras.backend.clear_session() didn't work for me.</p>",
      "rawMarkdown": "I had the same issue and as a workaround, I changed the Image size from 512 to 331.\n\nbut with 512 image size I was also getting ResourceExhaustedError after 2 folds and tf.keras.backend.clear_session() didn't work for me.",
      "votes": null
    },
    {
      "id": "750953",
      "postDate": "02/19/2020 21:41:02",
      "content": "<p>TPU memory is garbage collected when the model itself is garbage collected in your Python code.</p>\n\n<p>Here, you are trying to hold all of the models in TPU memory. However, your settings (image size 512x512px, batch_size=16*8=128 are optimized to use up as much TPU memory as is available to shoot for max accuracy and training speed. You will not be able to hold 5 copies of the model on the TPU.</p>\n\n<p>You can try to save the models to disk after training, then start the next fold with a  new model. There is code for saving/reloading a model in the TPU docs sample: <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">Five flowers with Keras and Xception on TPU</a></p>\n\n<p>When you reload the models, you can choose wether you reload them on TPU or CPU by doing the loading inside or outside of a <code>strategy.scope()</code></p>\n\n<p>(tf.keras.backend.clear_session() is not useful here)</p>",
      "rawMarkdown": "TPU memory is garbage collected when the model itself is garbage collected in your Python code.\n\nHere, you are trying to hold all of the models in TPU memory. However, your settings (image size 512x512px, batch_size=16*8=128 are optimized to use up as much TPU memory as is available to shoot for max accuracy and training speed. You will not be able to hold 5 copies of the model on the TPU.\n\nYou can try to save the models to disk after training, then start the next fold with a  new model. There is code for saving/reloading a model in the TPU docs sample: [Five flowers with Keras and Xception on TPU](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu)\n\nWhen you reload the models, you can choose wether you reload them on TPU or CPU by doing the loading inside or outside of a `strategy.scope()`\n\n(tf.keras.backend.clear_session() is not useful here)",
      "votes": null
    },
    {
      "id": "750954",
      "postDate": "02/19/2020 21:41:23",
      "content": "<p>I posted the answer above.</p>",
      "rawMarkdown": "I posted the answer above.",
      "votes": null
    },
    {
      "id": "751038",
      "postDate": "02/20/2020 00:22:20",
      "content": "<p>Hi Martin,</p>\n\n<p>Thanks for your answer. But I still want to know that if tensorflow has a API to release the TPU memory like pytorch ( torch.cuda.empty_cache())?  If you use pytorch ,you can very easily to delete unused objects and release the gpu memory.  </p>",
      "rawMarkdown": "Hi Martin,\n\nThanks for your answer. But I still want to know that if tensorflow has a API to release the TPU memory like pytorch ( torch.cuda.empty_cache())?  If you use pytorch ,you can very easily to delete unused objects and release the gpu memory.",
      "votes": null
    },
    {
      "id": "751100",
      "postDate": "02/20/2020 01:55:15",
      "content": "<p>In Tensorflow, to free TPU memory, release the model object.</p>",
      "rawMarkdown": "In Tensorflow, to free TPU memory, release the model object.",
      "votes": null
    },
    {
      "id": "751108",
      "postDate": "02/20/2020 02:07:59",
      "content": "<p>Thanks.  I will try it. </p>",
      "rawMarkdown": "Thanks.  I will try it.",
      "votes": null
    },
    {
      "id": "792877",
      "postDate": "03/31/2020 15:58:43",
      "content": "<p>hi , does it means  that  'del model, gc.collect()  '   can free tpu memory ? <a href=\"/qinhui1999\">@qinhui1999</a> </p>",
      "rawMarkdown": "hi , does it means  that  'del model, gc.collect()  '   can free tpu memory ? @qinhui1999",
      "votes": null
    },
    {
      "id": "793151",
      "postDate": "03/31/2020 20:47:23",
      "content": "<p>yes, that's how it is supposed to work</p>",
      "rawMarkdown": "yes, that's how it is supposed to work",
      "votes": null
    },
    {
      "id": "793298",
      "postDate": "03/31/2020 23:46:24",
      "content": "<p><a href=\"/bestpredict\">@bestpredict</a>  <a href=\"/mgornergoogle\">@mgornergoogle</a> No, it does not work.  :(</p>",
      "rawMarkdown": "bestpredict  @mgornergoogle No, it does not work.  :(",
      "votes": null
    },
    {
      "id": "793300",
      "postDate": "03/31/2020 23:51:04",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> You can see my kernel: <a href=\"https://www.kaggle.com/qinhui1999/tpu-enet-b7-incepention-b6\">https://www.kaggle.com/qinhui1999/tpu-enet-b7-incepention-b6</a>.  'del model and gc.collect()'  can not release the TPU memory.</p>",
      "rawMarkdown": "mgornergoogle You can see my kernel: https://www.kaggle.com/qinhui1999/tpu-enet-b7-incepention-b6.  'del model and gc.collect()'  can not release the TPU memory.",
      "votes": null
    },
    {
      "id": "793326",
      "postDate": "04/01/2020 00:27:45",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I agree with   <a href=\"/qinhui1999\">@qinhui1999</a>  sad and  I had  try  ' del model, gc.collect() '   today,  it seems did not work.  sad :(</p>",
      "rawMarkdown": "mgornergoogle I agree with   @qinhui1999  sad and  I had  try  ' del model, gc.collect() '   today,  it seems did not work.  sad :(",
      "votes": null
    },
    {
      "id": "793528",
      "postDate": "04/01/2020 05:06:03",
      "content": "<p>Thank you for the reproducible example. I'm submitting this to the TF team.</p>",
      "rawMarkdown": "Thank you for the reproducible example. I'm submitting this to the TF team.",
      "votes": null
    },
    {
      "id": "794181",
      "postDate": "04/01/2020 16:27:41",
      "content": "<p>I confirm there seems to be a problem with this sequence of models.</p>\n\n<p>There is an easy workaround:\nYou can call <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> between each model training to completely reinitialize the TPU. I've tested it in your notebook and it works.</p>",
      "rawMarkdown": "I confirm there seems to be a problem with this sequence of models.\n\nThere is an easy workaround:\nYou can call `tf.tpu.experimental.initialize_tpu_system(tpu)` between each model training to completely reinitialize the TPU. I've tested it in your notebook and it works.",
      "votes": null
    },
    {
      "id": "794185",
      "postDate": "04/01/2020 16:29:15",
      "content": "<p>TPU memory is garbage collected along with the Python references pointing to memory-consuming objects.\nIf for some reason it does not work as expected, you can totally reinitialze the TPU using:\n<code>tf.tpu.experimental.initialize_tpu_system(tpu)</code></p>",
      "rawMarkdown": "TPU memory is garbage collected along with the Python references pointing to memory-consuming objects.\nIf for some reason it does not work as expected, you can totally reinitialze the TPU using:\n`tf.tpu.experimental.initialize_tpu_system(tpu)`",
      "votes": null
    },
    {
      "id": "794607",
      "postDate": "04/01/2020 23:50:13",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> That's great ! I will test it. </p>",
      "rawMarkdown": "mgornergoogle That's great ! I will test it.",
      "votes": null
    },
    {
      "id": "794690",
      "postDate": "04/02/2020 02:36:14",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"/qinhui1999\">@qinhui1999</a>  thank you for help  , it help a lot</p>",
      "rawMarkdown": "mgornergoogle @qinhui1999  thank you for help  , it help a lot",
      "votes": null
    },
    {
      "id": "798043",
      "postDate": "04/05/2020 05:55:03",
      "content": "<p><code>When you reload the models, you can choose wether you reload them on TPU or CPU by doing the loading inside or outside of a strategy.scope()</code></p>\n\n<p>I'm getting following error ''NoneType' object has no attribute 'merge_call'' when trying to reload model inside <code>strategy.scope()</code> while it works fine on outside.</p>\n\n<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Do you have any idea regarding this.\nThanks in advance.</p>",
      "rawMarkdown": "`When you reload the models, you can choose wether you reload them on TPU or CPU by doing the loading inside or outside of a strategy.scope()`\n\nI'm getting following error ''NoneType' object has no attribute 'merge_call'' when trying to reload model inside `strategy.scope()` while it works fine on outside.\n\n@mgornergoogle Do you have any idea regarding this.\nThanks in advance.",
      "votes": null
    },
    {
      "id": "801670",
      "postDate": "04/08/2020 17:17:26",
      "content": "<p>It's a bit difficult to tell what's going on without seeing the code, but I found a couple of other forum posts as well as a notebook that touch on a similar issue - these should help you troubleshoot the error! Be sure to check out the links in the discussion.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132020\">Post 1</a></li>\n<li><a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131213\">Post 2</a></li>\n<li><a href=\"https://www.kaggle.com/yihdarshieh/custom-training-with-tpu\">Notebook</a></li>\n</ul>",
      "rawMarkdown": "It's a bit difficult to tell what's going on without seeing the code, but I found a couple of other forum posts as well as a notebook that touch on a similar issue - these should help you troubleshoot the error! Be sure to check out the links in the discussion.\n\n- [Post 1](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132020)\n- [Post 2](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131213)\n- [Notebook](https://www.kaggle.com/yihdarshieh/custom-training-with-tpu)",
      "votes": null
    },
    {
      "id": "802037",
      "postDate": "04/09/2020 04:39:31",
      "content": "<p>Can we use gc.collect() inside strategy scope?  Had to re-run kernel because of memory errors</p>",
      "rawMarkdown": "Can we use gc.collect() inside strategy scope?  Had to re-run kernel because of memory errors",
      "votes": null
    },
    {
      "id": "812553",
      "postDate": "04/18/2020 19:59:16",
      "content": "<p>Amazing.!! 🔥 <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> this trick actually works for me. Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> </p>",
      "rawMarkdown": "Amazing.!! 🔥 `tf.tpu.experimental.initialize_tpu_system(tpu)` this trick actually works for me. Thanks @mgornergoogle",
      "votes": null
    },
    {
      "id": "1065382",
      "postDate": "10/31/2020 08:15:35",
      "content": "<p>Where did you include <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> in your code? For me I still have the ResourceExhaustedError</p>",
      "rawMarkdown": "Where did you include `tf.tpu.experimental.initialize_tpu_system(tpu)` in your code? For me I still have the ResourceExhaustedError",
      "votes": null
    },
    {
      "id": "1505615",
      "postDate": "09/07/2021 12:10:41",
      "content": "<p>how can one check TPU memory usage like we see for GPU. It gets difficutl to set the batch size some times.</p>",
      "rawMarkdown": "how can one check TPU memory usage like we see for GPU. It gets difficutl to set the batch size some times.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 749774,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/18/2020 22:29:30",
      "content": "<p>Not sure what you are doing here. Could you explain further what you mean by \"5 fold\" or maybe share some code ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 749877,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "02/19/2020 00:48:06",
          "content": "<p>'5 fold' mean that we use kfold api(from sklearn ) to split the all train data into 5 parts.  Everytime we choose one of them for validation, four of them for train. Do the same thing 5 times to train 5 different models. You can see this <a href=\"https://www.kaggle.com/ragnar123/5-kfold-densenet201\">kernel</a> . If you replace the densenet201 with efficentnet b6, you will get the error infomation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749903,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/19/2020 01:44:21",
          "content": "<p>Got it. Let me check.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750954,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/19/2020 21:41:23",
          "content": "<p>I posted the answer above.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749923,
      "author_name": "gskdhiman",
      "author_url": "",
      "post_date": "02/19/2020 02:20:00",
      "content": "<p>I had the same issue and as a workaround, I changed the Image size from 512 to 331.</p>\n\n<p>but with 512 image size I was also getting ResourceExhaustedError after 2 folds and tf.keras.backend.clear_session() didn't work for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 750953,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/19/2020 21:41:02",
      "content": "<p>TPU memory is garbage collected when the model itself is garbage collected in your Python code.</p>\n\n<p>Here, you are trying to hold all of the models in TPU memory. However, your settings (image size 512x512px, batch_size=16*8=128 are optimized to use up as much TPU memory as is available to shoot for max accuracy and training speed. You will not be able to hold 5 copies of the model on the TPU.</p>\n\n<p>You can try to save the models to disk after training, then start the next fold with a  new model. There is code for saving/reloading a model in the TPU docs sample: <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">Five flowers with Keras and Xception on TPU</a></p>\n\n<p>When you reload the models, you can choose wether you reload them on TPU or CPU by doing the loading inside or outside of a <code>strategy.scope()</code></p>\n\n<p>(tf.keras.backend.clear_session() is not useful here)</p>",
      "votes": null,
      "replies": [
        {
          "id": 751038,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "02/20/2020 00:22:20",
          "content": "<p>Hi Martin,</p>\n\n<p>Thanks for your answer. But I still want to know that if tensorflow has a API to release the TPU memory like pytorch ( torch.cuda.empty_cache())?  If you use pytorch ,you can very easily to delete unused objects and release the gpu memory.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751100,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/20/2020 01:55:15",
          "content": "<p>In Tensorflow, to free TPU memory, release the model object.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751108,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "02/20/2020 02:07:59",
          "content": "<p>Thanks.  I will try it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 792877,
          "author_name": "bestpredict",
          "author_url": "",
          "post_date": "03/31/2020 15:58:43",
          "content": "<p>hi , does it means  that  'del model, gc.collect()  '   can free tpu memory ? <a href=\"/qinhui1999\">@qinhui1999</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 793151,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/31/2020 20:47:23",
          "content": "<p>yes, that's how it is supposed to work</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 793298,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "03/31/2020 23:46:24",
          "content": "<p><a href=\"/bestpredict\">@bestpredict</a>  <a href=\"/mgornergoogle\">@mgornergoogle</a> No, it does not work.  :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 793300,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "03/31/2020 23:51:04",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> You can see my kernel: <a href=\"https://www.kaggle.com/qinhui1999/tpu-enet-b7-incepention-b6\">https://www.kaggle.com/qinhui1999/tpu-enet-b7-incepention-b6</a>.  'del model and gc.collect()'  can not release the TPU memory.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 793326,
          "author_name": "bestpredict",
          "author_url": "",
          "post_date": "04/01/2020 00:27:45",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I agree with   <a href=\"/qinhui1999\">@qinhui1999</a>  sad and  I had  try  ' del model, gc.collect() '   today,  it seems did not work.  sad :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 793528,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/01/2020 05:06:03",
          "content": "<p>Thank you for the reproducible example. I'm submitting this to the TF team.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 794181,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/01/2020 16:27:41",
          "content": "<p>I confirm there seems to be a problem with this sequence of models.</p>\n\n<p>There is an easy workaround:\nYou can call <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> between each model training to completely reinitialize the TPU. I've tested it in your notebook and it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 794607,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "04/01/2020 23:50:13",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> That's great ! I will test it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 794690,
          "author_name": "bestpredict",
          "author_url": "",
          "post_date": "04/02/2020 02:36:14",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"/qinhui1999\">@qinhui1999</a>  thank you for help  , it help a lot</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 794185,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "04/01/2020 16:29:15",
      "content": "<p>TPU memory is garbage collected along with the Python references pointing to memory-consuming objects.\nIf for some reason it does not work as expected, you can totally reinitialze the TPU using:\n<code>tf.tpu.experimental.initialize_tpu_system(tpu)</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 798043,
          "author_name": "gaur128",
          "author_url": "",
          "post_date": "04/05/2020 05:55:03",
          "content": "<p><code>When you reload the models, you can choose wether you reload them on TPU or CPU by doing the loading inside or outside of a strategy.scope()</code></p>\n\n<p>I'm getting following error ''NoneType' object has no attribute 'merge_call'' when trying to reload model inside <code>strategy.scope()</code> while it works fine on outside.</p>\n\n<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Do you have any idea regarding this.\nThanks in advance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 801670,
          "author_name": "jessemostipak",
          "author_url": "",
          "post_date": "04/08/2020 17:17:26",
          "content": "<p>It's a bit difficult to tell what's going on without seeing the code, but I found a couple of other forum posts as well as a notebook that touch on a similar issue - these should help you troubleshoot the error! Be sure to check out the links in the discussion.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132020\">Post 1</a></li>\n<li><a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131213\">Post 2</a></li>\n<li><a href=\"https://www.kaggle.com/yihdarshieh/custom-training-with-tpu\">Notebook</a></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1505615,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/07/2021 12:10:41",
          "content": "<p>how can one check TPU memory usage like we see for GPU. It gets difficutl to set the batch size some times.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 802037,
      "author_name": "parmarsuraj99",
      "author_url": "",
      "post_date": "04/09/2020 04:39:31",
      "content": "<p>Can we use gc.collect() inside strategy scope?  Had to re-run kernel because of memory errors</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 812553,
      "author_name": "rhtsingh",
      "author_url": "",
      "post_date": "04/18/2020 19:59:16",
      "content": "<p>Amazing.!! 🔥 <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> this trick actually works for me. Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1065382,
          "author_name": "nikhiljohnk",
          "author_url": "",
          "post_date": "10/31/2020 08:15:35",
          "content": "<p>Where did you include <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> in your code? For me I still have the ResourceExhaustedError</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "748767": "Hi,\n\nI tried to train 5 fold Effientnet B6 in a kernel. But after two folds only , then I got the  TPU ResourceExhaustedError.  I tried to use 'K.clear_session()' to release the TPU memory, but it did not work.  : (",
    "749774": "Not sure what you are doing here. Could you explain further what you mean by \"5 fold\" or maybe share some code ?",
    "749877": "'5 fold' mean that we use kfold api(from sklearn ) to split the all train data into 5 parts.  Everytime we choose one of them for validation, four of them for train. Do the same thing 5 times to train 5 different models. You can see this [kernel](https://www.kaggle.com/ragnar123/5-kfold-densenet201) . If you replace the densenet201 with efficentnet b6, you will get the error infomation.",
    "749903": "Got it. Let me check.",
    "749923": "I had the same issue and as a workaround, I changed the Image size from 512 to 331.\n\nbut with 512 image size I was also getting ResourceExhaustedError after 2 folds and tf.keras.backend.clear_session() didn't work for me.",
    "750953": "TPU memory is garbage collected when the model itself is garbage collected in your Python code.\n\nHere, you are trying to hold all of the models in TPU memory. However, your settings (image size 512x512px, batch_size=16*8=128 are optimized to use up as much TPU memory as is available to shoot for max accuracy and training speed. You will not be able to hold 5 copies of the model on the TPU.\n\nYou can try to save the models to disk after training, then start the next fold with a  new model. There is code for saving/reloading a model in the TPU docs sample: [Five flowers with Keras and Xception on TPU](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu)\n\nWhen you reload the models, you can choose wether you reload them on TPU or CPU by doing the loading inside or outside of a `strategy.scope()`\n\n(tf.keras.backend.clear_session() is not useful here)",
    "750954": "I posted the answer above.",
    "751038": "Hi Martin,\n\nThanks for your answer. But I still want to know that if tensorflow has a API to release the TPU memory like pytorch ( torch.cuda.empty_cache())?  If you use pytorch ,you can very easily to delete unused objects and release the gpu memory.",
    "751100": "In Tensorflow, to free TPU memory, release the model object.",
    "751108": "Thanks.  I will try it.",
    "792877": "hi , does it means  that  'del model, gc.collect()  '   can free tpu memory ? @qinhui1999",
    "793151": "yes, that's how it is supposed to work",
    "793298": "bestpredict  @mgornergoogle No, it does not work.  :(",
    "793300": "mgornergoogle You can see my kernel: https://www.kaggle.com/qinhui1999/tpu-enet-b7-incepention-b6.  'del model and gc.collect()'  can not release the TPU memory.",
    "793326": "mgornergoogle I agree with   @qinhui1999  sad and  I had  try  ' del model, gc.collect() '   today,  it seems did not work.  sad :(",
    "793528": "Thank you for the reproducible example. I'm submitting this to the TF team.",
    "794181": "I confirm there seems to be a problem with this sequence of models.\n\nThere is an easy workaround:\nYou can call `tf.tpu.experimental.initialize_tpu_system(tpu)` between each model training to completely reinitialize the TPU. I've tested it in your notebook and it works.",
    "794185": "TPU memory is garbage collected along with the Python references pointing to memory-consuming objects.\nIf for some reason it does not work as expected, you can totally reinitialze the TPU using:\n`tf.tpu.experimental.initialize_tpu_system(tpu)`",
    "794607": "mgornergoogle That's great ! I will test it.",
    "794690": "mgornergoogle @qinhui1999  thank you for help  , it help a lot",
    "798043": "`When you reload the models, you can choose wether you reload them on TPU or CPU by doing the loading inside or outside of a strategy.scope()`\n\nI'm getting following error ''NoneType' object has no attribute 'merge_call'' when trying to reload model inside `strategy.scope()` while it works fine on outside.\n\n@mgornergoogle Do you have any idea regarding this.\nThanks in advance.",
    "801670": "It's a bit difficult to tell what's going on without seeing the code, but I found a couple of other forum posts as well as a notebook that touch on a similar issue - these should help you troubleshoot the error! Be sure to check out the links in the discussion.\n\n- [Post 1](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132020)\n- [Post 2](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131213)\n- [Notebook](https://www.kaggle.com/yihdarshieh/custom-training-with-tpu)",
    "802037": "Can we use gc.collect() inside strategy scope?  Had to re-run kernel because of memory errors",
    "812553": "Amazing.!! 🔥 `tf.tpu.experimental.initialize_tpu_system(tpu)` this trick actually works for me. Thanks @mgornergoogle",
    "1065382": "Where did you include `tf.tpu.experimental.initialize_tpu_system(tpu)` in your code? For me I still have the ResourceExhaustedError",
    "1505615": "how can one check TPU memory usage like we see for GPU. It gets difficutl to set the batch size some times."
  },
  "source": "meta"
}