{
  "id": 146242,
  "title": "K-Fold CV on TPU",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/146242",
  "author_name": "Shahules",
  "post_date": "2020-04-26T11:22:48.101000",
  "votes": 16,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I was trying to do a k fold cross-validation using XLM-Roberta (base/large) on TPU. I was able to train base model for 3 folds and large model for 1 fold after that it gave <code>OOM resource exhausted error</code>. Is there any way to optimize this? I did <code>K.clear_session()</code> after each fold but didn't help.</p>",
  "messages": [
    {
      "id": 821726,
      "postDate": "2020-04-26T11:22:48.103Z",
      "content": "<p>I was trying to do a k fold cross-validation using XLM-Roberta (base/large) on TPU. I was able to train base model for 3 folds and large model for 1 fold after that it gave <code>OOM resource exhausted error</code>. Is there any way to optimize this? I did <code>K.clear_session()</code> after each fold but didn't help.</p>",
      "rawMarkdown": "I was trying to do a k fold cross-validation using XLM-Roberta (base/large) on TPU. I was able to train base model for 3 folds and large model for 1 fold after that it gave `OOM resource exhausted error`. Is there any way to optimize this? I did `K.clear_session()` after each fold but didn't help.",
      "votes": 15
    },
    {
      "id": 821773,
      "postDate": "2020-04-26T12:13:34.480Z",
      "content": "<p>reinitializing the tpu resolver should clean up everything:\n<code>\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=tpu_id)\n        tf.config.experimental_connect_to_cluster(tpu)\n        tf.tpu.experimental.initialize_tpu_system(tpu)\n        strategy = tf.distribute.experimental.TPUStrategy(tpu)\n</code></p>",
      "rawMarkdown": "reinitializing the tpu resolver should clean up everything:\n```\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=tpu_id)\n        tf.config.experimental_connect_to_cluster(tpu)\n        tf.tpu.experimental.initialize_tpu_system(tpu)\n        strategy = tf.distribute.experimental.TPUStrategy(tpu)\n```",
      "votes": 7,
      "replies": [
        {
          "id": 821953,
          "postDate": "2020-04-26T14:25:13.390Z",
          "content": "<p>Thanks, but does this mean that we will have to re initialize everything associated such as batch sizze and connecting to dataset ? <a href=\"/hmendonca\">@hmendonca</a> </p>",
          "rawMarkdown": "Thanks, but does this mean that we will have to re initialize everything associated such as batch sizze and connecting to dataset ? @hmendonca "
        },
        {
          "id": 821970,
          "postDate": "2020-04-26T14:47:06.927Z",
          "content": "<p>Just the tpu resolver <a href=\"/shahules\">@shahules</a> \nif you're re-connecting to the same tpu, your batch size/tpu cores won't change.\nAnd you shouldn't need to restart the dataset either.</p>",
          "rawMarkdown": "Just the tpu resolver @shahules \nif you're re-connecting to the same tpu, your batch size/tpu cores won't change.\nAnd you shouldn't need to restart the dataset either.",
          "votes": 4
        },
        {
          "id": 822069,
          "postDate": "2020-04-26T16:30:05.633Z",
          "content": "<p>Thanks, buddy <a href=\"/hmendonca\">@hmendonca</a>  </p>",
          "rawMarkdown": "Thanks, buddy @hmendonca  "
        },
        {
          "id": 822170,
          "postDate": "2020-04-26T18:13:10.630Z",
          "content": "<p><a href=\"/hmendonca\">@hmendonca</a> can you tell me how to get the <code>tpu_id</code> ?Sorry,this is my first experience with TPUs</p>",
          "rawMarkdown": "@hmendonca can you tell me how to get the `tpu_id` ?Sorry,this is my first experience with TPUs"
        },
        {
          "id": 822902,
          "postDate": "2020-04-27T08:42:30.357Z",
          "content": "<p>it's the name of the tpu instance you've created on gcp (the first field in the form, if you're using the web interface)</p>",
          "rawMarkdown": "it's the name of the tpu instance you've created on gcp (the first field in the form, if you're using the web interface)",
          "votes": 4
        },
        {
          "id": 823729,
          "postDate": "2020-04-27T20:58:00.150Z",
          "content": "<p>Just this is brutal enough :-)</p>\n\n<p><code>\ntf.tpu.experimental.initialize_tpu_system(tpu)\n</code></p>",
          "rawMarkdown": "Just this is brutal enough :-)\n\n```\ntf.tpu.experimental.initialize_tpu_system(tpu)\n```",
          "votes": 4
        },
        {
          "id": 823974,
          "postDate": "2020-04-28T04:24:45.250Z",
          "content": "<p>thanks</p>",
          "rawMarkdown": "thanks",
          "votes": -1
        },
        {
          "id": 838076,
          "postDate": "2020-05-08T09:18:36.967Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a>, I am getting loss and val_loss as nan from 2nd epoch of the 1st fold when I tried run Kfold on TPU with the above line before calling the model.</p>",
          "rawMarkdown": "@mgornergoogle, I am getting loss and val_loss as nan from 2nd epoch of the 1st fold when I tried run Kfold on TPU with the above line before calling the model."
        },
        {
          "id": 855145,
          "postDate": "2020-05-20T16:25:24.173Z",
          "content": "<p>Does it also clear the RAM? It seems the ram is not release with TF 2.1 between each call of fit</p>",
          "rawMarkdown": "Does it also clear the RAM? It seems the ram is not release with TF 2.1 between each call of fit"
        }
      ]
    },
    {
      "id": 864705,
      "postDate": "2020-05-28T06:31:16.860Z",
      "content": "<p>Hi, I'm a little bit late to this question, but I also ran into resource exhaustion issues with keras. For me the biggest problem was that multiple fit/predict calls were eating RAM, and a 5fold CV sometimes died when it ran out of RAM.</p>\n\n<p>For me using a custom training and prediction loop completely solved the issue, and I can easily run a 10fold CV.\nI shared the code for the custom training loop here: <a href=\"https://www.kaggle.com/riblidezso/tpu-custom-tensoflow2-training-loop\">https://www.kaggle.com/riblidezso/tpu-custom-tensoflow2-training-loop</a></p>",
      "rawMarkdown": "Hi, I'm a little bit late to this question, but I also ran into resource exhaustion issues with keras. For me the biggest problem was that multiple fit/predict calls were eating RAM, and a 5fold CV sometimes died when it ran out of RAM.\n\nFor me using a custom training and prediction loop completely solved the issue, and I can easily run a 10fold CV.\nI shared the code for the custom training loop here: https://www.kaggle.com/riblidezso/tpu-custom-tensoflow2-training-loop",
      "votes": 2,
      "replies": [
        {
          "id": 889808,
          "postDate": "2020-06-17T06:51:51.040Z",
          "content": "<p><code>NotFoundError: {{function_node __inference_distributed_train_step_127624}} Resource worker/Adam/iter/replica_1_52180/N10tensorflow3VarE does not exist.\n     [[{{node cluster_distributed_train_step/_execute_1_0}}]]</code>\nhey <a href=\"/riblidezso\">@riblidezso</a> ,I am getting this error while trying to do Kfold with abv kernel.Can you tell me the reason for this?</p>",
          "rawMarkdown": "`NotFoundError: {{function_node __inference_distributed_train_step_127624}} Resource worker/Adam/iter/replica_1_52180/N10tensorflow3VarE does not exist.\n\t [[{{node cluster_distributed_train_step/_execute_1_0}}]]`\nhey @riblidezso ,I am getting this error while trying to do Kfold with abv kernel.Can you tell me the reason for this?"
        },
        {
          "id": 896593,
          "postDate": "2020-06-22T09:49:23.650Z",
          "content": "<p>Yes, you should not reinitialize anything if you use my code, just redefine the training/validation datasets and reload the model weights.</p>\n\n<p>See this discussion: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/158202\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/158202</a></p>",
          "rawMarkdown": "Yes, you should not reinitialize anything if you use my code, just redefine the training/validation datasets and reload the model weights.\n\nSee this discussion: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/158202"
        }
      ]
    },
    {
      "id": 853139,
      "postDate": "2020-05-19T01:53:48.340Z",
      "content": "<p>I list some tips for TPU 5 fold <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/143869#810012\">here</a>. The most important is to execute. </p>\n\n<pre><code>tf.tpu.experimental.initialize_tpu_system(tpu)\n</code></pre>\n\n<p>At the beginning of each new 5 fold loop iteration to clear memory</p>",
      "rawMarkdown": "I list some tips for TPU 5 fold [here][1]. The most important is to execute. \n    \n    tf.tpu.experimental.initialize_tpu_system(tpu)\n\nAt the beginning of each new 5 fold loop iteration to clear memory\n\n[1]: https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/143869#810012",
      "votes": 2,
      "replies": [
        {
          "id": 853894,
          "postDate": "2020-05-19T15:24:35.090Z",
          "content": "<p>Thanks for sharing. I want to highlight about <code>drop_remainder=True</code> during training. Without this, it can cause problems like NaN loss because of the partial batch in some folds. </p>",
          "rawMarkdown": "Thanks for sharing. I want to highlight about `drop_remainder=True ` during training. Without this, it can cause problems like NaN loss because of the partial batch in some folds. ",
          "votes": 2
        },
        {
          "id": 854000,
          "postDate": "2020-05-19T16:57:55.300Z",
          "content": "<p>Yes, nice reminder. I also noticed that training is faster if you <code>drop_remainder=True</code> however for inference you must use <code>drop_remainder=False</code>.</p>",
          "rawMarkdown": "Yes, nice reminder. I also noticed that training is faster if you `drop_remainder=True` however for inference you must use `drop_remainder=False`.",
          "votes": 1
        },
        {
          "id": 854174,
          "postDate": "2020-05-19T20:16:42.807Z",
          "content": "<p>Yes last partial batches are supported on TPU (<code>drop_remainder=False</code>), but processing them will require a little bit if extra time. That's the performance hit your are seeing.</p>\n\n<p>And in some last partial batch cases, for example a batch size of 1, but there are others, certain mathematical operations like batch norm are not defined and output NaN. This however should never happen during training if the dataset is repeated across all epochs because in that case all batches are of the same nominal size.</p>",
          "rawMarkdown": "Yes last partial batches are supported on TPU (`drop_remainder=False`), but processing them will require a little bit if extra time. That's the performance hit your are seeing.\n\nAnd in some last partial batch cases, for example a batch size of 1, but there are others, certain mathematical operations like batch norm are not defined and output NaN. This however should never happen during training if the dataset is repeated across all epochs because in that case all batches are of the same nominal size.",
          "votes": 2
        },
        {
          "id": 854185,
          "postDate": "2020-05-19T20:34:06.933Z",
          "content": "<p>Thanks for explaining, I faced this issue  when I used StratifiedKfold <a href=\"https://github.com/tensorflow/tpu/issues/766\">issue</a></p>",
          "rawMarkdown": "Thanks for explaining, I faced this issue  when I used StratifiedKfold [issue](https://github.com/tensorflow/tpu/issues/766)"
        }
      ]
    },
    {
      "id": 821807,
      "postDate": "2020-04-26T12:41:38.947Z",
      "content": "<p>It's kinda dump way, but I just save the initial model weights and load it in each fold without model build.</p>",
      "rawMarkdown": "It's kinda dump way, but I just save the initial model weights and load it in each fold without model build.",
      "votes": 1,
      "replies": [
        {
          "id": 821947,
          "postDate": "2020-04-26T14:21:56.990Z",
          "content": "<p>haha, hope someone will help us get better solution :)</p>",
          "rawMarkdown": "haha, hope someone will help us get better solution :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 822187,
      "postDate": "2020-04-26T18:23:54.207Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 822765,
          "postDate": "2020-04-27T05:46:02.197Z",
          "content": "<p>yes,kaggle need to update their TPU env to support TF2.2 functionalities.I hope they will do this before the competition ends.</p>",
          "rawMarkdown": "yes,kaggle need to update their TPU env to support TF2.2 functionalities.I hope they will do this before the competition ends."
        },
        {
          "id": 844874,
          "postDate": "2020-05-12T22:42:09.987Z",
          "content": "<p><a href=\"/kurianbenoy\">@kurianbenoy</a> I don't remember this conversation. What exactly did I say about TF 2.2 and concurrent k-fold cross-validation in Tensorflow?</p>",
          "rawMarkdown": "@kurianbenoy I don't remember this conversation. What exactly did I say about TF 2.2 and concurrent k-fold cross-validation in Tensorflow?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 821773,
      "author_name": "Henrique Mendonça",
      "author_url": "",
      "post_date": "2020-04-26T12:13:34.480000",
      "content": "<p>reinitializing the tpu resolver should clean up everything:\n<code>\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=tpu_id)\n        tf.config.experimental_connect_to_cluster(tpu)\n        tf.tpu.experimental.initialize_tpu_system(tpu)\n        strategy = tf.distribute.experimental.TPUStrategy(tpu)\n</code></p>",
      "votes": 7,
      "replies": [
        {
          "id": 821953,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-04-26T14:25:13.390000",
          "content": "<p>Thanks, but does this mean that we will have to re initialize everything associated such as batch sizze and connecting to dataset ? <a href=\"/hmendonca\">@hmendonca</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 821970,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-04-26T14:47:06.927000",
          "content": "<p>Just the tpu resolver <a href=\"/shahules\">@shahules</a> \nif you're re-connecting to the same tpu, your batch size/tpu cores won't change.\nAnd you shouldn't need to restart the dataset either.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 822069,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-04-26T16:30:05.633000",
          "content": "<p>Thanks, buddy <a href=\"/hmendonca\">@hmendonca</a>  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 822170,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-04-26T18:13:10.630000",
          "content": "<p><a href=\"/hmendonca\">@hmendonca</a> can you tell me how to get the <code>tpu_id</code> ?Sorry,this is my first experience with TPUs</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 822902,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-04-27T08:42:30.357000",
          "content": "<p>it's the name of the tpu instance you've created on gcp (the first field in the form, if you're using the web interface)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 823729,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-27T20:58:00.150000",
          "content": "<p>Just this is brutal enough :-)</p>\n\n<p><code>\ntf.tpu.experimental.initialize_tpu_system(tpu)\n</code></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 823974,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-04-28T04:24:45.250000",
          "content": "<p>thanks</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 838076,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2020-05-08T09:18:36.967000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a>, I am getting loss and val_loss as nan from 2nd epoch of the 1st fold when I tried run Kfold on TPU with the above line before calling the model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 855145,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2020-05-20T16:25:24.173000",
          "content": "<p>Does it also clear the RAM? It seems the ram is not release with TF 2.1 between each call of fit</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 864705,
      "author_name": "Dezső Ribli",
      "author_url": "",
      "post_date": "2020-05-28T06:31:16.860000",
      "content": "<p>Hi, I'm a little bit late to this question, but I also ran into resource exhaustion issues with keras. For me the biggest problem was that multiple fit/predict calls were eating RAM, and a 5fold CV sometimes died when it ran out of RAM.</p>\n\n<p>For me using a custom training and prediction loop completely solved the issue, and I can easily run a 10fold CV.\nI shared the code for the custom training loop here: <a href=\"https://www.kaggle.com/riblidezso/tpu-custom-tensoflow2-training-loop\">https://www.kaggle.com/riblidezso/tpu-custom-tensoflow2-training-loop</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 889808,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-06-17T06:51:51.040000",
          "content": "<p><code>NotFoundError: {{function_node __inference_distributed_train_step_127624}} Resource worker/Adam/iter/replica_1_52180/N10tensorflow3VarE does not exist.\n     [[{{node cluster_distributed_train_step/_execute_1_0}}]]</code>\nhey <a href=\"/riblidezso\">@riblidezso</a> ,I am getting this error while trying to do Kfold with abv kernel.Can you tell me the reason for this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896593,
          "author_name": "Dezső Ribli",
          "author_url": "",
          "post_date": "2020-06-22T09:49:23.650000",
          "content": "<p>Yes, you should not reinitialize anything if you use my code, just redefine the training/validation datasets and reload the model weights.</p>\n\n<p>See this discussion: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/158202\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/158202</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 853139,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-05-19T01:53:48.340000",
      "content": "<p>I list some tips for TPU 5 fold <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/143869#810012\">here</a>. The most important is to execute. </p>\n\n<pre><code>tf.tpu.experimental.initialize_tpu_system(tpu)\n</code></pre>\n\n<p>At the beginning of each new 5 fold loop iteration to clear memory</p>",
      "votes": 2,
      "replies": [
        {
          "id": 853894,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2020-05-19T15:24:35.090000",
          "content": "<p>Thanks for sharing. I want to highlight about <code>drop_remainder=True</code> during training. Without this, it can cause problems like NaN loss because of the partial batch in some folds. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 854000,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-05-19T16:57:55.300000",
          "content": "<p>Yes, nice reminder. I also noticed that training is faster if you <code>drop_remainder=True</code> however for inference you must use <code>drop_remainder=False</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 854174,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-05-19T20:16:42.807000",
          "content": "<p>Yes last partial batches are supported on TPU (<code>drop_remainder=False</code>), but processing them will require a little bit if extra time. That's the performance hit your are seeing.</p>\n\n<p>And in some last partial batch cases, for example a batch size of 1, but there are others, certain mathematical operations like batch norm are not defined and output NaN. This however should never happen during training if the dataset is repeated across all epochs because in that case all batches are of the same nominal size.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 854185,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2020-05-19T20:34:06.933000",
          "content": "<p>Thanks for explaining, I faced this issue  when I used StratifiedKfold <a href=\"https://github.com/tensorflow/tpu/issues/766\">issue</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 821807,
      "author_name": "soomiles",
      "author_url": "",
      "post_date": "2020-04-26T12:41:38.947000",
      "content": "<p>It's kinda dump way, but I just save the initial model weights and load it in each fold without model build.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 821947,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-04-26T14:21:56.990000",
          "content": "<p>haha, hope someone will help us get better solution :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 822187,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T18:23:54.207000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 822765,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-04-27T05:46:02.197000",
          "content": "<p>yes,kaggle need to update their TPU env to support TF2.2 functionalities.I hope they will do this before the competition ends.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 844874,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-05-12T22:42:09.987000",
          "content": "<p><a href=\"/kurianbenoy\">@kurianbenoy</a> I don't remember this conversation. What exactly did I say about TF 2.2 and concurrent k-fold cross-validation in Tensorflow?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "821726": "I was trying to do a k fold cross-validation using XLM-Roberta (base/large) on TPU. I was able to train base model for 3 folds and large model for 1 fold after that it gave `OOM resource exhausted error`. Is there any way to optimize this? I did `K.clear_session()` after each fold but didn't help.",
    "821773": "reinitializing the tpu resolver should clean up everything:\n```\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=tpu_id)\n        tf.config.experimental_connect_to_cluster(tpu)\n        tf.tpu.experimental.initialize_tpu_system(tpu)\n        strategy = tf.distribute.experimental.TPUStrategy(tpu)\n```",
    "864705": "Hi, I'm a little bit late to this question, but I also ran into resource exhaustion issues with keras. For me the biggest problem was that multiple fit/predict calls were eating RAM, and a 5fold CV sometimes died when it ran out of RAM.\n\nFor me using a custom training and prediction loop completely solved the issue, and I can easily run a 10fold CV.\nI shared the code for the custom training loop here: https://www.kaggle.com/riblidezso/tpu-custom-tensoflow2-training-loop",
    "853139": "I list some tips for TPU 5 fold [here][1]. The most important is to execute. \n    \n    tf.tpu.experimental.initialize_tpu_system(tpu)\n\nAt the beginning of each new 5 fold loop iteration to clear memory\n\n[1]: https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/143869#810012",
    "821807": "It's kinda dump way, but I just save the initial model weights and load it in each fold without model build.",
    "822187": ""
  }
}