{
  "id": 174762,
  "title": "Google Colab TPU: OOM issues?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174762",
  "author_name": "",
  "post_date": "2020-08-15T07:44:54.508228600Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Dear fellows,</p>\n<p>I experience very frequently OOM issues in Google Colab with specs (like IMG- size: 512, Batch Size: 32) which can easily be handled by Kaggle TPU. Do you experience the same? Any insights, why this happens?</p>",
  "messages": [
    {
      "id": "971124",
      "postDate": "08/15/2020 07:44:54",
      "content": "<p>Dear fellows,</p>\n<p>I experience very frequently OOM issues in Google Colab with specs (like IMG- size: 512, Batch Size: 32) which can easily be handled by Kaggle TPU. Do you experience the same? Any insights, why this happens?</p>",
      "rawMarkdown": "Dear fellows,\n\nI experience very frequently OOM issues in Google Colab with specs (like IMG- size: 512, Batch Size: 32) which can easily be handled by Kaggle TPU. Do you experience the same? Any insights, why this happens?",
      "votes": null
    },
    {
      "id": "971125",
      "postDate": "08/15/2020 07:48:32",
      "content": "<p><a href=\"https://www.kaggle.com/gdonchyts\" target=\"_blank\">@gdonchyts</a> , I read somewhere that you also experience this?</p>",
      "rawMarkdown": "gdonchyts , I read somewhere that you also experience this?",
      "votes": null
    },
    {
      "id": "971139",
      "postDate": "08/15/2020 07:58:08",
      "content": "<p>Kaggle = TPU_v3</p>\n<p>Colab = TPU_v2</p>",
      "rawMarkdown": "Kaggle = TPU_v3\n\nColab = TPU_v2",
      "votes": null
    },
    {
      "id": "971516",
      "postDate": "08/15/2020 16:05:45",
      "content": "<p>Colab TPUs are v2 while Kaggle are v3.<br>\nReduce batch size to about 1/4th and retry - It should work, but slower.</p>",
      "rawMarkdown": "Colab TPUs are v2 while Kaggle are v3.\nReduce batch size to about 1/4th and retry - It should work, but slower.",
      "votes": null
    },
    {
      "id": "971535",
      "postDate": "08/15/2020 16:35:30",
      "content": "<p>halve the batch size when you utilize colab<br>\nBased on my trial and error, it worked well</p>",
      "rawMarkdown": "halve the batch size when you utilize colab\nBased on my trial and error, it worked well",
      "votes": null
    },
    {
      "id": "971589",
      "postDate": "08/15/2020 17:33:30",
      "content": "<p>There were issues when Colab team has upgraded to TF 2.3, the memory consumption grew like at least twice. Google responded very promptly to the issue (<a href=\"https://github.com/googlecolab/colabtools/issues/1470\" target=\"_blank\">https://github.com/googlecolab/colabtools/issues/1470</a>) and posted a workaround allowing to switch TPU to TF 2.2 mode, which works perfectly.</p>\n<p>Another trick to use is mixed precision with bfloat16, this allows TPU to fit larger batches into memory, but it has to be tested on per model basis, usually requires the explicit cast to float32 at the end.</p>",
      "rawMarkdown": "There were issues when Colab team has upgraded to TF 2.3, the memory consumption grew like at least twice. Google responded very promptly to the issue (https://github.com/googlecolab/colabtools/issues/1470) and posted a workaround allowing to switch TPU to TF 2.2 mode, which works perfectly.\n\nAnother trick to use is mixed precision with bfloat16, this allows TPU to fit larger batches into memory, but it has to be tested on per model basis, usually requires the explicit cast to float32 at the end.",
      "votes": null
    },
    {
      "id": "971799",
      "postDate": "08/15/2020 22:36:37",
      "content": "<p>colab use TPU v2 which have 8 gb of memory for each core. TPUv3 on kaggle has 16 gb per core, so it is normal that you can't use the same batch size between kaggle and colab ☹️</p>",
      "rawMarkdown": "colab use TPU v2 which have 8 gb of memory for each core. TPUv3 on kaggle has 16 gb per core, so it is normal that you can't use the same batch size between kaggle and colab ☹️",
      "votes": null
    },
    {
      "id": "971871",
      "postDate": "08/16/2020 02:56:15",
      "content": "<p>Colab's TPU is v2, while kaggle's TPU is v3. They are different in the memory size. Just reduce the batch size. </p>",
      "rawMarkdown": "Colab's TPU is v2, while kaggle's TPU is v3. They are different in the memory size. Just reduce the batch size.",
      "votes": null
    },
    {
      "id": "971876",
      "postDate": "08/16/2020 03:17:27",
      "content": "<p>What version TensorFlow does Colab have? Using TF2.3 can only handle half the batch size that TF2.2 can handle when using EfficientNet TPU.</p>",
      "rawMarkdown": "What version TensorFlow does Colab have? Using TF2.3 can only handle half the batch size that TF2.2 can handle when using EfficientNet TPU.",
      "votes": null
    },
    {
      "id": "971899",
      "postDate": "08/16/2020 04:02:24",
      "content": "<p><a href=\"https://www.kaggle.com/chrisden\" target=\"_blank\">@chrisden</a>, as others have mentioned, the TPU in colab is older. Using the same script I had on Kaggle in colab doubled my training time per epoch until I used the following script. Basically when they upgraded to TF 2.3 something broke in colab so you need to install 2.2.0 and all you need is the following:-</p>\n<pre><code>!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n</code></pre>",
      "rawMarkdown": "chrisden, as others have mentioned, the TPU in colab is older. Using the same script I had on Kaggle in colab doubled my training time per epoch until I used the following script. Basically when they upgraded to TF 2.3 something broke in colab so you need to install 2.2.0 and all you need is the following:-\n\n```\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n```",
      "votes": null
    },
    {
      "id": "971941",
      "postDate": "08/16/2020 05:20:08",
      "content": "<p>Thanks to sharing this with everyone. There are bugs in TF2.3 so yes TF2.2 is better.</p>",
      "rawMarkdown": "Thanks to sharing this with everyone. There are bugs in TF2.3 so yes TF2.2 is better.",
      "votes": null
    },
    {
      "id": "972141",
      "postDate": "08/16/2020 09:09:27",
      "content": "<p>Thanks for clearing this up guys!</p>",
      "rawMarkdown": "Thanks for clearing this up guys!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 971125,
      "author_name": "ChristianDenich",
      "author_url": "",
      "post_date": "08/15/2020 07:48:32",
      "content": "<p><a href=\"https://www.kaggle.com/gdonchyts\" target=\"_blank\">@gdonchyts</a> , I read somewhere that you also experience this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 971589,
          "author_name": "gdonchyts",
          "author_url": "",
          "post_date": "08/15/2020 17:33:30",
          "content": "<p>There were issues when Colab team has upgraded to TF 2.3, the memory consumption grew like at least twice. Google responded very promptly to the issue (<a href=\"https://github.com/googlecolab/colabtools/issues/1470\" target=\"_blank\">https://github.com/googlecolab/colabtools/issues/1470</a>) and posted a workaround allowing to switch TPU to TF 2.2 mode, which works perfectly.</p>\n<p>Another trick to use is mixed precision with bfloat16, this allows TPU to fit larger batches into memory, but it has to be tested on per model basis, usually requires the explicit cast to float32 at the end.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 971139,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "08/15/2020 07:58:08",
      "content": "<p>Kaggle = TPU_v3</p>\n<p>Colab = TPU_v2</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971516,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "08/15/2020 16:05:45",
      "content": "<p>Colab TPUs are v2 while Kaggle are v3.<br>\nReduce batch size to about 1/4th and retry - It should work, but slower.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971535,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "08/15/2020 16:35:30",
      "content": "<p>halve the batch size when you utilize colab<br>\nBased on my trial and error, it worked well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971799,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "08/15/2020 22:36:37",
      "content": "<p>colab use TPU v2 which have 8 gb of memory for each core. TPUv3 on kaggle has 16 gb per core, so it is normal that you can't use the same batch size between kaggle and colab ☹️</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971876,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/16/2020 03:17:27",
      "content": "<p>What version TensorFlow does Colab have? Using TF2.3 can only handle half the batch size that TF2.2 can handle when using EfficientNet TPU.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971899,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "08/16/2020 04:02:24",
      "content": "<p><a href=\"https://www.kaggle.com/chrisden\" target=\"_blank\">@chrisden</a>, as others have mentioned, the TPU in colab is older. Using the same script I had on Kaggle in colab doubled my training time per epoch until I used the following script. Basically when they upgraded to TF 2.3 something broke in colab so you need to install 2.2.0 and all you need is the following:-</p>\n<pre><code>!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 971941,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/16/2020 05:20:08",
          "content": "<p>Thanks to sharing this with everyone. There are bugs in TF2.3 so yes TF2.2 is better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 972141,
      "author_name": "ChristianDenich",
      "author_url": "",
      "post_date": "08/16/2020 09:09:27",
      "content": "<p>Thanks for clearing this up guys!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971871,
      "author_name": "tianyu5",
      "author_url": "",
      "post_date": "08/16/2020 02:56:15",
      "content": "<p>Colab's TPU is v2, while kaggle's TPU is v3. They are different in the memory size. Just reduce the batch size. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "971124": "Dear fellows,\n\nI experience very frequently OOM issues in Google Colab with specs (like IMG- size: 512, Batch Size: 32) which can easily be handled by Kaggle TPU. Do you experience the same? Any insights, why this happens?",
    "971125": "gdonchyts , I read somewhere that you also experience this?",
    "971139": "Kaggle = TPU_v3\n\nColab = TPU_v2",
    "971516": "Colab TPUs are v2 while Kaggle are v3.\nReduce batch size to about 1/4th and retry - It should work, but slower.",
    "971535": "halve the batch size when you utilize colab\nBased on my trial and error, it worked well",
    "971589": "There were issues when Colab team has upgraded to TF 2.3, the memory consumption grew like at least twice. Google responded very promptly to the issue (https://github.com/googlecolab/colabtools/issues/1470) and posted a workaround allowing to switch TPU to TF 2.2 mode, which works perfectly.\n\nAnother trick to use is mixed precision with bfloat16, this allows TPU to fit larger batches into memory, but it has to be tested on per model basis, usually requires the explicit cast to float32 at the end.",
    "971799": "colab use TPU v2 which have 8 gb of memory for each core. TPUv3 on kaggle has 16 gb per core, so it is normal that you can't use the same batch size between kaggle and colab ☹️",
    "971871": "Colab's TPU is v2, while kaggle's TPU is v3. They are different in the memory size. Just reduce the batch size.",
    "971876": "What version TensorFlow does Colab have? Using TF2.3 can only handle half the batch size that TF2.2 can handle when using EfficientNet TPU.",
    "971899": "chrisden, as others have mentioned, the TPU in colab is older. Using the same script I had on Kaggle in colab doubled my training time per epoch until I used the following script. Basically when they upgraded to TF 2.3 something broke in colab so you need to install 2.2.0 and all you need is the following:-\n\n```\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n```",
    "971941": "Thanks to sharing this with everyone. There are bugs in TF2.3 so yes TF2.2 is better.",
    "972141": "Thanks for clearing this up guys!"
  },
  "source": "meta"
}