{
  "id": 174299,
  "title": "ResourceExhaustedError when using TPU",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174299",
  "author_name": "kwang",
  "post_date": "2020-08-13T04:10:31.848000",
  "votes": 15,
  "comment_count": 45,
  "views": 0,
  "content": "<p>since kaggle has updated tensorflow version to 2.3, the notebook cann't run in the new version.</p>\n<p>it would cause  ResourceExhaustedError , something like <a href=\"https://github.com/googlecolab/colabtools/issues/1470\" target=\"_blank\">this</a>.</p>\n<p>And downgrade tf to 2.2 didn't work. </p>",
  "messages": [
    {
      "id": 968488,
      "postDate": "2020-08-13T04:10:31.850Z",
      "content": "<p>since kaggle has updated tensorflow version to 2.3, the notebook cann't run in the new version.</p>\n<p>it would cause  ResourceExhaustedError , something like <a href=\"https://github.com/googlecolab/colabtools/issues/1470\" target=\"_blank\">this</a>.</p>\n<p>And downgrade tf to 2.2 didn't work. </p>",
      "rawMarkdown": "since kaggle has updated tensorflow version to 2.3, the notebook cann't run in the new version.\n\nit would cause  ResourceExhaustedError , something like [this](https://github.com/googlecolab/colabtools/issues/1470).\n\nAnd downgrade tf to 2.2 didn't work. \n",
      "votes": 15
    },
    {
      "id": 968664,
      "postDate": "2020-08-13T06:58:33.253Z",
      "content": "<p><strong>USE COLAB! Colab is working well</strong></p>\n<p><strong>just add these lines to colab before importing tensorflow module (downgrade to TF 2.2)</strong></p>\n<p>!pip install tensorflow~=2.2.0 tensorflow<em>gcs</em>config~=2.2.0<br>\nimport tensorflow as tf<br>\nimport requests<br>\nimport os<br>\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB<em>TPU</em>ADDR\"].split(\":\")[0], tf.<strong>version</strong>))<br>\nif resp.status_code != 200:<br>\n    print(\"Failed to switch the TPU to TF {}\".format(version))</p>\n<p><strong>To access data from kaggle to COLAB, you need to manually input GCS PATHs like this on colab</strong></p>\n<p>gcs_path = 'gs://kds-cfb16ed5f55adb5ab35a75a2fe74c3dc11a4b869dc8cfb3d9b1759e1';</p>\n<p>gcs_path2 = 'gs://kds-79d9c56dcae7978f98a6bbe8318ac6866ba4778fc95c5871fcfd0f03';</p>\n<p>gcs_path3 = 'gs://kds-032edf427ad6e338f5392f73759ff404a1a3f25c2abdc4995136de30';</p>\n<p>GCS<em>PATH = [None]*FOLDS; GCS</em>PATH2 = [None]*FOLDS; GCS_PATH3 = [None]*FOLDS</p>\n<p>for i,k in enumerate(IMG_SIZES):</p>\n<pre><code>GCS_PATH[i] = gcs_path;\n\n\nGCS_PATH2[i] = gcs_path2;\n\n\nGCS_PATH3[i] = gcs_path3;\n</code></pre>\n<p>files<em>train = np.sort(np.array(tf.io.gfile.glob(GCS</em>PATH[0] + '/train*.tfrec')));</p>\n<p>files<em>test  = np.sort(np.array(tf.io.gfile.glob(GCS</em>PATH[0] + '/test*.tfrec')));  </p>\n<p><strong>You can get directly GCS PATHs FROM KAGGLE ENVIRONMENT(In kaggle to get GCS PATHs)</strong></p>\n<p>GCS<em>PATH = [None]*FOLDS; GCS</em>PATH2 = [None]*FOLDS; GCS_PATH3 = [None]*FOLDS</p>\n<p>for i,k in enumerate(IMG_SIZES[:FOLDS]):</p>\n<pre><code>GCS_PATH[i] = KaggleDatasets().get_gcs_path('melanoma-%ix%i'%(k,k))\n\n\nGCS_PATH2[i] = KaggleDatasets().get_gcs_path('isic2019-%ix%i'%(k,k))\n\n\nGCS_PATH3[i] = KaggleDatasets().get_gcs_path('malignant-v2-%ix%i'%(k,k))\n</code></pre>\n<p>files<em>train = np.sort(np.array(tf.io.gfile.glob(GCS</em>PATH[0] + '/train*.tfrec')))</p>\n<p>files<em>test  = np.sort(np.array(tf.io.gfile.glob(GCS</em>PATH[0] + '/test*.tfrec')))</p>\n<p>print(GCS<em>PATH, GCS</em>PATH2, GCS_PATH3)</p>",
      "rawMarkdown": "**USE COLAB! Colab is working well**\n\n**just add these lines to colab before importing tensorflow module (downgrade to TF 2.2)**\n\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n    print(\"Failed to switch the TPU to TF {}\".format(version))\n\n\n**To access data from kaggle to COLAB, you need to manually input GCS PATHs like this on colab**\n\n\ngcs_path = 'gs://kds-cfb16ed5f55adb5ab35a75a2fe74c3dc11a4b869dc8cfb3d9b1759e1';\n\n\ngcs_path2 = 'gs://kds-79d9c56dcae7978f98a6bbe8318ac6866ba4778fc95c5871fcfd0f03';\n\n\ngcs_path3 = 'gs://kds-032edf427ad6e338f5392f73759ff404a1a3f25c2abdc4995136de30';\n\n\n\n\nGCS_PATH = [None]\\*FOLDS; GCS_PATH2 = [None]\\*FOLDS; GCS_PATH3 = [None]\\*FOLDS\n\n\nfor i,k in enumerate(IMG_SIZES):\n\n\n    GCS_PATH[i] = gcs_path;\n\n\n    GCS_PATH2[i] = gcs_path2;\n\n\n    GCS_PATH3[i] = gcs_path3;\n\n\nfiles_train = np.sort(np.array(tf.io.gfile.glob(GCS_PATH[0] + '/train\\*.tfrec')));\n\n  \nfiles_test  = np.sort(np.array(tf.io.gfile.glob(GCS_PATH[0] + '/test\\*.tfrec')));  \n\n**You can get directly GCS PATHs FROM KAGGLE ENVIRONMENT(In kaggle to get GCS PATHs)**\n\nGCS_PATH = [None]\\*FOLDS; GCS_PATH2 = [None]\\*FOLDS; GCS_PATH3 = [None]\\*FOLDS\n\n\nfor i,k in enumerate(IMG_SIZES[:FOLDS]):\n\n\n    GCS_PATH[i] = KaggleDatasets().get_gcs_path('melanoma-%ix%i'%(k,k))\n\n\n    GCS_PATH2[i] = KaggleDatasets().get_gcs_path('isic2019-%ix%i'%(k,k))\n\n\n    GCS_PATH3[i] = KaggleDatasets().get_gcs_path('malignant-v2-%ix%i'%(k,k))\n\n\nfiles_train = np.sort(np.array(tf.io.gfile.glob(GCS_PATH[0] + '/train\\*.tfrec')))\n\n\nfiles_test  = np.sort(np.array(tf.io.gfile.glob(GCS_PATH[0] + '/test\\*.tfrec')))\n\nprint(GCS_PATH, GCS_PATH2, GCS_PATH3)",
      "votes": 7,
      "replies": [
        {
          "id": 968910,
          "postDate": "2020-08-13T10:34:08.100Z",
          "content": "<p>I tried your method and it seems to be working, I can train on Colab now. However, I still get OOM when using the same batch size as yesterday, which worked well on Kaggle then. <br>\nDo you know if the initialisation of the TPU should be any different on Colab than on Kaggle?<br>\nI use the initialisation from Chris Deotte's popular notebook now. </p>",
          "rawMarkdown": "I tried your method and it seems to be working, I can train on Colab now. However, I still get OOM when using the same batch size as yesterday, which worked well on Kaggle then. \nDo you know if the initialisation of the TPU should be any different on Colab than on Kaggle?\nI use the initialisation from Chris Deotte's popular notebook now. ",
          "votes": 1
        },
        {
          "id": 968923,
          "postDate": "2020-08-13T10:50:43.880Z",
          "content": "<p>halve the batch size! <br>\ncolab tpu memory is much smaller than kaggle's<br>\nif you set a batch size at 32 on kaggle machine,  you should set it to 16 on colab( this is based on my trial and error)</p>",
          "rawMarkdown": "halve the batch size! \ncolab tpu memory is much smaller than kaggle's\nif you set a batch size at 32 on kaggle machine,  you should set it to 16 on colab( this is based on my trial and error)",
          "votes": 2
        },
        {
          "id": 968940,
          "postDate": "2020-08-13T11:07:17.007Z",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> were you able to load/access private kaggle datasets in colab? <br>\nI saved all of my tpu resources for training at last and this happened.</p>",
          "rawMarkdown": "@deepkim were you able to load/access private kaggle datasets in colab? \nI saved all of my tpu resources for training at last and this happened.",
          "votes": 1
        },
        {
          "id": 968975,
          "postDate": "2020-08-13T11:31:05.450Z",
          "content": "<p><a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a>  i haven tried accessing private data yet. I have been just using public data. </p>",
          "rawMarkdown": "@ankitsajwan  i haven tried accessing private data yet. I have been just using public data. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 969021,
      "postDate": "2020-08-13T12:09:52.600Z",
      "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> Do you have any idea why TPU batch sizes that worked before August 12th don't work after August 12th? Before August 12th we could train 384x384 batch size 256 = 32 x 8 on EffNetB6 <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a>. Now we cannot.</p>",
      "rawMarkdown": "@mgornergoogle Do you have any idea why TPU batch sizes that worked before August 12th don't work after August 12th? Before August 12th we could train 384x384 batch size 256 = 32 x 8 on EffNetB6 [here][1]. Now we cannot.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": 5,
      "replies": [
        {
          "id": 969436,
          "postDate": "2020-08-13T17:18:22.773Z",
          "content": "<p>Yes, I saw this issue with TF 2.3 as well. Investigating. Did you encounter this with other models besides EfficientNet ?<br>\n(The next thing I am going to try is to load EfficientNet from TFHub directly - that works now in TF 2.3 with the /job:localhost LoadOption - instead of the pip-installed lib. Maybe there was a bug fix in there…)</p>",
          "rawMarkdown": "Yes, I saw this issue with TF 2.3 as well. Investigating. Did you encounter this with other models besides EfficientNet ?\n(The next thing I am going to try is to load EfficientNet from TFHub directly - that works now in TF 2.3 with the /job:localhost LoadOption - instead of the pip-installed lib. Maybe there was a bug fix in there...)"
        },
        {
          "id": 969605,
          "postDate": "2020-08-13T19:58:57.210Z",
          "content": "<p>I tried tf.keras.applications.EfficinetNetBx and it did not work any better under TF 2.3. I have not been able to load EfficientNet from TF Hub yes because the tensorflow-hub library has not yet been updated with the new hub.KerasLayer(load_options) parameter. The TF team is investigating…</p>",
          "rawMarkdown": "I tried tf.keras.applications.EfficinetNetBx and it did not work any better under TF 2.3. I have not been able to load EfficientNet from TF Hub yes because the tensorflow-hub library has not yet been updated with the new hub.KerasLayer(load_options) parameter. The TF team is investigating...",
          "replies": [
            {
              "id": 969742,
              "postDate": "2020-08-13T22:36:10.527Z",
              "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> AFAIK, about <code>tf.keras.applicaiton.EfficientNetBx</code> - This API is new and only available in <code>tf-nightly</code>. </p>",
              "rawMarkdown": "@mgornergoogle AFAIK, about `tf.keras.applicaiton.EfficientNetBx` - This API is new and only available in `tf-nightly`. "
            },
            {
              "id": 970768,
              "postDate": "2020-08-14T18:55:01.490Z",
              "content": "<p>no tf.keras.applicaiton.EfficientNetBx is there in TF 2.3</p>",
              "rawMarkdown": "no tf.keras.applicaiton.EfficientNetBx is there in TF 2.3"
            }
          ]
        },
        {
          "id": 969608,
          "postDate": "2020-08-13T19:59:35.660Z",
          "content": "<p>Everyone, have you seen regressions with other models than EfficientNetBx ?</p>",
          "rawMarkdown": "Everyone, have you seen regressions with other models than EfficientNetBx ?"
        },
        {
          "id": 969668,
          "postDate": "2020-08-13T20:51:39.003Z",
          "content": "<p>I have tried all models on TF 2.2 but I only tried EfficientNet on TF 2.3 so far.</p>",
          "rawMarkdown": "I have tried all models on TF 2.2 but I only tried EfficientNet on TF 2.3 so far."
        }
      ]
    },
    {
      "id": 969299,
      "postDate": "2020-08-13T15:47:36.867Z",
      "content": "<p>The code snippet to downgrade to TF 2.2 for now should be as follows:</p>\n<pre><code>!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"TPU_NAME\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n</code></pre>\n<p>Let us know if this unblocks you, and we will be looking at the issue in the meantime. </p>",
      "rawMarkdown": "The code snippet to downgrade to TF 2.2 for now should be as follows:\n```\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"TPU_NAME\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n```\nLet us know if this unblocks you, and we will be looking at the issue in the meantime. ",
      "votes": 4,
      "replies": [
        {
          "id": 969327,
          "postDate": "2020-08-13T16:03:19.987Z",
          "content": "<p>Hi llya, Thanks for sharing this code snippets. I tried it but still couldn't successful to setup it and i found the following error:</p>\n<p>ConnectionError: HTTPConnectionPool(host='grpc', port=8475): Max retries exceeded with url: /requestversion/2.2.0 (Caused by NewConnectionError(': Failed to establish a new connection: [Errno -2] Name or service not known'))</p>",
          "rawMarkdown": "Hi llya, Thanks for sharing this code snippets. I tried it but still couldn't successful to setup it and i found the following error:\n\nConnectionError: HTTPConnectionPool(host='grpc', port=8475): Max retries exceeded with url: /requestversion/2.2.0 (Caused by NewConnectionError('<urllib3.connection.HTTPConnection object at 0x7f22f6a51610>: Failed to establish a new connection: [Errno -2] Name or service not known'))",
          "votes": 3
        },
        {
          "id": 969345,
          "postDate": "2020-08-13T16:10:20.657Z",
          "content": "<p><a href=\"https://www.kaggle.com/osamir\" target=\"_blank\">@osamir</a> this looks like an intermittent error. I assume your Internet was ON and that was a TPU enabled notebook, right? If so, please retry. I did work at least in a couple of notebooks I tried.</p>",
          "rawMarkdown": "@osamir this looks like an intermittent error. I assume your Internet was ON and that was a TPU enabled notebook, right? If so, please retry. I did work at least in a couple of notebooks I tried.",
          "votes": 1
        },
        {
          "id": 969352,
          "postDate": "2020-08-13T16:18:50.900Z",
          "content": "<p>I double checked that TPU is enabled and internet is on but still facing this error. I will try it again and see if it still exists or not.</p>",
          "rawMarkdown": "I double checked that TPU is enabled and internet is on but still facing this error. I will try it again and see if it still exists or not.",
          "votes": 1
        },
        {
          "id": 969381,
          "postDate": "2020-08-13T16:39:33.330Z",
          "content": "<p>We are planning to downgrade back to TF 2.2 for the TPU machines. So, hopefully, you will not need the workaround soon.</p>",
          "rawMarkdown": "We are planning to downgrade back to TF 2.2 for the TPU machines. So, hopefully, you will not need the workaround soon.",
          "votes": 2,
          "replies": [
            {
              "id": 969387,
              "postDate": "2020-08-13T16:42:12.087Z",
              "content": "<p>Upgrade to TF brings a mess many times. IMHO, authority should double check before launch, one day waste at the end time; very lossy. </p>",
              "rawMarkdown": "Upgrade to TF brings a mess many times. IMHO, authority should double check before launch, one day waste at the end time; very lossy. ",
              "votes": 1
            },
            {
              "id": 969405,
              "postDate": "2020-08-13T16:53:04.030Z",
              "content": "<p>actually the workaround doesn't work for me as it gives the same error as reported, and I've tried many times</p>",
              "rawMarkdown": "actually the workaround doesn't work for me as it gives the same error as reported, and I've tried many times",
              "votes": 1
            }
          ]
        },
        {
          "id": 969481,
          "postDate": "2020-08-13T17:56:41.110Z",
          "content": "<p>Workaround does not work for me. Using workaround gives new weird errors.</p>",
          "rawMarkdown": "Workaround does not work for me. Using workaround gives new weird errors."
        },
        {
          "id": 969696,
          "postDate": "2020-08-13T21:27:12.607Z",
          "content": "<p>Same for me, in Kaggle and on Colab</p>",
          "rawMarkdown": "Same for me, in Kaggle and on Colab"
        }
      ]
    },
    {
      "id": 968490,
      "postDate": "2020-08-13T04:12:23.387Z",
      "content": "<p>I agree. I cannot use TPU tonight. Everything that worked yesterday does not work today</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F1a0fcdf9beca61fe3fc18c03ec13cbc5%2Fe1.png?generation=1597291913411797&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F25b4e74365ae124d7e6689411cdc9f7f%2Fe2.png?generation=1597291925498803&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F7f6528504e281e6355766caf4ab6a086%2Fe3.png?generation=1597291935745289&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I agree. I cannot use TPU tonight. Everything that worked yesterday does not work today\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F1a0fcdf9beca61fe3fc18c03ec13cbc5%2Fe1.png?generation=1597291913411797&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F25b4e74365ae124d7e6689411cdc9f7f%2Fe2.png?generation=1597291925498803&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F7f6528504e281e6355766caf4ab6a086%2Fe3.png?generation=1597291935745289&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 968497,
          "postDate": "2020-08-13T04:17:38.053Z",
          "content": "<p>hope it will be fixed soon.</p>",
          "rawMarkdown": "hope it will be fixed soon.",
          "votes": 5
        },
        {
          "id": 969492,
          "postDate": "2020-08-13T18:07:41.563Z",
          "content": "<p>The rollback to TF 2.2 is currently in progress. If you restart your session or create a new one, it is likely that you'll get yesterday's TF 2.2.</p>",
          "rawMarkdown": "The rollback to TF 2.2 is currently in progress. If you restart your session or create a new one, it is likely that you'll get yesterday's TF 2.2.",
          "votes": 1
        },
        {
          "id": 969844,
          "postDate": "2020-08-14T01:55:17.390Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ifigotin\" target=\"_blank\">@ifigotin</a> , after the rollback everything works like it used to.</p>",
          "rawMarkdown": "Thanks @ifigotin , after the rollback everything works like it used to.",
          "votes": 2
        },
        {
          "id": 969857,
          "postDate": "2020-08-14T02:29:53.523Z",
          "content": "<p>Sounds great. As <a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> mentioned in this or the other discussion, we filed several issues with the TF team internally.  </p>",
          "rawMarkdown": "Sounds great. As @mgornergoogle mentioned in this or the other discussion, we filed several issues with the TF team internally.  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 968548,
      "postDate": "2020-08-13T05:13:23.233Z",
      "content": "<p>I was facing this <code>Resource Exhausted Error</code> for quite a while with big image and model size, I don't know if you are facing it due to the same reasons, but lowering the batch size worked for me.</p>",
      "rawMarkdown": "I was facing this `Resource Exhausted Error` for quite a while with big image and model size, I don't know if you are facing it due to the same reasons, but lowering the batch size worked for me.",
      "votes": 1,
      "replies": [
        {
          "id": 968567,
          "postDate": "2020-08-13T05:42:08.507Z",
          "content": "<p>Yes, that is the normal fix. But something changed at Kaggle. My popular notebook which ran yesterday does not run today. Something is different about the TPU</p>",
          "rawMarkdown": "Yes, that is the normal fix. But something changed at Kaggle. My popular notebook which ran yesterday does not run today. Something is different about the TPU",
          "votes": 2
        },
        {
          "id": 968597,
          "postDate": "2020-08-13T06:15:26.527Z",
          "content": "<p>Yeah, it definitely has changed, it didn't happen before the last week or so! :/</p>",
          "rawMarkdown": "Yeah, it definitely has changed, it didn't happen before the last week or so! :/",
          "votes": 1
        }
      ]
    },
    {
      "id": 968510,
      "postDate": "2020-08-13T04:37:30.203Z",
      "content": "<p>I have the same issue here</p>",
      "rawMarkdown": "I have the same issue here",
      "votes": 1,
      "replies": [
        {
          "id": 968512,
          "postDate": "2020-08-13T04:41:59.347Z",
          "content": "<p>It's frustrating. I wonder what happened. Did TF 2.3 cause this?</p>",
          "rawMarkdown": "It's frustrating. I wonder what happened. Did TF 2.3 cause this?",
          "votes": 1
        },
        {
          "id": 969125,
          "postDate": "2020-08-13T13:26:55.713Z",
          "content": "<p>Yes it is very frustrating. I was trying a new idea yesterday and blocked by this problem. I think it is related to TF 2.3. I tried to install TF 2.2.0 on kaggle but i haven't succeeded yet.</p>",
          "rawMarkdown": "Yes it is very frustrating. I was trying a new idea yesterday and blocked by this problem. I think it is related to TF 2.3. I tried to install TF 2.2.0 on kaggle but i haven't succeeded yet."
        }
      ]
    },
    {
      "id": 969503,
      "postDate": "2020-08-13T18:17:53.790Z",
      "content": "<p>[Joking aside] After getting back TF 2.2, I think the TPU users can demand session limit [at least one hour extra, which means 3 hours + 1 ] - due to lossy one-day. 😏</p>",
      "rawMarkdown": "[Joking aside] After getting back TF 2.2, I think the TPU users can demand session limit [at least one hour extra, which means 3 hours + 1 ] - due to lossy one-day. 😏",
      "votes": 2
    },
    {
      "id": 968984,
      "postDate": "2020-08-13T11:39:29.660Z",
      "content": "<p>It's the same problem as in colab 1 week ago, the TPU memory seems to be halved. Can kaggle please solve this, as training time is also very important in 3-hour kernels ???</p>",
      "rawMarkdown": "It's the same problem as in colab 1 week ago, the TPU memory seems to be halved. Can kaggle please solve this, as training time is also very important in 3-hour kernels ???",
      "votes": 2
    },
    {
      "id": 969841,
      "postDate": "2020-08-14T01:51:44.050Z",
      "content": "<p>For info, if anyone is trying to access private datasets in Colab, the following authentication code needs to be added in the colab notebook.</p>\n<pre><code>from google.colab import auth\nauth.authenticate_user()\n</code></pre>\n<p>The rest remains same as <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> explained.</p>",
      "rawMarkdown": "For info, if anyone is trying to access private datasets in Colab, the following authentication code needs to be added in the colab notebook.\n```\nfrom google.colab import auth\nauth.authenticate_user()\n```\nThe rest remains same as @deepkim explained."
    },
    {
      "id": 969449,
      "postDate": "2020-08-13T17:33:29.123Z",
      "content": "<p>What a mess !! Doesn't work on colab or kaggle for hundreds of users. Doesn't Tensorflow team check basic stuff before upgrades ?</p>",
      "rawMarkdown": "What a mess !! Doesn't work on colab or kaggle for hundreds of users. Doesn't Tensorflow team check basic stuff before upgrades ?"
    },
    {
      "id": 969243,
      "postDate": "2020-08-13T15:09:17.867Z",
      "content": "<p>I encountered the same error.<br>\nSo I used google colab TPU and the auc dropped 0.01 :(</p>",
      "rawMarkdown": "I encountered the same error.\nSo I used google colab TPU and the auc dropped 0.01 :("
    },
    {
      "id": 969044,
      "postDate": "2020-08-13T12:29:42.260Z",
      "content": "<p>So, that's the issue, causing OOM -_-<br>\nThe competition is about to end, and most of us rely on TPU and now it's giving this error at this endpoint. Damn! </p>\n<p><a href=\"https://www.kaggle.com/tahsin\" target=\"_blank\">@tahsin</a> vai, I think we're heading this issue. </p>",
      "rawMarkdown": "So, that's the issue, causing OOM -_-\nThe competition is about to end, and most of us rely on TPU and now it's giving this error at this endpoint. Damn! \n\n@tahsin vai, I think we're heading this issue. "
    },
    {
      "id": 968784,
      "postDate": "2020-08-13T08:30:53.393Z",
      "content": "<p>Can anyone select \"Pin to original environment (2020-07-29)\" on option Environment preferences? I tried it to solve the problem, but it's disabled on my end.</p>",
      "rawMarkdown": "Can anyone select \"Pin to original environment (2020-07-29)\" on option Environment preferences? I tried it to solve the problem, but it's disabled on my end."
    },
    {
      "id": 968782,
      "postDate": "2020-08-13T08:28:13.177Z",
      "content": "<p>Am currently using both Kaggle and Colab TPUs. Both are working fine as of writing this.<br>\nColab TPU needs the downgrade to 2.2 as mentioned in another thread though.</p>",
      "rawMarkdown": "Am currently using both Kaggle and Colab TPUs. Both are working fine as of writing this.\nColab TPU needs the downgrade to 2.2 as mentioned in another thread though."
    },
    {
      "id": 968553,
      "postDate": "2020-08-13T05:27:39.953Z",
      "content": "<p>Hmm. not sure what causes it… But for your information in Colab downgrading to TF 2.2 is working fine.</p>",
      "rawMarkdown": "Hmm. not sure what causes it... But for your information in Colab downgrading to TF 2.2 is working fine.",
      "replies": [
        {
          "id": 968575,
          "postDate": "2020-08-13T05:49:56.727Z",
          "content": "<p>Yes, downgrading to TF 2.2 works in Colab but doesn't work in kaggle notebook. </p>",
          "rawMarkdown": "Yes, downgrading to TF 2.2 works in Colab but doesn't work in kaggle notebook. ",
          "replies": [
            {
              "id": 968628,
              "postDate": "2020-08-13T06:33:16.470Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/gmhost\" target=\"_blank\">@gmhost</a> ! could you tell me how you downgrade to TF 2.2 in Colab.<br>\n is it pip install tensorflow-gpu==2.2.0 ? </p>",
              "rawMarkdown": "Hi @gmhost ! could you tell me how you downgrade to TF 2.2 in Colab.\n is it pip install tensorflow-gpu==2.2.0 ? "
            },
            {
              "id": 968650,
              "postDate": "2020-08-13T06:48:42.927Z",
              "content": "<p>!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0</p>",
              "rawMarkdown": "!pip install tensorflow~=2.2.0 tensorflow\\_gcs\\_config~=2.2.0"
            }
          ]
        }
      ]
    },
    {
      "id": 969189,
      "postDate": "2020-08-13T14:30:24.313Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 968664,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2020-08-13T06:58:33.253000",
      "content": "<p><strong>USE COLAB! Colab is working well</strong></p>\n<p><strong>just add these lines to colab before importing tensorflow module (downgrade to TF 2.2)</strong></p>\n<p>!pip install tensorflow~=2.2.0 tensorflow<em>gcs</em>config~=2.2.0<br>\nimport tensorflow as tf<br>\nimport requests<br>\nimport os<br>\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB<em>TPU</em>ADDR\"].split(\":\")[0], tf.<strong>version</strong>))<br>\nif resp.status_code != 200:<br>\n    print(\"Failed to switch the TPU to TF {}\".format(version))</p>\n<p><strong>To access data from kaggle to COLAB, you need to manually input GCS PATHs like this on colab</strong></p>\n<p>gcs_path = 'gs://kds-cfb16ed5f55adb5ab35a75a2fe74c3dc11a4b869dc8cfb3d9b1759e1';</p>\n<p>gcs_path2 = 'gs://kds-79d9c56dcae7978f98a6bbe8318ac6866ba4778fc95c5871fcfd0f03';</p>\n<p>gcs_path3 = 'gs://kds-032edf427ad6e338f5392f73759ff404a1a3f25c2abdc4995136de30';</p>\n<p>GCS<em>PATH = [None]*FOLDS; GCS</em>PATH2 = [None]*FOLDS; GCS_PATH3 = [None]*FOLDS</p>\n<p>for i,k in enumerate(IMG_SIZES):</p>\n<pre><code>GCS_PATH[i] = gcs_path;\n\n\nGCS_PATH2[i] = gcs_path2;\n\n\nGCS_PATH3[i] = gcs_path3;\n</code></pre>\n<p>files<em>train = np.sort(np.array(tf.io.gfile.glob(GCS</em>PATH[0] + '/train*.tfrec')));</p>\n<p>files<em>test  = np.sort(np.array(tf.io.gfile.glob(GCS</em>PATH[0] + '/test*.tfrec')));  </p>\n<p><strong>You can get directly GCS PATHs FROM KAGGLE ENVIRONMENT(In kaggle to get GCS PATHs)</strong></p>\n<p>GCS<em>PATH = [None]*FOLDS; GCS</em>PATH2 = [None]*FOLDS; GCS_PATH3 = [None]*FOLDS</p>\n<p>for i,k in enumerate(IMG_SIZES[:FOLDS]):</p>\n<pre><code>GCS_PATH[i] = KaggleDatasets().get_gcs_path('melanoma-%ix%i'%(k,k))\n\n\nGCS_PATH2[i] = KaggleDatasets().get_gcs_path('isic2019-%ix%i'%(k,k))\n\n\nGCS_PATH3[i] = KaggleDatasets().get_gcs_path('malignant-v2-%ix%i'%(k,k))\n</code></pre>\n<p>files<em>train = np.sort(np.array(tf.io.gfile.glob(GCS</em>PATH[0] + '/train*.tfrec')))</p>\n<p>files<em>test  = np.sort(np.array(tf.io.gfile.glob(GCS</em>PATH[0] + '/test*.tfrec')))</p>\n<p>print(GCS<em>PATH, GCS</em>PATH2, GCS_PATH3)</p>",
      "votes": 7,
      "replies": [
        {
          "id": 968910,
          "author_name": "Joeran",
          "author_url": "",
          "post_date": "2020-08-13T10:34:08.100000",
          "content": "<p>I tried your method and it seems to be working, I can train on Colab now. However, I still get OOM when using the same batch size as yesterday, which worked well on Kaggle then. <br>\nDo you know if the initialisation of the TPU should be any different on Colab than on Kaggle?<br>\nI use the initialisation from Chris Deotte's popular notebook now. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 968923,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2020-08-13T10:50:43.880000",
          "content": "<p>halve the batch size! <br>\ncolab tpu memory is much smaller than kaggle's<br>\nif you set a batch size at 32 on kaggle machine,  you should set it to 16 on colab( this is based on my trial and error)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 968940,
          "author_name": "sajwankit",
          "author_url": "",
          "post_date": "2020-08-13T11:07:17.007000",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> were you able to load/access private kaggle datasets in colab? <br>\nI saved all of my tpu resources for training at last and this happened.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 968975,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2020-08-13T11:31:05.450000",
          "content": "<p><a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a>  i haven tried accessing private data yet. I have been just using public data. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 969021,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-13T12:09:52.600000",
      "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> Do you have any idea why TPU batch sizes that worked before August 12th don't work after August 12th? Before August 12th we could train 384x384 batch size 256 = 32 x 8 on EffNetB6 <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a>. Now we cannot.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 969436,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-08-13T17:18:22.773000",
          "content": "<p>Yes, I saw this issue with TF 2.3 as well. Investigating. Did you encounter this with other models besides EfficientNet ?<br>\n(The next thing I am going to try is to load EfficientNet from TFHub directly - that works now in TF 2.3 with the /job:localhost LoadOption - instead of the pip-installed lib. Maybe there was a bug fix in there…)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 969605,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-08-13T19:58:57.210000",
          "content": "<p>I tried tf.keras.applications.EfficinetNetBx and it did not work any better under TF 2.3. I have not been able to load EfficientNet from TF Hub yes because the tensorflow-hub library has not yet been updated with the new hub.KerasLayer(load_options) parameter. The TF team is investigating…</p>",
          "votes": 0,
          "replies": [
            {
              "id": 969742,
              "author_name": "Innat",
              "author_url": "",
              "post_date": "2020-08-13T22:36:10.527000",
              "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> AFAIK, about <code>tf.keras.applicaiton.EfficientNetBx</code> - This API is new and only available in <code>tf-nightly</code>. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 970768,
              "author_name": "Martin Görner",
              "author_url": "",
              "post_date": "2020-08-14T18:55:01.490000",
              "content": "<p>no tf.keras.applicaiton.EfficientNetBx is there in TF 2.3</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 969608,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-08-13T19:59:35.660000",
          "content": "<p>Everyone, have you seen regressions with other models than EfficientNetBx ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 969668,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-13T20:51:39.003000",
          "content": "<p>I have tried all models on TF 2.2 but I only tried EfficientNet on TF 2.3 so far.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 969299,
      "author_name": "Ilya Figotin",
      "author_url": "",
      "post_date": "2020-08-13T15:47:36.867000",
      "content": "<p>The code snippet to downgrade to TF 2.2 for now should be as follows:</p>\n<pre><code>!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"TPU_NAME\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n</code></pre>\n<p>Let us know if this unblocks you, and we will be looking at the issue in the meantime. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 969327,
          "author_name": "omarsamir",
          "author_url": "",
          "post_date": "2020-08-13T16:03:19.987000",
          "content": "<p>Hi llya, Thanks for sharing this code snippets. I tried it but still couldn't successful to setup it and i found the following error:</p>\n<p>ConnectionError: HTTPConnectionPool(host='grpc', port=8475): Max retries exceeded with url: /requestversion/2.2.0 (Caused by NewConnectionError(': Failed to establish a new connection: [Errno -2] Name or service not known'))</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 969345,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-13T16:10:20.657000",
          "content": "<p><a href=\"https://www.kaggle.com/osamir\" target=\"_blank\">@osamir</a> this looks like an intermittent error. I assume your Internet was ON and that was a TPU enabled notebook, right? If so, please retry. I did work at least in a couple of notebooks I tried.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 969352,
          "author_name": "omarsamir",
          "author_url": "",
          "post_date": "2020-08-13T16:18:50.900000",
          "content": "<p>I double checked that TPU is enabled and internet is on but still facing this error. I will try it again and see if it still exists or not.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 969381,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-13T16:39:33.330000",
          "content": "<p>We are planning to downgrade back to TF 2.2 for the TPU machines. So, hopefully, you will not need the workaround soon.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 969387,
              "author_name": "Innat",
              "author_url": "",
              "post_date": "2020-08-13T16:42:12.087000",
              "content": "<p>Upgrade to TF brings a mess many times. IMHO, authority should double check before launch, one day waste at the end time; very lossy. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 969405,
              "author_name": "Phi",
              "author_url": "",
              "post_date": "2020-08-13T16:53:04.030000",
              "content": "<p>actually the workaround doesn't work for me as it gives the same error as reported, and I've tried many times</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 969481,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-13T17:56:41.110000",
          "content": "<p>Workaround does not work for me. Using workaround gives new weird errors.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 969696,
          "author_name": "Ronaldo S.A. Batista",
          "author_url": "",
          "post_date": "2020-08-13T21:27:12.607000",
          "content": "<p>Same for me, in Kaggle and on Colab</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 968490,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-13T04:12:23.387000",
      "content": "<p>I agree. I cannot use TPU tonight. Everything that worked yesterday does not work today</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F1a0fcdf9beca61fe3fc18c03ec13cbc5%2Fe1.png?generation=1597291913411797&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F25b4e74365ae124d7e6689411cdc9f7f%2Fe2.png?generation=1597291925498803&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F7f6528504e281e6355766caf4ab6a086%2Fe3.png?generation=1597291935745289&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 968497,
          "author_name": "kwang",
          "author_url": "",
          "post_date": "2020-08-13T04:17:38.053000",
          "content": "<p>hope it will be fixed soon.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 969492,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-13T18:07:41.563000",
          "content": "<p>The rollback to TF 2.2 is currently in progress. If you restart your session or create a new one, it is likely that you'll get yesterday's TF 2.2.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 969844,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-14T01:55:17.390000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ifigotin\" target=\"_blank\">@ifigotin</a> , after the rollback everything works like it used to.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 969857,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-14T02:29:53.523000",
          "content": "<p>Sounds great. As <a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> mentioned in this or the other discussion, we filed several issues with the TF team internally.  </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 968548,
      "author_name": "Gajendra Saraswat",
      "author_url": "",
      "post_date": "2020-08-13T05:13:23.233000",
      "content": "<p>I was facing this <code>Resource Exhausted Error</code> for quite a while with big image and model size, I don't know if you are facing it due to the same reasons, but lowering the batch size worked for me.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 968567,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-13T05:42:08.507000",
          "content": "<p>Yes, that is the normal fix. But something changed at Kaggle. My popular notebook which ran yesterday does not run today. Something is different about the TPU</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 968597,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-08-13T06:15:26.527000",
          "content": "<p>Yeah, it definitely has changed, it didn't happen before the last week or so! :/</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 968510,
      "author_name": "omarsamir",
      "author_url": "",
      "post_date": "2020-08-13T04:37:30.203000",
      "content": "<p>I have the same issue here</p>",
      "votes": 1,
      "replies": [
        {
          "id": 968512,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-13T04:41:59.347000",
          "content": "<p>It's frustrating. I wonder what happened. Did TF 2.3 cause this?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 969125,
          "author_name": "omarsamir",
          "author_url": "",
          "post_date": "2020-08-13T13:26:55.713000",
          "content": "<p>Yes it is very frustrating. I was trying a new idea yesterday and blocked by this problem. I think it is related to TF 2.3. I tried to install TF 2.2.0 on kaggle but i haven't succeeded yet.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 969503,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-08-13T18:17:53.790000",
      "content": "<p>[Joking aside] After getting back TF 2.2, I think the TPU users can demand session limit [at least one hour extra, which means 3 hours + 1 ] - due to lossy one-day. 😏</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 968984,
      "author_name": "Phi",
      "author_url": "",
      "post_date": "2020-08-13T11:39:29.660000",
      "content": "<p>It's the same problem as in colab 1 week ago, the TPU memory seems to be halved. Can kaggle please solve this, as training time is also very important in 3-hour kernels ???</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 969841,
      "author_name": "sajwankit",
      "author_url": "",
      "post_date": "2020-08-14T01:51:44.050000",
      "content": "<p>For info, if anyone is trying to access private datasets in Colab, the following authentication code needs to be added in the colab notebook.</p>\n<pre><code>from google.colab import auth\nauth.authenticate_user()\n</code></pre>\n<p>The rest remains same as <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> explained.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 969449,
      "author_name": "VSR",
      "author_url": "",
      "post_date": "2020-08-13T17:33:29.123000",
      "content": "<p>What a mess !! Doesn't work on colab or kaggle for hundreds of users. Doesn't Tensorflow team check basic stuff before upgrades ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 969243,
      "author_name": "fantastic_hirarin",
      "author_url": "",
      "post_date": "2020-08-13T15:09:17.867000",
      "content": "<p>I encountered the same error.<br>\nSo I used google colab TPU and the auc dropped 0.01 :(</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 969044,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-08-13T12:29:42.260000",
      "content": "<p>So, that's the issue, causing OOM -_-<br>\nThe competition is about to end, and most of us rely on TPU and now it's giving this error at this endpoint. Damn! </p>\n<p><a href=\"https://www.kaggle.com/tahsin\" target=\"_blank\">@tahsin</a> vai, I think we're heading this issue. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 968784,
      "author_name": "Sandy Khosasi",
      "author_url": "",
      "post_date": "2020-08-13T08:30:53.393000",
      "content": "<p>Can anyone select \"Pin to original environment (2020-07-29)\" on option Environment preferences? I tried it to solve the problem, but it's disabled on my end.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 968782,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-08-13T08:28:13.177000",
      "content": "<p>Am currently using both Kaggle and Colab TPUs. Both are working fine as of writing this.<br>\nColab TPU needs the downgrade to 2.2 as mentioned in another thread though.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 968553,
      "author_name": "Bayartsogt Yadamsuren",
      "author_url": "",
      "post_date": "2020-08-13T05:27:39.953000",
      "content": "<p>Hmm. not sure what causes it… But for your information in Colab downgrading to TF 2.2 is working fine.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 968575,
          "author_name": "kwang",
          "author_url": "",
          "post_date": "2020-08-13T05:49:56.727000",
          "content": "<p>Yes, downgrading to TF 2.2 works in Colab but doesn't work in kaggle notebook. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 968628,
              "author_name": "Mehul Sampat",
              "author_url": "",
              "post_date": "2020-08-13T06:33:16.470000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/gmhost\" target=\"_blank\">@gmhost</a> ! could you tell me how you downgrade to TF 2.2 in Colab.<br>\n is it pip install tensorflow-gpu==2.2.0 ? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 968650,
              "author_name": "kwang",
              "author_url": "",
              "post_date": "2020-08-13T06:48:42.927000",
              "content": "<p>!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 969189,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-13T14:30:24.313000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "968488": "since kaggle has updated tensorflow version to 2.3, the notebook cann't run in the new version.\n\nit would cause  ResourceExhaustedError , something like [this](https://github.com/googlecolab/colabtools/issues/1470).\n\nAnd downgrade tf to 2.2 didn't work. \n",
    "968664": "**USE COLAB! Colab is working well**\n\n**just add these lines to colab before importing tensorflow module (downgrade to TF 2.2)**\n\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n    print(\"Failed to switch the TPU to TF {}\".format(version))\n\n\n**To access data from kaggle to COLAB, you need to manually input GCS PATHs like this on colab**\n\n\ngcs_path = 'gs://kds-cfb16ed5f55adb5ab35a75a2fe74c3dc11a4b869dc8cfb3d9b1759e1';\n\n\ngcs_path2 = 'gs://kds-79d9c56dcae7978f98a6bbe8318ac6866ba4778fc95c5871fcfd0f03';\n\n\ngcs_path3 = 'gs://kds-032edf427ad6e338f5392f73759ff404a1a3f25c2abdc4995136de30';\n\n\n\n\nGCS_PATH = [None]\\*FOLDS; GCS_PATH2 = [None]\\*FOLDS; GCS_PATH3 = [None]\\*FOLDS\n\n\nfor i,k in enumerate(IMG_SIZES):\n\n\n    GCS_PATH[i] = gcs_path;\n\n\n    GCS_PATH2[i] = gcs_path2;\n\n\n    GCS_PATH3[i] = gcs_path3;\n\n\nfiles_train = np.sort(np.array(tf.io.gfile.glob(GCS_PATH[0] + '/train\\*.tfrec')));\n\n  \nfiles_test  = np.sort(np.array(tf.io.gfile.glob(GCS_PATH[0] + '/test\\*.tfrec')));  \n\n**You can get directly GCS PATHs FROM KAGGLE ENVIRONMENT(In kaggle to get GCS PATHs)**\n\nGCS_PATH = [None]\\*FOLDS; GCS_PATH2 = [None]\\*FOLDS; GCS_PATH3 = [None]\\*FOLDS\n\n\nfor i,k in enumerate(IMG_SIZES[:FOLDS]):\n\n\n    GCS_PATH[i] = KaggleDatasets().get_gcs_path('melanoma-%ix%i'%(k,k))\n\n\n    GCS_PATH2[i] = KaggleDatasets().get_gcs_path('isic2019-%ix%i'%(k,k))\n\n\n    GCS_PATH3[i] = KaggleDatasets().get_gcs_path('malignant-v2-%ix%i'%(k,k))\n\n\nfiles_train = np.sort(np.array(tf.io.gfile.glob(GCS_PATH[0] + '/train\\*.tfrec')))\n\n\nfiles_test  = np.sort(np.array(tf.io.gfile.glob(GCS_PATH[0] + '/test\\*.tfrec')))\n\nprint(GCS_PATH, GCS_PATH2, GCS_PATH3)",
    "969021": "@mgornergoogle Do you have any idea why TPU batch sizes that worked before August 12th don't work after August 12th? Before August 12th we could train 384x384 batch size 256 = 32 x 8 on EffNetB6 [here][1]. Now we cannot.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "969299": "The code snippet to downgrade to TF 2.2 for now should be as follows:\n```\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"TPU_NAME\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n```\nLet us know if this unblocks you, and we will be looking at the issue in the meantime. ",
    "968490": "I agree. I cannot use TPU tonight. Everything that worked yesterday does not work today\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F1a0fcdf9beca61fe3fc18c03ec13cbc5%2Fe1.png?generation=1597291913411797&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F25b4e74365ae124d7e6689411cdc9f7f%2Fe2.png?generation=1597291925498803&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F7f6528504e281e6355766caf4ab6a086%2Fe3.png?generation=1597291935745289&alt=media)",
    "968548": "I was facing this `Resource Exhausted Error` for quite a while with big image and model size, I don't know if you are facing it due to the same reasons, but lowering the batch size worked for me.",
    "968510": "I have the same issue here",
    "969503": "[Joking aside] After getting back TF 2.2, I think the TPU users can demand session limit [at least one hour extra, which means 3 hours + 1 ] - due to lossy one-day. 😏",
    "968984": "It's the same problem as in colab 1 week ago, the TPU memory seems to be halved. Can kaggle please solve this, as training time is also very important in 3-hour kernels ???",
    "969841": "For info, if anyone is trying to access private datasets in Colab, the following authentication code needs to be added in the colab notebook.\n```\nfrom google.colab import auth\nauth.authenticate_user()\n```\nThe rest remains same as @deepkim explained.",
    "969449": "What a mess !! Doesn't work on colab or kaggle for hundreds of users. Doesn't Tensorflow team check basic stuff before upgrades ?",
    "969243": "I encountered the same error.\nSo I used google colab TPU and the auc dropped 0.01 :(",
    "969044": "So, that's the issue, causing OOM -_-\nThe competition is about to end, and most of us rely on TPU and now it's giving this error at this endpoint. Damn! \n\n@tahsin vai, I think we're heading this issue. ",
    "968784": "Can anyone select \"Pin to original environment (2020-07-29)\" on option Environment preferences? I tried it to solve the problem, but it's disabled on my end.",
    "968782": "Am currently using both Kaggle and Colab TPUs. Both are working fine as of writing this.\nColab TPU needs the downgrade to 2.2 as mentioned in another thread though.",
    "968553": "Hmm. not sure what causes it... But for your information in Colab downgrading to TF 2.2 is working fine.",
    "969189": ""
  }
}