{
  "id": 172357,
  "title": "Colab update to tensorflow 2.3 memory problem",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172357",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2020-08-04T18:42:27.664000",
  "votes": 30,
  "comment_count": 40,
  "views": 0,
  "content": "<p>Sup, i know a lot of competitors are using colab to experiment. Yesterday colab update tensorflow to version 2.3. For some reason the script that i was using previously now is dying, the error is related to the memory of the tpu. If someone with more experince in this field could help us to resolve that problem i would appreciate it.</p>\n\n<p>Cheers, and good luck with the comp </p>",
  "messages": [
    {
      "id": 958072,
      "postDate": "2020-08-04T18:42:27.663Z",
      "content": "<p>Sup, i know a lot of competitors are using colab to experiment. Yesterday colab update tensorflow to version 2.3. For some reason the script that i was using previously now is dying, the error is related to the memory of the tpu. If someone with more experince in this field could help us to resolve that problem i would appreciate it.</p>\n\n<p>Cheers, and good luck with the comp </p>",
      "rawMarkdown": "Sup, i know a lot of competitors are using colab to experiment. Yesterday colab update tensorflow to version 2.3. For some reason the script that i was using previously now is dying, the error is related to the memory of the tpu. If someone with more experince in this field could help us to resolve that problem i would appreciate it.\n\nCheers, and good luck with the comp ",
      "votes": 30
    },
    {
      "id": 958437,
      "postDate": "2020-08-05T01:45:39.160Z",
      "content": "<p>A working workaround is posted on the colab issue tracker - works perfectly!</p>\n\n<p><code>\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n</code></p>",
      "rawMarkdown": "A working workaround is posted on the colab issue tracker - works perfectly!\n\n```\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n```",
      "votes": 27,
      "replies": [
        {
          "id": 958484,
          "postDate": "2020-08-05T02:44:24.337Z",
          "content": "<p>Nice, this work perfectly. Ty.</p>",
          "rawMarkdown": "Nice, this work perfectly. Ty."
        },
        {
          "id": 958650,
          "postDate": "2020-08-05T05:20:30.067Z",
          "content": "<p>thanks for the tip</p>",
          "rawMarkdown": "thanks for the tip"
        },
        {
          "id": 958890,
          "postDate": "2020-08-05T07:47:03.370Z",
          "content": "<p>Thank you. It works for me.</p>",
          "rawMarkdown": "Thank you. It works for me."
        },
        {
          "id": 959080,
          "postDate": "2020-08-05T10:21:27.807Z",
          "content": "<p>This works great and fixes all the issues. Thanks a lot for sharing!</p>",
          "rawMarkdown": "This works great and fixes all the issues. Thanks a lot for sharing!"
        },
        {
          "id": 959104,
          "postDate": "2020-08-05T10:50:02.287Z",
          "content": "<p>Thanks for sharing.</p>",
          "rawMarkdown": "Thanks for sharing."
        },
        {
          "id": 959861,
          "postDate": "2020-08-06T01:24:26.013Z",
          "content": "<p>Thanks for sharing !! It works for me.</p>",
          "rawMarkdown": "Thanks for sharing !! It works for me."
        },
        {
          "id": 962073,
          "postDate": "2020-08-07T19:18:24.650Z",
          "content": "<p>Thank You for the workaround</p>",
          "rawMarkdown": "Thank You for the workaround"
        },
        {
          "id": 962668,
          "postDate": "2020-08-08T10:25:26.553Z",
          "content": "<p>Do you have the same problem today? I'm facing the same problem today, and I tried the workaround but that didn't work…</p>",
          "rawMarkdown": "Do you have the same problem today? I'm facing the same problem today, and I tried the workaround but that didn't work..."
        },
        {
          "id": 962860,
          "postDate": "2020-08-08T13:52:28.280Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 958274,
      "postDate": "2020-08-04T22:27:31.900Z",
      "content": "<p>I am in the same boat. It took me some time to figure out that the Tensor Flow version had been changed on Colab. So, there were a few minutes during which I was questioning my sanity and was on the verge of quitting programming 😂.</p>",
      "rawMarkdown": "I am in the same boat. It took me some time to figure out that the Tensor Flow version had been changed on Colab. So, there were a few minutes during which I was questioning my sanity and was on the verge of quitting programming 😂.",
      "votes": 5
    },
    {
      "id": 958234,
      "postDate": "2020-08-04T21:21:43.507Z",
      "content": "<p>The same here, the only workaround that worked for me was to decrease the batch size like more than 2x times, resulting in at least 2-3x times the training speed decrease. I've submitted the issue here - please add your complaints as well to make it more visible: <a href=\"https://github.com/googlecolab/colabtools/issues/1470\">https://github.com/googlecolab/colabtools/issues/1470</a></p>",
      "rawMarkdown": "The same here, the only workaround that worked for me was to decrease the batch size like more than 2x times, resulting in at least 2-3x times the training speed decrease. I've submitted the issue here - please add your complaints as well to make it more visible: https://github.com/googlecolab/colabtools/issues/1470",
      "votes": 4,
      "replies": [
        {
          "id": 958395,
          "postDate": "2020-08-05T00:47:44.020Z",
          "content": "<p>Thanks for posting the issue. It's be solved.\n<a href=\"https://github.com/googlecolab/colabtools/issues/1470#issuecomment-668897876\">https://github.com/googlecolab/colabtools/issues/1470#issuecomment-668897876</a></p>",
          "rawMarkdown": "Thanks for posting the issue. It's be solved.\nhttps://github.com/googlecolab/colabtools/issues/1470#issuecomment-668897876",
          "votes": 2
        },
        {
          "id": 958450,
          "postDate": "2020-08-05T02:02:01.223Z",
          "content": "<p><a href=\"/yuanlin08\">@yuanlin08</a> Thank you for sharing this great news! I installed TF 2.2 on Colab following the instructions found at your link and it seems to be working -- I am training my second epoch with no memory overload!  It is back to normal!</p>",
          "rawMarkdown": "@yuanlin08 Thank you for sharing this great news! I installed TF 2.2 on Colab following the instructions found at your link and it seems to be working -- I am training my second epoch with no memory overload!  It is back to normal!",
          "votes": 2
        },
        {
          "id": 958473,
          "postDate": "2020-08-05T02:33:03.467Z",
          "content": "<p><a href=\"/graf10a\">@graf10a</a> Great! Good to know!😄 </p>",
          "rawMarkdown": "@graf10a Great! Good to know!😄 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 958254,
      "postDate": "2020-08-04T21:57:12.413Z",
      "content": "<p>I face the same issue! I spent the whole morning looking for a solution, but in the end I found nothing... Towards the end of the game, I was really very frustrated..</p>",
      "rawMarkdown": "I face the same issue! I spent the whole morning looking for a solution, but in the end I found nothing... Towards the end of the game, I was really very frustrated..",
      "votes": 1
    },
    {
      "id": 958172,
      "postDate": "2020-08-04T20:13:35.897Z",
      "content": "<p>I dont have memory problem yet, but mine has weird warning logs that never seen before and the training time is almost doubled ☹️ </p>",
      "rawMarkdown": "I dont have memory problem yet, but mine has weird warning logs that never seen before and the training time is almost doubled ☹️ ",
      "votes": 1
    },
    {
      "id": 958108,
      "postDate": "2020-08-04T19:17:24.107Z",
      "content": "<p>Hey <a href=\"/ragnar123\">@ragnar123</a> , I was having this same issue but I did not know it was related to the Tensorflow version, probably using the previous version may solve the issue, I will try it later.</p>",
      "rawMarkdown": "Hey @ragnar123 , I was having this same issue but I did not know it was related to the Tensorflow version, probably using the previous version may solve the issue, I will try it later.",
      "votes": 1,
      "replies": [
        {
          "id": 958136,
          "postDate": "2020-08-04T19:48:45.033Z",
          "content": "<p>The problem is that you cannot just downgrade the tensorflow on colab because the TPU are compiled for the latest version only.</p>\n\n<p>I'm having the exact same issues.</p>",
          "rawMarkdown": "The problem is that you cannot just downgrade the tensorflow on colab because the TPU are compiled for the latest version only.\n\nI'm having the exact same issues.",
          "votes": 3
        },
        {
          "id": 958167,
          "postDate": "2020-08-04T20:06:29.850Z",
          "content": "<p>Yes, and here is the <a href=\"https://github.com/tensorflow/tensorflow/issues/40622#issuecomment-654488935\">comment</a> on a similar issue on GitHub.</p>",
          "rawMarkdown": "Yes, and here is the [comment](https://github.com/tensorflow/tensorflow/issues/40622#issuecomment-654488935) on a similar issue on GitHub.",
          "votes": 1
        },
        {
          "id": 958210,
          "postDate": "2020-08-04T21:04:06.377Z",
          "content": "<p>damn.. bad time for us, so close to the competition end.</p>",
          "rawMarkdown": "damn.. bad time for us, so close to the competition end.",
          "votes": 1
        },
        {
          "id": 958250,
          "postDate": "2020-08-04T21:49:34.177Z",
          "content": "<p>Yea horrible timing!</p>",
          "rawMarkdown": "Yea horrible timing!",
          "votes": 1
        }
      ]
    },
    {
      "id": 958599,
      "postDate": "2020-08-05T04:33:12.760Z",
      "content": "<p><code>python\n!pip install tensorflow==2.2 --force\n</code>\nthis to install Tensorflow 2.2</p>\n\n<p><code>python\n%tensorflow_version 2.2\n</code>\nthis to switch 2.2\nand don't restart the colab notebook </p>",
      "rawMarkdown": "```python\n!pip install tensorflow==2.2 --force\n```\nthis to install Tensorflow 2.2\n\n```python\n%tensorflow_version 2.2\n```\nthis to switch 2.2\nand don't restart the colab notebook ",
      "votes": 2,
      "replies": [
        {
          "id": 958649,
          "postDate": "2020-08-05T05:20:09.970Z",
          "content": "<p>thanks for the tip</p>",
          "rawMarkdown": "thanks for the tip",
          "votes": 1
        },
        {
          "id": 958805,
          "postDate": "2020-08-05T06:37:02.310Z",
          "content": "<p>Still have got no success with this.</p>\n\n<p>I thought you could only switch between major versions with this magic command?</p>\n\n<p>I.e.</p>\n\n<p><code>%tensorflow_version 1.x</code></p>\n\n<p>or</p>\n\n<p><code>%tensorflow_version 2.x</code></p>",
          "rawMarkdown": "Still have got no success with this.\n\nI thought you could only switch between major versions with this magic command?\n\nI.e.\n\n`%tensorflow_version 1.x`\n\nor\n\n`%tensorflow_version 2.x`"
        },
        {
          "id": 958818,
          "postDate": "2020-08-05T06:43:54.250Z",
          "content": "<p><a href=\"/group16\">@group16</a> Try Gena's suggestion. It works for me.</p>",
          "rawMarkdown": "@group16 Try Gena's suggestion. It works for me.",
          "votes": 1
        },
        {
          "id": 958820,
          "postDate": "2020-08-05T06:45:41.207Z",
          "content": "<p>Thanks Alexey! </p>\n\n<p>I'm still getting <code>TypeError: catching classes that do not inherit from BaseException is not allowed</code> with <a href=\"/gdonchyts\">@gdonchyts</a> his solution.</p>",
          "rawMarkdown": "Thanks Alexey! \n\nI'm still getting `TypeError: catching classes that do not inherit from BaseException is not allowed` with @gdonchyts his solution.",
          "votes": 1
        },
        {
          "id": 958849,
          "postDate": "2020-08-05T07:11:41.333Z",
          "content": "<p><a href=\"/group16\">@group16</a> This is interesting. I have not seen this error before. Have you made any changes in your code recently?</p>",
          "rawMarkdown": "@group16 This is interesting. I have not seen this error before. Have you made any changes in your code recently?",
          "votes": 1
        },
        {
          "id": 958850,
          "postDate": "2020-08-05T07:12:23.813Z",
          "content": "<p>I get this error when I try to initialize my TPU's/clear caches. Let me see if I made some mistake there...</p>\n\n<p>EDIT: It works now. I was reading the 2.3 changelog yesterday and saw that <code>tf.distribute.experimental.TPUStrategy(tpu)</code> was renamed to <code>tf.distribute.TPUStrategy(tpu)</code> but forgot to revert that change.</p>\n\n<p>Many thanks!!!</p>",
          "rawMarkdown": "I get this error when I try to initialize my TPU's/clear caches. Let me see if I made some mistake there...\n\nEDIT: It works now. I was reading the 2.3 changelog yesterday and saw that `tf.distribute.experimental.TPUStrategy(tpu)` was renamed to `tf.distribute.TPUStrategy(tpu)` but forgot to revert that change.\n\nMany thanks!!!",
          "votes": 4
        },
        {
          "id": 958867,
          "postDate": "2020-08-05T07:29:34.257Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1614008%2Fb4dc809390d8b9b0ab19eb340da8b464%2Fscreencapture-colab-research-google-drive-15toH1bOGjdr68bVMDx7hmYGAo-avdoV4-2020-08-05-13_02_48.png?generation=1596612902310561&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1614008%2Fb4dc809390d8b9b0ab19eb340da8b464%2Fscreencapture-colab-research-google-drive-15toH1bOGjdr68bVMDx7hmYGAo-avdoV4-2020-08-05-13_02_48.png?generation=1596612902310561&amp;alt=media)\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 958138,
      "postDate": "2020-08-04T19:49:39.737Z",
      "content": "<p>Same issue. And I tried using pip install tensorflow==2.2 didn't work. \nWith this new version I have to reduce the batch size which will take me much more time in training. Any idea ?</p>",
      "rawMarkdown": "Same issue. And I tried using pip install tensorflow==2.2 didn't work. \nWith this new version I have to reduce the batch size which will take me much more time in training. Any idea ?",
      "votes": 2,
      "replies": [
        {
          "id": 958249,
          "postDate": "2020-08-04T21:48:56.667Z",
          "content": "<p>Also tried to install tf 2.2 but i believe that the tpu comes with a special driver for each version.</p>",
          "rawMarkdown": "Also tried to install tf 2.2 but i believe that the tpu comes with a special driver for each version.",
          "votes": 1
        }
      ]
    },
    {
      "id": 959673,
      "postDate": "2020-08-05T19:54:19.083Z",
      "content": "<p>A related question here didn't want to start new topic just for it. Does colab TPU's has less memory than kaggle ones or you just have to be lucky to connect decent one on colab?</p>",
      "rawMarkdown": "A related question here didn't want to start new topic just for it. Does colab TPU's has less memory than kaggle ones or you just have to be lucky to connect decent one on colab?",
      "replies": [
        {
          "id": 959710,
          "postDate": "2020-08-05T20:50:13.397Z",
          "content": "<p>Yes <a href=\"/datafan07\">@datafan07</a> Kaggle's TPU are v3 and Colab are v2, so you may notice that on Kaggle, you may use larger batches and it will also run the epochs faster!</p>",
          "rawMarkdown": "Yes @datafan07 Kaggle's TPU are v3 and Colab are v2, so you may notice that on Kaggle, you may use larger batches and it will also run the epochs faster!",
          "votes": 2
        },
        {
          "id": 959737,
          "postDate": "2020-08-05T21:14:41.567Z",
          "content": "<p>Aaah that's what I was suspicious about thanks for clarifying! Btw do you think batch size affect the final results? In my previous experiences had different results based on batch sizes, even with tuned learning rate according to batch size, is it the case in general?</p>",
          "rawMarkdown": "Aaah that's what I was suspicious about thanks for clarifying! Btw do you think batch size affect the final results? In my previous experiences had different results based on batch sizes, even with tuned learning rate according to batch size, is it the case in general?",
          "votes": 1
        },
        {
          "id": 959755,
          "postDate": "2020-08-05T21:36:32.590Z",
          "content": "<p>Hey <a href=\"/datafan07\">@datafan07</a> , that can be a tricky question, at one side smaller batch sizes can reduce overfitting, but larger batches will have more samples from each class and overall more representative samples from the complete set. \nIn my experience adjusting/scalling learning rate with batch size is not so trivial, in my current configuration I am having better results with smaller sizes, but I am not so sure of the reasons, what about you?</p>",
          "rawMarkdown": "Hey @datafan07 , that can be a tricky question, at one side smaller batch sizes can reduce overfitting, but larger batches will have more samples from each class and overall more representative samples from the complete set. \nIn my experience adjusting/scalling learning rate with batch size is not so trivial, in my current configuration I am having better results with smaller sizes, but I am not so sure of the reasons, what about you?",
          "votes": 1
        },
        {
          "id": 959764,
          "postDate": "2020-08-05T21:47:08.883Z",
          "content": "<p>Same here when I go bigger batches my score decreases in my config, but have been thinking about this a while, not found exact answer to that yet, everyone has different experiences. Not even limited to this competition...</p>",
          "rawMarkdown": "Same here when I go bigger batches my score decreases in my config, but have been thinking about this a while, not found exact answer to that yet, everyone has different experiences. Not even limited to this competition...",
          "votes": 1
        },
        {
          "id": 959789,
          "postDate": "2020-08-05T22:22:33.027Z",
          "content": "<p>Agreed, batch sizes are usually very dependent on task and architectures.</p>",
          "rawMarkdown": "Agreed, batch sizes are usually very dependent on task and architectures.",
          "votes": 1
        }
      ]
    },
    {
      "id": 959059,
      "postDate": "2020-08-05T09:58:01.577Z",
      "content": "<p>I am using gcp and have the same problem, so using tensorflow 2.2.</p>",
      "rawMarkdown": "I am using gcp and have the same problem, so using tensorflow 2.2."
    },
    {
      "id": 974390,
      "postDate": "2020-08-18T00:04:09.583Z",
      "content": "<p>thanks a lot.</p>",
      "rawMarkdown": "thanks a lot."
    }
  ],
  "comments": [
    {
      "id": 958437,
      "author_name": "Gena",
      "author_url": "",
      "post_date": "2020-08-05T01:45:39.160000",
      "content": "<p>A working workaround is posted on the colab issue tracker - works perfectly!</p>\n\n<p><code>\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n</code></p>",
      "votes": 27,
      "replies": [
        {
          "id": 958484,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-08-05T02:44:24.337000",
          "content": "<p>Nice, this work perfectly. Ty.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958650,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-05T05:20:30.067000",
          "content": "<p>thanks for the tip</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958890,
          "author_name": "Changyi",
          "author_url": "",
          "post_date": "2020-08-05T07:47:03.370000",
          "content": "<p>Thank you. It works for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 959080,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-08-05T10:21:27.807000",
          "content": "<p>This works great and fixes all the issues. Thanks a lot for sharing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 959104,
          "author_name": "Jose",
          "author_url": "",
          "post_date": "2020-08-05T10:50:02.287000",
          "content": "<p>Thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 959861,
          "author_name": "daishin",
          "author_url": "",
          "post_date": "2020-08-06T01:24:26.013000",
          "content": "<p>Thanks for sharing !! It works for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 962073,
          "author_name": "omarsamir",
          "author_url": "",
          "post_date": "2020-08-07T19:18:24.650000",
          "content": "<p>Thank You for the workaround</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 962668,
          "author_name": "DWGGG",
          "author_url": "",
          "post_date": "2020-08-08T10:25:26.553000",
          "content": "<p>Do you have the same problem today? I'm facing the same problem today, and I tried the workaround but that didn't work…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 962860,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-08T13:52:28.280000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 958274,
      "author_name": "Alexey Pronin",
      "author_url": "",
      "post_date": "2020-08-04T22:27:31.900000",
      "content": "<p>I am in the same boat. It took me some time to figure out that the Tensor Flow version had been changed on Colab. So, there were a few minutes during which I was questioning my sanity and was on the verge of quitting programming 😂.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 958234,
      "author_name": "Gena",
      "author_url": "",
      "post_date": "2020-08-04T21:21:43.507000",
      "content": "<p>The same here, the only workaround that worked for me was to decrease the batch size like more than 2x times, resulting in at least 2-3x times the training speed decrease. I've submitted the issue here - please add your complaints as well to make it more visible: <a href=\"https://github.com/googlecolab/colabtools/issues/1470\">https://github.com/googlecolab/colabtools/issues/1470</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 958395,
          "author_name": "Helen",
          "author_url": "",
          "post_date": "2020-08-05T00:47:44.020000",
          "content": "<p>Thanks for posting the issue. It's be solved.\n<a href=\"https://github.com/googlecolab/colabtools/issues/1470#issuecomment-668897876\">https://github.com/googlecolab/colabtools/issues/1470#issuecomment-668897876</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 958450,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-05T02:02:01.223000",
          "content": "<p><a href=\"/yuanlin08\">@yuanlin08</a> Thank you for sharing this great news! I installed TF 2.2 on Colab following the instructions found at your link and it seems to be working -- I am training my second epoch with no memory overload!  It is back to normal!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 958473,
          "author_name": "Helen",
          "author_url": "",
          "post_date": "2020-08-05T02:33:03.467000",
          "content": "<p><a href=\"/graf10a\">@graf10a</a> Great! Good to know!😄 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 958254,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-08-04T21:57:12.413000",
      "content": "<p>I face the same issue! I spent the whole morning looking for a solution, but in the end I found nothing... Towards the end of the game, I was really very frustrated..</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 958172,
      "author_name": "Phi",
      "author_url": "",
      "post_date": "2020-08-04T20:13:35.897000",
      "content": "<p>I dont have memory problem yet, but mine has weird warning logs that never seen before and the training time is almost doubled ☹️ </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 958108,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-08-04T19:17:24.107000",
      "content": "<p>Hey <a href=\"/ragnar123\">@ragnar123</a> , I was having this same issue but I did not know it was related to the Tensorflow version, probably using the previous version may solve the issue, I will try it later.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 958136,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-04T19:48:45.033000",
          "content": "<p>The problem is that you cannot just downgrade the tensorflow on colab because the TPU are compiled for the latest version only.</p>\n\n<p>I'm having the exact same issues.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 958167,
          "author_name": "Meet Ranoliya",
          "author_url": "",
          "post_date": "2020-08-04T20:06:29.850000",
          "content": "<p>Yes, and here is the <a href=\"https://github.com/tensorflow/tensorflow/issues/40622#issuecomment-654488935\">comment</a> on a similar issue on GitHub.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958210,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-04T21:04:06.377000",
          "content": "<p>damn.. bad time for us, so close to the competition end.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958250,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-08-04T21:49:34.177000",
          "content": "<p>Yea horrible timing!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 958599,
      "author_name": "Hitesh Gorana",
      "author_url": "",
      "post_date": "2020-08-05T04:33:12.760000",
      "content": "<p><code>python\n!pip install tensorflow==2.2 --force\n</code>\nthis to install Tensorflow 2.2</p>\n\n<p><code>python\n%tensorflow_version 2.2\n</code>\nthis to switch 2.2\nand don't restart the colab notebook </p>",
      "votes": 2,
      "replies": [
        {
          "id": 958649,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-05T05:20:09.970000",
          "content": "<p>thanks for the tip</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958805,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-05T06:37:02.310000",
          "content": "<p>Still have got no success with this.</p>\n\n<p>I thought you could only switch between major versions with this magic command?</p>\n\n<p>I.e.</p>\n\n<p><code>%tensorflow_version 1.x</code></p>\n\n<p>or</p>\n\n<p><code>%tensorflow_version 2.x</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958818,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-05T06:43:54.250000",
          "content": "<p><a href=\"/group16\">@group16</a> Try Gena's suggestion. It works for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958820,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-05T06:45:41.207000",
          "content": "<p>Thanks Alexey! </p>\n\n<p>I'm still getting <code>TypeError: catching classes that do not inherit from BaseException is not allowed</code> with <a href=\"/gdonchyts\">@gdonchyts</a> his solution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958849,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-05T07:11:41.333000",
          "content": "<p><a href=\"/group16\">@group16</a> This is interesting. I have not seen this error before. Have you made any changes in your code recently?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958850,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-05T07:12:23.813000",
          "content": "<p>I get this error when I try to initialize my TPU's/clear caches. Let me see if I made some mistake there...</p>\n\n<p>EDIT: It works now. I was reading the 2.3 changelog yesterday and saw that <code>tf.distribute.experimental.TPUStrategy(tpu)</code> was renamed to <code>tf.distribute.TPUStrategy(tpu)</code> but forgot to revert that change.</p>\n\n<p>Many thanks!!!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 958867,
          "author_name": "Hitesh Gorana",
          "author_url": "",
          "post_date": "2020-08-05T07:29:34.257000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1614008%2Fb4dc809390d8b9b0ab19eb340da8b464%2Fscreencapture-colab-research-google-drive-15toH1bOGjdr68bVMDx7hmYGAo-avdoV4-2020-08-05-13_02_48.png?generation=1596612902310561&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 958138,
      "author_name": "Changyi",
      "author_url": "",
      "post_date": "2020-08-04T19:49:39.737000",
      "content": "<p>Same issue. And I tried using pip install tensorflow==2.2 didn't work. \nWith this new version I have to reduce the batch size which will take me much more time in training. Any idea ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 958249,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-08-04T21:48:56.667000",
          "content": "<p>Also tried to install tf 2.2 but i believe that the tpu comes with a special driver for each version.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 959673,
      "author_name": "Ertuğrul Demir",
      "author_url": "",
      "post_date": "2020-08-05T19:54:19.083000",
      "content": "<p>A related question here didn't want to start new topic just for it. Does colab TPU's has less memory than kaggle ones or you just have to be lucky to connect decent one on colab?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 959710,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-05T20:50:13.397000",
          "content": "<p>Yes <a href=\"/datafan07\">@datafan07</a> Kaggle's TPU are v3 and Colab are v2, so you may notice that on Kaggle, you may use larger batches and it will also run the epochs faster!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 959737,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-05T21:14:41.567000",
          "content": "<p>Aaah that's what I was suspicious about thanks for clarifying! Btw do you think batch size affect the final results? In my previous experiences had different results based on batch sizes, even with tuned learning rate according to batch size, is it the case in general?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959755,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-05T21:36:32.590000",
          "content": "<p>Hey <a href=\"/datafan07\">@datafan07</a> , that can be a tricky question, at one side smaller batch sizes can reduce overfitting, but larger batches will have more samples from each class and overall more representative samples from the complete set. \nIn my experience adjusting/scalling learning rate with batch size is not so trivial, in my current configuration I am having better results with smaller sizes, but I am not so sure of the reasons, what about you?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959764,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-05T21:47:08.883000",
          "content": "<p>Same here when I go bigger batches my score decreases in my config, but have been thinking about this a while, not found exact answer to that yet, everyone has different experiences. Not even limited to this competition...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959789,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-05T22:22:33.027000",
          "content": "<p>Agreed, batch sizes are usually very dependent on task and architectures.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 959059,
      "author_name": "yeonmin",
      "author_url": "",
      "post_date": "2020-08-05T09:58:01.577000",
      "content": "<p>I am using gcp and have the same problem, so using tensorflow 2.2.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974390,
      "author_name": "Noone",
      "author_url": "",
      "post_date": "2020-08-18T00:04:09.583000",
      "content": "<p>thanks a lot.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "958072": "Sup, i know a lot of competitors are using colab to experiment. Yesterday colab update tensorflow to version 2.3. For some reason the script that i was using previously now is dying, the error is related to the memory of the tpu. If someone with more experince in this field could help us to resolve that problem i would appreciate it.\n\nCheers, and good luck with the comp ",
    "958437": "A working workaround is posted on the colab issue tracker - works perfectly!\n\n```\n!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0\nimport tensorflow as tf\nimport requests\nimport os\nresp = requests.post(\"http://{}:8475/requestversion/{}\".format(os.environ[\"COLAB_TPU_ADDR\"].split(\":\")[0], tf.__version__))\nif resp.status_code != 200:\n  print(\"Failed to switch the TPU to TF {}\".format(version))\n```",
    "958274": "I am in the same boat. It took me some time to figure out that the Tensor Flow version had been changed on Colab. So, there were a few minutes during which I was questioning my sanity and was on the verge of quitting programming 😂.",
    "958234": "The same here, the only workaround that worked for me was to decrease the batch size like more than 2x times, resulting in at least 2-3x times the training speed decrease. I've submitted the issue here - please add your complaints as well to make it more visible: https://github.com/googlecolab/colabtools/issues/1470",
    "958254": "I face the same issue! I spent the whole morning looking for a solution, but in the end I found nothing... Towards the end of the game, I was really very frustrated..",
    "958172": "I dont have memory problem yet, but mine has weird warning logs that never seen before and the training time is almost doubled ☹️ ",
    "958108": "Hey @ragnar123 , I was having this same issue but I did not know it was related to the Tensorflow version, probably using the previous version may solve the issue, I will try it later.",
    "958599": "```python\n!pip install tensorflow==2.2 --force\n```\nthis to install Tensorflow 2.2\n\n```python\n%tensorflow_version 2.2\n```\nthis to switch 2.2\nand don't restart the colab notebook ",
    "958138": "Same issue. And I tried using pip install tensorflow==2.2 didn't work. \nWith this new version I have to reduce the batch size which will take me much more time in training. Any idea ?",
    "959673": "A related question here didn't want to start new topic just for it. Does colab TPU's has less memory than kaggle ones or you just have to be lucky to connect decent one on colab?",
    "959059": "I am using gcp and have the same problem, so using tensorflow 2.2.",
    "974390": "thanks a lot."
  }
}