{
  "id": 140814,
  "title": "Have anyone tried on Colab TPU? ",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/140814",
  "author_name": "",
  "post_date": "2020-04-03T08:38:08.769970500Z",
  "votes": 9,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I've encountered weird behavior of clolab TPU, it is significantly slower than colab GPU. Have anyone faced such an issue? And also I was following their <a href=\"https://colab.research.google.com/github/tensorflow/docs/blob/master/site/en/guide/tpu.ipynb\">tutorials here</a>, is the setup ok or do I need to consider some other things? ☹️ </p>",
  "messages": [
    {
      "id": "796050",
      "postDate": "04/03/2020 08:38:08",
      "content": "<p>I've encountered weird behavior of clolab TPU, it is significantly slower than colab GPU. Have anyone faced such an issue? And also I was following their <a href=\"https://colab.research.google.com/github/tensorflow/docs/blob/master/site/en/guide/tpu.ipynb\">tutorials here</a>, is the setup ok or do I need to consider some other things? ☹️ </p>",
      "rawMarkdown": "I've encountered weird behavior of clolab TPU, it is significantly slower than colab GPU. Have anyone faced such an issue? And also I was following their [tutorials here](https://colab.research.google.com/github/tensorflow/docs/blob/master/site/en/guide/tpu.ipynb), is the setup ok or do I need to consider some other things? ☹️",
      "votes": null
    },
    {
      "id": "796055",
      "postDate": "04/03/2020 08:40:54",
      "content": "<p>Colab tpu is terribly slow\nPlease stick with kaggle kernels.\nIt will help you a lot for this competition.\nIf we can spend less time on model training over and over again and more time on exploring the data problem and designing good model then this tpu quota of kaggle SHOULD be enough to do well in this competition. Good luck buddy</p>",
      "rawMarkdown": "Colab tpu is terribly slow\nPlease stick with kaggle kernels.\nIt will help you a lot for this competition.\nIf we can spend less time on model training over and over again and more time on exploring the data problem and designing good model then this tpu quota of kaggle SHOULD be enough to do well in this competition. Good luck buddy",
      "votes": null
    },
    {
      "id": "796060",
      "postDate": "04/03/2020 08:50:16",
      "content": "<p>wow, \nthank you :) </p>",
      "rawMarkdown": "wow, \nthank you :)",
      "votes": null
    },
    {
      "id": "796575",
      "postDate": "04/03/2020 17:48:33",
      "content": "<p>However you can always use colab to experiment on <strong>base</strong> variants of models when you run out of kaggle weekly tpu hours.</p>",
      "rawMarkdown": "However you can always use colab to experiment on **base** variants of models when you run out of kaggle weekly tpu hours.",
      "votes": null
    },
    {
      "id": "796593",
      "postDate": "04/03/2020 18:12:32",
      "content": "<p>Yes, I know. Thank you. 🙂 </p>",
      "rawMarkdown": "Yes, I know. Thank you. 🙂",
      "votes": null
    },
    {
      "id": "804768",
      "postDate": "04/11/2020 23:07:34",
      "content": "<p>I asked the same question on twitter: <a href=\"https://twitter.com/JerryQu2/status/1243267971247718401?s=19\">https://twitter.com/JerryQu2/status/1243267971247718401?s=19</a></p>\n\n<p>Victor from Kaggle said:</p>\n\n<p>Kaggle uses: TPU v3-8\nColab uses: TPU v2s</p>\n\n<p>TPU v3-8: 8 cores, 128 GB HBM, 420 TFLOPs\nTPU v2-8: 8 cores, 63 GB HBM, 180 TFLOPs</p>\n\n<p>Hope that quantifies it a bit!</p>",
      "rawMarkdown": "I asked the same question on twitter: https://twitter.com/JerryQu2/status/1243267971247718401?s=19\n\nVictor from Kaggle said:\n\nKaggle uses: TPU v3-8\nColab uses: TPU v2s\n\nTPU v3-8: 8 cores, 128 GB HBM, 420 TFLOPs\nTPU v2-8: 8 cores, 63 GB HBM, 180 TFLOPs\n\nHope that quantifies it a bit!",
      "votes": null
    },
    {
      "id": "806797",
      "postDate": "04/14/2020 04:37:55",
      "content": "<p>I'm also using colab TPU. I'm struggling to put xlm-roberta-large on memory, but other than that there is no problem for me.</p>",
      "rawMarkdown": "I'm also using colab TPU. I'm struggling to put xlm-roberta-large on memory, but other than that there is no problem for me.",
      "votes": null
    },
    {
      "id": "806854",
      "postDate": "04/14/2020 06:20:29",
      "content": "<p>wow, really! I tried but it was very slow training. How much time does it take to end one epoch? </p>",
      "rawMarkdown": "wow, really! I tried but it was very slow training. How much time does it take to end one epoch?",
      "votes": null
    },
    {
      "id": "807288",
      "postDate": "04/14/2020 14:47:09",
      "content": "<p><a href=\"/ipythonx\">@ipythonx</a> I haven't use whole training set for 1 epoch, it takes around 40sec per 10000 sample. (bs=64, so 156 steps). I found somehow inference is very slow. I still don't know why and not sure it's colab only problem or not.</p>",
      "rawMarkdown": "ipythonx I haven't use whole training set for 1 epoch, it takes around 40sec per 10000 sample. (bs=64, so 156 steps). I found somehow inference is very slow. I still don't know why and not sure it's colab only problem or not.",
      "votes": null
    },
    {
      "id": "816524",
      "postDate": "04/22/2020 11:54:21",
      "content": "<p><a href=\"/bamps53\">@bamps53</a> i hope you are using all the cores in the TPU and training the model parallelly instead of only the first core..</p>",
      "rawMarkdown": "bamps53 i hope you are using all the cores in the TPU and training the model parallelly instead of only the first core..",
      "votes": null
    },
    {
      "id": "817742",
      "postDate": "04/23/2020 11:40:35",
      "content": "<p><a href=\"/bamps53\">@bamps53</a>  Have you succeed to train the large XLM roberta on Colab ? Don't you get an memory error ?</p>",
      "rawMarkdown": "bamps53  Have you succeed to train the large XLM roberta on Colab ? Don't you get an memory error ?",
      "votes": null
    },
    {
      "id": "818326",
      "postDate": "04/23/2020 19:42:41",
      "content": "<p><a href=\"/devanshc13\">@devanshc13</a> I think I'm using all the cores when training, but I'm not sure it works correctly when eval..</p>",
      "rawMarkdown": "devanshc13 I think I'm using all the cores when training, but I'm not sure it works correctly when eval..",
      "votes": null
    },
    {
      "id": "818328",
      "postDate": "04/23/2020 19:45:00",
      "content": "<p><a href=\"/ludovick\">@ludovick</a> Yes, I could. But it seems it uses almost all available memory. It's very unstable, when I do something fancier, it suddenly returns OOM. We should use TPU v3 when using roberta-large.</p>",
      "rawMarkdown": "ludovick Yes, I could. But it seems it uses almost all available memory. It's very unstable, when I do something fancier, it suddenly returns OOM. We should use TPU v3 when using roberta-large.",
      "votes": null
    },
    {
      "id": "826614",
      "postDate": "04/29/2020 18:28:19",
      "content": "<p><a href=\"/devanshc13\">@devanshc13</a> <a href=\"/bamps53\">@bamps53</a> Could you help us with some code on how to ensure the model is parallelizing over all cores on Colab TPUs? I have read about how sharding works on TF but haven't found code to make it work on Colab. Would be helpful!</p>",
      "rawMarkdown": "devanshc13 @bamps53 Could you help us with some code on how to ensure the model is parallelizing over all cores on Colab TPUs? I have read about how sharding works on TF but haven't found code to make it work on Colab. Would be helpful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 796055,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "04/03/2020 08:40:54",
      "content": "<p>Colab tpu is terribly slow\nPlease stick with kaggle kernels.\nIt will help you a lot for this competition.\nIf we can spend less time on model training over and over again and more time on exploring the data problem and designing good model then this tpu quota of kaggle SHOULD be enough to do well in this competition. Good luck buddy</p>",
      "votes": null,
      "replies": [
        {
          "id": 796060,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/03/2020 08:50:16",
          "content": "<p>wow, \nthank you :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 796575,
          "author_name": "sachinprabhu",
          "author_url": "",
          "post_date": "04/03/2020 17:48:33",
          "content": "<p>However you can always use colab to experiment on <strong>base</strong> variants of models when you run out of kaggle weekly tpu hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 796593,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/03/2020 18:12:32",
          "content": "<p>Yes, I know. Thank you. 🙂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 804768,
      "author_name": "jerryqu",
      "author_url": "",
      "post_date": "04/11/2020 23:07:34",
      "content": "<p>I asked the same question on twitter: <a href=\"https://twitter.com/JerryQu2/status/1243267971247718401?s=19\">https://twitter.com/JerryQu2/status/1243267971247718401?s=19</a></p>\n\n<p>Victor from Kaggle said:</p>\n\n<p>Kaggle uses: TPU v3-8\nColab uses: TPU v2s</p>\n\n<p>TPU v3-8: 8 cores, 128 GB HBM, 420 TFLOPs\nTPU v2-8: 8 cores, 63 GB HBM, 180 TFLOPs</p>\n\n<p>Hope that quantifies it a bit!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 806797,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "04/14/2020 04:37:55",
      "content": "<p>I'm also using colab TPU. I'm struggling to put xlm-roberta-large on memory, but other than that there is no problem for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 806854,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/14/2020 06:20:29",
          "content": "<p>wow, really! I tried but it was very slow training. How much time does it take to end one epoch? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 807288,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/14/2020 14:47:09",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> I haven't use whole training set for 1 epoch, it takes around 40sec per 10000 sample. (bs=64, so 156 steps). I found somehow inference is very slow. I still don't know why and not sure it's colab only problem or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817742,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "04/23/2020 11:40:35",
          "content": "<p><a href=\"/bamps53\">@bamps53</a>  Have you succeed to train the large XLM roberta on Colab ? Don't you get an memory error ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 818328,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/23/2020 19:45:00",
          "content": "<p><a href=\"/ludovick\">@ludovick</a> Yes, I could. But it seems it uses almost all available memory. It's very unstable, when I do something fancier, it suddenly returns OOM. We should use TPU v3 when using roberta-large.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 816524,
      "author_name": "devanshc13",
      "author_url": "",
      "post_date": "04/22/2020 11:54:21",
      "content": "<p><a href=\"/bamps53\">@bamps53</a> i hope you are using all the cores in the TPU and training the model parallelly instead of only the first core..</p>",
      "votes": null,
      "replies": [
        {
          "id": 818326,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/23/2020 19:42:41",
          "content": "<p><a href=\"/devanshc13\">@devanshc13</a> I think I'm using all the cores when training, but I'm not sure it works correctly when eval..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 826614,
          "author_name": "anjeer",
          "author_url": "",
          "post_date": "04/29/2020 18:28:19",
          "content": "<p><a href=\"/devanshc13\">@devanshc13</a> <a href=\"/bamps53\">@bamps53</a> Could you help us with some code on how to ensure the model is parallelizing over all cores on Colab TPUs? I have read about how sharding works on TF but haven't found code to make it work on Colab. Would be helpful!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "796050": "I've encountered weird behavior of clolab TPU, it is significantly slower than colab GPU. Have anyone faced such an issue? And also I was following their [tutorials here](https://colab.research.google.com/github/tensorflow/docs/blob/master/site/en/guide/tpu.ipynb), is the setup ok or do I need to consider some other things? ☹️",
    "796055": "Colab tpu is terribly slow\nPlease stick with kaggle kernels.\nIt will help you a lot for this competition.\nIf we can spend less time on model training over and over again and more time on exploring the data problem and designing good model then this tpu quota of kaggle SHOULD be enough to do well in this competition. Good luck buddy",
    "796060": "wow, \nthank you :)",
    "796575": "However you can always use colab to experiment on **base** variants of models when you run out of kaggle weekly tpu hours.",
    "796593": "Yes, I know. Thank you. 🙂",
    "804768": "I asked the same question on twitter: https://twitter.com/JerryQu2/status/1243267971247718401?s=19\n\nVictor from Kaggle said:\n\nKaggle uses: TPU v3-8\nColab uses: TPU v2s\n\nTPU v3-8: 8 cores, 128 GB HBM, 420 TFLOPs\nTPU v2-8: 8 cores, 63 GB HBM, 180 TFLOPs\n\nHope that quantifies it a bit!",
    "806797": "I'm also using colab TPU. I'm struggling to put xlm-roberta-large on memory, but other than that there is no problem for me.",
    "806854": "wow, really! I tried but it was very slow training. How much time does it take to end one epoch?",
    "807288": "ipythonx I haven't use whole training set for 1 epoch, it takes around 40sec per 10000 sample. (bs=64, so 156 steps). I found somehow inference is very slow. I still don't know why and not sure it's colab only problem or not.",
    "816524": "bamps53 i hope you are using all the cores in the TPU and training the model parallelly instead of only the first core..",
    "817742": "bamps53  Have you succeed to train the large XLM roberta on Colab ? Don't you get an memory error ?",
    "818326": "devanshc13 I think I'm using all the cores when training, but I'm not sure it works correctly when eval..",
    "818328": "ludovick Yes, I could. But it seems it uses almost all available memory. It's very unstable, when I do something fancier, it suddenly returns OOM. We should use TPU v3 when using roberta-large.",
    "826614": "devanshc13 @bamps53 Could you help us with some code on how to ensure the model is parallelizing over all cores on Colab TPUs? I have read about how sharding works on TF but haven't found code to make it work on Colab. Would be helpful!"
  },
  "source": "meta"
}