{
  "id": 154295,
  "title": "Training XLM-R large on Colab with TPU v2-8",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/154295",
  "author_name": "",
  "post_date": "2020-05-28T00:14:18.684056100Z",
  "votes": 56,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I guess I'm not the only one who always uses their weekly 30 TPU hours too early :)</p>\n\n<p>I figured out how to train XLM-R large with Colab TPU v2-8 with 8gb memory per core. Maybe it will be useful for others too.\n<a href=\"https://www.kaggle.com/riblidezso/colab-train-xlm-r-large-with-tpu-v2-8-on-colab\">https://www.kaggle.com/riblidezso/colab-train-xlm-r-large-with-tpu-v2-8-on-colab</a></p>\n\n<p>And link for shared colab notebook: <a href=\"https://colab.research.google.com/drive/1IMzxLB8kuH-ER6wNsZNHxcn7rmYU3tMe?usp=sharing\">https://colab.research.google.com/drive/1IMzxLB8kuH-ER6wNsZNHxcn7rmYU3tMe?usp=sharing</a></p>",
  "messages": [
    {
      "id": "864313",
      "postDate": "05/28/2020 00:14:18",
      "content": "<p>I guess I'm not the only one who always uses their weekly 30 TPU hours too early :)</p>\n\n<p>I figured out how to train XLM-R large with Colab TPU v2-8 with 8gb memory per core. Maybe it will be useful for others too.\n<a href=\"https://www.kaggle.com/riblidezso/colab-train-xlm-r-large-with-tpu-v2-8-on-colab\">https://www.kaggle.com/riblidezso/colab-train-xlm-r-large-with-tpu-v2-8-on-colab</a></p>\n\n<p>And link for shared colab notebook: <a href=\"https://colab.research.google.com/drive/1IMzxLB8kuH-ER6wNsZNHxcn7rmYU3tMe?usp=sharing\">https://colab.research.google.com/drive/1IMzxLB8kuH-ER6wNsZNHxcn7rmYU3tMe?usp=sharing</a></p>",
      "rawMarkdown": "I guess I'm not the only one who always uses their weekly 30 TPU hours too early :)\n\nI figured out how to train XLM-R large with Colab TPU v2-8 with 8gb memory per core. Maybe it will be useful for others too.\nhttps://www.kaggle.com/riblidezso/colab-train-xlm-r-large-with-tpu-v2-8-on-colab\n\nAnd link for shared colab notebook: https://colab.research.google.com/drive/1IMzxLB8kuH-ER6wNsZNHxcn7rmYU3tMe?usp=sharing",
      "votes": null
    },
    {
      "id": "864435",
      "postDate": "05/28/2020 02:07:20",
      "content": "<p>Amazing effort. Thanks!</p>",
      "rawMarkdown": "Amazing effort. Thanks!",
      "votes": null
    },
    {
      "id": "867102",
      "postDate": "05/30/2020 02:19:28",
      "content": "<p>Thank you <a href=\"/riblidezso\">@riblidezso</a>,  awesome👍  work!!</p>",
      "rawMarkdown": "Thank you @riblidezso,  awesome👍  work!!",
      "votes": null
    },
    {
      "id": "871157",
      "postDate": "06/02/2020 07:16:10",
      "content": "<p>Nice work. Thanks.</p>",
      "rawMarkdown": "Nice work. Thanks.",
      "votes": null
    },
    {
      "id": "873044",
      "postDate": "06/03/2020 18:11:30",
      "content": "<p><a href=\"/riblidezso\">@riblidezso</a> thanks for this wonderful starter. Just one (off-topic perhaps) things I wanted to know, do you think there can be some performance gap between the same setup from the different frameworks (tf, pt)?</p>",
      "rawMarkdown": "riblidezso thanks for this wonderful starter. Just one (off-topic perhaps) things I wanted to know, do you think there can be some performance gap between the same setup from the different frameworks (tf, pt)?",
      "votes": null
    },
    {
      "id": "875018",
      "postDate": "06/05/2020 13:10:19",
      "content": "<p>I haven't tried pytorch yet. I have seen from the shared notebooks that using TPUs is much easier with TF so I stuck with that. The custom loop is just as accurate as using the keras built in loop (it probably does the same things.)  And I have also reached very similar local validation scores with clipped SGD and Adam if I use clipped SGD for the transformer and Adam for the head layer, that still works on Colab too.</p>",
      "rawMarkdown": "I haven't tried pytorch yet. I have seen from the shared notebooks that using TPUs is much easier with TF so I stuck with that. The custom loop is just as accurate as using the keras built in loop (it probably does the same things.)  And I have also reached very similar local validation scores with clipped SGD and Adam if I use clipped SGD for the transformer and Adam for the head layer, that still works on Colab too.",
      "votes": null
    },
    {
      "id": "875037",
      "postDate": "06/05/2020 13:27:26",
      "content": "<p>Yes, thank you. I've seen your colab share. A lot of things to learn about the custom loop, thanks for sharing. However,  concerning the new <code>tf 2.x</code> and about its new functionality, I was trying to define a model with a newly sub-classing method with pre-trained weights. But got an error, <a href=\"https://stackoverflow.com/questions/62198201/tf-model-subclassing-attributionerror-nonetype-object-has-no-attribute-co\">similar to this asked question</a> in so. Is it the proper way or model sub-classing is not designed with that?</p>",
      "rawMarkdown": "Yes, thank you. I've seen your colab share. A lot of things to learn about the custom loop, thanks for sharing. However,  concerning the new `tf 2.x` and about its new functionality, I was trying to define a model with a newly sub-classing method with pre-trained weights. But got an error, [similar to this asked question](https://stackoverflow.com/questions/62198201/tf-model-subclassing-attributionerror-nonetype-object-has-no-attribute-co) in so. Is it the proper way or model sub-classing is not designed with that?",
      "votes": null
    },
    {
      "id": "875085",
      "postDate": "06/05/2020 14:08:01",
      "content": "<p>Sorry I can't help with that, I haven't tried that functionality yet.</p>",
      "rawMarkdown": "Sorry I can't help with that, I haven't tried that functionality yet.",
      "votes": null
    },
    {
      "id": "875098",
      "postDate": "06/05/2020 14:13:21",
      "content": "<p>ok, no worries. thanks anyway : ) </p>",
      "rawMarkdown": "ok, no worries. thanks anyway : )",
      "votes": null
    },
    {
      "id": "893975",
      "postDate": "06/20/2020 05:16:25",
      "content": "<p>nice one </p>",
      "rawMarkdown": "nice one",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 864435,
      "author_name": "namanj27",
      "author_url": "",
      "post_date": "05/28/2020 02:07:20",
      "content": "<p>Amazing effort. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 867102,
      "author_name": "hsinwenchang",
      "author_url": "",
      "post_date": "05/30/2020 02:19:28",
      "content": "<p>Thank you <a href=\"/riblidezso\">@riblidezso</a>,  awesome👍  work!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 871157,
      "author_name": "pratik1120",
      "author_url": "",
      "post_date": "06/02/2020 07:16:10",
      "content": "<p>Nice work. Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 873044,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "06/03/2020 18:11:30",
      "content": "<p><a href=\"/riblidezso\">@riblidezso</a> thanks for this wonderful starter. Just one (off-topic perhaps) things I wanted to know, do you think there can be some performance gap between the same setup from the different frameworks (tf, pt)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 875018,
          "author_name": "riblidezso",
          "author_url": "",
          "post_date": "06/05/2020 13:10:19",
          "content": "<p>I haven't tried pytorch yet. I have seen from the shared notebooks that using TPUs is much easier with TF so I stuck with that. The custom loop is just as accurate as using the keras built in loop (it probably does the same things.)  And I have also reached very similar local validation scores with clipped SGD and Adam if I use clipped SGD for the transformer and Adam for the head layer, that still works on Colab too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 875037,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "06/05/2020 13:27:26",
          "content": "<p>Yes, thank you. I've seen your colab share. A lot of things to learn about the custom loop, thanks for sharing. However,  concerning the new <code>tf 2.x</code> and about its new functionality, I was trying to define a model with a newly sub-classing method with pre-trained weights. But got an error, <a href=\"https://stackoverflow.com/questions/62198201/tf-model-subclassing-attributionerror-nonetype-object-has-no-attribute-co\">similar to this asked question</a> in so. Is it the proper way or model sub-classing is not designed with that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 875085,
          "author_name": "riblidezso",
          "author_url": "",
          "post_date": "06/05/2020 14:08:01",
          "content": "<p>Sorry I can't help with that, I haven't tried that functionality yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 875098,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "06/05/2020 14:13:21",
          "content": "<p>ok, no worries. thanks anyway : ) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 893975,
      "author_name": "muralidhar123",
      "author_url": "",
      "post_date": "06/20/2020 05:16:25",
      "content": "<p>nice one </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "864313": "I guess I'm not the only one who always uses their weekly 30 TPU hours too early :)\n\nI figured out how to train XLM-R large with Colab TPU v2-8 with 8gb memory per core. Maybe it will be useful for others too.\nhttps://www.kaggle.com/riblidezso/colab-train-xlm-r-large-with-tpu-v2-8-on-colab\n\nAnd link for shared colab notebook: https://colab.research.google.com/drive/1IMzxLB8kuH-ER6wNsZNHxcn7rmYU3tMe?usp=sharing",
    "864435": "Amazing effort. Thanks!",
    "867102": "Thank you @riblidezso,  awesome👍  work!!",
    "871157": "Nice work. Thanks.",
    "873044": "riblidezso thanks for this wonderful starter. Just one (off-topic perhaps) things I wanted to know, do you think there can be some performance gap between the same setup from the different frameworks (tf, pt)?",
    "875018": "I haven't tried pytorch yet. I have seen from the shared notebooks that using TPUs is much easier with TF so I stuck with that. The custom loop is just as accurate as using the keras built in loop (it probably does the same things.)  And I have also reached very similar local validation scores with clipped SGD and Adam if I use clipped SGD for the transformer and Adam for the head layer, that still works on Colab too.",
    "875037": "Yes, thank you. I've seen your colab share. A lot of things to learn about the custom loop, thanks for sharing. However,  concerning the new `tf 2.x` and about its new functionality, I was trying to define a model with a newly sub-classing method with pre-trained weights. But got an error, [similar to this asked question](https://stackoverflow.com/questions/62198201/tf-model-subclassing-attributionerror-nonetype-object-has-no-attribute-co) in so. Is it the proper way or model sub-classing is not designed with that?",
    "875085": "Sorry I can't help with that, I haven't tried that functionality yet.",
    "875098": "ok, no worries. thanks anyway : )",
    "893975": "nice one"
  },
  "source": "meta"
}