{
  "id": 124662,
  "title": "Fine-tuning pipeline for large model",
  "url": "/competitions/tensorflow2-question-answering/discussion/124662",
  "author_name": "",
  "post_date": "2020-01-05T19:02:16.979356100Z",
  "votes": 2,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Has anyone managed to get a fine-tuning pipeline going for a bert-large model or similar? I've spent considerable effort to set this up but even on the Kaggle (or Colab) GPUs it takes 20min to fine-tune on 100 examples (questions) and the maximum batch size that fits in memory is two (2)!</p>\n\n<p>The pipeline I've attempted is PyTorch / HuggingFace / bert-large-cased-whole-word-masking-finetuned-squad.</p>\n\n<p>There used to be a kernel showing how to fine-tune with distilbert, but it went OOM for me after a few hundred batches and it seems the author has taken it down.</p>\n\n<p>I guess the options are either use bert-base or use a TPU? Any 3rd option?</p>\n\n<p>Regards,\n-Alon</p>",
  "messages": [
    {
      "id": "711159",
      "postDate": "01/05/2020 19:02:16",
      "content": "<p>Has anyone managed to get a fine-tuning pipeline going for a bert-large model or similar? I've spent considerable effort to set this up but even on the Kaggle (or Colab) GPUs it takes 20min to fine-tune on 100 examples (questions) and the maximum batch size that fits in memory is two (2)!</p>\n\n<p>The pipeline I've attempted is PyTorch / HuggingFace / bert-large-cased-whole-word-masking-finetuned-squad.</p>\n\n<p>There used to be a kernel showing how to fine-tune with distilbert, but it went OOM for me after a few hundred batches and it seems the author has taken it down.</p>\n\n<p>I guess the options are either use bert-base or use a TPU? Any 3rd option?</p>\n\n<p>Regards,\n-Alon</p>",
      "rawMarkdown": "Has anyone managed to get a fine-tuning pipeline going for a bert-large model or similar? I've spent considerable effort to set this up but even on the Kaggle (or Colab) GPUs it takes 20min to fine-tune on 100 examples (questions) and the maximum batch size that fits in memory is two (2)!\n\nThe pipeline I've attempted is PyTorch / HuggingFace / bert-large-cased-whole-word-masking-finetuned-squad.\n\nThere used to be a kernel showing how to fine-tune with distilbert, but it went OOM for me after a few hundred batches and it seems the author has taken it down.\n\nI guess the options are either use bert-base or use a TPU? Any 3rd option?\n\nRegards,\n-Alon",
      "votes": null
    },
    {
      "id": "711165",
      "postDate": "01/05/2020 19:10:04",
      "content": "<p>tried XLNet large via V100 and took 70h for 1 miliion samples (1 epoch) with only batch=2, so I stopped trying because i dont have much V100 access and batch=2 might harm accuracy... </p>",
      "rawMarkdown": "tried XLNet large via V100 and took 70h for 1 miliion samples (1 epoch) with only batch=2, so I stopped trying because i dont have much V100 access and batch=2 might harm accuracy...",
      "votes": null
    },
    {
      "id": "711175",
      "postDate": "01/05/2020 19:18:33",
      "content": "<p>So - that didn't work. Did you fall back to a smaller model like bert-base? Or did you go with a TPU?</p>",
      "rawMarkdown": "So - that didn't work. Did you fall back to a smaller model like bert-base? Or did you go with a TPU?",
      "votes": null
    },
    {
      "id": "711252",
      "postDate": "01/05/2020 21:05:56",
      "content": "<p>I use Google Colab's TPU and for the storage buckets, a Google Cloud account on which I've 300$ credits in a free trial period. Using this setup, it's really easy to train BERT-joint, that is fine-tuning uncased BERT-Large on 494670 NQ examples.</p>",
      "rawMarkdown": "I use Google Colab's TPU and for the storage buckets, a Google Cloud account on which I've 300$ credits in a free trial period. Using this setup, it's really easy to train BERT-joint, that is fine-tuning uncased BERT-Large on 494670 NQ examples.",
      "votes": null
    },
    {
      "id": "711720",
      "postDate": "01/06/2020 13:26:58",
      "content": "<p><a href=\"/msheriey\">@msheriey</a>  but Google Colab's TPU memory is not enough for to train BERT-joint model right?</p>",
      "rawMarkdown": "@msheriey  but Google Colab's TPU memory is not enough for to train BERT-joint model right?",
      "votes": null
    },
    {
      "id": "711844",
      "postDate": "01/06/2020 15:19:55",
      "content": "<p>If what you mean by training BERT-joint is fine-tuning uncased BERT-Large then it's enough. Note that though the notebook itself provides you with ~12GB RAM, it gives you access to a <a href=\"https://cloud.google.com/tpu/docs/tpus\">TPU-v2-8</a> with 64 GiB of TPU memory.</p>",
      "rawMarkdown": "If what you mean by training BERT-joint is fine-tuning uncased BERT-Large then it's enough. Note that though the notebook itself provides you with ~12GB RAM, it gives you access to a [TPU-v2-8](https://cloud.google.com/tpu/docs/tpus) with 64 GiB of TPU memory.",
      "votes": null
    },
    {
      "id": "712615",
      "postDate": "01/07/2020 12:57:25",
      "content": "<p>I just made my training on GCP + TPU working. It looks like 1h for 1 epoch for bert-large-cased-whole-word-masking-finetuned-squad.</p>",
      "rawMarkdown": "I just made my training on GCP + TPU working. It looks like 1h for 1 epoch for bert-large-cased-whole-word-masking-finetuned-squad.",
      "votes": null
    },
    {
      "id": "712617",
      "postDate": "01/07/2020 12:58:50",
      "content": "<p><a href=\"/alonbochman\">@alonbochman</a> </p>\n\n<p>If you are talking about my training kernel, for OOM, please check your batch size. I also made an update for it one or two weeks ago, so please check the latest version. I am going to publish my TPU python code now.</p>\n\n<p>By the way, I didn't take any of my publish kernel down. But I don't know which kernel you are talking about.</p>",
      "rawMarkdown": "alonbochman \n\nIf you are talking about my training kernel, for OOM, please check your batch size. I also made an update for it one or two weeks ago, so please check the latest version. I am going to publish my TPU python code now.\n\nBy the way, I didn't take any of my publish kernel down. But I don't know which kernel you are talking about.",
      "votes": null
    },
    {
      "id": "712620",
      "postDate": "01/07/2020 13:00:01",
      "content": "<p>1 hour for 1 epoch of bert-large is too fast. It took me 5-6 hours to train 1 epoch for bert-large</p>",
      "rawMarkdown": "1 hour for 1 epoch of bert-large is too fast. It took me 5-6 hours to train 1 epoch for bert-large",
      "votes": null
    },
    {
      "id": "712621",
      "postDate": "01/07/2020 13:02:02",
      "content": "<p><a href=\"/alonbochman\">@alonbochman</a> </p>\n\n<p>Sorry, I did took my training down accidentally. It's up now!</p>",
      "rawMarkdown": "alonbochman \n\nSorry, I did took my training down accidentally. It's up now!",
      "votes": null
    },
    {
      "id": "712622",
      "postDate": "01/07/2020 13:03:37",
      "content": "<p>what's your GLOBAL batch size? I used TPU v3-8, and my batch_size is 8 * 8 (one 8 is the no. of replicat) = 64. Also, what's your no. of training example in TF record file? I have only about 500K training examples.</p>",
      "rawMarkdown": "what's your GLOBAL batch size? I used TPU v3-8, and my batch_size is 8 * 8 (one 8 is the no. of replicat) = 64. Also, what's your no. of training example in TF record file? I have only about 500K training examples.",
      "votes": null
    },
    {
      "id": "712628",
      "postDate": "01/07/2020 13:11:02",
      "content": "<p>I just published my TPU script.</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952\">https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952</a></p>",
      "rawMarkdown": "I just published my TPU script.\n\n[https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952](https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952)",
      "votes": null
    },
    {
      "id": "712631",
      "postDate": "01/07/2020 13:13:34",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "712635",
      "postDate": "01/07/2020 13:16:06",
      "content": "<p><a href=\"/axel81\">@axel81</a> , are you using my kernel with batch accumulation? In my TPU script, I removed this part, and my batch accumulation code is not very well optimized. </p>",
      "rawMarkdown": "axel81 , are you using my kernel with batch accumulation? In my TPU script, I removed this part, and my batch accumulation code is not very well optimized.",
      "votes": null
    },
    {
      "id": "712637",
      "postDate": "01/07/2020 13:19:28",
      "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> I am using without batch accumulation and I've modified some part of it. If you can suggest a optimization it would be great. Btw the max batch size I can use is 8 on TPU</p>",
      "rawMarkdown": "yihdarshieh I am using without batch accumulation and I've modified some part of it. If you can suggest a optimization it would be great. Btw the max batch size I can use is 8 on TPU",
      "votes": null
    },
    {
      "id": "712643",
      "postDate": "01/07/2020 13:22:09",
      "content": "<p>You can first look my tpu python script. If really not working for you, we can discuss later here.</p>",
      "rawMarkdown": "You can first look my tpu python script. If really not working for you, we can discuss later here.",
      "votes": null
    },
    {
      "id": "712644",
      "postDate": "01/07/2020 13:23:32",
      "content": "<p>Have you made it public?</p>",
      "rawMarkdown": "Have you made it public?",
      "votes": null
    },
    {
      "id": "712647",
      "postDate": "01/07/2020 13:26:34",
      "content": "<p>yes. </p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952\">https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952</a></p>",
      "rawMarkdown": "yes. \n\n\n[https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952](https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952)",
      "votes": null
    },
    {
      "id": "712651",
      "postDate": "01/07/2020 13:27:42",
      "content": "<p>Here model_name is <code>distill-bert</code>, are you sure you ran bert-large on TPU?</p>",
      "rawMarkdown": "Here model_name is `distill-bert`, are you sure you ran bert-large on TPU?",
      "votes": null
    },
    {
      "id": "712655",
      "postDate": "01/07/2020 13:35:09",
      "content": "<p>Yes</p>",
      "rawMarkdown": "Yes",
      "votes": null
    },
    {
      "id": "713009",
      "postDate": "01/07/2020 20:08:46",
      "content": "<p><a href=\"/axel81\">@axel81</a> ,</p>\n\n<p>Did you use <code>train_dist_dataset = tpu_strategy.experimental_distribute_dataset(train_dataset)</code> in your own tpu code? If not, it's probably the reason why your training is slow, and why you can only have batch_size = 8.</p>",
      "rawMarkdown": "axel81 ,\n\nDid you use `train_dist_dataset = tpu_strategy.experimental_distribute_dataset(train_dataset)` in your own tpu code? If not, it's probably the reason why your training is slow, and why you can only have batch_size = 8.",
      "votes": null
    },
    {
      "id": "713250",
      "postDate": "01/08/2020 04:30:55",
      "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> yes I am using <code>experimental_distribute_dataset</code>. My TPU code is mostly similar to your public kernel</p>",
      "rawMarkdown": "yihdarshieh yes I am using `experimental_distribute_dataset`. My TPU code is mostly similar to your public kernel",
      "votes": null
    },
    {
      "id": "713363",
      "postDate": "01/08/2020 07:42:50",
      "content": "<p>Ok, that's weird. Anyway, I am using GCP  VM with tpu node v3-8</p>",
      "rawMarkdown": "Ok, that's weird. Anyway, I am using GCP  VM with tpu node v3-8",
      "votes": null
    },
    {
      "id": "713395",
      "postDate": "01/08/2020 08:40:56",
      "content": "<p>I think public TPU on colab free tier is different. It is not helping in finetuning :(</p>",
      "rawMarkdown": "I think public TPU on colab free tier is different. It is not helping in finetuning :(",
      "votes": null
    },
    {
      "id": "717038",
      "postDate": "01/12/2020 16:28:18",
      "content": "<p>Quick question: can I get Bert-Large trained on your TPU script? Thank you!</p>",
      "rawMarkdown": "Quick question: can I get Bert-Large trained on your TPU script? Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 711165,
      "author_name": "httpwwwfszyc",
      "author_url": "",
      "post_date": "01/05/2020 19:10:04",
      "content": "<p>tried XLNet large via V100 and took 70h for 1 miliion samples (1 epoch) with only batch=2, so I stopped trying because i dont have much V100 access and batch=2 might harm accuracy... </p>",
      "votes": null,
      "replies": [
        {
          "id": 711175,
          "author_name": "alonbochman",
          "author_url": "",
          "post_date": "01/05/2020 19:18:33",
          "content": "<p>So - that didn't work. Did you fall back to a smaller model like bert-base? Or did you go with a TPU?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 711252,
      "author_name": "msheriey",
      "author_url": "",
      "post_date": "01/05/2020 21:05:56",
      "content": "<p>I use Google Colab's TPU and for the storage buckets, a Google Cloud account on which I've 300$ credits in a free trial period. Using this setup, it's really easy to train BERT-joint, that is fine-tuning uncased BERT-Large on 494670 NQ examples.</p>",
      "votes": null,
      "replies": [
        {
          "id": 711720,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/06/2020 13:26:58",
          "content": "<p><a href=\"/msheriey\">@msheriey</a>  but Google Colab's TPU memory is not enough for to train BERT-joint model right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 711844,
          "author_name": "msheriey",
          "author_url": "",
          "post_date": "01/06/2020 15:19:55",
          "content": "<p>If what you mean by training BERT-joint is fine-tuning uncased BERT-Large then it's enough. Note that though the notebook itself provides you with ~12GB RAM, it gives you access to a <a href=\"https://cloud.google.com/tpu/docs/tpus\">TPU-v2-8</a> with 64 GiB of TPU memory.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 712615,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/07/2020 12:57:25",
      "content": "<p>I just made my training on GCP + TPU working. It looks like 1h for 1 epoch for bert-large-cased-whole-word-masking-finetuned-squad.</p>",
      "votes": null,
      "replies": [
        {
          "id": 712620,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/07/2020 13:00:01",
          "content": "<p>1 hour for 1 epoch of bert-large is too fast. It took me 5-6 hours to train 1 epoch for bert-large</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712622,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/07/2020 13:03:37",
          "content": "<p>what's your GLOBAL batch size? I used TPU v3-8, and my batch_size is 8 * 8 (one 8 is the no. of replicat) = 64. Also, what's your no. of training example in TF record file? I have only about 500K training examples.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712635,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/07/2020 13:16:06",
          "content": "<p><a href=\"/axel81\">@axel81</a> , are you using my kernel with batch accumulation? In my TPU script, I removed this part, and my batch accumulation code is not very well optimized. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712637,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/07/2020 13:19:28",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> I am using without batch accumulation and I've modified some part of it. If you can suggest a optimization it would be great. Btw the max batch size I can use is 8 on TPU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712643,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/07/2020 13:22:09",
          "content": "<p>You can first look my tpu python script. If really not working for you, we can discuss later here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712644,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/07/2020 13:23:32",
          "content": "<p>Have you made it public?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712647,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/07/2020 13:26:34",
          "content": "<p>yes. </p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952\">https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712651,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/07/2020 13:27:42",
          "content": "<p>Here model_name is <code>distill-bert</code>, are you sure you ran bert-large on TPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712655,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/07/2020 13:35:09",
          "content": "<p>Yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713009,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/07/2020 20:08:46",
          "content": "<p><a href=\"/axel81\">@axel81</a> ,</p>\n\n<p>Did you use <code>train_dist_dataset = tpu_strategy.experimental_distribute_dataset(train_dataset)</code> in your own tpu code? If not, it's probably the reason why your training is slow, and why you can only have batch_size = 8.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713250,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/08/2020 04:30:55",
          "content": "<p><a href=\"/yihdarshieh\">@yihdarshieh</a> yes I am using <code>experimental_distribute_dataset</code>. My TPU code is mostly similar to your public kernel</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713363,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/08/2020 07:42:50",
          "content": "<p>Ok, that's weird. Anyway, I am using GCP  VM with tpu node v3-8</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713395,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/08/2020 08:40:56",
          "content": "<p>I think public TPU on colab free tier is different. It is not helping in finetuning :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 712617,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/07/2020 12:58:50",
      "content": "<p><a href=\"/alonbochman\">@alonbochman</a> </p>\n\n<p>If you are talking about my training kernel, for OOM, please check your batch size. I also made an update for it one or two weeks ago, so please check the latest version. I am going to publish my TPU python code now.</p>\n\n<p>By the way, I didn't take any of my publish kernel down. But I don't know which kernel you are talking about.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 712621,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/07/2020 13:02:02",
      "content": "<p><a href=\"/alonbochman\">@alonbochman</a> </p>\n\n<p>Sorry, I did took my training down accidentally. It's up now!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 712628,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/07/2020 13:11:02",
      "content": "<p>I just published my TPU script.</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952\">https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 712631,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/07/2020 13:13:34",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 717038,
          "author_name": "alluxia",
          "author_url": "",
          "post_date": "01/12/2020 16:28:18",
          "content": "<p>Quick question: can I get Bert-Large trained on your TPU script? Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "711159": "Has anyone managed to get a fine-tuning pipeline going for a bert-large model or similar? I've spent considerable effort to set this up but even on the Kaggle (or Colab) GPUs it takes 20min to fine-tune on 100 examples (questions) and the maximum batch size that fits in memory is two (2)!\n\nThe pipeline I've attempted is PyTorch / HuggingFace / bert-large-cased-whole-word-masking-finetuned-squad.\n\nThere used to be a kernel showing how to fine-tune with distilbert, but it went OOM for me after a few hundred batches and it seems the author has taken it down.\n\nI guess the options are either use bert-base or use a TPU? Any 3rd option?\n\nRegards,\n-Alon",
    "711165": "tried XLNet large via V100 and took 70h for 1 miliion samples (1 epoch) with only batch=2, so I stopped trying because i dont have much V100 access and batch=2 might harm accuracy...",
    "711175": "So - that didn't work. Did you fall back to a smaller model like bert-base? Or did you go with a TPU?",
    "711252": "I use Google Colab's TPU and for the storage buckets, a Google Cloud account on which I've 300$ credits in a free trial period. Using this setup, it's really easy to train BERT-joint, that is fine-tuning uncased BERT-Large on 494670 NQ examples.",
    "711720": "@msheriey  but Google Colab's TPU memory is not enough for to train BERT-joint model right?",
    "711844": "If what you mean by training BERT-joint is fine-tuning uncased BERT-Large then it's enough. Note that though the notebook itself provides you with ~12GB RAM, it gives you access to a [TPU-v2-8](https://cloud.google.com/tpu/docs/tpus) with 64 GiB of TPU memory.",
    "712615": "I just made my training on GCP + TPU working. It looks like 1h for 1 epoch for bert-large-cased-whole-word-masking-finetuned-squad.",
    "712617": "alonbochman \n\nIf you are talking about my training kernel, for OOM, please check your batch size. I also made an update for it one or two weeks ago, so please check the latest version. I am going to publish my TPU python code now.\n\nBy the way, I didn't take any of my publish kernel down. But I don't know which kernel you are talking about.",
    "712620": "1 hour for 1 epoch of bert-large is too fast. It took me 5-6 hours to train 1 epoch for bert-large",
    "712621": "alonbochman \n\nSorry, I did took my training down accidentally. It's up now!",
    "712622": "what's your GLOBAL batch size? I used TPU v3-8, and my batch_size is 8 * 8 (one 8 is the no. of replicat) = 64. Also, what's your no. of training example in TF record file? I have only about 500K training examples.",
    "712628": "I just published my TPU script.\n\n[https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952](https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952)",
    "712631": "",
    "712635": "axel81 , are you using my kernel with batch accumulation? In my TPU script, I removed this part, and my batch accumulation code is not very well optimized.",
    "712637": "yihdarshieh I am using without batch accumulation and I've modified some part of it. If you can suggest a optimization it would be great. Btw the max batch size I can use is 8 on TPU",
    "712643": "You can first look my tpu python script. If really not working for you, we can discuss later here.",
    "712644": "Have you made it public?",
    "712647": "yes. \n\n\n[https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952](https://www.kaggle.com/yihdarshieh/tf2-training-on-gcp-tpu?scriptVersionId=26464952)",
    "712651": "Here model_name is `distill-bert`, are you sure you ran bert-large on TPU?",
    "712655": "Yes",
    "713009": "axel81 ,\n\nDid you use `train_dist_dataset = tpu_strategy.experimental_distribute_dataset(train_dataset)` in your own tpu code? If not, it's probably the reason why your training is slow, and why you can only have batch_size = 8.",
    "713250": "yihdarshieh yes I am using `experimental_distribute_dataset`. My TPU code is mostly similar to your public kernel",
    "713363": "Ok, that's weird. Anyway, I am using GCP  VM with tpu node v3-8",
    "713395": "I think public TPU on colab free tier is different. It is not helping in finetuning :(",
    "717038": "Quick question: can I get Bert-Large trained on your TPU script? Thank you!"
  },
  "source": "meta"
}