{
  "id": 114835,
  "title": "GPU vs TPU for BERT",
  "url": "/competitions/tensorflow2-question-answering/discussion/114835",
  "author_name": "",
  "post_date": "2019-10-29T15:30:32.678246Z",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p>BERT performance compared for two different architectures.</p>\n\n<p><a href=\"https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert\">https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert</a></p>\n\n<p>Conclusion from the article</p>\n\n<blockquote>\n  <p>TPUs are about 32% to 54% faster for training BERT-like models. One can expect to replicate BERT base on an 8 GPU machine within about 10 to 17 days. On a standard, affordable GPU machine with 4 GPUs one can expect to train BERT base for about 34 days using 16-bit or about 11 days using 8-bit.</p>\n</blockquote>",
  "messages": [
    {
      "id": "660779",
      "postDate": "10/29/2019 15:30:32",
      "content": "<p>BERT performance compared for two different architectures.</p>\n\n<p><a href=\"https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert\">https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert</a></p>\n\n<p>Conclusion from the article</p>\n\n<blockquote>\n  <p>TPUs are about 32% to 54% faster for training BERT-like models. One can expect to replicate BERT base on an 8 GPU machine within about 10 to 17 days. On a standard, affordable GPU machine with 4 GPUs one can expect to train BERT base for about 34 days using 16-bit or about 11 days using 8-bit.</p>\n</blockquote>",
      "rawMarkdown": "BERT performance compared for two different architectures.\n\nhttps://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert\n\nConclusion from the article\n&gt;TPUs are about 32% to 54% faster for training BERT-like models. One can expect to replicate BERT base on an 8 GPU machine within about 10 to 17 days. On a standard, affordable GPU machine with 4 GPUs one can expect to train BERT base for about 34 days using 16-bit or about 11 days using 8-bit.",
      "votes": null
    },
    {
      "id": "660782",
      "postDate": "10/29/2019 15:37:08",
      "content": "<p>But TF2.0 doesn't support for TPU now.</p>",
      "rawMarkdown": "But TF2.0 doesn't support for TPU now.",
      "votes": null
    },
    {
      "id": "660818",
      "postDate": "10/29/2019 16:42:27",
      "content": "<p>Supporting version would be under development, so better to know how it may perform.</p>",
      "rawMarkdown": "Supporting version would be under development, so better to know how it may perform.",
      "votes": null
    },
    {
      "id": "660827",
      "postDate": "10/29/2019 16:57:23",
      "content": "<p>Based on what I know, the support for TPU will come out in the next version, Tensorflow 2.1, but the specified version in this competition is 2.0, I'm not sure if TF2.1 is allowed.</p>",
      "rawMarkdown": "Based on what I know, the support for TPU will come out in the next version, Tensorflow 2.1, but the specified version in this competition is 2.0, I'm not sure if TF2.1 is allowed.",
      "votes": null
    },
    {
      "id": "662634",
      "postDate": "10/31/2019 19:32:20",
      "content": "<p>When TensorFlow 2.1 is released with TPU support, you are certainly welcome to use it.</p>",
      "rawMarkdown": "When TensorFlow 2.1 is released with TPU support, you are certainly welcome to use it.",
      "votes": null
    },
    {
      "id": "675686",
      "postDate": "11/18/2019 12:22:51",
      "content": "<p>But for this competition, I think a few hours on a GPU is enough to fine tune BERT. So we not worry about training the BERT LM from scratch.</p>",
      "rawMarkdown": "But for this competition, I think a few hours on a GPU is enough to fine tune BERT. So we not worry about training the BERT LM from scratch.",
      "votes": null
    },
    {
      "id": "675832",
      "postDate": "11/18/2019 16:18:10",
      "content": "<p>I agree it's good that we don't need to train it from scratch, but fine-tuning in a few hours... that sounds optimistic :)</p>",
      "rawMarkdown": "I agree it's good that we don't need to train it from scratch, but fine-tuning in a few hours... that sounds optimistic :)",
      "votes": null
    },
    {
      "id": "676351",
      "postDate": "11/19/2019 05:26:02",
      "content": "<p>I got this information from the original BERT paper, sec. 3.2, </p>\n\n<blockquote>\n  <p>Compared  to  pre-training,  fine-tuning  is  relatively  inexpensive.   All  of  the  results  in  the  paper can be replicated in at most 1 hour on a single Cloud TPU, or a few hours on a GPU, starting from the exact same pre-trained model.</p>\n</blockquote>",
      "rawMarkdown": "I got this information from the original BERT paper, sec. 3.2, \n&gt; Compared  to  pre-training,  fine-tuning  is  relatively  inexpensive.   All  of  the  results  in  the  paper can be replicated in at most 1 hour on a single Cloud TPU, or a few hours on a GPU, starting from the exact same pre-trained model.",
      "votes": null
    },
    {
      "id": "676477",
      "postDate": "11/19/2019 08:11:06",
      "content": "<p>The Dataset used in the original paper BERT such as SQUAD is much smaller than the dataset used in this competition, I have tried the fine-tuning experiment in <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">bert-joint</a>, one epoch spends about 2 hours with TPU-v2. Although you can get the BERT fine-tuned checkpoint for this QA task, but if you want to get a good rank in this competition, you should try to use other models such as XLnet, Roberta, ALBERT. So if you can use TPU, it will be very helpful.</p>",
      "rawMarkdown": "The Dataset used in the original paper BERT such as SQUAD is much smaller than the dataset used in this competition, I have tried the fine-tuning experiment in [bert-joint](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint), one epoch spends about 2 hours with TPU-v2. Although you can get the BERT fine-tuned checkpoint for this QA task, but if you want to get a good rank in this competition, you should try to use other models such as XLnet, Roberta, ALBERT. So if you can use TPU, it will be very helpful.",
      "votes": null
    },
    {
      "id": "676538",
      "postDate": "11/19/2019 09:24:29",
      "content": "<p>Thanks for the clarification. Useful info.</p>",
      "rawMarkdown": "Thanks for the clarification. Useful info.",
      "votes": null
    },
    {
      "id": "676558",
      "postDate": "11/19/2019 09:51:23",
      "content": "<p>It may be not too bad: with bert-base-uncased one epoch (going over all questions, with just one candidate answer for each question) takes 1h 15m on 2080 ti, using pytorch + apex O1, predicting only long answers. So it may be possible to train something reasonable in a few hours indeed, but probably achieving best performance even for bert-base-uncased would take more time.</p>",
      "rawMarkdown": "It may be not too bad: with bert-base-uncased one epoch (going over all questions, with just one candidate answer for each question) takes 1h 15m on 2080 ti, using pytorch + apex O1, predicting only long answers. So it may be possible to train something reasonable in a few hours indeed, but probably achieving best performance even for bert-base-uncased would take more time.",
      "votes": null
    },
    {
      "id": "676658",
      "postDate": "11/19/2019 11:53:58",
      "content": "<p>Thanks for the info.</p>",
      "rawMarkdown": "Thanks for the info.",
      "votes": null
    },
    {
      "id": "695360",
      "postDate": "12/15/2019 02:37:31",
      "content": "<p>Would we still be eligible for the special prize if we use Tensorflow 2.1 instead of 2.0\n&gt; TensorFlow 2.0 Prizes:\nFirst Prize: $12,000\nSecond Prize: $8,000\nThird Prize: $5,000\nTo be eligible for the TensorFlow 2.0 prizes, your code must use TensorFlow 2.0. More specifically:\nit should run with a public release of TensorFlow 2.0 installed, and\nshould not use any tf.compat.v1 module symbols and should not raise deprecation warnings</p>",
      "rawMarkdown": "Would we still be eligible for the special prize if we use Tensorflow 2.1 instead of 2.0\n&gt; TensorFlow 2.0 Prizes:\nFirst Prize: $12,000\nSecond Prize: $8,000\nThird Prize: $5,000\nTo be eligible for the TensorFlow 2.0 prizes, your code must use TensorFlow 2.0. More specifically:\nit should run with a public release of TensorFlow 2.0 installed, and\nshould not use any tf.compat.v1 module symbols and should not raise deprecation warnings",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 660782,
      "author_name": "guozhiyu0914",
      "author_url": "",
      "post_date": "10/29/2019 15:37:08",
      "content": "<p>But TF2.0 doesn't support for TPU now.</p>",
      "votes": null,
      "replies": [
        {
          "id": 660818,
          "author_name": "cyberia",
          "author_url": "",
          "post_date": "10/29/2019 16:42:27",
          "content": "<p>Supporting version would be under development, so better to know how it may perform.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 660827,
          "author_name": "guozhiyu0914",
          "author_url": "",
          "post_date": "10/29/2019 16:57:23",
          "content": "<p>Based on what I know, the support for TPU will come out in the next version, Tensorflow 2.1, but the specified version in this competition is 2.0, I'm not sure if TF2.1 is allowed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662634,
          "author_name": "dynamicwebpaige",
          "author_url": "",
          "post_date": "10/31/2019 19:32:20",
          "content": "<p>When TensorFlow 2.1 is released with TPU support, you are certainly welcome to use it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 695360,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "12/15/2019 02:37:31",
          "content": "<p>Would we still be eligible for the special prize if we use Tensorflow 2.1 instead of 2.0\n&gt; TensorFlow 2.0 Prizes:\nFirst Prize: $12,000\nSecond Prize: $8,000\nThird Prize: $5,000\nTo be eligible for the TensorFlow 2.0 prizes, your code must use TensorFlow 2.0. More specifically:\nit should run with a public release of TensorFlow 2.0 installed, and\nshould not use any tf.compat.v1 module symbols and should not raise deprecation warnings</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 675686,
      "author_name": "arvindpdmn",
      "author_url": "",
      "post_date": "11/18/2019 12:22:51",
      "content": "<p>But for this competition, I think a few hours on a GPU is enough to fine tune BERT. So we not worry about training the BERT LM from scratch.</p>",
      "votes": null,
      "replies": [
        {
          "id": 675832,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "11/18/2019 16:18:10",
          "content": "<p>I agree it's good that we don't need to train it from scratch, but fine-tuning in a few hours... that sounds optimistic :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 676351,
          "author_name": "arvindpdmn",
          "author_url": "",
          "post_date": "11/19/2019 05:26:02",
          "content": "<p>I got this information from the original BERT paper, sec. 3.2, </p>\n\n<blockquote>\n  <p>Compared  to  pre-training,  fine-tuning  is  relatively  inexpensive.   All  of  the  results  in  the  paper can be replicated in at most 1 hour on a single Cloud TPU, or a few hours on a GPU, starting from the exact same pre-trained model.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 676477,
          "author_name": "guozhiyu0914",
          "author_url": "",
          "post_date": "11/19/2019 08:11:06",
          "content": "<p>The Dataset used in the original paper BERT such as SQUAD is much smaller than the dataset used in this competition, I have tried the fine-tuning experiment in <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">bert-joint</a>, one epoch spends about 2 hours with TPU-v2. Although you can get the BERT fine-tuned checkpoint for this QA task, but if you want to get a good rank in this competition, you should try to use other models such as XLnet, Roberta, ALBERT. So if you can use TPU, it will be very helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 676538,
          "author_name": "arvindpdmn",
          "author_url": "",
          "post_date": "11/19/2019 09:24:29",
          "content": "<p>Thanks for the clarification. Useful info.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 676558,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "11/19/2019 09:51:23",
          "content": "<p>It may be not too bad: with bert-base-uncased one epoch (going over all questions, with just one candidate answer for each question) takes 1h 15m on 2080 ti, using pytorch + apex O1, predicting only long answers. So it may be possible to train something reasonable in a few hours indeed, but probably achieving best performance even for bert-base-uncased would take more time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 676658,
      "author_name": "rahulloha",
      "author_url": "",
      "post_date": "11/19/2019 11:53:58",
      "content": "<p>Thanks for the info.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "660779": "BERT performance compared for two different architectures.\n\nhttps://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert\n\nConclusion from the article\n&gt;TPUs are about 32% to 54% faster for training BERT-like models. One can expect to replicate BERT base on an 8 GPU machine within about 10 to 17 days. On a standard, affordable GPU machine with 4 GPUs one can expect to train BERT base for about 34 days using 16-bit or about 11 days using 8-bit.",
    "660782": "But TF2.0 doesn't support for TPU now.",
    "660818": "Supporting version would be under development, so better to know how it may perform.",
    "660827": "Based on what I know, the support for TPU will come out in the next version, Tensorflow 2.1, but the specified version in this competition is 2.0, I'm not sure if TF2.1 is allowed.",
    "662634": "When TensorFlow 2.1 is released with TPU support, you are certainly welcome to use it.",
    "675686": "But for this competition, I think a few hours on a GPU is enough to fine tune BERT. So we not worry about training the BERT LM from scratch.",
    "675832": "I agree it's good that we don't need to train it from scratch, but fine-tuning in a few hours... that sounds optimistic :)",
    "676351": "I got this information from the original BERT paper, sec. 3.2, \n&gt; Compared  to  pre-training,  fine-tuning  is  relatively  inexpensive.   All  of  the  results  in  the  paper can be replicated in at most 1 hour on a single Cloud TPU, or a few hours on a GPU, starting from the exact same pre-trained model.",
    "676477": "The Dataset used in the original paper BERT such as SQUAD is much smaller than the dataset used in this competition, I have tried the fine-tuning experiment in [bert-joint](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint), one epoch spends about 2 hours with TPU-v2. Although you can get the BERT fine-tuned checkpoint for this QA task, but if you want to get a good rank in this competition, you should try to use other models such as XLnet, Roberta, ALBERT. So if you can use TPU, it will be very helpful.",
    "676538": "Thanks for the clarification. Useful info.",
    "676558": "It may be not too bad: with bert-base-uncased one epoch (going over all questions, with just one candidate answer for each question) takes 1h 15m on 2080 ti, using pytorch + apex O1, predicting only long answers. So it may be possible to train something reasonable in a few hours indeed, but probably achieving best performance even for bert-base-uncased would take more time.",
    "676658": "Thanks for the info.",
    "695360": "Would we still be eligible for the special prize if we use Tensorflow 2.1 instead of 2.0\n&gt; TensorFlow 2.0 Prizes:\nFirst Prize: $12,000\nSecond Prize: $8,000\nThird Prize: $5,000\nTo be eligible for the TensorFlow 2.0 prizes, your code must use TensorFlow 2.0. More specifically:\nit should run with a public release of TensorFlow 2.0 installed, and\nshould not use any tf.compat.v1 module symbols and should not raise deprecation warnings"
  },
  "source": "meta"
}