{
  "id": 140022,
  "title": "Need for speed",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/140022",
  "author_name": "Martin Görner",
  "post_date": "2020-03-31T00:31:50.307000",
  "votes": 34,
  "comment_count": 53,
  "views": 0,
  "content": "<p>BERT is a large model, requiring powerful hardware to iterate fast.\nI'll be posting various benchmarks numbers here to put some numbers behind \"fast\".</p>",
  "messages": [
    {
      "id": 792177,
      "postDate": "2020-03-31T00:31:50.307Z",
      "content": "<p>BERT is a large model, requiring powerful hardware to iterate fast.\nI'll be posting various benchmarks numbers here to put some numbers behind \"fast\".</p>",
      "rawMarkdown": "BERT is a large model, requiring powerful hardware to iterate fast.\nI'll be posting various benchmarks numbers here to put some numbers behind \"fast\".",
      "votes": 34
    },
    {
      "id": 792675,
      "postDate": "2020-03-31T12:58:50.560Z",
      "content": "<p>The speed gains are indeed impressive! What bothers me is that again there are many abstraction layers that don't let me understand the details of the running execution in a level I would love to. I regularly get some \"process died\" messages without any stacktrace, or some weird memory issues. It is very cumbersome to make it work.</p>\n\n<p>That said though, I only tried pytorch and I know that it is not optimized for that yet. Is debugging way easier when using TF / Keras + TPUs?</p>",
      "rawMarkdown": "The speed gains are indeed impressive! What bothers me is that again there are many abstraction layers that don't let me understand the details of the running execution in a level I would love to. I regularly get some \"process died\" messages without any stacktrace, or some weird memory issues. It is very cumbersome to make it work.\n\nThat said though, I only tried pytorch and I know that it is not optimized for that yet. Is debugging way easier when using TF / Keras + TPUs?",
      "votes": 6,
      "replies": [
        {
          "id": 792712,
          "postDate": "2020-03-31T13:56:55.013Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I've made many experiments on the this and the Flowers competitions and I can say that it seems to be easier to use TPUs with TF, but the error message doesn't really help, I got a lot of memory issues, but the log almost never points that, I've been able to solve most with the help of Martin and trial &amp; error.</p>",
          "rawMarkdown": "@philippsinger I've made many experiments on the this and the Flowers competitions and I can say that it seems to be easier to use TPUs with TF, but the error message doesn't really help, I got a lot of memory issues, but the log almost never points that, I've been able to solve most with the help of Martin and trial &amp; error.",
          "votes": 1
        },
        {
          "id": 792918,
          "postDate": "2020-03-31T16:35:19.053Z",
          "content": "<p>With PyTorch, are you initializing the TPU as recommended through</p>\n\n<p><code>\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code></p>\n\n<p>I have seem samples lying around that are missing this. If you dig into the script, you will notice a call to a specific port on the TPU (8475). This switches the TPU firmware to the latest nightly version for PyTorch. If you miss this step, your setup will still work but with bugs. This might explain some of the instability you have been experiencing.</p>",
          "rawMarkdown": "With PyTorch, are you initializing the TPU as recommended through\n\n```\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n```\n\nI have seem samples lying around that are missing this. If you dig into the script, you will notice a call to a specific port on the TPU (8475). This switches the TPU firmware to the latest nightly version for PyTorch. If you miss this step, your setup will still work but with bugs. This might explain some of the instability you have been experiencing.",
          "votes": 3
        },
        {
          "id": 792921,
          "postDate": "2020-03-31T16:42:14.760Z",
          "content": "<p>Yeah, am exactly running these commands.</p>",
          "rawMarkdown": "Yeah, am exactly running these commands."
        },
        {
          "id": 793107,
          "postDate": "2020-03-31T20:00:15.123Z",
          "content": "<p>If you can put together a reproducible crash in PyTorch, please do so, post it on the <a href=\"https://github.com/pytorch/xla\">PyTorch/XLA </a>GitHub and link it here so that I see it too.</p>",
          "rawMarkdown": "If you can put together a reproducible crash in PyTorch, please do so, post it on the [PyTorch/XLA ](https://github.com/pytorch/xla)GitHub and link it here so that I see it too."
        }
      ]
    },
    {
      "id": 794058,
      "postDate": "2020-04-01T14:20:44.027Z",
      "content": "<p>Very nice benchmarks! Have you tried 32gb GPUs, and nvlinked gpus?</p>",
      "rawMarkdown": "Very nice benchmarks! Have you tried 32gb GPUs, and nvlinked gpus?",
      "votes": 1,
      "replies": [
        {
          "id": 794075,
          "postDate": "2020-04-01T14:34:26.433Z",
          "content": "<p>Is nvlink really useful? Was thinking for a while to set it up.</p>",
          "rawMarkdown": "Is nvlink really useful? Was thinking for a while to set it up."
        },
        {
          "id": 794138,
          "postDate": "2020-04-01T15:39:02.260Z",
          "content": "<p>I never tried, but i thought that nvlink would let you run a model on both GPUs without any need for parallel controls?</p>",
          "rawMarkdown": "I never tried, but i thought that nvlink would let you run a model on both GPUs without any need for parallel controls?"
        },
        {
          "id": 794142,
          "postDate": "2020-04-01T15:43:22.767Z",
          "content": "<p>Yes, as I said above I tried 4xV100 and 8xV100 machines on GCP. They have NVLink (launch <a href=\"https://cloud.google.com/blog/products/compute/tesla-v100-gpus-are-now-generally-available\">announcement here</a>). But so far I got catastrophic results there. The model ran slower than on a single V100. Probably a mistake somewhere. If you guys have access to these kinds of machines, you can try if you get better results. The <a href=\"https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert\">benchmark notebook</a> is configured for multi-GPU with MirroredStrategy.</p>",
          "rawMarkdown": "Yes, as I said above I tried 4xV100 and 8xV100 machines on GCP. They have NVLink (launch [announcement here](https://cloud.google.com/blog/products/compute/tesla-v100-gpus-are-now-generally-available)). But so far I got catastrophic results there. The model ran slower than on a single V100. Probably a mistake somewhere. If you guys have access to these kinds of machines, you can try if you get better results. The [benchmark notebook](https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert) is configured for multi-GPU with MirroredStrategy."
        },
        {
          "id": 794641,
          "postDate": "2020-04-02T01:13:42.523Z",
          "content": "<p>That's weird... Could it be related to this? <a href=\"https://github.com/NVIDIA/apex/issues/282\">https://github.com/NVIDIA/apex/issues/282</a></p>",
          "rawMarkdown": "That's weird... Could it be related to this? https://github.com/NVIDIA/apex/issues/282"
        },
        {
          "id": 794650,
          "postDate": "2020-04-02T01:31:57.127Z",
          "content": "<p>I don't think GCP has 32GB V100s. They only have 16GB V100s. Using eight V100 16GB is not as efficient as using four V100 32GB. </p>\n\n<p>Using four NVLinked V100s 32GB I get approximately the same time per epoch as one TPUv3-8 in Flower Comp.</p>\n\n<p>(Note that a TPUv3-8 is actually four chips where each chip is 32GB. So four V100 32GB is an equal comparison to a TPUv3-8. Comparing a TPUv3-8 to a single P100 16GB is silly. The P100s don't have Tensor cores and are 3 times slower than V100).</p>",
          "rawMarkdown": "I don't think GCP has 32GB V100s. They only have 16GB V100s. Using eight V100 16GB is not as efficient as using four V100 32GB. \n\nUsing four NVLinked V100s 32GB I get approximately the same time per epoch as one TPUv3-8 in Flower Comp.\n\n(Note that a TPUv3-8 is actually four chips where each chip is 32GB. So four V100 32GB is an equal comparison to a TPUv3-8. Comparing a TPUv3-8 to a single P100 16GB is silly. The P100s don't have Tensor cores and are 3 times slower than V100).",
          "votes": 3
        },
        {
          "id": 794694,
          "postDate": "2020-04-02T02:38:52.173Z",
          "content": "<p>Yes <a href=\"/cdeotte\">@cdeotte</a> is right. Here is a short discussion about that:\n<a href=\"https://github.com/pytorch/xla/issues/1580\">https://github.com/pytorch/xla/issues/1580</a></p>",
          "rawMarkdown": "Yes @cdeotte is right. Here is a short discussion about that:\nhttps://github.com/pytorch/xla/issues/1580",
          "votes": 3
        },
        {
          "id": 794758,
          "postDate": "2020-04-02T04:20:21.817Z",
          "content": "<p>Any references guys for somone who had not a big background about architectures TPUv3-8, V100 32GB, NVLink,.. \nThanks</p>",
          "rawMarkdown": "Any references guys for somone who had not a big background about architectures TPUv3-8, V100 32GB, NVLink,.. \nThanks",
          "votes": 1
        },
        {
          "id": 795299,
          "postDate": "2020-04-02T15:48:57.907Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> if you had the time to run this BERT sample on 4 V100s, that would be nice. I cannot get it to work. My training times are slower than on a single V100.\n<a href=\"https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert\">https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert</a></p>",
          "rawMarkdown": "@cdeotte if you had the time to run this BERT sample on 4 V100s, that would be nice. I cannot get it to work. My training times are slower than on a single V100.\nhttps://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert",
          "votes": 2
        },
        {
          "id": 796854,
          "postDate": "2020-04-04T02:08:54.867Z",
          "content": "<p>Your TensorFlow code runs on 4x GPU V100 32GB with NVLInk. If I don't use XLA and only use mixed precision it is fastest and accommodates largest batch size. The four linked GPUs can use batch size 1024 just like TPUv3-8. The times I got are faster than your GPU benchmarks here. </p>\n\n<p>There is still something weird going on though. When i view GPU utilization, the 4 GPUs are not being utilized fully. The first GPU is mainly maxed out but the other 3 aren't working very hard. Also they cycle from working hard to not so hard. Before I publish times, I want to figure out how to balance the work load better so all GPUs are fully utilized.</p>",
          "rawMarkdown": "Your TensorFlow code runs on 4x GPU V100 32GB with NVLInk. If I don't use XLA and only use mixed precision it is fastest and accommodates largest batch size. The four linked GPUs can use batch size 1024 just like TPUv3-8. The times I got are faster than your GPU benchmarks here. \n\nThere is still something weird going on though. When i view GPU utilization, the 4 GPUs are not being utilized fully. The first GPU is mainly maxed out but the other 3 aren't working very hard. Also they cycle from working hard to not so hard. Before I publish times, I want to figure out how to balance the work load better so all GPUs are fully utilized.",
          "votes": 1
        },
        {
          "id": 799781,
          "postDate": "2020-04-06T18:45:47.110Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> The TF team got good results under TF 2.2. There was something fishy with MirroredStrategy + TF 2.1 + BERT apparently. I am re-running to confirm.</p>",
          "rawMarkdown": "@cdeotte The TF team got good results under TF 2.2. There was something fishy with MirroredStrategy + TF 2.1 + BERT apparently. I am re-running to confirm.",
          "votes": 1
        },
        {
          "id": 803234,
          "postDate": "2020-04-10T09:12:42.253Z",
          "content": "<p>You are too rich and have too much cards👀 </p>",
          "rawMarkdown": "You are too rich and have too much cards👀 "
        }
      ]
    },
    {
      "id": 800123,
      "postDate": "2020-04-07T05:51:06.450Z",
      "content": "<p>I managed to run this on multi-GPU configurations on GCP. There seems to be something off with MirroredStrategy + BERT in TF 2.1. I was getting slower training times than a single V100. In TF 2.2 it works though. BERT scales nicely to multi-GPU configs (3.7x faster on 4 V100s, 6x faster on 8 V100s) using MirroredStrategy in TF 2.2. The cost per training remains the same or more than a single GPU.</p>",
      "rawMarkdown": "I managed to run this on multi-GPU configurations on GCP. There seems to be something off with MirroredStrategy + BERT in TF 2.1. I was getting slower training times than a single V100. In TF 2.2 it works though. BERT scales nicely to multi-GPU configs (3.7x faster on 4 V100s, 6x faster on 8 V100s) using MirroredStrategy in TF 2.2. The cost per training remains the same or more than a single GPU.",
      "votes": 2
    },
    {
      "id": 796004,
      "postDate": "2020-04-03T07:39:00.443Z",
      "content": "<p>I published a <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#793801\">response in a forum nobody reads anymore</a>, replicating it below.  I see you welcome help in running these benchmarks on GPU, this is great.  We'll come back ASAP.  I also see NVLink is available in your experiments, this is also great.  It makes some of my previous answer irrelevant but I publish it as is still.</p>\n\n<hr>\n\n<p>Thanks for documenting the host CPU for TPU and for including multi V100 in the benchmark.</p>\n\n<p>First, let me start by saying that what follows is my opinion, and just my opinion. It is not an official NVIDIA statement nor does it represent NVIDIA position on this topic. It may well be that my colleagues at NVIDIA will disagree with some of what I say. After all, I have been at NVIDIA for only one month, and I am certainly not a GPU benchmarking expert.</p>\n\n<p>This said, let's proceed with some items.</p>\n\n<p>First, there is obviously a conflict of interest at play here given both Kaggle and TPU are owned by Google. A fair benchmarking would be that Kaggle/Google optimise their code settings for TPU and NVIDIA optimises code settings for GPU. Having you be the judge and one party is biased.</p>\n\n<p>Fair benchmarks are defined by a spec agreed upon by all parties, then each party performing their best independently. For this reason I would trust mlperf benchmarks way more than yours.</p>\n\n<p>Here are the latest mlperf results for training <a href=\"https://mlperf.org/training-results-0-6\">https://mlperf.org/training-results-0-6</a>\nThe comparison between TPU and GPU is way more balanced than what you show.</p>\n\n<p>Second, there are a number of issues here that can explain why your benchmark is at odd with other benchmarks like mlperf. Correct me if I am wrong.</p>\n\n<ul>\n<li>You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.</li>\n<li>You don't have NVLink between your GPUs</li>\n<li><p>You don't use mixed precision on V100 although it is available. glancing the notebook I found this comment:</p>\n\n<p>On GPU, specifically V100, mixed precision must be enabled for hardware TensorCores to be used.</p></li>\n<li><p>You don't use TPU v3 yet this is what started this whole topic</p></li>\n<li>The GPU cost is set by Google. We can probably find other providers that run V100 at a lower cost.</li>\n<li>I am also puzzled by the OOM you document. Why is TF using more RAM when using GPU than TPU? This has nothing to do with the GPU at first sight.</li>\n<li>Did you optimise GPU network bandwidth as recommended on <a href=\"https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth\">https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth</a> ?</li>\n</ul>\n\n<p>The good news here is that users get better and better options to train their model over time ;)</p>",
      "rawMarkdown": "I published a [response in a forum nobody reads anymore](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#793801), replicating it below.  I see you welcome help in running these benchmarks on GPU, this is great.  We'll come back ASAP.  I also see NVLink is available in your experiments, this is also great.  It makes some of my previous answer irrelevant but I publish it as is still.\n\n-----------\n\nThanks for documenting the host CPU for TPU and for including multi V100 in the benchmark.\n\nFirst, let me start by saying that what follows is my opinion, and just my opinion. It is not an official NVIDIA statement nor does it represent NVIDIA position on this topic. It may well be that my colleagues at NVIDIA will disagree with some of what I say. After all, I have been at NVIDIA for only one month, and I am certainly not a GPU benchmarking expert.\n\nThis said, let's proceed with some items.\n\nFirst, there is obviously a conflict of interest at play here given both Kaggle and TPU are owned by Google. A fair benchmarking would be that Kaggle/Google optimise their code settings for TPU and NVIDIA optimises code settings for GPU. Having you be the judge and one party is biased.\n\nFair benchmarks are defined by a spec agreed upon by all parties, then each party performing their best independently. For this reason I would trust mlperf benchmarks way more than yours.\n\nHere are the latest mlperf results for training https://mlperf.org/training-results-0-6\nThe comparison between TPU and GPU is way more balanced than what you show.\n\nSecond, there are a number of issues here that can explain why your benchmark is at odd with other benchmarks like mlperf. Correct me if I am wrong.\n\n-   You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.\n-  You don't have NVLink between your GPUs\n-  You don't use mixed precision on V100 although it is available. glancing the notebook I found this comment:\n\n    On GPU, specifically V100, mixed precision must be enabled for hardware TensorCores to be used.\n\n-   You don't use TPU v3 yet this is what started this whole topic\n-   The GPU cost is set by Google. We can probably find other providers that run V100 at a lower cost.\n-   I am also puzzled by the OOM you document. Why is TF using more RAM when using GPU than TPU? This has nothing to do with the GPU at first sight.\n-  Did you optimise GPU network bandwidth as recommended on https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth ?\n\nThe good news here is that users get better and better options to train their model over time ;)\n\n",
      "votes": -2,
      "replies": [
        {
          "id": 796684,
          "postDate": "2020-04-03T20:09:03.223Z",
          "content": "<p>&gt; You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.</p>\n\n<p>That's all I have on GCP, yes. Feel free to run this on bigger hardware and add your results.</p>\n\n<p>&gt; You don't use mixed precision</p>\n\n<p>??? line \"V100 mixed precision XLA\", also a whole table with only mixed precision results in the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#793801\">flowers topic</a>.</p>\n\n<p>&gt; I am also puzzled by the OOM you document.</p>\n\n<p>That's to show that I am running all models at the max batch size memory allows. An obvious way of cheating on the benchmark would be to run the TPU at batch size 1024 and the GPU at batch size 2.</p>\n\n<p>&gt; You don't have NVLink between your GPUs</p>\n\n<p>??? I think I do. All multi-GPU <a href=\"https://cloud.google.com/blog/products/compute/tesla-v100-gpus-are-now-generally-available\">v100 configs on GCP have it</a>.\n&gt; Did you optimize GPU network bandwidth</p>\n\n<p>I'm using the images provided out of the in Cloud AI Platform Notebooks. I'm not sure they have it. I'll ask the team.</p>\n\n<p>Side note about the multi-GPU setup: it will improve on speed, but not on the price/training metric. 2 GPUs are 2x more expensive than one and at best 2x faster. So the cost per training is at best the same as one GPU.</p>\n\n<p>Also, I should have said this earlier, the focus in these benchmarks is usability and productivity. The performance and cost numbers have to be easy to achieve by normal people.</p>\n\n<p>&gt; Here are the latest mlperf results for training <a href=\"https://mlperf.org/training-results-0-6\">https://mlperf.org/training-results-0-6</a></p>\n\n<p>The smallest TPU config there is TPUv3-32, The smallest GPU is a DGX1 machine. Both out of reach for the Kaggle crowd even if they have a little money to spend.</p>\n\n<p>Please feel free to run this on additional hardware and discuss the findings. I welcome the debate. I ran these benchmarks myself to make sure I know the code and can answer your question. </p>",
          "rawMarkdown": "&gt; You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.\n\nThat's all I have on GCP, yes. Feel free to run this on bigger hardware and add your results.\n\n&gt; You don't use mixed precision\n\n??? line \"V100 mixed precision XLA\", also a whole table with only mixed precision results in the [flowers topic](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#793801).\n\n&gt; I am also puzzled by the OOM you document.\n\nThat's to show that I am running all models at the max batch size memory allows. An obvious way of cheating on the benchmark would be to run the TPU at batch size 1024 and the GPU at batch size 2.\n\n&gt; You don't have NVLink between your GPUs\n\n??? I think I do. All multi-GPU [v100 configs on GCP have it](https://cloud.google.com/blog/products/compute/tesla-v100-gpus-are-now-generally-available).\n&gt; Did you optimize GPU network bandwidth\n\nI'm using the images provided out of the in Cloud AI Platform Notebooks. I'm not sure they have it. I'll ask the team.\n\nSide note about the multi-GPU setup: it will improve on speed, but not on the price/training metric. 2 GPUs are 2x more expensive than one and at best 2x faster. So the cost per training is at best the same as one GPU.\n\nAlso, I should have said this earlier, the focus in these benchmarks is usability and productivity. The performance and cost numbers have to be easy to achieve by normal people.\n\n&gt; Here are the latest mlperf results for training https://mlperf.org/training-results-0-6\n\nThe smallest TPU config there is TPUv3-32, The smallest GPU is a DGX1 machine. Both out of reach for the Kaggle crowd even if they have a little money to spend.\n\nPlease feel free to run this on additional hardware and discuss the findings. I welcome the debate. I ran these benchmarks myself to make sure I know the code and can answer your question. ",
          "votes": 1
        },
        {
          "id": 797196,
          "postDate": "2020-04-04T10:05:44.517Z",
          "content": "<p>What I as a customer care about is:</p>\n\n<ul>\n<li>Runtime per $</li>\n<li>Accuracy</li>\n</ul>\n\n<p>I have a budget, and want to train my models as fast and accurate as possible. All other benchmarks are quite useless to me. It doesn't make sense to compare a certain TPU configuration with a larger GPU cluster having similar TFlops if it costs twice or more than the TPUs. At the same time it doesn't make sense to compare cheaper TPU setup with way more TFlops if the accuracy/model performance is worse. So having a table like this (for different kind of ML models) would be great:</p>\n\n<p>| Hardware | TFlops | $/Hour | Runtime | Accuracy|\n| --- | --- | --- | --- | --- |\n| x | x | x | x | x |</p>",
          "rawMarkdown": "What I as a customer care about is:\n\n- Runtime per $\n- Accuracy\n\nI have a budget, and want to train my models as fast and accurate as possible. All other benchmarks are quite useless to me. It doesn't make sense to compare a certain TPU configuration with a larger GPU cluster having similar TFlops if it costs twice or more than the TPUs. At the same time it doesn't make sense to compare cheaper TPU setup with way more TFlops if the accuracy/model performance is worse. So having a table like this (for different kind of ML models) would be great:\n\n| Hardware | TFlops | $/Hour | Runtime | Accuracy|\n| --- | --- | --- | --- | --- |\n| x | x | x | x | x |"
        },
        {
          "id": 798258,
          "postDate": "2020-04-05T10:56:50.437Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> You should have  a look at <a href=\"https://vast.ai/\">https://vast.ai/</a> to get GPU at a much better price.</p>\n\n<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I should not have re posted my answer as is given you addressed some of the issues since I posted it.  Adding a caveat before it wasn't enough obviously.  I should have edited it.</p>",
          "rawMarkdown": "@philippsinger You should have  a look at https://vast.ai/ to get GPU at a much better price.\n\n@mgornergoogle I should not have re posted my answer as is given you addressed some of the issues since I posted it.  Adding a caveat before it wasn't enough obviously.  I should have edited it.\n\n"
        },
        {
          "id": 798315,
          "postDate": "2020-04-05T11:54:06.350Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I am using vast.ai for a while, but the platform has its own issues, far worse than any other cloud platform.</p>",
          "rawMarkdown": "@cpmpml I am using vast.ai for a while, but the platform has its own issues, far worse than any other cloud platform."
        },
        {
          "id": 798425,
          "postDate": "2020-04-05T13:38:13.813Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> You didn't list reliability or other criteria, just this:</p>\n\n<blockquote>\n  <p>What I as a customer care about is:</p>\n  \n  <p>Runtime per $\n     Accuracy</p>\n  \n  <p>I have a budget, and want to train my models as fast and accurate as possible. All other benchmarks are quite useless to me.</p>\n</blockquote>\n\n<p>Reliability, user experience, etc are also valid criteria.</p>",
          "rawMarkdown": "@philippsinger You didn't list reliability or other criteria, just this:\n\n&gt; What I as a customer care about is:\n&gt;\n&gt;    Runtime per $\n&gt;    Accuracy\n&gt;\n&gt; I have a budget, and want to train my models as fast and accurate as possible. All other benchmarks are quite useless to me.\n\nReliability, user experience, etc are also valid criteria.\n"
        },
        {
          "id": 798427,
          "postDate": "2020-04-05T13:40:20.303Z",
          "content": "<p>Here are tools that were recommended to me to be used with vast.ai . I pass the info but I have not used them.</p>\n\n<p><a href=\"https://pypi.org/project/remote_ikernel/\">https://pypi.org/project/remote_ikernel/</a></p>\n\n<p><a href=\"https://github.com/ml-tooling/ml-workspace\">https://github.com/ml-tooling/ml-workspace</a></p>",
          "rawMarkdown": "Here are tools that were recommended to me to be used with vast.ai . I pass the info but I have not used them.\n\nhttps://pypi.org/project/remote_ikernel/\n\nhttps://github.com/ml-tooling/ml-workspace",
          "votes": 1
        },
        {
          "id": 798481,
          "postDate": "2020-04-05T14:34:58.703Z",
          "content": "<p>vast.ai should not really be a discussion here as it is very unreliable and as said, has many issues. It is basically just some random guys hosting vms, you can imagine the issues that come with it. That said, I  occasionally use it, but it is not a good solution.</p>",
          "rawMarkdown": "vast.ai should not really be a discussion here as it is very unreliable and as said, has many issues. It is basically just some random guys hosting vms, you can imagine the issues that come with it. That said, I  occasionally use it, but it is not a good solution."
        },
        {
          "id": 798873,
          "postDate": "2020-04-06T00:12:21.180Z",
          "content": "<p>I don't have a dog in the Nvidia vs. Google compute fight but I have to say that Google's been very generous with giving away TPU compute time to students, Kagglers, and researchers. It's kinda crazy that Kagglers get 30 hours of TPUv3 usage per week. </p>",
          "rawMarkdown": "I don't have a dog in the Nvidia vs. Google compute fight but I have to say that Google's been very generous with giving away TPU compute time to students, Kagglers, and researchers. It's kinda crazy that Kagglers get 30 hours of TPUv3 usage per week. ",
          "votes": 8
        },
        {
          "id": 798877,
          "postDate": "2020-04-06T00:19:09.620Z",
          "content": "<blockquote>\n  <p>30 * 8.8 * 4$ per month on each kaggler! </p>\n</blockquote>\n\n<p>But not every kaggler can use it to it's fullest capacity as well but still it's great for people who want to take TPUs for a date maybe haha !</p>",
          "rawMarkdown": "&gt;30 * 8.8 * 4$ per month on each kaggler! \n\nBut not every kaggler can use it to it's fullest capacity as well but still it's great for people who want to take TPUs for a date maybe haha !"
        },
        {
          "id": 799793,
          "postDate": "2020-04-06T18:58:51.390Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Accuray has been a top worry for me as well but with more experience it has gone away. Just increasing the batch size when you jump on TPU often lowers the accuracy, because the learning rate is no longer adequate. But I have always been able to hp-tune my model to get back to the original accuracy. When I tried a tpu-v3-128 pod (128 cores), which requires really large batch sizes to fully utilize, getting the learning rate right was more challenging. I got there in the end but had to train for a bit longer. Depends on the model. On a TPU-v3-8 (8 cores, the ones available on Kaggle) I am now confident I can always get the same accuracy than GPUs.</p>\n\n<p>In the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983\">other benchmarking thread</a>, I did measure accuracy but I did not publish it to save space as it was the same for all models: 0.96 or 0.97. In this run of BERT benchmarks, I should have measured accuracy too (I didn't)</p>",
          "rawMarkdown": "@philippsinger Accuray has been a top worry for me as well but with more experience it has gone away. Just increasing the batch size when you jump on TPU often lowers the accuracy, because the learning rate is no longer adequate. But I have always been able to hp-tune my model to get back to the original accuracy. When I tried a tpu-v3-128 pod (128 cores), which requires really large batch sizes to fully utilize, getting the learning rate right was more challenging. I got there in the end but had to train for a bit longer. Depends on the model. On a TPU-v3-8 (8 cores, the ones available on Kaggle) I am now confident I can always get the same accuracy than GPUs.\n\nIn the [other benchmarking thread](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983), I did measure accuracy but I did not publish it to save space as it was the same for all models: 0.96 or 0.97. In this run of BERT benchmarks, I should have measured accuracy too (I didn't)",
          "votes": 1
        }
      ]
    },
    {
      "id": 811547,
      "postDate": "2020-04-18T03:53:44.513Z",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> for the benchmark  notebook. At first it used tf 2.1 if we change to use tf-nightly tpu init will fail.  Second if still using tf 2.1 but change a bit to use subclass api then tpu can not work. Can you help me for this Thanks! <br>\n<a href=\"https://github.com/tensorflow/tensorflow/issues/38538\">https://github.com/tensorflow/tensorflow/issues/38538</a></p>",
      "rawMarkdown": "@mgornergoogle for the benchmark  notebook. At first it used tf 2.1 if we change to use tf-nightly tpu init will fail.  Second if still using tf 2.1 but change a bit to use subclass api then tpu can not work. Can you help me for this Thanks!  \nhttps://github.com/tensorflow/tensorflow/issues/38538",
      "replies": [
        {
          "id": 814675,
          "postDate": "2020-04-20T21:25:11.643Z",
          "content": "<p>You say \"I tried to use latest tf in kaggle kernel\".\nThis is not possible on Kaggle with TPUs. Only the current version of Tensorflow is supported. Installing a new version of Tensorflow in your notebook does not upgrade the TF firmware on the TPU side and you end up with communication problems between the notebook and the TPU.</p>",
          "rawMarkdown": "You say \"I tried to use latest tf in kaggle kernel\".\nThis is not possible on Kaggle with TPUs. Only the current version of Tensorflow is supported. Installing a new version of Tensorflow in your notebook does not upgrade the TF firmware on the TPU side and you end up with communication problems between the notebook and the TPU."
        },
        {
          "id": 814755,
          "postDate": "2020-04-21T00:06:38.643Z",
          "content": "<p>Thanks Martin!</p>",
          "rawMarkdown": "Thanks Martin!"
        }
      ]
    },
    {
      "id": 793524,
      "postDate": "2020-04-01T04:58:57.813Z",
      "rawMarkdown": "",
      "votes": 14,
      "isDeleted": true,
      "replies": [
        {
          "id": 793594,
          "postDate": "2020-04-01T06:17:02.187Z",
          "content": "<p>TODO next: enable mixed precision on TPU and GPU. This should improve the V100 numbers a lot\nThen use a <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\">custom training loop version for TPU</a>. This should speed up the TPU by approx. 5%.</p>",
          "rawMarkdown": "TODO next: enable mixed precision on TPU and GPU. This should improve the V100 numbers a lot\nThen use a [custom training loop version for TPU](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455). This should speed up the TPU by approx. 5%.",
          "votes": 2
        },
        {
          "id": 793600,
          "postDate": "2020-04-01T06:25:33.493Z",
          "content": "<p>I completely forgot about this! Thanks for reminding us that this is an option! We can do this in Kaggle and with PyTorch XLA as well? I assume we can just set <code>XLA_USE_BF16</code> to True in our notebook?</p>",
          "rawMarkdown": "I completely forgot about this! Thanks for reminding us that this is an option! We can do this in Kaggle and with PyTorch XLA as well? I assume we can just set `XLA_USE_BF16` to True in our notebook?"
        },
        {
          "id": 794136,
          "postDate": "2020-04-01T15:37:44.880Z",
          "content": "<p>It won't have much effect on Kaggle P100s. Only V100s have proper dedicated mixed precision hardware.</p>",
          "rawMarkdown": "It won't have much effect on Kaggle P100s. Only V100s have proper dedicated mixed precision hardware."
        },
        {
          "id": 794437,
          "postDate": "2020-04-01T20:23:53.887Z",
          "content": "<p>No, I mean for the Kaggle TPUs.</p>",
          "rawMarkdown": "No, I mean for the Kaggle TPUs."
        },
        {
          "id": 794570,
          "postDate": "2020-04-01T22:41:48.890Z",
          "content": "<p>Kaggle GPUs are P100s.\nIf you try it, tell us what you found. I just ran on V100 with mixed precision and XLA and I got a speedup of 13%. I'll be updating the chart above shortly.</p>",
          "rawMarkdown": "Kaggle GPUs are P100s.\nIf you try it, tell us what you found. I just ran on V100 with mixed precision and XLA and I got a speedup of 13%. I'll be updating the chart above shortly."
        },
        {
          "id": 794575,
          "postDate": "2020-04-01T22:54:38.187Z",
          "content": "<p>Hmm ok maybe I don't understand how the half-precision works with TPUs. I thought it I use bfloat16, which is a TPU-specific 16-bit data type, it will speed up TPU computation? Why does it matter what GPU is being used?</p>",
          "rawMarkdown": "Hmm ok maybe I don't understand how the half-precision works with TPUs. I thought it I use bfloat16, which is a TPU-specific 16-bit data type, it will speed up TPU computation? Why does it matter what GPU is being used?"
        },
        {
          "id": 794640,
          "postDate": "2020-04-02T01:10:35.920Z",
          "content": "<p>Sorry, I meant \"Kaggle <strong>GPUs</strong> are P100s\". I fixed it above.</p>",
          "rawMarkdown": "Sorry, I meant \"Kaggle **GPUs** are P100s\". I fixed it above.",
          "votes": 1
        },
        {
          "id": 807481,
          "postDate": "2020-04-14T17:00:30.390Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Have you tried to apply mixed precision to transformer with tf.keras? I used below setting from tutorial but failed. It seems inside the transformer some data type is hard coded... If you already did and got successful result, could you please release the benchmark code? Thanks!\n<code>policy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)</code></p>",
          "rawMarkdown": "@mgornergoogle Have you tried to apply mixed precision to transformer with tf.keras? I used below setting from tutorial but failed. It seems inside the transformer some data type is hard coded... If you already did and got successful result, could you please release the benchmark code? Thanks!\n`policy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)`"
        },
        {
          "id": 808869,
          "postDate": "2020-04-15T17:24:23.380Z",
          "content": "<p>Yes, in the three rows in the table marked \"mixed precision XLA\". The <a href=\"https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert\">benchmark code</a> has this setting too. You can play with it. On a V100 GPU and with this model, mixed precision has a noticeable but not a dramatic effect. It had a bigger effect in the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#784972\">image classification benchmark</a>.</p>",
          "rawMarkdown": "Yes, in the three rows in the table marked \"mixed precision XLA\". The [benchmark code](https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert) has this setting too. You can play with it. On a V100 GPU and with this model, mixed precision has a noticeable but not a dramatic effect. It had a bigger effect in the [image classification benchmark](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#784972).",
          "votes": 1
        },
        {
          "id": 808952,
          "postDate": "2020-04-15T18:34:02.723Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Oh I missed that one! Thanks, I'll check it:)</p>",
          "rawMarkdown": "@mgornergoogle Oh I missed that one! Thanks, I'll check it:)"
        },
        {
          "id": 809227,
          "postDate": "2020-04-16T00:45:02.567Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I've tried, but failed... I made a notebook which explains where I'm failing. I believe it's because data type is hardcoded in the huggiung face bert implementation. If so, I should rather ask at hugging face repository, but can you please take a look at what is wrong? Thanks,\n<a href=\"https://www.kaggle.com/bamps53/hugging-face-transformer-tpu-xla?rvi=1\">https://www.kaggle.com/bamps53/hugging-face-transformer-tpu-xla?rvi=1</a></p>",
          "rawMarkdown": "@mgornergoogle I've tried, but failed... I made a notebook which explains where I'm failing. I believe it's because data type is hardcoded in the huggiung face bert implementation. If so, I should rather ask at hugging face repository, but can you please take a look at what is wrong? Thanks,\nhttps://www.kaggle.com/bamps53/hugging-face-transformer-tpu-xla?rvi=1"
        },
        {
          "id": 810133,
          "postDate": "2020-04-16T18:41:26.583Z",
          "content": "<p>I looked at the code but I'm not sure what is wrong. It is either something in the implementation of LayerNorm. I'm not sure if this function is from Tensorflow or HuggingFace (in Keras, it's called LayerNormalization). Or maybe it has to do with HuggingFace's embeddings code.</p>",
          "rawMarkdown": "I looked at the code but I'm not sure what is wrong. It is either something in the implementation of LayerNorm. I'm not sure if this function is from Tensorflow or HuggingFace (in Keras, it's called LayerNormalization). Or maybe it has to do with HuggingFace's embeddings code."
        },
        {
          "id": 810145,
          "postDate": "2020-04-16T18:48:04.217Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for help. I also checked hugging face implementation and played around by casting bfloat16 to every tensor which is failed to convert automatically, but they are keep coming... There was no code which specify float32 directly, but somehow some tensor stay float32. <br>\nAnyway thanks, if you couldn't figure out why, definitely me either!😃  I gave up and just post it at hugging face issue.</p>",
          "rawMarkdown": "@mgornergoogle Thanks for help. I also checked hugging face implementation and played around by casting bfloat16 to every tensor which is failed to convert automatically, but they are keep coming... There was no code which specify float32 directly, but somehow some tensor stay float32.  \nAnyway thanks, if you couldn't figure out why, definitely me either!😃  I gave up and just post it at hugging face issue."
        },
        {
          "id": 810193,
          "postDate": "2020-04-16T19:11:40.620Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> One more thing, I thinks it doesn't come from LayerNormalization, but <a href=\"https://github.com/tensorflow/tensorflow/blob/e5bf8de410005de06a7ff5393fafdf832ef1d4ad/tensorflow/python/keras/layers/embeddings.py#L107\">tf.keras.layers.Embedding</a>. Here kwargs['autocast'] is set to False. Does it related bfloat16 casting??</p>",
          "rawMarkdown": "@mgornergoogle One more thing, I thinks it doesn't come from LayerNormalization, but [tf.keras.layers.Embedding](https://github.com/tensorflow/tensorflow/blob/e5bf8de410005de06a7ff5393fafdf832ef1d4ad/tensorflow/python/keras/layers/embeddings.py#L107). Here kwargs['autocast'] is set to False. Does it related bfloat16 casting??"
        }
      ]
    },
    {
      "id": 792187,
      "postDate": "2020-03-31T00:46:13.323Z",
      "rawMarkdown": "",
      "votes": 5,
      "isDeleted": true,
      "replies": [
        {
          "id": 792489,
          "postDate": "2020-03-31T08:37:42.443Z",
          "content": "<p>Maybe we can help with providing numbers with 8 x V100 and enough cpu.  What you do not document is the cpu used on the TPU host.  The fact that the P100 is starving with a 2 core cpu certainly plays a role too.</p>",
          "rawMarkdown": "Maybe we can help with providing numbers with 8 x V100 and enough cpu.  What you do not document is the cpu used on the TPU host.  The fact that the P100 is starving with a 2 core cpu certainly plays a role too.",
          "votes": 5
        },
        {
          "id": 792491,
          "postDate": "2020-03-31T08:39:13.277Z",
          "content": "<p>I don't think the first notebook link is inaccessible. It is telling me to sign in through Google?</p>",
          "rawMarkdown": "I don't think the first notebook link is inaccessible. It is telling me to sign in through Google?",
          "votes": 1
        },
        {
          "id": 792903,
          "postDate": "2020-03-31T16:25:09.113Z",
          "content": "<p>Sorry for the bad link. fixed!</p>",
          "rawMarkdown": "Sorry for the bad link. fixed!"
        },
        {
          "id": 792905,
          "postDate": "2020-03-31T16:27:06.827Z",
          "content": "<p>Not sure the GPU is starving since in this example, data is entirely in memory. I'll be running these numbers today, we'll see.</p>",
          "rawMarkdown": "Not sure the GPU is starving since in this example, data is entirely in memory. I'll be running these numbers today, we'll see."
        }
      ]
    },
    {
      "id": 794805,
      "postDate": "2020-04-02T05:29:44.210Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 792675,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-03-31T12:58:50.560000",
      "content": "<p>The speed gains are indeed impressive! What bothers me is that again there are many abstraction layers that don't let me understand the details of the running execution in a level I would love to. I regularly get some \"process died\" messages without any stacktrace, or some weird memory issues. It is very cumbersome to make it work.</p>\n\n<p>That said though, I only tried pytorch and I know that it is not optimized for that yet. Is debugging way easier when using TF / Keras + TPUs?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 792712,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-03-31T13:56:55.013000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I've made many experiments on the this and the Flowers competitions and I can say that it seems to be easier to use TPUs with TF, but the error message doesn't really help, I got a lot of memory issues, but the log almost never points that, I've been able to solve most with the help of Martin and trial &amp; error.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 792918,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-03-31T16:35:19.053000",
          "content": "<p>With PyTorch, are you initializing the TPU as recommended through</p>\n\n<p><code>\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code></p>\n\n<p>I have seem samples lying around that are missing this. If you dig into the script, you will notice a call to a specific port on the TPU (8475). This switches the TPU firmware to the latest nightly version for PyTorch. If you miss this step, your setup will still work but with bugs. This might explain some of the instability you have been experiencing.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 792921,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-31T16:42:14.760000",
          "content": "<p>Yeah, am exactly running these commands.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793107,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-03-31T20:00:15.123000",
          "content": "<p>If you can put together a reproducible crash in PyTorch, please do so, post it on the <a href=\"https://github.com/pytorch/xla\">PyTorch/XLA </a>GitHub and link it here so that I see it too.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 794058,
      "author_name": "xhlulu",
      "author_url": "",
      "post_date": "2020-04-01T14:20:44.027000",
      "content": "<p>Very nice benchmarks! Have you tried 32gb GPUs, and nvlinked gpus?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 794075,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-04-01T14:34:26.433000",
          "content": "<p>Is nvlink really useful? Was thinking for a while to set it up.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794138,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-04-01T15:39:02.260000",
          "content": "<p>I never tried, but i thought that nvlink would let you run a model on both GPUs without any need for parallel controls?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794142,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-01T15:43:22.767000",
          "content": "<p>Yes, as I said above I tried 4xV100 and 8xV100 machines on GCP. They have NVLink (launch <a href=\"https://cloud.google.com/blog/products/compute/tesla-v100-gpus-are-now-generally-available\">announcement here</a>). But so far I got catastrophic results there. The model ran slower than on a single V100. Probably a mistake somewhere. If you guys have access to these kinds of machines, you can try if you get better results. The <a href=\"https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert\">benchmark notebook</a> is configured for multi-GPU with MirroredStrategy.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794641,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-04-02T01:13:42.523000",
          "content": "<p>That's weird... Could it be related to this? <a href=\"https://github.com/NVIDIA/apex/issues/282\">https://github.com/NVIDIA/apex/issues/282</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794650,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-04-02T01:31:57.127000",
          "content": "<p>I don't think GCP has 32GB V100s. They only have 16GB V100s. Using eight V100 16GB is not as efficient as using four V100 32GB. </p>\n\n<p>Using four NVLinked V100s 32GB I get approximately the same time per epoch as one TPUv3-8 in Flower Comp.</p>\n\n<p>(Note that a TPUv3-8 is actually four chips where each chip is 32GB. So four V100 32GB is an equal comparison to a TPUv3-8. Comparing a TPUv3-8 to a single P100 16GB is silly. The P100s don't have Tensor cores and are 3 times slower than V100).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 794694,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-04-02T02:38:52.173000",
          "content": "<p>Yes <a href=\"/cdeotte\">@cdeotte</a> is right. Here is a short discussion about that:\n<a href=\"https://github.com/pytorch/xla/issues/1580\">https://github.com/pytorch/xla/issues/1580</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 794758,
          "author_name": "Tarek Hamdi",
          "author_url": "",
          "post_date": "2020-04-02T04:20:21.817000",
          "content": "<p>Any references guys for somone who had not a big background about architectures TPUv3-8, V100 32GB, NVLink,.. \nThanks</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 795299,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-02T15:48:57.907000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> if you had the time to run this BERT sample on 4 V100s, that would be nice. I cannot get it to work. My training times are slower than on a single V100.\n<a href=\"https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert\">https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 796854,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-04-04T02:08:54.867000",
          "content": "<p>Your TensorFlow code runs on 4x GPU V100 32GB with NVLInk. If I don't use XLA and only use mixed precision it is fastest and accommodates largest batch size. The four linked GPUs can use batch size 1024 just like TPUv3-8. The times I got are faster than your GPU benchmarks here. </p>\n\n<p>There is still something weird going on though. When i view GPU utilization, the 4 GPUs are not being utilized fully. The first GPU is mainly maxed out but the other 3 aren't working very hard. Also they cycle from working hard to not so hard. Before I publish times, I want to figure out how to balance the work load better so all GPUs are fully utilized.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 799781,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-06T18:45:47.110000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> The TF team got good results under TF 2.2. There was something fishy with MirroredStrategy + TF 2.1 + BERT apparently. I am re-running to confirm.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 803234,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-04-10T09:12:42.253000",
          "content": "<p>You are too rich and have too much cards👀 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 800123,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-04-07T05:51:06.450000",
      "content": "<p>I managed to run this on multi-GPU configurations on GCP. There seems to be something off with MirroredStrategy + BERT in TF 2.1. I was getting slower training times than a single V100. In TF 2.2 it works though. BERT scales nicely to multi-GPU configs (3.7x faster on 4 V100s, 6x faster on 8 V100s) using MirroredStrategy in TF 2.2. The cost per training remains the same or more than a single GPU.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 796004,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-04-03T07:39:00.443000",
      "content": "<p>I published a <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#793801\">response in a forum nobody reads anymore</a>, replicating it below.  I see you welcome help in running these benchmarks on GPU, this is great.  We'll come back ASAP.  I also see NVLink is available in your experiments, this is also great.  It makes some of my previous answer irrelevant but I publish it as is still.</p>\n\n<hr>\n\n<p>Thanks for documenting the host CPU for TPU and for including multi V100 in the benchmark.</p>\n\n<p>First, let me start by saying that what follows is my opinion, and just my opinion. It is not an official NVIDIA statement nor does it represent NVIDIA position on this topic. It may well be that my colleagues at NVIDIA will disagree with some of what I say. After all, I have been at NVIDIA for only one month, and I am certainly not a GPU benchmarking expert.</p>\n\n<p>This said, let's proceed with some items.</p>\n\n<p>First, there is obviously a conflict of interest at play here given both Kaggle and TPU are owned by Google. A fair benchmarking would be that Kaggle/Google optimise their code settings for TPU and NVIDIA optimises code settings for GPU. Having you be the judge and one party is biased.</p>\n\n<p>Fair benchmarks are defined by a spec agreed upon by all parties, then each party performing their best independently. For this reason I would trust mlperf benchmarks way more than yours.</p>\n\n<p>Here are the latest mlperf results for training <a href=\"https://mlperf.org/training-results-0-6\">https://mlperf.org/training-results-0-6</a>\nThe comparison between TPU and GPU is way more balanced than what you show.</p>\n\n<p>Second, there are a number of issues here that can explain why your benchmark is at odd with other benchmarks like mlperf. Correct me if I am wrong.</p>\n\n<ul>\n<li>You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.</li>\n<li>You don't have NVLink between your GPUs</li>\n<li><p>You don't use mixed precision on V100 although it is available. glancing the notebook I found this comment:</p>\n\n<p>On GPU, specifically V100, mixed precision must be enabled for hardware TensorCores to be used.</p></li>\n<li><p>You don't use TPU v3 yet this is what started this whole topic</p></li>\n<li>The GPU cost is set by Google. We can probably find other providers that run V100 at a lower cost.</li>\n<li>I am also puzzled by the OOM you document. Why is TF using more RAM when using GPU than TPU? This has nothing to do with the GPU at first sight.</li>\n<li>Did you optimise GPU network bandwidth as recommended on <a href=\"https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth\">https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth</a> ?</li>\n</ul>\n\n<p>The good news here is that users get better and better options to train their model over time ;)</p>",
      "votes": -2,
      "replies": [
        {
          "id": 796684,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-03T20:09:03.223000",
          "content": "<p>&gt; You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.</p>\n\n<p>That's all I have on GCP, yes. Feel free to run this on bigger hardware and add your results.</p>\n\n<p>&gt; You don't use mixed precision</p>\n\n<p>??? line \"V100 mixed precision XLA\", also a whole table with only mixed precision results in the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#793801\">flowers topic</a>.</p>\n\n<p>&gt; I am also puzzled by the OOM you document.</p>\n\n<p>That's to show that I am running all models at the max batch size memory allows. An obvious way of cheating on the benchmark would be to run the TPU at batch size 1024 and the GPU at batch size 2.</p>\n\n<p>&gt; You don't have NVLink between your GPUs</p>\n\n<p>??? I think I do. All multi-GPU <a href=\"https://cloud.google.com/blog/products/compute/tesla-v100-gpus-are-now-generally-available\">v100 configs on GCP have it</a>.\n&gt; Did you optimize GPU network bandwidth</p>\n\n<p>I'm using the images provided out of the in Cloud AI Platform Notebooks. I'm not sure they have it. I'll ask the team.</p>\n\n<p>Side note about the multi-GPU setup: it will improve on speed, but not on the price/training metric. 2 GPUs are 2x more expensive than one and at best 2x faster. So the cost per training is at best the same as one GPU.</p>\n\n<p>Also, I should have said this earlier, the focus in these benchmarks is usability and productivity. The performance and cost numbers have to be easy to achieve by normal people.</p>\n\n<p>&gt; Here are the latest mlperf results for training <a href=\"https://mlperf.org/training-results-0-6\">https://mlperf.org/training-results-0-6</a></p>\n\n<p>The smallest TPU config there is TPUv3-32, The smallest GPU is a DGX1 machine. Both out of reach for the Kaggle crowd even if they have a little money to spend.</p>\n\n<p>Please feel free to run this on additional hardware and discuss the findings. I welcome the debate. I ran these benchmarks myself to make sure I know the code and can answer your question. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 797196,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-04-04T10:05:44.517000",
          "content": "<p>What I as a customer care about is:</p>\n\n<ul>\n<li>Runtime per $</li>\n<li>Accuracy</li>\n</ul>\n\n<p>I have a budget, and want to train my models as fast and accurate as possible. All other benchmarks are quite useless to me. It doesn't make sense to compare a certain TPU configuration with a larger GPU cluster having similar TFlops if it costs twice or more than the TPUs. At the same time it doesn't make sense to compare cheaper TPU setup with way more TFlops if the accuracy/model performance is worse. So having a table like this (for different kind of ML models) would be great:</p>\n\n<p>| Hardware | TFlops | $/Hour | Runtime | Accuracy|\n| --- | --- | --- | --- | --- |\n| x | x | x | x | x |</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 798258,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-04-05T10:56:50.437000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> You should have  a look at <a href=\"https://vast.ai/\">https://vast.ai/</a> to get GPU at a much better price.</p>\n\n<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I should not have re posted my answer as is given you addressed some of the issues since I posted it.  Adding a caveat before it wasn't enough obviously.  I should have edited it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 798315,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-04-05T11:54:06.350000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I am using vast.ai for a while, but the platform has its own issues, far worse than any other cloud platform.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 798425,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-04-05T13:38:13.813000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> You didn't list reliability or other criteria, just this:</p>\n\n<blockquote>\n  <p>What I as a customer care about is:</p>\n  \n  <p>Runtime per $\n     Accuracy</p>\n  \n  <p>I have a budget, and want to train my models as fast and accurate as possible. All other benchmarks are quite useless to me.</p>\n</blockquote>\n\n<p>Reliability, user experience, etc are also valid criteria.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 798427,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-04-05T13:40:20.303000",
          "content": "<p>Here are tools that were recommended to me to be used with vast.ai . I pass the info but I have not used them.</p>\n\n<p><a href=\"https://pypi.org/project/remote_ikernel/\">https://pypi.org/project/remote_ikernel/</a></p>\n\n<p><a href=\"https://github.com/ml-tooling/ml-workspace\">https://github.com/ml-tooling/ml-workspace</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 798481,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-04-05T14:34:58.703000",
          "content": "<p>vast.ai should not really be a discussion here as it is very unreliable and as said, has many issues. It is basically just some random guys hosting vms, you can imagine the issues that come with it. That said, I  occasionally use it, but it is not a good solution.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 798873,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-04-06T00:12:21.180000",
          "content": "<p>I don't have a dog in the Nvidia vs. Google compute fight but I have to say that Google's been very generous with giving away TPU compute time to students, Kagglers, and researchers. It's kinda crazy that Kagglers get 30 hours of TPUv3 usage per week. </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 798877,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-04-06T00:19:09.620000",
          "content": "<blockquote>\n  <p>30 * 8.8 * 4$ per month on each kaggler! </p>\n</blockquote>\n\n<p>But not every kaggler can use it to it's fullest capacity as well but still it's great for people who want to take TPUs for a date maybe haha !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 799793,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-06T18:58:51.390000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Accuray has been a top worry for me as well but with more experience it has gone away. Just increasing the batch size when you jump on TPU often lowers the accuracy, because the learning rate is no longer adequate. But I have always been able to hp-tune my model to get back to the original accuracy. When I tried a tpu-v3-128 pod (128 cores), which requires really large batch sizes to fully utilize, getting the learning rate right was more challenging. I got there in the end but had to train for a bit longer. Depends on the model. On a TPU-v3-8 (8 cores, the ones available on Kaggle) I am now confident I can always get the same accuracy than GPUs.</p>\n\n<p>In the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983\">other benchmarking thread</a>, I did measure accuracy but I did not publish it to save space as it was the same for all models: 0.96 or 0.97. In this run of BERT benchmarks, I should have measured accuracy too (I didn't)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 811547,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2020-04-18T03:53:44.513000",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> for the benchmark  notebook. At first it used tf 2.1 if we change to use tf-nightly tpu init will fail.  Second if still using tf 2.1 but change a bit to use subclass api then tpu can not work. Can you help me for this Thanks! <br>\n<a href=\"https://github.com/tensorflow/tensorflow/issues/38538\">https://github.com/tensorflow/tensorflow/issues/38538</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 814675,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-20T21:25:11.643000",
          "content": "<p>You say \"I tried to use latest tf in kaggle kernel\".\nThis is not possible on Kaggle with TPUs. Only the current version of Tensorflow is supported. Installing a new version of Tensorflow in your notebook does not upgrade the TF firmware on the TPU side and you end up with communication problems between the notebook and the TPU.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 814755,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2020-04-21T00:06:38.643000",
          "content": "<p>Thanks Martin!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793524,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-01T04:58:57.813000",
      "content": "",
      "votes": 14,
      "replies": [
        {
          "id": 793594,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-01T06:17:02.187000",
          "content": "<p>TODO next: enable mixed precision on TPU and GPU. This should improve the V100 numbers a lot\nThen use a <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\">custom training loop version for TPU</a>. This should speed up the TPU by approx. 5%.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 793600,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-04-01T06:25:33.493000",
          "content": "<p>I completely forgot about this! Thanks for reminding us that this is an option! We can do this in Kaggle and with PyTorch XLA as well? I assume we can just set <code>XLA_USE_BF16</code> to True in our notebook?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794136,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-01T15:37:44.880000",
          "content": "<p>It won't have much effect on Kaggle P100s. Only V100s have proper dedicated mixed precision hardware.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794437,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-04-01T20:23:53.887000",
          "content": "<p>No, I mean for the Kaggle TPUs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794570,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-01T22:41:48.890000",
          "content": "<p>Kaggle GPUs are P100s.\nIf you try it, tell us what you found. I just ran on V100 with mixed precision and XLA and I got a speedup of 13%. I'll be updating the chart above shortly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794575,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-04-01T22:54:38.187000",
          "content": "<p>Hmm ok maybe I don't understand how the half-precision works with TPUs. I thought it I use bfloat16, which is a TPU-specific 16-bit data type, it will speed up TPU computation? Why does it matter what GPU is being used?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794640,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-02T01:10:35.920000",
          "content": "<p>Sorry, I meant \"Kaggle <strong>GPUs</strong> are P100s\". I fixed it above.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 807481,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-04-14T17:00:30.390000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Have you tried to apply mixed precision to transformer with tf.keras? I used below setting from tutorial but failed. It seems inside the transformer some data type is hard coded... If you already did and got successful result, could you please release the benchmark code? Thanks!\n<code>policy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 808869,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-15T17:24:23.380000",
          "content": "<p>Yes, in the three rows in the table marked \"mixed precision XLA\". The <a href=\"https://www.kaggle.com/mgornergoogle/benchmark-jigsaw-multilingual-bert\">benchmark code</a> has this setting too. You can play with it. On a V100 GPU and with this model, mixed precision has a noticeable but not a dramatic effect. It had a bigger effect in the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#784972\">image classification benchmark</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 808952,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-04-15T18:34:02.723000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Oh I missed that one! Thanks, I'll check it:)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 809227,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-04-16T00:45:02.567000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I've tried, but failed... I made a notebook which explains where I'm failing. I believe it's because data type is hardcoded in the huggiung face bert implementation. If so, I should rather ask at hugging face repository, but can you please take a look at what is wrong? Thanks,\n<a href=\"https://www.kaggle.com/bamps53/hugging-face-transformer-tpu-xla?rvi=1\">https://www.kaggle.com/bamps53/hugging-face-transformer-tpu-xla?rvi=1</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 810133,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-16T18:41:26.583000",
          "content": "<p>I looked at the code but I'm not sure what is wrong. It is either something in the implementation of LayerNorm. I'm not sure if this function is from Tensorflow or HuggingFace (in Keras, it's called LayerNormalization). Or maybe it has to do with HuggingFace's embeddings code.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 810145,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-04-16T18:48:04.217000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for help. I also checked hugging face implementation and played around by casting bfloat16 to every tensor which is failed to convert automatically, but they are keep coming... There was no code which specify float32 directly, but somehow some tensor stay float32. <br>\nAnyway thanks, if you couldn't figure out why, definitely me either!😃  I gave up and just post it at hugging face issue.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 810193,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-04-16T19:11:40.620000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> One more thing, I thinks it doesn't come from LayerNormalization, but <a href=\"https://github.com/tensorflow/tensorflow/blob/e5bf8de410005de06a7ff5393fafdf832ef1d4ad/tensorflow/python/keras/layers/embeddings.py#L107\">tf.keras.layers.Embedding</a>. Here kwargs['autocast'] is set to False. Does it related bfloat16 casting??</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 792187,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-31T00:46:13.323000",
      "content": "",
      "votes": 5,
      "replies": [
        {
          "id": 792489,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-31T08:37:42.443000",
          "content": "<p>Maybe we can help with providing numbers with 8 x V100 and enough cpu.  What you do not document is the cpu used on the TPU host.  The fact that the P100 is starving with a 2 core cpu certainly plays a role too.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 792491,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-03-31T08:39:13.277000",
          "content": "<p>I don't think the first notebook link is inaccessible. It is telling me to sign in through Google?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 792903,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-03-31T16:25:09.113000",
          "content": "<p>Sorry for the bad link. fixed!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 792905,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-03-31T16:27:06.827000",
          "content": "<p>Not sure the GPU is starving since in this example, data is entirely in memory. I'll be running these numbers today, we'll see.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 794805,
      "author_name": "Dr. Hemanth Kumar",
      "author_url": "",
      "post_date": "2020-04-02T05:29:44.210000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "792177": "BERT is a large model, requiring powerful hardware to iterate fast.\nI'll be posting various benchmarks numbers here to put some numbers behind \"fast\".",
    "792675": "The speed gains are indeed impressive! What bothers me is that again there are many abstraction layers that don't let me understand the details of the running execution in a level I would love to. I regularly get some \"process died\" messages without any stacktrace, or some weird memory issues. It is very cumbersome to make it work.\n\nThat said though, I only tried pytorch and I know that it is not optimized for that yet. Is debugging way easier when using TF / Keras + TPUs?",
    "794058": "Very nice benchmarks! Have you tried 32gb GPUs, and nvlinked gpus?",
    "800123": "I managed to run this on multi-GPU configurations on GCP. There seems to be something off with MirroredStrategy + BERT in TF 2.1. I was getting slower training times than a single V100. In TF 2.2 it works though. BERT scales nicely to multi-GPU configs (3.7x faster on 4 V100s, 6x faster on 8 V100s) using MirroredStrategy in TF 2.2. The cost per training remains the same or more than a single GPU.",
    "796004": "I published a [response in a forum nobody reads anymore](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/137983#793801), replicating it below.  I see you welcome help in running these benchmarks on GPU, this is great.  We'll come back ASAP.  I also see NVLink is available in your experiments, this is also great.  It makes some of my previous answer irrelevant but I publish it as is still.\n\n-----------\n\nThanks for documenting the host CPU for TPU and for including multi V100 in the benchmark.\n\nFirst, let me start by saying that what follows is my opinion, and just my opinion. It is not an official NVIDIA statement nor does it represent NVIDIA position on this topic. It may well be that my colleagues at NVIDIA will disagree with some of what I say. After all, I have been at NVIDIA for only one month, and I am certainly not a GPU benchmarking expert.\n\nThis said, let's proceed with some items.\n\nFirst, there is obviously a conflict of interest at play here given both Kaggle and TPU are owned by Google. A fair benchmarking would be that Kaggle/Google optimise their code settings for TPU and NVIDIA optimises code settings for GPU. Having you be the judge and one party is biased.\n\nFair benchmarks are defined by a spec agreed upon by all parties, then each party performing their best independently. For this reason I would trust mlperf benchmarks way more than yours.\n\nHere are the latest mlperf results for training https://mlperf.org/training-results-0-6\nThe comparison between TPU and GPU is way more balanced than what you show.\n\nSecond, there are a number of issues here that can explain why your benchmark is at odd with other benchmarks like mlperf. Correct me if I am wrong.\n\n-   You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.\n-  You don't have NVLink between your GPUs\n-  You don't use mixed precision on V100 although it is available. glancing the notebook I found this comment:\n\n    On GPU, specifically V100, mixed precision must be enabled for hardware TensorCores to be used.\n\n-   You don't use TPU v3 yet this is what started this whole topic\n-   The GPU cost is set by Google. We can probably find other providers that run V100 at a lower cost.\n-   I am also puzzled by the OOM you document. Why is TF using more RAM when using GPU than TPU? This has nothing to do with the GPU at first sight.\n-  Did you optimise GPU network bandwidth as recommended on https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth ?\n\nThe good news here is that users get better and better options to train their model over time ;)\n\n",
    "811547": "@mgornergoogle for the benchmark  notebook. At first it used tf 2.1 if we change to use tf-nightly tpu init will fail.  Second if still using tf 2.1 but change a bit to use subclass api then tpu can not work. Can you help me for this Thanks!  \nhttps://github.com/tensorflow/tensorflow/issues/38538",
    "793524": "",
    "792187": "",
    "794805": "Thanks for sharing."
  }
}