{
  "id": 137983,
  "title": "Ok I get it... TPUs are awesome! So what are the main reasons they're not ubiquitous?",
  "url": "/competitions/flower-classification-with-tpus/discussion/137983",
  "author_name": "Alexander Soare",
  "post_date": "2020-03-23T09:58:58.702000",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Wow these TPUs are incredible! To me it looks like everyone else is out in the fields plowing with shovels (GPUs) when there are bulldozers (TPUs) available.</p>\n\n<p>I exaggerate a little on purpose 😏 . I don't doubt that I'm missing out on some crucial piece of context, and would love for someone to enlighten me.</p>\n\n<p>So what's the TPU's role in AI today and what will it be tomorrow? What about GPU's?</p>\n\n<p>PS: Thank you Kaggle for giving me a chance to work with this technology. What a time to be alive!</p>",
  "messages": [
    {
      "id": 783417,
      "postDate": "2020-03-23T09:58:58.703Z",
      "content": "<p>Wow these TPUs are incredible! To me it looks like everyone else is out in the fields plowing with shovels (GPUs) when there are bulldozers (TPUs) available.</p>\n\n<p>I exaggerate a little on purpose 😏 . I don't doubt that I'm missing out on some crucial piece of context, and would love for someone to enlighten me.</p>\n\n<p>So what's the TPU's role in AI today and what will it be tomorrow? What about GPU's?</p>\n\n<p>PS: Thank you Kaggle for giving me a chance to work with this technology. What a time to be alive!</p>",
      "rawMarkdown": "Wow these TPUs are incredible! To me it looks like everyone else is out in the fields plowing with shovels (GPUs) when there are bulldozers (TPUs) available.\n\nI exaggerate a little on purpose 😏 . I don't doubt that I'm missing out on some crucial piece of context, and would love for someone to enlighten me.\n\nSo what's the TPU's role in AI today and what will it be tomorrow? What about GPU's?\n\nPS: Thank you Kaggle for giving me a chance to work with this technology. What a time to be alive!",
      "votes": 6
    },
    {
      "id": 784524,
      "postDate": "2020-03-24T09:33:14.493Z",
      "content": "<p>The main reason is that you are comparing a machine with 8 cores recent TPU plus 4 cpu cores against a machine with a single P100, a 4 years old gpu, and only 2 cpu cores.  Cost is not the same at all either, TPU is way more expensive to run than the P100.  </p>\n\n<p>If you get for free a much more recent and powerful machine then no wonder you prefer the much more recent and powerful machine.</p>\n\n<p>The right comparison should be between a 4 cores V100 (eg a DGX Station) and the 8 cores TPU v3.8.   That's the ballpark: a V100 is worth 2 TPU v3 cores roughly.</p>\n\n<p>It would also be interesting to know why TPU aren't available for scoring in kernel only competitions.</p>",
      "rawMarkdown": "The main reason is that you are comparing a machine with 8 cores recent TPU plus 4 cpu cores against a machine with a single P100, a 4 years old gpu, and only 2 cpu cores.  Cost is not the same at all either, TPU is way more expensive to run than the P100.  \n\nIf you get for free a much more recent and powerful machine then no wonder you prefer the much more recent and powerful machine.\n\nThe right comparison should be between a 4 cores V100 (eg a DGX Station) and the 8 cores TPU v3.8.   That's the ballpark: a V100 is worth 2 TPU v3 cores roughly.\n\nIt would also be interesting to know why TPU aren't available for scoring in kernel only competitions.",
      "votes": 4,
      "replies": [
        {
          "id": 784532,
          "postDate": "2020-03-24T09:40:32.217Z",
          "content": "<p>Great answer! Thanks that's perfectly clear.</p>",
          "rawMarkdown": "Great answer! Thanks that's perfectly clear.",
          "votes": 1
        },
        {
          "id": 784566,
          "postDate": "2020-03-24T10:27:32.250Z",
          "content": "<blockquote>\n  <p>It would also be interesting to know why TPU aren't available for scoring in kernel only competitions.</p>\n</blockquote>\n\n<p>I see they seem available in the Toxic comment comp.  They were not available in Bengali, maybe  because  internet access was disabled for scoring there.</p>",
          "rawMarkdown": "&gt; It would also be interesting to know why TPU aren't available for scoring in kernel only competitions.\n\nI see they seem available in the Toxic comment comp.  They were not available in Bengali, maybe  because  internet access was disabled for scoring there.",
          "votes": 1
        },
        {
          "id": 784972,
          "postDate": "2020-03-24T16:47:36.190Z",
          "rawMarkdown": "",
          "votes": 6,
          "isDeleted": true
        },
        {
          "id": 792422,
          "postDate": "2020-03-31T07:28:37.353Z",
          "content": "<p>It's always interesting to see Nvidia and google folks competing their ml hardware at conferences, but did not expect to see this at Kaggle! \nIts a pleasure to get the availability of TPUs for free; would be great to test out DGXs as cloud infrastructure.</p>",
          "rawMarkdown": "It's always interesting to see Nvidia and google folks competing their ml hardware at conferences, but did not expect to see this at Kaggle! \nIts a pleasure to get the availability of TPUs for free; would be great to test out DGXs as cloud infrastructure.",
          "votes": 2
        },
        {
          "id": 793801,
          "postDate": "2020-04-01T09:42:56.997Z",
          "content": "<p>Thanks for documenting the host CPU for TPU and for including multi V100 in the benchmark.</p>\n\n<p>First, let me start by saying that what follows is my opinion, and just my opinion.  It is not an official NVIDIA statement nor does it represent NVIDIA position on this topic.  It may well be that my colleagues at NVIDIA will disagree with some of what I say.  After all, I have been at NVIDIA for only one month, and I am certainly not a GPU benchmarking expert.</p>\n\n<p>This said, let's proceed with some items.</p>\n\n<p>First, there is obviously a conflict of interest at play here given both Kaggle and TPU are owned by Google.  A fair benchmarking would be that Kaggle/Google optimise their code settings for TPU and NVIDIA optimises code settings for GPU.  Having you be the judge and one party is biased.  </p>\n\n<p>Fair benchmarks are defined by a spec agreed upon by all parties, then each party performing their best independently.  For this reason I would trust mlperf benchmarks way more than yours.  </p>\n\n<p>Here are the latest mlperf results for training <a href=\"https://mlperf.org/training-results-0-6\">https://mlperf.org/training-results-0-6</a>\nThe comparison between TPU and GPU is way more balanced than what you show.</p>\n\n<p>Second, there are a number of issues here that can explain why your benchmark is at odd with other benchmarks like mlperf.  Correct me if I am wrong.</p>\n\n<ul>\n<li>You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.</li>\n<li>You don't have NVLink between your GPUs</li>\n<li><p>You don't use mixed precision on V100 although it is available.  glancing the notebook I found this comment:</p>\n\n<blockquote>\n  <p>On GPU, specifically V100, mixed precision must be enabled for hardware TensorCores to be used.</p>\n</blockquote></li>\n<li><p>You don't use TPU v3 yet this is what started this whole topic</p></li>\n<li>The GPU cost is set by Google.  We can probably find other providers that run V100 at a lower cost.</li>\n<li>I am also puzzled by the OOM you document.  Why is TF using more RAM when using GPU than TPU?  This has nothing to do with the GPU at first sight.  </li>\n<li>Did you optimise GPU network bandwidth as recommended on <a href=\"https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth\">https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth</a> ?</li>\n</ul>\n\n<p>The good news here is that users get better and better options to train their model over time ;)</p>",
          "rawMarkdown": "Thanks for documenting the host CPU for TPU and for including multi V100 in the benchmark.\n\nFirst, let me start by saying that what follows is my opinion, and just my opinion.  It is not an official NVIDIA statement nor does it represent NVIDIA position on this topic.  It may well be that my colleagues at NVIDIA will disagree with some of what I say.  After all, I have been at NVIDIA for only one month, and I am certainly not a GPU benchmarking expert.\n\nThis said, let's proceed with some items.\n\nFirst, there is obviously a conflict of interest at play here given both Kaggle and TPU are owned by Google.  A fair benchmarking would be that Kaggle/Google optimise their code settings for TPU and NVIDIA optimises code settings for GPU.  Having you be the judge and one party is biased.  \n\nFair benchmarks are defined by a spec agreed upon by all parties, then each party performing their best independently.  For this reason I would trust mlperf benchmarks way more than yours.  \n\nHere are the latest mlperf results for training https://mlperf.org/training-results-0-6\nThe comparison between TPU and GPU is way more balanced than what you show.\n\nSecond, there are a number of issues here that can explain why your benchmark is at odd with other benchmarks like mlperf.  Correct me if I am wrong.\n\n- You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.\n- You don't have NVLink between your GPUs\n- You don't use mixed precision on V100 although it is available.  glancing the notebook I found this comment:\n&gt; On GPU, specifically V100, mixed precision must be enabled for hardware TensorCores to be used.\n\n- You don't use TPU v3 yet this is what started this whole topic\n- The GPU cost is set by Google.  We can probably find other providers that run V100 at a lower cost.\n- I am also puzzled by the OOM you document.  Why is TF using more RAM when using GPU than TPU?  This has nothing to do with the GPU at first sight.  \n- Did you optimise GPU network bandwidth as recommended on https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth ?\n\nThe good news here is that users get better and better options to train their model over time ;)",
          "votes": 3
        }
      ]
    },
    {
      "id": 784059,
      "postDate": "2020-03-23T23:03:47.167Z",
      "content": "<p>I think the main reason is that before Tensorflow 2.1 introduced Keras support for TPUs, they were a bit hard to use.</p>",
      "rawMarkdown": "I think the main reason is that before Tensorflow 2.1 introduced Keras support for TPUs, they were a bit hard to use.",
      "votes": 1,
      "replies": [
        {
          "id": 784449,
          "postDate": "2020-03-24T08:14:19.187Z",
          "content": "<p>Do you think we'll say good bye to gpu's once tpu's become just as usable?</p>",
          "rawMarkdown": "Do you think we'll say good bye to gpu's once tpu's become just as usable?",
          "votes": 1
        },
        {
          "id": 785380,
          "postDate": "2020-03-25T02:29:08.383Z",
          "content": "<p>They are just as usable now. The TPU and multi-GPU APIs are actually the same in Tensorflow, both for Keras model.fit (TPUStrategy / MirrorredStrategy are both a DistributionStrategy) and <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">custom training loop</a>.</p>",
          "rawMarkdown": "They are just as usable now. The TPU and multi-GPU APIs are actually the same in Tensorflow, both for Keras model.fit (TPUStrategy / MirrorredStrategy are both a DistributionStrategy) and [custom training loop](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443).",
          "votes": 1
        },
        {
          "id": 785636,
          "postDate": "2020-03-25T08:18:28.253Z",
          "content": "<p>Interesting. And thanks for sharing that table! So if usability/price is equal/better at a similar performance level, all I can think of that's holding GPU's up is momentum, awareness and that <strong>single</strong> gpus are easier to use.</p>",
          "rawMarkdown": "Interesting. And thanks for sharing that table! So if usability/price is equal/better at a similar performance level, all I can think of that's holding GPU's up is momentum, awareness and that **single** gpus are easier to use.",
          "votes": 1
        },
        {
          "id": 801552,
          "postDate": "2020-04-08T15:28:00.857Z",
          "content": "<p>I think nwe they are. Just that the availability for everyone will still be very unlikely. \nTPUs on Kaggle seems faster than one on Colab.</p>",
          "rawMarkdown": "I think nwe they are. Just that the availability for everyone will still be very unlikely. \nTPUs on Kaggle seems faster than one on Colab."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 784524,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-03-24T09:33:14.493000",
      "content": "<p>The main reason is that you are comparing a machine with 8 cores recent TPU plus 4 cpu cores against a machine with a single P100, a 4 years old gpu, and only 2 cpu cores.  Cost is not the same at all either, TPU is way more expensive to run than the P100.  </p>\n\n<p>If you get for free a much more recent and powerful machine then no wonder you prefer the much more recent and powerful machine.</p>\n\n<p>The right comparison should be between a 4 cores V100 (eg a DGX Station) and the 8 cores TPU v3.8.   That's the ballpark: a V100 is worth 2 TPU v3 cores roughly.</p>\n\n<p>It would also be interesting to know why TPU aren't available for scoring in kernel only competitions.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 784532,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2020-03-24T09:40:32.217000",
          "content": "<p>Great answer! Thanks that's perfectly clear.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 784566,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-24T10:27:32.250000",
          "content": "<blockquote>\n  <p>It would also be interesting to know why TPU aren't available for scoring in kernel only competitions.</p>\n</blockquote>\n\n<p>I see they seem available in the Toxic comment comp.  They were not available in Bengali, maybe  because  internet access was disabled for scoring there.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 784972,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-24T16:47:36.190000",
          "content": "",
          "votes": 6,
          "replies": []
        },
        {
          "id": 792422,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2020-03-31T07:28:37.353000",
          "content": "<p>It's always interesting to see Nvidia and google folks competing their ml hardware at conferences, but did not expect to see this at Kaggle! \nIts a pleasure to get the availability of TPUs for free; would be great to test out DGXs as cloud infrastructure.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 793801,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-04-01T09:42:56.997000",
          "content": "<p>Thanks for documenting the host CPU for TPU and for including multi V100 in the benchmark.</p>\n\n<p>First, let me start by saying that what follows is my opinion, and just my opinion.  It is not an official NVIDIA statement nor does it represent NVIDIA position on this topic.  It may well be that my colleagues at NVIDIA will disagree with some of what I say.  After all, I have been at NVIDIA for only one month, and I am certainly not a GPU benchmarking expert.</p>\n\n<p>This said, let's proceed with some items.</p>\n\n<p>First, there is obviously a conflict of interest at play here given both Kaggle and TPU are owned by Google.  A fair benchmarking would be that Kaggle/Google optimise their code settings for TPU and NVIDIA optimises code settings for GPU.  Having you be the judge and one party is biased.  </p>\n\n<p>Fair benchmarks are defined by a spec agreed upon by all parties, then each party performing their best independently.  For this reason I would trust mlperf benchmarks way more than yours.  </p>\n\n<p>Here are the latest mlperf results for training <a href=\"https://mlperf.org/training-results-0-6\">https://mlperf.org/training-results-0-6</a>\nThe comparison between TPU and GPU is way more balanced than what you show.</p>\n\n<p>Second, there are a number of issues here that can explain why your benchmark is at odd with other benchmarks like mlperf.  Correct me if I am wrong.</p>\n\n<ul>\n<li>You use old V100 with 16 GB memory, and not the more recent v100 with 32 GB.</li>\n<li>You don't have NVLink between your GPUs</li>\n<li><p>You don't use mixed precision on V100 although it is available.  glancing the notebook I found this comment:</p>\n\n<blockquote>\n  <p>On GPU, specifically V100, mixed precision must be enabled for hardware TensorCores to be used.</p>\n</blockquote></li>\n<li><p>You don't use TPU v3 yet this is what started this whole topic</p></li>\n<li>The GPU cost is set by Google.  We can probably find other providers that run V100 at a lower cost.</li>\n<li>I am also puzzled by the OOM you document.  Why is TF using more RAM when using GPU than TPU?  This has nothing to do with the GPU at first sight.  </li>\n<li>Did you optimise GPU network bandwidth as recommended on <a href=\"https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth\">https://cloud.google.com/compute/docs/gpus/optimize-gpus#high-bandwidth</a> ?</li>\n</ul>\n\n<p>The good news here is that users get better and better options to train their model over time ;)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 784059,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-03-23T23:03:47.167000",
      "content": "<p>I think the main reason is that before Tensorflow 2.1 introduced Keras support for TPUs, they were a bit hard to use.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 784449,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2020-03-24T08:14:19.187000",
          "content": "<p>Do you think we'll say good bye to gpu's once tpu's become just as usable?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 785380,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-03-25T02:29:08.383000",
          "content": "<p>They are just as usable now. The TPU and multi-GPU APIs are actually the same in Tensorflow, both for Keras model.fit (TPUStrategy / MirrorredStrategy are both a DistributionStrategy) and <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">custom training loop</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 785636,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2020-03-25T08:18:28.253000",
          "content": "<p>Interesting. And thanks for sharing that table! So if usability/price is equal/better at a similar performance level, all I can think of that's holding GPU's up is momentum, awareness and that <strong>single</strong> gpus are easier to use.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 801552,
          "author_name": "Suraj Parmar",
          "author_url": "",
          "post_date": "2020-04-08T15:28:00.857000",
          "content": "<p>I think nwe they are. Just that the availability for everyone will still be very unlikely. \nTPUs on Kaggle seems faster than one on Colab.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "783417": "Wow these TPUs are incredible! To me it looks like everyone else is out in the fields plowing with shovels (GPUs) when there are bulldozers (TPUs) available.\n\nI exaggerate a little on purpose 😏 . I don't doubt that I'm missing out on some crucial piece of context, and would love for someone to enlighten me.\n\nSo what's the TPU's role in AI today and what will it be tomorrow? What about GPU's?\n\nPS: Thank you Kaggle for giving me a chance to work with this technology. What a time to be alive!",
    "784524": "The main reason is that you are comparing a machine with 8 cores recent TPU plus 4 cpu cores against a machine with a single P100, a 4 years old gpu, and only 2 cpu cores.  Cost is not the same at all either, TPU is way more expensive to run than the P100.  \n\nIf you get for free a much more recent and powerful machine then no wonder you prefer the much more recent and powerful machine.\n\nThe right comparison should be between a 4 cores V100 (eg a DGX Station) and the 8 cores TPU v3.8.   That's the ballpark: a V100 is worth 2 TPU v3 cores roughly.\n\nIt would also be interesting to know why TPU aren't available for scoring in kernel only competitions.",
    "784059": "I think the main reason is that before Tensorflow 2.1 introduced Keras support for TPUs, they were a bit hard to use."
  }
}