{
  "id": 131158,
  "title": "GPUs vs TPUs performance",
  "url": "/competitions/flower-classification-with-tpus/discussion/131158",
  "author_name": "",
  "post_date": "2020-02-18T16:02:43.710328500Z",
  "votes": 7,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I've seen on recent competition some people complaining that they got worse results training on TPUs, of course, that can be caused by many factors including hardware and randomness, but now that we got TPUs on Kaggle maybe this can be better evaluated, I haven't done anything to test that but anyone got results?</p>",
  "messages": [
    {
      "id": "749346",
      "postDate": "02/18/2020 16:02:43",
      "content": "<p>I've seen on recent competition some people complaining that they got worse results training on TPUs, of course, that can be caused by many factors including hardware and randomness, but now that we got TPUs on Kaggle maybe this can be better evaluated, I haven't done anything to test that but anyone got results?</p>",
      "rawMarkdown": "I've seen on recent competition some people complaining that they got worse results training on TPUs, of course, that can be caused by many factors including hardware and randomness, but now that we got TPUs on Kaggle maybe this can be better evaluated, I haven't done anything to test that but anyone got results?",
      "votes": null
    },
    {
      "id": "749444",
      "postDate": "02/18/2020 17:40:22",
      "content": "<p>One overlooked kernel is <a href=\"/mgornergoogle\">@mgornergoogle</a> running my kernel on GPU. Check out the results over <a href=\"https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule\">here</a>. 🤓 </p>",
      "rawMarkdown": "One overlooked kernel is @mgornergoogle running my kernel on GPU. Check out the results over [here](https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule). 🤓",
      "votes": null
    },
    {
      "id": "749487",
      "postDate": "02/18/2020 18:24:29",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> Also check this out <a href=\"https://maelfabien.github.io/bigdata/ColabTPU\">TPU survival guide on Google Colaboratory</a>, it's both minimal and appealing with concepts briefed nicely.</p>",
      "rawMarkdown": "dimitreoliveira Also check this out [TPU survival guide on Google Colaboratory](https://maelfabien.github.io/bigdata/ColabTPU), it's both minimal and appealing with concepts briefed nicely.",
      "votes": null
    },
    {
      "id": "749607",
      "postDate": "02/18/2020 20:00:18",
      "content": "<p>On TPU and GPU, you do not train with the same batch size and therefore you also have to adapt your learning rate to the batch size you use.</p>\n\n<p>So yes, if you take a GPU model and just run it on a TPU with the same settings, the accuracy should be the same but the speed of training will be poor. At small batch sizes, the TPU is not being used to it's full potential.</p>\n\n<p>If you take a GPU model and increase the batch size (x8, a TPU has 8 cores) as you run it on a TPU, training will be faster but now you are training with a learning rate that was optimized for different parameters.</p>\n\n<p>The rule of thumb is to multiply batch size and learning rate by 8 when to switch from GPU to TPU and start from there. Ideal hyperparameters are usually not too far off from that point.</p>\n\n<p>It's about the same type of work as going from a single GPU to a 4x or 8x GPU configuration using MirroredStrategy in tensorflow.</p>",
      "rawMarkdown": "On TPU and GPU, you do not train with the same batch size and therefore you also have to adapt your learning rate to the batch size you use.\n\nSo yes, if you take a GPU model and just run it on a TPU with the same settings, the accuracy should be the same but the speed of training will be poor. At small batch sizes, the TPU is not being used to it's full potential.\n\nIf you take a GPU model and increase the batch size (x8, a TPU has 8 cores) as you run it on a TPU, training will be faster but now you are training with a learning rate that was optimized for different parameters.\n\nThe rule of thumb is to multiply batch size and learning rate by 8 when to switch from GPU to TPU and start from there. Ideal hyperparameters are usually not too far off from that point.\n\nIt's about the same type of work as going from a single GPU to a 4x or 8x GPU configuration using MirroredStrategy in tensorflow.",
      "votes": null
    },
    {
      "id": "749608",
      "postDate": "02/18/2020 20:03:08",
      "content": "<p>As you will notice, I have had to reduce the image size from 512x512px to 224x224px in that notebook to fit in the memory of the GPU. This alone explains a large part of the drop in accuracy.</p>\n\n<p>Also, my goal was to time the runtime, not produce an accurate model so I left the learning rate untouched when i modified the batch size. That is not ideal either.</p>",
      "rawMarkdown": "As you will notice, I have had to reduce the image size from 512x512px to 224x224px in that notebook to fit in the memory of the GPU. This alone explains a large part of the drop in accuracy.\n\nAlso, my goal was to time the runtime, not produce an accurate model so I left the learning rate untouched when i modified the batch size. That is not ideal either.",
      "votes": null
    },
    {
      "id": "749919",
      "postDate": "02/19/2020 02:14:43",
      "content": "<p>Nice!, thanks for the reference.</p>",
      "rawMarkdown": "Nice!, thanks for the reference.",
      "votes": null
    },
    {
      "id": "749922",
      "postDate": "02/19/2020 02:17:22",
      "content": "<p>Those are very great points, thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> </p>",
      "rawMarkdown": "Those are very great points, thanks @mgornergoogle",
      "votes": null
    },
    {
      "id": "750452",
      "postDate": "02/19/2020 11:47:58",
      "content": "<p>It is very much possible that TPUs can give bad results in some cases compared to GPU. Think of it like how some programs can utilize CUDA structure, while some cannot - so the ones which cannot utilize it, don't always get the benefit of an existing GPU, and can only utilize the CPU power. You can probably apply this analogy to GPU vs TPU!</p>",
      "rawMarkdown": "It is very much possible that TPUs can give bad results in some cases compared to GPU. Think of it like how some programs can utilize CUDA structure, while some cannot - so the ones which cannot utilize it, don't always get the benefit of an existing GPU, and can only utilize the CPU power. You can probably apply this analogy to GPU vs TPU!",
      "votes": null
    },
    {
      "id": "750491",
      "postDate": "02/19/2020 12:43:38",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> If I input <code>float32</code> image data into a TPU DenseNet201 compared to into a GPU DenseNet201, do both carry out all operations in <code>float32</code>? Or does the TPU do some operations in <code>bfloat16</code>? (i.e. what are the dtypes of intermediate weights, activations, etc?)</p>",
      "rawMarkdown": "mgornergoogle If I input `float32` image data into a TPU DenseNet201 compared to into a GPU DenseNet201, do both carry out all operations in `float32`? Or does the TPU do some operations in `bfloat16`? (i.e. what are the dtypes of intermediate weights, activations, etc?)",
      "votes": null
    },
    {
      "id": "750516",
      "postDate": "02/19/2020 13:16:42",
      "content": "<p>This is a great question Dimitre. It's hard to compare TPU vs GPU at Kaggle because we don't have equal hardware. Kaggle's TPUv3-8 probably costs around $35,000 whereas Kaggle's GPU P100 costs around $5,000.</p>\n\n<p>The TPUv3-8 is actually 4 chips connected together where each chip is physically similar to the GPU V100 ($8,500), i.e. they both have 32GB memory and similar compute. If you connect 4xGPU V100, then you have a similar comparison to the TPUv3-8 and you will observe that they both operate at similar speeds and similar batch size capacity.</p>\n\n<p>Regarding model accuracy, this is where it gets interesting. If you run the same code on both the TPUv3 and 4xGPU V100, although the training times are similar, the models are dissimilar (beyond expected randomness). I'm still investigating why but I think that the TPU may use mixed precision and therefore calculate different weights during training than the GPU. (Or the distributed gradient descent algorithm could be different).</p>\n\n<p>If anyone knows the mathematical differences between training CNN on TPU versus GPU, I'd love to hear.</p>",
      "rawMarkdown": "This is a great question Dimitre. It's hard to compare TPU vs GPU at Kaggle because we don't have equal hardware. Kaggle's TPUv3-8 probably costs around $35,000 whereas Kaggle's GPU P100 costs around $5,000.\n\nThe TPUv3-8 is actually 4 chips connected together where each chip is physically similar to the GPU V100 ($8,500), i.e. they both have 32GB memory and similar compute. If you connect 4xGPU V100, then you have a similar comparison to the TPUv3-8 and you will observe that they both operate at similar speeds and similar batch size capacity.\n\nRegarding model accuracy, this is where it gets interesting. If you run the same code on both the TPUv3 and 4xGPU V100, although the training times are similar, the models are dissimilar (beyond expected randomness). I'm still investigating why but I think that the TPU may use mixed precision and therefore calculate different weights during training than the GPU. (Or the distributed gradient descent algorithm could be different).\n\nIf anyone knows the mathematical differences between training CNN on TPU versus GPU, I'd love to hear.",
      "votes": null
    },
    {
      "id": "750648",
      "postDate": "02/19/2020 15:15:03",
      "content": "<p><a href=\"/shantanufeb17\">@shantanufeb17</a> , this is a nice point, buy I'm supposing we are using Tensorflow and running the same model structure, like the starting kernel, then we can suppose that Tensorflow uses GPUs as well as TPUs.</p>",
      "rawMarkdown": "shantanufeb17 , this is a nice point, buy I'm supposing we are using Tensorflow and running the same model structure, like the starting kernel, then we can suppose that Tensorflow uses GPUs as well as TPUs.",
      "votes": null
    },
    {
      "id": "750651",
      "postDate": "02/19/2020 15:17:59",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> , this makes things clearer, I mainly ask this because as you said TPUs are a huge improvement, so from now on I'll try to use it whenever I can, but I want to be sure that I'm not losing accuracy points 😄 </p>",
      "rawMarkdown": "Thanks @cdeotte , this makes things clearer, I mainly ask this because as you said TPUs are a huge improvement, so from now on I'll try to use it whenever I can, but I want to be sure that I'm not losing accuracy points 😄",
      "votes": null
    },
    {
      "id": "750977",
      "postDate": "02/19/2020 22:08:06",
      "content": "<p>The TPU will perform matrix multiplications (i.e. dense and convolutional layers) on the hadrware matrix multiplication unit (MXU).</p>\n\n<p>The MXU works with bfloat16 inputs, float32 accumulators and float32 outputs.</p>\n\n<p>In more detail: conversion of your matrices from float32 to bfloat16 is performed by the MXU automatically. Then multiplying two bfloat16 numbers together naturally produces a float32 result which is kept as float32. All additions happen in float32 and the final result is float32.</p>",
      "rawMarkdown": "The TPU will perform matrix multiplications (i.e. dense and convolutional layers) on the hadrware matrix multiplication unit (MXU).\n\nThe MXU works with bfloat16 inputs, float32 accumulators and float32 outputs.\n\nIn more detail: conversion of your matrices from float32 to bfloat16 is performed by the MXU automatically. Then multiplying two bfloat16 numbers together naturally produces a float32 result which is kept as float32. All additions happen in float32 and the final result is float32.",
      "votes": null
    },
    {
      "id": "2088559",
      "postDate": "01/06/2023 12:36:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> where is this mentioned about 4 GPU V100 being roughly equal to 1 TPUv3?</p>",
      "rawMarkdown": "Hi @cdeotte where is this mentioned about 4 GPU V100 being roughly equal to 1 TPUv3?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 749444,
      "author_name": "msheriey",
      "author_url": "",
      "post_date": "02/18/2020 17:40:22",
      "content": "<p>One overlooked kernel is <a href=\"/mgornergoogle\">@mgornergoogle</a> running my kernel on GPU. Check out the results over <a href=\"https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule\">here</a>. 🤓 </p>",
      "votes": null,
      "replies": [
        {
          "id": 749608,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 20:03:08",
          "content": "<p>As you will notice, I have had to reduce the image size from 512x512px to 224x224px in that notebook to fit in the memory of the GPU. This alone explains a large part of the drop in accuracy.</p>\n\n<p>Also, my goal was to time the runtime, not produce an accurate model so I left the learning rate untouched when i modified the batch size. That is not ideal either.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749487,
      "author_name": "msheriey",
      "author_url": "",
      "post_date": "02/18/2020 18:24:29",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> Also check this out <a href=\"https://maelfabien.github.io/bigdata/ColabTPU\">TPU survival guide on Google Colaboratory</a>, it's both minimal and appealing with concepts briefed nicely.</p>",
      "votes": null,
      "replies": [
        {
          "id": 749919,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "02/19/2020 02:14:43",
          "content": "<p>Nice!, thanks for the reference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749607,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/18/2020 20:00:18",
      "content": "<p>On TPU and GPU, you do not train with the same batch size and therefore you also have to adapt your learning rate to the batch size you use.</p>\n\n<p>So yes, if you take a GPU model and just run it on a TPU with the same settings, the accuracy should be the same but the speed of training will be poor. At small batch sizes, the TPU is not being used to it's full potential.</p>\n\n<p>If you take a GPU model and increase the batch size (x8, a TPU has 8 cores) as you run it on a TPU, training will be faster but now you are training with a learning rate that was optimized for different parameters.</p>\n\n<p>The rule of thumb is to multiply batch size and learning rate by 8 when to switch from GPU to TPU and start from there. Ideal hyperparameters are usually not too far off from that point.</p>\n\n<p>It's about the same type of work as going from a single GPU to a 4x or 8x GPU configuration using MirroredStrategy in tensorflow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 749922,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "02/19/2020 02:17:22",
          "content": "<p>Those are very great points, thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750491,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2020 12:43:38",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> If I input <code>float32</code> image data into a TPU DenseNet201 compared to into a GPU DenseNet201, do both carry out all operations in <code>float32</code>? Or does the TPU do some operations in <code>bfloat16</code>? (i.e. what are the dtypes of intermediate weights, activations, etc?)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750977,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/19/2020 22:08:06",
          "content": "<p>The TPU will perform matrix multiplications (i.e. dense and convolutional layers) on the hadrware matrix multiplication unit (MXU).</p>\n\n<p>The MXU works with bfloat16 inputs, float32 accumulators and float32 outputs.</p>\n\n<p>In more detail: conversion of your matrices from float32 to bfloat16 is performed by the MXU automatically. Then multiplying two bfloat16 numbers together naturally produces a float32 result which is kept as float32. All additions happen in float32 and the final result is float32.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 750452,
      "author_name": "shantanufeb17",
      "author_url": "",
      "post_date": "02/19/2020 11:47:58",
      "content": "<p>It is very much possible that TPUs can give bad results in some cases compared to GPU. Think of it like how some programs can utilize CUDA structure, while some cannot - so the ones which cannot utilize it, don't always get the benefit of an existing GPU, and can only utilize the CPU power. You can probably apply this analogy to GPU vs TPU!</p>",
      "votes": null,
      "replies": [
        {
          "id": 750648,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "02/19/2020 15:15:03",
          "content": "<p><a href=\"/shantanufeb17\">@shantanufeb17</a> , this is a nice point, buy I'm supposing we are using Tensorflow and running the same model structure, like the starting kernel, then we can suppose that Tensorflow uses GPUs as well as TPUs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 750516,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/19/2020 13:16:42",
      "content": "<p>This is a great question Dimitre. It's hard to compare TPU vs GPU at Kaggle because we don't have equal hardware. Kaggle's TPUv3-8 probably costs around $35,000 whereas Kaggle's GPU P100 costs around $5,000.</p>\n\n<p>The TPUv3-8 is actually 4 chips connected together where each chip is physically similar to the GPU V100 ($8,500), i.e. they both have 32GB memory and similar compute. If you connect 4xGPU V100, then you have a similar comparison to the TPUv3-8 and you will observe that they both operate at similar speeds and similar batch size capacity.</p>\n\n<p>Regarding model accuracy, this is where it gets interesting. If you run the same code on both the TPUv3 and 4xGPU V100, although the training times are similar, the models are dissimilar (beyond expected randomness). I'm still investigating why but I think that the TPU may use mixed precision and therefore calculate different weights during training than the GPU. (Or the distributed gradient descent algorithm could be different).</p>\n\n<p>If anyone knows the mathematical differences between training CNN on TPU versus GPU, I'd love to hear.</p>",
      "votes": null,
      "replies": [
        {
          "id": 750651,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "02/19/2020 15:17:59",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> , this makes things clearer, I mainly ask this because as you said TPUs are a huge improvement, so from now on I'll try to use it whenever I can, but I want to be sure that I'm not losing accuracy points 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2088559,
          "author_name": "eashish",
          "author_url": "",
          "post_date": "01/06/2023 12:36:54",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> where is this mentioned about 4 GPU V100 being roughly equal to 1 TPUv3?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "749346": "I've seen on recent competition some people complaining that they got worse results training on TPUs, of course, that can be caused by many factors including hardware and randomness, but now that we got TPUs on Kaggle maybe this can be better evaluated, I haven't done anything to test that but anyone got results?",
    "749444": "One overlooked kernel is @mgornergoogle running my kernel on GPU. Check out the results over [here](https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule). 🤓",
    "749487": "dimitreoliveira Also check this out [TPU survival guide on Google Colaboratory](https://maelfabien.github.io/bigdata/ColabTPU), it's both minimal and appealing with concepts briefed nicely.",
    "749607": "On TPU and GPU, you do not train with the same batch size and therefore you also have to adapt your learning rate to the batch size you use.\n\nSo yes, if you take a GPU model and just run it on a TPU with the same settings, the accuracy should be the same but the speed of training will be poor. At small batch sizes, the TPU is not being used to it's full potential.\n\nIf you take a GPU model and increase the batch size (x8, a TPU has 8 cores) as you run it on a TPU, training will be faster but now you are training with a learning rate that was optimized for different parameters.\n\nThe rule of thumb is to multiply batch size and learning rate by 8 when to switch from GPU to TPU and start from there. Ideal hyperparameters are usually not too far off from that point.\n\nIt's about the same type of work as going from a single GPU to a 4x or 8x GPU configuration using MirroredStrategy in tensorflow.",
    "749608": "As you will notice, I have had to reduce the image size from 512x512px to 224x224px in that notebook to fit in the memory of the GPU. This alone explains a large part of the drop in accuracy.\n\nAlso, my goal was to time the runtime, not produce an accurate model so I left the learning rate untouched when i modified the batch size. That is not ideal either.",
    "749919": "Nice!, thanks for the reference.",
    "749922": "Those are very great points, thanks @mgornergoogle",
    "750452": "It is very much possible that TPUs can give bad results in some cases compared to GPU. Think of it like how some programs can utilize CUDA structure, while some cannot - so the ones which cannot utilize it, don't always get the benefit of an existing GPU, and can only utilize the CPU power. You can probably apply this analogy to GPU vs TPU!",
    "750491": "mgornergoogle If I input `float32` image data into a TPU DenseNet201 compared to into a GPU DenseNet201, do both carry out all operations in `float32`? Or does the TPU do some operations in `bfloat16`? (i.e. what are the dtypes of intermediate weights, activations, etc?)",
    "750516": "This is a great question Dimitre. It's hard to compare TPU vs GPU at Kaggle because we don't have equal hardware. Kaggle's TPUv3-8 probably costs around $35,000 whereas Kaggle's GPU P100 costs around $5,000.\n\nThe TPUv3-8 is actually 4 chips connected together where each chip is physically similar to the GPU V100 ($8,500), i.e. they both have 32GB memory and similar compute. If you connect 4xGPU V100, then you have a similar comparison to the TPUv3-8 and you will observe that they both operate at similar speeds and similar batch size capacity.\n\nRegarding model accuracy, this is where it gets interesting. If you run the same code on both the TPUv3 and 4xGPU V100, although the training times are similar, the models are dissimilar (beyond expected randomness). I'm still investigating why but I think that the TPU may use mixed precision and therefore calculate different weights during training than the GPU. (Or the distributed gradient descent algorithm could be different).\n\nIf anyone knows the mathematical differences between training CNN on TPU versus GPU, I'd love to hear.",
    "750648": "shantanufeb17 , this is a nice point, buy I'm supposing we are using Tensorflow and running the same model structure, like the starting kernel, then we can suppose that Tensorflow uses GPUs as well as TPUs.",
    "750651": "Thanks @cdeotte , this makes things clearer, I mainly ask this because as you said TPUs are a huge improvement, so from now on I'll try to use it whenever I can, but I want to be sure that I'm not losing accuracy points 😄",
    "750977": "The TPU will perform matrix multiplications (i.e. dense and convolutional layers) on the hadrware matrix multiplication unit (MXU).\n\nThe MXU works with bfloat16 inputs, float32 accumulators and float32 outputs.\n\nIn more detail: conversion of your matrices from float32 to bfloat16 is performed by the MXU automatically. Then multiplying two bfloat16 numbers together naturally produces a float32 result which is kept as float32. All additions happen in float32 and the final result is float32.",
    "2088559": "Hi @cdeotte where is this mentioned about 4 GPU V100 being roughly equal to 1 TPUv3?"
  },
  "source": "meta"
}