{
  "id": 130199,
  "title": "How does TensorFlow use multiple devices?",
  "url": "/competitions/flower-classification-with-tpus/discussion/130199",
  "author_name": "",
  "post_date": "2020-02-12T17:22:55.524236Z",
  "votes": 24,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I have a technical question. Does anyone know how TensorFlow uses multiple TPU and/or multiple GPU during back propagation training of neural networks?</p>\n\n<p>For example, with TPU when we say <code>batch_size = 16 * 8</code> which equals 128, does each TPU get a batch of 16 and the model weights <code>W0</code> and activations <code>A0</code> at time = <code>T0</code>.  Then does each core compute the gradient with <code>W0</code> and <code>A0</code> and then every core updates the original weights <code>W0</code> by adding their gradient updates?</p>\n\n<p>Note that this is different than training on 1 core where if we perform 8 batches, then first batch uses <code>W0, A0</code> and second batch uses <code>W1, A1</code> (the updated weights and activations after first time step), and third batch uses <code>W2, A2</code> etc.</p>",
  "messages": [
    {
      "id": "744250",
      "postDate": "02/12/2020 17:22:55",
      "content": "<p>I have a technical question. Does anyone know how TensorFlow uses multiple TPU and/or multiple GPU during back propagation training of neural networks?</p>\n\n<p>For example, with TPU when we say <code>batch_size = 16 * 8</code> which equals 128, does each TPU get a batch of 16 and the model weights <code>W0</code> and activations <code>A0</code> at time = <code>T0</code>.  Then does each core compute the gradient with <code>W0</code> and <code>A0</code> and then every core updates the original weights <code>W0</code> by adding their gradient updates?</p>\n\n<p>Note that this is different than training on 1 core where if we perform 8 batches, then first batch uses <code>W0, A0</code> and second batch uses <code>W1, A1</code> (the updated weights and activations after first time step), and third batch uses <code>W2, A2</code> etc.</p>",
      "rawMarkdown": "I have a technical question. Does anyone know how TensorFlow uses multiple TPU and/or multiple GPU during back propagation training of neural networks?\n\nFor example, with TPU when we say `batch_size = 16 * 8` which equals 128, does each TPU get a batch of 16 and the model weights `W0` and activations `A0` at time = `T0`.  Then does each core compute the gradient with `W0` and `A0` and then every core updates the original weights `W0` by adding their gradient updates?\n\nNote that this is different than training on 1 core where if we perform 8 batches, then first batch uses `W0, A0` and second batch uses `W1, A1` (the updated weights and activations after first time step), and third batch uses `W2, A2` etc.",
      "votes": null
    },
    {
      "id": "744272",
      "postDate": "02/12/2020 17:34:35",
      "content": "<p>Not sure if this is gonna help but this a great <a href=\"https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255\">article</a> by Hugging-Face!</p>",
      "rawMarkdown": "Not sure if this is gonna help but this a great [article](https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255) by Hugging-Face!",
      "votes": null
    },
    {
      "id": "744295",
      "postDate": "02/12/2020 17:47:42",
      "content": "<p>Thanks I'll read that. How models use multiple TPUs or GPUs will affect how we think about hyperparameter choices.</p>",
      "rawMarkdown": "Thanks I'll read that. How models use multiple TPUs or GPUs will affect how we think about hyperparameter choices.",
      "votes": null
    },
    {
      "id": "744352",
      "postDate": "02/12/2020 18:59:13",
      "content": "<p>Yes I think that's what is happening. If you have 8 replicas (e.g. using TPU v3-8) then the total batch_size = 16 * 8.  Each TPU core gets 16 samples. You can choose how you <a href=\"https://github.com/see--/natural-question-answering/blob/master/train_eval.py#L126\">reduce</a> the loss along multiple cores. This will determine the gradients: <code>tf.distribute.ReduceOp.MEAN</code> or <code>tf.distribute.ReduceOp.SUM</code>. To add <a href=\"https://www.tensorflow.org/guide/distributed_training#tpustrategy\">TPUStrategy is same as MirroredStrategy</a>.</p>\n\n<p>So far, using <code>MEAN</code> and linearly scaling the learning rate as if I had one big batch_size of 16 * 8 works well.</p>",
      "rawMarkdown": "Yes I think that's what is happening. If you have 8 replicas (e.g. using TPU v3-8) then the total batch_size = 16 * 8.  Each TPU core gets 16 samples. You can choose how you [reduce](https://github.com/see--/natural-question-answering/blob/master/train_eval.py#L126) the loss along multiple cores. This will determine the gradients: `tf.distribute.ReduceOp.MEAN` or `tf.distribute.ReduceOp.SUM`. To add [TPUStrategy is same as MirroredStrategy](https://www.tensorflow.org/guide/distributed_training#tpustrategy).\n\nSo far, using `MEAN` and linearly scaling the learning rate as if I had one big batch_size of 16 * 8 works well.",
      "votes": null
    },
    {
      "id": "744375",
      "postDate": "02/12/2020 19:21:04",
      "content": "<p>Glad you asked! This is all implemented in Tensorflow distribution strategies. For TPU strategy specifically:</p>\n\n<ul>\n<li>data batches are split between the 8 TPU cores</li>\n<li>gradients are computed on each batch in parallel and merged</li>\n<li>all model weights are updated</li>\n</ul>\n\n<p>There is actually more work performed if you are running on a TPU pod (TPU = 1 board/4 chips/8cores, TPU pod = 2 to 256 TPU boards on a high-performance interconnect exposed to the user as a single accelerator). On a TPU pod:</p>\n\n<ul>\n<li>the model is replicated across all TPU boards</li>\n<li>the datset (tf.data.Dataset) is split at the file level between all TPU boards. Here it becomes important that your dataset is sharded across multiple files. TPU boards that do not need certain shards will never load them. If you have fewer files (shards) than TPU boards, you will actually get an error message on a TPU pod.</li>\n<li>batches of data are loaded on each TPU board and split between TPU cores.</li>\n<li>gradients are computed and merged across cores and boards and all model weights are updated (all reduce algorithm - this is where the fast interconnect between TPU boards comes into play)</li>\n</ul>\n\n<p>All this is happening automatically in tf.data.Dataset and the TPUStrategy code. The user just supplies a tf.data.Dataset with a larger batch size. The more TPU cores, the larger the batch size.</p>",
      "rawMarkdown": "Glad you asked! This is all implemented in Tensorflow distribution strategies. For TPU strategy specifically:\n\n- data batches are split between the 8 TPU cores\n- gradients are computed on each batch in parallel and merged\n- all model weights are updated\n\nThere is actually more work performed if you are running on a TPU pod (TPU = 1 board/4 chips/8cores, TPU pod = 2 to 256 TPU boards on a high-performance interconnect exposed to the user as a single accelerator). On a TPU pod:\n\n- the model is replicated across all TPU boards\n- the datset (tf.data.Dataset) is split at the file level between all TPU boards. Here it becomes important that your dataset is sharded across multiple files. TPU boards that do not need certain shards will never load them. If you have fewer files (shards) than TPU boards, you will actually get an error message on a TPU pod.\n- batches of data are loaded on each TPU board and split between TPU cores.\n- gradients are computed and merged across cores and boards and all model weights are updated (all reduce algorithm - this is where the fast interconnect between TPU boards comes into play)\n\nAll this is happening automatically in tf.data.Dataset and the TPUStrategy code. The user just supplies a tf.data.Dataset with a larger batch size. The more TPU cores, the larger the batch size.",
      "votes": null
    },
    {
      "id": "744379",
      "postDate": "02/12/2020 19:24:27",
      "content": "<p>And yes, if you use larger batches, you will need a bigger learning rate too. Start scaling the batch size linearly with core count but then expect some hyperparam tuning work to find the best LR schedule for a given batch size and TPU configuration.</p>",
      "rawMarkdown": "And yes, if you use larger batches, you will need a bigger learning rate too. Start scaling the batch size linearly with core count but then expect some hyperparam tuning work to find the best LR schedule for a given batch size and TPU configuration.",
      "votes": null
    },
    {
      "id": "744556",
      "postDate": "02/12/2020 23:40:51",
      "content": "<p>Thanks for clarification</p>",
      "rawMarkdown": "Thanks for clarification",
      "votes": null
    },
    {
      "id": "744557",
      "postDate": "02/12/2020 23:41:51",
      "content": "<p>Thanks for explanation, this is helpful. I've been playing around with <code>MirroredStrategy</code> too. TensorFlow makes it easy to train with different device configurations.</p>",
      "rawMarkdown": "Thanks for explanation, this is helpful. I've been playing around with `MirroredStrategy` too. TensorFlow makes it easy to train with different device configurations.",
      "votes": null
    },
    {
      "id": "745813",
      "postDate": "02/14/2020 08:37:11",
      "content": "<p>Thanks so much <a href=\"/adityaecdrid\">@adityaecdrid</a> for this article. I just read it and Thomas did a fantastic job of describing what is entailed i.e. HW considerations in how to best model your NN without OOM issues. The DataParallelModel reminded me of my signal processing days (pre GPU's) when we used specialized digital signal processing boards to do the same thing. Only in our case it is more like a cross between the DataParallelModel and the distributed training. Oh good old GNU brings back memories :-)</p>\n\n<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for asking this question.</p>",
      "rawMarkdown": "Thanks so much @adityaecdrid for this article. I just read it and Thomas did a fantastic job of describing what is entailed i.e. HW considerations in how to best model your NN without OOM issues. The DataParallelModel reminded me of my signal processing days (pre GPU's) when we used specialized digital signal processing boards to do the same thing. Only in our case it is more like a cross between the DataParallelModel and the distributed training. Oh good old GNU brings back memories :-)\n\nThanks @cdeotte for asking this question.",
      "votes": null
    },
    {
      "id": "753246",
      "postDate": "02/21/2020 23:20:05",
      "content": "<p>Hmmm! Who is the downvoter?</p>",
      "rawMarkdown": "Hmmm! Who is the downvoter?",
      "votes": null
    },
    {
      "id": "760270",
      "postDate": "03/01/2020 03:51:34",
      "content": "<p>Ｈ𝐀𝑷𝑷𝓎 🇰𝗮𝘨𝘨🇱𝖎Ｎɢ  💯</p>",
      "rawMarkdown": "Ｈ𝐀𝑷𝑷𝓎 🇰𝗮𝘨𝘨🇱𝖎Ｎɢ  💯",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 744272,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "02/12/2020 17:34:35",
      "content": "<p>Not sure if this is gonna help but this a great <a href=\"https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255\">article</a> by Hugging-Face!</p>",
      "votes": null,
      "replies": [
        {
          "id": 744295,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/12/2020 17:47:42",
          "content": "<p>Thanks I'll read that. How models use multiple TPUs or GPUs will affect how we think about hyperparameter choices.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745813,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/14/2020 08:37:11",
          "content": "<p>Thanks so much <a href=\"/adityaecdrid\">@adityaecdrid</a> for this article. I just read it and Thomas did a fantastic job of describing what is entailed i.e. HW considerations in how to best model your NN without OOM issues. The DataParallelModel reminded me of my signal processing days (pre GPU's) when we used specialized digital signal processing boards to do the same thing. Only in our case it is more like a cross between the DataParallelModel and the distributed training. Oh good old GNU brings back memories :-)</p>\n\n<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for asking this question.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 753246,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/21/2020 23:20:05",
          "content": "<p>Hmmm! Who is the downvoter?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 744352,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "02/12/2020 18:59:13",
      "content": "<p>Yes I think that's what is happening. If you have 8 replicas (e.g. using TPU v3-8) then the total batch_size = 16 * 8.  Each TPU core gets 16 samples. You can choose how you <a href=\"https://github.com/see--/natural-question-answering/blob/master/train_eval.py#L126\">reduce</a> the loss along multiple cores. This will determine the gradients: <code>tf.distribute.ReduceOp.MEAN</code> or <code>tf.distribute.ReduceOp.SUM</code>. To add <a href=\"https://www.tensorflow.org/guide/distributed_training#tpustrategy\">TPUStrategy is same as MirroredStrategy</a>.</p>\n\n<p>So far, using <code>MEAN</code> and linearly scaling the learning rate as if I had one big batch_size of 16 * 8 works well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 744557,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/12/2020 23:41:51",
          "content": "<p>Thanks for explanation, this is helpful. I've been playing around with <code>MirroredStrategy</code> too. TensorFlow makes it easy to train with different device configurations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 744375,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/12/2020 19:21:04",
      "content": "<p>Glad you asked! This is all implemented in Tensorflow distribution strategies. For TPU strategy specifically:</p>\n\n<ul>\n<li>data batches are split between the 8 TPU cores</li>\n<li>gradients are computed on each batch in parallel and merged</li>\n<li>all model weights are updated</li>\n</ul>\n\n<p>There is actually more work performed if you are running on a TPU pod (TPU = 1 board/4 chips/8cores, TPU pod = 2 to 256 TPU boards on a high-performance interconnect exposed to the user as a single accelerator). On a TPU pod:</p>\n\n<ul>\n<li>the model is replicated across all TPU boards</li>\n<li>the datset (tf.data.Dataset) is split at the file level between all TPU boards. Here it becomes important that your dataset is sharded across multiple files. TPU boards that do not need certain shards will never load them. If you have fewer files (shards) than TPU boards, you will actually get an error message on a TPU pod.</li>\n<li>batches of data are loaded on each TPU board and split between TPU cores.</li>\n<li>gradients are computed and merged across cores and boards and all model weights are updated (all reduce algorithm - this is where the fast interconnect between TPU boards comes into play)</li>\n</ul>\n\n<p>All this is happening automatically in tf.data.Dataset and the TPUStrategy code. The user just supplies a tf.data.Dataset with a larger batch size. The more TPU cores, the larger the batch size.</p>",
      "votes": null,
      "replies": [
        {
          "id": 744379,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/12/2020 19:24:27",
          "content": "<p>And yes, if you use larger batches, you will need a bigger learning rate too. Start scaling the batch size linearly with core count but then expect some hyperparam tuning work to find the best LR schedule for a given batch size and TPU configuration.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 744556,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/12/2020 23:40:51",
          "content": "<p>Thanks for clarification</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 760270,
      "author_name": "jazivxt",
      "author_url": "",
      "post_date": "03/01/2020 03:51:34",
      "content": "<p>Ｈ𝐀𝑷𝑷𝓎 🇰𝗮𝘨𝘨🇱𝖎Ｎɢ  💯</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "744250": "I have a technical question. Does anyone know how TensorFlow uses multiple TPU and/or multiple GPU during back propagation training of neural networks?\n\nFor example, with TPU when we say `batch_size = 16 * 8` which equals 128, does each TPU get a batch of 16 and the model weights `W0` and activations `A0` at time = `T0`.  Then does each core compute the gradient with `W0` and `A0` and then every core updates the original weights `W0` by adding their gradient updates?\n\nNote that this is different than training on 1 core where if we perform 8 batches, then first batch uses `W0, A0` and second batch uses `W1, A1` (the updated weights and activations after first time step), and third batch uses `W2, A2` etc.",
    "744272": "Not sure if this is gonna help but this a great [article](https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255) by Hugging-Face!",
    "744295": "Thanks I'll read that. How models use multiple TPUs or GPUs will affect how we think about hyperparameter choices.",
    "744352": "Yes I think that's what is happening. If you have 8 replicas (e.g. using TPU v3-8) then the total batch_size = 16 * 8.  Each TPU core gets 16 samples. You can choose how you [reduce](https://github.com/see--/natural-question-answering/blob/master/train_eval.py#L126) the loss along multiple cores. This will determine the gradients: `tf.distribute.ReduceOp.MEAN` or `tf.distribute.ReduceOp.SUM`. To add [TPUStrategy is same as MirroredStrategy](https://www.tensorflow.org/guide/distributed_training#tpustrategy).\n\nSo far, using `MEAN` and linearly scaling the learning rate as if I had one big batch_size of 16 * 8 works well.",
    "744375": "Glad you asked! This is all implemented in Tensorflow distribution strategies. For TPU strategy specifically:\n\n- data batches are split between the 8 TPU cores\n- gradients are computed on each batch in parallel and merged\n- all model weights are updated\n\nThere is actually more work performed if you are running on a TPU pod (TPU = 1 board/4 chips/8cores, TPU pod = 2 to 256 TPU boards on a high-performance interconnect exposed to the user as a single accelerator). On a TPU pod:\n\n- the model is replicated across all TPU boards\n- the datset (tf.data.Dataset) is split at the file level between all TPU boards. Here it becomes important that your dataset is sharded across multiple files. TPU boards that do not need certain shards will never load them. If you have fewer files (shards) than TPU boards, you will actually get an error message on a TPU pod.\n- batches of data are loaded on each TPU board and split between TPU cores.\n- gradients are computed and merged across cores and boards and all model weights are updated (all reduce algorithm - this is where the fast interconnect between TPU boards comes into play)\n\nAll this is happening automatically in tf.data.Dataset and the TPUStrategy code. The user just supplies a tf.data.Dataset with a larger batch size. The more TPU cores, the larger the batch size.",
    "744379": "And yes, if you use larger batches, you will need a bigger learning rate too. Start scaling the batch size linearly with core count but then expect some hyperparam tuning work to find the best LR schedule for a given batch size and TPU configuration.",
    "744556": "Thanks for clarification",
    "744557": "Thanks for explanation, this is helpful. I've been playing around with `MirroredStrategy` too. TensorFlow makes it easy to train with different device configurations.",
    "745813": "Thanks so much @adityaecdrid for this article. I just read it and Thomas did a fantastic job of describing what is entailed i.e. HW considerations in how to best model your NN without OOM issues. The DataParallelModel reminded me of my signal processing days (pre GPU's) when we used specialized digital signal processing boards to do the same thing. Only in our case it is more like a cross between the DataParallelModel and the distributed training. Oh good old GNU brings back memories :-)\n\nThanks @cdeotte for asking this question.",
    "753246": "Hmmm! Who is the downvoter?",
    "760270": "Ｈ𝐀𝑷𝑷𝓎 🇰𝗮𝘨𝘨🇱𝖎Ｎɢ  💯"
  },
  "source": "meta"
}