{
  "id": 137798,
  "title": "TPU - Gradient accumulation",
  "url": "/competitions/flower-classification-with-tpus/discussion/137798",
  "author_name": "",
  "post_date": "2020-03-22T14:11:21.000199100Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi, I published a kernel <a href=\"https://www.kaggle.com/yihdarshieh/tpu-gradient-accumulation?scriptVersionId=30622480\">TPU - Gradient accumulation</a>, which demonstrates how to do gradient accumulation with TPU. Even with EfficientNet, it's possible to have effective batch size 2048, which is 16 times larger than the batch size in common kernels in this competition.</p>",
  "messages": [
    {
      "id": "782637",
      "postDate": "03/22/2020 14:11:21",
      "content": "<p>Hi, I published a kernel <a href=\"https://www.kaggle.com/yihdarshieh/tpu-gradient-accumulation?scriptVersionId=30622480\">TPU - Gradient accumulation</a>, which demonstrates how to do gradient accumulation with TPU. Even with EfficientNet, it's possible to have effective batch size 2048, which is 16 times larger than the batch size in common kernels in this competition.</p>",
      "rawMarkdown": "Hi, I published a kernel [TPU - Gradient accumulation](https://www.kaggle.com/yihdarshieh/tpu-gradient-accumulation?scriptVersionId=30622480), which demonstrates how to do gradient accumulation with TPU. Even with EfficientNet, it's possible to have effective batch size 2048, which is 16 times larger than the batch size in common kernels in this competition.",
      "votes": null
    },
    {
      "id": "784040",
      "postDate": "03/23/2020 22:40:34",
      "content": "<p>Is 2048 the per-replica batch size or the global batch size ?</p>\n\n<p>The <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">Getting started notebook</a> uses a global batch size of 16*8=128.</p>",
      "rawMarkdown": "Is 2048 the per-replica batch size or the global batch size ?\n\nThe [Getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/) uses a global batch size of 16*8=128.",
      "votes": null
    },
    {
      "id": "784050",
      "postDate": "03/23/2020 22:53:47",
      "content": "<p>I read through the notebook. Great work btw. A couple of comments\n- I am not sure the <code>with strategy.scope():</code> inside of <code>def set_batch_configuration(...):</code> does anything.\n- You say that <code>small_images = images[start_idx:end_idx]</code> does not work and that's correct. Did you try to follow with a tf.reshape(small_images, [BATCH_SIZE_PER_REPLICA, ...]) ? The batch size is constant but Tensorflow needs to prove it in order to continue. Usually a reshape helps in these cases.</p>",
      "rawMarkdown": "I read through the notebook. Great work btw. A couple of comments\n- I am not sure the `with strategy.scope():` inside of `def set_batch_configuration(...):` does anything.\n- You say that `small_images = images[start_idx:end_idx]` does not work and that's correct. Did you try to follow with a tf.reshape(small_images, [BATCH_SIZE_PER_REPLICA, ...]) ? The batch size is constant but Tensorflow needs to prove it in order to continue. Usually a reshape helps in these cases.",
      "votes": null
    },
    {
      "id": "784235",
      "postDate": "03/24/2020 03:02:04",
      "content": "<p>2048 is the global batch size. Each replica receives 256 examples, and process them in batch size 8 to accumulate the gradients.</p>",
      "rawMarkdown": "2048 is the global batch size. Each replica receives 256 examples, and process them in batch size 8 to accumulate the gradients.",
      "votes": null
    },
    {
      "id": "784281",
      "postDate": "03/24/2020 04:35:57",
      "content": "<p>Thanks for clarifying. Another question:\nIn your experience, when do very large batches help ?</p>",
      "rawMarkdown": "Thanks for clarifying. Another question:\nIn your experience, when do very large batches help ?",
      "votes": null
    },
    {
      "id": "784687",
      "postDate": "03/24/2020 12:40:19",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I can't answer this question from my experience.</p>\n\n<p>I see people reporting the benefits from large batch size, but it requires good tuning of learning rate. \nBut with <code>very</code> large batch size, the final evaluation results tend to be a bit worse than <code>large</code> batch size. (It's a trade-off between training time and performance though.) The following picture is taken from <a href=\"https://arxiv.org/pdf/1904.00962.pdf\">Large Batch Optimization for Deep Learning: Training BERT in 76 minutes</a>.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F62fb7e32df5e778bc4f2d2d7fd0c9de4%2FCapture.PNG?generation=1585052912976606&amp;alt=media\" alt=\"\"></p>\n\n<p>Also some optimizers can't benefit from very large batch size. </p>\n\n<p>However, from a theoretical viewpoint, with an imbalanced training dataset, or a training dataset with many target labels, it's better to include all labels in each batch. But this probably doesn't require <code>very large</code> batch size (just being <code>large</code> might be enough).</p>\n\n<p>As you know, I published several kernels. And my last task on the list is to implement the idea <a href=\"https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d\">The Unknown Benefits of using a Soft-F1 Loss in Classification Systems</a>. Optimizing directly the F1 score benefits definitely (of course, to be confirmed) from <code>very large</code> batch (ideally, the whole epoch as 1 batch).</p>",
      "rawMarkdown": "mgornergoogle I can't answer this question from my experience.\n\nI see people reporting the benefits from large batch size, but it requires good tuning of learning rate. \nBut with `very` large batch size, the final evaluation results tend to be a bit worse than `large` batch size. (It's a trade-off between training time and performance though.) The following picture is taken from [Large Batch Optimization for Deep Learning: Training BERT in 76 minutes](https://arxiv.org/pdf/1904.00962.pdf).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F62fb7e32df5e778bc4f2d2d7fd0c9de4%2FCapture.PNG?generation=1585052912976606&amp;alt=media)\n\nAlso some optimizers can't benefit from very large batch size. \n\nHowever, from a theoretical viewpoint, with an imbalanced training dataset, or a training dataset with many target labels, it's better to include all labels in each batch. But this probably doesn't require `very large` batch size (just being `large` might be enough).\n\nAs you know, I published several kernels. And my last task on the list is to implement the idea [The Unknown Benefits of using a Soft-F1 Loss in Classification Systems](https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d). Optimizing directly the F1 score benefits definitely (of course, to be confirmed) from `very large` batch (ideally, the whole epoch as 1 batch).",
      "votes": null
    },
    {
      "id": "784890",
      "postDate": "03/24/2020 15:27:38",
      "content": "<p>Oh yes those are really large batch sizes. Also, this is using LARS/LAMB optimizers which are quite specific for very large batch sizes.</p>\n\n<p>My own experience is from working on TPUs and TPU pods, where large batch sizes are a requirement to get good performance. I found large batch sizes to be a bit of a struggle because the LR schedule must be retuned. Usually, if I can get to my original accuracy after re-tuning, I'm quite happy.</p>\n\n<p>The LAMB chart above reflects what I usually see: final score roughly flat across batch sizes (with good tuning) but getting more and more difficult to keep up as the batch size grows.</p>",
      "rawMarkdown": "Oh yes those are really large batch sizes. Also, this is using LARS/LAMB optimizers which are quite specific for very large batch sizes.\n\nMy own experience is from working on TPUs and TPU pods, where large batch sizes are a requirement to get good performance. I found large batch sizes to be a bit of a struggle because the LR schedule must be retuned. Usually, if I can get to my original accuracy after re-tuning, I'm quite happy.\n\nThe LAMB chart above reflects what I usually see: final score roughly flat across batch sizes (with good tuning) but getting more and more difficult to keep up as the batch size grows.",
      "votes": null
    },
    {
      "id": "785182",
      "postDate": "03/24/2020 21:22:34",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> ,</p>\n\n<p>Thank you for the comments.</p>\n\n<p>I tried your suggestion about <code>tf.reshape</code>, and it doesn't work in this case. Still get the following error</p>\n\n<blockquote>\n  <p>Compilation failure: XLA can't deduce compile time constant output shape for strided slice: [?], output shape must be a compile-time constant\n       [[{{node strided_slice_1}}]]\n       [[while/body/_1/while]]\n      TPU compilation failed\n       [[tpu_compile_succeeded_assert/_9911381903162691729/_3]]</p>\n</blockquote>\n\n<p>About <code>set_batch_configuration</code>, you are right. I tried to finish the kernel quickly at the end and didn't think about it ...</p>",
      "rawMarkdown": "mgornergoogle ,\n\nThank you for the comments.\n\nI tried your suggestion about `tf.reshape`, and it doesn't work in this case. Still get the following error\n\n&gt;    Compilation failure: XLA can't deduce compile time constant output shape for strided slice: [?], output shape must be a compile-time constant\n\t [[{{node strided_slice_1}}]]\n\t [[while/body/_1/while]]\n\tTPU compilation failed\n\t [[tpu_compile_succeeded_assert/_9911381903162691729/_3]]\n\nAbout `set_batch_configuration`, you are right. I tried to finish the kernel quickly at the end and didn't think about it ...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 784040,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/23/2020 22:40:34",
      "content": "<p>Is 2048 the per-replica batch size or the global batch size ?</p>\n\n<p>The <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">Getting started notebook</a> uses a global batch size of 16*8=128.</p>",
      "votes": null,
      "replies": [
        {
          "id": 784235,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/24/2020 03:02:04",
          "content": "<p>2048 is the global batch size. Each replica receives 256 examples, and process them in batch size 8 to accumulate the gradients.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784281,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/24/2020 04:35:57",
          "content": "<p>Thanks for clarifying. Another question:\nIn your experience, when do very large batches help ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784687,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/24/2020 12:40:19",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I can't answer this question from my experience.</p>\n\n<p>I see people reporting the benefits from large batch size, but it requires good tuning of learning rate. \nBut with <code>very</code> large batch size, the final evaluation results tend to be a bit worse than <code>large</code> batch size. (It's a trade-off between training time and performance though.) The following picture is taken from <a href=\"https://arxiv.org/pdf/1904.00962.pdf\">Large Batch Optimization for Deep Learning: Training BERT in 76 minutes</a>.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F62fb7e32df5e778bc4f2d2d7fd0c9de4%2FCapture.PNG?generation=1585052912976606&amp;alt=media\" alt=\"\"></p>\n\n<p>Also some optimizers can't benefit from very large batch size. </p>\n\n<p>However, from a theoretical viewpoint, with an imbalanced training dataset, or a training dataset with many target labels, it's better to include all labels in each batch. But this probably doesn't require <code>very large</code> batch size (just being <code>large</code> might be enough).</p>\n\n<p>As you know, I published several kernels. And my last task on the list is to implement the idea <a href=\"https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d\">The Unknown Benefits of using a Soft-F1 Loss in Classification Systems</a>. Optimizing directly the F1 score benefits definitely (of course, to be confirmed) from <code>very large</code> batch (ideally, the whole epoch as 1 batch).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784890,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/24/2020 15:27:38",
          "content": "<p>Oh yes those are really large batch sizes. Also, this is using LARS/LAMB optimizers which are quite specific for very large batch sizes.</p>\n\n<p>My own experience is from working on TPUs and TPU pods, where large batch sizes are a requirement to get good performance. I found large batch sizes to be a bit of a struggle because the LR schedule must be retuned. Usually, if I can get to my original accuracy after re-tuning, I'm quite happy.</p>\n\n<p>The LAMB chart above reflects what I usually see: final score roughly flat across batch sizes (with good tuning) but getting more and more difficult to keep up as the batch size grows.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 784050,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/23/2020 22:53:47",
      "content": "<p>I read through the notebook. Great work btw. A couple of comments\n- I am not sure the <code>with strategy.scope():</code> inside of <code>def set_batch_configuration(...):</code> does anything.\n- You say that <code>small_images = images[start_idx:end_idx]</code> does not work and that's correct. Did you try to follow with a tf.reshape(small_images, [BATCH_SIZE_PER_REPLICA, ...]) ? The batch size is constant but Tensorflow needs to prove it in order to continue. Usually a reshape helps in these cases.</p>",
      "votes": null,
      "replies": [
        {
          "id": 785182,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/24/2020 21:22:34",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> ,</p>\n\n<p>Thank you for the comments.</p>\n\n<p>I tried your suggestion about <code>tf.reshape</code>, and it doesn't work in this case. Still get the following error</p>\n\n<blockquote>\n  <p>Compilation failure: XLA can't deduce compile time constant output shape for strided slice: [?], output shape must be a compile-time constant\n       [[{{node strided_slice_1}}]]\n       [[while/body/_1/while]]\n      TPU compilation failed\n       [[tpu_compile_succeeded_assert/_9911381903162691729/_3]]</p>\n</blockquote>\n\n<p>About <code>set_batch_configuration</code>, you are right. I tried to finish the kernel quickly at the end and didn't think about it ...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "782637": "Hi, I published a kernel [TPU - Gradient accumulation](https://www.kaggle.com/yihdarshieh/tpu-gradient-accumulation?scriptVersionId=30622480), which demonstrates how to do gradient accumulation with TPU. Even with EfficientNet, it's possible to have effective batch size 2048, which is 16 times larger than the batch size in common kernels in this competition.",
    "784040": "Is 2048 the per-replica batch size or the global batch size ?\n\nThe [Getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/) uses a global batch size of 16*8=128.",
    "784050": "I read through the notebook. Great work btw. A couple of comments\n- I am not sure the `with strategy.scope():` inside of `def set_batch_configuration(...):` does anything.\n- You say that `small_images = images[start_idx:end_idx]` does not work and that's correct. Did you try to follow with a tf.reshape(small_images, [BATCH_SIZE_PER_REPLICA, ...]) ? The batch size is constant but Tensorflow needs to prove it in order to continue. Usually a reshape helps in these cases.",
    "784235": "2048 is the global batch size. Each replica receives 256 examples, and process them in batch size 8 to accumulate the gradients.",
    "784281": "Thanks for clarifying. Another question:\nIn your experience, when do very large batches help ?",
    "784687": "mgornergoogle I can't answer this question from my experience.\n\nI see people reporting the benefits from large batch size, but it requires good tuning of learning rate. \nBut with `very` large batch size, the final evaluation results tend to be a bit worse than `large` batch size. (It's a trade-off between training time and performance though.) The following picture is taken from [Large Batch Optimization for Deep Learning: Training BERT in 76 minutes](https://arxiv.org/pdf/1904.00962.pdf).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F62fb7e32df5e778bc4f2d2d7fd0c9de4%2FCapture.PNG?generation=1585052912976606&amp;alt=media)\n\nAlso some optimizers can't benefit from very large batch size. \n\nHowever, from a theoretical viewpoint, with an imbalanced training dataset, or a training dataset with many target labels, it's better to include all labels in each batch. But this probably doesn't require `very large` batch size (just being `large` might be enough).\n\nAs you know, I published several kernels. And my last task on the list is to implement the idea [The Unknown Benefits of using a Soft-F1 Loss in Classification Systems](https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d). Optimizing directly the F1 score benefits definitely (of course, to be confirmed) from `very large` batch (ideally, the whole epoch as 1 batch).",
    "784890": "Oh yes those are really large batch sizes. Also, this is using LARS/LAMB optimizers which are quite specific for very large batch sizes.\n\nMy own experience is from working on TPUs and TPU pods, where large batch sizes are a requirement to get good performance. I found large batch sizes to be a bit of a struggle because the LR schedule must be retuned. Usually, if I can get to my original accuracy after re-tuning, I'm quite happy.\n\nThe LAMB chart above reflects what I usually see: final score roughly flat across batch sizes (with good tuning) but getting more and more difficult to keep up as the batch size grows.",
    "785182": "mgornergoogle ,\n\nThank you for the comments.\n\nI tried your suggestion about `tf.reshape`, and it doesn't work in this case. Still get the following error\n\n&gt;    Compilation failure: XLA can't deduce compile time constant output shape for strided slice: [?], output shape must be a compile-time constant\n\t [[{{node strided_slice_1}}]]\n\t [[while/body/_1/while]]\n\tTPU compilation failed\n\t [[tpu_compile_succeeded_assert/_9911381903162691729/_3]]\n\nAbout `set_batch_configuration`, you are right. I tried to finish the kernel quickly at the end and didn't think about it ..."
  },
  "source": "meta"
}