{
  "id": 132191,
  "title": "How To - Rotation Augmentation GPU/TPU",
  "url": "/competitions/flower-classification-with-tpus/discussion/132191",
  "author_name": "",
  "post_date": "2020-02-24T19:33:17.575058100Z",
  "votes": 73,
  "comment_count": 39,
  "views": 0,
  "content": "<h1>Data Augmentation - Rotation, Shear, Zoom, Shift</h1>\n\n<p>I posted a starter kernel <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">here</a> showing how to apply rotation, shear, zoom, and shift augmentation when using <code>TensorFlow.data.Dataset()</code>. It performs augmentation on the GPU/TPU instead of CPU for maximum speed. Feel free to use this code in your own notebooks.</p>\n\n<h1>Speed Analysis</h1>\n\n<p>A GPU or TPU can process 200+ images (512x512x3) per second while training DenseNet201. That's incredibly fast! If we try to augment images in the CPU, then we may not be able to provide the GPU/TPU with images fast enough and thus we will slow down our training.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ffbb804fdbfe3a6bfb54acbf35039e36a%2Fcpu.jpg?generation=1582409535436227&amp;alt=media\" alt=\"\"></p>\n\n<h1>Preprocess on GPU/TPU</h1>\n\n<p>This is a great competition to appreciate the need to preprocess data on the GPU/TPU. If you wish to rotate a 512x512x3 image, that requires 5,000,000 multiplies and additions per image! </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fac5072f04b8e627020c439ab8326b038%2Frotate.JPG?generation=1582409737942526&amp;alt=media\" alt=\"\"></p>\n\n<p>For each pixel in your augmented image, you must find which pixel value in the original image to use. In the example above, we wish to determine which pixel to place in location <code>(1,3)</code>. So we must multiply the coordinate <code>(1,3)</code> by a rotation matrix to determine that we want pixel <code>(3,2)</code> from the original image which is color pink. We then place a pink pixel in destination image's pixel <code>(1,3)</code>.</p>\n\n<h1>Code Does Not Use For-Loops</h1>\n\n<p>When writing code for GPU/TPU (i.e. in TensorFlow, Numba, CuPy, CUDA, etc), we want to avoid using for-loops. With a 512x512 image, we must determine the pixel values for 262,144 = 512 x 512 pixels. We can either write a for-loop as <code>for i in range(262144):</code> or we can create a matrix <code>P</code> of size <code>(2,262144)</code> and multiply it by a <code>(2,2)</code> rotation matrix <code>R</code>. In the later method, this allows GPU/TPU to compute the 262,144 new coordinates in parallel since each GPU/TPU thread can multiply one column of matrix <code>P</code> by matrix <code>R</code> simultaneously. In the former method, the for-loop prevents us from parallelization (since each iteration must wait for previous iterations to complete).</p>\n\n<p>Once we have the coordinates of the 262,144 pixels that we want, we use <code>tensorflow.gather_nd()</code> which divides the job of finding the pixels in the original image among thousands of different GPU/TPU threads. Once again, we did not write <code>for i in range(262144):</code> to locate the 262,144 original pixels which would have prevented parallelism.</p>\n\n<h1>Example Rotated Image</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5042f32b75cd19ad71d5a35c5587ae79%2FScreen%20Shot%202020-02-24%20at%2011.32.02%20AM.png?generation=1582572736200811&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "755450",
      "postDate": "02/24/2020 19:33:17",
      "content": "<h1>Data Augmentation - Rotation, Shear, Zoom, Shift</h1>\n\n<p>I posted a starter kernel <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">here</a> showing how to apply rotation, shear, zoom, and shift augmentation when using <code>TensorFlow.data.Dataset()</code>. It performs augmentation on the GPU/TPU instead of CPU for maximum speed. Feel free to use this code in your own notebooks.</p>\n\n<h1>Speed Analysis</h1>\n\n<p>A GPU or TPU can process 200+ images (512x512x3) per second while training DenseNet201. That's incredibly fast! If we try to augment images in the CPU, then we may not be able to provide the GPU/TPU with images fast enough and thus we will slow down our training.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ffbb804fdbfe3a6bfb54acbf35039e36a%2Fcpu.jpg?generation=1582409535436227&amp;alt=media\" alt=\"\"></p>\n\n<h1>Preprocess on GPU/TPU</h1>\n\n<p>This is a great competition to appreciate the need to preprocess data on the GPU/TPU. If you wish to rotate a 512x512x3 image, that requires 5,000,000 multiplies and additions per image! </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fac5072f04b8e627020c439ab8326b038%2Frotate.JPG?generation=1582409737942526&amp;alt=media\" alt=\"\"></p>\n\n<p>For each pixel in your augmented image, you must find which pixel value in the original image to use. In the example above, we wish to determine which pixel to place in location <code>(1,3)</code>. So we must multiply the coordinate <code>(1,3)</code> by a rotation matrix to determine that we want pixel <code>(3,2)</code> from the original image which is color pink. We then place a pink pixel in destination image's pixel <code>(1,3)</code>.</p>\n\n<h1>Code Does Not Use For-Loops</h1>\n\n<p>When writing code for GPU/TPU (i.e. in TensorFlow, Numba, CuPy, CUDA, etc), we want to avoid using for-loops. With a 512x512 image, we must determine the pixel values for 262,144 = 512 x 512 pixels. We can either write a for-loop as <code>for i in range(262144):</code> or we can create a matrix <code>P</code> of size <code>(2,262144)</code> and multiply it by a <code>(2,2)</code> rotation matrix <code>R</code>. In the later method, this allows GPU/TPU to compute the 262,144 new coordinates in parallel since each GPU/TPU thread can multiply one column of matrix <code>P</code> by matrix <code>R</code> simultaneously. In the former method, the for-loop prevents us from parallelization (since each iteration must wait for previous iterations to complete).</p>\n\n<p>Once we have the coordinates of the 262,144 pixels that we want, we use <code>tensorflow.gather_nd()</code> which divides the job of finding the pixels in the original image among thousands of different GPU/TPU threads. Once again, we did not write <code>for i in range(262144):</code> to locate the 262,144 original pixels which would have prevented parallelism.</p>\n\n<h1>Example Rotated Image</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5042f32b75cd19ad71d5a35c5587ae79%2FScreen%20Shot%202020-02-24%20at%2011.32.02%20AM.png?generation=1582572736200811&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "# Data Augmentation - Rotation, Shear, Zoom, Shift\nI posted a starter kernel [here][1] showing how to apply rotation, shear, zoom, and shift augmentation when using `TensorFlow.data.Dataset()`. It performs augmentation on the GPU/TPU instead of CPU for maximum speed. Feel free to use this code in your own notebooks.\n# Speed Analysis\nA GPU or TPU can process 200+ images (512x512x3) per second while training DenseNet201. That's incredibly fast! If we try to augment images in the CPU, then we may not be able to provide the GPU/TPU with images fast enough and thus we will slow down our training.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ffbb804fdbfe3a6bfb54acbf35039e36a%2Fcpu.jpg?generation=1582409535436227&amp;alt=media)\n  \n# Preprocess on GPU/TPU\nThis is a great competition to appreciate the need to preprocess data on the GPU/TPU. If you wish to rotate a 512x512x3 image, that requires 5,000,000 multiplies and additions per image! \n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fac5072f04b8e627020c439ab8326b038%2Frotate.JPG?generation=1582409737942526&amp;alt=media)\n  \nFor each pixel in your augmented image, you must find which pixel value in the original image to use. In the example above, we wish to determine which pixel to place in location `(1,3)`. So we must multiply the coordinate `(1,3)` by a rotation matrix to determine that we want pixel `(3,2)` from the original image which is color pink. We then place a pink pixel in destination image's pixel `(1,3)`.\n\n# Code Does Not Use For-Loops\nWhen writing code for GPU/TPU (i.e. in TensorFlow, Numba, CuPy, CUDA, etc), we want to avoid using for-loops. With a 512x512 image, we must determine the pixel values for 262,144 = 512 x 512 pixels. We can either write a for-loop as `for i in range(262144):` or we can create a matrix `P` of size `(2,262144)` and multiply it by a `(2,2)` rotation matrix `R`. In the later method, this allows GPU/TPU to compute the 262,144 new coordinates in parallel since each GPU/TPU thread can multiply one column of matrix `P` by matrix `R` simultaneously. In the former method, the for-loop prevents us from parallelization (since each iteration must wait for previous iterations to complete).\n\nOnce we have the coordinates of the 262,144 pixels that we want, we use `tensorflow.gather_nd()` which divides the job of finding the pixels in the original image among thousands of different GPU/TPU threads. Once again, we did not write `for i in range(262144):` to locate the 262,144 original pixels which would have prevented parallelism.\n\n# Example Rotated Image\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5042f32b75cd19ad71d5a35c5587ae79%2FScreen%20Shot%202020-02-24%20at%2011.32.02%20AM.png?generation=1582572736200811&amp;alt=media)\n\n\n[1]: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96",
      "votes": null
    },
    {
      "id": "755524",
      "postDate": "02/24/2020 21:19:56",
      "content": "<p>Thanks, great to have rotation! Yesterday, I just tried to find if <code>tf.image</code> has arbitrary rotation method, and it doesn't. And tensorflow addons <code>tfa.image.rotate</code> is not working ...</p>",
      "rawMarkdown": "Thanks, great to have rotation! Yesterday, I just tried to find if `tf.image` has arbitrary rotation method, and it doesn't. And tensorflow addons `tfa.image.rotate` is not working ...",
      "votes": null
    },
    {
      "id": "755545",
      "postDate": "02/24/2020 21:40:59",
      "content": "<p>Yes augmentation is difficult in this competition because no libraries work. Now you can do rotation, shear, zoom, and shift. Enjoy!</p>",
      "rawMarkdown": "Yes augmentation is difficult in this competition because no libraries work. Now you can do rotation, shear, zoom, and shift. Enjoy!",
      "votes": null
    },
    {
      "id": "755560",
      "postDate": "02/24/2020 22:21:16",
      "content": "<p>You just shared a huge building block for many other augmentation methods that will need to be ported to tensorflow2 and TPU. Thank you!</p>\n\n<p>I am nominating you for the \"TPU Star\" Prize ;)</p>",
      "rawMarkdown": "You just shared a huge building block for many other augmentation methods that will need to be ported to tensorflow2 and TPU. Thank you!\n\nI am nominating you for the \"TPU Star\" Prize ;)",
      "votes": null
    },
    {
      "id": "755591",
      "postDate": "02/24/2020 23:15:02",
      "content": "<p>Thanks Ilu</p>",
      "rawMarkdown": "Thanks Ilu",
      "votes": null
    },
    {
      "id": "755617",
      "postDate": "02/25/2020 00:25:37",
      "content": "<p>Fantastic notebook. Thanks.\nCould you explain what the <code>tf.keras.mixed_precision.experimental.Policy</code>does and why it is needed ?</p>",
      "rawMarkdown": "Fantastic notebook. Thanks.\nCould you explain what the `tf.keras.mixed_precision.experimental.Policy `does and why it is needed ?",
      "votes": null
    },
    {
      "id": "755633",
      "postDate": "02/25/2020 01:01:37",
      "content": "<p>It is not needed. Currently the boolean switches are set to <code>False</code>. If you turn it on then Nvidia GPU V100 goes 40% faster and fits twice as large batch sizes. And TPUv3 goes 40% faster and fits twice as large batch sizes. (GPU P100 doesn't speed up but it does fit larger batch sizes).</p>\n\n<p>When <code>MIXED_PRECISION = True</code>, TensorFlow uses <code>fp16</code> instead of <code>fp32</code> wherever it can safely.</p>",
      "rawMarkdown": "It is not needed. Currently the boolean switches are set to `False`. If you turn it on then Nvidia GPU V100 goes 40% faster and fits twice as large batch sizes. And TPUv3 goes 40% faster and fits twice as large batch sizes. (GPU P100 doesn't speed up but it does fit larger batch sizes).\n\nWhen `MIXED_PRECISION = True`, TensorFlow uses `fp16` instead of `fp32` wherever it can safely.",
      "votes": null
    },
    {
      "id": "756235",
      "postDate": "02/25/2020 14:40:51",
      "content": "<p>Such a good notebook, thanks for share! 💯 </p>",
      "rawMarkdown": "Such a good notebook, thanks for share! 💯",
      "votes": null
    },
    {
      "id": "756625",
      "postDate": "02/25/2020 22:43:47",
      "content": "<p>About NVIDIA V100: does this setting enable the TensorCores on the V100 ?</p>\n\n<p>About the TPU v3: as you said, this setting is not needed for the TPU v3 to use mixed precision. But it does optimize memory usage because tensors are then stored in memory in bfloat16 format. I assume the 40% speed increase you mention is attributable solely to the memory optimization which allows bigger batches. Correct ?</p>",
      "rawMarkdown": "About NVIDIA V100: does this setting enable the TensorCores on the V100 ?\n\nAbout the TPU v3: as you said, this setting is not needed for the TPU v3 to use mixed precision. But it does optimize memory usage because tensors are then stored in memory in bfloat16 format. I assume the 40% speed increase you mention is attributable solely to the memory optimization which allows bigger batches. Correct ?",
      "votes": null
    },
    {
      "id": "756653",
      "postDate": "02/25/2020 23:19:56",
      "content": "<p>About NVIDIA V100: yes, <code>mixed_precision</code> on Nvidia V100 enables TensorCores.</p>\n\n<p>About the TPU v3: On the TPU, with batch size 128, one epoch (train and validate DenseNet201 with 512x512x3) is 70 seconds without <code>mixed_precision</code> enabled and 60 seconds with <code>mixed_precision</code> enabled. That's 20% increase. Then with <code>mixed_precision</code>, you can also double the batch size to 256 and do 50 second epochs gaining another 20%.</p>\n\n<p>4xGPU V100 (comparable hardware as TPUv3's 4 chips) with mixed precision at batch 192 (can't fit 256 with <code>tf.distribute.MirroredStrategy()</code>) does 60 second epochs. This is using TensorFlow. With Pytorch and Nvidia apex for mixed precision and distribution strategy, I believe 4xGPU V100 can do 50 second epochs or faster.</p>",
      "rawMarkdown": "About NVIDIA V100: yes, `mixed_precision` on Nvidia V100 enables TensorCores.\n\nAbout the TPU v3: On the TPU, with batch size 128, one epoch (train and validate DenseNet201 with 512x512x3) is 70 seconds without `mixed_precision` enabled and 60 seconds with `mixed_precision` enabled. That's 20% increase. Then with `mixed_precision`, you can also double the batch size to 256 and do 50 second epochs gaining another 20%.\n\n4xGPU V100 (comparable hardware as TPUv3's 4 chips) with mixed precision at batch 192 (can't fit 256 with `tf.distribute.MirroredStrategy()`) does 60 second epochs. This is using TensorFlow. With Pytorch and Nvidia apex for mixed precision and distribution strategy, I believe 4xGPU V100 can do 50 second epochs or faster.",
      "votes": null
    },
    {
      "id": "756666",
      "postDate": "02/25/2020 23:48:12",
      "content": "<p>Thank you for sharing. It is a really interesting notebook. \nWhen using the XLA_ACCELERATE flag you set the jit to True, what does this do? I have been unable to find any documentation on it other than the fact it exists.\nMy quick test causes me to run out of memory on my GPU.</p>",
      "rawMarkdown": "Thank you for sharing. It is a really interesting notebook. \nWhen using the XLA_ACCELERATE flag you set the jit to True, what does this do? I have been unable to find any documentation on it other than the fact it exists.\nMy quick test causes me to run out of memory on my GPU.",
      "votes": null
    },
    {
      "id": "756674",
      "postDate": "02/26/2020 00:01:31",
      "content": "<p>XLA_ACCELERATE is explained <a href=\"https://www.tensorflow.org/xla\">here</a>. When training some CNN architectures, it can speed up training by optimizing the mathematics (but as you noticed it can also use more memory). When using GPU V100 and DenseNet201, it makes each epoch go 2% faster but uses more memory. (So it's not worth it here).</p>",
      "rawMarkdown": "XLA_ACCELERATE is explained [here][1]. When training some CNN architectures, it can speed up training by optimizing the mathematics (but as you noticed it can also use more memory). When using GPU V100 and DenseNet201, it makes each epoch go 2% faster but uses more memory. (So it's not worth it here).\n\n[1]: https://www.tensorflow.org/xla",
      "votes": null
    },
    {
      "id": "756680",
      "postDate": "02/26/2020 00:08:33",
      "content": "<p>Thanks I'll give that a read. </p>\n\n<p>I've tried switching MIXED_PRECISION to True on an Nvidia 2080 Ti and the first epoch predicted runtime went to over 5hrs. I might try running it when I go to work to see if it really ends up that slow, but I'm a bit surprised as I expected a speed up. </p>",
      "rawMarkdown": "Thanks I'll give that a read. \n\nI've tried switching MIXED_PRECISION to True on an Nvidia 2080 Ti and the first epoch predicted runtime went to over 5hrs. I might try running it when I go to work to see if it really ends up that slow, but I'm a bit surprised as I expected a speed up.",
      "votes": null
    },
    {
      "id": "756692",
      "postDate": "02/26/2020 00:47:28",
      "content": "<p>If you use MIXED_PRECISION, you must force all softmax layers and the last layer to use <code>fp32</code>. So make sure you use this line (if you didn't copy my whole notebook). Mixed precision is explained by TensorFlow <a href=\"https://www.tensorflow.org/guide/keras/mixed_precision\">here</a></p>\n\n<pre><code>tf.keras.layers.Dense(len(CLASSES), activation='softmax',dtype='float32')\n</code></pre>\n\n<p>TensorFlow mixed precision is experimental, so perhaps it isn't working correctly. Alternatively you could use <a href=\"https://docs.nvidia.com/deeplearning/sdk/dali-developer-guide/docs/index.html\">Nvidia DALI</a> with PyTorch. That will definitely work.</p>",
      "rawMarkdown": "If you use MIXED_PRECISION, you must force all softmax layers and the last layer to use `fp32`. So make sure you use this line (if you didn't copy my whole notebook). Mixed precision is explained by TensorFlow [here][1]\n\n    tf.keras.layers.Dense(len(CLASSES), activation='softmax',dtype='float32')\n\nTensorFlow mixed precision is experimental, so perhaps it isn't working correctly. Alternatively you could use [Nvidia DALI][2] with PyTorch. That will definitely work.\n\n[1]: https://www.tensorflow.org/guide/keras/mixed_precision\n[2]: https://docs.nvidia.com/deeplearning/sdk/dali-developer-guide/docs/index.html",
      "votes": null
    },
    {
      "id": "756843",
      "postDate": "02/26/2020 06:16:53",
      "content": "<p>Thanks I'll give it a go</p>",
      "rawMarkdown": "Thanks I'll give it a go",
      "votes": null
    },
    {
      "id": "757195",
      "postDate": "02/26/2020 14:19:08",
      "content": "<p>But even faster (and cheaper way, comparing to TPU usage cost) is to pre-create tons of augmented images and store them at fast SSD storage. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1696976%2F904384ba972b5733bfb0337f406b5e58%2FUntitled.png?generation=1582726738946010&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "But even faster (and cheaper way, comparing to TPU usage cost) is to pre-create tons of augmented images and store them at fast SSD storage. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1696976%2F904384ba972b5733bfb0337f406b5e58%2FUntitled.png?generation=1582726738946010&amp;alt=media)",
      "votes": null
    },
    {
      "id": "757245",
      "postDate": "02/26/2020 15:06:04",
      "content": "<p>What you describe is slower and costs more money. </p>\n\n<p>What I describe takes 70 seconds per epoch on TPU <strong>without</strong> augmentation and 70 seconds on TPU <strong>with</strong> augmentation. (TPU has cores that were previously not being used which now get used).  In other words, adding fast TPU augmentation doesn't increase time nor money. </p>\n\n<p>In your suggestion, you now must use CPU time and money to save augmented images to disk beforehand (and wait for it to finish) thus costing time and money.</p>",
      "rawMarkdown": "What you describe is slower and costs more money. \n\nWhat I describe takes 70 seconds per epoch on TPU **without** augmentation and 70 seconds on TPU **with** augmentation. (TPU has cores that were previously not being used which now get used).  In other words, adding fast TPU augmentation doesn't increase time nor money. \n\nIn your suggestion, you now must use CPU time and money to save augmented images to disk beforehand (and wait for it to finish) thus costing time and money.",
      "votes": null
    },
    {
      "id": "757405",
      "postDate": "02/26/2020 18:27:34",
      "content": "<p>Yes, but I can do it in parallel, while my model is training. Also CPU time costs almost nothing (you can use up to 20 4-core CPU kernels on kaggle with no charge).</p>\n\n<p>Although I have to admit that there are some disadvantages, such as having the same dataset on every epoch, whilst with on-the-fly augmentation you have a new dataset on every epoch.</p>",
      "rawMarkdown": "Yes, but I can do it in parallel, while my model is training. Also CPU time costs almost nothing (you can use up to 20 4-core CPU kernels on kaggle with no charge).\n\nAlthough I have to admit that there are some disadvantages, such as having the same dataset on every epoch, whilst with on-the-fly augmentation you have a new dataset on every epoch.",
      "votes": null
    },
    {
      "id": "757449",
      "postDate": "02/26/2020 19:24:00",
      "content": "<p>How about loss scaling ? Isn't that something you need with fp16 on V100 ? </p>",
      "rawMarkdown": "How about loss scaling ? Isn't that something you need with fp16 on V100 ?",
      "votes": null
    },
    {
      "id": "757488",
      "postDate": "02/26/2020 20:26:47",
      "content": "<blockquote>\n  <p>How about loss scaling ? Isn't that something you need with fp16 on V100 ?</p>\n</blockquote>\n\n<p>After setting <code>MIXED_PRECISION=True</code>, if you use <code>model.fit()</code> then TensorFlow automatically handles loss scaling for you. If you make your own training loop using <code>tf.GradientTape()</code>, then you need to explicitly deal with loss scaling yourself.</p>",
      "rawMarkdown": "&gt;How about loss scaling ? Isn't that something you need with fp16 on V100 ?\n\nAfter setting `MIXED_PRECISION=True`, if you use `model.fit()` then TensorFlow automatically handles loss scaling for you. If you make your own training loop using `tf.GradientTape()`, then you need to explicitly deal with loss scaling yourself.",
      "votes": null
    },
    {
      "id": "757503",
      "postDate": "02/26/2020 21:05:09",
      "content": "<p>After a bit of googling it looks like I might need to use a specific Nvidia docker image to use the TensorCores on the RTX. At some point I will try it. \n<a href=\"/cdeotte\">@cdeotte</a> are you training locally? If so, are you using a Nvidia docker image? </p>\n\n<p>Nvida documentation seems to imply I have met the per-requisites \n<a href=\"https://docs.nvidia.com/deeplearning/sdk/mixed-precision-training/index.html#prereqs\">https://docs.nvidia.com/deeplearning/sdk/mixed-precision-training/index.html#prereqs</a></p>\n\n<p>Docker image with a test case\n<a href=\"https://github.com/NVIDIA/DeepLearningExamples/tree/master/TensorFlow/Classification/RN50v1.5\">https://github.com/NVIDIA/DeepLearningExamples/tree/master/TensorFlow/Classification/RN50v1.5</a></p>\n\n<p>I'm going to get my local improvement back on to the TPU so I can make submission and I'll revisit this later as I would quite like to crack this mixed precision training. </p>",
      "rawMarkdown": "After a bit of googling it looks like I might need to use a specific Nvidia docker image to use the TensorCores on the RTX. At some point I will try it. \n@cdeotte are you training locally? If so, are you using a Nvidia docker image? \n\nNvida documentation seems to imply I have met the per-requisites \nhttps://docs.nvidia.com/deeplearning/sdk/mixed-precision-training/index.html#prereqs\n\nDocker image with a test case\nhttps://github.com/NVIDIA/DeepLearningExamples/tree/master/TensorFlow/Classification/RN50v1.5\n\nI'm going to get my local improvement back on to the TPU so I can make submission and I'll revisit this later as I would quite like to crack this mixed precision training.",
      "votes": null
    },
    {
      "id": "757520",
      "postDate": "02/26/2020 21:30:54",
      "content": "<p>Thanks for the links. I'm training on a Nvidia GPU V100. I'm not using a docker, I just installed TensorFlow 2.1 via pip, and use <code>tensorflow.keras.applications.DenseNet201()</code>. I assumed that the following two lines at the beginning of my code took care of everything:</p>\n\n<pre><code>policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n</code></pre>\n\n<p>I observe a 20% speed increase with the same batch size. And then when I double batchsize I gain another 20% speed up. I have to admit, I'm new to this and after reading your links, perhaps I need to do more and perhaps I can achieve much greater speedups.</p>",
      "rawMarkdown": "Thanks for the links. I'm training on a Nvidia GPU V100. I'm not using a docker, I just installed TensorFlow 2.1 via pip, and use `tensorflow.keras.applications.DenseNet201()`. I assumed that the following two lines at the beginning of my code took care of everything:\n\n    policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\n    mixed_precision.set_policy(policy)\n\nI observe a 20% speed increase with the same batch size. And then when I double batchsize I gain another 20% speed up. I have to admit, I'm new to this and after reading your links, perhaps I need to do more and perhaps I can achieve much greater speedups.",
      "votes": null
    },
    {
      "id": "757526",
      "postDate": "02/26/2020 21:39:31",
      "content": "<p>That is my understanding is the same. I need dig into it more, something just doesn't seem right. It trains happily without the mixed precision.</p>",
      "rawMarkdown": "That is my understanding is the same. I need dig into it more, something just doesn't seem right. It trains happily without the mixed precision.",
      "votes": null
    },
    {
      "id": "757694",
      "postDate": "02/27/2020 02:54:13",
      "content": "<p>&gt; Yes, but I can do it in parallel, while my model is training</p>\n\n<p>That's the whole point of this discussion post. The CPU <strong>cannot</strong> do it in parallel. It is not fast enough. Let's say you're training 5 folds for 20 epochs each. That is a total of 100 epochs. Each epoch you need 12,000 flower images (512x512x3). In your scenario, you need many many hours to produce 1,200,000 rotated images and store those 943,718,400,000 bytes of data (943GB) to a 1TB SSD harddrive beforehand.</p>\n\n<p>But what you suggest - preprocessing data and saving to disk before training is very helpful in certain situations where you do not need fresh data each epoch. Thanks for pointing out that technique.</p>",
      "rawMarkdown": "&gt; Yes, but I can do it in parallel, while my model is training\n  \nThat's the whole point of this discussion post. The CPU **cannot** do it in parallel. It is not fast enough. Let's say you're training 5 folds for 20 epochs each. That is a total of 100 epochs. Each epoch you need 12,000 flower images (512x512x3). In your scenario, you need many many hours to produce 1,200,000 rotated images and store those 943,718,400,000 bytes of data (943GB) to a 1TB SSD harddrive beforehand.\n\nBut what you suggest - preprocessing data and saving to disk before training is very helpful in certain situations where you do not need fresh data each epoch. Thanks for pointing out that technique.",
      "votes": null
    },
    {
      "id": "758482",
      "postDate": "02/27/2020 20:02:39",
      "content": "<p>I created Augmentation Layer for keras + TF some times ago. But my TF knowledge was limited. Anyway in some cases it works better than CPU:</p>\n\n<p><a href=\"https://github.com/ZFTurbo/Keras-augmentation-layer\">https://github.com/ZFTurbo/Keras-augmentation-layer</a></p>\n\n<p>I believe your ideas could be implemented in similar Layer for easier re-usage.</p>",
      "rawMarkdown": "I created Augmentation Layer for keras + TF some times ago. But my TF knowledge was limited. Anyway in some cases it works better than CPU:\n\nhttps://github.com/ZFTurbo/Keras-augmentation-layer\n\nI believe your ideas could be implemented in similar Layer for easier re-usage.",
      "votes": null
    },
    {
      "id": "758493",
      "postDate": "02/27/2020 20:28:14",
      "content": "<p>That's awesome. Nice job. My team did a similar thing in Kaggle Molecule Comp. We were given 3D information about atoms and needed to calculate distances and angles. We added input layers to our CNN to do that quickly. It was fast.</p>",
      "rawMarkdown": "That's awesome. Nice job. My team did a similar thing in Kaggle Molecule Comp. We were given 3D information about atoms and needed to calculate distances and angles. We added input layers to our CNN to do that quickly. It was fast.",
      "votes": null
    },
    {
      "id": "758567",
      "postDate": "02/27/2020 23:32:25",
      "content": "<p>I tried with a V100 on GCP (<a href=\"https://console.cloud.google.com/ai-platform/notebooks\">Cloud AI Platform Notebooks</a>) with their pre-installed version of TF 2.1</p>\n\n<p>Mixed precision works fine there if I enable XLA with <code>set_jit(True)</code>. It gives a healthy speed boost to the V100. Without XLA however, enabling mixed precision seems to disable the GPU entirely. I'm back to CPU training speeds.</p>",
      "rawMarkdown": "I tried with a V100 on GCP ([Cloud AI Platform Notebooks](https://console.cloud.google.com/ai-platform/notebooks)) with their pre-installed version of TF 2.1\n\nMixed precision works fine there if I enable XLA with `set_jit(True)`. It gives a healthy speed boost to the V100. Without XLA however, enabling mixed precision seems to disable the GPU entirely. I'm back to CPU training speeds.",
      "votes": null
    },
    {
      "id": "759406",
      "postDate": "02/29/2020 02:01:22",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , I have a question mainly concering about the statement <code>Augmentation GPU/TPU</code>. What I understand from this statement is that you are saying the data augmentation is done on accelerators like GPU or TPU.</p>\n\n<p>However, when I read the <a href=\"https://www.tensorflow.org/guide/data_performance\">Better performance with the tf.data API</a>, it looks like the operations done in <code>tf.data.Dataset API</code> are always on CPU. Only when we use the outputs from the <code>tf.data.Dataset</code> to do something, for example, using a batch to train the model, we can make the computations on GPU / TPU.</p>\n\n<p>In your data augmentation kernel, you have</p>\n\n<pre><code>def get_training_dataset(dataset,do_aug=True):\n\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n\n    if do_aug: dataset = dataset.map(transform, num_parallel_calls=AUTO)\n\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n\n    return dataset\n</code></pre>\n\n<p>The <code>transform</code> method is called inside a <code>tf.data.Dataset API</code>, which I think it won't run on GPU / TPU. I haven't try and verify by myself though.</p>\n\n<p>In order to make your great data augumentation <code>transform</code> run on GPU / TPU, one way is not to do this in dataset API. Instead, after we get a batch, we call <code>transform</code> on that batch, and pass the result to model.</p>\n\n<p>However, I don't know how to do this when we use a sequential model + model.fit(). I can possibly make it work if we use subclassed model + custom training loop.</p>\n\n<p>In your reply to <a href=\"/zfturbo\">@zfturbo</a> comments, it seems that you have done somethings (in another competition) to make the processing actually running on GPU, right?</p>",
      "rawMarkdown": "cdeotte , I have a question mainly concering about the statement `Augmentation GPU/TPU`. What I understand from this statement is that you are saying the data augmentation is done on accelerators like GPU or TPU.\n\nHowever, when I read the [Better performance with the tf.data API](https://www.tensorflow.org/guide/data_performance), it looks like the operations done in `tf.data.Dataset API` are always on CPU. Only when we use the outputs from the `tf.data.Dataset` to do something, for example, using a batch to train the model, we can make the computations on GPU / TPU.\n\nIn your data augmentation kernel, you have\n\n    def get_training_dataset(dataset,do_aug=True):\n        \n        dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n        \n        if do_aug: dataset = dataset.map(transform, num_parallel_calls=AUTO)\n        \n        dataset = dataset.repeat() # the training dataset must repeat for several epochs\n        dataset = dataset.shuffle(2048)\n        dataset = dataset.batch(BATCH_SIZE)\n        dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n        \n        return dataset\n\nThe `transform` method is called inside a `tf.data.Dataset API`, which I think it won't run on GPU / TPU. I haven't try and verify by myself though.\n\nIn order to make your great data augumentation `transform` run on GPU / TPU, one way is not to do this in dataset API. Instead, after we get a batch, we call `transform` on that batch, and pass the result to model.\n\nHowever, I don't know how to do this when we use a sequential model + model.fit(). I can possibly make it work if we use subclassed model + custom training loop.\n\nIn your reply to @zfturbo comments, it seems that you have done somethings (in another competition) to make the processing actually running on GPU, right?",
      "votes": null
    },
    {
      "id": "759409",
      "postDate": "02/29/2020 02:08:31",
      "content": "<p>A small implementation detail is that even though tf.data.Datset runs on the \"CPU\", it is not the CPU of your Kaggle VM. There is another VM with cycles to spare which has the TPU attached to it through a PCI link. If you use <code>tf.data.Dataset.prefetch(AUTO)</code>, you can get a lot of data transformations for \"free\" while the TPU is doing its forward and backward pass.</p>",
      "rawMarkdown": "A small implementation detail is that even though tf.data.Datset runs on the \"CPU\", it is not the CPU of your Kaggle VM. There is another VM with cycles to spare which has the TPU attached to it through a PCI link. If you use `tf.data.Dataset.prefetch(AUTO)`, you can get a lot of data transformations for \"free\" while the TPU is doing its forward and backward pass.",
      "votes": null
    },
    {
      "id": "759419",
      "postDate": "02/29/2020 02:28:54",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> , if data processing runs on the CPU in another VM, then what the CPU in the main Kaggle VM is doing during the training?? Why not just using the CPU of the main Kaggle VM to do data processing and let TPU do training? It's not good enough?</p>",
      "rawMarkdown": "mgornergoogle , if data processing runs on the CPU in another VM, then what the CPU in the main Kaggle VM is doing during the training?? Why not just using the CPU of the main Kaggle VM to do data processing and let TPU do training? It's not good enough?",
      "votes": null
    },
    {
      "id": "759950",
      "postDate": "02/29/2020 16:43:07",
      "content": "<p>You are correct Yih-Dar SHIEH. Thanks for the clarification. My understanding of <code>TensorFlow.data.Dataset()</code> was incorrect. I will fix my posts and notebook soon.</p>\n\n<h3>When Using TensorFlow data Dataset</h3>\n\n<p>When adding <code>transform()</code> to <code>TensorFlow.data.Dataset()</code>, the documentation says that preprocessing occurs on the CPU and not the GPU/TPU</p>\n\n<pre><code>dataset = dataset.map(transform)\n</code></pre>\n\n<h3>When Using TensorFlow Model Layers</h3>\n\n<p>Below is how you would preprocess on the GPU/TPU. You can add the rotation transformation directly into your model which occurs on GPU/TPU.</p>\n\n<pre><code>input = tf.keras.layers.Input((512,512,3))\nx = tf.keras.layers.Lambda(transform)(input)\nx = DenseNet201(x)\n</code></pre>\n\n<p>However I would need to update my transform function so that it takes an entire batch as input and outputs an entire batch. Also it wouldn't take labels as input nor output.</p>",
      "rawMarkdown": "You are correct Yih-Dar SHIEH. Thanks for the clarification. My understanding of `TensorFlow.data.Dataset()` was incorrect. I will fix my posts and notebook soon.\n\n### When Using TensorFlow data Dataset\nWhen adding `transform()` to `TensorFlow.data.Dataset()`, the documentation says that preprocessing occurs on the CPU and not the GPU/TPU\n\n    dataset = dataset.map(transform)\n\n### When Using TensorFlow Model Layers\nBelow is how you would preprocess on the GPU/TPU. You can add the rotation transformation directly into your model which occurs on GPU/TPU.\n\n    input = tf.keras.layers.Input((512,512,3))\n    x = tf.keras.layers.Lambda(transform)(input)\n    x = DenseNet201(x)\n\n   However I would need to update my transform function so that it takes an entire batch as input and outputs an entire batch. Also it wouldn't take labels as input nor output.",
      "votes": null
    },
    {
      "id": "759952",
      "postDate": "02/29/2020 16:47:21",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for the clarification Martin. Do you know what CPU, clock speed and core count gets used with TPUv3? (I'm curious to understand time trials from my experiments).</p>",
      "rawMarkdown": "mgornergoogle Thanks for the clarification Martin. Do you know what CPU, clock speed and core count gets used with TPUv3? (I'm curious to understand time trials from my experiments).",
      "votes": null
    },
    {
      "id": "760086",
      "postDate": "02/29/2020 21:07:00",
      "content": "<p><code>However I would need to update my transform function so that it takes an entire batch as input and outputs an entire batch. Also it wouldn't take labels as input nor output.</code></p>\n\n<p>Completely true. I am also working on it, but quite slow :). You can go ahead. Thanks for answering.</p>",
      "rawMarkdown": "`    However I would need to update my transform function so that it takes an entire batch as input and outputs an entire batch. Also it wouldn't take labels as input nor output.`\n\nCompletely true. I am also working on it, but quite slow :). You can go ahead. Thanks for answering.",
      "votes": null
    },
    {
      "id": "761696",
      "postDate": "03/02/2020 21:12:48",
      "content": "<blockquote>\n  <p>Do you know what CPU, clock speed and core count gets used with TPUv3?</p>\n</blockquote>\n\n<p>Not precisely but I know it's a lot. This has been dimensioned to make sure the data-hungry TPUv3 can always be fed with enough data. </p>",
      "rawMarkdown": "&gt; Do you know what CPU, clock speed and core count gets used with TPUv3?\n\nNot precisely but I know it's a lot. This has been dimensioned to make sure the data-hungry TPUv3 can always be fed with enough data.",
      "votes": null
    },
    {
      "id": "764814",
      "postDate": "03/05/2020 23:46:53",
      "content": "<p>Are you considering upgrading the gpu too?  V100 with 32 GB are 1.5 years old, and all we have on Kaggle notebooks are P100.  Or is it a decision by Kaggle to not offer recent GPU and move people to tpu only for recent pretrained models?</p>",
      "rawMarkdown": "Are you considering upgrading the gpu too?  V100 with 32 GB are 1.5 years old, and all we have on Kaggle notebooks are P100.  Or is it a decision by Kaggle to not offer recent GPU and move people to tpu only for recent pretrained models?",
      "votes": null
    },
    {
      "id": "872963",
      "postDate": "06/03/2020 16:41:52",
      "content": "<p>Quick pointers to other data augmentation techniques that work well on TPU/GPU:\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935\">CutMix and MixUp on GPU/TPU</a>\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132986\">GridMask data augmentation on GPU/TPU</a>\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/133168\">batch implementation of rotation/shear/zoom/shift</a>\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134264\">batch implementations of CutMix, MixUp, GridMask</a>\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134557\">perspective transformations</a></p>",
      "rawMarkdown": "Quick pointers to other data augmentation techniques that work well on TPU/GPU:\n - [CutMix and MixUp on GPU/TPU](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935)\n - [GridMask data augmentation on GPU/TPU](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132986)\n - [batch implementation of rotation/shear/zoom/shift](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/133168)\n - [batch implementations of CutMix, MixUp, GridMask](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134264)\n - [perspective transformations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134557)",
      "votes": null
    },
    {
      "id": "953514",
      "postDate": "07/31/2020 21:03:53",
      "content": "<p>This is really helpful for beginners. Thank you for sharing your work. May I please ask how to decide which augmentation transformations to begin with (specially if beginners are attempting data augmentation)?</p>",
      "rawMarkdown": "This is really helpful for beginners. Thank you for sharing your work. May I please ask how to decide which augmentation transformations to begin with (specially if beginners are attempting data augmentation)?",
      "votes": null
    },
    {
      "id": "953558",
      "postDate": "07/31/2020 21:45:07",
      "content": "<p>The idea for augmentation is two fold. First you want to generate more data that looks like your train data. So for example, it is appropriate to rotate images any angle in Melanoma skin lesion detection. But it wouldn't be appropriate to rotate handwritten digits more than plus minus 15 degrees. Since the more the better, just do every type of augmentation that creates realistic additional images.</p>\n\n<p>The second reason for augmentation is to make learning more difficult for your model. That forces your model to learn the underlying idea more generally. This is accomplished with augmentation like coarse dropout or cutout which removes parts of the image. You should only obscure an image to the point that a human can still identify the image more or less. When applying augmentation for this reason, you will need to train for more epochs because your model learns more slowly. But it will become more intelligent at the end.</p>",
      "rawMarkdown": "The idea for augmentation is two fold. First you want to generate more data that looks like your train data. So for example, it is appropriate to rotate images any angle in Melanoma skin lesion detection. But it wouldn't be appropriate to rotate handwritten digits more than plus minus 15 degrees. Since the more the better, just do every type of augmentation that creates realistic additional images.\n\nThe second reason for augmentation is to make learning more difficult for your model. That forces your model to learn the underlying idea more generally. This is accomplished with augmentation like coarse dropout or cutout which removes parts of the image. You should only obscure an image to the point that a human can still identify the image more or less. When applying augmentation for this reason, you will need to train for more epochs because your model learns more slowly. But it will become more intelligent at the end.",
      "votes": null
    },
    {
      "id": "953568",
      "postDate": "07/31/2020 21:54:31",
      "content": "<p>Thank you. Very well explained. Keeping the images still realistic is perhaps the key while augmenting. </p>",
      "rawMarkdown": "Thank you. Very well explained. Keeping the images still realistic is perhaps the key while augmenting.",
      "votes": null
    },
    {
      "id": "953570",
      "postDate": "07/31/2020 22:08:48",
      "content": "<p>Yes. That is why it is always good to display your images after applying augmentation. </p>\n\n<p>Look at dozens of the augmented images yourself. Adjust the hyperparameters (like how much rotation, zoom, color contrast, etc) to the maximum you can while keeping the images realistic looking. Then apply the most different types of augmentation that you can.</p>\n\n<p>Deep learning needs lots of training data and augmentation makes it for us. Augmentation is one of the best ways to increase accuracy for deep learning.</p>",
      "rawMarkdown": "Yes. That is why it is always good to display your images after applying augmentation. \n\nLook at dozens of the augmented images yourself. Adjust the hyperparameters (like how much rotation, zoom, color contrast, etc) to the maximum you can while keeping the images realistic looking. Then apply the most different types of augmentation that you can.\n\nDeep learning needs lots of training data and augmentation makes it for us. Augmentation is one of the best ways to increase accuracy for deep learning.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 755524,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "02/24/2020 21:19:56",
      "content": "<p>Thanks, great to have rotation! Yesterday, I just tried to find if <code>tf.image</code> has arbitrary rotation method, and it doesn't. And tensorflow addons <code>tfa.image.rotate</code> is not working ...</p>",
      "votes": null,
      "replies": [
        {
          "id": 755545,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/24/2020 21:40:59",
          "content": "<p>Yes augmentation is difficult in this competition because no libraries work. Now you can do rotation, shear, zoom, and shift. Enjoy!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755560,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "02/24/2020 22:21:16",
      "content": "<p>You just shared a huge building block for many other augmentation methods that will need to be ported to tensorflow2 and TPU. Thank you!</p>\n\n<p>I am nominating you for the \"TPU Star\" Prize ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 755591,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/24/2020 23:15:02",
          "content": "<p>Thanks Ilu</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755617,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/25/2020 00:25:37",
      "content": "<p>Fantastic notebook. Thanks.\nCould you explain what the <code>tf.keras.mixed_precision.experimental.Policy</code>does and why it is needed ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 755633,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/25/2020 01:01:37",
          "content": "<p>It is not needed. Currently the boolean switches are set to <code>False</code>. If you turn it on then Nvidia GPU V100 goes 40% faster and fits twice as large batch sizes. And TPUv3 goes 40% faster and fits twice as large batch sizes. (GPU P100 doesn't speed up but it does fit larger batch sizes).</p>\n\n<p>When <code>MIXED_PRECISION = True</code>, TensorFlow uses <code>fp16</code> instead of <code>fp32</code> wherever it can safely.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756625,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/25/2020 22:43:47",
          "content": "<p>About NVIDIA V100: does this setting enable the TensorCores on the V100 ?</p>\n\n<p>About the TPU v3: as you said, this setting is not needed for the TPU v3 to use mixed precision. But it does optimize memory usage because tensors are then stored in memory in bfloat16 format. I assume the 40% speed increase you mention is attributable solely to the memory optimization which allows bigger batches. Correct ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756653,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/25/2020 23:19:56",
          "content": "<p>About NVIDIA V100: yes, <code>mixed_precision</code> on Nvidia V100 enables TensorCores.</p>\n\n<p>About the TPU v3: On the TPU, with batch size 128, one epoch (train and validate DenseNet201 with 512x512x3) is 70 seconds without <code>mixed_precision</code> enabled and 60 seconds with <code>mixed_precision</code> enabled. That's 20% increase. Then with <code>mixed_precision</code>, you can also double the batch size to 256 and do 50 second epochs gaining another 20%.</p>\n\n<p>4xGPU V100 (comparable hardware as TPUv3's 4 chips) with mixed precision at batch 192 (can't fit 256 with <code>tf.distribute.MirroredStrategy()</code>) does 60 second epochs. This is using TensorFlow. With Pytorch and Nvidia apex for mixed precision and distribution strategy, I believe 4xGPU V100 can do 50 second epochs or faster.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756666,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "02/25/2020 23:48:12",
          "content": "<p>Thank you for sharing. It is a really interesting notebook. \nWhen using the XLA_ACCELERATE flag you set the jit to True, what does this do? I have been unable to find any documentation on it other than the fact it exists.\nMy quick test causes me to run out of memory on my GPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756674,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/26/2020 00:01:31",
          "content": "<p>XLA_ACCELERATE is explained <a href=\"https://www.tensorflow.org/xla\">here</a>. When training some CNN architectures, it can speed up training by optimizing the mathematics (but as you noticed it can also use more memory). When using GPU V100 and DenseNet201, it makes each epoch go 2% faster but uses more memory. (So it's not worth it here).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756680,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "02/26/2020 00:08:33",
          "content": "<p>Thanks I'll give that a read. </p>\n\n<p>I've tried switching MIXED_PRECISION to True on an Nvidia 2080 Ti and the first epoch predicted runtime went to over 5hrs. I might try running it when I go to work to see if it really ends up that slow, but I'm a bit surprised as I expected a speed up. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756692,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/26/2020 00:47:28",
          "content": "<p>If you use MIXED_PRECISION, you must force all softmax layers and the last layer to use <code>fp32</code>. So make sure you use this line (if you didn't copy my whole notebook). Mixed precision is explained by TensorFlow <a href=\"https://www.tensorflow.org/guide/keras/mixed_precision\">here</a></p>\n\n<pre><code>tf.keras.layers.Dense(len(CLASSES), activation='softmax',dtype='float32')\n</code></pre>\n\n<p>TensorFlow mixed precision is experimental, so perhaps it isn't working correctly. Alternatively you could use <a href=\"https://docs.nvidia.com/deeplearning/sdk/dali-developer-guide/docs/index.html\">Nvidia DALI</a> with PyTorch. That will definitely work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756843,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "02/26/2020 06:16:53",
          "content": "<p>Thanks I'll give it a go</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757449,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/26/2020 19:24:00",
          "content": "<p>How about loss scaling ? Isn't that something you need with fp16 on V100 ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757488,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/26/2020 20:26:47",
          "content": "<blockquote>\n  <p>How about loss scaling ? Isn't that something you need with fp16 on V100 ?</p>\n</blockquote>\n\n<p>After setting <code>MIXED_PRECISION=True</code>, if you use <code>model.fit()</code> then TensorFlow automatically handles loss scaling for you. If you make your own training loop using <code>tf.GradientTape()</code>, then you need to explicitly deal with loss scaling yourself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757503,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "02/26/2020 21:05:09",
          "content": "<p>After a bit of googling it looks like I might need to use a specific Nvidia docker image to use the TensorCores on the RTX. At some point I will try it. \n<a href=\"/cdeotte\">@cdeotte</a> are you training locally? If so, are you using a Nvidia docker image? </p>\n\n<p>Nvida documentation seems to imply I have met the per-requisites \n<a href=\"https://docs.nvidia.com/deeplearning/sdk/mixed-precision-training/index.html#prereqs\">https://docs.nvidia.com/deeplearning/sdk/mixed-precision-training/index.html#prereqs</a></p>\n\n<p>Docker image with a test case\n<a href=\"https://github.com/NVIDIA/DeepLearningExamples/tree/master/TensorFlow/Classification/RN50v1.5\">https://github.com/NVIDIA/DeepLearningExamples/tree/master/TensorFlow/Classification/RN50v1.5</a></p>\n\n<p>I'm going to get my local improvement back on to the TPU so I can make submission and I'll revisit this later as I would quite like to crack this mixed precision training. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757520,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/26/2020 21:30:54",
          "content": "<p>Thanks for the links. I'm training on a Nvidia GPU V100. I'm not using a docker, I just installed TensorFlow 2.1 via pip, and use <code>tensorflow.keras.applications.DenseNet201()</code>. I assumed that the following two lines at the beginning of my code took care of everything:</p>\n\n<pre><code>policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n</code></pre>\n\n<p>I observe a 20% speed increase with the same batch size. And then when I double batchsize I gain another 20% speed up. I have to admit, I'm new to this and after reading your links, perhaps I need to do more and perhaps I can achieve much greater speedups.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757526,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "02/26/2020 21:39:31",
          "content": "<p>That is my understanding is the same. I need dig into it more, something just doesn't seem right. It trains happily without the mixed precision.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758567,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/27/2020 23:32:25",
          "content": "<p>I tried with a V100 on GCP (<a href=\"https://console.cloud.google.com/ai-platform/notebooks\">Cloud AI Platform Notebooks</a>) with their pre-installed version of TF 2.1</p>\n\n<p>Mixed precision works fine there if I enable XLA with <code>set_jit(True)</code>. It gives a healthy speed boost to the V100. Without XLA however, enabling mixed precision seems to disable the GPU entirely. I'm back to CPU training speeds.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 756235,
      "author_name": "dasmehdixtr",
      "author_url": "",
      "post_date": "02/25/2020 14:40:51",
      "content": "<p>Such a good notebook, thanks for share! 💯 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 757195,
      "author_name": "nroman",
      "author_url": "",
      "post_date": "02/26/2020 14:19:08",
      "content": "<p>But even faster (and cheaper way, comparing to TPU usage cost) is to pre-create tons of augmented images and store them at fast SSD storage. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1696976%2F904384ba972b5733bfb0337f406b5e58%2FUntitled.png?generation=1582726738946010&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 757245,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/26/2020 15:06:04",
          "content": "<p>What you describe is slower and costs more money. </p>\n\n<p>What I describe takes 70 seconds per epoch on TPU <strong>without</strong> augmentation and 70 seconds on TPU <strong>with</strong> augmentation. (TPU has cores that were previously not being used which now get used).  In other words, adding fast TPU augmentation doesn't increase time nor money. </p>\n\n<p>In your suggestion, you now must use CPU time and money to save augmented images to disk beforehand (and wait for it to finish) thus costing time and money.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757405,
          "author_name": "nroman",
          "author_url": "",
          "post_date": "02/26/2020 18:27:34",
          "content": "<p>Yes, but I can do it in parallel, while my model is training. Also CPU time costs almost nothing (you can use up to 20 4-core CPU kernels on kaggle with no charge).</p>\n\n<p>Although I have to admit that there are some disadvantages, such as having the same dataset on every epoch, whilst with on-the-fly augmentation you have a new dataset on every epoch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757694,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/27/2020 02:54:13",
          "content": "<p>&gt; Yes, but I can do it in parallel, while my model is training</p>\n\n<p>That's the whole point of this discussion post. The CPU <strong>cannot</strong> do it in parallel. It is not fast enough. Let's say you're training 5 folds for 20 epochs each. That is a total of 100 epochs. Each epoch you need 12,000 flower images (512x512x3). In your scenario, you need many many hours to produce 1,200,000 rotated images and store those 943,718,400,000 bytes of data (943GB) to a 1TB SSD harddrive beforehand.</p>\n\n<p>But what you suggest - preprocessing data and saving to disk before training is very helpful in certain situations where you do not need fresh data each epoch. Thanks for pointing out that technique.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 758482,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "02/27/2020 20:02:39",
      "content": "<p>I created Augmentation Layer for keras + TF some times ago. But my TF knowledge was limited. Anyway in some cases it works better than CPU:</p>\n\n<p><a href=\"https://github.com/ZFTurbo/Keras-augmentation-layer\">https://github.com/ZFTurbo/Keras-augmentation-layer</a></p>\n\n<p>I believe your ideas could be implemented in similar Layer for easier re-usage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 758493,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/27/2020 20:28:14",
          "content": "<p>That's awesome. Nice job. My team did a similar thing in Kaggle Molecule Comp. We were given 3D information about atoms and needed to calculate distances and angles. We added input layers to our CNN to do that quickly. It was fast.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759406,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "02/29/2020 02:01:22",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , I have a question mainly concering about the statement <code>Augmentation GPU/TPU</code>. What I understand from this statement is that you are saying the data augmentation is done on accelerators like GPU or TPU.</p>\n\n<p>However, when I read the <a href=\"https://www.tensorflow.org/guide/data_performance\">Better performance with the tf.data API</a>, it looks like the operations done in <code>tf.data.Dataset API</code> are always on CPU. Only when we use the outputs from the <code>tf.data.Dataset</code> to do something, for example, using a batch to train the model, we can make the computations on GPU / TPU.</p>\n\n<p>In your data augmentation kernel, you have</p>\n\n<pre><code>def get_training_dataset(dataset,do_aug=True):\n\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n\n    if do_aug: dataset = dataset.map(transform, num_parallel_calls=AUTO)\n\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n\n    return dataset\n</code></pre>\n\n<p>The <code>transform</code> method is called inside a <code>tf.data.Dataset API</code>, which I think it won't run on GPU / TPU. I haven't try and verify by myself though.</p>\n\n<p>In order to make your great data augumentation <code>transform</code> run on GPU / TPU, one way is not to do this in dataset API. Instead, after we get a batch, we call <code>transform</code> on that batch, and pass the result to model.</p>\n\n<p>However, I don't know how to do this when we use a sequential model + model.fit(). I can possibly make it work if we use subclassed model + custom training loop.</p>\n\n<p>In your reply to <a href=\"/zfturbo\">@zfturbo</a> comments, it seems that you have done somethings (in another competition) to make the processing actually running on GPU, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 759409,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/29/2020 02:08:31",
          "content": "<p>A small implementation detail is that even though tf.data.Datset runs on the \"CPU\", it is not the CPU of your Kaggle VM. There is another VM with cycles to spare which has the TPU attached to it through a PCI link. If you use <code>tf.data.Dataset.prefetch(AUTO)</code>, you can get a lot of data transformations for \"free\" while the TPU is doing its forward and backward pass.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759419,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/29/2020 02:28:54",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> , if data processing runs on the CPU in another VM, then what the CPU in the main Kaggle VM is doing during the training?? Why not just using the CPU of the main Kaggle VM to do data processing and let TPU do training? It's not good enough?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759950,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/29/2020 16:43:07",
          "content": "<p>You are correct Yih-Dar SHIEH. Thanks for the clarification. My understanding of <code>TensorFlow.data.Dataset()</code> was incorrect. I will fix my posts and notebook soon.</p>\n\n<h3>When Using TensorFlow data Dataset</h3>\n\n<p>When adding <code>transform()</code> to <code>TensorFlow.data.Dataset()</code>, the documentation says that preprocessing occurs on the CPU and not the GPU/TPU</p>\n\n<pre><code>dataset = dataset.map(transform)\n</code></pre>\n\n<h3>When Using TensorFlow Model Layers</h3>\n\n<p>Below is how you would preprocess on the GPU/TPU. You can add the rotation transformation directly into your model which occurs on GPU/TPU.</p>\n\n<pre><code>input = tf.keras.layers.Input((512,512,3))\nx = tf.keras.layers.Lambda(transform)(input)\nx = DenseNet201(x)\n</code></pre>\n\n<p>However I would need to update my transform function so that it takes an entire batch as input and outputs an entire batch. Also it wouldn't take labels as input nor output.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759952,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/29/2020 16:47:21",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for the clarification Martin. Do you know what CPU, clock speed and core count gets used with TPUv3? (I'm curious to understand time trials from my experiments).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760086,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/29/2020 21:07:00",
          "content": "<p><code>However I would need to update my transform function so that it takes an entire batch as input and outputs an entire batch. Also it wouldn't take labels as input nor output.</code></p>\n\n<p>Completely true. I am also working on it, but quite slow :). You can go ahead. Thanks for answering.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761696,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/02/2020 21:12:48",
          "content": "<blockquote>\n  <p>Do you know what CPU, clock speed and core count gets used with TPUv3?</p>\n</blockquote>\n\n<p>Not precisely but I know it's a lot. This has been dimensioned to make sure the data-hungry TPUv3 can always be fed with enough data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764814,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/05/2020 23:46:53",
          "content": "<p>Are you considering upgrading the gpu too?  V100 with 32 GB are 1.5 years old, and all we have on Kaggle notebooks are P100.  Or is it a decision by Kaggle to not offer recent GPU and move people to tpu only for recent pretrained models?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 872963,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "06/03/2020 16:41:52",
      "content": "<p>Quick pointers to other data augmentation techniques that work well on TPU/GPU:\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935\">CutMix and MixUp on GPU/TPU</a>\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132986\">GridMask data augmentation on GPU/TPU</a>\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/133168\">batch implementation of rotation/shear/zoom/shift</a>\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134264\">batch implementations of CutMix, MixUp, GridMask</a>\n - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134557\">perspective transformations</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953514,
      "author_name": "ektasharma",
      "author_url": "",
      "post_date": "07/31/2020 21:03:53",
      "content": "<p>This is really helpful for beginners. Thank you for sharing your work. May I please ask how to decide which augmentation transformations to begin with (specially if beginners are attempting data augmentation)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 953558,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/31/2020 21:45:07",
          "content": "<p>The idea for augmentation is two fold. First you want to generate more data that looks like your train data. So for example, it is appropriate to rotate images any angle in Melanoma skin lesion detection. But it wouldn't be appropriate to rotate handwritten digits more than plus minus 15 degrees. Since the more the better, just do every type of augmentation that creates realistic additional images.</p>\n\n<p>The second reason for augmentation is to make learning more difficult for your model. That forces your model to learn the underlying idea more generally. This is accomplished with augmentation like coarse dropout or cutout which removes parts of the image. You should only obscure an image to the point that a human can still identify the image more or less. When applying augmentation for this reason, you will need to train for more epochs because your model learns more slowly. But it will become more intelligent at the end.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 953568,
          "author_name": "ektasharma",
          "author_url": "",
          "post_date": "07/31/2020 21:54:31",
          "content": "<p>Thank you. Very well explained. Keeping the images still realistic is perhaps the key while augmenting. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 953570,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/31/2020 22:08:48",
          "content": "<p>Yes. That is why it is always good to display your images after applying augmentation. </p>\n\n<p>Look at dozens of the augmented images yourself. Adjust the hyperparameters (like how much rotation, zoom, color contrast, etc) to the maximum you can while keeping the images realistic looking. Then apply the most different types of augmentation that you can.</p>\n\n<p>Deep learning needs lots of training data and augmentation makes it for us. Augmentation is one of the best ways to increase accuracy for deep learning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "755450": "# Data Augmentation - Rotation, Shear, Zoom, Shift\nI posted a starter kernel [here][1] showing how to apply rotation, shear, zoom, and shift augmentation when using `TensorFlow.data.Dataset()`. It performs augmentation on the GPU/TPU instead of CPU for maximum speed. Feel free to use this code in your own notebooks.\n# Speed Analysis\nA GPU or TPU can process 200+ images (512x512x3) per second while training DenseNet201. That's incredibly fast! If we try to augment images in the CPU, then we may not be able to provide the GPU/TPU with images fast enough and thus we will slow down our training.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ffbb804fdbfe3a6bfb54acbf35039e36a%2Fcpu.jpg?generation=1582409535436227&amp;alt=media)\n  \n# Preprocess on GPU/TPU\nThis is a great competition to appreciate the need to preprocess data on the GPU/TPU. If you wish to rotate a 512x512x3 image, that requires 5,000,000 multiplies and additions per image! \n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fac5072f04b8e627020c439ab8326b038%2Frotate.JPG?generation=1582409737942526&amp;alt=media)\n  \nFor each pixel in your augmented image, you must find which pixel value in the original image to use. In the example above, we wish to determine which pixel to place in location `(1,3)`. So we must multiply the coordinate `(1,3)` by a rotation matrix to determine that we want pixel `(3,2)` from the original image which is color pink. We then place a pink pixel in destination image's pixel `(1,3)`.\n\n# Code Does Not Use For-Loops\nWhen writing code for GPU/TPU (i.e. in TensorFlow, Numba, CuPy, CUDA, etc), we want to avoid using for-loops. With a 512x512 image, we must determine the pixel values for 262,144 = 512 x 512 pixels. We can either write a for-loop as `for i in range(262144):` or we can create a matrix `P` of size `(2,262144)` and multiply it by a `(2,2)` rotation matrix `R`. In the later method, this allows GPU/TPU to compute the 262,144 new coordinates in parallel since each GPU/TPU thread can multiply one column of matrix `P` by matrix `R` simultaneously. In the former method, the for-loop prevents us from parallelization (since each iteration must wait for previous iterations to complete).\n\nOnce we have the coordinates of the 262,144 pixels that we want, we use `tensorflow.gather_nd()` which divides the job of finding the pixels in the original image among thousands of different GPU/TPU threads. Once again, we did not write `for i in range(262144):` to locate the 262,144 original pixels which would have prevented parallelism.\n\n# Example Rotated Image\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5042f32b75cd19ad71d5a35c5587ae79%2FScreen%20Shot%202020-02-24%20at%2011.32.02%20AM.png?generation=1582572736200811&amp;alt=media)\n\n\n[1]: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96",
    "755524": "Thanks, great to have rotation! Yesterday, I just tried to find if `tf.image` has arbitrary rotation method, and it doesn't. And tensorflow addons `tfa.image.rotate` is not working ...",
    "755545": "Yes augmentation is difficult in this competition because no libraries work. Now you can do rotation, shear, zoom, and shift. Enjoy!",
    "755560": "You just shared a huge building block for many other augmentation methods that will need to be ported to tensorflow2 and TPU. Thank you!\n\nI am nominating you for the \"TPU Star\" Prize ;)",
    "755591": "Thanks Ilu",
    "755617": "Fantastic notebook. Thanks.\nCould you explain what the `tf.keras.mixed_precision.experimental.Policy `does and why it is needed ?",
    "755633": "It is not needed. Currently the boolean switches are set to `False`. If you turn it on then Nvidia GPU V100 goes 40% faster and fits twice as large batch sizes. And TPUv3 goes 40% faster and fits twice as large batch sizes. (GPU P100 doesn't speed up but it does fit larger batch sizes).\n\nWhen `MIXED_PRECISION = True`, TensorFlow uses `fp16` instead of `fp32` wherever it can safely.",
    "756235": "Such a good notebook, thanks for share! 💯",
    "756625": "About NVIDIA V100: does this setting enable the TensorCores on the V100 ?\n\nAbout the TPU v3: as you said, this setting is not needed for the TPU v3 to use mixed precision. But it does optimize memory usage because tensors are then stored in memory in bfloat16 format. I assume the 40% speed increase you mention is attributable solely to the memory optimization which allows bigger batches. Correct ?",
    "756653": "About NVIDIA V100: yes, `mixed_precision` on Nvidia V100 enables TensorCores.\n\nAbout the TPU v3: On the TPU, with batch size 128, one epoch (train and validate DenseNet201 with 512x512x3) is 70 seconds without `mixed_precision` enabled and 60 seconds with `mixed_precision` enabled. That's 20% increase. Then with `mixed_precision`, you can also double the batch size to 256 and do 50 second epochs gaining another 20%.\n\n4xGPU V100 (comparable hardware as TPUv3's 4 chips) with mixed precision at batch 192 (can't fit 256 with `tf.distribute.MirroredStrategy()`) does 60 second epochs. This is using TensorFlow. With Pytorch and Nvidia apex for mixed precision and distribution strategy, I believe 4xGPU V100 can do 50 second epochs or faster.",
    "756666": "Thank you for sharing. It is a really interesting notebook. \nWhen using the XLA_ACCELERATE flag you set the jit to True, what does this do? I have been unable to find any documentation on it other than the fact it exists.\nMy quick test causes me to run out of memory on my GPU.",
    "756674": "XLA_ACCELERATE is explained [here][1]. When training some CNN architectures, it can speed up training by optimizing the mathematics (but as you noticed it can also use more memory). When using GPU V100 and DenseNet201, it makes each epoch go 2% faster but uses more memory. (So it's not worth it here).\n\n[1]: https://www.tensorflow.org/xla",
    "756680": "Thanks I'll give that a read. \n\nI've tried switching MIXED_PRECISION to True on an Nvidia 2080 Ti and the first epoch predicted runtime went to over 5hrs. I might try running it when I go to work to see if it really ends up that slow, but I'm a bit surprised as I expected a speed up.",
    "756692": "If you use MIXED_PRECISION, you must force all softmax layers and the last layer to use `fp32`. So make sure you use this line (if you didn't copy my whole notebook). Mixed precision is explained by TensorFlow [here][1]\n\n    tf.keras.layers.Dense(len(CLASSES), activation='softmax',dtype='float32')\n\nTensorFlow mixed precision is experimental, so perhaps it isn't working correctly. Alternatively you could use [Nvidia DALI][2] with PyTorch. That will definitely work.\n\n[1]: https://www.tensorflow.org/guide/keras/mixed_precision\n[2]: https://docs.nvidia.com/deeplearning/sdk/dali-developer-guide/docs/index.html",
    "756843": "Thanks I'll give it a go",
    "757195": "But even faster (and cheaper way, comparing to TPU usage cost) is to pre-create tons of augmented images and store them at fast SSD storage. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1696976%2F904384ba972b5733bfb0337f406b5e58%2FUntitled.png?generation=1582726738946010&amp;alt=media)",
    "757245": "What you describe is slower and costs more money. \n\nWhat I describe takes 70 seconds per epoch on TPU **without** augmentation and 70 seconds on TPU **with** augmentation. (TPU has cores that were previously not being used which now get used).  In other words, adding fast TPU augmentation doesn't increase time nor money. \n\nIn your suggestion, you now must use CPU time and money to save augmented images to disk beforehand (and wait for it to finish) thus costing time and money.",
    "757405": "Yes, but I can do it in parallel, while my model is training. Also CPU time costs almost nothing (you can use up to 20 4-core CPU kernels on kaggle with no charge).\n\nAlthough I have to admit that there are some disadvantages, such as having the same dataset on every epoch, whilst with on-the-fly augmentation you have a new dataset on every epoch.",
    "757449": "How about loss scaling ? Isn't that something you need with fp16 on V100 ?",
    "757488": "&gt;How about loss scaling ? Isn't that something you need with fp16 on V100 ?\n\nAfter setting `MIXED_PRECISION=True`, if you use `model.fit()` then TensorFlow automatically handles loss scaling for you. If you make your own training loop using `tf.GradientTape()`, then you need to explicitly deal with loss scaling yourself.",
    "757503": "After a bit of googling it looks like I might need to use a specific Nvidia docker image to use the TensorCores on the RTX. At some point I will try it. \n@cdeotte are you training locally? If so, are you using a Nvidia docker image? \n\nNvida documentation seems to imply I have met the per-requisites \nhttps://docs.nvidia.com/deeplearning/sdk/mixed-precision-training/index.html#prereqs\n\nDocker image with a test case\nhttps://github.com/NVIDIA/DeepLearningExamples/tree/master/TensorFlow/Classification/RN50v1.5\n\nI'm going to get my local improvement back on to the TPU so I can make submission and I'll revisit this later as I would quite like to crack this mixed precision training.",
    "757520": "Thanks for the links. I'm training on a Nvidia GPU V100. I'm not using a docker, I just installed TensorFlow 2.1 via pip, and use `tensorflow.keras.applications.DenseNet201()`. I assumed that the following two lines at the beginning of my code took care of everything:\n\n    policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\n    mixed_precision.set_policy(policy)\n\nI observe a 20% speed increase with the same batch size. And then when I double batchsize I gain another 20% speed up. I have to admit, I'm new to this and after reading your links, perhaps I need to do more and perhaps I can achieve much greater speedups.",
    "757526": "That is my understanding is the same. I need dig into it more, something just doesn't seem right. It trains happily without the mixed precision.",
    "757694": "&gt; Yes, but I can do it in parallel, while my model is training\n  \nThat's the whole point of this discussion post. The CPU **cannot** do it in parallel. It is not fast enough. Let's say you're training 5 folds for 20 epochs each. That is a total of 100 epochs. Each epoch you need 12,000 flower images (512x512x3). In your scenario, you need many many hours to produce 1,200,000 rotated images and store those 943,718,400,000 bytes of data (943GB) to a 1TB SSD harddrive beforehand.\n\nBut what you suggest - preprocessing data and saving to disk before training is very helpful in certain situations where you do not need fresh data each epoch. Thanks for pointing out that technique.",
    "758482": "I created Augmentation Layer for keras + TF some times ago. But my TF knowledge was limited. Anyway in some cases it works better than CPU:\n\nhttps://github.com/ZFTurbo/Keras-augmentation-layer\n\nI believe your ideas could be implemented in similar Layer for easier re-usage.",
    "758493": "That's awesome. Nice job. My team did a similar thing in Kaggle Molecule Comp. We were given 3D information about atoms and needed to calculate distances and angles. We added input layers to our CNN to do that quickly. It was fast.",
    "758567": "I tried with a V100 on GCP ([Cloud AI Platform Notebooks](https://console.cloud.google.com/ai-platform/notebooks)) with their pre-installed version of TF 2.1\n\nMixed precision works fine there if I enable XLA with `set_jit(True)`. It gives a healthy speed boost to the V100. Without XLA however, enabling mixed precision seems to disable the GPU entirely. I'm back to CPU training speeds.",
    "759406": "cdeotte , I have a question mainly concering about the statement `Augmentation GPU/TPU`. What I understand from this statement is that you are saying the data augmentation is done on accelerators like GPU or TPU.\n\nHowever, when I read the [Better performance with the tf.data API](https://www.tensorflow.org/guide/data_performance), it looks like the operations done in `tf.data.Dataset API` are always on CPU. Only when we use the outputs from the `tf.data.Dataset` to do something, for example, using a batch to train the model, we can make the computations on GPU / TPU.\n\nIn your data augmentation kernel, you have\n\n    def get_training_dataset(dataset,do_aug=True):\n        \n        dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n        \n        if do_aug: dataset = dataset.map(transform, num_parallel_calls=AUTO)\n        \n        dataset = dataset.repeat() # the training dataset must repeat for several epochs\n        dataset = dataset.shuffle(2048)\n        dataset = dataset.batch(BATCH_SIZE)\n        dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n        \n        return dataset\n\nThe `transform` method is called inside a `tf.data.Dataset API`, which I think it won't run on GPU / TPU. I haven't try and verify by myself though.\n\nIn order to make your great data augumentation `transform` run on GPU / TPU, one way is not to do this in dataset API. Instead, after we get a batch, we call `transform` on that batch, and pass the result to model.\n\nHowever, I don't know how to do this when we use a sequential model + model.fit(). I can possibly make it work if we use subclassed model + custom training loop.\n\nIn your reply to @zfturbo comments, it seems that you have done somethings (in another competition) to make the processing actually running on GPU, right?",
    "759409": "A small implementation detail is that even though tf.data.Datset runs on the \"CPU\", it is not the CPU of your Kaggle VM. There is another VM with cycles to spare which has the TPU attached to it through a PCI link. If you use `tf.data.Dataset.prefetch(AUTO)`, you can get a lot of data transformations for \"free\" while the TPU is doing its forward and backward pass.",
    "759419": "mgornergoogle , if data processing runs on the CPU in another VM, then what the CPU in the main Kaggle VM is doing during the training?? Why not just using the CPU of the main Kaggle VM to do data processing and let TPU do training? It's not good enough?",
    "759950": "You are correct Yih-Dar SHIEH. Thanks for the clarification. My understanding of `TensorFlow.data.Dataset()` was incorrect. I will fix my posts and notebook soon.\n\n### When Using TensorFlow data Dataset\nWhen adding `transform()` to `TensorFlow.data.Dataset()`, the documentation says that preprocessing occurs on the CPU and not the GPU/TPU\n\n    dataset = dataset.map(transform)\n\n### When Using TensorFlow Model Layers\nBelow is how you would preprocess on the GPU/TPU. You can add the rotation transformation directly into your model which occurs on GPU/TPU.\n\n    input = tf.keras.layers.Input((512,512,3))\n    x = tf.keras.layers.Lambda(transform)(input)\n    x = DenseNet201(x)\n\n   However I would need to update my transform function so that it takes an entire batch as input and outputs an entire batch. Also it wouldn't take labels as input nor output.",
    "759952": "mgornergoogle Thanks for the clarification Martin. Do you know what CPU, clock speed and core count gets used with TPUv3? (I'm curious to understand time trials from my experiments).",
    "760086": "`    However I would need to update my transform function so that it takes an entire batch as input and outputs an entire batch. Also it wouldn't take labels as input nor output.`\n\nCompletely true. I am also working on it, but quite slow :). You can go ahead. Thanks for answering.",
    "761696": "&gt; Do you know what CPU, clock speed and core count gets used with TPUv3?\n\nNot precisely but I know it's a lot. This has been dimensioned to make sure the data-hungry TPUv3 can always be fed with enough data.",
    "764814": "Are you considering upgrading the gpu too?  V100 with 32 GB are 1.5 years old, and all we have on Kaggle notebooks are P100.  Or is it a decision by Kaggle to not offer recent GPU and move people to tpu only for recent pretrained models?",
    "872963": "Quick pointers to other data augmentation techniques that work well on TPU/GPU:\n - [CutMix and MixUp on GPU/TPU](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935)\n - [GridMask data augmentation on GPU/TPU](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132986)\n - [batch implementation of rotation/shear/zoom/shift](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/133168)\n - [batch implementations of CutMix, MixUp, GridMask](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134264)\n - [perspective transformations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134557)",
    "953514": "This is really helpful for beginners. Thank you for sharing your work. May I please ask how to decide which augmentation transformations to begin with (specially if beginners are attempting data augmentation)?",
    "953558": "The idea for augmentation is two fold. First you want to generate more data that looks like your train data. So for example, it is appropriate to rotate images any angle in Melanoma skin lesion detection. But it wouldn't be appropriate to rotate handwritten digits more than plus minus 15 degrees. Since the more the better, just do every type of augmentation that creates realistic additional images.\n\nThe second reason for augmentation is to make learning more difficult for your model. That forces your model to learn the underlying idea more generally. This is accomplished with augmentation like coarse dropout or cutout which removes parts of the image. You should only obscure an image to the point that a human can still identify the image more or less. When applying augmentation for this reason, you will need to train for more epochs because your model learns more slowly. But it will become more intelligent at the end.",
    "953568": "Thank you. Very well explained. Keeping the images still realistic is perhaps the key while augmenting.",
    "953570": "Yes. That is why it is always good to display your images after applying augmentation. \n\nLook at dozens of the augmented images yourself. Adjust the hyperparameters (like how much rotation, zoom, color contrast, etc) to the maximum you can while keeping the images realistic looking. Then apply the most different types of augmentation that you can.\n\nDeep learning needs lots of training data and augmentation makes it for us. Augmentation is one of the best ways to increase accuracy for deep learning."
  },
  "source": "meta"
}