{
  "id": 201977,
  "title": "Ninja Techniques To Make Most Out of Google Cloud TPUs !",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201977",
  "author_name": "",
  "post_date": "2020-12-07T16:31:24.275978700Z",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As in this competition we are given liberty to explore TPUs I searched for ways to utilize TPUs effectively and found the following ways ! Please let me know in the comments that whether using some of this or all of this techniques provide some major performance improvements in your model which can be either reduction in Training Time or Better Accuracy or Both 😊😊. Also feel free to add some of the techniques which you might have discovered as it will help to make most out of TPUs 😜 !</p>\n<p><strong>A. Speed Up Your Augmentations</strong></p>\n<ol>\n<li><p>Augmenting Data particularly with Image Data becomes critical when training models on TPU as CPUs can become a bottleneck which are having very lower speeds than TPU . To avoid doing this preprocess your data before hand don't do preprocessing or augmentations while you are Training your models on TPU , otherwise your CPU will become a bottleneck and entire purpose of using TPUs will be lost.</p></li>\n<li><p>A much novel solution to this bottleneck problem is discussed by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\" target=\"_blank\">here </a> what he has done is he had written augmentations script in Tensorflow itself doing so tells TPU that augmentations are to be done on TPU and not CPU , this speeds up augmentation process as a whole and now there is  no problem of bottleneck.</p></li>\n</ol>\n<p>The image below as an attachment <code>tpuvscpu.JPG</code> from Chris Kernel's captures essence of this entire Problem.</p>\n<p><strong>B. Choose Your Batch Size Wisely to Speed Up TPU</strong></p>\n<ol>\n<li><p>TPU uses something called as <strong>sharding</strong> , it is nothing but it distributes the jobs over all its 8 cores and parallelize whole process of model training !</p></li>\n<li><p>In most of the public kernels I have seen this code segment is use <code>BATCH_SIZE = 16 * strategy.num_replicas_in_sync</code> , now this is correct but what it evaluates to is  , **128 ** but one should know that if you use batch size = n  , then in TPU the effective batch size is calculated as n/8 (Since we have 8 cores) , for our purpose effective batch size becomes 16 across each cores , which means we are underutilizing CPU terribly 😂.</p></li>\n<li><p>To utilize TPU at its peak we should try out Batch Sizes which are bigger than 128 untll we come across a batch_size which doesn't fits in TPU !</p></li>\n<li><p>As TPU core uses 128 * 128 memory cells for matrix processing it is better to use largest batch size which is evenly divisible by 128 ! We can try using 256 , 512 , 1024 and 2048 till we find big enough batch size which doesnt fits in the TPU memory ! </p></li>\n</ol>\n<p><strong>C. Speeding Up Training Using Mixed Precision and XLA</strong></p>\n<ol>\n<li><p>Using mixed precision and/or XLA allows TPU to handle larger batch sizes than it can fit on TPU ! This can be tremendously useful while training on bigger batch sizes .</p></li>\n<li><p>An example code of how to do this is given in Chris Deotte's Golden Notebook which can be find <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\" target=\"_blank\">here</a> , refer title \"Mixed Precision and/or XLA\" , to directly get the code :) .</p></li>\n</ol>\n<p><strong>D. Do Padding and Pooling Effectively</strong></p>\n<ol>\n<li>TPU uses something called tiled memory scheme , where efficient memory utilization is dependent on amount of memory wasted due to padding ,  to utilize memory effectively one must be careful using batch_size and padding dimensions since this are determining factors which decides tiling.</li>\n</ol>\n<p>2.Prefer <code>tf.nn.avg_pool</code> over <code>tf.nn.max_pool</code> , this might be counter intuitive in contrast to other image classifications task we have seen , what one needs to know is that in TPU the gradient calculations for max pool operation is slower than that of average pool operations .</p>\n<p><strong>E. Use Fusion To Ace Training !</strong></p>\n<ol>\n<li><p>TPU compilers use something called as Fusion internally . Benefits of Fusion includes reducing main memory accesses , increasing utilization of hardware devices , reducing memory requirements of the model .</p></li>\n<li><p>One way we can use Fusion is use <code>tf.nn.fused_batch_norm</code> over <code>tf.nn.batch_normalization</code> , as clear from the name itself it is fused variant of batch normalization ! **Note : For tf.layers.batch_normalization, set the \"fused\" argument to true. **</p></li>\n</ol>\n<p>That was all thanks to <a href=\"https://cloud.google.com/tpu/docs/performance-guide#:~:text=For%20optimum%20memory%20usage%2C%20use,effectively%20use%20the%20TPU%20memory.\" target=\"_blank\">this </a>documentation for letting me  know some tips to utilize TPU well , feel free to add more :) .</p>",
  "messages": [
    {
      "id": "1105225",
      "postDate": "12/07/2020 16:31:24",
      "content": "<p>As in this competition we are given liberty to explore TPUs I searched for ways to utilize TPUs effectively and found the following ways ! Please let me know in the comments that whether using some of this or all of this techniques provide some major performance improvements in your model which can be either reduction in Training Time or Better Accuracy or Both 😊😊. Also feel free to add some of the techniques which you might have discovered as it will help to make most out of TPUs 😜 !</p>\n<p><strong>A. Speed Up Your Augmentations</strong></p>\n<ol>\n<li><p>Augmenting Data particularly with Image Data becomes critical when training models on TPU as CPUs can become a bottleneck which are having very lower speeds than TPU . To avoid doing this preprocess your data before hand don't do preprocessing or augmentations while you are Training your models on TPU , otherwise your CPU will become a bottleneck and entire purpose of using TPUs will be lost.</p></li>\n<li><p>A much novel solution to this bottleneck problem is discussed by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\" target=\"_blank\">here </a> what he has done is he had written augmentations script in Tensorflow itself doing so tells TPU that augmentations are to be done on TPU and not CPU , this speeds up augmentation process as a whole and now there is  no problem of bottleneck.</p></li>\n</ol>\n<p>The image below as an attachment <code>tpuvscpu.JPG</code> from Chris Kernel's captures essence of this entire Problem.</p>\n<p><strong>B. Choose Your Batch Size Wisely to Speed Up TPU</strong></p>\n<ol>\n<li><p>TPU uses something called as <strong>sharding</strong> , it is nothing but it distributes the jobs over all its 8 cores and parallelize whole process of model training !</p></li>\n<li><p>In most of the public kernels I have seen this code segment is use <code>BATCH_SIZE = 16 * strategy.num_replicas_in_sync</code> , now this is correct but what it evaluates to is  , **128 ** but one should know that if you use batch size = n  , then in TPU the effective batch size is calculated as n/8 (Since we have 8 cores) , for our purpose effective batch size becomes 16 across each cores , which means we are underutilizing CPU terribly 😂.</p></li>\n<li><p>To utilize TPU at its peak we should try out Batch Sizes which are bigger than 128 untll we come across a batch_size which doesn't fits in TPU !</p></li>\n<li><p>As TPU core uses 128 * 128 memory cells for matrix processing it is better to use largest batch size which is evenly divisible by 128 ! We can try using 256 , 512 , 1024 and 2048 till we find big enough batch size which doesnt fits in the TPU memory ! </p></li>\n</ol>\n<p><strong>C. Speeding Up Training Using Mixed Precision and XLA</strong></p>\n<ol>\n<li><p>Using mixed precision and/or XLA allows TPU to handle larger batch sizes than it can fit on TPU ! This can be tremendously useful while training on bigger batch sizes .</p></li>\n<li><p>An example code of how to do this is given in Chris Deotte's Golden Notebook which can be find <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\" target=\"_blank\">here</a> , refer title \"Mixed Precision and/or XLA\" , to directly get the code :) .</p></li>\n</ol>\n<p><strong>D. Do Padding and Pooling Effectively</strong></p>\n<ol>\n<li>TPU uses something called tiled memory scheme , where efficient memory utilization is dependent on amount of memory wasted due to padding ,  to utilize memory effectively one must be careful using batch_size and padding dimensions since this are determining factors which decides tiling.</li>\n</ol>\n<p>2.Prefer <code>tf.nn.avg_pool</code> over <code>tf.nn.max_pool</code> , this might be counter intuitive in contrast to other image classifications task we have seen , what one needs to know is that in TPU the gradient calculations for max pool operation is slower than that of average pool operations .</p>\n<p><strong>E. Use Fusion To Ace Training !</strong></p>\n<ol>\n<li><p>TPU compilers use something called as Fusion internally . Benefits of Fusion includes reducing main memory accesses , increasing utilization of hardware devices , reducing memory requirements of the model .</p></li>\n<li><p>One way we can use Fusion is use <code>tf.nn.fused_batch_norm</code> over <code>tf.nn.batch_normalization</code> , as clear from the name itself it is fused variant of batch normalization ! **Note : For tf.layers.batch_normalization, set the \"fused\" argument to true. **</p></li>\n</ol>\n<p>That was all thanks to <a href=\"https://cloud.google.com/tpu/docs/performance-guide#:~:text=For%20optimum%20memory%20usage%2C%20use,effectively%20use%20the%20TPU%20memory.\" target=\"_blank\">this </a>documentation for letting me  know some tips to utilize TPU well , feel free to add more :) .</p>",
      "rawMarkdown": "As in this competition we are given liberty to explore TPUs I searched for ways to utilize TPUs effectively and found the following ways ! Please let me know in the comments that whether using some of this or all of this techniques provide some major performance improvements in your model which can be either reduction in Training Time or Better Accuracy or Both 😊😊. Also feel free to add some of the techniques which you might have discovered as it will help to make most out of TPUs 😜 !\n\n**A. Speed Up Your Augmentations**\n1. Augmenting Data particularly with Image Data becomes critical when training models on TPU as CPUs can become a bottleneck which are having very lower speeds than TPU . To avoid doing this preprocess your data before hand don't do preprocessing or augmentations while you are Training your models on TPU , otherwise your CPU will become a bottleneck and entire purpose of using TPUs will be lost.\n\n2. A much novel solution to this bottleneck problem is discussed by @cdeotte [here ](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96) what he has done is he had written augmentations script in Tensorflow itself doing so tells TPU that augmentations are to be done on TPU and not CPU , this speeds up augmentation process as a whole and now there is  no problem of bottleneck.\n\nThe image below as an attachment `tpuvscpu.JPG` from Chris Kernel's captures essence of this entire Problem.\n\n**B. Choose Your Batch Size Wisely to Speed Up TPU**\n1. TPU uses something called as **sharding** , it is nothing but it distributes the jobs over all its 8 cores and parallelize whole process of model training !\n\n2. In most of the public kernels I have seen this code segment is use `BATCH_SIZE = 16 * strategy.num_replicas_in_sync` , now this is correct but what it evaluates to is  , **128 ** but one should know that if you use batch size = n  , then in TPU the effective batch size is calculated as n/8 (Since we have 8 cores) , for our purpose effective batch size becomes 16 across each cores , which means we are underutilizing CPU terribly 😂.\n\n3. To utilize TPU at its peak we should try out Batch Sizes which are bigger than 128 untll we come across a batch_size which doesn't fits in TPU !\n\n4. As TPU core uses 128 * 128 memory cells for matrix processing it is better to use largest batch size which is evenly divisible by 128 ! We can try using 256 , 512 , 1024 and 2048 till we find big enough batch size which doesnt fits in the TPU memory ! \n\n**C. Speeding Up Training Using Mixed Precision and XLA**\n\n1. Using mixed precision and/or XLA allows TPU to handle larger batch sizes than it can fit on TPU ! This can be tremendously useful while training on bigger batch sizes .\n\n2. An example code of how to do this is given in Chris Deotte's Golden Notebook which can be find [here](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96) , refer title \"Mixed Precision and/or XLA\" , to directly get the code :) .\n\n**D. Do Padding and Pooling Effectively**\n\n1. TPU uses something called tiled memory scheme , where efficient memory utilization is dependent on amount of memory wasted due to padding ,  to utilize memory effectively one must be careful using batch_size and padding dimensions since this are determining factors which decides tiling.\n\n2.Prefer `tf.nn.avg_pool` over ` tf.nn.max_pool` , this might be counter intuitive in contrast to other image classifications task we have seen , what one needs to know is that in TPU the gradient calculations for max pool operation is slower than that of average pool operations .\n\n**E. Use Fusion To Ace Training !**\n\n1. TPU compilers use something called as Fusion internally . Benefits of Fusion includes reducing main memory accesses , increasing utilization of hardware devices , reducing memory requirements of the model .\n\n2. One way we can use Fusion is use `tf.nn.fused_batch_norm` over `tf.nn.batch_normalization` , as clear from the name itself it is fused variant of batch normalization ! **Note : For tf.layers.batch_normalization, set the \"fused\" argument to true. **\n\nThat was all thanks to [this ](https://cloud.google.com/tpu/docs/performance-guide#:~:text=For%20optimum%20memory%20usage%2C%20use,effectively%20use%20the%20TPU%20memory.)documentation for letting me  know some tips to utilize TPU well , feel free to add more :) .",
      "votes": null
    },
    {
      "id": "1105458",
      "postDate": "12/07/2020 22:25:26",
      "content": "<p>It may be a better idea to produce the augmented images beforehand and have them as tfrecords with the fold numbers. The fold numbers are critical because we don't want to use the augmented images as the validation data in the cross validation process.</p>",
      "rawMarkdown": "It may be a better idea to produce the augmented images beforehand and have them as tfrecords with the fold numbers. The fold numbers are critical because we don't want to use the augmented images as the validation data in the cross validation process.",
      "votes": null
    },
    {
      "id": "1105655",
      "postDate": "12/08/2020 04:03:59",
      "content": "<p><a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> ,yes it is better idea to prepare augmented images before training as a part of pre processing pipeline ! The Google Cloud TPU documentation also suggest to do preprocessing before training ! But writing augmentation is Tensorflow library and doing it on GPU is something new one should learn and try out !</p>",
      "rawMarkdown": "tolgadincer ,yes it is better idea to prepare augmented images before training as a part of pre processing pipeline ! The Google Cloud TPU documentation also suggest to do preprocessing before training ! But writing augmentation is Tensorflow library and doing it on GPU is something new one should learn and try out !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1105458,
      "author_name": "tolgadincer",
      "author_url": "",
      "post_date": "12/07/2020 22:25:26",
      "content": "<p>It may be a better idea to produce the augmented images beforehand and have them as tfrecords with the fold numbers. The fold numbers are critical because we don't want to use the augmented images as the validation data in the cross validation process.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1105655,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "12/08/2020 04:03:59",
          "content": "<p><a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> ,yes it is better idea to prepare augmented images before training as a part of pre processing pipeline ! The Google Cloud TPU documentation also suggest to do preprocessing before training ! But writing augmentation is Tensorflow library and doing it on GPU is something new one should learn and try out !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1105225": "As in this competition we are given liberty to explore TPUs I searched for ways to utilize TPUs effectively and found the following ways ! Please let me know in the comments that whether using some of this or all of this techniques provide some major performance improvements in your model which can be either reduction in Training Time or Better Accuracy or Both 😊😊. Also feel free to add some of the techniques which you might have discovered as it will help to make most out of TPUs 😜 !\n\n**A. Speed Up Your Augmentations**\n1. Augmenting Data particularly with Image Data becomes critical when training models on TPU as CPUs can become a bottleneck which are having very lower speeds than TPU . To avoid doing this preprocess your data before hand don't do preprocessing or augmentations while you are Training your models on TPU , otherwise your CPU will become a bottleneck and entire purpose of using TPUs will be lost.\n\n2. A much novel solution to this bottleneck problem is discussed by @cdeotte [here ](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96) what he has done is he had written augmentations script in Tensorflow itself doing so tells TPU that augmentations are to be done on TPU and not CPU , this speeds up augmentation process as a whole and now there is  no problem of bottleneck.\n\nThe image below as an attachment `tpuvscpu.JPG` from Chris Kernel's captures essence of this entire Problem.\n\n**B. Choose Your Batch Size Wisely to Speed Up TPU**\n1. TPU uses something called as **sharding** , it is nothing but it distributes the jobs over all its 8 cores and parallelize whole process of model training !\n\n2. In most of the public kernels I have seen this code segment is use `BATCH_SIZE = 16 * strategy.num_replicas_in_sync` , now this is correct but what it evaluates to is  , **128 ** but one should know that if you use batch size = n  , then in TPU the effective batch size is calculated as n/8 (Since we have 8 cores) , for our purpose effective batch size becomes 16 across each cores , which means we are underutilizing CPU terribly 😂.\n\n3. To utilize TPU at its peak we should try out Batch Sizes which are bigger than 128 untll we come across a batch_size which doesn't fits in TPU !\n\n4. As TPU core uses 128 * 128 memory cells for matrix processing it is better to use largest batch size which is evenly divisible by 128 ! We can try using 256 , 512 , 1024 and 2048 till we find big enough batch size which doesnt fits in the TPU memory ! \n\n**C. Speeding Up Training Using Mixed Precision and XLA**\n\n1. Using mixed precision and/or XLA allows TPU to handle larger batch sizes than it can fit on TPU ! This can be tremendously useful while training on bigger batch sizes .\n\n2. An example code of how to do this is given in Chris Deotte's Golden Notebook which can be find [here](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96) , refer title \"Mixed Precision and/or XLA\" , to directly get the code :) .\n\n**D. Do Padding and Pooling Effectively**\n\n1. TPU uses something called tiled memory scheme , where efficient memory utilization is dependent on amount of memory wasted due to padding ,  to utilize memory effectively one must be careful using batch_size and padding dimensions since this are determining factors which decides tiling.\n\n2.Prefer `tf.nn.avg_pool` over ` tf.nn.max_pool` , this might be counter intuitive in contrast to other image classifications task we have seen , what one needs to know is that in TPU the gradient calculations for max pool operation is slower than that of average pool operations .\n\n**E. Use Fusion To Ace Training !**\n\n1. TPU compilers use something called as Fusion internally . Benefits of Fusion includes reducing main memory accesses , increasing utilization of hardware devices , reducing memory requirements of the model .\n\n2. One way we can use Fusion is use `tf.nn.fused_batch_norm` over `tf.nn.batch_normalization` , as clear from the name itself it is fused variant of batch normalization ! **Note : For tf.layers.batch_normalization, set the \"fused\" argument to true. **\n\nThat was all thanks to [this ](https://cloud.google.com/tpu/docs/performance-guide#:~:text=For%20optimum%20memory%20usage%2C%20use,effectively%20use%20the%20TPU%20memory.)documentation for letting me  know some tips to utilize TPU well , feel free to add more :) .",
    "1105458": "It may be a better idea to produce the augmented images beforehand and have them as tfrecords with the fold numbers. The fold numbers are critical because we don't want to use the augmented images as the validation data in the cross validation process.",
    "1105655": "tolgadincer ,yes it is better idea to prepare augmented images before training as a part of pre processing pipeline ! The Google Cloud TPU documentation also suggest to do preprocessing before training ! But writing augmentation is Tensorflow library and doing it on GPU is something new one should learn and try out !"
  },
  "source": "meta"
}