{
  "id": 133168,
  "title": "Make Chris Deotte's data augmentation faster",
  "url": "/competitions/flower-classification-with-tpus/discussion/133168",
  "author_name": "",
  "post_date": "2020-03-01T03:21:54.995420300Z",
  "votes": 15,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I published a kernel <a href=\"https://www.kaggle.com/yihdarshieh/make-chris-deotte-s-data-augmentation-faster?scriptVersionId=29453906\">Make Chris Deotte's data augmentation faster</a>.</p>\n\n<p>Instead of running data processing with tf.data.Dataset on CPU, I create a subclass of <code>tf.keras.layers.Layer</code> which is doing the image transformation in GPU / TPU when available.</p>\n\n<p>Timing for 1 iteration over the whole training dataset (when TPU is enabled):</p>\n\n<pre><code>- No data agumentation: ~ 13.7 sec.\n- Chris Deotte's data augmentation with tf.data.Dataset (so on CPU): ~ 25.9 sec.\n- Chris Deotte's data augmentation on TPU: ~ 16.1 sec.\n</code></pre>\n\n<p>When measuring the timing with GPU being enabled, the difference will be much larger.</p>\n\n<p>Chris may also work on updating his kernel after our discussion. So I don't try to make a training kernel with these layers. Of course, if my code can save his time or anyone else's time, it's good.\nThe main functions are <code>batch_transform()</code> and <code>get_batch_transformatioin_matrix()</code> which are just the batch version of Chris Deotte's code.</p>",
  "messages": [
    {
      "id": "760253",
      "postDate": "03/01/2020 03:21:54",
      "content": "<p>I published a kernel <a href=\"https://www.kaggle.com/yihdarshieh/make-chris-deotte-s-data-augmentation-faster?scriptVersionId=29453906\">Make Chris Deotte's data augmentation faster</a>.</p>\n\n<p>Instead of running data processing with tf.data.Dataset on CPU, I create a subclass of <code>tf.keras.layers.Layer</code> which is doing the image transformation in GPU / TPU when available.</p>\n\n<p>Timing for 1 iteration over the whole training dataset (when TPU is enabled):</p>\n\n<pre><code>- No data agumentation: ~ 13.7 sec.\n- Chris Deotte's data augmentation with tf.data.Dataset (so on CPU): ~ 25.9 sec.\n- Chris Deotte's data augmentation on TPU: ~ 16.1 sec.\n</code></pre>\n\n<p>When measuring the timing with GPU being enabled, the difference will be much larger.</p>\n\n<p>Chris may also work on updating his kernel after our discussion. So I don't try to make a training kernel with these layers. Of course, if my code can save his time or anyone else's time, it's good.\nThe main functions are <code>batch_transform()</code> and <code>get_batch_transformatioin_matrix()</code> which are just the batch version of Chris Deotte's code.</p>",
      "rawMarkdown": "I published a kernel [Make Chris Deotte's data augmentation faster](https://www.kaggle.com/yihdarshieh/make-chris-deotte-s-data-augmentation-faster?scriptVersionId=29453906).\n\nInstead of running data processing with tf.data.Dataset on CPU, I create a subclass of `tf.keras.layers.Layer` which is doing the image transformation in GPU / TPU when available.\n\nTiming for 1 iteration over the whole training dataset (when TPU is enabled):\n\n    - No data agumentation: ~ 13.7 sec.\n    - Chris Deotte's data augmentation with tf.data.Dataset (so on CPU): ~ 25.9 sec.\n    - Chris Deotte's data augmentation on TPU: ~ 16.1 sec.\n\nWhen measuring the timing with GPU being enabled, the difference will be much larger.\n\nChris may also work on updating his kernel after our discussion. So I don't try to make a training kernel with these layers. Of course, if my code can save his time or anyone else's time, it's good.\nThe main functions are ` batch_transform()` and `get_batch_transformatioin_matrix()` which are just the batch version of Chris Deotte's code.",
      "votes": null
    },
    {
      "id": "760278",
      "postDate": "03/01/2020 04:06:18",
      "content": "<p>Wow. Thanks Yih-Dar SHIEH. This is great. On Monday, I'll read everything carefully and comment.</p>",
      "rawMarkdown": "Wow. Thanks Yih-Dar SHIEH. This is great. On Monday, I'll read everything carefully and comment.",
      "votes": null
    },
    {
      "id": "760320",
      "postDate": "03/01/2020 05:50:30",
      "content": "<p>Nice! So does this means our original data augmentation methods still work on CPU even we already merged them into tf.data.dataset pipeline ?</p>",
      "rawMarkdown": "Nice! So does this means our original data augmentation methods still work on CPU even we already merged them into tf.data.dataset pipeline ?",
      "votes": null
    },
    {
      "id": "760377",
      "postDate": "03/01/2020 07:54:40",
      "content": "<p>Yes, anything inside tf.data.Dataset is run on CPU.</p>",
      "rawMarkdown": "Yes, anything inside tf.data.Dataset is run on CPU.",
      "votes": null
    },
    {
      "id": "760386",
      "postDate": "03/01/2020 08:07:05",
      "content": "<p>On Monday, I'll compute benchmark times comparing different augmentations. </p>\n\n<p>Using <code>tf.data.Dataset()</code> on CPU isn't necessarily slower than <code>tf.keras.layers</code> on GPU/TPU because <code>tf.data.Dataset()</code> runs in parallel while the GPU/TPU is training. When using <code>tf.keras.layers</code>, it is probably run sequentially (augmentation then train).</p>\n\n<p>For example without data augmentation, DenseNet201 takes 70 seconds to complete an epoch on TPU. With data augmentation using <code>tf.data.Dataset()</code> and my rotation function, it still takes 70 seconds to complete an epoch on TPU because it is done simultaneously with training. </p>",
      "rawMarkdown": "On Monday, I'll compute benchmark times comparing different augmentations. \n\nUsing `tf.data.Dataset()` on CPU isn't necessarily slower than `tf.keras.layers` on GPU/TPU because `tf.data.Dataset()` runs in parallel while the GPU/TPU is training. When using `tf.keras.layers`, it is probably run sequentially (augmentation then train).\n\nFor example without data augmentation, DenseNet201 takes 70 seconds to complete an epoch on TPU. With data augmentation using `tf.data.Dataset()` and my rotation function, it still takes 70 seconds to complete an epoch on TPU because it is done simultaneously with training.",
      "votes": null
    },
    {
      "id": "760409",
      "postDate": "03/01/2020 08:40:15",
      "content": "<p>That's true. Only when the computation is too heavy on CPU, even with parallelization, the GPU / TPU approach will actually help.</p>",
      "rawMarkdown": "That's true. Only when the computation is too heavy on CPU, even with parallelization, the GPU / TPU approach will actually help.",
      "votes": null
    },
    {
      "id": "760413",
      "postDate": "03/01/2020 08:48:26",
      "content": "<p>Thank you both for the explanation!</p>",
      "rawMarkdown": "Thank you both for the explanation!",
      "votes": null
    },
    {
      "id": "761737",
      "postDate": "03/02/2020 22:10:32",
      "content": "<p>Thank you for sharing. This is great.</p>",
      "rawMarkdown": "Thank you for sharing. This is great.",
      "votes": null
    },
    {
      "id": "762931",
      "postDate": "03/04/2020 00:10:05",
      "content": "<p>I benchmarked your augmentation routine. It's super fast! </p>\n\n<p>Here are the results on TPU. When you use my routine with <code>tf.data.Dataset</code> i.e. on CPU, one epoch of training and validation takes 70 seconds (batch size 128, image size 512x512x3, image count 12000). When you use your routine on CPU, it takes 70 seconds, and when you don't use data augmentation it takes 70 seconds. Therefore both your and my routine complete in under 70 seconds on CPU in parallel with training.</p>\n\n<p>Here's where it gets interesting. If you use <code>MIXED_PRECISION=True</code> on TPU and increase batch size to 256, then without augmentation, one epoch takes 50 seconds. If you use my routine on CPU, then one epoch takes 60 seconds. If you use your routine on CPU, then one epoch takes 50 seconds! So your routine is faster than 50 seconds on CPU in parallel with training while mine is not.</p>\n\n<p>When we use your routine on TPU as a <code>tf.keras.layers()</code>, then it occurs sequentially before each batch train. As such <code>MIXED_PRECISION=False</code> which normally takes 70 seconds becomes 90 seconds.</p>",
      "rawMarkdown": "I benchmarked your augmentation routine. It's super fast! \n\nHere are the results on TPU. When you use my routine with `tf.data.Dataset` i.e. on CPU, one epoch of training and validation takes 70 seconds (batch size 128, image size 512x512x3, image count 12000). When you use your routine on CPU, it takes 70 seconds, and when you don't use data augmentation it takes 70 seconds. Therefore both your and my routine complete in under 70 seconds on CPU in parallel with training.\n\nHere's where it gets interesting. If you use `MIXED_PRECISION=True` on TPU and increase batch size to 256, then without augmentation, one epoch takes 50 seconds. If you use my routine on CPU, then one epoch takes 60 seconds. If you use your routine on CPU, then one epoch takes 50 seconds! So your routine is faster than 50 seconds on CPU in parallel with training while mine is not.\n\nWhen we use your routine on TPU as a `tf.keras.layers()`, then it occurs sequentially before each batch train. As such `MIXED_PRECISION=False` which normally takes 70 seconds becomes 90 seconds.",
      "votes": null
    },
    {
      "id": "762957",
      "postDate": "03/04/2020 01:18:40",
      "content": "<p>Thanks for the detailed benchmark. So the batch processing speeds the computation on CPU when we use mixed precision.</p>",
      "rawMarkdown": "Thanks for the detailed benchmark. So the batch processing speeds the computation on CPU when we use mixed precision.",
      "votes": null
    },
    {
      "id": "762976",
      "postDate": "03/04/2020 01:47:59",
      "content": "<p>Yes. But I don't think the <code>MIXED_PRECISION</code> speeds up your augmentation nor mine. (I think <code>tf.data.Dataset</code> still uses 32 bit precision for both). </p>\n\n<p>The <code>MIXED_PRECISION</code> speeds up the TPU training to 50 seconds. And then 50 seconds becomes too fast for my augmentation to keep up but your augmentation can still keep up.</p>",
      "rawMarkdown": "Yes. But I don't think the `MIXED_PRECISION` speeds up your augmentation nor mine. (I think `tf.data.Dataset` still uses 32 bit precision for both). \n\nThe `MIXED_PRECISION` speeds up the TPU training to 50 seconds. And then 50 seconds becomes too fast for my augmentation to keep up but your augmentation can still keep up.",
      "votes": null
    },
    {
      "id": "763656",
      "postDate": "03/04/2020 17:57:17",
      "content": "<p>What is the different between <a href=\"/cdeotte\">@cdeotte</a> 's and <a href=\"/yihdarshieh\">@yihdarshieh</a> 's routines ? Is it that the latter works directly on batches ?</p>",
      "rawMarkdown": "What is the different between @cdeotte 's and @yihdarshieh 's routines ? Is it that the latter works directly on batches ?",
      "votes": null
    },
    {
      "id": "763670",
      "postDate": "03/04/2020 18:06:33",
      "content": "<p>Yes. Mine works on individual images and thus performs <code>batch_size</code> number of matrix mulitplies, and Yih-Dar SHIEH's works on the entire batch and thus performs <strong>one</strong> (large) matrix multiply. </p>\n\n<p>Another advantage of Yih-Dar SHIEH's is that since it operates on batches, you can directly embed it into a <code>tf.keras.layer()</code> in your model. (And that's what the above benchmarks are testing).</p>",
      "rawMarkdown": "Yes. Mine works on individual images and thus performs `batch_size` number of matrix mulitplies, and Yih-Dar SHIEH's works on the entire batch and thus performs **one** (large) matrix multiply. \n\nAnother advantage of Yih-Dar SHIEH's is that since it operates on batches, you can directly embed it into a `tf.keras.layer()` in your model. (And that's what the above benchmarks are testing).",
      "votes": null
    },
    {
      "id": "765077",
      "postDate": "03/06/2020 08:33:33",
      "content": "<p>Thanks again for sharing this. If you don't mind, I have a question would like to ask. I noticed that you use <code>reduce_sum</code> to get the <code>batch_size</code> information in augmentation layer. In my understanding, since the data will be distributed into several processors for calculation while training, so the <code>batch_size</code> in each processor might not be the same as we set in the beginning. Thus the <code>batch_size</code> we get from <code>reduce_sum</code> is the batch size in each processor. Is this understanding correct?</p>\n\n<p>Here is my question, if I already knew how many processor will be used while training, can I hard code the batch size for each processor? </p>\n\n<p>For example:\nI set the initial batch size to 128, and I have 8 processors for training. So in augmentation layer, I can hard code the batch_size to 128/8=16. Is this correct? </p>",
      "rawMarkdown": "Thanks again for sharing this. If you don't mind, I have a question would like to ask. I noticed that you use `reduce_sum` to get the `batch_size` information in augmentation layer. In my understanding, since the data will be distributed into several processors for calculation while training, so the `batch_size` in each processor might not be the same as we set in the beginning. Thus the `batch_size` we get from `reduce_sum` is the batch size in each processor. Is this understanding correct?\n\nHere is my question, if I already knew how many processor will be used while training, can I hard code the batch size for each processor? \n\nFor example:\nI set the initial batch size to 128, and I have 8 processors for training. So in augmentation layer, I can hard code the batch_size to 128/8=16. Is this correct?",
      "votes": null
    },
    {
      "id": "765083",
      "postDate": "03/06/2020 08:39:11",
      "content": "<p>BTW, I already tried it, and it works. But I am not sure whether this will cause some further problems or not.</p>",
      "rawMarkdown": "BTW, I already tried it, and it works. But I am not sure whether this will cause some further problems or not.",
      "votes": null
    },
    {
      "id": "765141",
      "postDate": "03/06/2020 09:41:46",
      "content": "<p>Yes, each tpu core gets a subset of the batch, and the batch size calculated is the number of examples in that subset. It won't cause any training issues, because the augmentation layer is only used for transforming the images.</p>\n\n<p>As Chris pointed, using augmentation layer is not a good approach, so avoid using it.</p>",
      "rawMarkdown": "Yes, each tpu core gets a subset of the batch, and the batch size calculated is the number of examples in that subset. It won't cause any training issues, because the augmentation layer is only used for transforming the images.\n\nAs Chris pointed, using augmentation layer is not a good approach, so avoid using it.",
      "votes": null
    },
    {
      "id": "765169",
      "postDate": "03/06/2020 10:13:02",
      "content": "<p>Thanks for the answer, but could you tell me where can I see Chris's comments about avoiding augmentation layer? Or could you please tell me the reason? Thanks in advances!</p>",
      "rawMarkdown": "Thanks for the answer, but could you tell me where can I see Chris's comments about avoiding augmentation layer? Or could you please tell me the reason? Thanks in advances!",
      "votes": null
    },
    {
      "id": "765204",
      "postDate": "03/06/2020 11:20:25",
      "content": "<p>It is in a comment this post. Search  the following text</p>\n\n<p>I benchmarked your augmentation routine. It's super fast</p>",
      "rawMarkdown": "It is in a comment this post. Search  the following text\n\nI benchmarked your augmentation routine. It's super fast",
      "votes": null
    },
    {
      "id": "765419",
      "postDate": "03/06/2020 15:52:59",
      "content": "<p>I ran the same benchmarks on Kaggle's GPU. The result is the same as TPU. Using <code>tf.data.Dataset</code> is better than <code>tf.keras.layers</code>.</p>\n\n<p>Using <code>tf.data.Dataset()</code> with either no augmentation, chris augmentation, or Yih-Dar SHIEH augmentation all have 190 second epochs (with 224x224x3 images batch size 16). Using <code>tf.keras.layers()</code> with Yih-Dar SHIEH augmentation takes 220 seconds.</p>\n\n<p>(MIXED_PRECISION doesn't speed up the P100 GPU, so i didn't do those benchmarks)</p>",
      "rawMarkdown": "I ran the same benchmarks on Kaggle's GPU. The result is the same as TPU. Using `tf.data.Dataset` is better than `tf.keras.layers`.\n\nUsing `tf.data.Dataset()` with either no augmentation, chris augmentation, or Yih-Dar SHIEH augmentation all have 190 second epochs (with 224x224x3 images batch size 16). Using `tf.keras.layers()` with Yih-Dar SHIEH augmentation takes 220 seconds.\n\n(MIXED_PRECISION doesn't speed up the P100 GPU, so i didn't do those benchmarks)",
      "votes": null
    },
    {
      "id": "765670",
      "postDate": "03/07/2020 00:08:09",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks for the informative benchmarks!</p>",
      "rawMarkdown": "cdeotte Thanks for the informative benchmarks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 760278,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/01/2020 04:06:18",
      "content": "<p>Wow. Thanks Yih-Dar SHIEH. This is great. On Monday, I'll read everything carefully and comment.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 760320,
      "author_name": "xiejialun",
      "author_url": "",
      "post_date": "03/01/2020 05:50:30",
      "content": "<p>Nice! So does this means our original data augmentation methods still work on CPU even we already merged them into tf.data.dataset pipeline ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 760377,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/01/2020 07:54:40",
          "content": "<p>Yes, anything inside tf.data.Dataset is run on CPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760386,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/01/2020 08:07:05",
          "content": "<p>On Monday, I'll compute benchmark times comparing different augmentations. </p>\n\n<p>Using <code>tf.data.Dataset()</code> on CPU isn't necessarily slower than <code>tf.keras.layers</code> on GPU/TPU because <code>tf.data.Dataset()</code> runs in parallel while the GPU/TPU is training. When using <code>tf.keras.layers</code>, it is probably run sequentially (augmentation then train).</p>\n\n<p>For example without data augmentation, DenseNet201 takes 70 seconds to complete an epoch on TPU. With data augmentation using <code>tf.data.Dataset()</code> and my rotation function, it still takes 70 seconds to complete an epoch on TPU because it is done simultaneously with training. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760409,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/01/2020 08:40:15",
          "content": "<p>That's true. Only when the computation is too heavy on CPU, even with parallelization, the GPU / TPU approach will actually help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760413,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/01/2020 08:48:26",
          "content": "<p>Thank you both for the explanation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761737,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/02/2020 22:10:32",
      "content": "<p>Thank you for sharing. This is great.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 762931,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/04/2020 00:10:05",
      "content": "<p>I benchmarked your augmentation routine. It's super fast! </p>\n\n<p>Here are the results on TPU. When you use my routine with <code>tf.data.Dataset</code> i.e. on CPU, one epoch of training and validation takes 70 seconds (batch size 128, image size 512x512x3, image count 12000). When you use your routine on CPU, it takes 70 seconds, and when you don't use data augmentation it takes 70 seconds. Therefore both your and my routine complete in under 70 seconds on CPU in parallel with training.</p>\n\n<p>Here's where it gets interesting. If you use <code>MIXED_PRECISION=True</code> on TPU and increase batch size to 256, then without augmentation, one epoch takes 50 seconds. If you use my routine on CPU, then one epoch takes 60 seconds. If you use your routine on CPU, then one epoch takes 50 seconds! So your routine is faster than 50 seconds on CPU in parallel with training while mine is not.</p>\n\n<p>When we use your routine on TPU as a <code>tf.keras.layers()</code>, then it occurs sequentially before each batch train. As such <code>MIXED_PRECISION=False</code> which normally takes 70 seconds becomes 90 seconds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 762957,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/04/2020 01:18:40",
          "content": "<p>Thanks for the detailed benchmark. So the batch processing speeds the computation on CPU when we use mixed precision.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 762976,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/04/2020 01:47:59",
          "content": "<p>Yes. But I don't think the <code>MIXED_PRECISION</code> speeds up your augmentation nor mine. (I think <code>tf.data.Dataset</code> still uses 32 bit precision for both). </p>\n\n<p>The <code>MIXED_PRECISION</code> speeds up the TPU training to 50 seconds. And then 50 seconds becomes too fast for my augmentation to keep up but your augmentation can still keep up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763656,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/04/2020 17:57:17",
          "content": "<p>What is the different between <a href=\"/cdeotte\">@cdeotte</a> 's and <a href=\"/yihdarshieh\">@yihdarshieh</a> 's routines ? Is it that the latter works directly on batches ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763670,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/04/2020 18:06:33",
          "content": "<p>Yes. Mine works on individual images and thus performs <code>batch_size</code> number of matrix mulitplies, and Yih-Dar SHIEH's works on the entire batch and thus performs <strong>one</strong> (large) matrix multiply. </p>\n\n<p>Another advantage of Yih-Dar SHIEH's is that since it operates on batches, you can directly embed it into a <code>tf.keras.layer()</code> in your model. (And that's what the above benchmarks are testing).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765419,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/06/2020 15:52:59",
          "content": "<p>I ran the same benchmarks on Kaggle's GPU. The result is the same as TPU. Using <code>tf.data.Dataset</code> is better than <code>tf.keras.layers</code>.</p>\n\n<p>Using <code>tf.data.Dataset()</code> with either no augmentation, chris augmentation, or Yih-Dar SHIEH augmentation all have 190 second epochs (with 224x224x3 images batch size 16). Using <code>tf.keras.layers()</code> with Yih-Dar SHIEH augmentation takes 220 seconds.</p>\n\n<p>(MIXED_PRECISION doesn't speed up the P100 GPU, so i didn't do those benchmarks)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765670,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/07/2020 00:08:09",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks for the informative benchmarks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 765077,
      "author_name": "xiejialun",
      "author_url": "",
      "post_date": "03/06/2020 08:33:33",
      "content": "<p>Thanks again for sharing this. If you don't mind, I have a question would like to ask. I noticed that you use <code>reduce_sum</code> to get the <code>batch_size</code> information in augmentation layer. In my understanding, since the data will be distributed into several processors for calculation while training, so the <code>batch_size</code> in each processor might not be the same as we set in the beginning. Thus the <code>batch_size</code> we get from <code>reduce_sum</code> is the batch size in each processor. Is this understanding correct?</p>\n\n<p>Here is my question, if I already knew how many processor will be used while training, can I hard code the batch size for each processor? </p>\n\n<p>For example:\nI set the initial batch size to 128, and I have 8 processors for training. So in augmentation layer, I can hard code the batch_size to 128/8=16. Is this correct? </p>",
      "votes": null,
      "replies": [
        {
          "id": 765083,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/06/2020 08:39:11",
          "content": "<p>BTW, I already tried it, and it works. But I am not sure whether this will cause some further problems or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765141,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/06/2020 09:41:46",
          "content": "<p>Yes, each tpu core gets a subset of the batch, and the batch size calculated is the number of examples in that subset. It won't cause any training issues, because the augmentation layer is only used for transforming the images.</p>\n\n<p>As Chris pointed, using augmentation layer is not a good approach, so avoid using it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765169,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/06/2020 10:13:02",
          "content": "<p>Thanks for the answer, but could you tell me where can I see Chris's comments about avoiding augmentation layer? Or could you please tell me the reason? Thanks in advances!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765204,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "03/06/2020 11:20:25",
          "content": "<p>It is in a comment this post. Search  the following text</p>\n\n<p>I benchmarked your augmentation routine. It's super fast</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "760253": "I published a kernel [Make Chris Deotte's data augmentation faster](https://www.kaggle.com/yihdarshieh/make-chris-deotte-s-data-augmentation-faster?scriptVersionId=29453906).\n\nInstead of running data processing with tf.data.Dataset on CPU, I create a subclass of `tf.keras.layers.Layer` which is doing the image transformation in GPU / TPU when available.\n\nTiming for 1 iteration over the whole training dataset (when TPU is enabled):\n\n    - No data agumentation: ~ 13.7 sec.\n    - Chris Deotte's data augmentation with tf.data.Dataset (so on CPU): ~ 25.9 sec.\n    - Chris Deotte's data augmentation on TPU: ~ 16.1 sec.\n\nWhen measuring the timing with GPU being enabled, the difference will be much larger.\n\nChris may also work on updating his kernel after our discussion. So I don't try to make a training kernel with these layers. Of course, if my code can save his time or anyone else's time, it's good.\nThe main functions are ` batch_transform()` and `get_batch_transformatioin_matrix()` which are just the batch version of Chris Deotte's code.",
    "760278": "Wow. Thanks Yih-Dar SHIEH. This is great. On Monday, I'll read everything carefully and comment.",
    "760320": "Nice! So does this means our original data augmentation methods still work on CPU even we already merged them into tf.data.dataset pipeline ?",
    "760377": "Yes, anything inside tf.data.Dataset is run on CPU.",
    "760386": "On Monday, I'll compute benchmark times comparing different augmentations. \n\nUsing `tf.data.Dataset()` on CPU isn't necessarily slower than `tf.keras.layers` on GPU/TPU because `tf.data.Dataset()` runs in parallel while the GPU/TPU is training. When using `tf.keras.layers`, it is probably run sequentially (augmentation then train).\n\nFor example without data augmentation, DenseNet201 takes 70 seconds to complete an epoch on TPU. With data augmentation using `tf.data.Dataset()` and my rotation function, it still takes 70 seconds to complete an epoch on TPU because it is done simultaneously with training.",
    "760409": "That's true. Only when the computation is too heavy on CPU, even with parallelization, the GPU / TPU approach will actually help.",
    "760413": "Thank you both for the explanation!",
    "761737": "Thank you for sharing. This is great.",
    "762931": "I benchmarked your augmentation routine. It's super fast! \n\nHere are the results on TPU. When you use my routine with `tf.data.Dataset` i.e. on CPU, one epoch of training and validation takes 70 seconds (batch size 128, image size 512x512x3, image count 12000). When you use your routine on CPU, it takes 70 seconds, and when you don't use data augmentation it takes 70 seconds. Therefore both your and my routine complete in under 70 seconds on CPU in parallel with training.\n\nHere's where it gets interesting. If you use `MIXED_PRECISION=True` on TPU and increase batch size to 256, then without augmentation, one epoch takes 50 seconds. If you use my routine on CPU, then one epoch takes 60 seconds. If you use your routine on CPU, then one epoch takes 50 seconds! So your routine is faster than 50 seconds on CPU in parallel with training while mine is not.\n\nWhen we use your routine on TPU as a `tf.keras.layers()`, then it occurs sequentially before each batch train. As such `MIXED_PRECISION=False` which normally takes 70 seconds becomes 90 seconds.",
    "762957": "Thanks for the detailed benchmark. So the batch processing speeds the computation on CPU when we use mixed precision.",
    "762976": "Yes. But I don't think the `MIXED_PRECISION` speeds up your augmentation nor mine. (I think `tf.data.Dataset` still uses 32 bit precision for both). \n\nThe `MIXED_PRECISION` speeds up the TPU training to 50 seconds. And then 50 seconds becomes too fast for my augmentation to keep up but your augmentation can still keep up.",
    "763656": "What is the different between @cdeotte 's and @yihdarshieh 's routines ? Is it that the latter works directly on batches ?",
    "763670": "Yes. Mine works on individual images and thus performs `batch_size` number of matrix mulitplies, and Yih-Dar SHIEH's works on the entire batch and thus performs **one** (large) matrix multiply. \n\nAnother advantage of Yih-Dar SHIEH's is that since it operates on batches, you can directly embed it into a `tf.keras.layer()` in your model. (And that's what the above benchmarks are testing).",
    "765077": "Thanks again for sharing this. If you don't mind, I have a question would like to ask. I noticed that you use `reduce_sum` to get the `batch_size` information in augmentation layer. In my understanding, since the data will be distributed into several processors for calculation while training, so the `batch_size` in each processor might not be the same as we set in the beginning. Thus the `batch_size` we get from `reduce_sum` is the batch size in each processor. Is this understanding correct?\n\nHere is my question, if I already knew how many processor will be used while training, can I hard code the batch size for each processor? \n\nFor example:\nI set the initial batch size to 128, and I have 8 processors for training. So in augmentation layer, I can hard code the batch_size to 128/8=16. Is this correct?",
    "765083": "BTW, I already tried it, and it works. But I am not sure whether this will cause some further problems or not.",
    "765141": "Yes, each tpu core gets a subset of the batch, and the batch size calculated is the number of examples in that subset. It won't cause any training issues, because the augmentation layer is only used for transforming the images.\n\nAs Chris pointed, using augmentation layer is not a good approach, so avoid using it.",
    "765169": "Thanks for the answer, but could you tell me where can I see Chris's comments about avoiding augmentation layer? Or could you please tell me the reason? Thanks in advances!",
    "765204": "It is in a comment this post. Search  the following text\n\nI benchmarked your augmentation routine. It's super fast",
    "765419": "I ran the same benchmarks on Kaggle's GPU. The result is the same as TPU. Using `tf.data.Dataset` is better than `tf.keras.layers`.\n\nUsing `tf.data.Dataset()` with either no augmentation, chris augmentation, or Yih-Dar SHIEH augmentation all have 190 second epochs (with 224x224x3 images batch size 16). Using `tf.keras.layers()` with Yih-Dar SHIEH augmentation takes 220 seconds.\n\n(MIXED_PRECISION doesn't speed up the P100 GPU, so i didn't do those benchmarks)",
    "765670": "cdeotte Thanks for the informative benchmarks!"
  },
  "source": "meta"
}