{
  "id": 132127,
  "title": "Augmentation question",
  "url": "/competitions/flower-classification-with-tpus/discussion/132127",
  "author_name": "",
  "post_date": "2020-02-24T11:40:43.519795600Z",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hallo,\nI am rather new to tf and I have a question about augmentation.\nWhen I have a function like</p>\n\n<p>def get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset</p>\n\n<p>...which is used in the model.fit() method, is the dataset calculated new for each epoch and does it make a difference if I call the map() function after repeat.\nI would like to get new augmentations for every epoch, otherwise the variance of the training data would be very poor.</p>\n\n<p>regards</p>",
  "messages": [
    {
      "id": "755053",
      "postDate": "02/24/2020 11:40:43",
      "content": "<p>Hallo,\nI am rather new to tf and I have a question about augmentation.\nWhen I have a function like</p>\n\n<p>def get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset</p>\n\n<p>...which is used in the model.fit() method, is the dataset calculated new for each epoch and does it make a difference if I call the map() function after repeat.\nI would like to get new augmentations for every epoch, otherwise the variance of the training data would be very poor.</p>\n\n<p>regards</p>",
      "rawMarkdown": "Hallo,\nI am rather new to tf and I have a question about augmentation.\nWhen I have a function like\n\ndef get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n\n...which is used in the model.fit() method, is the dataset calculated new for each epoch and does it make a difference if I call the map() function after repeat.\nI would like to get new augmentations for every epoch, otherwise the variance of the training data would be very poor.\n\nregards",
      "votes": null
    },
    {
      "id": "755369",
      "postDate": "02/24/2020 17:57:36",
      "content": "<p>As much as I know, the data is augmented and feed to the model.fit() every epoch and so is will keep changing every epoch. The get_training_data() function returns the new random variant of augmented data for each epoch.</p>",
      "rawMarkdown": "As much as I know, the data is augmented and feed to the model.fit() every epoch and so is will keep changing every epoch. The get_training_data() function returns the new random variant of augmented data for each epoch.",
      "votes": null
    },
    {
      "id": "755421",
      "postDate": "02/24/2020 19:02:07",
      "content": "<p>This is a great question. The best way to answer this it to make a small dataset of 10 images, and display what is happening. Because the TensorFlow documentation may say one thing but what is actually happening may be different.</p>\n\n<p>I think you may be right, that <code>dataset = dataset.repeat()</code> may need to be before <code>dataset = dataset.map(dataaugment)</code>. Because the repeat command turns the dataset from size 12,000 to size infinite. We want to do that before augmentation. As opposed to taking and repeating 12,000 augmented images.</p>\n\n<p>EDIT: The order doesn't matter. I posted a test above.</p>",
      "rawMarkdown": "This is a great question. The best way to answer this it to make a small dataset of 10 images, and display what is happening. Because the TensorFlow documentation may say one thing but what is actually happening may be different.\n\nI think you may be right, that `dataset = dataset.repeat()` may need to be before `dataset = dataset.map(dataaugment)`. Because the repeat command turns the dataset from size 12,000 to size infinite. We want to do that before augmentation. As opposed to taking and repeating 12,000 augmented images.\n\nEDIT: The order doesn't matter. I posted a test above.",
      "votes": null
    },
    {
      "id": "755436",
      "postDate": "02/24/2020 19:20:10",
      "content": "<p>CORRECTION: The TF team confirmed that <code>map</code> and <code>repeat</code> are commutative. Repeat repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.</p>\n\n<hr>\n\n<p>[original (wrong) comment]</p>\n\n<p>You are correct, as defined, the dataset is augmented then repeated. Entropy will be much better if you repeat and then augment.</p>\n\n<p>The tf.data.Dataset API does not create a set of data as such, it creates a set of transformations to be applied to data. These transformations are then applied as you train. For training data, it is recommended to use an infinitely repeating dataset that will cover all epochs of training. That's what the <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/#Datasets\">getting started notebook</a> does. But there are indeed two ways of getting an infinitely repeating dataset: augmenting the dataset, then repeating it and the other way around. </p>",
      "rawMarkdown": "CORRECTION: The TF team confirmed that `map` and `repeat` are commutative. Repeat repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.\n\n------\n[original (wrong) comment]\n\nYou are correct, as defined, the dataset is augmented then repeated. Entropy will be much better if you repeat and then augment.\n\nThe tf.data.Dataset API does not create a set of data as such, it creates a set of transformations to be applied to data. These transformations are then applied as you train. For training data, it is recommended to use an infinitely repeating dataset that will cover all epochs of training. That's what the [getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/#Datasets) does. But there are indeed two ways of getting an infinitely repeating dataset: augmenting the dataset, then repeating it and the other way around.",
      "votes": null
    },
    {
      "id": "755437",
      "postDate": "02/24/2020 19:22:14",
      "content": "<p>CORRECTION: The TF team confirmed that map and repeat are commutative. Repeat repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.</p>\n\n<hr>\n\n<p>[original (wrong) comment]</p>\n\n<p>Not in the <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/#Datasets\">getting started notebook</a>. Since the dataset is first augmented, then repeated indefinitely to cover all epochs, the same initially augmented images are being looped through at each epoch. Not ideal indeed.</p>",
      "rawMarkdown": "CORRECTION: The TF team confirmed that map and repeat are commutative. Repeat repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.\n\n---\n[original (wrong) comment]\n\nNot in the [getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/#Datasets). Since the dataset is first augmented, then repeated indefinitely to cover all epochs, the same initially augmented images are being looped through at each epoch. Not ideal indeed.",
      "votes": null
    },
    {
      "id": "755445",
      "postDate": "02/24/2020 19:26:50",
      "content": "<p>I just ran some tests. We get different augmentation each epoch. It doesn't matter where we make the call <code>map(data_augment)</code> in the chain. Upon reflection, this makes sense. TensorFlow doesn't compute anything until when it is needed. Augmentation includes a random number tensor. So every time augmentation is asked for, TensorFlow produces a random number Tensor first.</p>\n\n<h3>No Augmentation</h3>\n\n<pre><code>all = get_training_dataset(load_dataset(TRAINING_FILENAMES)).unbatch().batch(4)\nbatch = next(iter(all))\nfour_images = tf.data.Dataset.from_tensors( batch ).unbatch().repeat()\nrow = 2; col = 4;\nplt.figure(figsize=(15,int(15*row/col))\nfor j,(img,label) in enumerate(four_images):\n    plt.subplot(row,col,j+1)\n    plt.axis('off'); plt.imshow(img)\n    if j&amp;gt;=7: break\nplt.show()\n</code></pre>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff1fc243fb7c4c2373902a9e26f653f22%2FScreen%20Shot%202020-02-24%20at%2011.19.15%20AM.png?generation=1582571973827383&amp;alt=media\" alt=\"\"></p>\n\n<h3>Augment AFTER Repeat</h3>\n\n<p>Look closely and you will see that some images are flipped left to right.</p>\n\n<pre><code>four_images = tf.data.Dataset.from_tensors( batch ).unbatch().repeat().map(data_augment)\n</code></pre>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F83d6c6badc4c6c6f5749fd91b7d54b25%2FScreen%20Shot%202020-02-24%20at%2011.22.32%20AM.png?generation=1582572192907792&amp;alt=media\" alt=\"\"></p>\n\n<h3>Augment BEFORE Repeat</h3>\n\n<p>Look closely and you will see that some images are flipped left to right.</p>\n\n<pre><code>four_images = tf.data.Dataset.from_tensors( batch ).unbatch().map(data_augment).repeat()\n</code></pre>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F3a99ea476c156a9e5fbda2a73ea9f6ff%2FScreen%20Shot%202020-02-24%20at%2011.24.01%20AM.png?generation=1582572256025671&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I just ran some tests. We get different augmentation each epoch. It doesn't matter where we make the call `map(data_augment)` in the chain. Upon reflection, this makes sense. TensorFlow doesn't compute anything until when it is needed. Augmentation includes a random number tensor. So every time augmentation is asked for, TensorFlow produces a random number Tensor first.\n\n### No Augmentation\n    all = get_training_dataset(load_dataset(TRAINING_FILENAMES)).unbatch().batch(4)\n    batch = next(iter(all))\n    four_images = tf.data.Dataset.from_tensors( batch ).unbatch().repeat()\n    row = 2; col = 4;\n    plt.figure(figsize=(15,int(15*row/col))\n    for j,(img,label) in enumerate(four_images):\n        plt.subplot(row,col,j+1)\n        plt.axis('off'); plt.imshow(img)\n        if j&gt;=7: break\n    plt.show()\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff1fc243fb7c4c2373902a9e26f653f22%2FScreen%20Shot%202020-02-24%20at%2011.19.15%20AM.png?generation=1582571973827383&amp;alt=media)\n\n### Augment AFTER Repeat\nLook closely and you will see that some images are flipped left to right.\n\n    four_images = tf.data.Dataset.from_tensors( batch ).unbatch().repeat().map(data_augment)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F83d6c6badc4c6c6f5749fd91b7d54b25%2FScreen%20Shot%202020-02-24%20at%2011.22.32%20AM.png?generation=1582572192907792&amp;alt=media)\n\n### Augment BEFORE Repeat\nLook closely and you will see that some images are flipped left to right.\n\n    four_images = tf.data.Dataset.from_tensors( batch ).unbatch().map(data_augment).repeat()\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F3a99ea476c156a9e5fbda2a73ea9f6ff%2FScreen%20Shot%202020-02-24%20at%2011.24.01%20AM.png?generation=1582572256025671&amp;alt=media)",
      "votes": null
    },
    {
      "id": "755448",
      "postDate": "02/24/2020 19:27:56",
      "content": "<p>Hey Martin, I just did a test. It doesn't appear to matter whether you augment or repeat first.</p>",
      "rawMarkdown": "Hey Martin, I just did a test. It doesn't appear to matter whether you augment or repeat first.",
      "votes": null
    },
    {
      "id": "755467",
      "postDate": "02/24/2020 19:51:38",
      "content": "<p>Thanks for the input guys.\nHmm. I would say entropy matters :)\nAnyway - can one tell from 2 epochs, if there is a difference?</p>\n\n<p>I guess one can :)</p>",
      "rawMarkdown": "Thanks for the input guys.\nHmm. I would say entropy matters :)\nAnyway - can one tell from 2 epochs, if there is a difference?\n\nI guess one can :)",
      "votes": null
    },
    {
      "id": "755472",
      "postDate": "02/24/2020 19:55:47",
      "content": "<p>Yes, 2 epochs is enough. If TensorFlow augmented all the images once and then repeated them without augmenting the new epoch, we would see that in my experiment below using only 2 epochs. (But my experiments show that each epoch is getting new images regardless of the location of <code>dataaugment</code> in chain of calls).</p>",
      "rawMarkdown": "Yes, 2 epochs is enough. If TensorFlow augmented all the images once and then repeated them without augmenting the new epoch, we would see that in my experiment below using only 2 epochs. (But my experiments show that each epoch is getting new images regardless of the location of `dataaugment` in chain of calls).",
      "votes": null
    },
    {
      "id": "755610",
      "postDate": "02/25/2020 00:09:35",
      "content": "<p>Indeed. I tried as well in this <a href=\"https://colab.research.google.com/drive/1dw2M866LyaiKWTDamgrGPBsYIEMWUW_U\">Colab notebook</a>.\nYou seem to be right, the dataset is augmented correctly, with fresh data augmentation at each epoch, no matter the order of augment and repeat.</p>",
      "rawMarkdown": "Indeed. I tried as well in this [Colab notebook](https://colab.research.google.com/drive/1dw2M866LyaiKWTDamgrGPBsYIEMWUW_U).\nYou seem to be right, the dataset is augmented correctly, with fresh data augmentation at each epoch, no matter the order of augment and repeat.",
      "votes": null
    },
    {
      "id": "755612",
      "postDate": "02/25/2020 00:09:56",
      "content": "<p>Let me ask the TF team wether this is normal.</p>",
      "rawMarkdown": "Let me ask the TF team wether this is normal.",
      "votes": null
    },
    {
      "id": "756379",
      "postDate": "02/25/2020 17:08:14",
      "content": "<p>The TF team confirmed that <code>map</code> and <code>repeat</code> are indeed commutative. <code>Repeat</code> repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.</p>",
      "rawMarkdown": "The TF team confirmed that `map` and `repeat` are indeed commutative. `Repeat` repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.",
      "votes": null
    },
    {
      "id": "756405",
      "postDate": "02/25/2020 17:43:08",
      "content": "<p>Great you could find out - thx</p>",
      "rawMarkdown": "Great you could find out - thx",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 755369,
      "author_name": "shank885",
      "author_url": "",
      "post_date": "02/24/2020 17:57:36",
      "content": "<p>As much as I know, the data is augmented and feed to the model.fit() every epoch and so is will keep changing every epoch. The get_training_data() function returns the new random variant of augmented data for each epoch.</p>",
      "votes": null,
      "replies": [
        {
          "id": 755437,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/24/2020 19:22:14",
          "content": "<p>CORRECTION: The TF team confirmed that map and repeat are commutative. Repeat repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.</p>\n\n<hr>\n\n<p>[original (wrong) comment]</p>\n\n<p>Not in the <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/#Datasets\">getting started notebook</a>. Since the dataset is first augmented, then repeated indefinitely to cover all epochs, the same initially augmented images are being looped through at each epoch. Not ideal indeed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755421,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/24/2020 19:02:07",
      "content": "<p>This is a great question. The best way to answer this it to make a small dataset of 10 images, and display what is happening. Because the TensorFlow documentation may say one thing but what is actually happening may be different.</p>\n\n<p>I think you may be right, that <code>dataset = dataset.repeat()</code> may need to be before <code>dataset = dataset.map(dataaugment)</code>. Because the repeat command turns the dataset from size 12,000 to size infinite. We want to do that before augmentation. As opposed to taking and repeating 12,000 augmented images.</p>\n\n<p>EDIT: The order doesn't matter. I posted a test above.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755436,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/24/2020 19:20:10",
      "content": "<p>CORRECTION: The TF team confirmed that <code>map</code> and <code>repeat</code> are commutative. Repeat repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.</p>\n\n<hr>\n\n<p>[original (wrong) comment]</p>\n\n<p>You are correct, as defined, the dataset is augmented then repeated. Entropy will be much better if you repeat and then augment.</p>\n\n<p>The tf.data.Dataset API does not create a set of data as such, it creates a set of transformations to be applied to data. These transformations are then applied as you train. For training data, it is recommended to use an infinitely repeating dataset that will cover all epochs of training. That's what the <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/#Datasets\">getting started notebook</a> does. But there are indeed two ways of getting an infinitely repeating dataset: augmenting the dataset, then repeating it and the other way around. </p>",
      "votes": null,
      "replies": [
        {
          "id": 755448,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/24/2020 19:27:56",
          "content": "<p>Hey Martin, I just did a test. It doesn't appear to matter whether you augment or repeat first.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755610,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/25/2020 00:09:35",
          "content": "<p>Indeed. I tried as well in this <a href=\"https://colab.research.google.com/drive/1dw2M866LyaiKWTDamgrGPBsYIEMWUW_U\">Colab notebook</a>.\nYou seem to be right, the dataset is augmented correctly, with fresh data augmentation at each epoch, no matter the order of augment and repeat.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755612,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/25/2020 00:09:56",
          "content": "<p>Let me ask the TF team wether this is normal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756379,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/25/2020 17:08:14",
          "content": "<p>The TF team confirmed that <code>map</code> and <code>repeat</code> are indeed commutative. <code>Repeat</code> repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756405,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "02/25/2020 17:43:08",
          "content": "<p>Great you could find out - thx</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755445,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/24/2020 19:26:50",
      "content": "<p>I just ran some tests. We get different augmentation each epoch. It doesn't matter where we make the call <code>map(data_augment)</code> in the chain. Upon reflection, this makes sense. TensorFlow doesn't compute anything until when it is needed. Augmentation includes a random number tensor. So every time augmentation is asked for, TensorFlow produces a random number Tensor first.</p>\n\n<h3>No Augmentation</h3>\n\n<pre><code>all = get_training_dataset(load_dataset(TRAINING_FILENAMES)).unbatch().batch(4)\nbatch = next(iter(all))\nfour_images = tf.data.Dataset.from_tensors( batch ).unbatch().repeat()\nrow = 2; col = 4;\nplt.figure(figsize=(15,int(15*row/col))\nfor j,(img,label) in enumerate(four_images):\n    plt.subplot(row,col,j+1)\n    plt.axis('off'); plt.imshow(img)\n    if j&amp;gt;=7: break\nplt.show()\n</code></pre>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff1fc243fb7c4c2373902a9e26f653f22%2FScreen%20Shot%202020-02-24%20at%2011.19.15%20AM.png?generation=1582571973827383&amp;alt=media\" alt=\"\"></p>\n\n<h3>Augment AFTER Repeat</h3>\n\n<p>Look closely and you will see that some images are flipped left to right.</p>\n\n<pre><code>four_images = tf.data.Dataset.from_tensors( batch ).unbatch().repeat().map(data_augment)\n</code></pre>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F83d6c6badc4c6c6f5749fd91b7d54b25%2FScreen%20Shot%202020-02-24%20at%2011.22.32%20AM.png?generation=1582572192907792&amp;alt=media\" alt=\"\"></p>\n\n<h3>Augment BEFORE Repeat</h3>\n\n<p>Look closely and you will see that some images are flipped left to right.</p>\n\n<pre><code>four_images = tf.data.Dataset.from_tensors( batch ).unbatch().map(data_augment).repeat()\n</code></pre>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F3a99ea476c156a9e5fbda2a73ea9f6ff%2FScreen%20Shot%202020-02-24%20at%2011.24.01%20AM.png?generation=1582572256025671&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755467,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "02/24/2020 19:51:38",
      "content": "<p>Thanks for the input guys.\nHmm. I would say entropy matters :)\nAnyway - can one tell from 2 epochs, if there is a difference?</p>\n\n<p>I guess one can :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 755472,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/24/2020 19:55:47",
          "content": "<p>Yes, 2 epochs is enough. If TensorFlow augmented all the images once and then repeated them without augmenting the new epoch, we would see that in my experiment below using only 2 epochs. (But my experiments show that each epoch is getting new images regardless of the location of <code>dataaugment</code> in chain of calls).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "755053": "Hallo,\nI am rather new to tf and I have a question about augmentation.\nWhen I have a function like\n\ndef get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n\n...which is used in the model.fit() method, is the dataset calculated new for each epoch and does it make a difference if I call the map() function after repeat.\nI would like to get new augmentations for every epoch, otherwise the variance of the training data would be very poor.\n\nregards",
    "755369": "As much as I know, the data is augmented and feed to the model.fit() every epoch and so is will keep changing every epoch. The get_training_data() function returns the new random variant of augmented data for each epoch.",
    "755421": "This is a great question. The best way to answer this it to make a small dataset of 10 images, and display what is happening. Because the TensorFlow documentation may say one thing but what is actually happening may be different.\n\nI think you may be right, that `dataset = dataset.repeat()` may need to be before `dataset = dataset.map(dataaugment)`. Because the repeat command turns the dataset from size 12,000 to size infinite. We want to do that before augmentation. As opposed to taking and repeating 12,000 augmented images.\n\nEDIT: The order doesn't matter. I posted a test above.",
    "755436": "CORRECTION: The TF team confirmed that `map` and `repeat` are commutative. Repeat repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.\n\n------\n[original (wrong) comment]\n\nYou are correct, as defined, the dataset is augmented then repeated. Entropy will be much better if you repeat and then augment.\n\nThe tf.data.Dataset API does not create a set of data as such, it creates a set of transformations to be applied to data. These transformations are then applied as you train. For training data, it is recommended to use an infinitely repeating dataset that will cover all epochs of training. That's what the [getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/#Datasets) does. But there are indeed two ways of getting an infinitely repeating dataset: augmenting the dataset, then repeating it and the other way around.",
    "755437": "CORRECTION: The TF team confirmed that map and repeat are commutative. Repeat repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.\n\n---\n[original (wrong) comment]\n\nNot in the [getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/#Datasets). Since the dataset is first augmented, then repeated indefinitely to cover all epochs, the same initially augmented images are being looped through at each epoch. Not ideal indeed.",
    "755445": "I just ran some tests. We get different augmentation each epoch. It doesn't matter where we make the call `map(data_augment)` in the chain. Upon reflection, this makes sense. TensorFlow doesn't compute anything until when it is needed. Augmentation includes a random number tensor. So every time augmentation is asked for, TensorFlow produces a random number Tensor first.\n\n### No Augmentation\n    all = get_training_dataset(load_dataset(TRAINING_FILENAMES)).unbatch().batch(4)\n    batch = next(iter(all))\n    four_images = tf.data.Dataset.from_tensors( batch ).unbatch().repeat()\n    row = 2; col = 4;\n    plt.figure(figsize=(15,int(15*row/col))\n    for j,(img,label) in enumerate(four_images):\n        plt.subplot(row,col,j+1)\n        plt.axis('off'); plt.imshow(img)\n        if j&gt;=7: break\n    plt.show()\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff1fc243fb7c4c2373902a9e26f653f22%2FScreen%20Shot%202020-02-24%20at%2011.19.15%20AM.png?generation=1582571973827383&amp;alt=media)\n\n### Augment AFTER Repeat\nLook closely and you will see that some images are flipped left to right.\n\n    four_images = tf.data.Dataset.from_tensors( batch ).unbatch().repeat().map(data_augment)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F83d6c6badc4c6c6f5749fd91b7d54b25%2FScreen%20Shot%202020-02-24%20at%2011.22.32%20AM.png?generation=1582572192907792&amp;alt=media)\n\n### Augment BEFORE Repeat\nLook closely and you will see that some images are flipped left to right.\n\n    four_images = tf.data.Dataset.from_tensors( batch ).unbatch().map(data_augment).repeat()\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F3a99ea476c156a9e5fbda2a73ea9f6ff%2FScreen%20Shot%202020-02-24%20at%2011.24.01%20AM.png?generation=1582572256025671&amp;alt=media)",
    "755448": "Hey Martin, I just did a test. It doesn't appear to matter whether you augment or repeat first.",
    "755467": "Thanks for the input guys.\nHmm. I would say entropy matters :)\nAnyway - can one tell from 2 epochs, if there is a difference?\n\nI guess one can :)",
    "755472": "Yes, 2 epochs is enough. If TensorFlow augmented all the images once and then repeated them without augmenting the new epoch, we would see that in my experiment below using only 2 epochs. (But my experiments show that each epoch is getting new images regardless of the location of `dataaugment` in chain of calls).",
    "755610": "Indeed. I tried as well in this [Colab notebook](https://colab.research.google.com/drive/1dw2M866LyaiKWTDamgrGPBsYIEMWUW_U).\nYou seem to be right, the dataset is augmented correctly, with fresh data augmentation at each epoch, no matter the order of augment and repeat.",
    "755612": "Let me ask the TF team wether this is normal.",
    "756379": "The TF team confirmed that `map` and `repeat` are indeed commutative. `Repeat` repeats the dataset, including whatever transformations are applied to it, wherever it is placed in the dataset definition.",
    "756405": "Great you could find out - thx"
  },
  "source": "meta"
}