{
  "id": 156946,
  "title": "TFRecords Data Augmentation",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156946",
  "author_name": "",
  "post_date": "2020-06-08T16:07:45.633540700Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Can anyone please let me know (and possibly explain) is there any way of doing image augmentation after every epoch while using TFRecords (so something like using ImageDataGenerator)?</p>\n\n<p>I have been looking for a solution for a while and as I had seen so far data augmentation epoch after epoch is only possible with ImageDataGenerator, but it uses numpy arrays so it is not usable with tensors.</p>\n\n<p>In the lack of solution I decided to take a shoot at a different approach with using data in .npy files, but now I have a problem with RAM which I also haven't been able to sort out.</p>",
  "messages": [
    {
      "id": "878527",
      "postDate": "06/08/2020 16:07:45",
      "content": "<p>Can anyone please let me know (and possibly explain) is there any way of doing image augmentation after every epoch while using TFRecords (so something like using ImageDataGenerator)?</p>\n\n<p>I have been looking for a solution for a while and as I had seen so far data augmentation epoch after epoch is only possible with ImageDataGenerator, but it uses numpy arrays so it is not usable with tensors.</p>\n\n<p>In the lack of solution I decided to take a shoot at a different approach with using data in .npy files, but now I have a problem with RAM which I also haven't been able to sort out.</p>",
      "rawMarkdown": "Can anyone please let me know (and possibly explain) is there any way of doing image augmentation after every epoch while using TFRecords (so something like using ImageDataGenerator)?\n\nI have been looking for a solution for a while and as I had seen so far data augmentation epoch after epoch is only possible with ImageDataGenerator, but it uses numpy arrays so it is not usable with tensors.\n\nIn the lack of solution I decided to take a shoot at a different approach with using data in .npy files, but now I have a problem with RAM which I also haven't been able to sort out.",
      "votes": null
    },
    {
      "id": "878547",
      "postDate": "06/08/2020 16:17:04",
      "content": "<p>You can start here: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191</a>\nAnd then you can check tf.image module documentation.</p>",
      "rawMarkdown": "You can start here: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\nAnd then you can check tf.image module documentation.",
      "votes": null
    },
    {
      "id": "878573",
      "postDate": "06/08/2020 16:35:40",
      "content": "<p>You can use the dataset.map(function) where function has as parameter (image, label) and return the image and the label. In that function you can do all kinds of data augmentation. Flower classification has good scripts where you will find this.</p>",
      "rawMarkdown": "You can use the dataset.map(function) where function has as parameter (image, label) and return the image and the label. In that function you can do all kinds of data augmentation. Flower classification has good scripts where you will find this.",
      "votes": null
    },
    {
      "id": "878626",
      "postDate": "06/08/2020 17:19:21",
      "content": "<p>I have another question. How implement (using TFRecords) that version of mixup data augmentation where mixup is done between two random batches (not between images from one batch how was implemented in <a href=\"/cdeotte\">@cdeotte</a> notebook).</p>",
      "rawMarkdown": "I have another question. How implement (using TFRecords) that version of mixup data augmentation where mixup is done between two random batches (not between images from one batch how was implemented in @cdeotte notebook).",
      "votes": null
    },
    {
      "id": "878866",
      "postDate": "06/09/2020 00:46:37",
      "content": "<p>You need to write the augmentations operations in TensorFlow. There are many examples in the discussions and notebooks of Flower Comp. For example <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\">here</a> and <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935\">here</a></p>",
      "rawMarkdown": "You need to write the augmentations operations in TensorFlow. There are many examples in the discussions and notebooks of Flower Comp. For example [here][1] and [here][2]\n\n[1]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\n[2]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935",
      "votes": null
    },
    {
      "id": "879238",
      "postDate": "06/09/2020 10:23:33",
      "content": "<p>Thanks Chris, I took a look at your transform function for rotations, shear, zoom,... Great work!\nBut my question is also this... If we write code like this:</p>\n\n<p><code>\ndef get_training_dataset():\n           dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n           dataset = dataset.map(transform, num_parallel_calls=AUTO) #your transform function\n           ...\n</code></p>\n\n<p>Is augment and transform applied only one time (when we call get_training_dataset function) or is it applied after every epoch (or possibly batch) when we train the network with that data?</p>",
      "rawMarkdown": "Thanks Chris, I took a look at your transform function for rotations, shear, zoom,... Great work!\nBut my question is also this... If we write code like this:\n\n```\ndef get_training_dataset():\n           dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n           dataset = dataset.map(transform, num_parallel_calls=AUTO) #your transform function\n           ...\n```\n\nIs augment and transform applied only one time (when we call get_training_dataset function) or is it applied after every epoch (or possibly batch) when we train the network with that data?",
      "votes": null
    },
    {
      "id": "879542",
      "postDate": "06/09/2020 14:45:13",
      "content": "<p>It is called every epoch. So every epoch you get new unique images.</p>",
      "rawMarkdown": "It is called every epoch. So every epoch you get new unique images.",
      "votes": null
    },
    {
      "id": "879549",
      "postDate": "06/09/2020 14:48:29",
      "content": "<p>For example in the notebook <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">here</a> which demonstrates this augmentation, i make a dataset of size 1, repeat, and apply map with</p>\n\n<pre><code>augmented_element = one_element.repeat().map(transform).batch(row*col)\n</code></pre>\n\n<p>So each epoch has 1 image and each epoch the image is different as shown below\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F579656c5e0f8752fe453cfeeccebd279%2FScreen%20Shot%202020-06-09%20at%207.46.24%20AM.png?generation=1591714066384477&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "For example in the notebook [here][1] which demonstrates this augmentation, i make a dataset of size 1, repeat, and apply map with\n\n    augmented_element = one_element.repeat().map(transform).batch(row*col)\n\nSo each epoch has 1 image and each epoch the image is different as shown below\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F579656c5e0f8752fe453cfeeccebd279%2FScreen%20Shot%202020-06-09%20at%207.46.24%20AM.png?generation=1591714066384477&amp;alt=media)\n\n[1]: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96",
      "votes": null
    },
    {
      "id": "879614",
      "postDate": "06/09/2020 15:29:50",
      "content": "<p>Great, thanks!</p>",
      "rawMarkdown": "Great, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 878547,
      "author_name": "stanislavblinov",
      "author_url": "",
      "post_date": "06/08/2020 16:17:04",
      "content": "<p>You can start here: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191</a>\nAnd then you can check tf.image module documentation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 878573,
      "author_name": "ragnar123",
      "author_url": "",
      "post_date": "06/08/2020 16:35:40",
      "content": "<p>You can use the dataset.map(function) where function has as parameter (image, label) and return the image and the label. In that function you can do all kinds of data augmentation. Flower classification has good scripts where you will find this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 878626,
      "author_name": "andreyzotov",
      "author_url": "",
      "post_date": "06/08/2020 17:19:21",
      "content": "<p>I have another question. How implement (using TFRecords) that version of mixup data augmentation where mixup is done between two random batches (not between images from one batch how was implemented in <a href=\"/cdeotte\">@cdeotte</a> notebook).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 878866,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/09/2020 00:46:37",
      "content": "<p>You need to write the augmentations operations in TensorFlow. There are many examples in the discussions and notebooks of Flower Comp. For example <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\">here</a> and <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 879238,
          "author_name": "frankosikic",
          "author_url": "",
          "post_date": "06/09/2020 10:23:33",
          "content": "<p>Thanks Chris, I took a look at your transform function for rotations, shear, zoom,... Great work!\nBut my question is also this... If we write code like this:</p>\n\n<p><code>\ndef get_training_dataset():\n           dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n           dataset = dataset.map(transform, num_parallel_calls=AUTO) #your transform function\n           ...\n</code></p>\n\n<p>Is augment and transform applied only one time (when we call get_training_dataset function) or is it applied after every epoch (or possibly batch) when we train the network with that data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 879542,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/09/2020 14:45:13",
          "content": "<p>It is called every epoch. So every epoch you get new unique images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 879549,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/09/2020 14:48:29",
          "content": "<p>For example in the notebook <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">here</a> which demonstrates this augmentation, i make a dataset of size 1, repeat, and apply map with</p>\n\n<pre><code>augmented_element = one_element.repeat().map(transform).batch(row*col)\n</code></pre>\n\n<p>So each epoch has 1 image and each epoch the image is different as shown below\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F579656c5e0f8752fe453cfeeccebd279%2FScreen%20Shot%202020-06-09%20at%207.46.24%20AM.png?generation=1591714066384477&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 879614,
          "author_name": "frankosikic",
          "author_url": "",
          "post_date": "06/09/2020 15:29:50",
          "content": "<p>Great, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "878527": "Can anyone please let me know (and possibly explain) is there any way of doing image augmentation after every epoch while using TFRecords (so something like using ImageDataGenerator)?\n\nI have been looking for a solution for a while and as I had seen so far data augmentation epoch after epoch is only possible with ImageDataGenerator, but it uses numpy arrays so it is not usable with tensors.\n\nIn the lack of solution I decided to take a shoot at a different approach with using data in .npy files, but now I have a problem with RAM which I also haven't been able to sort out.",
    "878547": "You can start here: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\nAnd then you can check tf.image module documentation.",
    "878573": "You can use the dataset.map(function) where function has as parameter (image, label) and return the image and the label. In that function you can do all kinds of data augmentation. Flower classification has good scripts where you will find this.",
    "878626": "I have another question. How implement (using TFRecords) that version of mixup data augmentation where mixup is done between two random batches (not between images from one batch how was implemented in @cdeotte notebook).",
    "878866": "You need to write the augmentations operations in TensorFlow. There are many examples in the discussions and notebooks of Flower Comp. For example [here][1] and [here][2]\n\n[1]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\n[2]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935",
    "879238": "Thanks Chris, I took a look at your transform function for rotations, shear, zoom,... Great work!\nBut my question is also this... If we write code like this:\n\n```\ndef get_training_dataset():\n           dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n           dataset = dataset.map(transform, num_parallel_calls=AUTO) #your transform function\n           ...\n```\n\nIs augment and transform applied only one time (when we call get_training_dataset function) or is it applied after every epoch (or possibly batch) when we train the network with that data?",
    "879542": "It is called every epoch. So every epoch you get new unique images.",
    "879549": "For example in the notebook [here][1] which demonstrates this augmentation, i make a dataset of size 1, repeat, and apply map with\n\n    augmented_element = one_element.repeat().map(transform).batch(row*col)\n\nSo each epoch has 1 image and each epoch the image is different as shown below\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F579656c5e0f8752fe453cfeeccebd279%2FScreen%20Shot%202020-06-09%20at%207.46.24%20AM.png?generation=1591714066384477&amp;alt=media)\n\n[1]: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96",
    "879614": "Great, thanks!"
  },
  "source": "meta"
}