{
  "id": 201990,
  "title": "Question regarding tensorflow datasets and the map function",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201990",
  "author_name": "",
  "post_date": "2020-12-07T18:05:05.692262500Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello Everyone,</p>\n<p>I've been searching around and after reading the tensorflow documentation and stack overflow i couldn't find this out:<br>\nIs everytime a new batch is called from tf.Dataset execute the .map function or just on the line it is executed?<br>\nMy goal is to perform a data augmentation step between epochs, so that the images the model is seeing are never exactly the same, hopefully increasing generalization. If i just call model.fit is it already doing that? or should i implement it explicitly?</p>",
  "messages": [
    {
      "id": "1105283",
      "postDate": "12/07/2020 18:05:05",
      "content": "<p>Hello Everyone,</p>\n<p>I've been searching around and after reading the tensorflow documentation and stack overflow i couldn't find this out:<br>\nIs everytime a new batch is called from tf.Dataset execute the .map function or just on the line it is executed?<br>\nMy goal is to perform a data augmentation step between epochs, so that the images the model is seeing are never exactly the same, hopefully increasing generalization. If i just call model.fit is it already doing that? or should i implement it explicitly?</p>",
      "rawMarkdown": "Hello Everyone,\n\nI've been searching around and after reading the tensorflow documentation and stack overflow i couldn't find this out:\nIs everytime a new batch is called from tf.Dataset execute the .map function or just on the line it is executed?\nMy goal is to perform a data augmentation step between epochs, so that the images the model is seeing are never exactly the same, hopefully increasing generalization. If i just call model.fit is it already doing that? or should i implement it explicitly?",
      "votes": null
    },
    {
      "id": "1105512",
      "postDate": "12/08/2020 00:21:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> ,</p>\n<p>About <code>Is every time a new batch is called from tf.Dataset execute the .map function or just on the line it is executed?</code> it depends on how you write your Tensorflow pipeline, if you are writing code following the tutorial the data augmentation should be different every time, <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training\" target=\"_blank\">here</a> is an example, in practice, the code that I wrote will apply the data augmentation function to each image, so each one will have a different probability of being augmented.</p>",
      "rawMarkdown": "Hi @capiru ,\n\nAbout `Is every time a new batch is called from tf.Dataset execute the .map function or just on the line it is executed?` it depends on how you write your Tensorflow pipeline, if you are writing code following the tutorial the data augmentation should be different every time, [here](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training) is an example, in practice, the code that I wrote will apply the data augmentation function to each image, so each one will have a different probability of being augmented.",
      "votes": null
    },
    {
      "id": "1105973",
      "postDate": "12/08/2020 11:33:53",
      "content": "<p>Thank you for your reply, so having the dataset repeated ensures that when the same image is shown in a different epoch, it has a chance of being augmented. Am i understanding your code correctly?</p>",
      "rawMarkdown": "Thank you for your reply, so having the dataset repeated ensures that when the same image is shown in a different epoch, it has a chance of being augmented. Am i understanding your code correctly?",
      "votes": null
    },
    {
      "id": "1106067",
      "postDate": "12/08/2020 13:31:06",
      "content": "<p>The intuition behind <code>.repeat()</code> is that it makes your data pipeline look like a continuous stream of data, it won't finish after the data ends, this makes the training faster. The reason why each image will have different augmentation each time is that each time a batch is sampled the data augmentation function will be applied again, resulting in different augmentations.</p>",
      "rawMarkdown": "The intuition behind `.repeat()` is that it makes your data pipeline look like a continuous stream of data, it won't finish after the data ends, this makes the training faster. The reason why each image will have different augmentation each time is that each time a batch is sampled the data augmentation function will be applied again, resulting in different augmentations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1105512,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "12/08/2020 00:21:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> ,</p>\n<p>About <code>Is every time a new batch is called from tf.Dataset execute the .map function or just on the line it is executed?</code> it depends on how you write your Tensorflow pipeline, if you are writing code following the tutorial the data augmentation should be different every time, <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training\" target=\"_blank\">here</a> is an example, in practice, the code that I wrote will apply the data augmentation function to each image, so each one will have a different probability of being augmented.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1105973,
          "author_name": "capiru",
          "author_url": "",
          "post_date": "12/08/2020 11:33:53",
          "content": "<p>Thank you for your reply, so having the dataset repeated ensures that when the same image is shown in a different epoch, it has a chance of being augmented. Am i understanding your code correctly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1106067,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/08/2020 13:31:06",
          "content": "<p>The intuition behind <code>.repeat()</code> is that it makes your data pipeline look like a continuous stream of data, it won't finish after the data ends, this makes the training faster. The reason why each image will have different augmentation each time is that each time a batch is sampled the data augmentation function will be applied again, resulting in different augmentations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1105283": "Hello Everyone,\n\nI've been searching around and after reading the tensorflow documentation and stack overflow i couldn't find this out:\nIs everytime a new batch is called from tf.Dataset execute the .map function or just on the line it is executed?\nMy goal is to perform a data augmentation step between epochs, so that the images the model is seeing are never exactly the same, hopefully increasing generalization. If i just call model.fit is it already doing that? or should i implement it explicitly?",
    "1105512": "Hi @capiru ,\n\nAbout `Is every time a new batch is called from tf.Dataset execute the .map function or just on the line it is executed?` it depends on how you write your Tensorflow pipeline, if you are writing code following the tutorial the data augmentation should be different every time, [here](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training) is an example, in practice, the code that I wrote will apply the data augmentation function to each image, so each one will have a different probability of being augmented.",
    "1105973": "Thank you for your reply, so having the dataset repeated ensures that when the same image is shown in a different epoch, it has a chance of being augmented. Am i understanding your code correctly?",
    "1106067": "The intuition behind `.repeat()` is that it makes your data pipeline look like a continuous stream of data, it won't finish after the data ends, this makes the training faster. The reason why each image will have different augmentation each time is that each time a batch is sampled the data augmentation function will be applied again, resulting in different augmentations."
  },
  "source": "meta"
}