{
  "id": 125154,
  "title": "Data augmentation and Pre-Processing",
  "url": "/competitions/bengaliai-cv19/discussion/125154",
  "author_name": "",
  "post_date": "2020-01-08T20:41:03.621881600Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I am trying to upgrade my score using some data augmentation. At the moment without data augmentation I am using some pre-processing to clean up the original images, I use code from (<a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a> ), and fit the model using Keras.\nJust want to make a discussion about my two possible approaches now:</p>\n\nA\n\n<ol>\n<li>Load original images</li>\n<li>Pre-processing algorithms</li>\n<li>Use ImageDataGenerator and model.fit_generator() to train the model.\n<ul><li>This is the obvious method, it is fast and easy.</li></ul></li>\n</ol>\n\nB\n\n<ol>\n<li>Load original images</li>\n<li>I have a custom function that generate the batches used in model.fil_generator(<em>_</em>, ) first argument (works like .flow() function). Inside this function the original images are picked, augmented and only than pre-processed. \n<ul><li>In my opinion this approach makes a lot more sense, but takes a lot more time to run, as it needs to compute all batch images when a new step is initialized.</li></ul></li>\n</ol>\n\n<p>How are you guys handling the data augmentation and pre-processing order?</p>",
  "messages": [
    {
      "id": "713944",
      "postDate": "01/08/2020 20:41:03",
      "content": "<p>I am trying to upgrade my score using some data augmentation. At the moment without data augmentation I am using some pre-processing to clean up the original images, I use code from (<a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a> ), and fit the model using Keras.\nJust want to make a discussion about my two possible approaches now:</p>\n\nA\n\n<ol>\n<li>Load original images</li>\n<li>Pre-processing algorithms</li>\n<li>Use ImageDataGenerator and model.fit_generator() to train the model.\n<ul><li>This is the obvious method, it is fast and easy.</li></ul></li>\n</ol>\n\nB\n\n<ol>\n<li>Load original images</li>\n<li>I have a custom function that generate the batches used in model.fil_generator(<em>_</em>, ) first argument (works like .flow() function). Inside this function the original images are picked, augmented and only than pre-processed. \n<ul><li>In my opinion this approach makes a lot more sense, but takes a lot more time to run, as it needs to compute all batch images when a new step is initialized.</li></ul></li>\n</ol>\n\n<p>How are you guys handling the data augmentation and pre-processing order?</p>",
      "rawMarkdown": "I am trying to upgrade my score using some data augmentation. At the moment without data augmentation I am using some pre-processing to clean up the original images, I use code from (https://www.kaggle.com/iafoss/image-preprocessing-128x128 ), and fit the model using Keras.\nJust want to make a discussion about my two possible approaches now:\n#### A \n1. Load original images\n2. Pre-processing algorithms\n3. Use ImageDataGenerator and model.fit_generator() to train the model.\n- This is the obvious method, it is fast and easy.\n\n#### B\n1. Load original images\n2. I have a custom function that generate the batches used in model.fil_generator(___, ) first argument (works like .flow() function). Inside this function the original images are picked, augmented and only than pre-processed. \n- In my opinion this approach makes a lot more sense, but takes a lot more time to run, as it needs to compute all batch images when a new step is initialized.\n\nHow are you guys handling the data augmentation and pre-processing order?",
      "votes": null
    },
    {
      "id": "714122",
      "postDate": "01/09/2020 04:54:52",
      "content": "<p>You preprocess and save a trainable version of your data, then in the DataGenerator, you use a library such as albumentations to augment the data.\n<a href=\"/cdeotte\">@cdeotte</a> used it to great effect here: <a href=\"https://www.kaggle.com/cdeotte/25-million-images-0-99757-mnist\">25 Million Images! [0.99757] MNIST</a></p>",
      "rawMarkdown": "You preprocess and save a trainable version of your data, then in the DataGenerator, you use a library such as albumentations to augment the data.\n@cdeotte used it to great effect here: [25 Million Images! [0.99757] MNIST](https://www.kaggle.com/cdeotte/25-million-images-0-99757-mnist)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 714122,
      "author_name": "mightyrains",
      "author_url": "",
      "post_date": "01/09/2020 04:54:52",
      "content": "<p>You preprocess and save a trainable version of your data, then in the DataGenerator, you use a library such as albumentations to augment the data.\n<a href=\"/cdeotte\">@cdeotte</a> used it to great effect here: <a href=\"https://www.kaggle.com/cdeotte/25-million-images-0-99757-mnist\">25 Million Images! [0.99757] MNIST</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "713944": "I am trying to upgrade my score using some data augmentation. At the moment without data augmentation I am using some pre-processing to clean up the original images, I use code from (https://www.kaggle.com/iafoss/image-preprocessing-128x128 ), and fit the model using Keras.\nJust want to make a discussion about my two possible approaches now:\n#### A \n1. Load original images\n2. Pre-processing algorithms\n3. Use ImageDataGenerator and model.fit_generator() to train the model.\n- This is the obvious method, it is fast and easy.\n\n#### B\n1. Load original images\n2. I have a custom function that generate the batches used in model.fil_generator(___, ) first argument (works like .flow() function). Inside this function the original images are picked, augmented and only than pre-processed. \n- In my opinion this approach makes a lot more sense, but takes a lot more time to run, as it needs to compute all batch images when a new step is initialized.\n\nHow are you guys handling the data augmentation and pre-processing order?",
    "714122": "You preprocess and save a trainable version of your data, then in the DataGenerator, you use a library such as albumentations to augment the data.\n@cdeotte used it to great effect here: [25 Million Images! [0.99757] MNIST](https://www.kaggle.com/cdeotte/25-million-images-0-99757-mnist)"
  },
  "source": "meta"
}