{
  "id": 108185,
  "title": "Possible to use ImageDataGenerator method from_from_directory with this Kaggle dataset?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/108185",
  "author_name": "",
  "post_date": "2019-09-09T19:09:41.653412300Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Q: Is it possible to use <code>ImageDataGenerator</code> method <code>flow_from_directory</code> with this Kaggle dataset?</strong>\n- For the APTOS 2019 Blindness Detection competition, I was attempting to use the Keras <code>ImageDataGenerator</code> method <code>flow_from_directory</code> to pull the images from the dataset for training, validation, testing.  [<a href=\"https://keras.io/preprocessing/image/#flow_from_directory\">Reference Link</a>]\n- However, that method expects a specific directory structure, with the training and validation images in subdirectories named for their associated classes, and the testing images placed in a single subdirectory for all images.  I was able to implement this on my local PC and perform the training there, but would have preferred to be able to do this directly in the Kaggle kernal.\n- I ended up doing all of the training/validation on my local PC and then performing only the inference via the Kaggle kernel.  (I <em>was</em> able to use flow_from_directory for inference by specifying the full input directory as the directory and indicating the testing folder as the only \"class\", but this was a fluke in my view and not usable for the training and validation tasks.</p>\n\n<p><strong>Q: Is my only option to create my own version of <code>flow_from_directory</code> that doesn't require the segmented directory structure and pulls the label info from the <code>test.csv</code> source instead?</strong>\n- This appears to be what at least some on the leaderboard did...</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "622521",
      "postDate": "09/09/2019 19:09:41",
      "content": "<p><strong>Q: Is it possible to use <code>ImageDataGenerator</code> method <code>flow_from_directory</code> with this Kaggle dataset?</strong>\n- For the APTOS 2019 Blindness Detection competition, I was attempting to use the Keras <code>ImageDataGenerator</code> method <code>flow_from_directory</code> to pull the images from the dataset for training, validation, testing.  [<a href=\"https://keras.io/preprocessing/image/#flow_from_directory\">Reference Link</a>]\n- However, that method expects a specific directory structure, with the training and validation images in subdirectories named for their associated classes, and the testing images placed in a single subdirectory for all images.  I was able to implement this on my local PC and perform the training there, but would have preferred to be able to do this directly in the Kaggle kernal.\n- I ended up doing all of the training/validation on my local PC and then performing only the inference via the Kaggle kernel.  (I <em>was</em> able to use flow_from_directory for inference by specifying the full input directory as the directory and indicating the testing folder as the only \"class\", but this was a fluke in my view and not usable for the training and validation tasks.</p>\n\n<p><strong>Q: Is my only option to create my own version of <code>flow_from_directory</code> that doesn't require the segmented directory structure and pulls the label info from the <code>test.csv</code> source instead?</strong>\n- This appears to be what at least some on the leaderboard did...</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "**Q: Is it possible to use `ImageDataGenerator` method `flow_from_directory` with this Kaggle dataset?**\n- For the APTOS 2019 Blindness Detection competition, I was attempting to use the Keras `ImageDataGenerator` method `flow_from_directory` to pull the images from the dataset for training, validation, testing.  [[Reference Link](https://keras.io/preprocessing/image/#flow_from_directory)]\n- However, that method expects a specific directory structure, with the training and validation images in subdirectories named for their associated classes, and the testing images placed in a single subdirectory for all images.  I was able to implement this on my local PC and perform the training there, but would have preferred to be able to do this directly in the Kaggle kernal.\n- I ended up doing all of the training/validation on my local PC and then performing only the inference via the Kaggle kernel.  (I *was* able to use flow_from_directory for inference by specifying the full input directory as the directory and indicating the testing folder as the only \"class\", but this was a fluke in my view and not usable for the training and validation tasks.\n\n**Q: Is my only option to create my own version of `flow_from_directory` that doesn't require the segmented directory structure and pulls the label info from the `test.csv` source instead?**\n- This appears to be what at least some on the leaderboard did...\n\nThanks!",
      "votes": null
    },
    {
      "id": "623236",
      "postDate": "09/10/2019 16:11:57",
      "content": "<p>use <code>.flow_from_dataframe</code> method.\neg.\n```\ntrain_datagen = ImageDataGenerator(rotation_range=360,\n                                   horizontal_flip=True,\n                                   vertical_flip=True,\n                                   width_shift_range=0.2,\n                                   height_shift_range=0.2,\n                                   preprocessing_function=preprocess_image, \n                                   rescale=1 / 255.)</p>\n\n<p>train_generator = train_datagen.flow_from_dataframe(train_df, \n                                                  x_col='id_code', \n                                                  y_col='diagnosis',\n                                                  directory = '../input/aptos2019-blindness-detection/train_images',\n                                                  target_size=(IMG_WIDTH, IMG_HEIGHT),\n                                                  batch_size=BATCH_SIZE,\n                                                  class_mode='other', \n                                                  )\n```</p>",
      "rawMarkdown": "use ` .flow_from_dataframe` method.\neg.\n```\ntrain_datagen = ImageDataGenerator(rotation_range=360,\n                                   horizontal_flip=True,\n                                   vertical_flip=True,\n                                   width_shift_range=0.2,\n                                   height_shift_range=0.2,\n                                   preprocessing_function=preprocess_image, \n                                   rescale=1 / 255.)\n\ntrain_generator = train_datagen.flow_from_dataframe(train_df, \n                                                  x_col='id_code', \n                                                  y_col='diagnosis',\n                                                  directory = '../input/aptos2019-blindness-detection/train_images',\n                                                  target_size=(IMG_WIDTH, IMG_HEIGHT),\n                                                  batch_size=BATCH_SIZE,\n                                                  class_mode='other', \n                                                  )\n```",
      "votes": null
    },
    {
      "id": "626605",
      "postDate": "09/14/2019 14:39:28",
      "content": "<p>Great, thank you Neeraj -- I will try that approach!  Didn't realize that <code>flow_from_dataframe</code> would support this -- very helpful.</p>",
      "rawMarkdown": "Great, thank you Neeraj -- I will try that approach!  Didn't realize that `flow_from_dataframe` would support this -- very helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 623236,
      "author_name": "neeraj17",
      "author_url": "",
      "post_date": "09/10/2019 16:11:57",
      "content": "<p>use <code>.flow_from_dataframe</code> method.\neg.\n```\ntrain_datagen = ImageDataGenerator(rotation_range=360,\n                                   horizontal_flip=True,\n                                   vertical_flip=True,\n                                   width_shift_range=0.2,\n                                   height_shift_range=0.2,\n                                   preprocessing_function=preprocess_image, \n                                   rescale=1 / 255.)</p>\n\n<p>train_generator = train_datagen.flow_from_dataframe(train_df, \n                                                  x_col='id_code', \n                                                  y_col='diagnosis',\n                                                  directory = '../input/aptos2019-blindness-detection/train_images',\n                                                  target_size=(IMG_WIDTH, IMG_HEIGHT),\n                                                  batch_size=BATCH_SIZE,\n                                                  class_mode='other', \n                                                  )\n```</p>",
      "votes": null,
      "replies": [
        {
          "id": 626605,
          "author_name": "daddyjab",
          "author_url": "",
          "post_date": "09/14/2019 14:39:28",
          "content": "<p>Great, thank you Neeraj -- I will try that approach!  Didn't realize that <code>flow_from_dataframe</code> would support this -- very helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "622521": "**Q: Is it possible to use `ImageDataGenerator` method `flow_from_directory` with this Kaggle dataset?**\n- For the APTOS 2019 Blindness Detection competition, I was attempting to use the Keras `ImageDataGenerator` method `flow_from_directory` to pull the images from the dataset for training, validation, testing.  [[Reference Link](https://keras.io/preprocessing/image/#flow_from_directory)]\n- However, that method expects a specific directory structure, with the training and validation images in subdirectories named for their associated classes, and the testing images placed in a single subdirectory for all images.  I was able to implement this on my local PC and perform the training there, but would have preferred to be able to do this directly in the Kaggle kernal.\n- I ended up doing all of the training/validation on my local PC and then performing only the inference via the Kaggle kernel.  (I *was* able to use flow_from_directory for inference by specifying the full input directory as the directory and indicating the testing folder as the only \"class\", but this was a fluke in my view and not usable for the training and validation tasks.\n\n**Q: Is my only option to create my own version of `flow_from_directory` that doesn't require the segmented directory structure and pulls the label info from the `test.csv` source instead?**\n- This appears to be what at least some on the leaderboard did...\n\nThanks!",
    "623236": "use ` .flow_from_dataframe` method.\neg.\n```\ntrain_datagen = ImageDataGenerator(rotation_range=360,\n                                   horizontal_flip=True,\n                                   vertical_flip=True,\n                                   width_shift_range=0.2,\n                                   height_shift_range=0.2,\n                                   preprocessing_function=preprocess_image, \n                                   rescale=1 / 255.)\n\ntrain_generator = train_datagen.flow_from_dataframe(train_df, \n                                                  x_col='id_code', \n                                                  y_col='diagnosis',\n                                                  directory = '../input/aptos2019-blindness-detection/train_images',\n                                                  target_size=(IMG_WIDTH, IMG_HEIGHT),\n                                                  batch_size=BATCH_SIZE,\n                                                  class_mode='other', \n                                                  )\n```",
    "626605": "Great, thank you Neeraj -- I will try that approach!  Didn't realize that `flow_from_dataframe` would support this -- very helpful."
  },
  "source": "meta"
}