{
  "id": 305531,
  "title": "Loading images for training",
  "url": "/competitions/happy-whale-and-dolphin/discussion/305531",
  "author_name": "",
  "post_date": "2022-02-05T18:31:43.385943900Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>How can I load these images for training without running out of memory?  I tried the following:</p>\n<pre><code>from keras.preprocessing.image import ImageDataGenerator\n\ndatagen = ImageDataGenerator()\ntraingenerator = datagen.flow_from_directory(train_images_dir)\n</code></pre>\n<p>But that gives me the following message:</p>\n<p><code>Found 0 images belonging to 0 classes.</code></p>\n<p>So then I tried copying the images to directories according to the individual IDs:</p>\n<pre><code>import os\nimport shutil\n\ntrain_images_dir = \"../input/happy-whale-and-dolphin/train_images/\"\nfor index, row in train.iterrows():\n    individual_id = row['individual_id']\n    if not os.path.exists(individual_id):\n        os.makedirs(individual_id)\n    src = os.path.join(train_images_dir, row['image'])\n    dst = os.path.join(individual_id, row['image'])\n    shutil.copyfile(src, dst)\n</code></pre>\n<p>But that led to an out-of-space error.</p>\n<p>How can this be this done?</p>",
  "messages": [
    {
      "id": "1677451",
      "postDate": "02/05/2022 18:31:43",
      "content": "<p>How can I load these images for training without running out of memory?  I tried the following:</p>\n<pre><code>from keras.preprocessing.image import ImageDataGenerator\n\ndatagen = ImageDataGenerator()\ntraingenerator = datagen.flow_from_directory(train_images_dir)\n</code></pre>\n<p>But that gives me the following message:</p>\n<p><code>Found 0 images belonging to 0 classes.</code></p>\n<p>So then I tried copying the images to directories according to the individual IDs:</p>\n<pre><code>import os\nimport shutil\n\ntrain_images_dir = \"../input/happy-whale-and-dolphin/train_images/\"\nfor index, row in train.iterrows():\n    individual_id = row['individual_id']\n    if not os.path.exists(individual_id):\n        os.makedirs(individual_id)\n    src = os.path.join(train_images_dir, row['image'])\n    dst = os.path.join(individual_id, row['image'])\n    shutil.copyfile(src, dst)\n</code></pre>\n<p>But that led to an out-of-space error.</p>\n<p>How can this be this done?</p>",
      "rawMarkdown": "How can I load these images for training without running out of memory?  I tried the following:\n\n```\nfrom keras.preprocessing.image import ImageDataGenerator\n\ndatagen = ImageDataGenerator()\ntraingenerator = datagen.flow_from_directory(train_images_dir)\n\n```\nBut that gives me the following message:\n\n`Found 0 images belonging to 0 classes.`\n\nSo then I tried copying the images to directories according to the individual IDs:\n\n```\nimport os\nimport shutil\n\ntrain_images_dir = \"../input/happy-whale-and-dolphin/train_images/\"\nfor index, row in train.iterrows():\n    individual_id = row['individual_id']\n    if not os.path.exists(individual_id):\n        os.makedirs(individual_id)\n    src = os.path.join(train_images_dir, row['image'])\n    dst = os.path.join(individual_id, row['image'])\n    shutil.copyfile(src, dst)\n\n```\nBut that led to an out-of-space error.\n\nHow can this be this done?",
      "votes": null
    },
    {
      "id": "1678086",
      "postDate": "02/06/2022 08:51:42",
      "content": "<p>Take a look at <a href=\"https://keras.io/api/preprocessing/image/\" target=\"_blank\">image_dataset_from_directory</a> and <code>label</code> param:</p>\n<blockquote>\n  <p>labels: Either \"inferred\" (labels are generated from the directory structure), None (no labels), or a list/tuple of integer labels of the same size as the number of image files found in the directory. Labels should be sorted according to the alphanumeric order of the image file paths (obtained via os.walk(directory) in Python).</p>\n</blockquote>",
      "rawMarkdown": "Take a look at [image_dataset_from_directory](https://keras.io/api/preprocessing/image/) and `label` param:\n> labels: Either \"inferred\" (labels are generated from the directory structure), None (no labels), or a list/tuple of integer labels of the same size as the number of image files found in the directory. Labels should be sorted according to the alphanumeric order of the image file paths (obtained via os.walk(directory) in Python).",
      "votes": null
    },
    {
      "id": "1679469",
      "postDate": "02/07/2022 08:46:23",
      "content": "<p>A few days ago I have published a <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow/species-classification-starter-tf-data-pipeline\" target=\"_blank\">notebook with a species classification that contains an advance efficient data loading pipeline </a>using tf.data from TensorFlow. Using that approach you will avoid a bottleneck of your data loading.</p>\n<p>Hope, you will find it helpful</p>",
      "rawMarkdown": "A few days ago I have published a [notebook with a species classification that contains an advance efficient data loading pipeline ](https://www.kaggle.com/meowmeowmeowmeowmeow/species-classification-starter-tf-data-pipeline)using tf.data from TensorFlow. Using that approach you will avoid a bottleneck of your data loading.\n\nHope, you will find it helpful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1678086,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "02/06/2022 08:51:42",
      "content": "<p>Take a look at <a href=\"https://keras.io/api/preprocessing/image/\" target=\"_blank\">image_dataset_from_directory</a> and <code>label</code> param:</p>\n<blockquote>\n  <p>labels: Either \"inferred\" (labels are generated from the directory structure), None (no labels), or a list/tuple of integer labels of the same size as the number of image files found in the directory. Labels should be sorted according to the alphanumeric order of the image file paths (obtained via os.walk(directory) in Python).</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1679469,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "02/07/2022 08:46:23",
      "content": "<p>A few days ago I have published a <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow/species-classification-starter-tf-data-pipeline\" target=\"_blank\">notebook with a species classification that contains an advance efficient data loading pipeline </a>using tf.data from TensorFlow. Using that approach you will avoid a bottleneck of your data loading.</p>\n<p>Hope, you will find it helpful</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1677451": "How can I load these images for training without running out of memory?  I tried the following:\n\n```\nfrom keras.preprocessing.image import ImageDataGenerator\n\ndatagen = ImageDataGenerator()\ntraingenerator = datagen.flow_from_directory(train_images_dir)\n\n```\nBut that gives me the following message:\n\n`Found 0 images belonging to 0 classes.`\n\nSo then I tried copying the images to directories according to the individual IDs:\n\n```\nimport os\nimport shutil\n\ntrain_images_dir = \"../input/happy-whale-and-dolphin/train_images/\"\nfor index, row in train.iterrows():\n    individual_id = row['individual_id']\n    if not os.path.exists(individual_id):\n        os.makedirs(individual_id)\n    src = os.path.join(train_images_dir, row['image'])\n    dst = os.path.join(individual_id, row['image'])\n    shutil.copyfile(src, dst)\n\n```\nBut that led to an out-of-space error.\n\nHow can this be this done?",
    "1678086": "Take a look at [image_dataset_from_directory](https://keras.io/api/preprocessing/image/) and `label` param:\n> labels: Either \"inferred\" (labels are generated from the directory structure), None (no labels), or a list/tuple of integer labels of the same size as the number of image files found in the directory. Labels should be sorted according to the alphanumeric order of the image file paths (obtained via os.walk(directory) in Python).",
    "1679469": "A few days ago I have published a [notebook with a species classification that contains an advance efficient data loading pipeline ](https://www.kaggle.com/meowmeowmeowmeowmeow/species-classification-starter-tf-data-pipeline)using tf.data from TensorFlow. Using that approach you will avoid a bottleneck of your data loading.\n\nHope, you will find it helpful"
  },
  "source": "meta"
}