{
  "id": 232907,
  "title": "Tensorflow image_dataset_from_directory cannot find image files",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/232907",
  "author_name": "",
  "post_date": "2021-04-16T02:21:05.706817200Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I hope some kind soul help me understand how to use this method.</p>\n<p>I am trying to build a dataset using the <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image_dataset_from_directory\" target=\"_blank\">tf.keras.preprocessing.image_dataset_from_directory </a>.</p>\n<p>In my local machine, I cannot build the dataset in 'train_images' folder. However, I can solve it when I move all image files from 'train_images' folder to a subdirectory '/im'.<br>\nIn Kaggle notebook, I cannot move the files from ./kaggle/input' folder to solve it.(at least I did not figure it out yet).</p>\n<p>My doubts:<br>\na) The name of the image files has anything to do? The filenames seems to be HexBase numbers<br>\nb) Should I use another form to input directory (os, pathlib…)?</p>\n<p>Thanks in advance for any suggestions.</p>\n<p><strong>Code and error message below</strong></p>\n<p>However, I receive the message </p>\n<blockquote>\n  <p>ValueError: Expected the lengths of <code>labels</code> to match the number of files in the target directory. len(labels) is 18632 while we found 0 files in /kaggle/input/plant-pathology-2021-fgvc8/train_images.</p>\n</blockquote>\n<p>The code is below:</p>\n<pre><code>import tensorflow as tf\n\nbatch_size = 32\nimg_height = 100\nimg_width =  100\nchannels = 3\n\ny_train = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv')\npath_img = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\nlabels_list = y_train['labels'].tolist()\n\ntrain_ds = tf.keras.preprocessing.image_dataset_from_directory (\n        path_img\n        , validation_split = 0.1\n        , subset = 'training'\n        , seed=123\n        , shuffle = False\n        , image_size = (img_height, img_width)\n        , batch_size = batch_size \n        , color_mode = 'grayscale'\n        , labels = labels_list\n        , label_mode = 'categorical'\n        )       \n</code></pre>",
  "messages": [
    {
      "id": "1275152",
      "postDate": "04/16/2021 02:21:05",
      "content": "<p>I hope some kind soul help me understand how to use this method.</p>\n<p>I am trying to build a dataset using the <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image_dataset_from_directory\" target=\"_blank\">tf.keras.preprocessing.image_dataset_from_directory </a>.</p>\n<p>In my local machine, I cannot build the dataset in 'train_images' folder. However, I can solve it when I move all image files from 'train_images' folder to a subdirectory '/im'.<br>\nIn Kaggle notebook, I cannot move the files from ./kaggle/input' folder to solve it.(at least I did not figure it out yet).</p>\n<p>My doubts:<br>\na) The name of the image files has anything to do? The filenames seems to be HexBase numbers<br>\nb) Should I use another form to input directory (os, pathlib…)?</p>\n<p>Thanks in advance for any suggestions.</p>\n<p><strong>Code and error message below</strong></p>\n<p>However, I receive the message </p>\n<blockquote>\n  <p>ValueError: Expected the lengths of <code>labels</code> to match the number of files in the target directory. len(labels) is 18632 while we found 0 files in /kaggle/input/plant-pathology-2021-fgvc8/train_images.</p>\n</blockquote>\n<p>The code is below:</p>\n<pre><code>import tensorflow as tf\n\nbatch_size = 32\nimg_height = 100\nimg_width =  100\nchannels = 3\n\ny_train = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv')\npath_img = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\nlabels_list = y_train['labels'].tolist()\n\ntrain_ds = tf.keras.preprocessing.image_dataset_from_directory (\n        path_img\n        , validation_split = 0.1\n        , subset = 'training'\n        , seed=123\n        , shuffle = False\n        , image_size = (img_height, img_width)\n        , batch_size = batch_size \n        , color_mode = 'grayscale'\n        , labels = labels_list\n        , label_mode = 'categorical'\n        )       \n</code></pre>",
      "rawMarkdown": "I hope some kind soul help me understand how to use this method.\n\nI am trying to build a dataset using the [tf.keras.preprocessing.image_dataset_from_directory ](https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image_dataset_from_directory).\n\nIn my local machine, I cannot build the dataset in 'train_images' folder. However, I can solve it when I move all image files from 'train_images' folder to a subdirectory '/im'.\nIn Kaggle notebook, I cannot move the files from ./kaggle/input' folder to solve it.(at least I did not figure it out yet).\n\nMy doubts:\na) The name of the image files has anything to do? The filenames seems to be HexBase numbers\nb) Should I use another form to input directory (os, pathlib...)?\n\nThanks in advance for any suggestions.\n\n**Code and error message below**\n\nHowever, I receive the message \n> ValueError: Expected the lengths of `labels` to match the number of files in the target directory. len(labels) is 18632 while we found 0 files in /kaggle/input/plant-pathology-2021-fgvc8/train_images.\n\nThe code is below:\n\n\n\n```\nimport tensorflow as tf\n\nbatch_size = 32\nimg_height = 100\nimg_width =  100\nchannels = 3\n\ny_train = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv')\npath_img = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\nlabels_list = y_train['labels'].tolist()\n\ntrain_ds = tf.keras.preprocessing.image_dataset_from_directory (\n        path_img\n        , validation_split = 0.1\n        , subset = 'training'\n        , seed=123\n        , shuffle = False\n        , image_size = (img_height, img_width)\n        , batch_size = batch_size \n        , color_mode = 'grayscale'\n        , labels = labels_list\n        , label_mode = 'categorical'\n        )       \n```",
      "votes": null
    },
    {
      "id": "1275292",
      "postDate": "04/16/2021 06:54:41",
      "content": "<p>You cannot use <code>image_dataset_from_directory</code> here as the images are not divided in respective directories of their classes.</p>\n<p>You can try using <code>flow_from_dataframe</code> here.</p>\n<p>Checkout my notebook for reference : <a href=\"https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-gpu-training\" target=\"_blank\">https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-gpu-training</a></p>",
      "rawMarkdown": "You cannot use `image_dataset_from_directory` here as the images are not divided in respective directories of their classes.\n\nYou can try using `flow_from_dataframe` here.\n\nCheckout my notebook for reference : https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-gpu-training",
      "votes": null
    },
    {
      "id": "1279224",
      "postDate": "04/20/2021 18:01:54",
      "content": "<p>Thanks for your notebook. Very helpful.</p>\n<p>Do you know if it it possible to send tuples or list as labels?</p>\n<p>For example, besides 'healthy' or 'scab complex' label, I want use [1,0,0,0,0,0] and [0,1,0,1,0,0] as there are 6 possible unique conditions (complex, healthy, scab, rust, frog_eyed and powdery). </p>\n<p>Then it will become a 6 multi-label classification, and not 12 multi-class classification.</p>",
      "rawMarkdown": "Thanks for your notebook. Very helpful.\n\nDo you know if it it possible to send tuples or list as labels?\n\nFor example, besides 'healthy' or 'scab complex' label, I want use [1,0,0,0,0,0] and [0,1,0,1,0,0] as there are 6 possible unique conditions (complex, healthy, scab, rust, frog_eyed and powdery). \n\n Then it will become a 6 multi-label classification, and not 12 multi-class classification.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1275292,
      "author_name": "shanmukh05",
      "author_url": "",
      "post_date": "04/16/2021 06:54:41",
      "content": "<p>You cannot use <code>image_dataset_from_directory</code> here as the images are not divided in respective directories of their classes.</p>\n<p>You can try using <code>flow_from_dataframe</code> here.</p>\n<p>Checkout my notebook for reference : <a href=\"https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-gpu-training\" target=\"_blank\">https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-gpu-training</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1279224,
          "author_name": "thalesgaluchi",
          "author_url": "",
          "post_date": "04/20/2021 18:01:54",
          "content": "<p>Thanks for your notebook. Very helpful.</p>\n<p>Do you know if it it possible to send tuples or list as labels?</p>\n<p>For example, besides 'healthy' or 'scab complex' label, I want use [1,0,0,0,0,0] and [0,1,0,1,0,0] as there are 6 possible unique conditions (complex, healthy, scab, rust, frog_eyed and powdery). </p>\n<p>Then it will become a 6 multi-label classification, and not 12 multi-class classification.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1275152": "I hope some kind soul help me understand how to use this method.\n\nI am trying to build a dataset using the [tf.keras.preprocessing.image_dataset_from_directory ](https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image_dataset_from_directory).\n\nIn my local machine, I cannot build the dataset in 'train_images' folder. However, I can solve it when I move all image files from 'train_images' folder to a subdirectory '/im'.\nIn Kaggle notebook, I cannot move the files from ./kaggle/input' folder to solve it.(at least I did not figure it out yet).\n\nMy doubts:\na) The name of the image files has anything to do? The filenames seems to be HexBase numbers\nb) Should I use another form to input directory (os, pathlib...)?\n\nThanks in advance for any suggestions.\n\n**Code and error message below**\n\nHowever, I receive the message \n> ValueError: Expected the lengths of `labels` to match the number of files in the target directory. len(labels) is 18632 while we found 0 files in /kaggle/input/plant-pathology-2021-fgvc8/train_images.\n\nThe code is below:\n\n\n\n```\nimport tensorflow as tf\n\nbatch_size = 32\nimg_height = 100\nimg_width =  100\nchannels = 3\n\ny_train = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv')\npath_img = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\nlabels_list = y_train['labels'].tolist()\n\ntrain_ds = tf.keras.preprocessing.image_dataset_from_directory (\n        path_img\n        , validation_split = 0.1\n        , subset = 'training'\n        , seed=123\n        , shuffle = False\n        , image_size = (img_height, img_width)\n        , batch_size = batch_size \n        , color_mode = 'grayscale'\n        , labels = labels_list\n        , label_mode = 'categorical'\n        )       \n```",
    "1275292": "You cannot use `image_dataset_from_directory` here as the images are not divided in respective directories of their classes.\n\nYou can try using `flow_from_dataframe` here.\n\nCheckout my notebook for reference : https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-gpu-training",
    "1279224": "Thanks for your notebook. Very helpful.\n\nDo you know if it it possible to send tuples or list as labels?\n\nFor example, besides 'healthy' or 'scab complex' label, I want use [1,0,0,0,0,0] and [0,1,0,1,0,0] as there are 6 possible unique conditions (complex, healthy, scab, rust, frog_eyed and powdery). \n\n Then it will become a 6 multi-label classification, and not 12 multi-class classification."
  },
  "source": "meta"
}