{
  "id": 211725,
  "title": "How can I split the images in the 'train_images' folder to separate folders for each class? ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/211725",
  "author_name": "",
  "post_date": "2021-01-16T10:29:47.164529100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Please could you help me understand how to split the images available in the 'train_images' folder to separate folders for each class based on 'train.csv'? </p>",
  "messages": [
    {
      "id": "1155256",
      "postDate": "01/16/2021 10:29:47",
      "content": "<p>Please could you help me understand how to split the images available in the 'train_images' folder to separate folders for each class based on 'train.csv'? </p>",
      "rawMarkdown": "Please could you help me understand how to split the images available in the 'train_images' folder to separate folders for each class based on 'train.csv'?",
      "votes": null
    },
    {
      "id": "1155486",
      "postDate": "01/16/2021 12:22:58",
      "content": "<p>Try keras preprocessing <strong>ImageDataGenerator.flow_from_dataframe</strong> method.<br>\nMust visit <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator#flow_from_directory\" target=\"_blank\">Tensorflow doc</a> and <a href=\"https://vijayabhaskar96.medium.com/tutorial-on-keras-imagedatagenerator-with-flow-from-dataframe-8bd5776e45c1\" target=\"_blank\">this tutorial article</a> for same.</p>\n<pre><code># This the quick implementation of ImageDataGenerator.flow_from_dataframe\n\nimport pandas as pd\nfrom keras_preprocessing.image import ImageDataGenerator\n\ndf = pd.read_csv('train.csv')\ndatagen = ImageDataGenerator(rescale=1./255)\n\ntrain_generator = datagen.flow_from_dataframe(dataframe=df, directory='.\\train_imgs', x_col='image_id', y_col='label', class_mode=\"categorical\", target_size=(256,256), batch_size=256)\n</code></pre>\n<p>This train_generator returns A <strong>DataFrameIterator</strong> yielding tuples of <strong>(x, y)</strong> where <strong>x</strong> is a NumPy array containing a batch of images with shape <strong>(batch_size, *target_size, channels)</strong> and <strong>y</strong> is a NumPy array of corresponding <strong>labels</strong>.</p>",
      "rawMarkdown": "Try keras preprocessing **ImageDataGenerator.flow_from_dataframe** method.\nMust visit [Tensorflow doc](https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator#flow_from_directory) and [this tutorial article](https://vijayabhaskar96.medium.com/tutorial-on-keras-imagedatagenerator-with-flow-from-dataframe-8bd5776e45c1) for same.\n\n```\n# This the quick implementation of ImageDataGenerator.flow_from_dataframe\n\nimport pandas as pd\nfrom keras_preprocessing.image import ImageDataGenerator\n\ndf = pd.read_csv('train.csv')\ndatagen = ImageDataGenerator(rescale=1./255)\n\ntrain_generator = datagen.flow_from_dataframe(dataframe=df, directory='.\\train_imgs', x_col='image_id', y_col='label', class_mode=\"categorical\", target_size=(256,256), batch_size=256)\n```\n\nThis train_generator returns A **DataFrameIterator** yielding tuples of **(x, y)** where **x** is a NumPy array containing a batch of images with shape **(batch_size, *target_size, channels)** and **y** is a NumPy array of corresponding **labels**.",
      "votes": null
    },
    {
      "id": "1156149",
      "postDate": "01/17/2021 00:38:59",
      "content": "<p>There are many ways to split the images in train.csv. One of the basic ways is to simply split using sklearn's<br>\n<code>train_test_split()</code>. </p>\n<p>Let say you have 1000 images in your dataset. You do test_train_split set to 0.8. X_train will have 800 images and Valid will have 800 labels corresponding to the 800 images in it, X_test will have 200 images and Y_test will have 200 labels. That's your train and test set.</p>\n<p>There are many other ways such as stratified kfolding and such, but these are usually the basics of splitting your data. Hope this helps!</p>",
      "rawMarkdown": "There are many ways to split the images in train.csv. One of the basic ways is to simply split using sklearn's\n```train_test_split()```. \n\nLet say you have 1000 images in your dataset. You do test_train_split set to 0.8. X_train will have 800 images and Valid will have 800 labels corresponding to the 800 images in it, X_test will have 200 images and Y_test will have 200 labels. That's your train and test set.\n\nThere are many other ways such as stratified kfolding and such, but these are usually the basics of splitting your data. Hope this helps!",
      "votes": null
    },
    {
      "id": "1156732",
      "postDate": "01/17/2021 11:22:14",
      "content": "<p>You basically have to traverse through the 'train.csv' file and read the labels (from 0 to 4) .Create 5 different folders corresponding to each of the five categories in your directory. Then copy the images from the 'train_images' folder to the folders that you have created by comparing the labels using the code:</p>\n<blockquote>\n  <p>shutil.copy(src_path,dest_path)</p>\n</blockquote>\n<p>After you have done this, images would be split into five folders. Then to split train and test images run the code:</p>\n<blockquote>\n  <p>splitfolders.ratio(src_path, output=dest_path, seed=1337, ratio=(.8, .2), group_prefix=None)</p>\n</blockquote>\n<p>Here the source path is where you have split your 'train_images' folder images into five folders . The destination path is of your choice. Hope this helps!!</p>",
      "rawMarkdown": "You basically have to traverse through the 'train.csv' file and read the labels (from 0 to 4) .Create 5 different folders corresponding to each of the five categories in your directory. Then copy the images from the 'train_images' folder to the folders that you have created by comparing the labels using the code:\n> shutil.copy(src_path,dest_path)\n\n\nAfter you have done this, images would be split into five folders. Then to split train and test images run the code:\n> splitfolders.ratio(src_path, output=dest_path, seed=1337, ratio=(.8, .2), group_prefix=None)\n\nHere the source path is where you have split your 'train_images' folder images into five folders . The destination path is of your choice. Hope this helps!!",
      "votes": null
    },
    {
      "id": "1161118",
      "postDate": "01/20/2021 10:42:23",
      "content": "<p>Thank you all, you guys are really awesome!!!</p>\n<p>I was able to read the 'train.csv' file using a custom function and was able to create separate folders for each class.<br>\nAlso, tried using flow_from_dataframe() and it works as well.</p>\n<p>Then, used splitfolders.ratio() to split the images into train, val, and test.<br>\nI've used training_test_split() for non-image data before. Maybe, I'll try this for image data next time.</p>\n<p>Have just completed preprocessing and started to train the model. Good luck with your models. :)</p>\n<p>Thanks again,<br>\nArun</p>",
      "rawMarkdown": "Thank you all, you guys are really awesome!!!\n\nI was able to read the 'train.csv' file using a custom function and was able to create separate folders for each class.\nAlso, tried using flow_from_dataframe() and it works as well.\n\nThen, used splitfolders.ratio() to split the images into train, val, and test.\nI've used training_test_split() for non-image data before. Maybe, I'll try this for image data next time.\n\nHave just completed preprocessing and started to train the model. Good luck with your models. :)\n\nThanks again,\nArun",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1155486,
      "author_name": "vatsalmavani",
      "author_url": "",
      "post_date": "01/16/2021 12:22:58",
      "content": "<p>Try keras preprocessing <strong>ImageDataGenerator.flow_from_dataframe</strong> method.<br>\nMust visit <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator#flow_from_directory\" target=\"_blank\">Tensorflow doc</a> and <a href=\"https://vijayabhaskar96.medium.com/tutorial-on-keras-imagedatagenerator-with-flow-from-dataframe-8bd5776e45c1\" target=\"_blank\">this tutorial article</a> for same.</p>\n<pre><code># This the quick implementation of ImageDataGenerator.flow_from_dataframe\n\nimport pandas as pd\nfrom keras_preprocessing.image import ImageDataGenerator\n\ndf = pd.read_csv('train.csv')\ndatagen = ImageDataGenerator(rescale=1./255)\n\ntrain_generator = datagen.flow_from_dataframe(dataframe=df, directory='.\\train_imgs', x_col='image_id', y_col='label', class_mode=\"categorical\", target_size=(256,256), batch_size=256)\n</code></pre>\n<p>This train_generator returns A <strong>DataFrameIterator</strong> yielding tuples of <strong>(x, y)</strong> where <strong>x</strong> is a NumPy array containing a batch of images with shape <strong>(batch_size, *target_size, channels)</strong> and <strong>y</strong> is a NumPy array of corresponding <strong>labels</strong>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156149,
      "author_name": "andyjianzhou",
      "author_url": "",
      "post_date": "01/17/2021 00:38:59",
      "content": "<p>There are many ways to split the images in train.csv. One of the basic ways is to simply split using sklearn's<br>\n<code>train_test_split()</code>. </p>\n<p>Let say you have 1000 images in your dataset. You do test_train_split set to 0.8. X_train will have 800 images and Valid will have 800 labels corresponding to the 800 images in it, X_test will have 200 images and Y_test will have 200 labels. That's your train and test set.</p>\n<p>There are many other ways such as stratified kfolding and such, but these are usually the basics of splitting your data. Hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156732,
      "author_name": "shameemarmy",
      "author_url": "",
      "post_date": "01/17/2021 11:22:14",
      "content": "<p>You basically have to traverse through the 'train.csv' file and read the labels (from 0 to 4) .Create 5 different folders corresponding to each of the five categories in your directory. Then copy the images from the 'train_images' folder to the folders that you have created by comparing the labels using the code:</p>\n<blockquote>\n  <p>shutil.copy(src_path,dest_path)</p>\n</blockquote>\n<p>After you have done this, images would be split into five folders. Then to split train and test images run the code:</p>\n<blockquote>\n  <p>splitfolders.ratio(src_path, output=dest_path, seed=1337, ratio=(.8, .2), group_prefix=None)</p>\n</blockquote>\n<p>Here the source path is where you have split your 'train_images' folder images into five folders . The destination path is of your choice. Hope this helps!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1161118,
      "author_name": "arun728",
      "author_url": "",
      "post_date": "01/20/2021 10:42:23",
      "content": "<p>Thank you all, you guys are really awesome!!!</p>\n<p>I was able to read the 'train.csv' file using a custom function and was able to create separate folders for each class.<br>\nAlso, tried using flow_from_dataframe() and it works as well.</p>\n<p>Then, used splitfolders.ratio() to split the images into train, val, and test.<br>\nI've used training_test_split() for non-image data before. Maybe, I'll try this for image data next time.</p>\n<p>Have just completed preprocessing and started to train the model. Good luck with your models. :)</p>\n<p>Thanks again,<br>\nArun</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1155256": "Please could you help me understand how to split the images available in the 'train_images' folder to separate folders for each class based on 'train.csv'?",
    "1155486": "Try keras preprocessing **ImageDataGenerator.flow_from_dataframe** method.\nMust visit [Tensorflow doc](https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator#flow_from_directory) and [this tutorial article](https://vijayabhaskar96.medium.com/tutorial-on-keras-imagedatagenerator-with-flow-from-dataframe-8bd5776e45c1) for same.\n\n```\n# This the quick implementation of ImageDataGenerator.flow_from_dataframe\n\nimport pandas as pd\nfrom keras_preprocessing.image import ImageDataGenerator\n\ndf = pd.read_csv('train.csv')\ndatagen = ImageDataGenerator(rescale=1./255)\n\ntrain_generator = datagen.flow_from_dataframe(dataframe=df, directory='.\\train_imgs', x_col='image_id', y_col='label', class_mode=\"categorical\", target_size=(256,256), batch_size=256)\n```\n\nThis train_generator returns A **DataFrameIterator** yielding tuples of **(x, y)** where **x** is a NumPy array containing a batch of images with shape **(batch_size, *target_size, channels)** and **y** is a NumPy array of corresponding **labels**.",
    "1156149": "There are many ways to split the images in train.csv. One of the basic ways is to simply split using sklearn's\n```train_test_split()```. \n\nLet say you have 1000 images in your dataset. You do test_train_split set to 0.8. X_train will have 800 images and Valid will have 800 labels corresponding to the 800 images in it, X_test will have 200 images and Y_test will have 200 labels. That's your train and test set.\n\nThere are many other ways such as stratified kfolding and such, but these are usually the basics of splitting your data. Hope this helps!",
    "1156732": "You basically have to traverse through the 'train.csv' file and read the labels (from 0 to 4) .Create 5 different folders corresponding to each of the five categories in your directory. Then copy the images from the 'train_images' folder to the folders that you have created by comparing the labels using the code:\n> shutil.copy(src_path,dest_path)\n\n\nAfter you have done this, images would be split into five folders. Then to split train and test images run the code:\n> splitfolders.ratio(src_path, output=dest_path, seed=1337, ratio=(.8, .2), group_prefix=None)\n\nHere the source path is where you have split your 'train_images' folder images into five folders . The destination path is of your choice. Hope this helps!!",
    "1161118": "Thank you all, you guys are really awesome!!!\n\nI was able to read the 'train.csv' file using a custom function and was able to create separate folders for each class.\nAlso, tried using flow_from_dataframe() and it works as well.\n\nThen, used splitfolders.ratio() to split the images into train, val, and test.\nI've used training_test_split() for non-image data before. Maybe, I'll try this for image data next time.\n\nHave just completed preprocessing and started to train the model. Good luck with your models. :)\n\nThanks again,\nArun"
  },
  "source": "meta"
}