{
  "id": 206487,
  "title": "How can I make train.csv(like a cassava train.csv) file from this folders?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/206487",
  "author_name": "",
  "post_date": "2020-12-24T22:06:00.115763500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, All.</p>\n<p>I have a dataset in which image files are given separately in folders.</p>\n<p>Each folders have own class name.</p>\n<p>How can I make train.csv(like a cassava train.csv) file from this folders?</p>\n<p>I wonder.</p>\n<p>Thanks.<br>\nBest,</p>\n<p><a href=\"https://www.kaggle.com/bemoregt\" target=\"_blank\">@bemoregt</a>.</p>",
  "messages": [
    {
      "id": "1125629",
      "postDate": "12/24/2020 22:06:00",
      "content": "<p>Hi, All.</p>\n<p>I have a dataset in which image files are given separately in folders.</p>\n<p>Each folders have own class name.</p>\n<p>How can I make train.csv(like a cassava train.csv) file from this folders?</p>\n<p>I wonder.</p>\n<p>Thanks.<br>\nBest,</p>\n<p><a href=\"https://www.kaggle.com/bemoregt\" target=\"_blank\">@bemoregt</a>.</p>",
      "rawMarkdown": "Hi, All.\n\nI have a dataset in which image files are given separately in folders.\n\nEach folders have own class name.\n\nHow can I make train.csv(like a cassava train.csv) file from this folders?\n\nI wonder.\n\nThanks.\nBest,\n\n@bemoregt.",
      "votes": null
    },
    {
      "id": "1126474",
      "postDate": "12/25/2020 16:33:14",
      "content": "<ul>\n<li>image_folder<br>\n&nbsp; &nbsp; &nbsp; -- \\0<br>\n&nbsp; &nbsp; &nbsp; -- \\1<br>\n&nbsp; &nbsp; &nbsp; -- \\2<br>\n&nbsp; &nbsp; &nbsp; -- \\3</li>\n</ul>\n<p>If this is the structure, I would suggest to use python <strong>os</strong> module. If there are only 3/4 classes, I would do this like the following way -</p>\n<pre><code>import os\nimport pandas as pd\n\nroot_dir = '/image_folder'\nclasses = ['0', '1', '2', '3'] # if there are more classes, I would use os.listdir() to get this list.\n\nimage_names= []\nlabels = []\nfor class in classes:\n    # getting folder path for each label\n    class_image_dir = os.path.join(root_dir, class)\n\n    # get the list of all image names in the folder\n    class_specific_image_names = os.listdir(class_image_dir)\n\n    # extending image_names list\n    image_names  = image_names + class_specific_image_names\n\n    # extending labels list with new labels\n    labels = labels + ([class]*len(class_specific_image_names))\n\ndf = pd.Dataframe({'image_name':image_names, 'label': labels})\n\ndf.to_csv(\"train.csv\")\n</code></pre>\n<p>Does this answer help you?</p>",
      "rawMarkdown": "image_folder\n      -- \\0\n      -- \\1\n      -- \\2\n      -- \\3\n\nIf this is the structure, I would suggest to use python **os** module. If there are only 3/4 classes, I would do this like the following way -\n```\nimport os\nimport pandas as pd\n\nroot_dir = '/image_folder'\nclasses = ['0', '1', '2', '3'] # if there are more classes, I would use os.listdir() to get this list.\n\nimage_names= []\nlabels = []\nfor class in classes:\n    # getting folder path for each label\n    class_image_dir = os.path.join(root_dir, class)\n\n    # get the list of all image names in the folder\n    class_specific_image_names = os.listdir(class_image_dir)\n    \n    # extending image_names list\n    image_names  = image_names + class_specific_image_names\n\n    # extending labels list with new labels\n    labels = labels + ([class]*len(class_specific_image_names))\n\ndf = pd.Dataframe({'image_name':image_names, 'label': labels})\n\ndf.to_csv(\"train.csv\")\n```\nDoes this answer help you?",
      "votes": null
    },
    {
      "id": "1126526",
      "postDate": "12/25/2020 17:18:24",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/jabertuhin\" target=\"_blank\">@jabertuhin</a> </p>\n<p>I'll try that.</p>\n<p><a href=\"https://www.kaggle.com/bemoregt\" target=\"_blank\">@bemoregt</a>.</p>",
      "rawMarkdown": "Thank you, @jabertuhin \n\nI'll try that.\n\n@bemoregt.",
      "votes": null
    },
    {
      "id": "1126760",
      "postDate": "12/25/2020 22:43:02",
      "content": "<pre><code>from pathlib import Path\nimport pandas as pd\n\nBASE_DIR = Path(\"base_img_folder\")\n\nall_images = []\nall_labels = []\nfor folder in BASE_DIR.iterdir():\n    current_class_name = folder.name\n    current_class_img_names = list(folder.iterdir())\n    # You can use only the filenames as image_id. In this case,\n    # make sure that the image filenames are all unique.\n    # Because we extending to the list in the below.\n    # current_class_img_names  = [file.name for file in current_class_img_names]\n    all_images.extend(current_class_img_names)\n    all_labels.extend([current_class_name] * len(current_class_img_names ))\n\ndf_data = list(zip(all_images, all_labels))\ntrain_df = pd.DataFrame(df_data, columns=[\"image_id\", \"label\"])\n</code></pre>\n<p>base_img_folder<br>\n|--class_0<br>\n    | -- class_0_img_1.png<br>\n    | -- class_0_img_2.png<br>\n    …<br>\n|--class_1<br>\n    | -- class_1_img_1.png<br>\n    | -- class_1_img_2.png<br>\n    …<br>\n|--class_2<br>\n    | -- class_2_img_1.png<br>\n    | -- class_2_img_2.png<br>\n    …<br>\n|--class_3<br>\n    | -- class_3_img_1.png<br>\n    | -- class_3_img_2.png<br>\n    …</p>\n<p>If your structure is like that, you can use this code as well.</p>",
      "rawMarkdown": "```python\nfrom pathlib import Path\nimport pandas as pd\n\nBASE_DIR = Path(\"base_img_folder\")\n\nall_images = []\nall_labels = []\nfor folder in BASE_DIR.iterdir():\n    current_class_name = folder.name\n    current_class_img_names = list(folder.iterdir())\n    # You can use only the filenames as image_id. In this case,\n    # make sure that the image filenames are all unique.\n    # Because we extending to the list in the below.\n    # current_class_img_names  = [file.name for file in current_class_img_names]\n    all_images.extend(current_class_img_names)\n    all_labels.extend([current_class_name] * len(current_class_img_names ))\n\ndf_data = list(zip(all_images, all_labels))\ntrain_df = pd.DataFrame(df_data, columns=[\"image_id\", \"label\"])\n\n```\n\nbase_img_folder\n|--class_0\n    | -- class_0_img_1.png\n    | -- class_0_img_2.png\n    ...\n|--class_1\n    | -- class_1_img_1.png\n    | -- class_1_img_2.png\n    ...\n|--class_2\n    | -- class_2_img_1.png\n    | -- class_2_img_2.png\n    ...\n|--class_3\n    | -- class_3_img_1.png\n    | -- class_3_img_2.png\n    ...\n\n\nIf your structure is like that, you can use this code as well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1126474,
      "author_name": "jabertuhin",
      "author_url": "",
      "post_date": "12/25/2020 16:33:14",
      "content": "<ul>\n<li>image_folder<br>\n&nbsp; &nbsp; &nbsp; -- \\0<br>\n&nbsp; &nbsp; &nbsp; -- \\1<br>\n&nbsp; &nbsp; &nbsp; -- \\2<br>\n&nbsp; &nbsp; &nbsp; -- \\3</li>\n</ul>\n<p>If this is the structure, I would suggest to use python <strong>os</strong> module. If there are only 3/4 classes, I would do this like the following way -</p>\n<pre><code>import os\nimport pandas as pd\n\nroot_dir = '/image_folder'\nclasses = ['0', '1', '2', '3'] # if there are more classes, I would use os.listdir() to get this list.\n\nimage_names= []\nlabels = []\nfor class in classes:\n    # getting folder path for each label\n    class_image_dir = os.path.join(root_dir, class)\n\n    # get the list of all image names in the folder\n    class_specific_image_names = os.listdir(class_image_dir)\n\n    # extending image_names list\n    image_names  = image_names + class_specific_image_names\n\n    # extending labels list with new labels\n    labels = labels + ([class]*len(class_specific_image_names))\n\ndf = pd.Dataframe({'image_name':image_names, 'label': labels})\n\ndf.to_csv(\"train.csv\")\n</code></pre>\n<p>Does this answer help you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1126526,
          "author_name": "bemorekgg",
          "author_url": "",
          "post_date": "12/25/2020 17:18:24",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/jabertuhin\" target=\"_blank\">@jabertuhin</a> </p>\n<p>I'll try that.</p>\n<p><a href=\"https://www.kaggle.com/bemoregt\" target=\"_blank\">@bemoregt</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1126760,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "12/25/2020 22:43:02",
      "content": "<pre><code>from pathlib import Path\nimport pandas as pd\n\nBASE_DIR = Path(\"base_img_folder\")\n\nall_images = []\nall_labels = []\nfor folder in BASE_DIR.iterdir():\n    current_class_name = folder.name\n    current_class_img_names = list(folder.iterdir())\n    # You can use only the filenames as image_id. In this case,\n    # make sure that the image filenames are all unique.\n    # Because we extending to the list in the below.\n    # current_class_img_names  = [file.name for file in current_class_img_names]\n    all_images.extend(current_class_img_names)\n    all_labels.extend([current_class_name] * len(current_class_img_names ))\n\ndf_data = list(zip(all_images, all_labels))\ntrain_df = pd.DataFrame(df_data, columns=[\"image_id\", \"label\"])\n</code></pre>\n<p>base_img_folder<br>\n|--class_0<br>\n    | -- class_0_img_1.png<br>\n    | -- class_0_img_2.png<br>\n    …<br>\n|--class_1<br>\n    | -- class_1_img_1.png<br>\n    | -- class_1_img_2.png<br>\n    …<br>\n|--class_2<br>\n    | -- class_2_img_1.png<br>\n    | -- class_2_img_2.png<br>\n    …<br>\n|--class_3<br>\n    | -- class_3_img_1.png<br>\n    | -- class_3_img_2.png<br>\n    …</p>\n<p>If your structure is like that, you can use this code as well.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1125629": "Hi, All.\n\nI have a dataset in which image files are given separately in folders.\n\nEach folders have own class name.\n\nHow can I make train.csv(like a cassava train.csv) file from this folders?\n\nI wonder.\n\nThanks.\nBest,\n\n@bemoregt.",
    "1126474": "image_folder\n      -- \\0\n      -- \\1\n      -- \\2\n      -- \\3\n\nIf this is the structure, I would suggest to use python **os** module. If there are only 3/4 classes, I would do this like the following way -\n```\nimport os\nimport pandas as pd\n\nroot_dir = '/image_folder'\nclasses = ['0', '1', '2', '3'] # if there are more classes, I would use os.listdir() to get this list.\n\nimage_names= []\nlabels = []\nfor class in classes:\n    # getting folder path for each label\n    class_image_dir = os.path.join(root_dir, class)\n\n    # get the list of all image names in the folder\n    class_specific_image_names = os.listdir(class_image_dir)\n    \n    # extending image_names list\n    image_names  = image_names + class_specific_image_names\n\n    # extending labels list with new labels\n    labels = labels + ([class]*len(class_specific_image_names))\n\ndf = pd.Dataframe({'image_name':image_names, 'label': labels})\n\ndf.to_csv(\"train.csv\")\n```\nDoes this answer help you?",
    "1126526": "Thank you, @jabertuhin \n\nI'll try that.\n\n@bemoregt.",
    "1126760": "```python\nfrom pathlib import Path\nimport pandas as pd\n\nBASE_DIR = Path(\"base_img_folder\")\n\nall_images = []\nall_labels = []\nfor folder in BASE_DIR.iterdir():\n    current_class_name = folder.name\n    current_class_img_names = list(folder.iterdir())\n    # You can use only the filenames as image_id. In this case,\n    # make sure that the image filenames are all unique.\n    # Because we extending to the list in the below.\n    # current_class_img_names  = [file.name for file in current_class_img_names]\n    all_images.extend(current_class_img_names)\n    all_labels.extend([current_class_name] * len(current_class_img_names ))\n\ndf_data = list(zip(all_images, all_labels))\ntrain_df = pd.DataFrame(df_data, columns=[\"image_id\", \"label\"])\n\n```\n\nbase_img_folder\n|--class_0\n    | -- class_0_img_1.png\n    | -- class_0_img_2.png\n    ...\n|--class_1\n    | -- class_1_img_1.png\n    | -- class_1_img_2.png\n    ...\n|--class_2\n    | -- class_2_img_1.png\n    | -- class_2_img_2.png\n    ...\n|--class_3\n    | -- class_3_img_1.png\n    | -- class_3_img_2.png\n    ...\n\n\nIf your structure is like that, you can use this code as well."
  },
  "source": "meta"
}