{
  "id": 134493,
  "title": "Image IO performance",
  "url": "/competitions/bengaliai-cv19/discussion/134493",
  "author_name": "Nedomas",
  "post_date": "2020-03-08T15:23:11.383000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi! </p>\n\n<p>Just wanted to ask if you found a performant way to load the images? I've converted <code>.parquet</code> files into <code>.feather</code> and now concat those. I'm using the Pytorch <code>DataLoader</code> and giving it concat'ed dataset, but the training takes ages (1 epoch ~30mins, on Floydhub's Tesla K80). I'm guessing it's slow in IO.</p>\n\n<p>Here's how I concat it:\n<code>\nimages_dataset = pd.concat([\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_0.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_1.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_2.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_3.feather\"),\n], ignore_index=True)\n</code></p>\n\n<p>I'm wondering - maybe its better to split those <code>feather</code> files into separate image <code>.png</code> files and load them as-required? I've read up on hdf5, but not sure if that would help here.</p>\n\n<p>What was the best image loading method for this dataset for you?</p>",
  "messages": [
    {
      "id": 766699,
      "postDate": "2020-03-08T15:23:11.383Z",
      "content": "<p>Hi! </p>\n\n<p>Just wanted to ask if you found a performant way to load the images? I've converted <code>.parquet</code> files into <code>.feather</code> and now concat those. I'm using the Pytorch <code>DataLoader</code> and giving it concat'ed dataset, but the training takes ages (1 epoch ~30mins, on Floydhub's Tesla K80). I'm guessing it's slow in IO.</p>\n\n<p>Here's how I concat it:\n<code>\nimages_dataset = pd.concat([\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_0.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_1.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_2.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_3.feather\"),\n], ignore_index=True)\n</code></p>\n\n<p>I'm wondering - maybe its better to split those <code>feather</code> files into separate image <code>.png</code> files and load them as-required? I've read up on hdf5, but not sure if that would help here.</p>\n\n<p>What was the best image loading method for this dataset for you?</p>",
      "rawMarkdown": "Hi! \n\nJust wanted to ask if you found a performant way to load the images? I've converted `.parquet` files into `.feather` and now concat those. I'm using the Pytorch `DataLoader` and giving it concat'ed dataset, but the training takes ages (1 epoch ~30mins, on Floydhub's Tesla K80). I'm guessing it's slow in IO.\n\nHere's how I concat it:\n```\nimages_dataset = pd.concat([\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_0.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_1.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_2.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_3.feather\"),\n], ignore_index=True)\n```\n\nI'm wondering - maybe its better to split those `feather` files into separate image `.png` files and load them as-required? I've read up on hdf5, but not sure if that would help here.\n\nWhat was the best image loading method for this dataset for you?"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "766699": "Hi! \n\nJust wanted to ask if you found a performant way to load the images? I've converted `.parquet` files into `.feather` and now concat those. I'm using the Pytorch `DataLoader` and giving it concat'ed dataset, but the training takes ages (1 epoch ~30mins, on Floydhub's Tesla K80). I'm guessing it's slow in IO.\n\nHere's how I concat it:\n```\nimages_dataset = pd.concat([\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_0.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_1.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_2.feather\"),\n    pd.read_feather(f\"{args.images_input_dir}/train_image_data_3.feather\"),\n], ignore_index=True)\n```\n\nI'm wondering - maybe its better to split those `feather` files into separate image `.png` files and load them as-required? I've read up on hdf5, but not sure if that would help here.\n\nWhat was the best image loading method for this dataset for you?"
  }
}