{
  "id": 151940,
  "title": "Iteration Efficiency Techniques",
  "url": "/competitions/herbarium-2020-fgvc7/discussion/151940",
  "author_name": "",
  "post_date": "2020-05-17T18:55:13.596493100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>What techniques are you using to make iterating faster?\nAt the moment, my bottleneck are the dataloaders. I'm making different sized datasets (both total number and image resolution), so I can iterate faster. I currently keep everything in jpeg, but was wondering if anyone is using something faster? Is anyone storing in numpy arrays?</p>\n\n<p>Thanks,\nDaniel</p>",
  "messages": [
    {
      "id": "851614",
      "postDate": "05/17/2020 18:55:13",
      "content": "<p>Hi all,</p>\n\n<p>What techniques are you using to make iterating faster?\nAt the moment, my bottleneck are the dataloaders. I'm making different sized datasets (both total number and image resolution), so I can iterate faster. I currently keep everything in jpeg, but was wondering if anyone is using something faster? Is anyone storing in numpy arrays?</p>\n\n<p>Thanks,\nDaniel</p>",
      "rawMarkdown": "Hi all,\n\nWhat techniques are you using to make iterating faster?\nAt the moment, my bottleneck are the dataloaders. I'm making different sized datasets (both total number and image resolution), so I can iterate faster. I currently keep everything in jpeg, but was wondering if anyone is using something faster? Is anyone storing in numpy arrays?\n\nThanks,\nDaniel",
      "votes": null
    },
    {
      "id": "854247",
      "postDate": "05/19/2020 22:18:10",
      "content": "<p>you can try to change the num_workers in the data loader, apart from that, if you are running in local,then it should be you read speed of the disk/ssd</p>",
      "rawMarkdown": "you can try to change the num_workers in the data loader, apart from that, if you are running in local,then it should be you read speed of the disk/ssd",
      "votes": null
    },
    {
      "id": "854352",
      "postDate": "05/20/2020 01:27:52",
      "content": "<p>I'm not in this competition so i don't know specific details. Regarding image competitions, the bottle neck i usually see is if i'm trying to do fancy data augmentation like non linear transformations. If you're not doing that, they you shouldn't have a bottle neck with dataloaders.</p>",
      "rawMarkdown": "I'm not in this competition so i don't know specific details. Regarding image competitions, the bottle neck i usually see is if i'm trying to do fancy data augmentation like non linear transformations. If you're not doing that, they you shouldn't have a bottle neck with dataloaders.",
      "votes": null
    },
    {
      "id": "855460",
      "postDate": "05/21/2020 00:19:16",
      "content": "<p>That's exactly where my bottleneck is. The data augmentations are slowing me down. Do you happen to have any tips for that? Currently, I made a data set of smaller images, and I data augment off those. Then I'll do one final epoch on the final test set images to realign any distribution differences.</p>\n\n<p>Thanks for the help.</p>",
      "rawMarkdown": "That's exactly where my bottleneck is. The data augmentations are slowing me down. Do you happen to have any tips for that? Currently, I made a data set of smaller images, and I data augment off those. Then I'll do one final epoch on the final test set images to realign any distribution differences.\n\nThanks for the help.",
      "votes": null
    },
    {
      "id": "856618",
      "postDate": "05/21/2020 22:41:45",
      "content": "<p>I confirm the problem. If you train a reasonably small model, just reading and resizing 60 GB of large images is the bottle neck, unless you read from SSD and have like 8 cpu's per GPU. Reducing number of images is problematic as there are 38k classes to learn. </p>\n\n<p>I think your scheme is good. I use cv2.imread/resize, much faster than PIL. </p>\n\n<p>More advanced approaches might be to use the progressive feature of jpgs (or a preview set), get a ROI at low-res, crop high-res before augmentation. I won't have time to try this, though.</p>\n\n<p>Also, does anyone augment on the GPU? Normally not the way to go, but here...</p>",
      "rawMarkdown": "I confirm the problem. If you train a reasonably small model, just reading and resizing 60 GB of large images is the bottle neck, unless you read from SSD and have like 8 cpu's per GPU. Reducing number of images is problematic as there are 38k classes to learn. \n\nI think your scheme is good. I use cv2.imread/resize, much faster than PIL. \n\nMore advanced approaches might be to use the progressive feature of jpgs (or a preview set), get a ROI at low-res, crop high-res before augmentation. I won't have time to try this, though.\n\nAlso, does anyone augment on the GPU? Normally not the way to go, but here...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 854247,
      "author_name": "marcelosanchezortega",
      "author_url": "",
      "post_date": "05/19/2020 22:18:10",
      "content": "<p>you can try to change the num_workers in the data loader, apart from that, if you are running in local,then it should be you read speed of the disk/ssd</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 854352,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/20/2020 01:27:52",
      "content": "<p>I'm not in this competition so i don't know specific details. Regarding image competitions, the bottle neck i usually see is if i'm trying to do fancy data augmentation like non linear transformations. If you're not doing that, they you shouldn't have a bottle neck with dataloaders.</p>",
      "votes": null,
      "replies": [
        {
          "id": 855460,
          "author_name": "datadote",
          "author_url": "",
          "post_date": "05/21/2020 00:19:16",
          "content": "<p>That's exactly where my bottleneck is. The data augmentations are slowing me down. Do you happen to have any tips for that? Currently, I made a data set of smaller images, and I data augment off those. Then I'll do one final epoch on the final test set images to realign any distribution differences.</p>\n\n<p>Thanks for the help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 856618,
      "author_name": "greendolphin",
      "author_url": "",
      "post_date": "05/21/2020 22:41:45",
      "content": "<p>I confirm the problem. If you train a reasonably small model, just reading and resizing 60 GB of large images is the bottle neck, unless you read from SSD and have like 8 cpu's per GPU. Reducing number of images is problematic as there are 38k classes to learn. </p>\n\n<p>I think your scheme is good. I use cv2.imread/resize, much faster than PIL. </p>\n\n<p>More advanced approaches might be to use the progressive feature of jpgs (or a preview set), get a ROI at low-res, crop high-res before augmentation. I won't have time to try this, though.</p>\n\n<p>Also, does anyone augment on the GPU? Normally not the way to go, but here...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "851614": "Hi all,\n\nWhat techniques are you using to make iterating faster?\nAt the moment, my bottleneck are the dataloaders. I'm making different sized datasets (both total number and image resolution), so I can iterate faster. I currently keep everything in jpeg, but was wondering if anyone is using something faster? Is anyone storing in numpy arrays?\n\nThanks,\nDaniel",
    "854247": "you can try to change the num_workers in the data loader, apart from that, if you are running in local,then it should be you read speed of the disk/ssd",
    "854352": "I'm not in this competition so i don't know specific details. Regarding image competitions, the bottle neck i usually see is if i'm trying to do fancy data augmentation like non linear transformations. If you're not doing that, they you shouldn't have a bottle neck with dataloaders.",
    "855460": "That's exactly where my bottleneck is. The data augmentations are slowing me down. Do you happen to have any tips for that? Currently, I made a data set of smaller images, and I data augment off those. Then I'll do one final epoch on the final test set images to realign any distribution differences.\n\nThanks for the help.",
    "856618": "I confirm the problem. If you train a reasonably small model, just reading and resizing 60 GB of large images is the bottle neck, unless you read from SSD and have like 8 cpu's per GPU. Reducing number of images is problematic as there are 38k classes to learn. \n\nI think your scheme is good. I use cv2.imread/resize, much faster than PIL. \n\nMore advanced approaches might be to use the progressive feature of jpgs (or a preview set), get a ROI at low-res, crop high-res before augmentation. I won't have time to try this, though.\n\nAlso, does anyone augment on the GPU? Normally not the way to go, but here..."
  },
  "source": "meta"
}