{
  "id": 171513,
  "title": "PyTorch: What are you all using for fast dataloading? LMDB, HFS5, Cache, etc?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171513",
  "author_name": "",
  "post_date": "2020-08-01T07:03:19.787896Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>One of the largest speed bumps in deep learning can be the dataloading.  There exists many ways to speed up dataloading.  LMDB and HFS5 are very common.  </p>\n\n<p>What are you all using to speed up your data loading?</p>\n\n<p>I do not use either.  </p>\n\n<p>Please describe your epochs.  Each Epoch of mine takes about 2min 35sec.  I read each image, perform over a dozen augmentations, then do validation, then at the end of the fold I do validation on the OOF w/TTA 25.  That's ALOT of loading of the same data over and over (well, almost the same, for validation I have my own dataloader/dataset which is a subset of train).  I do about 15 Epochs and then I am doing 5 Folds as well, so again, LOTS of loading of the same data.</p>\n\n<p>What I do, is a basic \"cache\" in my dataset, like so:</p>\n\n<p>`class MelanomaDataset(Dataset):\n    def <strong>init</strong>(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):</p>\n\n<pre><code>    self.df = df\n    self.imfolder = imfolder\n    self.transforms = transforms\n    self.train = train\n    self.meta_features = meta_features\n</code></pre>\n\n<p>**        shared_array_base = mp.Array(ctypes.c_ubyte, len(self.df)*3*image_size*image_size)\n        shared_array = np.ctypeslib.as_array(shared_array_base.get_obj())\n        self.shared_array = shared_array.reshape(len(self.df), image_size, image_size, 3)\n        self.use_cache = False**</p>\n\n<pre><code>def __getitem__(self, index):\n  **  if not self.use_cache:\n        im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n        image = cv2.imread(im_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        self.shared_array[index] = image\n    image = self.shared_array[index] **\n\n    if self.transforms:\n        augmented = self.transforms(image=image)\n        image = augmented['image']\n\n    meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n    if self.train:\n        y = self.df.iloc[index]['target']\n        return (image, meta), y\n    else:\n        return (image, meta)\n</code></pre>\n\n<p>**    def set_use_cache(self, use_cache):\n        self.use_cache = use_cache**</p>\n\n<pre><code>def __len__(self):\n    return len(self.df)\n</code></pre>\n\n<p><code>\nI then prime the cache(s) at the start of each fold.  This is not necessary, as you could just set</code>dataloader.dataset.use_cache=True` at the end of the first Epoch.  I just like to get it out of the way.</p>\n\n<p><code>\ndef prime_cache(dataloader, name):\n    for idx, (x, y) in enumerate(dataloader, start=1):\n        print(f\"Loading Cache using {name} {idx*dataloader.batch_size}\\r\", end='')\n    print(\"\\n\")\n    dataloader.dataset.set_use_cache(use_cache=True)\n</code></p>\n\n<p><code>\n       if(epoch == 0):\n            prime_cache(train_loader, \"train_loader\")\n        print(\"Training............\\r\", end='')\n</code></p>\n\n<p>This is possible because of how Numpy stores its data.  So no copy is done when going to Numpy.  It simply uses the same contiguous pre-allocated space that the c-type created.  You could do the same with a torch Tensor, but it seems Numpy is a better format before you do your augmentations.  So the first time an image is requested, it fills the cache location.  At the end of an epoch, the cache is 100%, and you can switch the flag <code>use_cache</code> to <code>True</code> and then all subsequent requests come direct from CPU memory.</p>\n\n<p>Obviously this is taking up num_images * image_size space in your memory, and furthermore each num_worker gets a copy so that's num_images * image_size * num_workers.......just something to keep an eye on if you don't have much memory.  </p>\n\n<p>I use three loaders: train, val and test, so I have three caches.</p>\n\n<p>I am especially interested in what your basic flow is per epoch and how fast you can complete an epoch and what type of things you have done to speed up dataloading.</p>",
  "messages": [
    {
      "id": "953838",
      "postDate": "08/01/2020 07:03:19",
      "content": "<p>One of the largest speed bumps in deep learning can be the dataloading.  There exists many ways to speed up dataloading.  LMDB and HFS5 are very common.  </p>\n\n<p>What are you all using to speed up your data loading?</p>\n\n<p>I do not use either.  </p>\n\n<p>Please describe your epochs.  Each Epoch of mine takes about 2min 35sec.  I read each image, perform over a dozen augmentations, then do validation, then at the end of the fold I do validation on the OOF w/TTA 25.  That's ALOT of loading of the same data over and over (well, almost the same, for validation I have my own dataloader/dataset which is a subset of train).  I do about 15 Epochs and then I am doing 5 Folds as well, so again, LOTS of loading of the same data.</p>\n\n<p>What I do, is a basic \"cache\" in my dataset, like so:</p>\n\n<p>`class MelanomaDataset(Dataset):\n    def <strong>init</strong>(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):</p>\n\n<pre><code>    self.df = df\n    self.imfolder = imfolder\n    self.transforms = transforms\n    self.train = train\n    self.meta_features = meta_features\n</code></pre>\n\n<p>**        shared_array_base = mp.Array(ctypes.c_ubyte, len(self.df)*3*image_size*image_size)\n        shared_array = np.ctypeslib.as_array(shared_array_base.get_obj())\n        self.shared_array = shared_array.reshape(len(self.df), image_size, image_size, 3)\n        self.use_cache = False**</p>\n\n<pre><code>def __getitem__(self, index):\n  **  if not self.use_cache:\n        im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n        image = cv2.imread(im_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        self.shared_array[index] = image\n    image = self.shared_array[index] **\n\n    if self.transforms:\n        augmented = self.transforms(image=image)\n        image = augmented['image']\n\n    meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n    if self.train:\n        y = self.df.iloc[index]['target']\n        return (image, meta), y\n    else:\n        return (image, meta)\n</code></pre>\n\n<p>**    def set_use_cache(self, use_cache):\n        self.use_cache = use_cache**</p>\n\n<pre><code>def __len__(self):\n    return len(self.df)\n</code></pre>\n\n<p><code>\nI then prime the cache(s) at the start of each fold.  This is not necessary, as you could just set</code>dataloader.dataset.use_cache=True` at the end of the first Epoch.  I just like to get it out of the way.</p>\n\n<p><code>\ndef prime_cache(dataloader, name):\n    for idx, (x, y) in enumerate(dataloader, start=1):\n        print(f\"Loading Cache using {name} {idx*dataloader.batch_size}\\r\", end='')\n    print(\"\\n\")\n    dataloader.dataset.set_use_cache(use_cache=True)\n</code></p>\n\n<p><code>\n       if(epoch == 0):\n            prime_cache(train_loader, \"train_loader\")\n        print(\"Training............\\r\", end='')\n</code></p>\n\n<p>This is possible because of how Numpy stores its data.  So no copy is done when going to Numpy.  It simply uses the same contiguous pre-allocated space that the c-type created.  You could do the same with a torch Tensor, but it seems Numpy is a better format before you do your augmentations.  So the first time an image is requested, it fills the cache location.  At the end of an epoch, the cache is 100%, and you can switch the flag <code>use_cache</code> to <code>True</code> and then all subsequent requests come direct from CPU memory.</p>\n\n<p>Obviously this is taking up num_images * image_size space in your memory, and furthermore each num_worker gets a copy so that's num_images * image_size * num_workers.......just something to keep an eye on if you don't have much memory.  </p>\n\n<p>I use three loaders: train, val and test, so I have three caches.</p>\n\n<p>I am especially interested in what your basic flow is per epoch and how fast you can complete an epoch and what type of things you have done to speed up dataloading.</p>",
      "rawMarkdown": "One of the largest speed bumps in deep learning can be the dataloading.  There exists many ways to speed up dataloading.  LMDB and HFS5 are very common.  \n\nWhat are you all using to speed up your data loading?\n\nI do not use either.  \n\nPlease describe your epochs.  Each Epoch of mine takes about 2min 35sec.  I read each image, perform over a dozen augmentations, then do validation, then at the end of the fold I do validation on the OOF w/TTA 25.  That's ALOT of loading of the same data over and over (well, almost the same, for validation I have my own dataloader/dataset which is a subset of train).  I do about 15 Epochs and then I am doing 5 Folds as well, so again, LOTS of loading of the same data.\n\nWhat I do, is a basic \"cache\" in my dataset, like so:\n\n`class MelanomaDataset(Dataset):\n    def __init__(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):\n        \n        self.df = df\n        self.imfolder = imfolder\n        self.transforms = transforms\n        self.train = train\n        self.meta_features = meta_features\n        \n**        shared_array_base = mp.Array(ctypes.c_ubyte, len(self.df)*3*image_size*image_size)\n        shared_array = np.ctypeslib.as_array(shared_array_base.get_obj())\n        self.shared_array = shared_array.reshape(len(self.df), image_size, image_size, 3)\n        self.use_cache = False**\n        \n    def __getitem__(self, index):\n      **  if not self.use_cache:\n            im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n            image = cv2.imread(im_path)\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            self.shared_array[index] = image\n        image = self.shared_array[index] **\n            \n        if self.transforms:\n            augmented = self.transforms(image=image)\n            image = augmented['image']\n            \n        meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n        if self.train:\n            y = self.df.iloc[index]['target']\n            return (image, meta), y\n        else:\n            return (image, meta)\n    \n**    def set_use_cache(self, use_cache):\n        self.use_cache = use_cache**\n    \n    def __len__(self):\n        return len(self.df)\n\n`\nI then prime the cache(s) at the start of each fold.  This is not necessary, as you could just set `dataloader.dataset.use_cache=True` at the end of the first Epoch.  I just like to get it out of the way.\n\n```\ndef prime_cache(dataloader, name):\n    for idx, (x, y) in enumerate(dataloader, start=1):\n        print(f\"Loading Cache using {name} {idx*dataloader.batch_size}\\r\", end='')\n    print(\"\\n\")\n    dataloader.dataset.set_use_cache(use_cache=True)\n```\n\n\n```\n       if(epoch == 0):\n            prime_cache(train_loader, \"train_loader\")\n        print(\"Training............\\r\", end='')\n```\n\nThis is possible because of how Numpy stores its data.  So no copy is done when going to Numpy.  It simply uses the same contiguous pre-allocated space that the c-type created.  You could do the same with a torch Tensor, but it seems Numpy is a better format before you do your augmentations.  So the first time an image is requested, it fills the cache location.  At the end of an epoch, the cache is 100%, and you can switch the flag `use_cache` to `True` and then all subsequent requests come direct from CPU memory.\n\nObviously this is taking up num_images * image_size space in your memory, and furthermore each num_worker gets a copy so that's num_images * image_size * num_workers.......just something to keep an eye on if you don't have much memory.  \n\nI use three loaders: train, val and test, so I have three caches.\n\nI am especially interested in what your basic flow is per epoch and how fast you can complete an epoch and what type of things you have done to speed up dataloading.",
      "votes": null
    },
    {
      "id": "953887",
      "postDate": "08/01/2020 08:08:58",
      "content": "<p>Do you really need to do 25 times tta on your validation set at every epoch?\nI personnally perform TTA once on validation set when my model has converged in order to get my OOF CV, I feel like you could save a lot of time by doing this, without hurting your early stoping scheme.</p>",
      "rawMarkdown": "Do you really need to do 25 times tta on your validation set at every epoch?\nI personnally perform TTA once on validation set when my model has converged in order to get my OOF CV, I feel like you could save a lot of time by doing this, without hurting your early stoping scheme.",
      "votes": null
    },
    {
      "id": "953901",
      "postDate": "08/01/2020 08:26:30",
      "content": "<p>Yes you are right. I do TTA 25 for OOF........I was doing every epoch to test some things, as I calculate AUC/ROC every epoch, so I can see how certain things are working.  But yes, normally I don't do TTA on each epoch, just on OOF (1 per fold) and test (1 per fold)</p>",
      "rawMarkdown": "Yes you are right. I do TTA 25 for OOF........I was doing every epoch to test some things, as I calculate AUC/ROC every epoch, so I can see how certain things are working.  But yes, normally I don't do TTA on each epoch, just on OOF (1 per fold) and test (1 per fold)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 953887,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "08/01/2020 08:08:58",
      "content": "<p>Do you really need to do 25 times tta on your validation set at every epoch?\nI personnally perform TTA once on validation set when my model has converged in order to get my OOF CV, I feel like you could save a lot of time by doing this, without hurting your early stoping scheme.</p>",
      "votes": null,
      "replies": [
        {
          "id": 953901,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/01/2020 08:26:30",
          "content": "<p>Yes you are right. I do TTA 25 for OOF........I was doing every epoch to test some things, as I calculate AUC/ROC every epoch, so I can see how certain things are working.  But yes, normally I don't do TTA on each epoch, just on OOF (1 per fold) and test (1 per fold)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "953838": "One of the largest speed bumps in deep learning can be the dataloading.  There exists many ways to speed up dataloading.  LMDB and HFS5 are very common.  \n\nWhat are you all using to speed up your data loading?\n\nI do not use either.  \n\nPlease describe your epochs.  Each Epoch of mine takes about 2min 35sec.  I read each image, perform over a dozen augmentations, then do validation, then at the end of the fold I do validation on the OOF w/TTA 25.  That's ALOT of loading of the same data over and over (well, almost the same, for validation I have my own dataloader/dataset which is a subset of train).  I do about 15 Epochs and then I am doing 5 Folds as well, so again, LOTS of loading of the same data.\n\nWhat I do, is a basic \"cache\" in my dataset, like so:\n\n`class MelanomaDataset(Dataset):\n    def __init__(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):\n        \n        self.df = df\n        self.imfolder = imfolder\n        self.transforms = transforms\n        self.train = train\n        self.meta_features = meta_features\n        \n**        shared_array_base = mp.Array(ctypes.c_ubyte, len(self.df)*3*image_size*image_size)\n        shared_array = np.ctypeslib.as_array(shared_array_base.get_obj())\n        self.shared_array = shared_array.reshape(len(self.df), image_size, image_size, 3)\n        self.use_cache = False**\n        \n    def __getitem__(self, index):\n      **  if not self.use_cache:\n            im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n            image = cv2.imread(im_path)\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            self.shared_array[index] = image\n        image = self.shared_array[index] **\n            \n        if self.transforms:\n            augmented = self.transforms(image=image)\n            image = augmented['image']\n            \n        meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n        if self.train:\n            y = self.df.iloc[index]['target']\n            return (image, meta), y\n        else:\n            return (image, meta)\n    \n**    def set_use_cache(self, use_cache):\n        self.use_cache = use_cache**\n    \n    def __len__(self):\n        return len(self.df)\n\n`\nI then prime the cache(s) at the start of each fold.  This is not necessary, as you could just set `dataloader.dataset.use_cache=True` at the end of the first Epoch.  I just like to get it out of the way.\n\n```\ndef prime_cache(dataloader, name):\n    for idx, (x, y) in enumerate(dataloader, start=1):\n        print(f\"Loading Cache using {name} {idx*dataloader.batch_size}\\r\", end='')\n    print(\"\\n\")\n    dataloader.dataset.set_use_cache(use_cache=True)\n```\n\n\n```\n       if(epoch == 0):\n            prime_cache(train_loader, \"train_loader\")\n        print(\"Training............\\r\", end='')\n```\n\nThis is possible because of how Numpy stores its data.  So no copy is done when going to Numpy.  It simply uses the same contiguous pre-allocated space that the c-type created.  You could do the same with a torch Tensor, but it seems Numpy is a better format before you do your augmentations.  So the first time an image is requested, it fills the cache location.  At the end of an epoch, the cache is 100%, and you can switch the flag `use_cache` to `True` and then all subsequent requests come direct from CPU memory.\n\nObviously this is taking up num_images * image_size space in your memory, and furthermore each num_worker gets a copy so that's num_images * image_size * num_workers.......just something to keep an eye on if you don't have much memory.  \n\nI use three loaders: train, val and test, so I have three caches.\n\nI am especially interested in what your basic flow is per epoch and how fast you can complete an epoch and what type of things you have done to speed up dataloading.",
    "953887": "Do you really need to do 25 times tta on your validation set at every epoch?\nI personnally perform TTA once on validation set when my model has converged in order to get my OOF CV, I feel like you could save a lot of time by doing this, without hurting your early stoping scheme.",
    "953901": "Yes you are right. I do TTA 25 for OOF........I was doing every epoch to test some things, as I calculate AUC/ROC every epoch, so I can see how certain things are working.  But yes, normally I don't do TTA on each epoch, just on OOF (1 per fold) and test (1 per fold)"
  },
  "source": "meta"
}