{
  "id": 172904,
  "title": "Turbo charging data loading using c-types cache",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172904",
  "author_name": "",
  "post_date": "2020-08-06T22:28:57.616118700Z",
  "votes": 7,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Data loading can be one of the more time-consuming parts of a deep learning workflow.  One technique I have been doing which has drastically reduced the time it takes in my workflows is loading data from memory cache instead of disk.</p>\n\n<p>Normally with data loading, you read from disk, each image, each time.  If you have 10 epochs,  5 folds training, validation, validation for OOF, and test:</p>\n\n<p>Example: 10,000 training images, 1,000 test images, 10 epochs, TTA 15</p>\n\n<p>Training:    Read 10,000 images x 10 epochs x 5 folds =   500,000 images from disk\nValidation: Read 10,000 images x 10 epochs x 5 folds =   500,000 images from disk\nOOF:          Read 10,000 images x TTA 15 x 5 folds      =    750,000 images from disk\nTest:          Read 1,000   images x  TTA 15 x 5 folds     =      75,000 images from disk\nTotal images read from disk:                                          = 1,825,000 images from disk                   </p>\n\n<p>If we can read the data into cache, we can re-read it lightening fast and save a lot of expensive I/O.  </p>\n\n<p>Example using cache: 10,000 training images, 1,000 test images, 10 epochs, TTA 15</p>\n\n<p>Training:    Read 10,000 images x 5 folds        =   50,000 images from disk\nValidation: Read 10,000 images x 5 folds        =   50,000 images from disk\nOOF:          Read 10,000 images x 5 folds        =   50,000 images from disk\nTest:          Read    1,000 images x 5 folds        =     5,000 images from disk\nTotal images read from disk:                             = 155,000 images from disk     </p>\n\n<p>How do we read from cache? It's actually not that difficult.  Here is a Dataset in PyTorch before caching:</p>\n\n<p>```\nclass MelanomaDataset(Dataset):\n    def <strong>init</strong>(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):</p>\n\n<pre><code>    self.df = df\n    self.imfolder = imfolder\n    self.transforms = transforms\n    self.train = train\n    self.meta_features = meta_features\n\ndef __getitem__(self, index):\n     im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n     image = cv2.imread(im_path)\n     image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n\n    if self.transforms:\n        augmented = self.transforms(image=image)\n        image = augmented['image']\n\n    meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n    if self.train:\n        y = self.df.iloc[index]['target']\n        return image, meta, y\n    else:\n        return image, meta\n\ndef __len__(self):\n    return len(self.df)\n</code></pre>\n\n<p>```</p>\n\n<p>Here is a Dataset with caching:</p>\n\n<p>```\nclass MelanomaDataset(Dataset):\n    def <strong>init</strong>(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):</p>\n\n<pre><code>    self.df = df\n    self.imfolder = imfolder\n    self.transforms = transforms\n    self.train = train\n    self.meta_features = meta_features\n\n    if(param['cache_on']):\n        shared_array_base = mp.Array(ctypes.c_ubyte, len(self.df)*3*image_size*image_size)\n        shared_array = np.ctypeslib.as_array(shared_array_base.get_obj())\n        self.shared_array = shared_array.reshape(len(self.df), image_size, image_size, 3)\n        self.use_cache = False\n        print(\"Cache initialized\")\n\ndef __getitem__(self, index):\n    if(param['cache_on']):\n        if not self.use_cache:\n            im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n            image = cv2.imread(im_path)\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            self.shared_array[index] = image\n        image = self.shared_array[index]\n    else:\n        im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n        image = cv2.imread(im_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n\n    if self.transforms:\n        augmented = self.transforms(image=image)\n        image = augmented['image']\n\n    meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n    if self.train:\n        y = self.df.iloc[index]['target']\n        return image, meta, y\n    else:\n        return (image, meta)\n\ndef set_use_cache(self, use_cache):\n    self.use_cache = use_cache\n\ndef __len__(self):\n    return len(self.df)\n</code></pre>\n\n<p>```</p>\n\n<p>A little bit about what is happening here.  We allocate memory using c-types.  This is because, under the hood, Numpy and Pytorch (and likely other frameworks) actually use c-types for storage.  This way, when you move the data from c-types storage to Numpy or a Pytorch tensor, there is no copying of data, it uses the data in the same exact location.  If anything is re-arranged, simply a new \"view\" into the data is created.  We pre-allocate this data structure, so we must know the image size beforehand.  </p>\n\n<p><code>use_cache</code> is defaulted to <code>False</code>.  So the first iteration of your train/val/oof/test loops it will not use the cache, and instead  it will be priming the cache.  On the next iterations forward (<code>epoch == 1</code>), you will then start to pull from the cache.</p>\n\n<p><code>\nfor epoch in range(param['epochs'][fold - 1]):\n        if(param['cache_on']):\n            if(epoch == 1):\n                 train_loader.dataset.set_use_cache(use_cache=True)\n</code></p>\n\n<p>Now you just saved more than 1.5M disk file requests at the cost of some CPU memory.  Every epoch will be faster.  Your images/sec numbers will go way up.  Obviously the first epoch, first TTA, etc.  Will be just as slow as it was before, as it primes the cache.  After that you are in Turbo mode.  </p>\n\n<p>Please post here if you used this and how it has speeded up your I/O!</p>",
  "messages": [
    {
      "id": "961078",
      "postDate": "08/06/2020 22:28:57",
      "content": "<p>Data loading can be one of the more time-consuming parts of a deep learning workflow.  One technique I have been doing which has drastically reduced the time it takes in my workflows is loading data from memory cache instead of disk.</p>\n\n<p>Normally with data loading, you read from disk, each image, each time.  If you have 10 epochs,  5 folds training, validation, validation for OOF, and test:</p>\n\n<p>Example: 10,000 training images, 1,000 test images, 10 epochs, TTA 15</p>\n\n<p>Training:    Read 10,000 images x 10 epochs x 5 folds =   500,000 images from disk\nValidation: Read 10,000 images x 10 epochs x 5 folds =   500,000 images from disk\nOOF:          Read 10,000 images x TTA 15 x 5 folds      =    750,000 images from disk\nTest:          Read 1,000   images x  TTA 15 x 5 folds     =      75,000 images from disk\nTotal images read from disk:                                          = 1,825,000 images from disk                   </p>\n\n<p>If we can read the data into cache, we can re-read it lightening fast and save a lot of expensive I/O.  </p>\n\n<p>Example using cache: 10,000 training images, 1,000 test images, 10 epochs, TTA 15</p>\n\n<p>Training:    Read 10,000 images x 5 folds        =   50,000 images from disk\nValidation: Read 10,000 images x 5 folds        =   50,000 images from disk\nOOF:          Read 10,000 images x 5 folds        =   50,000 images from disk\nTest:          Read    1,000 images x 5 folds        =     5,000 images from disk\nTotal images read from disk:                             = 155,000 images from disk     </p>\n\n<p>How do we read from cache? It's actually not that difficult.  Here is a Dataset in PyTorch before caching:</p>\n\n<p>```\nclass MelanomaDataset(Dataset):\n    def <strong>init</strong>(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):</p>\n\n<pre><code>    self.df = df\n    self.imfolder = imfolder\n    self.transforms = transforms\n    self.train = train\n    self.meta_features = meta_features\n\ndef __getitem__(self, index):\n     im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n     image = cv2.imread(im_path)\n     image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n\n    if self.transforms:\n        augmented = self.transforms(image=image)\n        image = augmented['image']\n\n    meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n    if self.train:\n        y = self.df.iloc[index]['target']\n        return image, meta, y\n    else:\n        return image, meta\n\ndef __len__(self):\n    return len(self.df)\n</code></pre>\n\n<p>```</p>\n\n<p>Here is a Dataset with caching:</p>\n\n<p>```\nclass MelanomaDataset(Dataset):\n    def <strong>init</strong>(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):</p>\n\n<pre><code>    self.df = df\n    self.imfolder = imfolder\n    self.transforms = transforms\n    self.train = train\n    self.meta_features = meta_features\n\n    if(param['cache_on']):\n        shared_array_base = mp.Array(ctypes.c_ubyte, len(self.df)*3*image_size*image_size)\n        shared_array = np.ctypeslib.as_array(shared_array_base.get_obj())\n        self.shared_array = shared_array.reshape(len(self.df), image_size, image_size, 3)\n        self.use_cache = False\n        print(\"Cache initialized\")\n\ndef __getitem__(self, index):\n    if(param['cache_on']):\n        if not self.use_cache:\n            im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n            image = cv2.imread(im_path)\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            self.shared_array[index] = image\n        image = self.shared_array[index]\n    else:\n        im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n        image = cv2.imread(im_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n\n    if self.transforms:\n        augmented = self.transforms(image=image)\n        image = augmented['image']\n\n    meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n    if self.train:\n        y = self.df.iloc[index]['target']\n        return image, meta, y\n    else:\n        return (image, meta)\n\ndef set_use_cache(self, use_cache):\n    self.use_cache = use_cache\n\ndef __len__(self):\n    return len(self.df)\n</code></pre>\n\n<p>```</p>\n\n<p>A little bit about what is happening here.  We allocate memory using c-types.  This is because, under the hood, Numpy and Pytorch (and likely other frameworks) actually use c-types for storage.  This way, when you move the data from c-types storage to Numpy or a Pytorch tensor, there is no copying of data, it uses the data in the same exact location.  If anything is re-arranged, simply a new \"view\" into the data is created.  We pre-allocate this data structure, so we must know the image size beforehand.  </p>\n\n<p><code>use_cache</code> is defaulted to <code>False</code>.  So the first iteration of your train/val/oof/test loops it will not use the cache, and instead  it will be priming the cache.  On the next iterations forward (<code>epoch == 1</code>), you will then start to pull from the cache.</p>\n\n<p><code>\nfor epoch in range(param['epochs'][fold - 1]):\n        if(param['cache_on']):\n            if(epoch == 1):\n                 train_loader.dataset.set_use_cache(use_cache=True)\n</code></p>\n\n<p>Now you just saved more than 1.5M disk file requests at the cost of some CPU memory.  Every epoch will be faster.  Your images/sec numbers will go way up.  Obviously the first epoch, first TTA, etc.  Will be just as slow as it was before, as it primes the cache.  After that you are in Turbo mode.  </p>\n\n<p>Please post here if you used this and how it has speeded up your I/O!</p>",
      "rawMarkdown": "Data loading can be one of the more time-consuming parts of a deep learning workflow.  One technique I have been doing which has drastically reduced the time it takes in my workflows is loading data from memory cache instead of disk.\n\nNormally with data loading, you read from disk, each image, each time.  If you have 10 epochs,  5 folds training, validation, validation for OOF, and test:\n\nExample: 10,000 training images, 1,000 test images, 10 epochs, TTA 15\n\nTraining:    Read 10,000 images x 10 epochs x 5 folds =   500,000 images from disk\nValidation: Read 10,000 images x 10 epochs x 5 folds =   500,000 images from disk\nOOF:          Read 10,000 images x TTA 15 x 5 folds      =    750,000 images from disk\nTest:          Read 1,000   images x  TTA 15 x 5 folds     =      75,000 images from disk\nTotal images read from disk:                                          = 1,825,000 images from disk                   \n\nIf we can read the data into cache, we can re-read it lightening fast and save a lot of expensive I/O.  \n\nExample using cache: 10,000 training images, 1,000 test images, 10 epochs, TTA 15\n\nTraining:    Read 10,000 images x 5 folds        =   50,000 images from disk\nValidation: Read 10,000 images x 5 folds        =   50,000 images from disk\nOOF:          Read 10,000 images x 5 folds        =   50,000 images from disk\nTest:          Read    1,000 images x 5 folds        =     5,000 images from disk\nTotal images read from disk:                             = 155,000 images from disk     \n\nHow do we read from cache? It's actually not that difficult.  Here is a Dataset in PyTorch before caching:\n\n```\nclass MelanomaDataset(Dataset):\n    def __init__(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):\n        \n        self.df = df\n        self.imfolder = imfolder\n        self.transforms = transforms\n        self.train = train\n        self.meta_features = meta_features\n                \n    def __getitem__(self, index):\n         im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n         image = cv2.imread(im_path)\n         image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            \n        if self.transforms:\n            augmented = self.transforms(image=image)\n            image = augmented['image']\n            \n        meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n        if self.train:\n            y = self.df.iloc[index]['target']\n            return image, meta, y\n        else:\n            return image, meta\n   \n    def __len__(self):\n        return len(self.df)\n```\n\nHere is a Dataset with caching:\n\n```\nclass MelanomaDataset(Dataset):\n    def __init__(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):\n        \n        self.df = df\n        self.imfolder = imfolder\n        self.transforms = transforms\n        self.train = train\n        self.meta_features = meta_features\n        \n        if(param['cache_on']):\n            shared_array_base = mp.Array(ctypes.c_ubyte, len(self.df)*3*image_size*image_size)\n            shared_array = np.ctypeslib.as_array(shared_array_base.get_obj())\n            self.shared_array = shared_array.reshape(len(self.df), image_size, image_size, 3)\n            self.use_cache = False\n            print(\"Cache initialized\")\n        \n    def __getitem__(self, index):\n        if(param['cache_on']):\n            if not self.use_cache:\n                im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n                image = cv2.imread(im_path)\n                image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n                self.shared_array[index] = image\n            image = self.shared_array[index]\n        else:\n            im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n            image = cv2.imread(im_path)\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            \n        if self.transforms:\n            augmented = self.transforms(image=image)\n            image = augmented['image']\n            \n        meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n        if self.train:\n            y = self.df.iloc[index]['target']\n            return image, meta, y\n        else:\n            return (image, meta)\n    \n    def set_use_cache(self, use_cache):\n        self.use_cache = use_cache\n    \n    def __len__(self):\n        return len(self.df)\n```\n\nA little bit about what is happening here.  We allocate memory using c-types.  This is because, under the hood, Numpy and Pytorch (and likely other frameworks) actually use c-types for storage.  This way, when you move the data from c-types storage to Numpy or a Pytorch tensor, there is no copying of data, it uses the data in the same exact location.  If anything is re-arranged, simply a new \"view\" into the data is created.  We pre-allocate this data structure, so we must know the image size beforehand.  \n\n`use_cache` is defaulted to `False`.  So the first iteration of your train/val/oof/test loops it will not use the cache, and instead  it will be priming the cache.  On the next iterations forward (`epoch == 1`), you will then start to pull from the cache.\n\n```\nfor epoch in range(param['epochs'][fold - 1]):\n        if(param['cache_on']):\n            if(epoch == 1):\n                 train_loader.dataset.set_use_cache(use_cache=True)\n```\n        \nNow you just saved more than 1.5M disk file requests at the cost of some CPU memory.  Every epoch will be faster.  Your images/sec numbers will go way up.  Obviously the first epoch, first TTA, etc.  Will be just as slow as it was before, as it primes the cache.  After that you are in Turbo mode.  \n\nPlease post here if you used this and how it has speeded up your I/O!",
      "votes": null
    },
    {
      "id": "961109",
      "postDate": "08/06/2020 23:25:27",
      "content": "<p>I calculate duration of each epoch and first epoch always takes much more time because data is read from disk, all other epochs are faster because they are read from cache of operating system.</p>",
      "rawMarkdown": "I calculate duration of each epoch and first epoch always takes much more time because data is read from disk, all other epochs are faster because they are read from cache of operating system.",
      "votes": null
    },
    {
      "id": "961111",
      "postDate": "08/06/2020 23:31:57",
      "content": "<p>Yes, first epoch, or first TTA, will take as long as it normally does reading from disk.  Also, the instantiation of the Dataset requires the allocation of the data structure, which takes a few seconds.........and of course takes a little longer depending on image size.  Because with Epochs, Folds and TTA we are re-iterating so many times, that is where the savings come in.  The more epochs, folds a and TTA you do, the more savings.</p>",
      "rawMarkdown": "Yes, first epoch, or first TTA, will take as long as it normally does reading from disk.  Also, the instantiation of the Dataset requires the allocation of the data structure, which takes a few seconds.........and of course takes a little longer depending on image size.  Because with Epochs, Folds and TTA we are re-iterating so many times, that is where the savings come in.  The more epochs, folds a and TTA you do, the more savings.",
      "votes": null
    },
    {
      "id": "961112",
      "postDate": "08/06/2020 23:38:09",
      "content": "<p>my point is - you don't need to store data in memory because it is already stored by operating system cache</p>",
      "rawMarkdown": "my point is - you don't need to store data in memory because it is already stored by operating system cache",
      "votes": null
    },
    {
      "id": "961145",
      "postDate": "08/07/2020 00:17:04",
      "content": "<p>Epoch 0 is 2min 25sec (loading cache)\nEpoch 1 - 18 (or more, depending on when Early Stopping hits) is 2min 10sec</p>\n\n<p>So a savings of 15 seconds per epoch.  So in my case that's:</p>\n\n<p>Training: 15 seconds per epoch x 17 epochs x 5 folds    = 1275 seconds\nValidation: 15 seconds per epoch x 17 epochs x 5 folds = 1275 seconds\nOOF: 3 seconds per TTA x TTA 14 x 5 folds                    =   210 seconds\nTest: 1.5 seconds per TTA x TTA 14 * 5 folds                   =   105 seconds \nTotal saved:                                                                         = 2865 seconds = 47.75 minutes</p>",
      "rawMarkdown": "Epoch 0 is 2min 25sec (loading cache)\nEpoch 1 - 18 (or more, depending on when Early Stopping hits) is 2min 10sec\n\nSo a savings of 15 seconds per epoch.  So in my case that's:\n\nTraining: 15 seconds per epoch x 17 epochs x 5 folds    = 1275 seconds\nValidation: 15 seconds per epoch x 17 epochs x 5 folds = 1275 seconds\nOOF: 3 seconds per TTA x TTA 14 x 5 folds                    =   210 seconds\nTest: 1.5 seconds per TTA x TTA 14 * 5 folds                   =   105 seconds \nTotal saved:                                                                         = 2865 seconds = 47.75 minutes",
      "votes": null
    },
    {
      "id": "961149",
      "postDate": "08/07/2020 00:26:34",
      "content": "<p>epoch 0<br>\n{'train<em>score': 0.5907124039945352, 'valid</em>score': 0.6284858871924628, 'train<em>loss': 0.19184398562075033, 'valid</em>loss': 0.08261917846606058, 'duration': 772.745526, 'lr': 0.00020902725541536234}<br>\nsaving model<br>\nepoch 1<br>\n{'train<em>score': 0.7463795335384227, 'valid</em>score': 0.8594380617329083, 'train<em>loss': 0.0882036334134919, 'valid</em>loss': 0.06833077350302655, 'duration': 449.452999, 'lr': 0.0005614256203280725}<br>\nsaving model</p>",
      "rawMarkdown": "epoch 0\n{'train_score': 0.5907124039945352, 'valid_score': 0.6284858871924628, 'train_loss': 0.19184398562075033, 'valid_loss': 0.08261917846606058, 'duration': 772.745526, 'lr': 0.00020902725541536234}\nsaving model\nepoch 1\n{'train_score': 0.7463795335384227, 'valid_score': 0.8594380617329083, 'train_loss': 0.0882036334134919, 'valid_loss': 0.06833077350302655, 'duration': 449.452999, 'lr': 0.0005614256203280725}\nsaving model",
      "votes": null
    },
    {
      "id": "961182",
      "postDate": "08/07/2020 01:23:57",
      "content": "<p>Is that 772 and 449 in seconds? I guess what you are showing is that without external caching, just your OS, your epoch time drops from the OS caching.  I am surprised the difference is over 300 seconds, I mean, you must have very slow image/sec ingest to have that kind of savings. </p>",
      "rawMarkdown": "Is that 772 and 449 in seconds? I guess what you are showing is that without external caching, just your OS, your epoch time drops from the OS caching.  I am surprised the difference is over 300 seconds, I mean, you must have very slow image/sec ingest to have that kind of savings.",
      "votes": null
    },
    {
      "id": "962669",
      "postDate": "08/08/2020 10:26:29",
      "content": "<p>Why not using a numpy array directly?  That's how I implement my data loader.  Just curious to know if you gain anything using ctype instead of a numpy array.</p>",
      "rawMarkdown": "Why not using a numpy array directly?  That's how I implement my data loader.  Just curious to know if you gain anything using ctype instead of a numpy array.",
      "votes": null
    },
    {
      "id": "962745",
      "postDate": "08/08/2020 12:01:30",
      "content": "<p>I am creating a single array, which can be accessed by PyTorches \"workers\".  PyTorches \"workers\" (set by <code>num_workers</code> &gt; 0 in the data loader) use Python multi-processing.  This allows you to setup  a data structure (<code>mp.Array</code>) which is shared between processes.  NumPy arrays would be instantiated for every worker, as each worker clones the dataset, and thus create each their own underlying storage in memory, which is not what you would want to do.</p>\n\n<p><a href=\"https://docs.python.org/2/library/multiprocessing.html\">https://docs.python.org/2/library/multiprocessing.html</a></p>",
      "rawMarkdown": "I am creating a single array, which can be accessed by PyTorches \"workers\".  PyTorches \"workers\" (set by `num_workers` &gt; 0 in the data loader) use Python multi-processing.  This allows you to setup  a data structure (`mp.Array`) which is shared between processes.  NumPy arrays would be instantiated for every worker, as each worker clones the dataset, and thus create each their own underlying storage in memory, which is not what you would want to do.\n\nhttps://docs.python.org/2/library/multiprocessing.html",
      "votes": null
    },
    {
      "id": "962748",
      "postDate": "08/08/2020 12:04:33",
      "content": "<p>these are numbers from Colab notebook, this competition, 256x256 images</p>",
      "rawMarkdown": "these are numbers from Colab notebook, this competition, 256x256 images",
      "votes": null
    },
    {
      "id": "964910",
      "postDate": "08/10/2020 09:00:19",
      "content": "<p>Awesome, thanks.  IO isn't better for me as I already cached values and saved the array to disk. However footprint is definitely smaller when sharing the array across workers.  It will pay off when I'll move beyond 256x256 images hopefully.  Thanks.</p>\n<p>Not sure why you don't get more upvotes on this.</p>",
      "rawMarkdown": "Awesome, thanks.  IO isn't better for me as I already cached values and saved the array to disk. However footprint is definitely smaller when sharing the array across workers.  It will pay off when I'll move beyond 256x256 images hopefully.  Thanks.\n\nNot sure why you don't get more upvotes on this.",
      "votes": null
    },
    {
      "id": "964990",
      "postDate": "08/10/2020 10:03:34",
      "content": "<p>I think either it's hard to understand what is happening, or I do a bad job of explaining it, lol. </p>",
      "rawMarkdown": "I think either it's hard to understand what is happening, or I do a bad job of explaining it, lol.",
      "votes": null
    },
    {
      "id": "964996",
      "postDate": "08/10/2020 10:09:00",
      "content": "<p>A nice side effect: I was thinking of using hdf5 as loading the array for each worker would not have been possible with large image sizes.  You made this irrelevant.  Thanks again.</p>",
      "rawMarkdown": "A nice side effect: I was thinking of using hdf5 as loading the array for each worker would not have been possible with large image sizes.  You made this irrelevant.  Thanks again.",
      "votes": null
    },
    {
      "id": "965005",
      "postDate": "08/10/2020 10:16:12",
      "content": "<p>I assumed requirement for your approach is to fit whole dataset into memory, this competition has large dataset</p>",
      "rawMarkdown": "I assumed requirement for your approach is to fit whole dataset into memory, this competition has large dataset",
      "votes": null
    },
    {
      "id": "965036",
      "postDate": "08/10/2020 10:43:13",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> yes it has to go into CPU memory.  So you would need about 1.4GB of CPU memory to store train data at 256x256.  6.9GB of CPU memory for train data at 768x768.  Most people have 32GB, or 64GB, or even 128GB of CPU memory, and PyTorch is probably only using a small fraction of that.  If you are running on a system with say 8-16GB of memory then you need to be careful for sure.</p>",
      "rawMarkdown": "jacekpoplawski yes it has to go into CPU memory.  So you would need about 1.4GB of CPU memory to store train data at 256x256.  6.9GB of CPU memory for train data at 768x768.  Most people have 32GB, or 64GB, or even 128GB of CPU memory, and PyTorch is probably only using a small fraction of that.  If you are running on a system with say 8-16GB of memory then you need to be careful for sure.",
      "votes": null
    },
    {
      "id": "965044",
      "postDate": "08/10/2020 10:52:36",
      "content": "<p>I'm afraid that storing train data at 256x256 uses 6.5GB.</p>\n<p>3×256×256×33126 = 6512836608</p>",
      "rawMarkdown": "I'm afraid that storing train data at 256x256 uses 6.5GB.\n\n3×256×256×33126 = 6512836608",
      "votes": null
    },
    {
      "id": "965086",
      "postDate": "08/10/2020 11:24:56",
      "content": "<p>you're right!  I naively just looked at the size of my train directory lol.....I have enough memory that I don't think about it.......but yeah, no doubt with increased image size it can become an issue, but also with increased image size so does disk I/O!  And of course, you have to store validation and test as well.  A lot of systems should have the memory to work with 256 or 384.......or beyond.........some don't.  But if you have the memory I say use it!</p>",
      "rawMarkdown": "you're right!  I naively just looked at the size of my train directory lol.....I have enough memory that I don't think about it.......but yeah, no doubt with increased image size it can become an issue, but also with increased image size so does disk I/O!  And of course, you have to store validation and test as well.  A lot of systems should have the memory to work with 256 or 384.......or beyond.........some don't.  But if you have the memory I say use it!",
      "votes": null
    },
    {
      "id": "965127",
      "postDate": "08/10/2020 12:00:52",
      "content": "<p>Probably I will not be able to get medal in this competition, but my way of working was to have more data than just 33126 images, that's why I am not able to use your idea in this case. </p>\n<p>Anyway, as I said operating system caches data from disk. Is your idea to resolve problem with multiple cores, do you mean that operating system caches each core independently, and your cache is one for all cores?</p>",
      "rawMarkdown": "Probably I will not be able to get medal in this competition, but my way of working was to have more data than just 33126 images, that's why I am not able to use your idea in this case. \n\nAnyway, as I said operating system caches data from disk. Is your idea to resolve problem with multiple cores, do you mean that operating system caches each core independently, and your cache is one for all cores?",
      "votes": null
    },
    {
      "id": "965132",
      "postDate": "08/10/2020 12:04:55",
      "content": "<p>OS file caching depends on a number of things.  I think you will see that your disk still has read access when loading files for each epoch.   In most OS's file cache is shared between processes.  In python data structures are not shared between processes, but if you instantiate them from the same data which resides in shared memory then they will effectively be sharing a data structure, which is what <code>mp.Array</code>() accomplishes.</p>",
      "rawMarkdown": "OS file caching depends on a number of things.  I think you will see that your disk still has read access when loading files for each epoch.   In most OS's file cache is shared between processes.  In python data structures are not shared between processes, but if you instantiate them from the same data which resides in shared memory then they will effectively be sharing a data structure, which is what `mp.Array`() accomplishes.",
      "votes": null
    },
    {
      "id": "965142",
      "postDate": "08/10/2020 12:11:18",
      "content": "<p>I remember when I was starting with machine learning i was always loading whole dataset, then later in Kaggle competition I learned that data can be huge and to avoid loading everything at once I learned how generator works. When I switched from Keras to PyTorch it was even easier to implemenent generator by using Dataset and DataLoader. I would probably need to make some tests to understand benefits of this idea, but have no time until this competition is over.</p>",
      "rawMarkdown": "I remember when I was starting with machine learning i was always loading whole dataset, then later in Kaggle competition I learned that data can be huge and to avoid loading everything at once I learned how generator works. When I switched from Keras to PyTorch it was even easier to implemenent generator by using Dataset and DataLoader. I would probably need to make some tests to understand benefits of this idea, but have no time until this competition is over.",
      "votes": null
    },
    {
      "id": "965872",
      "postDate": "08/10/2020 23:19:51",
      "content": "<p>Just a place holder so that I can find this post later! Thanks for sharing!</p>",
      "rawMarkdown": "Just a place holder so that I can find this post later! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "966519",
      "postDate": "08/11/2020 13:40:16",
      "content": "<p>Here is a link to my other post <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172888\">Using CUDA pin_memory and non_blocking to speed up workflows in PyTorch</a> which goes hand and hand with this post for making sure you are getting every last bit out of your GPU(s).</p>",
      "rawMarkdown": "Here is a link to my other post [Using CUDA pin_memory and non_blocking to speed up workflows in PyTorch](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172888) which goes hand and hand with this post for making sure you are getting every last bit out of your GPU(s).",
      "votes": null
    },
    {
      "id": "969746",
      "postDate": "08/13/2020 22:44:47",
      "content": "<p>The overall goal of doing this is to keep your GPU's at 100%.  If they are not at 100% then you are leaving resources on the table, you are not getting all you can out of your environment.  Once your GPU's are at 100% you can just stop trying to optimize ingest and pipeline, because you can't do any better.  I have done a number of tweaks, most of which I have posted to the forums, and now often my GPU's look like this:</p>\n<p>`root@1e5de513dd11:/workspace/data/kaggle/Melanoma/notebook# nvidia-smi<br>\nThu Aug 13 22:42:22 2020       <br>\n+-----------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 450.51.05    Driver Version: 450.51.05    CUDA Version: 11.0     |<br>\n|-------------------------------+----------------------+----------------------+<br>\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |<br>\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |<br>\n|                               |                      |               MIG M. |<br>\n|===============================+======================+======================|<br>\n|   0  GeForce GTX 108…  On   | 00000000:05:00.0  On |                  N/A |<br>\n| 53%   84C    P2    85W / 250W |  11165MiB / 11175MiB |    100%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+<br>\n|   1  GeForce GTX 108…  On   | 00000000:06:00.0 Off |                  N/A |<br>\n| 51%   85C    P2    85W / 250W |  11105MiB / 11178MiB |    100%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+<br>\n|   2  GeForce GTX 108…  On   | 00000000:09:00.0 Off |                  N/A |<br>\n| 59%   84C    P2    75W / 250W |  10309MiB / 11178MiB |    100%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+<br>\n|   3  GeForce GTX 108…  On   | 00000000:0A:00.0 Off |                  N/A |<br>\n| 48%   77C    P2    85W / 250W |   9707MiB / 11178MiB |    100%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+</p>\n<p>+-----------------------------------------------------------------------------+<br>\n| Processes:                                                                  |<br>\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |<br>\n|        ID   ID                                                   Usage      |<br>\n|=============================================================================|<br>\n+-----------------------------------------------------------------------------+`</p>\n<p>So 100% on all four.  There is some fluctuation, but its always up there.  A recap of things that will help:</p>\n<ol>\n<li>Fast disks/SSD.  Once it's cached it's irrelevant, but this helps for first epoch/TTA.</li>\n<li>Turbo mode cache as posted here</li>\n<li>Using <code>pin_memory</code> and <code>non_blocking</code> directives in CUDA <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172888\" target=\"_blank\">as described here</a></li>\n<li>Using AMP/Mixed Precision.  Why? Because the larger batch size, which means a larger/better sample which to compute certain metrics on, and not to mention its more efficient and faster when using larger batch sizes.</li>\n</ol>\n<p>Anything else to get the most out of your GPU's? Let us know!</p>",
      "rawMarkdown": "The overall goal of doing this is to keep your GPU's at 100%.  If they are not at 100% then you are leaving resources on the table, you are not getting all you can out of your environment.  Once your GPU's are at 100% you can just stop trying to optimize ingest and pipeline, because you can't do any better.  I have done a number of tweaks, most of which I have posted to the forums, and now often my GPU's look like this:\n\n`root@1e5de513dd11:/workspace/data/kaggle/Melanoma/notebook# nvidia-smi\nThu Aug 13 22:42:22 2020       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.05    Driver Version: 450.51.05    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  GeForce GTX 108...  On   | 00000000:05:00.0  On |                  N/A |\n| 53%   84C    P2    85W / 250W |  11165MiB / 11175MiB |    100%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   1  GeForce GTX 108...  On   | 00000000:06:00.0 Off |                  N/A |\n| 51%   85C    P2    85W / 250W |  11105MiB / 11178MiB |    100%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   2  GeForce GTX 108...  On   | 00000000:09:00.0 Off |                  N/A |\n| 59%   84C    P2    75W / 250W |  10309MiB / 11178MiB |    100%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   3  GeForce GTX 108...  On   | 00000000:0A:00.0 Off |                  N/A |\n| 48%   77C    P2    85W / 250W |   9707MiB / 11178MiB |    100%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n+-----------------------------------------------------------------------------+`\n\nSo 100% on all four.  There is some fluctuation, but its always up there.  A recap of things that will help:\n\n1. Fast disks/SSD.  Once it's cached it's irrelevant, but this helps for first epoch/TTA.\n2. Turbo mode cache as posted here\n3. Using `pin_memory` and `non_blocking` directives in CUDA [as described here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172888)\n4. Using AMP/Mixed Precision.  Why? Because the larger batch size, which means a larger/better sample which to compute certain metrics on, and not to mention its more efficient and faster when using larger batch sizes.\n\nAnything else to get the most out of your GPU's? Let us know!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961109,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "08/06/2020 23:25:27",
      "content": "<p>I calculate duration of each epoch and first epoch always takes much more time because data is read from disk, all other epochs are faster because they are read from cache of operating system.</p>",
      "votes": null,
      "replies": [
        {
          "id": 961111,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/06/2020 23:31:57",
          "content": "<p>Yes, first epoch, or first TTA, will take as long as it normally does reading from disk.  Also, the instantiation of the Dataset requires the allocation of the data structure, which takes a few seconds.........and of course takes a little longer depending on image size.  Because with Epochs, Folds and TTA we are re-iterating so many times, that is where the savings come in.  The more epochs, folds a and TTA you do, the more savings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961112,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/06/2020 23:38:09",
          "content": "<p>my point is - you don't need to store data in memory because it is already stored by operating system cache</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961145,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/07/2020 00:17:04",
          "content": "<p>Epoch 0 is 2min 25sec (loading cache)\nEpoch 1 - 18 (or more, depending on when Early Stopping hits) is 2min 10sec</p>\n\n<p>So a savings of 15 seconds per epoch.  So in my case that's:</p>\n\n<p>Training: 15 seconds per epoch x 17 epochs x 5 folds    = 1275 seconds\nValidation: 15 seconds per epoch x 17 epochs x 5 folds = 1275 seconds\nOOF: 3 seconds per TTA x TTA 14 x 5 folds                    =   210 seconds\nTest: 1.5 seconds per TTA x TTA 14 * 5 folds                   =   105 seconds \nTotal saved:                                                                         = 2865 seconds = 47.75 minutes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961149,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/07/2020 00:26:34",
          "content": "<p>epoch 0<br>\n{'train<em>score': 0.5907124039945352, 'valid</em>score': 0.6284858871924628, 'train<em>loss': 0.19184398562075033, 'valid</em>loss': 0.08261917846606058, 'duration': 772.745526, 'lr': 0.00020902725541536234}<br>\nsaving model<br>\nepoch 1<br>\n{'train<em>score': 0.7463795335384227, 'valid</em>score': 0.8594380617329083, 'train<em>loss': 0.0882036334134919, 'valid</em>loss': 0.06833077350302655, 'duration': 449.452999, 'lr': 0.0005614256203280725}<br>\nsaving model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961182,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/07/2020 01:23:57",
          "content": "<p>Is that 772 and 449 in seconds? I guess what you are showing is that without external caching, just your OS, your epoch time drops from the OS caching.  I am surprised the difference is over 300 seconds, I mean, you must have very slow image/sec ingest to have that kind of savings. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962748,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/08/2020 12:04:33",
          "content": "<p>these are numbers from Colab notebook, this competition, 256x256 images</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 962669,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/08/2020 10:26:29",
      "content": "<p>Why not using a numpy array directly?  That's how I implement my data loader.  Just curious to know if you gain anything using ctype instead of a numpy array.</p>",
      "votes": null,
      "replies": [
        {
          "id": 962745,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/08/2020 12:01:30",
          "content": "<p>I am creating a single array, which can be accessed by PyTorches \"workers\".  PyTorches \"workers\" (set by <code>num_workers</code> &gt; 0 in the data loader) use Python multi-processing.  This allows you to setup  a data structure (<code>mp.Array</code>) which is shared between processes.  NumPy arrays would be instantiated for every worker, as each worker clones the dataset, and thus create each their own underlying storage in memory, which is not what you would want to do.</p>\n\n<p><a href=\"https://docs.python.org/2/library/multiprocessing.html\">https://docs.python.org/2/library/multiprocessing.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964910,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 09:00:19",
          "content": "<p>Awesome, thanks.  IO isn't better for me as I already cached values and saved the array to disk. However footprint is definitely smaller when sharing the array across workers.  It will pay off when I'll move beyond 256x256 images hopefully.  Thanks.</p>\n<p>Not sure why you don't get more upvotes on this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964990,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 10:03:34",
          "content": "<p>I think either it's hard to understand what is happening, or I do a bad job of explaining it, lol. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964996,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 10:09:00",
          "content": "<p>A nice side effect: I was thinking of using hdf5 as loading the array for each worker would not have been possible with large image sizes.  You made this irrelevant.  Thanks again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965005,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 10:16:12",
          "content": "<p>I assumed requirement for your approach is to fit whole dataset into memory, this competition has large dataset</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965036,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 10:43:13",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> yes it has to go into CPU memory.  So you would need about 1.4GB of CPU memory to store train data at 256x256.  6.9GB of CPU memory for train data at 768x768.  Most people have 32GB, or 64GB, or even 128GB of CPU memory, and PyTorch is probably only using a small fraction of that.  If you are running on a system with say 8-16GB of memory then you need to be careful for sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965044,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 10:52:36",
          "content": "<p>I'm afraid that storing train data at 256x256 uses 6.5GB.</p>\n<p>3×256×256×33126 = 6512836608</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965086,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 11:24:56",
          "content": "<p>you're right!  I naively just looked at the size of my train directory lol.....I have enough memory that I don't think about it.......but yeah, no doubt with increased image size it can become an issue, but also with increased image size so does disk I/O!  And of course, you have to store validation and test as well.  A lot of systems should have the memory to work with 256 or 384.......or beyond.........some don't.  But if you have the memory I say use it!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965127,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 12:00:52",
          "content": "<p>Probably I will not be able to get medal in this competition, but my way of working was to have more data than just 33126 images, that's why I am not able to use your idea in this case. </p>\n<p>Anyway, as I said operating system caches data from disk. Is your idea to resolve problem with multiple cores, do you mean that operating system caches each core independently, and your cache is one for all cores?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965132,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 12:04:55",
          "content": "<p>OS file caching depends on a number of things.  I think you will see that your disk still has read access when loading files for each epoch.   In most OS's file cache is shared between processes.  In python data structures are not shared between processes, but if you instantiate them from the same data which resides in shared memory then they will effectively be sharing a data structure, which is what <code>mp.Array</code>() accomplishes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965142,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 12:11:18",
          "content": "<p>I remember when I was starting with machine learning i was always loading whole dataset, then later in Kaggle competition I learned that data can be huge and to avoid loading everything at once I learned how generator works. When I switched from Keras to PyTorch it was even easier to implemenent generator by using Dataset and DataLoader. I would probably need to make some tests to understand benefits of this idea, but have no time until this competition is over.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 965872,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "08/10/2020 23:19:51",
      "content": "<p>Just a place holder so that I can find this post later! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 969746,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "08/13/2020 22:44:47",
      "content": "<p>The overall goal of doing this is to keep your GPU's at 100%.  If they are not at 100% then you are leaving resources on the table, you are not getting all you can out of your environment.  Once your GPU's are at 100% you can just stop trying to optimize ingest and pipeline, because you can't do any better.  I have done a number of tweaks, most of which I have posted to the forums, and now often my GPU's look like this:</p>\n<p>`root@1e5de513dd11:/workspace/data/kaggle/Melanoma/notebook# nvidia-smi<br>\nThu Aug 13 22:42:22 2020       <br>\n+-----------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 450.51.05    Driver Version: 450.51.05    CUDA Version: 11.0     |<br>\n|-------------------------------+----------------------+----------------------+<br>\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |<br>\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |<br>\n|                               |                      |               MIG M. |<br>\n|===============================+======================+======================|<br>\n|   0  GeForce GTX 108…  On   | 00000000:05:00.0  On |                  N/A |<br>\n| 53%   84C    P2    85W / 250W |  11165MiB / 11175MiB |    100%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+<br>\n|   1  GeForce GTX 108…  On   | 00000000:06:00.0 Off |                  N/A |<br>\n| 51%   85C    P2    85W / 250W |  11105MiB / 11178MiB |    100%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+<br>\n|   2  GeForce GTX 108…  On   | 00000000:09:00.0 Off |                  N/A |<br>\n| 59%   84C    P2    75W / 250W |  10309MiB / 11178MiB |    100%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+<br>\n|   3  GeForce GTX 108…  On   | 00000000:0A:00.0 Off |                  N/A |<br>\n| 48%   77C    P2    85W / 250W |   9707MiB / 11178MiB |    100%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+</p>\n<p>+-----------------------------------------------------------------------------+<br>\n| Processes:                                                                  |<br>\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |<br>\n|        ID   ID                                                   Usage      |<br>\n|=============================================================================|<br>\n+-----------------------------------------------------------------------------+`</p>\n<p>So 100% on all four.  There is some fluctuation, but its always up there.  A recap of things that will help:</p>\n<ol>\n<li>Fast disks/SSD.  Once it's cached it's irrelevant, but this helps for first epoch/TTA.</li>\n<li>Turbo mode cache as posted here</li>\n<li>Using <code>pin_memory</code> and <code>non_blocking</code> directives in CUDA <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172888\" target=\"_blank\">as described here</a></li>\n<li>Using AMP/Mixed Precision.  Why? Because the larger batch size, which means a larger/better sample which to compute certain metrics on, and not to mention its more efficient and faster when using larger batch sizes.</li>\n</ol>\n<p>Anything else to get the most out of your GPU's? Let us know!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 966519,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "08/11/2020 13:40:16",
      "content": "<p>Here is a link to my other post <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172888\">Using CUDA pin_memory and non_blocking to speed up workflows in PyTorch</a> which goes hand and hand with this post for making sure you are getting every last bit out of your GPU(s).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "961078": "Data loading can be one of the more time-consuming parts of a deep learning workflow.  One technique I have been doing which has drastically reduced the time it takes in my workflows is loading data from memory cache instead of disk.\n\nNormally with data loading, you read from disk, each image, each time.  If you have 10 epochs,  5 folds training, validation, validation for OOF, and test:\n\nExample: 10,000 training images, 1,000 test images, 10 epochs, TTA 15\n\nTraining:    Read 10,000 images x 10 epochs x 5 folds =   500,000 images from disk\nValidation: Read 10,000 images x 10 epochs x 5 folds =   500,000 images from disk\nOOF:          Read 10,000 images x TTA 15 x 5 folds      =    750,000 images from disk\nTest:          Read 1,000   images x  TTA 15 x 5 folds     =      75,000 images from disk\nTotal images read from disk:                                          = 1,825,000 images from disk                   \n\nIf we can read the data into cache, we can re-read it lightening fast and save a lot of expensive I/O.  \n\nExample using cache: 10,000 training images, 1,000 test images, 10 epochs, TTA 15\n\nTraining:    Read 10,000 images x 5 folds        =   50,000 images from disk\nValidation: Read 10,000 images x 5 folds        =   50,000 images from disk\nOOF:          Read 10,000 images x 5 folds        =   50,000 images from disk\nTest:          Read    1,000 images x 5 folds        =     5,000 images from disk\nTotal images read from disk:                             = 155,000 images from disk     \n\nHow do we read from cache? It's actually not that difficult.  Here is a Dataset in PyTorch before caching:\n\n```\nclass MelanomaDataset(Dataset):\n    def __init__(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):\n        \n        self.df = df\n        self.imfolder = imfolder\n        self.transforms = transforms\n        self.train = train\n        self.meta_features = meta_features\n                \n    def __getitem__(self, index):\n         im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n         image = cv2.imread(im_path)\n         image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            \n        if self.transforms:\n            augmented = self.transforms(image=image)\n            image = augmented['image']\n            \n        meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n        if self.train:\n            y = self.df.iloc[index]['target']\n            return image, meta, y\n        else:\n            return image, meta\n   \n    def __len__(self):\n        return len(self.df)\n```\n\nHere is a Dataset with caching:\n\n```\nclass MelanomaDataset(Dataset):\n    def __init__(self, df: pd.DataFrame, imfolder: str, train: bool = True, transforms = None, meta_features = None, image_size = 256):\n        \n        self.df = df\n        self.imfolder = imfolder\n        self.transforms = transforms\n        self.train = train\n        self.meta_features = meta_features\n        \n        if(param['cache_on']):\n            shared_array_base = mp.Array(ctypes.c_ubyte, len(self.df)*3*image_size*image_size)\n            shared_array = np.ctypeslib.as_array(shared_array_base.get_obj())\n            self.shared_array = shared_array.reshape(len(self.df), image_size, image_size, 3)\n            self.use_cache = False\n            print(\"Cache initialized\")\n        \n    def __getitem__(self, index):\n        if(param['cache_on']):\n            if not self.use_cache:\n                im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n                image = cv2.imread(im_path)\n                image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n                self.shared_array[index] = image\n            image = self.shared_array[index]\n        else:\n            im_path = os.path.join(self.imfolder, self.df.iloc[index]['image_name'] + '.jpg')\n            image = cv2.imread(im_path)\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            \n        if self.transforms:\n            augmented = self.transforms(image=image)\n            image = augmented['image']\n            \n        meta = np.array(self.df.iloc[index][self.meta_features].values, dtype=np.float32)\n\n        if self.train:\n            y = self.df.iloc[index]['target']\n            return image, meta, y\n        else:\n            return (image, meta)\n    \n    def set_use_cache(self, use_cache):\n        self.use_cache = use_cache\n    \n    def __len__(self):\n        return len(self.df)\n```\n\nA little bit about what is happening here.  We allocate memory using c-types.  This is because, under the hood, Numpy and Pytorch (and likely other frameworks) actually use c-types for storage.  This way, when you move the data from c-types storage to Numpy or a Pytorch tensor, there is no copying of data, it uses the data in the same exact location.  If anything is re-arranged, simply a new \"view\" into the data is created.  We pre-allocate this data structure, so we must know the image size beforehand.  \n\n`use_cache` is defaulted to `False`.  So the first iteration of your train/val/oof/test loops it will not use the cache, and instead  it will be priming the cache.  On the next iterations forward (`epoch == 1`), you will then start to pull from the cache.\n\n```\nfor epoch in range(param['epochs'][fold - 1]):\n        if(param['cache_on']):\n            if(epoch == 1):\n                 train_loader.dataset.set_use_cache(use_cache=True)\n```\n        \nNow you just saved more than 1.5M disk file requests at the cost of some CPU memory.  Every epoch will be faster.  Your images/sec numbers will go way up.  Obviously the first epoch, first TTA, etc.  Will be just as slow as it was before, as it primes the cache.  After that you are in Turbo mode.  \n\nPlease post here if you used this and how it has speeded up your I/O!",
    "961109": "I calculate duration of each epoch and first epoch always takes much more time because data is read from disk, all other epochs are faster because they are read from cache of operating system.",
    "961111": "Yes, first epoch, or first TTA, will take as long as it normally does reading from disk.  Also, the instantiation of the Dataset requires the allocation of the data structure, which takes a few seconds.........and of course takes a little longer depending on image size.  Because with Epochs, Folds and TTA we are re-iterating so many times, that is where the savings come in.  The more epochs, folds a and TTA you do, the more savings.",
    "961112": "my point is - you don't need to store data in memory because it is already stored by operating system cache",
    "961145": "Epoch 0 is 2min 25sec (loading cache)\nEpoch 1 - 18 (or more, depending on when Early Stopping hits) is 2min 10sec\n\nSo a savings of 15 seconds per epoch.  So in my case that's:\n\nTraining: 15 seconds per epoch x 17 epochs x 5 folds    = 1275 seconds\nValidation: 15 seconds per epoch x 17 epochs x 5 folds = 1275 seconds\nOOF: 3 seconds per TTA x TTA 14 x 5 folds                    =   210 seconds\nTest: 1.5 seconds per TTA x TTA 14 * 5 folds                   =   105 seconds \nTotal saved:                                                                         = 2865 seconds = 47.75 minutes",
    "961149": "epoch 0\n{'train_score': 0.5907124039945352, 'valid_score': 0.6284858871924628, 'train_loss': 0.19184398562075033, 'valid_loss': 0.08261917846606058, 'duration': 772.745526, 'lr': 0.00020902725541536234}\nsaving model\nepoch 1\n{'train_score': 0.7463795335384227, 'valid_score': 0.8594380617329083, 'train_loss': 0.0882036334134919, 'valid_loss': 0.06833077350302655, 'duration': 449.452999, 'lr': 0.0005614256203280725}\nsaving model",
    "961182": "Is that 772 and 449 in seconds? I guess what you are showing is that without external caching, just your OS, your epoch time drops from the OS caching.  I am surprised the difference is over 300 seconds, I mean, you must have very slow image/sec ingest to have that kind of savings.",
    "962669": "Why not using a numpy array directly?  That's how I implement my data loader.  Just curious to know if you gain anything using ctype instead of a numpy array.",
    "962745": "I am creating a single array, which can be accessed by PyTorches \"workers\".  PyTorches \"workers\" (set by `num_workers` &gt; 0 in the data loader) use Python multi-processing.  This allows you to setup  a data structure (`mp.Array`) which is shared between processes.  NumPy arrays would be instantiated for every worker, as each worker clones the dataset, and thus create each their own underlying storage in memory, which is not what you would want to do.\n\nhttps://docs.python.org/2/library/multiprocessing.html",
    "962748": "these are numbers from Colab notebook, this competition, 256x256 images",
    "964910": "Awesome, thanks.  IO isn't better for me as I already cached values and saved the array to disk. However footprint is definitely smaller when sharing the array across workers.  It will pay off when I'll move beyond 256x256 images hopefully.  Thanks.\n\nNot sure why you don't get more upvotes on this.",
    "964990": "I think either it's hard to understand what is happening, or I do a bad job of explaining it, lol.",
    "964996": "A nice side effect: I was thinking of using hdf5 as loading the array for each worker would not have been possible with large image sizes.  You made this irrelevant.  Thanks again.",
    "965005": "I assumed requirement for your approach is to fit whole dataset into memory, this competition has large dataset",
    "965036": "jacekpoplawski yes it has to go into CPU memory.  So you would need about 1.4GB of CPU memory to store train data at 256x256.  6.9GB of CPU memory for train data at 768x768.  Most people have 32GB, or 64GB, or even 128GB of CPU memory, and PyTorch is probably only using a small fraction of that.  If you are running on a system with say 8-16GB of memory then you need to be careful for sure.",
    "965044": "I'm afraid that storing train data at 256x256 uses 6.5GB.\n\n3×256×256×33126 = 6512836608",
    "965086": "you're right!  I naively just looked at the size of my train directory lol.....I have enough memory that I don't think about it.......but yeah, no doubt with increased image size it can become an issue, but also with increased image size so does disk I/O!  And of course, you have to store validation and test as well.  A lot of systems should have the memory to work with 256 or 384.......or beyond.........some don't.  But if you have the memory I say use it!",
    "965127": "Probably I will not be able to get medal in this competition, but my way of working was to have more data than just 33126 images, that's why I am not able to use your idea in this case. \n\nAnyway, as I said operating system caches data from disk. Is your idea to resolve problem with multiple cores, do you mean that operating system caches each core independently, and your cache is one for all cores?",
    "965132": "OS file caching depends on a number of things.  I think you will see that your disk still has read access when loading files for each epoch.   In most OS's file cache is shared between processes.  In python data structures are not shared between processes, but if you instantiate them from the same data which resides in shared memory then they will effectively be sharing a data structure, which is what `mp.Array`() accomplishes.",
    "965142": "I remember when I was starting with machine learning i was always loading whole dataset, then later in Kaggle competition I learned that data can be huge and to avoid loading everything at once I learned how generator works. When I switched from Keras to PyTorch it was even easier to implemenent generator by using Dataset and DataLoader. I would probably need to make some tests to understand benefits of this idea, but have no time until this competition is over.",
    "965872": "Just a place holder so that I can find this post later! Thanks for sharing!",
    "966519": "Here is a link to my other post [Using CUDA pin_memory and non_blocking to speed up workflows in PyTorch](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172888) which goes hand and hand with this post for making sure you are getting every last bit out of your GPU(s).",
    "969746": "The overall goal of doing this is to keep your GPU's at 100%.  If they are not at 100% then you are leaving resources on the table, you are not getting all you can out of your environment.  Once your GPU's are at 100% you can just stop trying to optimize ingest and pipeline, because you can't do any better.  I have done a number of tweaks, most of which I have posted to the forums, and now often my GPU's look like this:\n\n`root@1e5de513dd11:/workspace/data/kaggle/Melanoma/notebook# nvidia-smi\nThu Aug 13 22:42:22 2020       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.05    Driver Version: 450.51.05    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  GeForce GTX 108...  On   | 00000000:05:00.0  On |                  N/A |\n| 53%   84C    P2    85W / 250W |  11165MiB / 11175MiB |    100%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   1  GeForce GTX 108...  On   | 00000000:06:00.0 Off |                  N/A |\n| 51%   85C    P2    85W / 250W |  11105MiB / 11178MiB |    100%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   2  GeForce GTX 108...  On   | 00000000:09:00.0 Off |                  N/A |\n| 59%   84C    P2    75W / 250W |  10309MiB / 11178MiB |    100%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   3  GeForce GTX 108...  On   | 00000000:0A:00.0 Off |                  N/A |\n| 48%   77C    P2    85W / 250W |   9707MiB / 11178MiB |    100%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n+-----------------------------------------------------------------------------+`\n\nSo 100% on all four.  There is some fluctuation, but its always up there.  A recap of things that will help:\n\n1. Fast disks/SSD.  Once it's cached it's irrelevant, but this helps for first epoch/TTA.\n2. Turbo mode cache as posted here\n3. Using `pin_memory` and `non_blocking` directives in CUDA [as described here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172888)\n4. Using AMP/Mixed Precision.  Why? Because the larger batch size, which means a larger/better sample which to compute certain metrics on, and not to mention its more efficient and faster when using larger batch sizes.\n\nAnything else to get the most out of your GPU's? Let us know!"
  },
  "source": "meta"
}