{
  "id": 172888,
  "title": "Using CUDA pin_memory and non_blocking to speed up workflows",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172888",
  "author_name": "",
  "post_date": "2020-08-06T21:24:07.767537700Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I am always trying to get the most efficiency out of my workflow.  One of the highest cost parts of deep learning is managing the I/O.  I am only familiar with how Pytorch handles this so this will be in relation to Pytorch but likely applies to other frameworks.</p>\n\n<p>A typical ingest in Pytorch involves reading the data from a dataset via a dataloader.  This comes from disk to the dataset, to the loader and into your loop.  You then typically move the data from CPU memory to CUDA memory (a copy).  This copy takes time.  <code>pin_memory</code> is used so that the memory is already allocated in the GPU at the same time it is allocated in the CPU, this saves some setup time.  The data still needs to be copied, however.</p>\n\n<p>CUDA allows you to copy data from CPU to GPU asynchronously.  Meaning while that is happening, you can do other things.  So the idea is you start copying <code>batch+1</code> from CPU to GPU, at the same time you are doing feed-forward/backprop on <code>batch</code>.  This is accomplished by setting <code>non_blocking=True</code></p>\n\n<p>Here is an example of how that is done.</p>\n\n<p>Our original code:</p>\n\n<p><code>\n       for idx, (image, meta, y) in enumerate(train_loader): \n            image = image.to(device=device, dtype=torch.float32)\n            meta = meta.to(device=device, dtype=torch.float32)\n            y = y.to(device=device, dtype=torch.float32)\n            .\n            .\n            .\n            training loop continues\n</code></p>\n\n<p>Code which takes advantage of <code>pin_memory</code> and <code>non_blocking</code>:\n<code>\ntrain_iter = iter(train_loader)\nnext_batch = train_iter.next()\nnext_batch = [_.cuda(non_blocking=True) for _ in next_batch ]\nfor i in range(len(train_loader)):\n    image, meta, y=next_batch\n    if i + 2 != len(train_loader): \n        # start copying data of next batch\n        next_batch = train_iter.next()\n        next_batch = [ _.cuda(non_blocking=True) for _ in next_batch]\n        .\n        .\n        .\n        training loop continues\n</code></p>\n\n<p>You should be setting <code>pin_memory=True</code> in your dataloader.  Do a test, see how it speeds up your workflow.  If your GPU's are already at 100% then this may not help you much.  But if you find your GPU's are waiting on I/O, then this is another method to keep them well fed.</p>",
  "messages": [
    {
      "id": "961034",
      "postDate": "08/06/2020 21:24:07",
      "content": "<p>I am always trying to get the most efficiency out of my workflow.  One of the highest cost parts of deep learning is managing the I/O.  I am only familiar with how Pytorch handles this so this will be in relation to Pytorch but likely applies to other frameworks.</p>\n\n<p>A typical ingest in Pytorch involves reading the data from a dataset via a dataloader.  This comes from disk to the dataset, to the loader and into your loop.  You then typically move the data from CPU memory to CUDA memory (a copy).  This copy takes time.  <code>pin_memory</code> is used so that the memory is already allocated in the GPU at the same time it is allocated in the CPU, this saves some setup time.  The data still needs to be copied, however.</p>\n\n<p>CUDA allows you to copy data from CPU to GPU asynchronously.  Meaning while that is happening, you can do other things.  So the idea is you start copying <code>batch+1</code> from CPU to GPU, at the same time you are doing feed-forward/backprop on <code>batch</code>.  This is accomplished by setting <code>non_blocking=True</code></p>\n\n<p>Here is an example of how that is done.</p>\n\n<p>Our original code:</p>\n\n<p><code>\n       for idx, (image, meta, y) in enumerate(train_loader): \n            image = image.to(device=device, dtype=torch.float32)\n            meta = meta.to(device=device, dtype=torch.float32)\n            y = y.to(device=device, dtype=torch.float32)\n            .\n            .\n            .\n            training loop continues\n</code></p>\n\n<p>Code which takes advantage of <code>pin_memory</code> and <code>non_blocking</code>:\n<code>\ntrain_iter = iter(train_loader)\nnext_batch = train_iter.next()\nnext_batch = [_.cuda(non_blocking=True) for _ in next_batch ]\nfor i in range(len(train_loader)):\n    image, meta, y=next_batch\n    if i + 2 != len(train_loader): \n        # start copying data of next batch\n        next_batch = train_iter.next()\n        next_batch = [ _.cuda(non_blocking=True) for _ in next_batch]\n        .\n        .\n        .\n        training loop continues\n</code></p>\n\n<p>You should be setting <code>pin_memory=True</code> in your dataloader.  Do a test, see how it speeds up your workflow.  If your GPU's are already at 100% then this may not help you much.  But if you find your GPU's are waiting on I/O, then this is another method to keep them well fed.</p>",
      "rawMarkdown": "I am always trying to get the most efficiency out of my workflow.  One of the highest cost parts of deep learning is managing the I/O.  I am only familiar with how Pytorch handles this so this will be in relation to Pytorch but likely applies to other frameworks.\n\nA typical ingest in Pytorch involves reading the data from a dataset via a dataloader.  This comes from disk to the dataset, to the loader and into your loop.  You then typically move the data from CPU memory to CUDA memory (a copy).  This copy takes time.  `pin_memory` is used so that the memory is already allocated in the GPU at the same time it is allocated in the CPU, this saves some setup time.  The data still needs to be copied, however.\n\nCUDA allows you to copy data from CPU to GPU asynchronously.  Meaning while that is happening, you can do other things.  So the idea is you start copying `batch+1` from CPU to GPU, at the same time you are doing feed-forward/backprop on `batch`.  This is accomplished by setting `non_blocking=True`\n\nHere is an example of how that is done.\n\nOur original code:\n\n```\n       for idx, (image, meta, y) in enumerate(train_loader): \n            image = image.to(device=device, dtype=torch.float32)\n            meta = meta.to(device=device, dtype=torch.float32)\n            y = y.to(device=device, dtype=torch.float32)\n            .\n            .\n            .\n            training loop continues\n```\n\n\nCode which takes advantage of `pin_memory` and `non_blocking`:\n```\ntrain_iter = iter(train_loader)\nnext_batch = train_iter.next()\nnext_batch = [_.cuda(non_blocking=True) for _ in next_batch ]\nfor i in range(len(train_loader)):\n    image, meta, y=next_batch\n    if i + 2 != len(train_loader): \n        # start copying data of next batch\n        next_batch = train_iter.next()\n        next_batch = [ _.cuda(non_blocking=True) for _ in next_batch]\n        .\n        .\n        .\n        training loop continues\n```\n\nYou should be setting `pin_memory=True` in your dataloader.  Do a test, see how it speeds up your workflow.  If your GPU's are already at 100% then this may not help you much.  But if you find your GPU's are waiting on I/O, then this is another method to keep them well fed.",
      "votes": null
    },
    {
      "id": "961482",
      "postDate": "08/07/2020 08:02:09",
      "content": "<p>Wow thats a great piece of advice. Another thing that I always face is that while dataloading happends inside my loop I find GPU usage is regularly shifting to 0 and say 87 . Does that happen due to my transforms?. It stays 0 for most of the time. Can you guide me in there? What is really happening. I used albumentations.Compose with hair augmention and some basic augmentations. I didn't even resized inside the transform.</p>",
      "rawMarkdown": "Wow thats a great piece of advice. Another thing that I always face is that while dataloading happends inside my loop I find GPU usage is regularly shifting to 0 and say 87 . Does that happen due to my transforms?. It stays 0 for most of the time. Can you guide me in there? What is really happening. I used albumentations.Compose with hair augmention and some basic augmentations. I didn't even resized inside the transform.",
      "votes": null
    },
    {
      "id": "961488",
      "postDate": "08/07/2020 08:20:13",
      "content": "<p>You are most likely cpu bottlenecked.</p>",
      "rawMarkdown": "You are most likely cpu bottlenecked.",
      "votes": null
    },
    {
      "id": "961573",
      "postDate": "08/07/2020 09:51:21",
      "content": "<p>When your GPU's are idle there could be a few reasons.  It could be waiting on file I/O, it could definitely be waiting on your transforms, it could be allocating memory for the copy.</p>\n\n<p>Using a very fast file access method could help epoch 1, such as LMDB or HFS5.  Someone pointed out that most OS's do a decent job of file caching so after epoch 1 much of the files should be in memory.  I use a memory cache and that saves me 47+ minutes per run and I break this down in a post I did.</p>\n\n<p>Make sure <code>pin_memory=True</code> in your dataloader.  Also  make sure you have optimized your <code>num_workers</code>.  Don't just set an arbitrary value, find the RIGHT value.  Using the actual train data, <a href=\"http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/\">Follow this article and run this code</a>.  I find just setting <code>num_workers</code> to the number of CPU cores, or 2x the cores, or some other magic number, is not optimal.</p>\n\n<p>You can actually profile PyTorch to see where its spending time, but I have not done that.  </p>",
      "rawMarkdown": "When your GPU's are idle there could be a few reasons.  It could be waiting on file I/O, it could definitely be waiting on your transforms, it could be allocating memory for the copy.\n\nUsing a very fast file access method could help epoch 1, such as LMDB or HFS5.  Someone pointed out that most OS's do a decent job of file caching so after epoch 1 much of the files should be in memory.  I use a memory cache and that saves me 47+ minutes per run and I break this down in a post I did.\n\nMake sure `pin_memory=True` in your dataloader.  Also  make sure you have optimized your `num_workers`.  Don't just set an arbitrary value, find the RIGHT value.  Using the actual train data, [Follow this article and run this code](http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/).  I find just setting `num_workers` to the number of CPU cores, or 2x the cores, or some other magic number, is not optimal.\n\nYou can actually profile PyTorch to see where its spending time, but I have not done that.",
      "votes": null
    },
    {
      "id": "962692",
      "postDate": "08/08/2020 10:59:47",
      "content": "<p>Thanks for sharing.  However your code didn't work with my loader code, maybe because I' using pytorch 1.6 and you use an older version?</p>\n<p>A little look at pytorch documentation gave me the solution:</p>\n<ol>\n<li>As you wrote, pass <code>pin_memory=True</code> to data loaders.</li>\n<li>Pass <code>non_blocking=True</code> argument to a <code>to()</code> or a <code>cuda()</code> call. </li>\n</ol>\n<p>My code now looks like:</p>\n<pre><code>def train_epoch(loader, model, optimizer, scaler, device, mixup):\n\n    model.train()\n    model.zero_grad()\n    for i,(batch) in enumerate(tqdm(loader)):\n\n        data = batch['image'].to(device, non_blocking=True )\n        target = batch['target'].to(device, non_blocking=True )\n        ...\n</code></pre>",
      "rawMarkdown": "Thanks for sharing.  However your code didn't work with my loader code, maybe because I' using pytorch 1.6 and you use an older version?\n\nA little look at pytorch documentation gave me the solution:\n\n1. As you wrote, pass `pin_memory=True` to data loaders.\n2. Pass `non_blocking=True` argument to a `to()` or a `cuda()` call. \n\nMy code now looks like:\n\n```\ndef train_epoch(loader, model, optimizer, scaler, device, mixup):\n\n    model.train()\n    model.zero_grad()\n    for i,(batch) in enumerate(tqdm(loader)):\n        \n        data = batch['image'].to(device, non_blocking=True )\n        target = batch['target'].to(device, non_blocking=True )\n        ...\n```",
      "votes": null
    },
    {
      "id": "962757",
      "postDate": "08/08/2020 12:16:54",
      "content": "<p>What error did you get?  The code you post above is valid, but it misses the point.  The code example I posted is moving one batch to memory asynchronously as its processing another batch.  The code you posted does set the async flag, but it really won't accomplish too much.</p>\n\n<p>Async calls only \"work\" if there is something to do between the time you start the async task and when it is needed.  Most likely after you flag <code>data</code> to move to cuda memory, your next line (or close to it) is something along the lines of:</p>\n\n<p>```\npred = model(data)</p>\n\n<p><code>``\nwhich means there was no time for PyTorch to go do something else while it was moving</code>data<code>to memory, it had to move</code>data<code>to cuda before it could execute</code>pred = model(data)<code>.  Now as for</code>target<code>being moved to cuda memory, yes that can totally happen asynchronously because PyTorch can go off and do that while its running</code>pred = model(data)<code>.  However, as you probably know, this doesn't buy you much, since</code>target<code>is nothing more than a Tensor of 0's and 1's of size</code>batch`.</p>\n\n<p>You can totally move a batch of <code>data</code> asynchronously to cuda memory while you are processing another batch.  My example shows the general architecture.  It is more convoluted looking than a normal iterator, but if you think about what is happening it has to be like that.  </p>",
      "rawMarkdown": "What error did you get?  The code you post above is valid, but it misses the point.  The code example I posted is moving one batch to memory asynchronously as its processing another batch.  The code you posted does set the async flag, but it really won't accomplish too much.\n\nAsync calls only \"work\" if there is something to do between the time you start the async task and when it is needed.  Most likely after you flag `data` to move to cuda memory, your next line (or close to it) is something along the lines of:\n\n```\npred = model(data)\n\n```\nwhich means there was no time for PyTorch to go do something else while it was moving `data` to memory, it had to move `data` to cuda before it could execute `pred = model(data)`.  Now as for `target` being moved to cuda memory, yes that can totally happen asynchronously because PyTorch can go off and do that while its running `pred = model(data)`.  However, as you probably know, this doesn't buy you much, since `target` is nothing more than a Tensor of 0's and 1's of size `batch`.\n\nYou can totally move a batch of `data` asynchronously to cuda memory while you are processing another batch.  My example shows the general architecture.  It is more convoluted looking than a normal iterator, but if you think about what is happening it has to be like that.",
      "votes": null
    },
    {
      "id": "963030",
      "postDate": "08/08/2020 16:01:43",
      "content": "<p>The error was something like 'string does not have cuda attribute'.  I get you point however, will dig further.</p>",
      "rawMarkdown": "The error was something like 'string does not have cuda attribute'.  I get you point however, will dig further.",
      "votes": null
    },
    {
      "id": "964976",
      "postDate": "08/10/2020 09:55:59",
      "content": "<p>My batch is a dictionary and the code was trying to move the keys to gpu rather than the values.  I modified the code to deal with dictionaries and it works like a charm.  Thanks a lot.  I see a minor speedup as my gpu was used at full capacity already, but I'm sure this will also pay off with larger images.  </p>\n<p>Shouldn't this condition:</p>\n<pre><code>    if i + 2 != len(train_loader): \n</code></pre>\n<p>be</p>\n<pre><code>    if i + 1 != len(train_loader): \n</code></pre>\n<p>Anyway, here is my code for reference.</p>\n<pre><code>bar = tqdm(range(len(loader)))\nload_iter = iter(loader)\nbatch = load_iter.next()\nbatch = { k:batch[k].cuda(non_blocking=True) for k in batch.keys() }\n\nfor i in bar:\n\n    data = batch['image']\n    target = batch['target']\n    if i + 1 &lt; len(loader):\n        batch = load_iter.next()\n        batch = { k:batch[k].cuda(non_blocking=True) for k in batch.keys() }\n    ...\n</code></pre>",
      "rawMarkdown": "My batch is a dictionary and the code was trying to move the keys to gpu rather than the values.  I modified the code to deal with dictionaries and it works like a charm.  Thanks a lot.  I see a minor speedup as my gpu was used at full capacity already, but I'm sure this will also pay off with larger images.  \n\nShouldn't this condition:\n\n```\n    if i + 2 != len(train_loader): \n\n```\nbe\n\n```\n    if i + 1 != len(train_loader): \n\n```\n\nAnyway, here is my code for reference.\n\n    bar = tqdm(range(len(loader)))\n    load_iter = iter(loader)\n    batch = load_iter.next()\n    batch = { k:batch[k].cuda(non_blocking=True) for k in batch.keys() }\n    \n    for i in bar:\n        \n        data = batch['image']\n        target = batch['target']\n        if i + 1 &lt; len(loader):\n            batch = load_iter.next()\n            batch = { k:batch[k].cuda(non_blocking=True) for k in batch.keys() }\n        ...",
      "votes": null
    },
    {
      "id": "965006",
      "postDate": "08/10/2020 10:16:54",
      "content": "<p>I am glad it is working.  As for the <code>i +2</code>, I believe its correct in my example, but may need to be changed depending on your iterator.  In my example, I use <code>range</code>().  </p>\n\n<p>Say <code>loader</code> is of size 10.  <code>range(len(loader))</code> would loop from 0 to 9 (which is 10 total elements).  We are holding one batch in <code>next_batch</code>.  Therefore, once we are at i=8 we should stop, because we already pre-iterated before the loop started.  Make sense?</p>",
      "rawMarkdown": "I am glad it is working.  As for the `i +2`, I believe its correct in my example, but may need to be changed depending on your iterator.  In my example, I use `range`().  \n\nSay `loader` is of size 10.  `range(len(loader))` would loop from 0 to 9 (which is 10 total elements).  We are holding one batch in `next_batch`.  Therefore, once we are at i=8 we should stop, because we already pre-iterated before the loop started.  Make sense?",
      "votes": null
    },
    {
      "id": "965040",
      "postDate": "08/10/2020 10:47:43",
      "content": "<p>I am using range as well as you can see from my code.</p>\n<p>When you are at 8 you have loaded 9 batches.  You still have one to go.</p>",
      "rawMarkdown": "I am using range as well as you can see from my code.\n\nWhen you are at 8 you have loaded 9 batches.  You still have one to go.",
      "votes": null
    },
    {
      "id": "965071",
      "postDate": "08/10/2020 11:16:54",
      "content": "<p>ahh yes, I believe you are correct, good catch. </p>",
      "rawMarkdown": "ahh yes, I believe you are correct, good catch.",
      "votes": null
    },
    {
      "id": "965080",
      "postDate": "08/10/2020 11:21:16",
      "content": "<p>No pb, it is always tricky, and I had to check several times to get it right ;)</p>",
      "rawMarkdown": "No pb, it is always tricky, and I had to check several times to get it right ;)",
      "votes": null
    },
    {
      "id": "966517",
      "postDate": "08/11/2020 13:38:34",
      "content": "<p>One of the main reasons for doing what I propose is to keep your GPU's 100% utilized, or more accurately at least make them more utilized.  Here is typical outputs of my GPU's with <code>nvidia-smi</code></p>\n\n<p>```\nroot@c4d2b64472f2:/workspace# nvidia-smi\nTue Aug 11 13:33:57 2020 <br>\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.05    Driver Version: 450.51.05    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  GeForce GTX 108...  On   | 00000000:05:00.0  On |                  N/A |\n| 44%   79C    P2   193W / 250W |  10907MiB / 11175MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   1  GeForce GTX 108...  On   | 00000000:06:00.0 Off |                  N/A |\n| 44%   79C    P2   186W / 250W |  10367MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   2  GeForce GTX 108...  On   | 00000000:09:00.0 Off |                  N/A |\n| 48%   83C    P2   146W / 250W |  10367MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   3  GeForce GTX 108...  On   | 00000000:0A:00.0 Off |                  N/A |\n| 39%   71C    P2   201W / 250W |  10355MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+</p>\n\n<p>+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n+-----------------------------------------------------------------------------+\n```</p>\n\n<p>If you look at \"on average\" your utilization before you do these changes, and after, you will see the increase, and remember these operations are async so they are hidden in between other CPU steps going on.</p>\n\n<p>Here is a link to my other post on <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172904\">Turbo charging PyTorch dataloading using a c-types cache</a> which goes hand in hand with this.</p>",
      "rawMarkdown": "One of the main reasons for doing what I propose is to keep your GPU's 100% utilized, or more accurately at least make them more utilized.  Here is typical outputs of my GPU's with `nvidia-smi `\n\n```\nroot@c4d2b64472f2:/workspace# nvidia-smi\nTue Aug 11 13:33:57 2020       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.05    Driver Version: 450.51.05    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  GeForce GTX 108...  On   | 00000000:05:00.0  On |                  N/A |\n| 44%   79C    P2   193W / 250W |  10907MiB / 11175MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   1  GeForce GTX 108...  On   | 00000000:06:00.0 Off |                  N/A |\n| 44%   79C    P2   186W / 250W |  10367MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   2  GeForce GTX 108...  On   | 00000000:09:00.0 Off |                  N/A |\n| 48%   83C    P2   146W / 250W |  10367MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   3  GeForce GTX 108...  On   | 00000000:0A:00.0 Off |                  N/A |\n| 39%   71C    P2   201W / 250W |  10355MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n+-----------------------------------------------------------------------------+\n```\n\nIf you look at \"on average\" your utilization before you do these changes, and after, you will see the increase, and remember these operations are async so they are hidden in between other CPU steps going on.\n\nHere is a link to my other post on [Turbo charging PyTorch dataloading using a c-types cache](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172904) which goes hand in hand with this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 962692,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/08/2020 10:59:47",
      "content": "<p>Thanks for sharing.  However your code didn't work with my loader code, maybe because I' using pytorch 1.6 and you use an older version?</p>\n<p>A little look at pytorch documentation gave me the solution:</p>\n<ol>\n<li>As you wrote, pass <code>pin_memory=True</code> to data loaders.</li>\n<li>Pass <code>non_blocking=True</code> argument to a <code>to()</code> or a <code>cuda()</code> call. </li>\n</ol>\n<p>My code now looks like:</p>\n<pre><code>def train_epoch(loader, model, optimizer, scaler, device, mixup):\n\n    model.train()\n    model.zero_grad()\n    for i,(batch) in enumerate(tqdm(loader)):\n\n        data = batch['image'].to(device, non_blocking=True )\n        target = batch['target'].to(device, non_blocking=True )\n        ...\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 962757,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/08/2020 12:16:54",
          "content": "<p>What error did you get?  The code you post above is valid, but it misses the point.  The code example I posted is moving one batch to memory asynchronously as its processing another batch.  The code you posted does set the async flag, but it really won't accomplish too much.</p>\n\n<p>Async calls only \"work\" if there is something to do between the time you start the async task and when it is needed.  Most likely after you flag <code>data</code> to move to cuda memory, your next line (or close to it) is something along the lines of:</p>\n\n<p>```\npred = model(data)</p>\n\n<p><code>``\nwhich means there was no time for PyTorch to go do something else while it was moving</code>data<code>to memory, it had to move</code>data<code>to cuda before it could execute</code>pred = model(data)<code>.  Now as for</code>target<code>being moved to cuda memory, yes that can totally happen asynchronously because PyTorch can go off and do that while its running</code>pred = model(data)<code>.  However, as you probably know, this doesn't buy you much, since</code>target<code>is nothing more than a Tensor of 0's and 1's of size</code>batch`.</p>\n\n<p>You can totally move a batch of <code>data</code> asynchronously to cuda memory while you are processing another batch.  My example shows the general architecture.  It is more convoluted looking than a normal iterator, but if you think about what is happening it has to be like that.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963030,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2020 16:01:43",
          "content": "<p>The error was something like 'string does not have cuda attribute'.  I get you point however, will dig further.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964976,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 09:55:59",
          "content": "<p>My batch is a dictionary and the code was trying to move the keys to gpu rather than the values.  I modified the code to deal with dictionaries and it works like a charm.  Thanks a lot.  I see a minor speedup as my gpu was used at full capacity already, but I'm sure this will also pay off with larger images.  </p>\n<p>Shouldn't this condition:</p>\n<pre><code>    if i + 2 != len(train_loader): \n</code></pre>\n<p>be</p>\n<pre><code>    if i + 1 != len(train_loader): \n</code></pre>\n<p>Anyway, here is my code for reference.</p>\n<pre><code>bar = tqdm(range(len(loader)))\nload_iter = iter(loader)\nbatch = load_iter.next()\nbatch = { k:batch[k].cuda(non_blocking=True) for k in batch.keys() }\n\nfor i in bar:\n\n    data = batch['image']\n    target = batch['target']\n    if i + 1 &lt; len(loader):\n        batch = load_iter.next()\n        batch = { k:batch[k].cuda(non_blocking=True) for k in batch.keys() }\n    ...\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965006,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 10:16:54",
          "content": "<p>I am glad it is working.  As for the <code>i +2</code>, I believe its correct in my example, but may need to be changed depending on your iterator.  In my example, I use <code>range</code>().  </p>\n\n<p>Say <code>loader</code> is of size 10.  <code>range(len(loader))</code> would loop from 0 to 9 (which is 10 total elements).  We are holding one batch in <code>next_batch</code>.  Therefore, once we are at i=8 we should stop, because we already pre-iterated before the loop started.  Make sense?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965040,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 10:47:43",
          "content": "<p>I am using range as well as you can see from my code.</p>\n<p>When you are at 8 you have loaded 9 batches.  You still have one to go.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965071,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 11:16:54",
          "content": "<p>ahh yes, I believe you are correct, good catch. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965080,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 11:21:16",
          "content": "<p>No pb, it is always tricky, and I had to check several times to get it right ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 961482,
      "author_name": "tawheedrony",
      "author_url": "",
      "post_date": "08/07/2020 08:02:09",
      "content": "<p>Wow thats a great piece of advice. Another thing that I always face is that while dataloading happends inside my loop I find GPU usage is regularly shifting to 0 and say 87 . Does that happen due to my transforms?. It stays 0 for most of the time. Can you guide me in there? What is really happening. I used albumentations.Compose with hair augmention and some basic augmentations. I didn't even resized inside the transform.</p>",
      "votes": null,
      "replies": [
        {
          "id": 961488,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/07/2020 08:20:13",
          "content": "<p>You are most likely cpu bottlenecked.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961573,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/07/2020 09:51:21",
          "content": "<p>When your GPU's are idle there could be a few reasons.  It could be waiting on file I/O, it could definitely be waiting on your transforms, it could be allocating memory for the copy.</p>\n\n<p>Using a very fast file access method could help epoch 1, such as LMDB or HFS5.  Someone pointed out that most OS's do a decent job of file caching so after epoch 1 much of the files should be in memory.  I use a memory cache and that saves me 47+ minutes per run and I break this down in a post I did.</p>\n\n<p>Make sure <code>pin_memory=True</code> in your dataloader.  Also  make sure you have optimized your <code>num_workers</code>.  Don't just set an arbitrary value, find the RIGHT value.  Using the actual train data, <a href=\"http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/\">Follow this article and run this code</a>.  I find just setting <code>num_workers</code> to the number of CPU cores, or 2x the cores, or some other magic number, is not optimal.</p>\n\n<p>You can actually profile PyTorch to see where its spending time, but I have not done that.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 966517,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "08/11/2020 13:38:34",
      "content": "<p>One of the main reasons for doing what I propose is to keep your GPU's 100% utilized, or more accurately at least make them more utilized.  Here is typical outputs of my GPU's with <code>nvidia-smi</code></p>\n\n<p>```\nroot@c4d2b64472f2:/workspace# nvidia-smi\nTue Aug 11 13:33:57 2020 <br>\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.05    Driver Version: 450.51.05    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  GeForce GTX 108...  On   | 00000000:05:00.0  On |                  N/A |\n| 44%   79C    P2   193W / 250W |  10907MiB / 11175MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   1  GeForce GTX 108...  On   | 00000000:06:00.0 Off |                  N/A |\n| 44%   79C    P2   186W / 250W |  10367MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   2  GeForce GTX 108...  On   | 00000000:09:00.0 Off |                  N/A |\n| 48%   83C    P2   146W / 250W |  10367MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   3  GeForce GTX 108...  On   | 00000000:0A:00.0 Off |                  N/A |\n| 39%   71C    P2   201W / 250W |  10355MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+</p>\n\n<p>+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n+-----------------------------------------------------------------------------+\n```</p>\n\n<p>If you look at \"on average\" your utilization before you do these changes, and after, you will see the increase, and remember these operations are async so they are hidden in between other CPU steps going on.</p>\n\n<p>Here is a link to my other post on <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172904\">Turbo charging PyTorch dataloading using a c-types cache</a> which goes hand in hand with this.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "961034": "I am always trying to get the most efficiency out of my workflow.  One of the highest cost parts of deep learning is managing the I/O.  I am only familiar with how Pytorch handles this so this will be in relation to Pytorch but likely applies to other frameworks.\n\nA typical ingest in Pytorch involves reading the data from a dataset via a dataloader.  This comes from disk to the dataset, to the loader and into your loop.  You then typically move the data from CPU memory to CUDA memory (a copy).  This copy takes time.  `pin_memory` is used so that the memory is already allocated in the GPU at the same time it is allocated in the CPU, this saves some setup time.  The data still needs to be copied, however.\n\nCUDA allows you to copy data from CPU to GPU asynchronously.  Meaning while that is happening, you can do other things.  So the idea is you start copying `batch+1` from CPU to GPU, at the same time you are doing feed-forward/backprop on `batch`.  This is accomplished by setting `non_blocking=True`\n\nHere is an example of how that is done.\n\nOur original code:\n\n```\n       for idx, (image, meta, y) in enumerate(train_loader): \n            image = image.to(device=device, dtype=torch.float32)\n            meta = meta.to(device=device, dtype=torch.float32)\n            y = y.to(device=device, dtype=torch.float32)\n            .\n            .\n            .\n            training loop continues\n```\n\n\nCode which takes advantage of `pin_memory` and `non_blocking`:\n```\ntrain_iter = iter(train_loader)\nnext_batch = train_iter.next()\nnext_batch = [_.cuda(non_blocking=True) for _ in next_batch ]\nfor i in range(len(train_loader)):\n    image, meta, y=next_batch\n    if i + 2 != len(train_loader): \n        # start copying data of next batch\n        next_batch = train_iter.next()\n        next_batch = [ _.cuda(non_blocking=True) for _ in next_batch]\n        .\n        .\n        .\n        training loop continues\n```\n\nYou should be setting `pin_memory=True` in your dataloader.  Do a test, see how it speeds up your workflow.  If your GPU's are already at 100% then this may not help you much.  But if you find your GPU's are waiting on I/O, then this is another method to keep them well fed.",
    "961482": "Wow thats a great piece of advice. Another thing that I always face is that while dataloading happends inside my loop I find GPU usage is regularly shifting to 0 and say 87 . Does that happen due to my transforms?. It stays 0 for most of the time. Can you guide me in there? What is really happening. I used albumentations.Compose with hair augmention and some basic augmentations. I didn't even resized inside the transform.",
    "961488": "You are most likely cpu bottlenecked.",
    "961573": "When your GPU's are idle there could be a few reasons.  It could be waiting on file I/O, it could definitely be waiting on your transforms, it could be allocating memory for the copy.\n\nUsing a very fast file access method could help epoch 1, such as LMDB or HFS5.  Someone pointed out that most OS's do a decent job of file caching so after epoch 1 much of the files should be in memory.  I use a memory cache and that saves me 47+ minutes per run and I break this down in a post I did.\n\nMake sure `pin_memory=True` in your dataloader.  Also  make sure you have optimized your `num_workers`.  Don't just set an arbitrary value, find the RIGHT value.  Using the actual train data, [Follow this article and run this code](http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/).  I find just setting `num_workers` to the number of CPU cores, or 2x the cores, or some other magic number, is not optimal.\n\nYou can actually profile PyTorch to see where its spending time, but I have not done that.",
    "962692": "Thanks for sharing.  However your code didn't work with my loader code, maybe because I' using pytorch 1.6 and you use an older version?\n\nA little look at pytorch documentation gave me the solution:\n\n1. As you wrote, pass `pin_memory=True` to data loaders.\n2. Pass `non_blocking=True` argument to a `to()` or a `cuda()` call. \n\nMy code now looks like:\n\n```\ndef train_epoch(loader, model, optimizer, scaler, device, mixup):\n\n    model.train()\n    model.zero_grad()\n    for i,(batch) in enumerate(tqdm(loader)):\n        \n        data = batch['image'].to(device, non_blocking=True )\n        target = batch['target'].to(device, non_blocking=True )\n        ...\n```",
    "962757": "What error did you get?  The code you post above is valid, but it misses the point.  The code example I posted is moving one batch to memory asynchronously as its processing another batch.  The code you posted does set the async flag, but it really won't accomplish too much.\n\nAsync calls only \"work\" if there is something to do between the time you start the async task and when it is needed.  Most likely after you flag `data` to move to cuda memory, your next line (or close to it) is something along the lines of:\n\n```\npred = model(data)\n\n```\nwhich means there was no time for PyTorch to go do something else while it was moving `data` to memory, it had to move `data` to cuda before it could execute `pred = model(data)`.  Now as for `target` being moved to cuda memory, yes that can totally happen asynchronously because PyTorch can go off and do that while its running `pred = model(data)`.  However, as you probably know, this doesn't buy you much, since `target` is nothing more than a Tensor of 0's and 1's of size `batch`.\n\nYou can totally move a batch of `data` asynchronously to cuda memory while you are processing another batch.  My example shows the general architecture.  It is more convoluted looking than a normal iterator, but if you think about what is happening it has to be like that.",
    "963030": "The error was something like 'string does not have cuda attribute'.  I get you point however, will dig further.",
    "964976": "My batch is a dictionary and the code was trying to move the keys to gpu rather than the values.  I modified the code to deal with dictionaries and it works like a charm.  Thanks a lot.  I see a minor speedup as my gpu was used at full capacity already, but I'm sure this will also pay off with larger images.  \n\nShouldn't this condition:\n\n```\n    if i + 2 != len(train_loader): \n\n```\nbe\n\n```\n    if i + 1 != len(train_loader): \n\n```\n\nAnyway, here is my code for reference.\n\n    bar = tqdm(range(len(loader)))\n    load_iter = iter(loader)\n    batch = load_iter.next()\n    batch = { k:batch[k].cuda(non_blocking=True) for k in batch.keys() }\n    \n    for i in bar:\n        \n        data = batch['image']\n        target = batch['target']\n        if i + 1 &lt; len(loader):\n            batch = load_iter.next()\n            batch = { k:batch[k].cuda(non_blocking=True) for k in batch.keys() }\n        ...",
    "965006": "I am glad it is working.  As for the `i +2`, I believe its correct in my example, but may need to be changed depending on your iterator.  In my example, I use `range`().  \n\nSay `loader` is of size 10.  `range(len(loader))` would loop from 0 to 9 (which is 10 total elements).  We are holding one batch in `next_batch`.  Therefore, once we are at i=8 we should stop, because we already pre-iterated before the loop started.  Make sense?",
    "965040": "I am using range as well as you can see from my code.\n\nWhen you are at 8 you have loaded 9 batches.  You still have one to go.",
    "965071": "ahh yes, I believe you are correct, good catch.",
    "965080": "No pb, it is always tricky, and I had to check several times to get it right ;)",
    "966517": "One of the main reasons for doing what I propose is to keep your GPU's 100% utilized, or more accurately at least make them more utilized.  Here is typical outputs of my GPU's with `nvidia-smi `\n\n```\nroot@c4d2b64472f2:/workspace# nvidia-smi\nTue Aug 11 13:33:57 2020       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.05    Driver Version: 450.51.05    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  GeForce GTX 108...  On   | 00000000:05:00.0  On |                  N/A |\n| 44%   79C    P2   193W / 250W |  10907MiB / 11175MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   1  GeForce GTX 108...  On   | 00000000:06:00.0 Off |                  N/A |\n| 44%   79C    P2   186W / 250W |  10367MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   2  GeForce GTX 108...  On   | 00000000:09:00.0 Off |                  N/A |\n| 48%   83C    P2   146W / 250W |  10367MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n|   3  GeForce GTX 108...  On   | 00000000:0A:00.0 Off |                  N/A |\n| 39%   71C    P2   201W / 250W |  10355MiB / 11178MiB |     99%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n+-----------------------------------------------------------------------------+\n```\n\nIf you look at \"on average\" your utilization before you do these changes, and after, you will see the increase, and remember these operations are async so they are hidden in between other CPU steps going on.\n\nHere is a link to my other post on [Turbo charging PyTorch dataloading using a c-types cache](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172904) which goes hand in hand with this."
  },
  "source": "meta"
}