{
  "id": 179951,
  "title": "Accelerating Computer Vision on PyTorch (Nvidia @ ECCV 2020)",
  "url": "/competitions/birdsong-recognition/discussion/179951",
  "author_name": "",
  "post_date": "2020-09-03T12:00:18.973579100Z",
  "votes": 55,
  "comment_count": 4,
  "views": 0,
  "content": "<p>For those who use PyTorch, here is a cool tutorial with a bunch of tricks to achieve better training and inference times with PyTorch : </p>\n<ul>\n<li>Site : <a href=\"https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/\" target=\"_blank\">https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/</a></li>\n<li>Presentation : <a href=\"https://www.youtube.com/watch?v=9mS1fIYj1So\" target=\"_blank\">https://www.youtube.com/watch?v=9mS1fIYj1So</a></li>\n<li>Slides pdf : <a href=\"https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/files/szymon_migacz-pytorch-performance-tuning-guide.pdf\" target=\"_blank\">https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/files/szymon_migacz-pytorch-performance-tuning-guide.pdf</a></li>\n</ul>\n<p>Some stuff can be useful here, even though the time bottleneck seems to come from the data generation.</p>\n<p>Here are a few points mentionned that can easily be implemented: </p>\n<ul>\n<li>Dataloader: Use <code>num_workers &gt; 0</code> and <code>pin_memory=True</code> for data loaders</li>\n<li>Use <code>torch.backends.cudnn.benchmark=True</code></li>\n<li>Increase your <code>batch_size</code> as much as you can</li>\n<li>Instead of <code>model.zero_grad()</code> use <code>for param in model.parameters(): param.grad = None</code></li>\n</ul>\n<p>And more, give a look at the presentation :)</p>",
  "messages": [
    {
      "id": "996577",
      "postDate": "09/03/2020 12:00:18",
      "content": "<p>For those who use PyTorch, here is a cool tutorial with a bunch of tricks to achieve better training and inference times with PyTorch : </p>\n<ul>\n<li>Site : <a href=\"https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/\" target=\"_blank\">https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/</a></li>\n<li>Presentation : <a href=\"https://www.youtube.com/watch?v=9mS1fIYj1So\" target=\"_blank\">https://www.youtube.com/watch?v=9mS1fIYj1So</a></li>\n<li>Slides pdf : <a href=\"https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/files/szymon_migacz-pytorch-performance-tuning-guide.pdf\" target=\"_blank\">https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/files/szymon_migacz-pytorch-performance-tuning-guide.pdf</a></li>\n</ul>\n<p>Some stuff can be useful here, even though the time bottleneck seems to come from the data generation.</p>\n<p>Here are a few points mentionned that can easily be implemented: </p>\n<ul>\n<li>Dataloader: Use <code>num_workers &gt; 0</code> and <code>pin_memory=True</code> for data loaders</li>\n<li>Use <code>torch.backends.cudnn.benchmark=True</code></li>\n<li>Increase your <code>batch_size</code> as much as you can</li>\n<li>Instead of <code>model.zero_grad()</code> use <code>for param in model.parameters(): param.grad = None</code></li>\n</ul>\n<p>And more, give a look at the presentation :)</p>",
      "rawMarkdown": "For those who use PyTorch, here is a cool tutorial with a bunch of tricks to achieve better training and inference times with PyTorch : \n- Site : https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/\n- Presentation : https://www.youtube.com/watch?v=9mS1fIYj1So\n- Slides pdf : https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/files/szymon_migacz-pytorch-performance-tuning-guide.pdf\n\n\nSome stuff can be useful here, even though the time bottleneck seems to come from the data generation.\n\nHere are a few points mentionned that can easily be implemented: \n\n- Dataloader: Use `num_workers > 0` and `pin_memory=True` for data loaders\n- Use `torch.backends.cudnn.benchmark=True`\n- Increase your `batch_size` as much as you can\n- Instead of `model.zero_grad()` use `for param in model.parameters(): param.grad = None`\n\nAnd more, give a look at the presentation :)",
      "votes": null
    },
    {
      "id": "996581",
      "postDate": "09/03/2020 12:05:16",
      "content": "<p>Thanks! Can one get the same features with Pytorch Lightning, sounds pretty similar at first glance? Have it on my todo list to implement.</p>",
      "rawMarkdown": "Thanks! Can one get the same features with Pytorch Lightning, sounds pretty similar at first glance? Have it on my todo list to implement.",
      "votes": null
    },
    {
      "id": "996634",
      "postDate": "09/03/2020 12:48:11",
      "content": "<p>Use asynchronous transfer with non_blocking=True</p>\n<p>Other additions discussed in comments ot this tweet: <a href=\"https://twitter.com/karpathy/status/1299921324333170689\" target=\"_blank\">https://twitter.com/karpathy/status/1299921324333170689</a></p>",
      "rawMarkdown": "Use asynchronous transfer with non_blocking=True\n\nOther additions discussed in comments ot this tweet: https://twitter.com/karpathy/status/1299921324333170689",
      "votes": null
    },
    {
      "id": "997137",
      "postDate": "09/03/2020 19:13:52",
      "content": "<p>I'm using the dataloader tip for num workers and pin memory in PyTorch lightning. </p>\n<p>I would guess the other tips also apply to PL. model.zero_grad() is hidden away but I'm sure it can be modified somehow. </p>",
      "rawMarkdown": "I'm using the dataloader tip for num workers and pin memory in PyTorch lightning. \n\nI would guess the other tips also apply to PL. model.zero_grad() is hidden away but I'm sure it can be modified somehow.",
      "votes": null
    },
    {
      "id": "998581",
      "postDate": "09/04/2020 20:56:10",
      "content": "<p>Thank you! Thats very intresting. I will take a deeper look.</p>",
      "rawMarkdown": "Thank you! Thats very intresting. I will take a deeper look.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 996581,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "09/03/2020 12:05:16",
      "content": "<p>Thanks! Can one get the same features with Pytorch Lightning, sounds pretty similar at first glance? Have it on my todo list to implement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 997137,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "09/03/2020 19:13:52",
          "content": "<p>I'm using the dataloader tip for num workers and pin memory in PyTorch lightning. </p>\n<p>I would guess the other tips also apply to PL. model.zero_grad() is hidden away but I'm sure it can be modified somehow. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 996634,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/03/2020 12:48:11",
      "content": "<p>Use asynchronous transfer with non_blocking=True</p>\n<p>Other additions discussed in comments ot this tweet: <a href=\"https://twitter.com/karpathy/status/1299921324333170689\" target=\"_blank\">https://twitter.com/karpathy/status/1299921324333170689</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 998581,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "09/04/2020 20:56:10",
      "content": "<p>Thank you! Thats very intresting. I will take a deeper look.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "996577": "For those who use PyTorch, here is a cool tutorial with a bunch of tricks to achieve better training and inference times with PyTorch : \n- Site : https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/\n- Presentation : https://www.youtube.com/watch?v=9mS1fIYj1So\n- Slides pdf : https://nvlabs.github.io/eccv2020-mixed-precision-tutorial/files/szymon_migacz-pytorch-performance-tuning-guide.pdf\n\n\nSome stuff can be useful here, even though the time bottleneck seems to come from the data generation.\n\nHere are a few points mentionned that can easily be implemented: \n\n- Dataloader: Use `num_workers > 0` and `pin_memory=True` for data loaders\n- Use `torch.backends.cudnn.benchmark=True`\n- Increase your `batch_size` as much as you can\n- Instead of `model.zero_grad()` use `for param in model.parameters(): param.grad = None`\n\nAnd more, give a look at the presentation :)",
    "996581": "Thanks! Can one get the same features with Pytorch Lightning, sounds pretty similar at first glance? Have it on my todo list to implement.",
    "996634": "Use asynchronous transfer with non_blocking=True\n\nOther additions discussed in comments ot this tweet: https://twitter.com/karpathy/status/1299921324333170689",
    "997137": "I'm using the dataloader tip for num workers and pin memory in PyTorch lightning. \n\nI would guess the other tips also apply to PL. model.zero_grad() is hidden away but I'm sure it can be modified somehow.",
    "998581": "Thank you! Thats very intresting. I will take a deeper look."
  },
  "source": "meta"
}