{
  "id": 112582,
  "title": "More Tricks to Train w/ Bigger Batches (pytorch)",
  "url": "/competitions/understanding_cloud_organization/discussion/112582",
  "author_name": "Hanke Chen",
  "post_date": "2019-10-14T03:11:37.416000",
  "votes": 30,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Menu: Saving GPU Memory w/ Zero Cost</h1>\n\n<ol>\n<li>0x00 Checklist of small tricks and tips</li>\n<li>0x01 FP16 &amp; Better Code Style</li>\n<li>0x02 Optimizing Dataloader</li>\n<li>0x03 Data-Parallel w/ APEX</li>\n</ol>\n\n<p>https://miro.medium.com/max/700/0*Cqwuv9s1gI_rqHSw\" title=\"\" /&gt;</p>\n\n<h2>0x00 Checklist of small tricks and tips</h2>\n\n<p>Here you are:\n - Using <code>inplace=True</code> when using activation functions like <code>ReLu</code>. This disables saving intermediate feature maps in your GPU.\n - Using <code>eval()</code> and <code>with torch.no_grad():</code> when appropriate. <code>eval()</code> tells the model to switch off the calculation of batch_norm and the dropout. \n - DO NOT USE <code>torch.cuda.emprty_cache()</code> Although this will make the free memory used by torch visible by <code>nvidia-smi</code>, it does not actually reduce any memory. PyTorch did it all for you automatically. Some users even reported a small <a href=\"https://discuss.pytorch.org/t/about-torch-cuda-empty-cache/34232/5\">delay</a>.\n - Careful: when the number of images in EACH GPU is smaller than 8, batch_norm can be unstable, unless you use <code>sync_bn</code> in <code>APEX</code> (described below)\n - Unlike some tutorial, <code>torch.backends.cudnn.deterministic = True</code> does not speed up calculation. It is conformed by multiple users that <code>torch.backends.cudnn.deterministic = False</code> is slower than <code>torch.backends.cudnn.deterministic = True</code> on MNIST dataset. It seems that the actual effect of this setting depends on the dataset. You can set it to <code>True</code> it if you like.\n - Do not take out the output mask when you do not need it for calculation. The flow of data between GPU and CPU cost time and CPU memory.</p>\n\n<h2>0x01 FP16 &amp; Better Code Style</h2>\n\n<p>Described in <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/111904\">the other thread</a>.</p>\n\n<h2>0x02 Optimizing Dataloader</h2>\n\n<p>Set your <code>train_loader</code> as following:\n - <code>num_worker</code>: the number of CPU threads when loading data. I usually set it to the amount of CPU I have on the machine. It is recommended that value should be greater than or equal to the amount of CPU on the machine.\n - <code>pin_memory</code>: whether to load data into RAM before loading it into GPU. If you are using a server, you can set it to True.\n - <code>drop_last</code>: whether to drop the last batch in an epoch if the last batch has fewer images than your <code>batch_size</code>. This can stabilize your training.</p>\n\n<p><code>\ndata_loader = data.DataLoader(YOUR_PYTORCH_DATASET,\n                              num_workers=THE_NUMBER_OF_CPU_I_HAVE,\n                              pin_memory=True,\n                              drop_last=True,  # Last batch will mess up with batch norm https://github.com/pytorch/pytorch/issues/4534\n                              ))\n</code>\nIf your <code>pin_memory = True</code>, you can also set <code>non_blocking=True</code> to reduce the waiting time for  GPUs. For more info: <a href=\"https://devblogs.nvidia.com/how-overlap-data-transfers-cuda-cc/\">https://devblogs.nvidia.com/how-overlap-data-transfers-cuda-cc/</a></p>\n\n<p>```\n\"\"\"Sync Point\"\"\"\nimage = image.cuda(non_blocking=True)\nlabels = labels.cuda(non_blocking=True).float()</p>\n\n<p>\"\"\"Async Point\"\"\"\nprediction = net(image)\n```</p>\n\n<p><img src=\"https://devblogs.nvidia.com/wp-content/uploads/2012/11/C1060Timeline-1024x679.png\" alt=\"\">\n<img src=\"https://devblogs.nvidia.com/wp-content/uploads/2012/11/C2050Timeline-1024x670.png\" alt=\"\"></p>\n\n<h2>0x03 Data-Parallel w/ APEX</h2>\n\n<p>If you have multiple GPUs, you can use this method.\nVery simple: change <code>torch.nn.parallel.DistributedDataParallel</code> to <code>apex.parallel.DistributedDataParallel</code></p>\n\n<p>An example provided by APEX: <a href=\"https://github.com/NVIDIA/apex/tree/master/examples/imagenet\">https://github.com/NVIDIA/apex/tree/master/examples/imagenet</a></p>\n\n<h3>Installation Can be Tricky Though</h3>\n\n<p>Here are some tricky parts:</p>\n\nCorrect Version\n\n<ol>\n<li>Your <code>nvcc --version</code> version should match <code>&gt;&gt;&gt; torch.version.cuda</code> version. Otherwise, it will cause errors.</li>\n<li>If you are using these GPUs, you should use <code>nvcc</code> and <code>pytorch.cuda</code> 10.0 instead of 9.x\n<code>\nGeForce GTX 1650\nGeForce GTX 1660\nGeForce GTX 1660 Ti\nGeForce RTX 2060\nGeForce RTX 2060 Super\nGeForce RTX 2070\nGeForce RTX 2070 Super\nGeForce RTX 2080\nGeForce RTX 2080 Super\nGeForce RTX 2080 Ti\nTitan RTX\nQuadro RTX 4000\nQuadro RTX 5000\nQuadro RTX 6000\nQuadro RTX 8000\nTesla T4\n</code>\nOtherwise, use <code>nvcc</code> and <code>pytorch.cuda</code> 9.2.\nIf you need to re-install pytorch.cuda: <a href=\"https://link.zhihu.com/?target=https%3A//pytorch.org/get-started/locally/\">HERE</a>\nIf you need to re-install nvcc: <a href=\"https://link.zhihu.com/?target=https%3A//www.pugetsystems.com/labs/hpc/How-to-install-CUDA-9-2-on-Ubuntu-18-04-1184/\">NVCC9.2</a>, <a href=\"https://link.zhihu.com/?target=https%3A//www.pugetsystems.com/labs/hpc/How-To-Install-CUDA-10-together-with-9-2-on-Ubuntu-18-04-with-support-for-NVIDIA-20XX-Turing-GPUs-1236/\">NVCC10.0</a>\nAnd then you have to re-install APEX all over again.</li>\n</ol>\n\n<p>.\n.\n.\n(Oh my God. Kaggle has a bug that the whole article will disappear when resizing the browser, and I have to write the entire article all over again =. =  And <code>Ctrl + Z</code> does not work either.)\n.\nHope it helps. Please upvote if you learned something here, thanks!\nFeel free to correct me or give us some extra suggestions :)</p>",
  "messages": [
    {
      "id": 648325,
      "postDate": "2019-10-14T03:11:37.417Z",
      "content": "<h1>Menu: Saving GPU Memory w/ Zero Cost</h1>\n\n<ol>\n<li>0x00 Checklist of small tricks and tips</li>\n<li>0x01 FP16 &amp; Better Code Style</li>\n<li>0x02 Optimizing Dataloader</li>\n<li>0x03 Data-Parallel w/ APEX</li>\n</ol>\n\n<p>https://miro.medium.com/max/700/0*Cqwuv9s1gI_rqHSw\" title=\"\" /&gt;</p>\n\n<h2>0x00 Checklist of small tricks and tips</h2>\n\n<p>Here you are:\n - Using <code>inplace=True</code> when using activation functions like <code>ReLu</code>. This disables saving intermediate feature maps in your GPU.\n - Using <code>eval()</code> and <code>with torch.no_grad():</code> when appropriate. <code>eval()</code> tells the model to switch off the calculation of batch_norm and the dropout. \n - DO NOT USE <code>torch.cuda.emprty_cache()</code> Although this will make the free memory used by torch visible by <code>nvidia-smi</code>, it does not actually reduce any memory. PyTorch did it all for you automatically. Some users even reported a small <a href=\"https://discuss.pytorch.org/t/about-torch-cuda-empty-cache/34232/5\">delay</a>.\n - Careful: when the number of images in EACH GPU is smaller than 8, batch_norm can be unstable, unless you use <code>sync_bn</code> in <code>APEX</code> (described below)\n - Unlike some tutorial, <code>torch.backends.cudnn.deterministic = True</code> does not speed up calculation. It is conformed by multiple users that <code>torch.backends.cudnn.deterministic = False</code> is slower than <code>torch.backends.cudnn.deterministic = True</code> on MNIST dataset. It seems that the actual effect of this setting depends on the dataset. You can set it to <code>True</code> it if you like.\n - Do not take out the output mask when you do not need it for calculation. The flow of data between GPU and CPU cost time and CPU memory.</p>\n\n<h2>0x01 FP16 &amp; Better Code Style</h2>\n\n<p>Described in <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/111904\">the other thread</a>.</p>\n\n<h2>0x02 Optimizing Dataloader</h2>\n\n<p>Set your <code>train_loader</code> as following:\n - <code>num_worker</code>: the number of CPU threads when loading data. I usually set it to the amount of CPU I have on the machine. It is recommended that value should be greater than or equal to the amount of CPU on the machine.\n - <code>pin_memory</code>: whether to load data into RAM before loading it into GPU. If you are using a server, you can set it to True.\n - <code>drop_last</code>: whether to drop the last batch in an epoch if the last batch has fewer images than your <code>batch_size</code>. This can stabilize your training.</p>\n\n<p><code>\ndata_loader = data.DataLoader(YOUR_PYTORCH_DATASET,\n                              num_workers=THE_NUMBER_OF_CPU_I_HAVE,\n                              pin_memory=True,\n                              drop_last=True,  # Last batch will mess up with batch norm https://github.com/pytorch/pytorch/issues/4534\n                              ))\n</code>\nIf your <code>pin_memory = True</code>, you can also set <code>non_blocking=True</code> to reduce the waiting time for  GPUs. For more info: <a href=\"https://devblogs.nvidia.com/how-overlap-data-transfers-cuda-cc/\">https://devblogs.nvidia.com/how-overlap-data-transfers-cuda-cc/</a></p>\n\n<p>```\n\"\"\"Sync Point\"\"\"\nimage = image.cuda(non_blocking=True)\nlabels = labels.cuda(non_blocking=True).float()</p>\n\n<p>\"\"\"Async Point\"\"\"\nprediction = net(image)\n```</p>\n\n<p><img src=\"https://devblogs.nvidia.com/wp-content/uploads/2012/11/C1060Timeline-1024x679.png\" alt=\"\">\n<img src=\"https://devblogs.nvidia.com/wp-content/uploads/2012/11/C2050Timeline-1024x670.png\" alt=\"\"></p>\n\n<h2>0x03 Data-Parallel w/ APEX</h2>\n\n<p>If you have multiple GPUs, you can use this method.\nVery simple: change <code>torch.nn.parallel.DistributedDataParallel</code> to <code>apex.parallel.DistributedDataParallel</code></p>\n\n<p>An example provided by APEX: <a href=\"https://github.com/NVIDIA/apex/tree/master/examples/imagenet\">https://github.com/NVIDIA/apex/tree/master/examples/imagenet</a></p>\n\n<h3>Installation Can be Tricky Though</h3>\n\n<p>Here are some tricky parts:</p>\n\nCorrect Version\n\n<ol>\n<li>Your <code>nvcc --version</code> version should match <code>&gt;&gt;&gt; torch.version.cuda</code> version. Otherwise, it will cause errors.</li>\n<li>If you are using these GPUs, you should use <code>nvcc</code> and <code>pytorch.cuda</code> 10.0 instead of 9.x\n<code>\nGeForce GTX 1650\nGeForce GTX 1660\nGeForce GTX 1660 Ti\nGeForce RTX 2060\nGeForce RTX 2060 Super\nGeForce RTX 2070\nGeForce RTX 2070 Super\nGeForce RTX 2080\nGeForce RTX 2080 Super\nGeForce RTX 2080 Ti\nTitan RTX\nQuadro RTX 4000\nQuadro RTX 5000\nQuadro RTX 6000\nQuadro RTX 8000\nTesla T4\n</code>\nOtherwise, use <code>nvcc</code> and <code>pytorch.cuda</code> 9.2.\nIf you need to re-install pytorch.cuda: <a href=\"https://link.zhihu.com/?target=https%3A//pytorch.org/get-started/locally/\">HERE</a>\nIf you need to re-install nvcc: <a href=\"https://link.zhihu.com/?target=https%3A//www.pugetsystems.com/labs/hpc/How-to-install-CUDA-9-2-on-Ubuntu-18-04-1184/\">NVCC9.2</a>, <a href=\"https://link.zhihu.com/?target=https%3A//www.pugetsystems.com/labs/hpc/How-To-Install-CUDA-10-together-with-9-2-on-Ubuntu-18-04-with-support-for-NVIDIA-20XX-Turing-GPUs-1236/\">NVCC10.0</a>\nAnd then you have to re-install APEX all over again.</li>\n</ol>\n\n<p>.\n.\n.\n(Oh my God. Kaggle has a bug that the whole article will disappear when resizing the browser, and I have to write the entire article all over again =. =  And <code>Ctrl + Z</code> does not work either.)\n.\nHope it helps. Please upvote if you learned something here, thanks!\nFeel free to correct me or give us some extra suggestions :)</p>",
      "rawMarkdown": "# Menu: Saving GPU Memory w/ Zero Cost\n1. 0x00 Checklist of small tricks and tips\n2. 0x01 FP16 &amp; Better Code Style\n3. 0x02 Optimizing Dataloader\n4. 0x03 Data-Parallel w/ APEX\n\n![GPU Memory Usage: https://miro.medium.com/max/700/0*Cqwuv9s1gI_rqHSw](https://pic4.zhimg.com/80/v2-022fb85ea6fd73d54e74e1d43f221093_hd.jpg)\n\n## 0x00 Checklist of small tricks and tips\nHere you are:\n - Using `inplace=True` when using activation functions like `ReLu`. This disables saving intermediate feature maps in your GPU.\n - Using `eval()` and `with torch.no_grad():` when appropriate. `eval()` tells the model to switch off the calculation of batch_norm and the dropout. \n - DO NOT USE `torch.cuda.emprty_cache()` Although this will make the free memory used by torch visible by `nvidia-smi`, it does not actually reduce any memory. PyTorch did it all for you automatically. Some users even reported a small [delay](https://discuss.pytorch.org/t/about-torch-cuda-empty-cache/34232/5).\n - Careful: when the number of images in EACH GPU is smaller than 8, batch_norm can be unstable, unless you use `sync_bn` in `APEX` (described below)\n - Unlike some tutorial, `torch.backends.cudnn.deterministic = True` does not speed up calculation. It is conformed by multiple users that `torch.backends.cudnn.deterministic = False` is slower than `torch.backends.cudnn.deterministic = True` on MNIST dataset. It seems that the actual effect of this setting depends on the dataset. You can set it to `True` it if you like.\n - Do not take out the output mask when you do not need it for calculation. The flow of data between GPU and CPU cost time and CPU memory.\n\n## 0x01 FP16 &amp; Better Code Style\nDescribed in [the other thread](https://www.kaggle.com/c/understanding_cloud_organization/discussion/111904).\n\n## 0x02 Optimizing Dataloader\nSet your `train_loader` as following:\n - `num_worker`: the number of CPU threads when loading data. I usually set it to the amount of CPU I have on the machine. It is recommended that value should be greater than or equal to the amount of CPU on the machine.\n - `pin_memory`: whether to load data into RAM before loading it into GPU. If you are using a server, you can set it to True.\n - `drop_last`: whether to drop the last batch in an epoch if the last batch has fewer images than your `batch_size`. This can stabilize your training.\n\n```\ndata_loader = data.DataLoader(YOUR_PYTORCH_DATASET,\n                              num_workers=THE_NUMBER_OF_CPU_I_HAVE,\n                              pin_memory=True,\n                              drop_last=True,  # Last batch will mess up with batch norm https://github.com/pytorch/pytorch/issues/4534\n                              ))\n```\nIf your `pin_memory = True`, you can also set `non_blocking=True` to reduce the waiting time for  GPUs. For more info: https://devblogs.nvidia.com/how-overlap-data-transfers-cuda-cc/\n\n```\n\"\"\"Sync Point\"\"\"\nimage = image.cuda(non_blocking=True)\nlabels = labels.cuda(non_blocking=True).float()\n\n\"\"\"Async Point\"\"\"\nprediction = net(image)\n```\n\n![](https://devblogs.nvidia.com/wp-content/uploads/2012/11/C1060Timeline-1024x679.png)\n![](https://devblogs.nvidia.com/wp-content/uploads/2012/11/C2050Timeline-1024x670.png)\n\n## 0x03 Data-Parallel w/ APEX\nIf you have multiple GPUs, you can use this method.\nVery simple: change `torch.nn.parallel.DistributedDataParallel` to `apex.parallel.DistributedDataParallel`\n\nAn example provided by APEX: https://github.com/NVIDIA/apex/tree/master/examples/imagenet\n\n### Installation Can be Tricky Though\nHere are some tricky parts:\n\n#### Correct Version\n1. Your `nvcc --version` version should match `&gt;&gt;&gt; torch.version.cuda` version. Otherwise, it will cause errors.\n2. If you are using these GPUs, you should use `nvcc` and `pytorch.cuda` 10.0 instead of 9.x\n```\nGeForce GTX 1650\nGeForce GTX 1660\nGeForce GTX 1660 Ti\nGeForce RTX 2060\nGeForce RTX 2060 Super\nGeForce RTX 2070\nGeForce RTX 2070 Super\nGeForce RTX 2080\nGeForce RTX 2080 Super\nGeForce RTX 2080 Ti\nTitan RTX\nQuadro RTX 4000\nQuadro RTX 5000\nQuadro RTX 6000\nQuadro RTX 8000\nTesla T4\n```\nOtherwise, use `nvcc` and `pytorch.cuda` 9.2.\nIf you need to re-install pytorch.cuda: [HERE](https://link.zhihu.com/?target=https%3A//pytorch.org/get-started/locally/)\nIf you need to re-install nvcc: [NVCC9.2](https://link.zhihu.com/?target=https%3A//www.pugetsystems.com/labs/hpc/How-to-install-CUDA-9-2-on-Ubuntu-18-04-1184/), [NVCC10.0](https://link.zhihu.com/?target=https%3A//www.pugetsystems.com/labs/hpc/How-To-Install-CUDA-10-together-with-9-2-on-Ubuntu-18-04-with-support-for-NVIDIA-20XX-Turing-GPUs-1236/)\nAnd then you have to re-install APEX all over again.\n\n.\n.\n.\n(Oh my God. Kaggle has a bug that the whole article will disappear when resizing the browser, and I have to write the entire article all over again =. =  And `Ctrl + Z` does not work either.)\n.\nHope it helps. Please upvote if you learned something here, thanks!\nFeel free to correct me or give us some extra suggestions :)",
      "votes": 29
    },
    {
      "id": 649094,
      "postDate": "2019-10-15T00:24:12.690Z",
      "content": "<p>Excellent references, thanks for posting it!!</p>",
      "rawMarkdown": "Excellent references, thanks for posting it!!"
    },
    {
      "id": 913696,
      "postDate": "2020-07-03T11:07:26.313Z",
      "content": "<p>When I set <code>num_worker</code> from 0 to 4, it would throw the OOM earlier.</p>",
      "rawMarkdown": "When I set `num_worker` from 0 to 4, it would throw the OOM earlier.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 649094,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-10-15T00:24:12.690000",
      "content": "<p>Excellent references, thanks for posting it!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 913696,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-03T11:07:26.313000",
      "content": "<p>When I set <code>num_worker</code> from 0 to 4, it would throw the OOM earlier.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "648325": "# Menu: Saving GPU Memory w/ Zero Cost\n1. 0x00 Checklist of small tricks and tips\n2. 0x01 FP16 &amp; Better Code Style\n3. 0x02 Optimizing Dataloader\n4. 0x03 Data-Parallel w/ APEX\n\n![GPU Memory Usage: https://miro.medium.com/max/700/0*Cqwuv9s1gI_rqHSw](https://pic4.zhimg.com/80/v2-022fb85ea6fd73d54e74e1d43f221093_hd.jpg)\n\n## 0x00 Checklist of small tricks and tips\nHere you are:\n - Using `inplace=True` when using activation functions like `ReLu`. This disables saving intermediate feature maps in your GPU.\n - Using `eval()` and `with torch.no_grad():` when appropriate. `eval()` tells the model to switch off the calculation of batch_norm and the dropout. \n - DO NOT USE `torch.cuda.emprty_cache()` Although this will make the free memory used by torch visible by `nvidia-smi`, it does not actually reduce any memory. PyTorch did it all for you automatically. Some users even reported a small [delay](https://discuss.pytorch.org/t/about-torch-cuda-empty-cache/34232/5).\n - Careful: when the number of images in EACH GPU is smaller than 8, batch_norm can be unstable, unless you use `sync_bn` in `APEX` (described below)\n - Unlike some tutorial, `torch.backends.cudnn.deterministic = True` does not speed up calculation. It is conformed by multiple users that `torch.backends.cudnn.deterministic = False` is slower than `torch.backends.cudnn.deterministic = True` on MNIST dataset. It seems that the actual effect of this setting depends on the dataset. You can set it to `True` it if you like.\n - Do not take out the output mask when you do not need it for calculation. The flow of data between GPU and CPU cost time and CPU memory.\n\n## 0x01 FP16 &amp; Better Code Style\nDescribed in [the other thread](https://www.kaggle.com/c/understanding_cloud_organization/discussion/111904).\n\n## 0x02 Optimizing Dataloader\nSet your `train_loader` as following:\n - `num_worker`: the number of CPU threads when loading data. I usually set it to the amount of CPU I have on the machine. It is recommended that value should be greater than or equal to the amount of CPU on the machine.\n - `pin_memory`: whether to load data into RAM before loading it into GPU. If you are using a server, you can set it to True.\n - `drop_last`: whether to drop the last batch in an epoch if the last batch has fewer images than your `batch_size`. This can stabilize your training.\n\n```\ndata_loader = data.DataLoader(YOUR_PYTORCH_DATASET,\n                              num_workers=THE_NUMBER_OF_CPU_I_HAVE,\n                              pin_memory=True,\n                              drop_last=True,  # Last batch will mess up with batch norm https://github.com/pytorch/pytorch/issues/4534\n                              ))\n```\nIf your `pin_memory = True`, you can also set `non_blocking=True` to reduce the waiting time for  GPUs. For more info: https://devblogs.nvidia.com/how-overlap-data-transfers-cuda-cc/\n\n```\n\"\"\"Sync Point\"\"\"\nimage = image.cuda(non_blocking=True)\nlabels = labels.cuda(non_blocking=True).float()\n\n\"\"\"Async Point\"\"\"\nprediction = net(image)\n```\n\n![](https://devblogs.nvidia.com/wp-content/uploads/2012/11/C1060Timeline-1024x679.png)\n![](https://devblogs.nvidia.com/wp-content/uploads/2012/11/C2050Timeline-1024x670.png)\n\n## 0x03 Data-Parallel w/ APEX\nIf you have multiple GPUs, you can use this method.\nVery simple: change `torch.nn.parallel.DistributedDataParallel` to `apex.parallel.DistributedDataParallel`\n\nAn example provided by APEX: https://github.com/NVIDIA/apex/tree/master/examples/imagenet\n\n### Installation Can be Tricky Though\nHere are some tricky parts:\n\n#### Correct Version\n1. Your `nvcc --version` version should match `&gt;&gt;&gt; torch.version.cuda` version. Otherwise, it will cause errors.\n2. If you are using these GPUs, you should use `nvcc` and `pytorch.cuda` 10.0 instead of 9.x\n```\nGeForce GTX 1650\nGeForce GTX 1660\nGeForce GTX 1660 Ti\nGeForce RTX 2060\nGeForce RTX 2060 Super\nGeForce RTX 2070\nGeForce RTX 2070 Super\nGeForce RTX 2080\nGeForce RTX 2080 Super\nGeForce RTX 2080 Ti\nTitan RTX\nQuadro RTX 4000\nQuadro RTX 5000\nQuadro RTX 6000\nQuadro RTX 8000\nTesla T4\n```\nOtherwise, use `nvcc` and `pytorch.cuda` 9.2.\nIf you need to re-install pytorch.cuda: [HERE](https://link.zhihu.com/?target=https%3A//pytorch.org/get-started/locally/)\nIf you need to re-install nvcc: [NVCC9.2](https://link.zhihu.com/?target=https%3A//www.pugetsystems.com/labs/hpc/How-to-install-CUDA-9-2-on-Ubuntu-18-04-1184/), [NVCC10.0](https://link.zhihu.com/?target=https%3A//www.pugetsystems.com/labs/hpc/How-To-Install-CUDA-10-together-with-9-2-on-Ubuntu-18-04-with-support-for-NVIDIA-20XX-Turing-GPUs-1236/)\nAnd then you have to re-install APEX all over again.\n\n.\n.\n.\n(Oh my God. Kaggle has a bug that the whole article will disappear when resizing the browser, and I have to write the entire article all over again =. =  And `Ctrl + Z` does not work either.)\n.\nHope it helps. Please upvote if you learned something here, thanks!\nFeel free to correct me or give us some extra suggestions :)",
    "649094": "Excellent references, thanks for posting it!!",
    "913696": "When I set `num_worker` from 0 to 4, it would throw the OOM earlier."
  }
}