{
  "id": 164309,
  "title": "Native Mixed Precision in PyTorch1.6+🔥",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164309",
  "author_name": "",
  "post_date": "2020-07-05T16:22:08.828674900Z",
  "votes": 21,
  "comment_count": 8,
  "views": 0,
  "content": "<p>PyTorch 1.6.0 is going to be officially released soon, but if you haven't seen it yet, it now finally supports Automatic Mixed Precision natively, through <code>torch.cuda.amp</code>\nI've updated my lightning kernel to use it:\n<a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\">https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp</a></p>\n\n<p>If you already use PyTorch Lightning, you don't really need to change anything in your code apart from setting <code>precision=16</code> and updating your pytorch version:\n<code>!pip install --pre torch==1.6.0.dev20200625+cu101 torchvision==0.7.0.dev20200625+cu101 -f https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html</code>\n(or any higher dev version)</p>\n\n<p>If you aren't familiar with AMP here is an article from NVIDIA <a href=\"https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html\">https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html</a></p>\n\n<p><strong>TL;DR: It allows you to double your max batch size per GPU and speed up the training</strong></p>\n\n<p><img src=\"https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/graphics/training-iteration.png\" alt=\"MP\"></p>\n\n<p>The same can be achieved in <strong>tf.keras</strong> by setting <code>auto_mixed_precision</code> and wrapping your optimizer with a LossScaleOptimizer\n<code>tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})</code>\n<code>opt = tf.keras.mixed_precision.experimental.LossScaleOptimizer(opt, \"dynamic\")</code></p>",
  "messages": [
    {
      "id": "916421",
      "postDate": "07/05/2020 16:22:08",
      "content": "<p>PyTorch 1.6.0 is going to be officially released soon, but if you haven't seen it yet, it now finally supports Automatic Mixed Precision natively, through <code>torch.cuda.amp</code>\nI've updated my lightning kernel to use it:\n<a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\">https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp</a></p>\n\n<p>If you already use PyTorch Lightning, you don't really need to change anything in your code apart from setting <code>precision=16</code> and updating your pytorch version:\n<code>!pip install --pre torch==1.6.0.dev20200625+cu101 torchvision==0.7.0.dev20200625+cu101 -f https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html</code>\n(or any higher dev version)</p>\n\n<p>If you aren't familiar with AMP here is an article from NVIDIA <a href=\"https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html\">https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html</a></p>\n\n<p><strong>TL;DR: It allows you to double your max batch size per GPU and speed up the training</strong></p>\n\n<p><img src=\"https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/graphics/training-iteration.png\" alt=\"MP\"></p>\n\n<p>The same can be achieved in <strong>tf.keras</strong> by setting <code>auto_mixed_precision</code> and wrapping your optimizer with a LossScaleOptimizer\n<code>tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})</code>\n<code>opt = tf.keras.mixed_precision.experimental.LossScaleOptimizer(opt, \"dynamic\")</code></p>",
      "rawMarkdown": "PyTorch 1.6.0 is going to be officially released soon, but if you haven't seen it yet, it now finally supports Automatic Mixed Precision natively, through `torch.cuda.amp`\nI've updated my lightning kernel to use it:\nhttps://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\n\nIf you already use PyTorch Lightning, you don't really need to change anything in your code apart from setting `precision=16` and updating your pytorch version:\n`!pip install --pre torch==1.6.0.dev20200625+cu101 torchvision==0.7.0.dev20200625+cu101 -f https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html`\n(or any higher dev version)\n\nIf you aren't familiar with AMP here is an article from NVIDIA https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html\n\n**TL;DR: It allows you to double your max batch size per GPU and speed up the training**\n\n![MP](https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/graphics/training-iteration.png)\n\nThe same can be achieved in **tf.keras** by setting `auto_mixed_precision` and wrapping your optimizer with a LossScaleOptimizer\n`tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})`\n`opt = tf.keras.mixed_precision.experimental.LossScaleOptimizer(opt, \"dynamic\")`",
      "votes": null
    },
    {
      "id": "916759",
      "postDate": "07/06/2020 02:19:57",
      "content": "<p>Is this feature also available in Tensorflow???</p>",
      "rawMarkdown": "Is this feature also available in Tensorflow???",
      "votes": null
    },
    {
      "id": "916764",
      "postDate": "07/06/2020 02:25:36",
      "content": "<p>Yes, TF supports NVIIDA AMP</p>",
      "rawMarkdown": "Yes, TF supports NVIIDA AMP",
      "votes": null
    },
    {
      "id": "916957",
      "postDate": "07/06/2020 06:37:00",
      "content": "<p>Yes <a href=\"/redwankarimsony\">@redwankarimsony</a> \nAs I was saying above you can do it by setting <code>auto_mixed_precision</code> and wrapping your optimizer with a LossScaleOptimizer\n<code>tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})</code>\n<code>opt = tf.keras.mixed_precision.experimental.LossScaleOptimizer(opt, \"dynamic\")</code></p>",
      "rawMarkdown": "Yes @redwankarimsony \nAs I was saying above you can do it by setting `auto_mixed_precision` and wrapping your optimizer with a LossScaleOptimizer\n`tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})`\n`opt = tf.keras.mixed_precision.experimental.LossScaleOptimizer(opt, \"dynamic\")`",
      "votes": null
    },
    {
      "id": "930421",
      "postDate": "07/15/2020 12:55:42",
      "content": "<p>Hello <a href=\"/hmendonca\">@hmendonca</a> ! Thanks for the great news. I have been wondering for days how to install PyTorch 1.6 on Kaggle environment. Quick question though: when the GPU usage is already of 98%, will AMP have any positive effect on training time? I could not find any clear answers.</p>",
      "rawMarkdown": "Hello @hmendonca ! Thanks for the great news. I have been wondering for days how to install PyTorch 1.6 on Kaggle environment. Quick question though: when the GPU usage is already of 98%, will AMP have any positive effect on training time? I could not find any clear answers.",
      "votes": null
    },
    {
      "id": "930574",
      "postDate": "07/15/2020 14:55:24",
      "content": "<p><a href=\"/rftexas\">@rftexas</a> to install, please look at <a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Install-modules\">https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Install-modules</a>\nIt installs 1.7.0 actually, if you want 1.6 find the correct version in <a href=\"https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html\">https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html</a></p>\n\n<p>Unfortunately, AMP doesn't speed up on Kaggle's P100, but it does (a lot) on newer architectures like Colab's T4, or V100, and will probably fly on the new A100 ;D</p>",
      "rawMarkdown": "rftexas to install, please look at https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Install-modules\nIt installs 1.7.0 actually, if you want 1.6 find the correct version in https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html\n\nUnfortunately, AMP doesn't speed up on Kaggle's P100, but it does (a lot) on newer architectures like Colab's T4, or V100, and will probably fly on the new A100 ;D",
      "votes": null
    },
    {
      "id": "931076",
      "postDate": "07/16/2020 00:30:22",
      "content": "<p>Your notebook can be used as a Tutorial itself for <strong>PyTorch Lightning</strong>. I really want to use it, i love the fact you don't need to change the code for running in GPU or TPU. Thanks for share.</p>",
      "rawMarkdown": "Your notebook can be used as a Tutorial itself for **PyTorch Lightning**. I really want to use it, i love the fact you don't need to change the code for running in GPU or TPU. Thanks for share.",
      "votes": null
    },
    {
      "id": "931357",
      "postDate": "07/16/2020 06:32:09",
      "content": "<p>You will need to be using GPU's that have good and stable Half Precision support (i.e. RTX, V100, Quadro RTX, etc) aka TensorCores.....if anyone tries this on a GPU that doesn't have TensorCores reply how it does.</p>",
      "rawMarkdown": "You will need to be using GPU's that have good and stable Half Precision support (i.e. RTX, V100, Quadro RTX, etc) aka TensorCores.....if anyone tries this on a GPU that doesn't have TensorCores reply how it does.",
      "votes": null
    },
    {
      "id": "931916",
      "postDate": "07/16/2020 14:50:59",
      "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> As I said below </p>\n\n<blockquote>\n  <p>AMP doesn't speed up on Kaggle's P100, but it does (a lot) on newer architectures like Colab's T4, or V100, and will probably fly on the new A100 ;D</p>\n</blockquote>\n\n<p>However, it still allows you to double your batch size in any GPU, which can by itself improve the model learning, accuracy, LB, etc</p>",
      "rawMarkdown": "brianfeeny As I said below \n&gt; AMP doesn't speed up on Kaggle's P100, but it does (a lot) on newer architectures like Colab's T4, or V100, and will probably fly on the new A100 ;D\n\nHowever, it still allows you to double your batch size in any GPU, which can by itself improve the model learning, accuracy, LB, etc",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 931076,
      "author_name": "hiramcho",
      "author_url": "",
      "post_date": "07/16/2020 00:30:22",
      "content": "<p>Your notebook can be used as a Tutorial itself for <strong>PyTorch Lightning</strong>. I really want to use it, i love the fact you don't need to change the code for running in GPU or TPU. Thanks for share.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 916759,
      "author_name": "redwankarimsony",
      "author_url": "",
      "post_date": "07/06/2020 02:19:57",
      "content": "<p>Is this feature also available in Tensorflow???</p>",
      "votes": null,
      "replies": [
        {
          "id": 916764,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/06/2020 02:25:36",
          "content": "<p>Yes, TF supports NVIIDA AMP</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916957,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/06/2020 06:37:00",
          "content": "<p>Yes <a href=\"/redwankarimsony\">@redwankarimsony</a> \nAs I was saying above you can do it by setting <code>auto_mixed_precision</code> and wrapping your optimizer with a LossScaleOptimizer\n<code>tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})</code>\n<code>opt = tf.keras.mixed_precision.experimental.LossScaleOptimizer(opt, \"dynamic\")</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 930421,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "07/15/2020 12:55:42",
      "content": "<p>Hello <a href=\"/hmendonca\">@hmendonca</a> ! Thanks for the great news. I have been wondering for days how to install PyTorch 1.6 on Kaggle environment. Quick question though: when the GPU usage is already of 98%, will AMP have any positive effect on training time? I could not find any clear answers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 930574,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/15/2020 14:55:24",
          "content": "<p><a href=\"/rftexas\">@rftexas</a> to install, please look at <a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Install-modules\">https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Install-modules</a>\nIt installs 1.7.0 actually, if you want 1.6 find the correct version in <a href=\"https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html\">https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html</a></p>\n\n<p>Unfortunately, AMP doesn't speed up on Kaggle's P100, but it does (a lot) on newer architectures like Colab's T4, or V100, and will probably fly on the new A100 ;D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 931357,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "07/16/2020 06:32:09",
      "content": "<p>You will need to be using GPU's that have good and stable Half Precision support (i.e. RTX, V100, Quadro RTX, etc) aka TensorCores.....if anyone tries this on a GPU that doesn't have TensorCores reply how it does.</p>",
      "votes": null,
      "replies": [
        {
          "id": 931916,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/16/2020 14:50:59",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> As I said below </p>\n\n<blockquote>\n  <p>AMP doesn't speed up on Kaggle's P100, but it does (a lot) on newer architectures like Colab's T4, or V100, and will probably fly on the new A100 ;D</p>\n</blockquote>\n\n<p>However, it still allows you to double your batch size in any GPU, which can by itself improve the model learning, accuracy, LB, etc</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "916421": "PyTorch 1.6.0 is going to be officially released soon, but if you haven't seen it yet, it now finally supports Automatic Mixed Precision natively, through `torch.cuda.amp`\nI've updated my lightning kernel to use it:\nhttps://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\n\nIf you already use PyTorch Lightning, you don't really need to change anything in your code apart from setting `precision=16` and updating your pytorch version:\n`!pip install --pre torch==1.6.0.dev20200625+cu101 torchvision==0.7.0.dev20200625+cu101 -f https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html`\n(or any higher dev version)\n\nIf you aren't familiar with AMP here is an article from NVIDIA https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html\n\n**TL;DR: It allows you to double your max batch size per GPU and speed up the training**\n\n![MP](https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/graphics/training-iteration.png)\n\nThe same can be achieved in **tf.keras** by setting `auto_mixed_precision` and wrapping your optimizer with a LossScaleOptimizer\n`tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})`\n`opt = tf.keras.mixed_precision.experimental.LossScaleOptimizer(opt, \"dynamic\")`",
    "916759": "Is this feature also available in Tensorflow???",
    "916764": "Yes, TF supports NVIIDA AMP",
    "916957": "Yes @redwankarimsony \nAs I was saying above you can do it by setting `auto_mixed_precision` and wrapping your optimizer with a LossScaleOptimizer\n`tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})`\n`opt = tf.keras.mixed_precision.experimental.LossScaleOptimizer(opt, \"dynamic\")`",
    "930421": "Hello @hmendonca ! Thanks for the great news. I have been wondering for days how to install PyTorch 1.6 on Kaggle environment. Quick question though: when the GPU usage is already of 98%, will AMP have any positive effect on training time? I could not find any clear answers.",
    "930574": "rftexas to install, please look at https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Install-modules\nIt installs 1.7.0 actually, if you want 1.6 find the correct version in https://download.pytorch.org/whl/nightly/cu101/torch_nightly.html\n\nUnfortunately, AMP doesn't speed up on Kaggle's P100, but it does (a lot) on newer architectures like Colab's T4, or V100, and will probably fly on the new A100 ;D",
    "931076": "Your notebook can be used as a Tutorial itself for **PyTorch Lightning**. I really want to use it, i love the fact you don't need to change the code for running in GPU or TPU. Thanks for share.",
    "931357": "You will need to be using GPU's that have good and stable Half Precision support (i.e. RTX, V100, Quadro RTX, etc) aka TensorCores.....if anyone tries this on a GPU that doesn't have TensorCores reply how it does.",
    "931916": "brianfeeny As I said below \n&gt; AMP doesn't speed up on Kaggle's P100, but it does (a lot) on newer architectures like Colab's T4, or V100, and will probably fly on the new A100 ;D\n\nHowever, it still allows you to double your batch size in any GPU, which can by itself improve the model learning, accuracy, LB, etc"
  },
  "source": "meta"
}