{
  "id": 174045,
  "title": "Don't forget to use your gpu tensor cores",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174045",
  "author_name": "",
  "post_date": "2020-08-12T05:09:18.501640Z",
  "votes": 22,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Some new gpu's (RTX and Volta) have tensor cores for mixed precision training (fp16 instead of fp32). <br>\n<a href=\"https://developer.nvidia.com/automatic-mixed-precision\" target=\"_blank\">Automatic Mixed Precision for Deep Learning</a></p>\n<p>It's really easy to implement (tensorflow, other frameworks in the article above) just wrap your optimizer:<br>\n<code>opt = tf.keras.optimizers.Adam(learning_rate=0.001)</code><br>\n<code>opt = tf.train.experimental.enable_mixed_precision_graph_rewrite(opt)</code></p>\n<p>EffNet training time faster for ~40% </p>",
  "messages": [
    {
      "id": "967212",
      "postDate": "08/12/2020 05:09:18",
      "content": "<p>Some new gpu's (RTX and Volta) have tensor cores for mixed precision training (fp16 instead of fp32). <br>\n<a href=\"https://developer.nvidia.com/automatic-mixed-precision\" target=\"_blank\">Automatic Mixed Precision for Deep Learning</a></p>\n<p>It's really easy to implement (tensorflow, other frameworks in the article above) just wrap your optimizer:<br>\n<code>opt = tf.keras.optimizers.Adam(learning_rate=0.001)</code><br>\n<code>opt = tf.train.experimental.enable_mixed_precision_graph_rewrite(opt)</code></p>\n<p>EffNet training time faster for ~40% </p>",
      "rawMarkdown": "Some new gpu's (RTX and Volta) have tensor cores for mixed precision training (fp16 instead of fp32). \n[Automatic Mixed Precision for Deep Learning](https://developer.nvidia.com/automatic-mixed-precision)\n\nIt's really easy to implement (tensorflow, other frameworks in the article above) just wrap your optimizer:\n`opt = tf.keras.optimizers.Adam(learning_rate=0.001)`\n`opt = tf.train.experimental.enable_mixed_precision_graph_rewrite(opt)`\n\nEffNet training time faster for ~40%",
      "votes": null
    },
    {
      "id": "967509",
      "postDate": "08/12/2020 09:55:18",
      "content": "<p>thanks for mentioning it, since 6 days left have to do it!</p>",
      "rawMarkdown": "thanks for mentioning it, since 6 days left have to do it!",
      "votes": null
    },
    {
      "id": "968258",
      "postDate": "08/12/2020 20:01:49",
      "content": "<p><a href=\"https://www.kaggle.com/uspangaliev\" target=\"_blank\">@uspangaliev</a> thanks for mentioning it. -)</p>\n<p>Here is a short summary of usages for all frameworks. (Direct for the above link).</p>\n<h2>TensorFlow</h2>\n<pre><code>opt = tf.keras.optimizers.Adam(learning_rate=0.001)\nopt = tf.train.experimental.enable_mixed_precision_graph_rewrite(opt)\n</code></pre>\n<h2>PyTorch</h2>\n<pre><code>model, optimizer = amp.initialize(model, optimizer, opt_level=\"O1\")\nwith amp.scale_loss(loss, optimizer) as scaled_loss:\n    scaled_loss.backward()\n</code></pre>\n<h2>MXNet</h2>\n<pre><code>amp.init()\namp.init_trainer(trainer)\nwith amp.scale_loss(loss, trainer) as scaled_loss:\n   autograd.backward(scaled_loss)\n</code></pre>\n<h2>PaddlePaddle</h2>\n<pre><code>sgd = SGDOptimizer()\nmp_sgd = fluid.contrib.mixed_precision.decorator.decorate(sgd)\nmp_sgd.minimize(loss)\n</code></pre>\n<hr>\n<p>Also, see this discussion, it might help. <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169897\" target=\"_blank\">mixed precision with TPU tf.distrubute</a><br>\n+<br>\nCheck this interesting discussion, might help. - <a href=\"https://github.com/tensorflow/tensorflow/issues/34406\" target=\"_blank\">Mixed precision training with tf.keras 2.0</a> </p>",
      "rawMarkdown": "uspangaliev thanks for mentioning it. -)\n\nHere is a short summary of usages for all frameworks. (Direct for the above link).\n\n## TensorFlow\n```\nopt = tf.keras.optimizers.Adam(learning_rate=0.001)\nopt = tf.train.experimental.enable_mixed_precision_graph_rewrite(opt)\n```\n\n## PyTorch\n```\nmodel, optimizer = amp.initialize(model, optimizer, opt_level=\"O1\")\nwith amp.scale_loss(loss, optimizer) as scaled_loss:\n    scaled_loss.backward()\n```\n## MXNet\n```\namp.init()\namp.init_trainer(trainer)\nwith amp.scale_loss(loss, trainer) as scaled_loss:\n   autograd.backward(scaled_loss)\n```\n\n## PaddlePaddle\n```\nsgd = SGDOptimizer()\nmp_sgd = fluid.contrib.mixed_precision.decorator.decorate(sgd)\nmp_sgd.minimize(loss)\n```\n\n---\n\nAlso, see this discussion, it might help. [mixed precision with TPU tf.distrubute](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169897)\n+\nCheck this interesting discussion, might help. - [Mixed precision training with tf.keras 2.0](https://github.com/tensorflow/tensorflow/issues/34406)",
      "votes": null
    },
    {
      "id": "968342",
      "postDate": "08/12/2020 22:53:18",
      "content": "<p>Efficientnet don't use tensorcores on gpu.  It is why they run faster on tpu.</p>",
      "rawMarkdown": "Efficientnet don't use tensorcores on gpu.  It is why they run faster on tpu.",
      "votes": null
    },
    {
      "id": "968357",
      "postDate": "08/12/2020 23:29:38",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "968981",
      "postDate": "08/13/2020 11:36:30",
      "content": "<p>With PyTorch 1.6, now you can use FP16 with Pascal architecture GPUs as well. You can have a look at my <a href=\"https://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments\">kernel</a> where I use FP16 with GPU (PyTorch 1.6) instead of TPU. It's really easy and stable as well. Works with GTX 10 series GPU like a charm.\n<a href=\"https://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments\">https://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments</a></p>",
      "rawMarkdown": "With PyTorch 1.6, now you can use FP16 with Pascal architecture GPUs as well. You can have a look at my [kernel](https://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments) where I use FP16 with GPU (PyTorch 1.6) instead of TPU. It's really easy and stable as well. Works with GTX 10 series GPU like a charm.\nhttps://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments",
      "votes": null
    },
    {
      "id": "969086",
      "postDate": "08/13/2020 12:56:09",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> is there a principle why that happens? </p>",
      "rawMarkdown": "cpmpml is there a principle why that happens?",
      "votes": null
    },
    {
      "id": "969200",
      "postDate": "08/13/2020 14:41:41",
      "content": "<p>Some fp16 kernels are missing in cuda.  We are looking into it. The P100 available in Kaggle kernels have no tensorcores anyway, hence the issue is not relevant.  But is is if you use V100 for instance.</p>\n<p>Good news is that using mixed precision helps GPU still.  I am using amp in pytorch, and TF users should also use mixed precision.  It helps, even if tensorcores aren't used.</p>",
      "rawMarkdown": "Some fp16 kernels are missing in cuda.  We are looking into it. The P100 available in Kaggle kernels have no tensorcores anyway, hence the issue is not relevant.  But is is if you use V100 for instance.\n\nGood news is that using mixed precision helps GPU still.  I am using amp in pytorch, and TF users should also use mixed precision.  It helps, even if tensorcores aren't used.",
      "votes": null
    },
    {
      "id": "969286",
      "postDate": "08/13/2020 15:35:28",
      "content": "<p>Exactly. Using AMP with even P100 GPUs we can literally use double the batch size of what we can without AMP. It helps a lot and the training is faster as well. It has become even easier with PyTorch 1.6.</p>",
      "rawMarkdown": "Exactly. Using AMP with even P100 GPUs we can literally use double the batch size of what we can without AMP. It helps a lot and the training is faster as well. It has become even easier with PyTorch 1.6.",
      "votes": null
    },
    {
      "id": "970458",
      "postDate": "08/14/2020 13:25:27",
      "content": "<p>Thanx for the information :)</p>",
      "rawMarkdown": "Thanx for the information :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 967509,
      "author_name": "yash612",
      "author_url": "",
      "post_date": "08/12/2020 09:55:18",
      "content": "<p>thanks for mentioning it, since 6 days left have to do it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 968258,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "08/12/2020 20:01:49",
      "content": "<p><a href=\"https://www.kaggle.com/uspangaliev\" target=\"_blank\">@uspangaliev</a> thanks for mentioning it. -)</p>\n<p>Here is a short summary of usages for all frameworks. (Direct for the above link).</p>\n<h2>TensorFlow</h2>\n<pre><code>opt = tf.keras.optimizers.Adam(learning_rate=0.001)\nopt = tf.train.experimental.enable_mixed_precision_graph_rewrite(opt)\n</code></pre>\n<h2>PyTorch</h2>\n<pre><code>model, optimizer = amp.initialize(model, optimizer, opt_level=\"O1\")\nwith amp.scale_loss(loss, optimizer) as scaled_loss:\n    scaled_loss.backward()\n</code></pre>\n<h2>MXNet</h2>\n<pre><code>amp.init()\namp.init_trainer(trainer)\nwith amp.scale_loss(loss, trainer) as scaled_loss:\n   autograd.backward(scaled_loss)\n</code></pre>\n<h2>PaddlePaddle</h2>\n<pre><code>sgd = SGDOptimizer()\nmp_sgd = fluid.contrib.mixed_precision.decorator.decorate(sgd)\nmp_sgd.minimize(loss)\n</code></pre>\n<hr>\n<p>Also, see this discussion, it might help. <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169897\" target=\"_blank\">mixed precision with TPU tf.distrubute</a><br>\n+<br>\nCheck this interesting discussion, might help. - <a href=\"https://github.com/tensorflow/tensorflow/issues/34406\" target=\"_blank\">Mixed precision training with tf.keras 2.0</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 968342,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/12/2020 22:53:18",
      "content": "<p>Efficientnet don't use tensorcores on gpu.  It is why they run faster on tpu.</p>",
      "votes": null,
      "replies": [
        {
          "id": 969086,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "08/13/2020 12:56:09",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> is there a principle why that happens? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 969200,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/13/2020 14:41:41",
          "content": "<p>Some fp16 kernels are missing in cuda.  We are looking into it. The P100 available in Kaggle kernels have no tensorcores anyway, hence the issue is not relevant.  But is is if you use V100 for instance.</p>\n<p>Good news is that using mixed precision helps GPU still.  I am using amp in pytorch, and TF users should also use mixed precision.  It helps, even if tensorcores aren't used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 969286,
          "author_name": "sovitrath",
          "author_url": "",
          "post_date": "08/13/2020 15:35:28",
          "content": "<p>Exactly. Using AMP with even P100 GPUs we can literally use double the batch size of what we can without AMP. It helps a lot and the training is faster as well. It has become even easier with PyTorch 1.6.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 968357,
      "author_name": "darkcore",
      "author_url": "",
      "post_date": "08/12/2020 23:29:38",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 970458,
      "author_name": "luisr55",
      "author_url": "",
      "post_date": "08/14/2020 13:25:27",
      "content": "<p>Thanx for the information :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 968981,
      "author_name": "sovitrath",
      "author_url": "",
      "post_date": "08/13/2020 11:36:30",
      "content": "<p>With PyTorch 1.6, now you can use FP16 with Pascal architecture GPUs as well. You can have a look at my <a href=\"https://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments\">kernel</a> where I use FP16 with GPU (PyTorch 1.6) instead of TPU. It's really easy and stable as well. Works with GTX 10 series GPU like a charm.\n<a href=\"https://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments\">https://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "967212": "Some new gpu's (RTX and Volta) have tensor cores for mixed precision training (fp16 instead of fp32). \n[Automatic Mixed Precision for Deep Learning](https://developer.nvidia.com/automatic-mixed-precision)\n\nIt's really easy to implement (tensorflow, other frameworks in the article above) just wrap your optimizer:\n`opt = tf.keras.optimizers.Adam(learning_rate=0.001)`\n`opt = tf.train.experimental.enable_mixed_precision_graph_rewrite(opt)`\n\nEffNet training time faster for ~40%",
    "967509": "thanks for mentioning it, since 6 days left have to do it!",
    "968258": "uspangaliev thanks for mentioning it. -)\n\nHere is a short summary of usages for all frameworks. (Direct for the above link).\n\n## TensorFlow\n```\nopt = tf.keras.optimizers.Adam(learning_rate=0.001)\nopt = tf.train.experimental.enable_mixed_precision_graph_rewrite(opt)\n```\n\n## PyTorch\n```\nmodel, optimizer = amp.initialize(model, optimizer, opt_level=\"O1\")\nwith amp.scale_loss(loss, optimizer) as scaled_loss:\n    scaled_loss.backward()\n```\n## MXNet\n```\namp.init()\namp.init_trainer(trainer)\nwith amp.scale_loss(loss, trainer) as scaled_loss:\n   autograd.backward(scaled_loss)\n```\n\n## PaddlePaddle\n```\nsgd = SGDOptimizer()\nmp_sgd = fluid.contrib.mixed_precision.decorator.decorate(sgd)\nmp_sgd.minimize(loss)\n```\n\n---\n\nAlso, see this discussion, it might help. [mixed precision with TPU tf.distrubute](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169897)\n+\nCheck this interesting discussion, might help. - [Mixed precision training with tf.keras 2.0](https://github.com/tensorflow/tensorflow/issues/34406)",
    "968342": "Efficientnet don't use tensorcores on gpu.  It is why they run faster on tpu.",
    "968357": "Thanks for sharing!",
    "968981": "With PyTorch 1.6, now you can use FP16 with Pascal architecture GPUs as well. You can have a look at my [kernel](https://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments) where I use FP16 with GPU (PyTorch 1.6) instead of TPU. It's really easy and stable as well. Works with GTX 10 series GPU like a charm.\nhttps://www.kaggle.com/sovitrath/caltech-256-amp-pytorch/comments",
    "969086": "cpmpml is there a principle why that happens?",
    "969200": "Some fp16 kernels are missing in cuda.  We are looking into it. The P100 available in Kaggle kernels have no tensorcores anyway, hence the issue is not relevant.  But is is if you use V100 for instance.\n\nGood news is that using mixed precision helps GPU still.  I am using amp in pytorch, and TF users should also use mixed precision.  It helps, even if tensorcores aren't used.",
    "969286": "Exactly. Using AMP with even P100 GPUs we can literally use double the batch size of what we can without AMP. It helps a lot and the training is faster as well. It has become even easier with PyTorch 1.6.",
    "970458": "Thanx for the information :)"
  },
  "source": "meta"
}