{
  "id": 37415,
  "title": "Let's learn fp16 (half float) and multi-GPU in pytorch here!",
  "url": "/competitions/carvana-image-masking-challenge/discussion/37415",
  "author_name": "",
  "post_date": "2017-08-02T00:50:09.647714800Z",
  "votes": 14,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Every kaggle competition solves a different problem and i learn a different thing. In this Carvana Image Masking Challenge, able to hange large input and output (e.g. prediction of mask at 1024x1024) is may be an advantage. </p>\n\n<p>This thread is for using fp16 (16-bit float) and multi-gpu training and inference. I hope experienced kagglers who had worked with fp16 or multi-gpu can answer our question here.</p>\n\n<p>For a start, I have probelm with BN layer.</p>\n\n<ol>\n<li><p>will BN layer be affected. (e.g. the moving vars needs high precision) by fp16?</p></li>\n<li><p>for multi gpu, how to compute BN statistics (moving mean and vars) over 4 gpu? The results doesn't seems to be stable? I read that if one solution is to freeze BN statistics (assuming that you have a pretrain model to begin with). Is that what you really do?</p></li>\n</ol>\n\n<p>There is a new paper on batch renormalisation. <a href=\"https://arxiv.org/abs/1702.03275\">https://arxiv.org/abs/1702.03275</a>.\nHas any try this and with this solve the issues above?</p>\n\n<p>I will post some code later, if the results are stable.</p>",
  "messages": [
    {
      "id": "209357",
      "postDate": "08/02/2017 00:50:09",
      "content": "<p>Every kaggle competition solves a different problem and i learn a different thing. In this Carvana Image Masking Challenge, able to hange large input and output (e.g. prediction of mask at 1024x1024) is may be an advantage. </p>\n\n<p>This thread is for using fp16 (16-bit float) and multi-gpu training and inference. I hope experienced kagglers who had worked with fp16 or multi-gpu can answer our question here.</p>\n\n<p>For a start, I have probelm with BN layer.</p>\n\n<ol>\n<li><p>will BN layer be affected. (e.g. the moving vars needs high precision) by fp16?</p></li>\n<li><p>for multi gpu, how to compute BN statistics (moving mean and vars) over 4 gpu? The results doesn't seems to be stable? I read that if one solution is to freeze BN statistics (assuming that you have a pretrain model to begin with). Is that what you really do?</p></li>\n</ol>\n\n<p>There is a new paper on batch renormalisation. <a href=\"https://arxiv.org/abs/1702.03275\">https://arxiv.org/abs/1702.03275</a>.\nHas any try this and with this solve the issues above?</p>\n\n<p>I will post some code later, if the results are stable.</p>",
      "rawMarkdown": "Every kaggle competition solves a different problem and i learn a different thing. In this Carvana Image Masking Challenge, able to hange large input and output (e.g. prediction of mask at 1024x1024) is may be an advantage. \n\nThis thread is for using fp16 (16-bit float) and multi-gpu training and inference. I hope experienced kagglers who had worked with fp16 or multi-gpu can answer our question here.\n\nFor a start, I have probelm with BN layer.\n\n 1. will BN layer be affected. (e.g. the moving vars needs high precision) by fp16?\n \n 2. for multi gpu, how to compute BN statistics (moving mean and vars) over 4 gpu? The results doesn't seems to be stable? I read that if one solution is to freeze BN statistics (assuming that you have a pretrain model to begin with). Is that what you really do?\n \nThere is a new paper on batch renormalisation. https://arxiv.org/abs/1702.03275.\nHas any try this and with this solve the issues above?\n\nI will post some code later, if the results are stable.",
      "votes": null
    },
    {
      "id": "209404",
      "postDate": "08/02/2017 05:44:02",
      "content": "<p>You don't want to use fp16 on usual GPUs, it is much slower on them (basically on everything except Teslas). Also, it is usually fine for inference but while training we still need fp32 precision. </p>\n\n<p>But I like how you are open to experiments, keep going!</p>",
      "rawMarkdown": "You don't want to use fp16 on usual GPUs, it is much slower on them (basically on everything except Teslas). Also, it is usually fine for inference but while training we still need fp32 precision. \n\nBut I like how you are open to experiments, keep going!",
      "votes": null
    },
    {
      "id": "209436",
      "postDate": "08/02/2017 07:58:46",
      "content": "<p>I agree with Sergey, for training I think fp32 is needed. </p>\n\n<p>However on inference you could try using int8. NVIDIA has a tool called TensorRT that allows to optimize a network for using int8 with a small loss on precision. <br>\n<a href=\"http://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf\">http://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf</a></p>\n\n<p>However it's a quite complicated process on this days. I believe it will become much easier in the future.</p>",
      "rawMarkdown": "I agree with Sergey, for training I think fp32 is needed. \n\nHowever on inference you could try using int8. NVIDIA has a tool called TensorRT that allows to optimize a network for using int8 with a small loss on precision.  \nhttp://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf\n\nHowever it's a quite complicated process on this days. I believe it will become much easier in the future.",
      "votes": null
    },
    {
      "id": "211087",
      "postDate": "08/08/2017 01:39:51",
      "content": "<p>sorry for this noob question. But what do you all mean by inference in this context? I see the word mentioned a lot. Do you mean predicting the test set only or something else?</p>",
      "rawMarkdown": "sorry for this noob question. But what do you all mean by inference in this context? I see the word mentioned a lot. Do you mean predicting the test set only or something else?",
      "votes": null
    },
    {
      "id": "211097",
      "postDate": "08/08/2017 02:29:05",
      "content": "<p>Yeah, in this particular context, it means predicting the bitmap of the crop mask for data we don't have the ground truth to. This competition has the additional requirement of packaging predictions as run-length encodings in order to limit submission sizes</p>",
      "rawMarkdown": "Yeah, in this particular context, it means predicting the bitmap of the crop mask for data we don't have the ground truth to. This competition has the additional requirement of packaging predictions as run-length encodings in order to limit submission sizes",
      "votes": null
    },
    {
      "id": "211108",
      "postDate": "08/08/2017 03:24:22",
      "content": "<p>it means test mode. back propagation is not required and intermediate values are not cached.</p>",
      "rawMarkdown": "it means test mode. back propagation is not required and intermediate values are not cached.",
      "votes": null
    },
    {
      "id": "218177",
      "postDate": "09/02/2017 19:36:33",
      "content": "<p>BN can be trained in fp16 without overflow and loss of accuracy here. Example training loop is shown below. it reduces GPU memory by 40% for the feature maps so that the network can be wider (i.e. using more channels). The conversion between fp16 and fp32 however does take some time.</p>\n\n<pre><code>        images  = Variable(images).cuda()\n        labels  = Variable(labels).cuda()\n\n\n        net.half()  ## \n        images = images.half() ##\n\n        #forward\n        logits = net(images)\n        logits = logits.float() ##\n\n        loss = criterion(logits, labels) \n\n        loss.backward() \n        net.float() ##\n        optimizer.step()\n</code></pre>",
      "rawMarkdown": "BN can be trained in fp16 without overflow and loss of accuracy here. Example training loop is shown below. it reduces GPU memory by 40% for the feature maps so that the network can be wider (i.e. using more channels). The conversion between fp16 and fp32 however does take some time.\n\n\n            images  = Variable(images).cuda()\n            labels  = Variable(labels).cuda()\n            \n   \n            net.half()  ## \n            images = images.half() ##\n\n            #forward\n            logits = net(images)\n            logits = logits.float() ##\n \n            loss = criterion(logits, labels) \n\n            loss.backward() \n            net.float() ##\n            optimizer.step()",
      "votes": null
    },
    {
      "id": "218181",
      "postDate": "09/02/2017 19:44:23",
      "content": "<p>Isn't fp16 much slower than fp32 on Geforce GPUs?</p>",
      "rawMarkdown": "Isn't fp16 much slower than fp32 on Geforce GPUs?",
      "votes": null
    },
    {
      "id": "218236",
      "postDate": "09/03/2017 03:18:14",
      "content": "<p>I tried using np.float16 in planet very briefly since my gpu wasn't able to compete so much on large batch sizes and resolutions. I think I faced an error that was taking too much time or maybe it was saying fp16 wasn't supported. Looks like the HalfTensor is a relatively new addition:\n<a href=\"https://github.com/pytorch/pytorch/commit/67f94557ff26428ac911d8c08c7f9b619a41950e\">https://github.com/pytorch/pytorch/commit/67f94557ff26428ac911d8c08c7f9b619a41950e</a></p>\n\n<p>Don't remember what the exact problem was but I think this would definitely be worth a little more time, just depends on how much more time. Luckily, you can likely see if it's worth it or not after a handful of epochs. The 1080 seems to handle half floats well for me. I think nvidia increased the fp16 tflops for the new series as a move to be more supportive of building larger networks, but that's just a guess.</p>",
      "rawMarkdown": "I tried using np.float16 in planet very briefly since my gpu wasn't able to compete so much on large batch sizes and resolutions. I think I faced an error that was taking too much time or maybe it was saying fp16 wasn't supported. Looks like the HalfTensor is a relatively new addition:\nhttps://github.com/pytorch/pytorch/commit/67f94557ff26428ac911d8c08c7f9b619a41950e\n\nDon't remember what the exact problem was but I think this would definitely be worth a little more time, just depends on how much more time. Luckily, you can likely see if it's worth it or not after a handful of epochs. The 1080 seems to handle half floats well for me. I think nvidia increased the fp16 tflops for the new series as a move to be more supportive of building larger networks, but that's just a guess.",
      "votes": null
    },
    {
      "id": "218239",
      "postDate": "09/03/2017 03:41:29",
      "content": "<p>it is slower. but you can have a bigger model</p>",
      "rawMarkdown": "it is slower. but you can have a bigger model",
      "votes": null
    },
    {
      "id": "432818",
      "postDate": "12/04/2018 11:13:10",
      "content": "<p>Hi, I was experimenting with this for a while and made a POC, so I've yet to test on Volta or Tesla GPU, but on Pascal architecture I got about 15s less per epoch, that is almost 50mins less than the training time in FP32. Best per till now is that the storage is exactly halved.\nYou can checkout my code here: <a href=\"https://github.com/suvojit-0x55aa/mixed-precision-pytorch\">suvojit-0x55aa/mixed-precision-pytorch</a> \nAny suggestion and opinions are most welcome.</p>",
      "rawMarkdown": "Hi, I was experimenting with this for a while and made a POC, so I've yet to test on Volta or Tesla GPU, but on Pascal architecture I got about 15s less per epoch, that is almost 50mins less than the training time in FP32. Best per till now is that the storage is exactly halved.\nYou can checkout my code here: [suvojit-0x55aa/mixed-precision-pytorch](https://github.com/suvojit-0x55aa/mixed-precision-pytorch) \nAny suggestion and opinions are most welcome.",
      "votes": null
    },
    {
      "id": "755090",
      "postDate": "02/24/2020 12:20:05",
      "content": "<p>I'm only commenting because this is a top result from Google, even 3 years later. \nFP16 (half float) is <strong>considerably faster</strong> on any up-to-date GPU (Pascal and later) and you can\neasily see this for your self by training using <code>cuda().half()</code> vs <code>cuda().float()</code>.\nHowever, <code>half</code> often leads to numerical instability, resulting in <code>nan</code> or other issues.\nThere are workarounds, but it depends on the network parameters and the optimizer.</p>",
      "rawMarkdown": "I'm only commenting because this is a top result from Google, even 3 years later. \nFP16 (half float) is **considerably faster** on any up-to-date GPU (Pascal and later) and you can\neasily see this for your self by training using `cuda().half()` vs `cuda().float()`.\nHowever, `half` often leads to numerical instability, resulting in `nan` or other issues.\nThere are workarounds, but it depends on the network parameters and the optimizer.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 209404,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "08/02/2017 05:44:02",
      "content": "<p>You don't want to use fp16 on usual GPUs, it is much slower on them (basically on everything except Teslas). Also, it is usually fine for inference but while training we still need fp32 precision. </p>\n\n<p>But I like how you are open to experiments, keep going!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 209436,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "08/02/2017 07:58:46",
      "content": "<p>I agree with Sergey, for training I think fp32 is needed. </p>\n\n<p>However on inference you could try using int8. NVIDIA has a tool called TensorRT that allows to optimize a network for using int8 with a small loss on precision. <br>\n<a href=\"http://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf\">http://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf</a></p>\n\n<p>However it's a quite complicated process on this days. I believe it will become much easier in the future.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 211087,
      "author_name": "cjansen",
      "author_url": "",
      "post_date": "08/08/2017 01:39:51",
      "content": "<p>sorry for this noob question. But what do you all mean by inference in this context? I see the word mentioned a lot. Do you mean predicting the test set only or something else?</p>",
      "votes": null,
      "replies": [
        {
          "id": 211097,
          "author_name": "cpruce",
          "author_url": "",
          "post_date": "08/08/2017 02:29:05",
          "content": "<p>Yeah, in this particular context, it means predicting the bitmap of the crop mask for data we don't have the ground truth to. This competition has the additional requirement of packaging predictions as run-length encodings in order to limit submission sizes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211108,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/08/2017 03:24:22",
          "content": "<p>it means test mode. back propagation is not required and intermediate values are not cached.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 218177,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/02/2017 19:36:33",
      "content": "<p>BN can be trained in fp16 without overflow and loss of accuracy here. Example training loop is shown below. it reduces GPU memory by 40% for the feature maps so that the network can be wider (i.e. using more channels). The conversion between fp16 and fp32 however does take some time.</p>\n\n<pre><code>        images  = Variable(images).cuda()\n        labels  = Variable(labels).cuda()\n\n\n        net.half()  ## \n        images = images.half() ##\n\n        #forward\n        logits = net(images)\n        logits = logits.float() ##\n\n        loss = criterion(logits, labels) \n\n        loss.backward() \n        net.float() ##\n        optimizer.step()\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 218181,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "09/02/2017 19:44:23",
          "content": "<p>Isn't fp16 much slower than fp32 on Geforce GPUs?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218236,
          "author_name": "cpruce",
          "author_url": "",
          "post_date": "09/03/2017 03:18:14",
          "content": "<p>I tried using np.float16 in planet very briefly since my gpu wasn't able to compete so much on large batch sizes and resolutions. I think I faced an error that was taking too much time or maybe it was saying fp16 wasn't supported. Looks like the HalfTensor is a relatively new addition:\n<a href=\"https://github.com/pytorch/pytorch/commit/67f94557ff26428ac911d8c08c7f9b619a41950e\">https://github.com/pytorch/pytorch/commit/67f94557ff26428ac911d8c08c7f9b619a41950e</a></p>\n\n<p>Don't remember what the exact problem was but I think this would definitely be worth a little more time, just depends on how much more time. Luckily, you can likely see if it's worth it or not after a handful of epochs. The 1080 seems to handle half floats well for me. I think nvidia increased the fp16 tflops for the new series as a move to be more supportive of building larger networks, but that's just a guess.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218239,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/03/2017 03:41:29",
          "content": "<p>it is slower. but you can have a bigger model</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 432818,
      "author_name": "shinmigami",
      "author_url": "",
      "post_date": "12/04/2018 11:13:10",
      "content": "<p>Hi, I was experimenting with this for a while and made a POC, so I've yet to test on Volta or Tesla GPU, but on Pascal architecture I got about 15s less per epoch, that is almost 50mins less than the training time in FP32. Best per till now is that the storage is exactly halved.\nYou can checkout my code here: <a href=\"https://github.com/suvojit-0x55aa/mixed-precision-pytorch\">suvojit-0x55aa/mixed-precision-pytorch</a> \nAny suggestion and opinions are most welcome.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755090,
      "author_name": "alexgiokas",
      "author_url": "",
      "post_date": "02/24/2020 12:20:05",
      "content": "<p>I'm only commenting because this is a top result from Google, even 3 years later. \nFP16 (half float) is <strong>considerably faster</strong> on any up-to-date GPU (Pascal and later) and you can\neasily see this for your self by training using <code>cuda().half()</code> vs <code>cuda().float()</code>.\nHowever, <code>half</code> often leads to numerical instability, resulting in <code>nan</code> or other issues.\nThere are workarounds, but it depends on the network parameters and the optimizer.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "209357": "Every kaggle competition solves a different problem and i learn a different thing. In this Carvana Image Masking Challenge, able to hange large input and output (e.g. prediction of mask at 1024x1024) is may be an advantage. \n\nThis thread is for using fp16 (16-bit float) and multi-gpu training and inference. I hope experienced kagglers who had worked with fp16 or multi-gpu can answer our question here.\n\nFor a start, I have probelm with BN layer.\n\n 1. will BN layer be affected. (e.g. the moving vars needs high precision) by fp16?\n \n 2. for multi gpu, how to compute BN statistics (moving mean and vars) over 4 gpu? The results doesn't seems to be stable? I read that if one solution is to freeze BN statistics (assuming that you have a pretrain model to begin with). Is that what you really do?\n \nThere is a new paper on batch renormalisation. https://arxiv.org/abs/1702.03275.\nHas any try this and with this solve the issues above?\n\nI will post some code later, if the results are stable.",
    "209404": "You don't want to use fp16 on usual GPUs, it is much slower on them (basically on everything except Teslas). Also, it is usually fine for inference but while training we still need fp32 precision. \n\nBut I like how you are open to experiments, keep going!",
    "209436": "I agree with Sergey, for training I think fp32 is needed. \n\nHowever on inference you could try using int8. NVIDIA has a tool called TensorRT that allows to optimize a network for using int8 with a small loss on precision.  \nhttp://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf\n\nHowever it's a quite complicated process on this days. I believe it will become much easier in the future.",
    "211087": "sorry for this noob question. But what do you all mean by inference in this context? I see the word mentioned a lot. Do you mean predicting the test set only or something else?",
    "211097": "Yeah, in this particular context, it means predicting the bitmap of the crop mask for data we don't have the ground truth to. This competition has the additional requirement of packaging predictions as run-length encodings in order to limit submission sizes",
    "211108": "it means test mode. back propagation is not required and intermediate values are not cached.",
    "218177": "BN can be trained in fp16 without overflow and loss of accuracy here. Example training loop is shown below. it reduces GPU memory by 40% for the feature maps so that the network can be wider (i.e. using more channels). The conversion between fp16 and fp32 however does take some time.\n\n\n            images  = Variable(images).cuda()\n            labels  = Variable(labels).cuda()\n            \n   \n            net.half()  ## \n            images = images.half() ##\n\n            #forward\n            logits = net(images)\n            logits = logits.float() ##\n \n            loss = criterion(logits, labels) \n\n            loss.backward() \n            net.float() ##\n            optimizer.step()",
    "218181": "Isn't fp16 much slower than fp32 on Geforce GPUs?",
    "218236": "I tried using np.float16 in planet very briefly since my gpu wasn't able to compete so much on large batch sizes and resolutions. I think I faced an error that was taking too much time or maybe it was saying fp16 wasn't supported. Looks like the HalfTensor is a relatively new addition:\nhttps://github.com/pytorch/pytorch/commit/67f94557ff26428ac911d8c08c7f9b619a41950e\n\nDon't remember what the exact problem was but I think this would definitely be worth a little more time, just depends on how much more time. Luckily, you can likely see if it's worth it or not after a handful of epochs. The 1080 seems to handle half floats well for me. I think nvidia increased the fp16 tflops for the new series as a move to be more supportive of building larger networks, but that's just a guess.",
    "218239": "it is slower. but you can have a bigger model",
    "432818": "Hi, I was experimenting with this for a while and made a POC, so I've yet to test on Volta or Tesla GPU, but on Pascal architecture I got about 15s less per epoch, that is almost 50mins less than the training time in FP32. Best per till now is that the storage is exactly halved.\nYou can checkout my code here: [suvojit-0x55aa/mixed-precision-pytorch](https://github.com/suvojit-0x55aa/mixed-precision-pytorch) \nAny suggestion and opinions are most welcome.",
    "755090": "I'm only commenting because this is a top result from Google, even 3 years later. \nFP16 (half float) is **considerably faster** on any up-to-date GPU (Pascal and later) and you can\neasily see this for your self by training using `cuda().half()` vs `cuda().float()`.\nHowever, `half` often leads to numerical instability, resulting in `nan` or other issues.\nThere are workarounds, but it depends on the network parameters and the optimizer."
  },
  "source": "meta"
}