{
  "id": 238546,
  "title": "FP16, transformer training, NaN",
  "url": "/competitions/bms-molecular-translation/discussion/238546",
  "author_name": "",
  "post_date": "2021-05-12T14:14:37.660518Z",
  "votes": 11,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I frequently encounter NAN in loss when training transformer in fp16.<br>\n(the problem disappears in fp32).</p>\n<p>Does anyone have a suggestion to avoid this problem?<br>\ne.g. </p>\n<ul>\n<li>adding clipping function to feed-forward layer?</li>\n<li>or should i modify the softmax weight computation (minus off the max value)</li>\n</ul>\n<p>It is difficult for me to catch the overflow and right now i  do not know where the overflow has occurred. Does anyone have a good suggestion?</p>\n<p>but i thought each layer is already normalized by layerNorm?</p>",
  "messages": [
    {
      "id": "1304266",
      "postDate": "05/12/2021 14:14:37",
      "content": "<p>I frequently encounter NAN in loss when training transformer in fp16.<br>\n(the problem disappears in fp32).</p>\n<p>Does anyone have a suggestion to avoid this problem?<br>\ne.g. </p>\n<ul>\n<li>adding clipping function to feed-forward layer?</li>\n<li>or should i modify the softmax weight computation (minus off the max value)</li>\n</ul>\n<p>It is difficult for me to catch the overflow and right now i  do not know where the overflow has occurred. Does anyone have a good suggestion?</p>\n<p>but i thought each layer is already normalized by layerNorm?</p>",
      "rawMarkdown": "I frequently encounter NAN in loss when training transformer in fp16.\n(the problem disappears in fp32).\n\nDoes anyone have a suggestion to avoid this problem?\ne.g. \n- adding clipping function to feed-forward layer?\n- or should i modify the softmax weight computation (minus off the max value)\n\nIt is difficult for me to catch the overflow and right now i  do not know where the overflow has occurred. Does anyone have a good suggestion?\n\nbut i thought each layer is already normalized by layerNorm?",
      "votes": null
    },
    {
      "id": "1304289",
      "postDate": "05/12/2021 14:35:27",
      "content": "<p>Do you mean transformer decoder or the transformer image encoder you are (apparently) using?<br>\nI have found that it helps to add LayerNorm before input into transformer decoder. Although it worsens performance.</p>\n<p>Have you tried clipping the gradient norm at optimizer level? That also helps to avoid NaNs (without using mentioned LayerNorm) in general.<br>\nOnce gain it worsens my performance instead of improving it, but I have the suspicion that the cause lies elsewhere.</p>\n<p>In my case there is also no problem with fp32 training</p>\n<p>You can determine cause of NaNs by checking intermediate tensors for NaNs. For PyTorch e.g. <code>not tensor.isnan().all()</code></p>",
      "rawMarkdown": "Do you mean transformer decoder or the transformer image encoder you are (apparently) using?\nI have found that it helps to add LayerNorm before input into transformer decoder. Although it worsens performance.\n\nHave you tried clipping the gradient norm at optimizer level? That also helps to avoid NaNs (without using mentioned LayerNorm) in general.\nOnce gain it worsens my performance instead of improving it, but I have the suspicion that the cause lies elsewhere.\n\nIn my case there is also no problem with fp32 training\n\nYou can determine cause of NaNs by checking intermediate tensors for NaNs. For PyTorch e.g. ` not tensor.isnan().all()`",
      "votes": null
    },
    {
      "id": "1304293",
      "postDate": "05/12/2021 14:40:51",
      "content": "<p>It depends on the encoder, some visual transformer will not show this behaviour :) The decoder will not have this problem. You can test that: freeze encoder and train with the half preicision.</p>",
      "rawMarkdown": "It depends on the encoder, some visual transformer will not show this behaviour :) The decoder will not have this problem. You can test that: freeze encoder and train with the half preicision.",
      "votes": null
    },
    {
      "id": "1304302",
      "postDate": "05/12/2021 14:43:36",
      "content": "<p>yes=)<br>\nfor some transformers decoders I had to clip to <code>0.5</code> for others to <code>1.0</code> plus very low <code>lr</code> and <code>batch_size</code> not lower than <code>32</code>. </p>",
      "rawMarkdown": "yes=)\nfor some transformers decoders I had to clip to `0.5` for others to `1.0` plus very low `lr` and `batch_size` not lower than `32`.",
      "votes": null
    },
    {
      "id": "1304330",
      "postDate": "05/12/2021 14:59:31",
      "content": "<p>I fight that, too.<br>\nWhen it happens, I try to just start from the last model that had no nan losses yet. Sometimes it runs without any issues after that, but there are times when it runs for only a few epochs and I start getting nans again, and in this case if I just try to continue as the last time, it will fail again very soon.<br>\nIn these cases, increasing the batch size can help, even with gradient accumulation. Somewhat equivalent to increasing the batch size is lowering the learning rate.<br>\nIf these doesn't help, I load back older model, like 5-10 epochs before the first nan was encountered and continue from there, usually with increased batch size.<br>\nNote that I don't save and load the optimizer state and don't use warmup after restarting - I act like nothing has happened.</p>\n<p>I'm not sure of it, as I didn't go after what causes these nans, but I think that it happens when the model encounters very rare cases of inchis on which the model performs poorly. I only tried training on 40k training images in which I put some images with inchis that were unique in some ways compared to all the other (40k-1) images. The model was training well, there was no sign of sickness in the training loss, but the validation was a rollercoaster. If I removed those unique images, the validation was normal.<br>\nIf these unique inchis are really the causes, then eliminating those can help.<br>\nOr just monitoring if there are too high losses in the current batch, and if so, you can skip the update. (I didn't try this, as I had success continuing training so far with what I wrote before.)</p>",
      "rawMarkdown": "I fight that, too.\nWhen it happens, I try to just start from the last model that had no nan losses yet. Sometimes it runs without any issues after that, but there are times when it runs for only a few epochs and I start getting nans again, and in this case if I just try to continue as the last time, it will fail again very soon.\nIn these cases, increasing the batch size can help, even with gradient accumulation. Somewhat equivalent to increasing the batch size is lowering the learning rate.\nIf these doesn't help, I load back older model, like 5-10 epochs before the first nan was encountered and continue from there, usually with increased batch size.\nNote that I don't save and load the optimizer state and don't use warmup after restarting - I act like nothing has happened.\n\nI'm not sure of it, as I didn't go after what causes these nans, but I think that it happens when the model encounters very rare cases of inchis on which the model performs poorly. I only tried training on 40k training images in which I put some images with inchis that were unique in some ways compared to all the other (40k-1) images. The model was training well, there was no sign of sickness in the training loss, but the validation was a rollercoaster. If I removed those unique images, the validation was normal.\nIf these unique inchis are really the causes, then eliminating those can help.\nOr just monitoring if there are too high losses in the current batch, and if so, you can skip the update. (I didn't try this, as I had success continuing training so far with what I wrote before.)",
      "votes": null
    },
    {
      "id": "1304345",
      "postDate": "05/12/2021 15:06:31",
      "content": "<p>I found out recently that there is more elegant way of detecting what causing <code>underflow</code> or <code>overflow</code> … here is a whole page from huggingface  its pretty easy to implement <a href=\"https://huggingface.co/transformers/master/debugging.html#underflow-and-overflow-detection\" target=\"_blank\">https://huggingface.co/transformers/master/debugging.html#underflow-and-overflow-detection</a></p>",
      "rawMarkdown": "I found out recently that there is more elegant way of detecting what causing `underflow` or `overflow` ... here is a whole page from huggingface  its pretty easy to implement https://huggingface.co/transformers/master/debugging.html#underflow-and-overflow-detection",
      "votes": null
    },
    {
      "id": "1304408",
      "postDate": "05/12/2021 15:51:29",
      "content": "<p>huggingface solution is good.</p>\n<p>i need to study it.<br>\nThanks a lot.</p>\n<p>they are using pytorch hook function for callback</p>",
      "rawMarkdown": "huggingface solution is good.\n\ni need to study it.\nThanks a lot.\n\nthey are using pytorch hook function for callback",
      "votes": null
    },
    {
      "id": "1305698",
      "postDate": "05/13/2021 12:35:04",
      "content": "<p>if you are using fp16, you can finetune in fp32. my experiments show that fp32 produces better results</p>",
      "rawMarkdown": "if you are using fp16, you can finetune in fp32. my experiments show that fp32 produces better results",
      "votes": null
    },
    {
      "id": "1306154",
      "postDate": "05/13/2021 16:28:59",
      "content": "<p>I'm running into this issue at the moment. I tried using fp32 for the optimizer only to rule out that being the problem - it still ended up with nans. Now trying with a combination of lower lr and lower value for <code>clip_grad_norm</code>. If that fails I'll switch to fp32 for the whole thing. But then it ends up being S___L___O___W…</p>",
      "rawMarkdown": "I'm running into this issue at the moment. I tried using fp32 for the optimizer only to rule out that being the problem - it still ended up with nans. Now trying with a combination of lower lr and lower value for `clip_grad_norm`. If that fails I'll switch to fp32 for the whole thing. But then it ends up being S___L___O___W...",
      "votes": null
    },
    {
      "id": "1306479",
      "postDate": "05/13/2021 19:34:14",
      "content": "<p>I met the same problem.  my forked TPU training notebook work fine on KAGGLE.  I manage to run it on Colab and continue training. Due to OOM problem,  I have to reduce batchsize, however, then after several epochs, NAN pops up and  training break.</p>",
      "rawMarkdown": "I met the same problem.  my forked TPU training notebook work fine on KAGGLE.  I manage to run it on Colab and continue training. Due to OOM problem,  I have to reduce batchsize, however, then after several epochs, NAN pops up and  training break.",
      "votes": null
    },
    {
      "id": "1306509",
      "postDate": "05/13/2021 20:33:01",
      "content": "<p>Beware if you use fp32 on GPU on Pytorch with models like EffNets </p>\n<p>The implementation of depthwise convolution is not optimal compared to tensorflow <a href=\"https://github.com/pytorch/pytorch/issues/18631\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/18631</a></p>",
      "rawMarkdown": "Beware if you use fp32 on GPU on Pytorch with models like EffNets \n\nThe implementation of depthwise convolution is not optimal compared to tensorflow [https://github.com/pytorch/pytorch/issues/18631](https://github.com/pytorch/pytorch/issues/18631)",
      "votes": null
    },
    {
      "id": "1306517",
      "postDate": "05/13/2021 20:42:42",
      "content": "<p>I just wanted to point out an overview of my observations (some again) at one place. It might help others to determine their cause. Others could post alike below?<br>\nFramework: Pytorch, AMP Mixed precision training</p>\n<ul>\n<li>NaNs occur at CNN-image-model -&gt; transformer transition</li>\n<li>Actually dropout itself determines if NaNs occur at all. Depends on concrete architecture when exactly of course. Higher dropout -&gt; NaNs</li>\n<li>If NaNs at all then during 20%-38% of first epoch with 80% training data. Never in later epochs. At least until epoch 5. (Almost) Never trained beyond that yet.</li>\n<li>Batch size influence -&gt; never looked into it.</li>\n<li>Help I: Gradient norm clipping, but decreases performance (Shouldn't as far as I am aware, Am I wrong?)</li>\n<li>Help II: Layer-norm after CNN-image-model feature reshaping. Decreases performance also, but nullified by dropout increase. Other norms work worse</li>\n</ul>\n<p>=&gt; Best solution so far: Keep dropout lower, no gradient norm, no NaNs, not best theoretical performance by a good margin :/<br>\n=&gt; Second best: LayerNorm after CNN. </p>",
      "rawMarkdown": "I just wanted to point out an overview of my observations (some again) at one place. It might help others to determine their cause. Others could post alike below?\nFramework: Pytorch, AMP Mixed precision training\n- NaNs occur at CNN-image-model -> transformer transition\n- Actually dropout itself determines if NaNs occur at all. Depends on concrete architecture when exactly of course. Higher dropout -> NaNs\n- If NaNs at all then during 20%-38% of first epoch with 80% training data. Never in later epochs. At least until epoch 5. (Almost) Never trained beyond that yet.\n- Batch size influence -> never looked into it.\n- Help I: Gradient norm clipping, but decreases performance (Shouldn't as far as I am aware, Am I wrong?)\n- Help II: Layer-norm after CNN-image-model feature reshaping. Decreases performance also, but nullified by dropout increase. Other norms work worse\n\n=> Best solution so far: Keep dropout lower, no gradient norm, no NaNs, not best theoretical performance by a good margin :/\n=> Second best: LayerNorm after CNN.",
      "votes": null
    },
    {
      "id": "1307403",
      "postDate": "05/14/2021 12:28:23",
      "content": "<p>I haven't played too much with visual transformers (as <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> suggests, I agree the encoder is the likely place where you are seeing the issues), but it seems they are generally trained with lower learning rates than we are used to with CNNs.</p>\n<p>Also, seems people devote quite a few epochs to warm up.</p>\n<p>I have been consistently getting <code>NaNs</code>  when training with a flat <code>LR</code> without warmup, but then I switched to this schedule:</p>\n<pre><code>from fastai.callback.schedule import *\n\np = torch.linspace(0.,1,100)\nf = combine_scheds([0.05, 0.95], [SchedLin(LR/1000, LR), SchedCos(LR, LR/100)])\nplt.plot(p, [f(o) for o in p]);\n</code></pre>\n<p><img src=\"https://i.imgur.com/jXvP3fW.png\" alt=\"\"></p>\n<p>I train with roughly 1/3 max <code>LR</code> how the arch was trained for the paper, and scale the <code>LR</code> linearly with batch size. Seems transformers really like low <code>LRs</code> so this might still be too high.</p>\n<p><code>fastai</code> is very modular, so I am grabbing the scheduler and have been using it with pytorch and ignite as well.</p>\n<p>Doing the above, I still see a weird jump in train loss some time into the training and subsequently the model trains very poorly. If I add gradient clipping, the problem goes away.</p>",
      "rawMarkdown": "I haven't played too much with visual transformers (as @tugstugi suggests, I agree the encoder is the likely place where you are seeing the issues), but it seems they are generally trained with lower learning rates than we are used to with CNNs.\n\nAlso, seems people devote quite a few epochs to warm up.\n\nI have been consistently getting `NaNs`  when training with a flat `LR` without warmup, but then I switched to this schedule:\n\n```\nfrom fastai.callback.schedule import *\n\np = torch.linspace(0.,1,100)\nf = combine_scheds([0.05, 0.95], [SchedLin(LR/1000, LR), SchedCos(LR, LR/100)])\nplt.plot(p, [f(o) for o in p]);\n```\n\n![](https://i.imgur.com/jXvP3fW.png)\n\nI train with roughly 1/3 max `LR` how the arch was trained for the paper, and scale the `LR` linearly with batch size. Seems transformers really like low `LRs` so this might still be too high.\n\n`fastai` is very modular, so I am grabbing the scheduler and have been using it with pytorch and ignite as well.\n\nDoing the above, I still see a weird jump in train loss some time into the training and subsequently the model trains very poorly. If I add gradient clipping, the problem goes away.",
      "votes": null
    },
    {
      "id": "1308757",
      "postDate": "05/15/2021 12:44:14",
      "content": "<p>I had this issue with the LSTM approach and I found my NaN's were coming from overflow in the attention calculation for one random sample in one random batch. The sample didn't seem particularly special in any way. Therefore grad clipping seemed to have nothing to do with it.</p>",
      "rawMarkdown": "I had this issue with the LSTM approach and I found my NaN's were coming from overflow in the attention calculation for one random sample in one random batch. The sample didn't seem particularly special in any way. Therefore grad clipping seemed to have nothing to do with it.",
      "votes": null
    },
    {
      "id": "1309498",
      "postDate": "05/16/2021 03:31:00",
      "content": "<p>Certain Transformer variants play poorly with FP16 - I've encountered this with the T5 architecture for example in an NLP task at the GeGLU layer.  </p>\n<p>Two common solutions :- </p>\n<ol>\n<li>Don't use FP16 and use gradient checkpointing instead if you originally were using it due to VRAM limits</li>\n<li>Pinpoint the layers that are blowing up (you can just monitor the gradients and see where the NaNs start) and white-list those specific layers to be run in FP32 with the rest kept in FP16. The last time I checked, this was easier with PyTorch's Autocast code rather than Nvidia AMP but that might have changed. </li>\n</ol>",
      "rawMarkdown": "Certain Transformer variants play poorly with FP16 - I've encountered this with the T5 architecture for example in an NLP task at the GeGLU layer.  \n\nTwo common solutions :- \n1. Don't use FP16 and use gradient checkpointing instead if you originally were using it due to VRAM limits\n2. Pinpoint the layers that are blowing up (you can just monitor the gradients and see where the NaNs start) and white-list those specific layers to be run in FP32 with the rest kept in FP16. The last time I checked, this was easier with PyTorch's Autocast code rather than Nvidia AMP but that might have changed.",
      "votes": null
    },
    {
      "id": "1328181",
      "postDate": "05/30/2021 02:38:59",
      "content": "<p>it is important to cross check your results.<br>\nNAN can happens in both training and inference</p>",
      "rawMarkdown": "it is important to cross check your results.\nNAN can happens in both training and inference",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1304289,
      "author_name": "cepheidq",
      "author_url": "",
      "post_date": "05/12/2021 14:35:27",
      "content": "<p>Do you mean transformer decoder or the transformer image encoder you are (apparently) using?<br>\nI have found that it helps to add LayerNorm before input into transformer decoder. Although it worsens performance.</p>\n<p>Have you tried clipping the gradient norm at optimizer level? That also helps to avoid NaNs (without using mentioned LayerNorm) in general.<br>\nOnce gain it worsens my performance instead of improving it, but I have the suspicion that the cause lies elsewhere.</p>\n<p>In my case there is also no problem with fp32 training</p>\n<p>You can determine cause of NaNs by checking intermediate tensors for NaNs. For PyTorch e.g. <code>not tensor.isnan().all()</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1304293,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "05/12/2021 14:40:51",
      "content": "<p>It depends on the encoder, some visual transformer will not show this behaviour :) The decoder will not have this problem. You can test that: freeze encoder and train with the half preicision.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1304302,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "05/12/2021 14:43:36",
          "content": "<p>yes=)<br>\nfor some transformers decoders I had to clip to <code>0.5</code> for others to <code>1.0</code> plus very low <code>lr</code> and <code>batch_size</code> not lower than <code>32</code>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1304330,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "05/12/2021 14:59:31",
      "content": "<p>I fight that, too.<br>\nWhen it happens, I try to just start from the last model that had no nan losses yet. Sometimes it runs without any issues after that, but there are times when it runs for only a few epochs and I start getting nans again, and in this case if I just try to continue as the last time, it will fail again very soon.<br>\nIn these cases, increasing the batch size can help, even with gradient accumulation. Somewhat equivalent to increasing the batch size is lowering the learning rate.<br>\nIf these doesn't help, I load back older model, like 5-10 epochs before the first nan was encountered and continue from there, usually with increased batch size.<br>\nNote that I don't save and load the optimizer state and don't use warmup after restarting - I act like nothing has happened.</p>\n<p>I'm not sure of it, as I didn't go after what causes these nans, but I think that it happens when the model encounters very rare cases of inchis on which the model performs poorly. I only tried training on 40k training images in which I put some images with inchis that were unique in some ways compared to all the other (40k-1) images. The model was training well, there was no sign of sickness in the training loss, but the validation was a rollercoaster. If I removed those unique images, the validation was normal.<br>\nIf these unique inchis are really the causes, then eliminating those can help.<br>\nOr just monitoring if there are too high losses in the current batch, and if so, you can skip the update. (I didn't try this, as I had success continuing training so far with what I wrote before.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1304345,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "05/12/2021 15:06:31",
          "content": "<p>I found out recently that there is more elegant way of detecting what causing <code>underflow</code> or <code>overflow</code> … here is a whole page from huggingface  its pretty easy to implement <a href=\"https://huggingface.co/transformers/master/debugging.html#underflow-and-overflow-detection\" target=\"_blank\">https://huggingface.co/transformers/master/debugging.html#underflow-and-overflow-detection</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1304408,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/12/2021 15:51:29",
          "content": "<p>huggingface solution is good.</p>\n<p>i need to study it.<br>\nThanks a lot.</p>\n<p>they are using pytorch hook function for callback</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1305698,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/13/2021 12:35:04",
      "content": "<p>if you are using fp16, you can finetune in fp32. my experiments show that fp32 produces better results</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1306154,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "05/13/2021 16:28:59",
      "content": "<p>I'm running into this issue at the moment. I tried using fp32 for the optimizer only to rule out that being the problem - it still ended up with nans. Now trying with a combination of lower lr and lower value for <code>clip_grad_norm</code>. If that fails I'll switch to fp32 for the whole thing. But then it ends up being S___L___O___W…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1306479,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "05/13/2021 19:34:14",
      "content": "<p>I met the same problem.  my forked TPU training notebook work fine on KAGGLE.  I manage to run it on Colab and continue training. Due to OOM problem,  I have to reduce batchsize, however, then after several epochs, NAN pops up and  training break.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1306509,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/13/2021 20:33:01",
      "content": "<p>Beware if you use fp32 on GPU on Pytorch with models like EffNets </p>\n<p>The implementation of depthwise convolution is not optimal compared to tensorflow <a href=\"https://github.com/pytorch/pytorch/issues/18631\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/18631</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1306517,
      "author_name": "cepheidq",
      "author_url": "",
      "post_date": "05/13/2021 20:42:42",
      "content": "<p>I just wanted to point out an overview of my observations (some again) at one place. It might help others to determine their cause. Others could post alike below?<br>\nFramework: Pytorch, AMP Mixed precision training</p>\n<ul>\n<li>NaNs occur at CNN-image-model -&gt; transformer transition</li>\n<li>Actually dropout itself determines if NaNs occur at all. Depends on concrete architecture when exactly of course. Higher dropout -&gt; NaNs</li>\n<li>If NaNs at all then during 20%-38% of first epoch with 80% training data. Never in later epochs. At least until epoch 5. (Almost) Never trained beyond that yet.</li>\n<li>Batch size influence -&gt; never looked into it.</li>\n<li>Help I: Gradient norm clipping, but decreases performance (Shouldn't as far as I am aware, Am I wrong?)</li>\n<li>Help II: Layer-norm after CNN-image-model feature reshaping. Decreases performance also, but nullified by dropout increase. Other norms work worse</li>\n</ul>\n<p>=&gt; Best solution so far: Keep dropout lower, no gradient norm, no NaNs, not best theoretical performance by a good margin :/<br>\n=&gt; Second best: LayerNorm after CNN. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1307403,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "05/14/2021 12:28:23",
      "content": "<p>I haven't played too much with visual transformers (as <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> suggests, I agree the encoder is the likely place where you are seeing the issues), but it seems they are generally trained with lower learning rates than we are used to with CNNs.</p>\n<p>Also, seems people devote quite a few epochs to warm up.</p>\n<p>I have been consistently getting <code>NaNs</code>  when training with a flat <code>LR</code> without warmup, but then I switched to this schedule:</p>\n<pre><code>from fastai.callback.schedule import *\n\np = torch.linspace(0.,1,100)\nf = combine_scheds([0.05, 0.95], [SchedLin(LR/1000, LR), SchedCos(LR, LR/100)])\nplt.plot(p, [f(o) for o in p]);\n</code></pre>\n<p><img src=\"https://i.imgur.com/jXvP3fW.png\" alt=\"\"></p>\n<p>I train with roughly 1/3 max <code>LR</code> how the arch was trained for the paper, and scale the <code>LR</code> linearly with batch size. Seems transformers really like low <code>LRs</code> so this might still be too high.</p>\n<p><code>fastai</code> is very modular, so I am grabbing the scheduler and have been using it with pytorch and ignite as well.</p>\n<p>Doing the above, I still see a weird jump in train loss some time into the training and subsequently the model trains very poorly. If I add gradient clipping, the problem goes away.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1308757,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "05/15/2021 12:44:14",
      "content": "<p>I had this issue with the LSTM approach and I found my NaN's were coming from overflow in the attention calculation for one random sample in one random batch. The sample didn't seem particularly special in any way. Therefore grad clipping seemed to have nothing to do with it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1309498,
      "author_name": "leecming",
      "author_url": "",
      "post_date": "05/16/2021 03:31:00",
      "content": "<p>Certain Transformer variants play poorly with FP16 - I've encountered this with the T5 architecture for example in an NLP task at the GeGLU layer.  </p>\n<p>Two common solutions :- </p>\n<ol>\n<li>Don't use FP16 and use gradient checkpointing instead if you originally were using it due to VRAM limits</li>\n<li>Pinpoint the layers that are blowing up (you can just monitor the gradients and see where the NaNs start) and white-list those specific layers to be run in FP32 with the rest kept in FP16. The last time I checked, this was easier with PyTorch's Autocast code rather than Nvidia AMP but that might have changed. </li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1328181,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/30/2021 02:38:59",
      "content": "<p>it is important to cross check your results.<br>\nNAN can happens in both training and inference</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1304266": "I frequently encounter NAN in loss when training transformer in fp16.\n(the problem disappears in fp32).\n\nDoes anyone have a suggestion to avoid this problem?\ne.g. \n- adding clipping function to feed-forward layer?\n- or should i modify the softmax weight computation (minus off the max value)\n\nIt is difficult for me to catch the overflow and right now i  do not know where the overflow has occurred. Does anyone have a good suggestion?\n\nbut i thought each layer is already normalized by layerNorm?",
    "1304289": "Do you mean transformer decoder or the transformer image encoder you are (apparently) using?\nI have found that it helps to add LayerNorm before input into transformer decoder. Although it worsens performance.\n\nHave you tried clipping the gradient norm at optimizer level? That also helps to avoid NaNs (without using mentioned LayerNorm) in general.\nOnce gain it worsens my performance instead of improving it, but I have the suspicion that the cause lies elsewhere.\n\nIn my case there is also no problem with fp32 training\n\nYou can determine cause of NaNs by checking intermediate tensors for NaNs. For PyTorch e.g. ` not tensor.isnan().all()`",
    "1304293": "It depends on the encoder, some visual transformer will not show this behaviour :) The decoder will not have this problem. You can test that: freeze encoder and train with the half preicision.",
    "1304302": "yes=)\nfor some transformers decoders I had to clip to `0.5` for others to `1.0` plus very low `lr` and `batch_size` not lower than `32`.",
    "1304330": "I fight that, too.\nWhen it happens, I try to just start from the last model that had no nan losses yet. Sometimes it runs without any issues after that, but there are times when it runs for only a few epochs and I start getting nans again, and in this case if I just try to continue as the last time, it will fail again very soon.\nIn these cases, increasing the batch size can help, even with gradient accumulation. Somewhat equivalent to increasing the batch size is lowering the learning rate.\nIf these doesn't help, I load back older model, like 5-10 epochs before the first nan was encountered and continue from there, usually with increased batch size.\nNote that I don't save and load the optimizer state and don't use warmup after restarting - I act like nothing has happened.\n\nI'm not sure of it, as I didn't go after what causes these nans, but I think that it happens when the model encounters very rare cases of inchis on which the model performs poorly. I only tried training on 40k training images in which I put some images with inchis that were unique in some ways compared to all the other (40k-1) images. The model was training well, there was no sign of sickness in the training loss, but the validation was a rollercoaster. If I removed those unique images, the validation was normal.\nIf these unique inchis are really the causes, then eliminating those can help.\nOr just monitoring if there are too high losses in the current batch, and if so, you can skip the update. (I didn't try this, as I had success continuing training so far with what I wrote before.)",
    "1304345": "I found out recently that there is more elegant way of detecting what causing `underflow` or `overflow` ... here is a whole page from huggingface  its pretty easy to implement https://huggingface.co/transformers/master/debugging.html#underflow-and-overflow-detection",
    "1304408": "huggingface solution is good.\n\ni need to study it.\nThanks a lot.\n\nthey are using pytorch hook function for callback",
    "1305698": "if you are using fp16, you can finetune in fp32. my experiments show that fp32 produces better results",
    "1306154": "I'm running into this issue at the moment. I tried using fp32 for the optimizer only to rule out that being the problem - it still ended up with nans. Now trying with a combination of lower lr and lower value for `clip_grad_norm`. If that fails I'll switch to fp32 for the whole thing. But then it ends up being S___L___O___W...",
    "1306479": "I met the same problem.  my forked TPU training notebook work fine on KAGGLE.  I manage to run it on Colab and continue training. Due to OOM problem,  I have to reduce batchsize, however, then after several epochs, NAN pops up and  training break.",
    "1306509": "Beware if you use fp32 on GPU on Pytorch with models like EffNets \n\nThe implementation of depthwise convolution is not optimal compared to tensorflow [https://github.com/pytorch/pytorch/issues/18631](https://github.com/pytorch/pytorch/issues/18631)",
    "1306517": "I just wanted to point out an overview of my observations (some again) at one place. It might help others to determine their cause. Others could post alike below?\nFramework: Pytorch, AMP Mixed precision training\n- NaNs occur at CNN-image-model -> transformer transition\n- Actually dropout itself determines if NaNs occur at all. Depends on concrete architecture when exactly of course. Higher dropout -> NaNs\n- If NaNs at all then during 20%-38% of first epoch with 80% training data. Never in later epochs. At least until epoch 5. (Almost) Never trained beyond that yet.\n- Batch size influence -> never looked into it.\n- Help I: Gradient norm clipping, but decreases performance (Shouldn't as far as I am aware, Am I wrong?)\n- Help II: Layer-norm after CNN-image-model feature reshaping. Decreases performance also, but nullified by dropout increase. Other norms work worse\n\n=> Best solution so far: Keep dropout lower, no gradient norm, no NaNs, not best theoretical performance by a good margin :/\n=> Second best: LayerNorm after CNN.",
    "1307403": "I haven't played too much with visual transformers (as @tugstugi suggests, I agree the encoder is the likely place where you are seeing the issues), but it seems they are generally trained with lower learning rates than we are used to with CNNs.\n\nAlso, seems people devote quite a few epochs to warm up.\n\nI have been consistently getting `NaNs`  when training with a flat `LR` without warmup, but then I switched to this schedule:\n\n```\nfrom fastai.callback.schedule import *\n\np = torch.linspace(0.,1,100)\nf = combine_scheds([0.05, 0.95], [SchedLin(LR/1000, LR), SchedCos(LR, LR/100)])\nplt.plot(p, [f(o) for o in p]);\n```\n\n![](https://i.imgur.com/jXvP3fW.png)\n\nI train with roughly 1/3 max `LR` how the arch was trained for the paper, and scale the `LR` linearly with batch size. Seems transformers really like low `LRs` so this might still be too high.\n\n`fastai` is very modular, so I am grabbing the scheduler and have been using it with pytorch and ignite as well.\n\nDoing the above, I still see a weird jump in train loss some time into the training and subsequently the model trains very poorly. If I add gradient clipping, the problem goes away.",
    "1308757": "I had this issue with the LSTM approach and I found my NaN's were coming from overflow in the attention calculation for one random sample in one random batch. The sample didn't seem particularly special in any way. Therefore grad clipping seemed to have nothing to do with it.",
    "1309498": "Certain Transformer variants play poorly with FP16 - I've encountered this with the T5 architecture for example in an NLP task at the GeGLU layer.  \n\nTwo common solutions :- \n1. Don't use FP16 and use gradient checkpointing instead if you originally were using it due to VRAM limits\n2. Pinpoint the layers that are blowing up (you can just monitor the gradients and see where the NaNs start) and white-list those specific layers to be run in FP32 with the rest kept in FP16. The last time I checked, this was easier with PyTorch's Autocast code rather than Nvidia AMP but that might have changed.",
    "1328181": "it is important to cross check your results.\nNAN can happens in both training and inference"
  },
  "source": "meta"
}