{
  "id": 158006,
  "title": "Loss Decreases to `nan` after ~14 epochs in fp16",
  "url": "/competitions/alaska2-image-steganalysis/discussion/158006",
  "author_name": "",
  "post_date": "2020-06-13T00:53:58.560657100Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am using half-precision training with RAdam optimizer and keep getting a <code>nan</code> loss especially after 14 epochs. It tried opt 01 and opt 02 levels but it did not make a difference. I tested the input data and no <code>nan</code>'s occur there, so it must be something in the optimizer/precision. I tried increasing the <code>eps</code> in RAdam, too, but to no avail. Has anybody dealt with such issue?</p>",
  "messages": [
    {
      "id": "883814",
      "postDate": "06/13/2020 00:53:58",
      "content": "<p>I am using half-precision training with RAdam optimizer and keep getting a <code>nan</code> loss especially after 14 epochs. It tried opt 01 and opt 02 levels but it did not make a difference. I tested the input data and no <code>nan</code>'s occur there, so it must be something in the optimizer/precision. I tried increasing the <code>eps</code> in RAdam, too, but to no avail. Has anybody dealt with such issue?</p>",
      "rawMarkdown": "I am using half-precision training with RAdam optimizer and keep getting a `nan` loss especially after 14 epochs. It tried opt 01 and opt 02 levels but it did not make a difference. I tested the input data and no `nan`'s occur there, so it must be something in the optimizer/precision. I tried increasing the `eps` in RAdam, too, but to no avail. Has anybody dealt with such issue?",
      "votes": null
    },
    {
      "id": "883927",
      "postDate": "06/13/2020 05:00:29",
      "content": "<p>Training loss or validation loss? Is the loss generally decreasing up until that point? What learning rate are you using? Have you tried lowering it?</p>",
      "rawMarkdown": "Training loss or validation loss? Is the loss generally decreasing up until that point? What learning rate are you using? Have you tried lowering it?",
      "votes": null
    },
    {
      "id": "884739",
      "postDate": "06/13/2020 15:32:33",
      "content": "<p>what is your current eps value? I faced the same problem. Solved it by limiting minimum lr with 5e-5.</p>",
      "rawMarkdown": "what is your current eps value? I faced the same problem. Solved it by limiting minimum lr with 5e-5.",
      "votes": null
    },
    {
      "id": "885182",
      "postDate": "06/14/2020 01:56:36",
      "content": "<p>I lowered from 1e-8 (default) to 1e-6 and then to 1e-4, all times i had loss degrade to <code>nan</code>. I also notice that this happens when i continue training while loading a checkpoint, not sure if this is terribly important</p>",
      "rawMarkdown": "I lowered from 1e-8 (default) to 1e-6 and then to 1e-4, all times i had loss degrade to `nan`. I also notice that this happens when i continue training while loading a checkpoint, not sure if this is terribly important",
      "votes": null
    },
    {
      "id": "885184",
      "postDate": "06/14/2020 02:00:47",
      "content": "<p>Training loss. In some experiments, <code>apex</code> would scale down to 0 and the training never finishes, in other ones it does but then the loss is all <code>nan</code> during validation.</p>\n\n<p>To answer your other questions:\n- the loss does decrease but during that faulty epoch  I think i see it jumping higher than usual (say, from 0.5 to 1.4).\n- learning rate is 0.005 initially which i decrease during plateau by a factor 0.5. Now that i think about it, my LR was 0.0025 before i started getting <code>nan</code>. \n- I haven't tried lowering it manually.</p>",
      "rawMarkdown": "Training loss. In some experiments, `apex` would scale down to 0 and the training never finishes, in other ones it does but then the loss is all `nan` during validation.\n\nTo answer your other questions:\n- the loss does decrease but during that faulty epoch  I think i see it jumping higher than usual (say, from 0.5 to 1.4).\n- learning rate is 0.005 initially which i decrease during plateau by a factor 0.5. Now that i think about it, my LR was 0.0025 before i started getting `nan`. \n- I haven't tried lowering it manually.",
      "votes": null
    },
    {
      "id": "885474",
      "postDate": "06/14/2020 08:47:55",
      "content": "<p>are you sure that fp32 training runs as expected?</p>",
      "rawMarkdown": "are you sure that fp32 training runs as expected?",
      "votes": null
    },
    {
      "id": "886227",
      "postDate": "06/14/2020 20:10:51",
      "content": "<p>Yep. I just loaded the fp16 pre-trained model and started same training script without half-precision and it works. Can't figure out what the source of the problem is</p>",
      "rawMarkdown": "Yep. I just loaded the fp16 pre-trained model and started same training script without half-precision and it works. Can't figure out what the source of the problem is",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 883927,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "06/13/2020 05:00:29",
      "content": "<p>Training loss or validation loss? Is the loss generally decreasing up until that point? What learning rate are you using? Have you tried lowering it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 885184,
          "author_name": "iilmer",
          "author_url": "",
          "post_date": "06/14/2020 02:00:47",
          "content": "<p>Training loss. In some experiments, <code>apex</code> would scale down to 0 and the training never finishes, in other ones it does but then the loss is all <code>nan</code> during validation.</p>\n\n<p>To answer your other questions:\n- the loss does decrease but during that faulty epoch  I think i see it jumping higher than usual (say, from 0.5 to 1.4).\n- learning rate is 0.005 initially which i decrease during plateau by a factor 0.5. Now that i think about it, my LR was 0.0025 before i started getting <code>nan</code>. \n- I haven't tried lowering it manually.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 884739,
      "author_name": "vovanf98",
      "author_url": "",
      "post_date": "06/13/2020 15:32:33",
      "content": "<p>what is your current eps value? I faced the same problem. Solved it by limiting minimum lr with 5e-5.</p>",
      "votes": null,
      "replies": [
        {
          "id": 885182,
          "author_name": "iilmer",
          "author_url": "",
          "post_date": "06/14/2020 01:56:36",
          "content": "<p>I lowered from 1e-8 (default) to 1e-6 and then to 1e-4, all times i had loss degrade to <code>nan</code>. I also notice that this happens when i continue training while loading a checkpoint, not sure if this is terribly important</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 885474,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "06/14/2020 08:47:55",
          "content": "<p>are you sure that fp32 training runs as expected?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 886227,
          "author_name": "iilmer",
          "author_url": "",
          "post_date": "06/14/2020 20:10:51",
          "content": "<p>Yep. I just loaded the fp16 pre-trained model and started same training script without half-precision and it works. Can't figure out what the source of the problem is</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "883814": "I am using half-precision training with RAdam optimizer and keep getting a `nan` loss especially after 14 epochs. It tried opt 01 and opt 02 levels but it did not make a difference. I tested the input data and no `nan`'s occur there, so it must be something in the optimizer/precision. I tried increasing the `eps` in RAdam, too, but to no avail. Has anybody dealt with such issue?",
    "883927": "Training loss or validation loss? Is the loss generally decreasing up until that point? What learning rate are you using? Have you tried lowering it?",
    "884739": "what is your current eps value? I faced the same problem. Solved it by limiting minimum lr with 5e-5.",
    "885182": "I lowered from 1e-8 (default) to 1e-6 and then to 1e-4, all times i had loss degrade to `nan`. I also notice that this happens when i continue training while loading a checkpoint, not sure if this is terribly important",
    "885184": "Training loss. In some experiments, `apex` would scale down to 0 and the training never finishes, in other ones it does but then the loss is all `nan` during validation.\n\nTo answer your other questions:\n- the loss does decrease but during that faulty epoch  I think i see it jumping higher than usual (say, from 0.5 to 1.4).\n- learning rate is 0.005 initially which i decrease during plateau by a factor 0.5. Now that i think about it, my LR was 0.0025 before i started getting `nan`. \n- I haven't tried lowering it manually.",
    "885474": "are you sure that fp32 training runs as expected?",
    "886227": "Yep. I just loaded the fp16 pre-trained model and started same training script without half-precision and it works. Can't figure out what the source of the problem is"
  },
  "source": "meta"
}