{
  "id": 258706,
  "title": "A strange \"Loss to NaN\" training issue",
  "url": "/competitions/seti-breakthrough-listen/discussion/258706",
  "author_name": "Vadim Timakin",
  "post_date": "2021-08-03T12:15:45.159000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi everyone! I've always been using my own pipeline for all the CV tasks, this task isn't an exception. <br>\nBut here I faced a strange problem - my loss goes to NaN after a couple of epochs. But the strange thing is that this doesn't happen with all the models. For example EfficientNet B0 trains fine and achieves good accuracy at the end. But ResNet34d, EffNet B4, NfNet l0 just don't train, first one or two epochs loss almost doesn't change and after that it goes to nan value.<br>\nI'm not using any unique approaches or tools for training yet, all my parameters are kinda default.</p>\n<p>I've tried:</p>\n<ul>\n<li>Turn OFF gradient accumulation</li>\n<li>Turn ON gradient clipping</li>\n<li>Turn OFF mixed precision</li>\n<li>Turn OFF warm up</li>\n<li>Different optimizers: Adam, Ranger, AdamW</li>\n</ul>\n<p>Nothing of that has fixed the issue. What might be the reason?</p>",
  "messages": [
    {
      "id": 1431201,
      "postDate": "2021-08-03T12:15:45.160Z",
      "content": "<p>Hi everyone! I've always been using my own pipeline for all the CV tasks, this task isn't an exception. <br>\nBut here I faced a strange problem - my loss goes to NaN after a couple of epochs. But the strange thing is that this doesn't happen with all the models. For example EfficientNet B0 trains fine and achieves good accuracy at the end. But ResNet34d, EffNet B4, NfNet l0 just don't train, first one or two epochs loss almost doesn't change and after that it goes to nan value.<br>\nI'm not using any unique approaches or tools for training yet, all my parameters are kinda default.</p>\n<p>I've tried:</p>\n<ul>\n<li>Turn OFF gradient accumulation</li>\n<li>Turn ON gradient clipping</li>\n<li>Turn OFF mixed precision</li>\n<li>Turn OFF warm up</li>\n<li>Different optimizers: Adam, Ranger, AdamW</li>\n</ul>\n<p>Nothing of that has fixed the issue. What might be the reason?</p>",
      "rawMarkdown": "Hi everyone! I've always been using my own pipeline for all the CV tasks, this task isn't an exception. \nBut here I faced a strange problem - my loss goes to NaN after a couple of epochs. But the strange thing is that this doesn't happen with all the models. For example EfficientNet B0 trains fine and achieves good accuracy at the end. But ResNet34d, EffNet B4, NfNet l0 just don't train, first one or two epochs loss almost doesn't change and after that it goes to nan value.\nI'm not using any unique approaches or tools for training yet, all my parameters are kinda default.\n\nI've tried:\n- Turn OFF gradient accumulation\n- Turn ON gradient clipping\n- Turn OFF mixed precision\n- Turn OFF warm up\n- Different optimizers: Adam, Ranger, AdamW\n\nNothing of that has fixed the issue. What might be the reason?",
      "votes": 1
    },
    {
      "id": 1440617,
      "postDate": "2021-08-03T18:35:09.843Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> , that was the key!</p>",
      "rawMarkdown": "Thank you @reighns @cpmpml , that was the key!",
      "votes": 2
    },
    {
      "id": 1442819,
      "postDate": "2021-08-04T00:31:59.027Z",
      "content": "<p>I ran into the same problem，my initial learning rate is 1e-4. gradient clipping does not work. but NAN will not appear when I don’t use mixed precision training.</p>",
      "rawMarkdown": "I ran into the same problem，my initial learning rate is 1e-4. gradient clipping does not work. but NAN will not appear when I don’t use mixed precision training."
    },
    {
      "id": 1441168,
      "postDate": "2021-08-03T19:46:33.013Z",
      "content": "<p>Depends on your initial learning rate and the type of scheduler and its params</p>",
      "rawMarkdown": "Depends on your initial learning rate and the type of scheduler and its params"
    },
    {
      "id": 1436779,
      "postDate": "2021-08-03T15:19:16.200Z",
      "content": "<p>decrease your learning rate</p>",
      "rawMarkdown": "decrease your learning rate"
    },
    {
      "id": 1436167,
      "postDate": "2021-08-03T14:58:40.913Z",
      "content": "<p>I am not sure but you should check whether your inputs/outputs contain NaN. </p>",
      "rawMarkdown": "I am not sure but you should check whether your inputs/outputs contain NaN. "
    },
    {
      "id": 1431353,
      "postDate": "2021-08-03T12:23:56.680Z",
      "content": "<p>I highly suggest being less aggressive on your initial learning rate - depending on the scheduler you use.</p>\n<p>Try lowering the initial lr and let me know if it works.</p>",
      "rawMarkdown": "I highly suggest being less aggressive on your initial learning rate - depending on the scheduler you use.\n\nTry lowering the initial lr and let me know if it works."
    },
    {
      "id": 1436163,
      "postDate": "2021-08-03T14:58:40.913Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1440617,
      "author_name": "Vadim Timakin",
      "author_url": "",
      "post_date": "2021-08-03T18:35:09.843000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> , that was the key!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1442819,
      "author_name": "Vvvvia",
      "author_url": "",
      "post_date": "2021-08-04T00:31:59.027000",
      "content": "<p>I ran into the same problem，my initial learning rate is 1e-4. gradient clipping does not work. but NAN will not appear when I don’t use mixed precision training.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1441168,
      "author_name": "Parth Dhameliya",
      "author_url": "",
      "post_date": "2021-08-03T19:46:33.013000",
      "content": "<p>Depends on your initial learning rate and the type of scheduler and its params</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1436779,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-08-03T15:19:16.200000",
      "content": "<p>decrease your learning rate</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1436167,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-08-03T14:58:40.913000",
      "content": "<p>I am not sure but you should check whether your inputs/outputs contain NaN. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1431353,
      "author_name": "gao-hongnan",
      "author_url": "",
      "post_date": "2021-08-03T12:23:56.680000",
      "content": "<p>I highly suggest being less aggressive on your initial learning rate - depending on the scheduler you use.</p>\n<p>Try lowering the initial lr and let me know if it works.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1436163,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-03T14:58:40.913000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1431201": "Hi everyone! I've always been using my own pipeline for all the CV tasks, this task isn't an exception. \nBut here I faced a strange problem - my loss goes to NaN after a couple of epochs. But the strange thing is that this doesn't happen with all the models. For example EfficientNet B0 trains fine and achieves good accuracy at the end. But ResNet34d, EffNet B4, NfNet l0 just don't train, first one or two epochs loss almost doesn't change and after that it goes to nan value.\nI'm not using any unique approaches or tools for training yet, all my parameters are kinda default.\n\nI've tried:\n- Turn OFF gradient accumulation\n- Turn ON gradient clipping\n- Turn OFF mixed precision\n- Turn OFF warm up\n- Different optimizers: Adam, Ranger, AdamW\n\nNothing of that has fixed the issue. What might be the reason?",
    "1440617": "Thank you @reighns @cpmpml , that was the key!",
    "1442819": "I ran into the same problem，my initial learning rate is 1e-4. gradient clipping does not work. but NAN will not appear when I don’t use mixed precision training.",
    "1441168": "Depends on your initial learning rate and the type of scheduler and its params",
    "1436779": "decrease your learning rate",
    "1436167": "I am not sure but you should check whether your inputs/outputs contain NaN. ",
    "1431353": "I highly suggest being less aggressive on your initial learning rate - depending on the scheduler you use.\n\nTry lowering the initial lr and let me know if it works.",
    "1436163": ""
  }
}