{
  "id": 125438,
  "title": "how to solve sudden \"validation loss : nan\" problem?",
  "url": "/competitions/pku-autonomous-driving/discussion/125438",
  "author_name": "",
  "post_date": "2020-01-10T15:21:49.293795900Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>i was training  a model with cyclic learning rate and after 7th epoch i get nan validation loss..isn't it \"exploding gradient problem\"? will gradient accumulation be able to solve this issue? i don't get such errors when i try adam or adamW,</p>\n\n<h1>training statistics so far of the model i am trying :</h1>\n\n<p>Train Epoch: 0  LR: 0.0049400000    Loss: 1.878861\nDev loss: 2.1991\nTrain Epoch: 1  LR: 0.0086800000    Loss: 1.840539\nDev loss: 1.7849\nTrain Epoch: 2  LR: 0.0075800000    Loss: 1.847198\nDev loss: 1.9127\nTrain Epoch: 3  LR: 0.0038400000    Loss: 1.287447\nDev loss: 1.3331\nTrain Epoch: 4  LR: 0.0023000000    Loss: 1.416327\nDev loss: 1.2588\nTrain Epoch: 5  LR: 0.0060400000    Loss: 1.299999\nDev loss: 1.4838\nTrain Epoch: 6  LR: 0.0097800000    Loss: 1.540868\nDev loss: 1.5280\nTrain Epoch: 7  LR: 0.0064800000    Loss: 1.790969\nDev loss: 1.2738\nTrain Epoch: 8  LR: 0.0027400000    Loss: 1.092477\nDev loss: nan</p>",
  "messages": [
    {
      "id": "715479",
      "postDate": "01/10/2020 15:21:49",
      "content": "<p>i was training  a model with cyclic learning rate and after 7th epoch i get nan validation loss..isn't it \"exploding gradient problem\"? will gradient accumulation be able to solve this issue? i don't get such errors when i try adam or adamW,</p>\n\n<h1>training statistics so far of the model i am trying :</h1>\n\n<p>Train Epoch: 0  LR: 0.0049400000    Loss: 1.878861\nDev loss: 2.1991\nTrain Epoch: 1  LR: 0.0086800000    Loss: 1.840539\nDev loss: 1.7849\nTrain Epoch: 2  LR: 0.0075800000    Loss: 1.847198\nDev loss: 1.9127\nTrain Epoch: 3  LR: 0.0038400000    Loss: 1.287447\nDev loss: 1.3331\nTrain Epoch: 4  LR: 0.0023000000    Loss: 1.416327\nDev loss: 1.2588\nTrain Epoch: 5  LR: 0.0060400000    Loss: 1.299999\nDev loss: 1.4838\nTrain Epoch: 6  LR: 0.0097800000    Loss: 1.540868\nDev loss: 1.5280\nTrain Epoch: 7  LR: 0.0064800000    Loss: 1.790969\nDev loss: 1.2738\nTrain Epoch: 8  LR: 0.0027400000    Loss: 1.092477\nDev loss: nan</p>",
      "rawMarkdown": "i was training  a model with cyclic learning rate and after 7th epoch i get nan validation loss..isn't it \"exploding gradient problem\"? will gradient accumulation be able to solve this issue? i don't get such errors when i try adam or adamW,\n\n#training statistics so far of the model i am trying : \n\nTrain Epoch: 0 \tLR: 0.0049400000\tLoss: 1.878861\nDev loss: 2.1991\nTrain Epoch: 1 \tLR: 0.0086800000\tLoss: 1.840539\nDev loss: 1.7849\nTrain Epoch: 2 \tLR: 0.0075800000\tLoss: 1.847198\nDev loss: 1.9127\nTrain Epoch: 3 \tLR: 0.0038400000\tLoss: 1.287447\nDev loss: 1.3331\nTrain Epoch: 4 \tLR: 0.0023000000\tLoss: 1.416327\nDev loss: 1.2588\nTrain Epoch: 5 \tLR: 0.0060400000\tLoss: 1.299999\nDev loss: 1.4838\nTrain Epoch: 6 \tLR: 0.0097800000\tLoss: 1.540868\nDev loss: 1.5280\nTrain Epoch: 7 \tLR: 0.0064800000\tLoss: 1.790969\nDev loss: 1.2738\nTrain Epoch: 8 \tLR: 0.0027400000\tLoss: 1.092477\nDev loss: nan",
      "votes": null
    },
    {
      "id": "715733",
      "postDate": "01/10/2020 19:05:46",
      "content": "<p>I have the same issue as well, using Adam</p>",
      "rawMarkdown": "I have the same issue as well, using Adam",
      "votes": null
    },
    {
      "id": "715738",
      "postDate": "01/10/2020 19:08:19",
      "content": "<p>Maybe  It's because of those learning rate schedulers? ReduceLrOnPlateau didn’t help as well.</p>",
      "rawMarkdown": "Maybe  It's because of those learning rate schedulers? ReduceLrOnPlateau didn’t help as well.",
      "votes": null
    },
    {
      "id": "715791",
      "postDate": "01/10/2020 21:06:04",
      "content": "<p>Does the error effect training? Are results hindered?</p>",
      "rawMarkdown": "Does the error effect training? Are results hindered?",
      "votes": null
    },
    {
      "id": "715830",
      "postDate": "01/10/2020 21:46:49",
      "content": "<p>What about decreasing the learning rate? I have an experience with NaN loss in the past. Although it was for another problem, in my case, the problem was that the learning rate was too big. After I decreased the learning rate, everything was fine. </p>\n\n<p>If you google about \"NaN loss\", you might find more suggestions (e.g., <a href=\"https://stackoverflow.com/questions/40050397/deep-learning-nan-loss-reasons\">https://stackoverflow.com/questions/40050397/deep-learning-nan-loss-reasons</a>). Changing the optimizer might solve the problem in some cases. Some people also suggest to check whether there is NaN value in the training data. </p>",
      "rawMarkdown": "What about decreasing the learning rate? I have an experience with NaN loss in the past. Although it was for another problem, in my case, the problem was that the learning rate was too big. After I decreased the learning rate, everything was fine. \n\nIf you google about \"NaN loss\", you might find more suggestions (e.g., https://stackoverflow.com/questions/40050397/deep-learning-nan-loss-reasons). Changing the optimizer might solve the problem in some cases. Some people also suggest to check whether there is NaN value in the training data.",
      "votes": null
    },
    {
      "id": "715931",
      "postDate": "01/11/2020 01:53:17",
      "content": "<p>no nan in our training data,i am using default 1e-3 for AdamW and still getting nan</p>",
      "rawMarkdown": "no nan in our training data,i am using default 1e-3 for AdamW and still getting nan",
      "votes": null
    },
    {
      "id": "715932",
      "postDate": "01/11/2020 01:55:30",
      "content": "<p>error effecting training</p>",
      "rawMarkdown": "error effecting training",
      "votes": null
    },
    {
      "id": "716145",
      "postDate": "01/11/2020 10:11:20",
      "content": "<p>If I remember it correctly, when I got this issue, I used Adam with learning rate 1e-3. Then I changed the learning rate into either 5e-4 or 5e-5.</p>\n\n<p>In general, NaN is produced when there is a floating-point operation that produces an undefined result, e.g., division by 0, etc (you might want to look more about when NaN is produced here: <a href=\"https://stackoverflow.com/questions/25506281/what-are-all-the-possible-calculations-that-could-cause-a-nan-in-python?answertab=votes#tab-top\">https://stackoverflow.com/questions/25506281/what-are-all-the-possible-calculations-that-could-cause-a-nan-in-python?answertab=votes#tab-top</a> ). So, I think you might need to find which operation that could cause NaN. In my experience, another reason that I remember was that there was a bug in my code that makes some training batches have size 0. If you suspect that the problem is related to the gradient, maybe you could also try to use gradient clipping (I think it is also a common approach).</p>\n\n<p>In the end, when I looked into this problem, it seems that there could be many possible reasons for this. </p>",
      "rawMarkdown": "If I remember it correctly, when I got this issue, I used Adam with learning rate 1e-3. Then I changed the learning rate into either 5e-4 or 5e-5.\n\nIn general, NaN is produced when there is a floating-point operation that produces an undefined result, e.g., division by 0, etc (you might want to look more about when NaN is produced here: https://stackoverflow.com/questions/25506281/what-are-all-the-possible-calculations-that-could-cause-a-nan-in-python?answertab=votes#tab-top ). So, I think you might need to find which operation that could cause NaN. In my experience, another reason that I remember was that there was a bug in my code that makes some training batches have size 0. If you suspect that the problem is related to the gradient, maybe you could also try to use gradient clipping (I think it is also a common approach).\n\nIn the end, when I looked into this problem, it seems that there could be many possible reasons for this.",
      "votes": null
    },
    {
      "id": "716170",
      "postDate": "01/11/2020 10:58:39",
      "content": "<p>i just figured out why i got nan but don't know how to solve this issue!\ni just increased kernel size of maxpooling from 2 to 3 and i got validation loss = nan but when i use kernel_size = 2 i don't get this nan error in validation loss! <a href=\"/kagglarsa\">@kagglarsa</a> </p>",
      "rawMarkdown": "i just figured out why i got nan but don't know how to solve this issue!\ni just increased kernel size of maxpooling from 2 to 3 and i got validation loss = nan but when i use kernel_size = 2 i don't get this nan error in validation loss! @kagglarsa",
      "votes": null
    },
    {
      "id": "716620",
      "postDate": "01/12/2020 02:21:45",
      "content": "<p>I did not have problems about max-pool. you can check the tensor changes before and after max-pool.\nMaybe some values  changed wrongly.</p>",
      "rawMarkdown": "I did not have problems about max-pool. you can check the tensor changes before and after max-pool.\nMaybe some values  changed wrongly.",
      "votes": null
    },
    {
      "id": "716978",
      "postDate": "01/12/2020 14:46:05",
      "content": "<p>I lowered lr and it seems to have fixed the error, an added bonus is my model converges better and faster too.</p>",
      "rawMarkdown": "I lowered lr and it seems to have fixed the error, an added bonus is my model converges better and faster too.",
      "votes": null
    },
    {
      "id": "717119",
      "postDate": "01/12/2020 18:58:58",
      "content": "<p>Did you also change the stride? Someone reported that: if the stride is larger than the kernel size of the pooling, it may cause NaN. See the link below:</p>\n\n<p><a href=\"https://stackoverflow.com/questions/33962226/common-causes-of-nans-during-training?answertab=votes#tab-top\">https://stackoverflow.com/questions/33962226/common-causes-of-nans-during-training?answertab=votes#tab-top</a></p>\n\n<p>However, theoretically speaking, I still don't really get why this is the case. I have to recall how the derivatives are propagated in max pooling. So far I couldn't think about an example that causes this situation. Probably the reason of NaN is related to the result after pooling, and maybe it is related to the change of the size of the pooling output. Notice that, when you change the max pool kernel size, it will affect the size of the pooling output. Since you make the max pool size bigger, the output dimension should become smaller. Could it be the case that the next layer requires an input with bigger dimension, and when it got an input with smaller dimension, a problem occurred? But this is just a wild guess.</p>",
      "rawMarkdown": "Did you also change the stride? Someone reported that: if the stride is larger than the kernel size of the pooling, it may cause NaN. See the link below:\n\nhttps://stackoverflow.com/questions/33962226/common-causes-of-nans-during-training?answertab=votes#tab-top\n\nHowever, theoretically speaking, I still don't really get why this is the case. I have to recall how the derivatives are propagated in max pooling. So far I couldn't think about an example that causes this situation. Probably the reason of NaN is related to the result after pooling, and maybe it is related to the change of the size of the pooling output. Notice that, when you change the max pool kernel size, it will affect the size of the pooling output. Since you make the max pool size bigger, the output dimension should become smaller. Could it be the case that the next layer requires an input with bigger dimension, and when it got an input with smaller dimension, a problem occurred? But this is just a wild guess.",
      "votes": null
    },
    {
      "id": "719312",
      "postDate": "01/15/2020 11:21:06",
      "content": "<p>log(0) gives NaN . </p>",
      "rawMarkdown": "log(0) gives NaN .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 715733,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "01/10/2020 19:05:46",
      "content": "<p>I have the same issue as well, using Adam</p>",
      "votes": null,
      "replies": [
        {
          "id": 715738,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/10/2020 19:08:19",
          "content": "<p>Maybe  It's because of those learning rate schedulers? ReduceLrOnPlateau didn’t help as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715791,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "01/10/2020 21:06:04",
          "content": "<p>Does the error effect training? Are results hindered?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715932,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/11/2020 01:55:30",
          "content": "<p>error effecting training</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 715830,
      "author_name": "kagglarsa",
      "author_url": "",
      "post_date": "01/10/2020 21:46:49",
      "content": "<p>What about decreasing the learning rate? I have an experience with NaN loss in the past. Although it was for another problem, in my case, the problem was that the learning rate was too big. After I decreased the learning rate, everything was fine. </p>\n\n<p>If you google about \"NaN loss\", you might find more suggestions (e.g., <a href=\"https://stackoverflow.com/questions/40050397/deep-learning-nan-loss-reasons\">https://stackoverflow.com/questions/40050397/deep-learning-nan-loss-reasons</a>). Changing the optimizer might solve the problem in some cases. Some people also suggest to check whether there is NaN value in the training data. </p>",
      "votes": null,
      "replies": [
        {
          "id": 715931,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/11/2020 01:53:17",
          "content": "<p>no nan in our training data,i am using default 1e-3 for AdamW and still getting nan</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716145,
          "author_name": "kagglarsa",
          "author_url": "",
          "post_date": "01/11/2020 10:11:20",
          "content": "<p>If I remember it correctly, when I got this issue, I used Adam with learning rate 1e-3. Then I changed the learning rate into either 5e-4 or 5e-5.</p>\n\n<p>In general, NaN is produced when there is a floating-point operation that produces an undefined result, e.g., division by 0, etc (you might want to look more about when NaN is produced here: <a href=\"https://stackoverflow.com/questions/25506281/what-are-all-the-possible-calculations-that-could-cause-a-nan-in-python?answertab=votes#tab-top\">https://stackoverflow.com/questions/25506281/what-are-all-the-possible-calculations-that-could-cause-a-nan-in-python?answertab=votes#tab-top</a> ). So, I think you might need to find which operation that could cause NaN. In my experience, another reason that I remember was that there was a bug in my code that makes some training batches have size 0. If you suspect that the problem is related to the gradient, maybe you could also try to use gradient clipping (I think it is also a common approach).</p>\n\n<p>In the end, when I looked into this problem, it seems that there could be many possible reasons for this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716170,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/11/2020 10:58:39",
          "content": "<p>i just figured out why i got nan but don't know how to solve this issue!\ni just increased kernel size of maxpooling from 2 to 3 and i got validation loss = nan but when i use kernel_size = 2 i don't get this nan error in validation loss! <a href=\"/kagglarsa\">@kagglarsa</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716620,
          "author_name": "guobaozi",
          "author_url": "",
          "post_date": "01/12/2020 02:21:45",
          "content": "<p>I did not have problems about max-pool. you can check the tensor changes before and after max-pool.\nMaybe some values  changed wrongly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716978,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "01/12/2020 14:46:05",
          "content": "<p>I lowered lr and it seems to have fixed the error, an added bonus is my model converges better and faster too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 717119,
          "author_name": "kagglarsa",
          "author_url": "",
          "post_date": "01/12/2020 18:58:58",
          "content": "<p>Did you also change the stride? Someone reported that: if the stride is larger than the kernel size of the pooling, it may cause NaN. See the link below:</p>\n\n<p><a href=\"https://stackoverflow.com/questions/33962226/common-causes-of-nans-during-training?answertab=votes#tab-top\">https://stackoverflow.com/questions/33962226/common-causes-of-nans-during-training?answertab=votes#tab-top</a></p>\n\n<p>However, theoretically speaking, I still don't really get why this is the case. I have to recall how the derivatives are propagated in max pooling. So far I couldn't think about an example that causes this situation. Probably the reason of NaN is related to the result after pooling, and maybe it is related to the change of the size of the pooling output. Notice that, when you change the max pool kernel size, it will affect the size of the pooling output. Since you make the max pool size bigger, the output dimension should become smaller. Could it be the case that the next layer requires an input with bigger dimension, and when it got an input with smaller dimension, a problem occurred? But this is just a wild guess.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 719312,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "01/15/2020 11:21:06",
      "content": "<p>log(0) gives NaN . </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "715479": "i was training  a model with cyclic learning rate and after 7th epoch i get nan validation loss..isn't it \"exploding gradient problem\"? will gradient accumulation be able to solve this issue? i don't get such errors when i try adam or adamW,\n\n#training statistics so far of the model i am trying : \n\nTrain Epoch: 0 \tLR: 0.0049400000\tLoss: 1.878861\nDev loss: 2.1991\nTrain Epoch: 1 \tLR: 0.0086800000\tLoss: 1.840539\nDev loss: 1.7849\nTrain Epoch: 2 \tLR: 0.0075800000\tLoss: 1.847198\nDev loss: 1.9127\nTrain Epoch: 3 \tLR: 0.0038400000\tLoss: 1.287447\nDev loss: 1.3331\nTrain Epoch: 4 \tLR: 0.0023000000\tLoss: 1.416327\nDev loss: 1.2588\nTrain Epoch: 5 \tLR: 0.0060400000\tLoss: 1.299999\nDev loss: 1.4838\nTrain Epoch: 6 \tLR: 0.0097800000\tLoss: 1.540868\nDev loss: 1.5280\nTrain Epoch: 7 \tLR: 0.0064800000\tLoss: 1.790969\nDev loss: 1.2738\nTrain Epoch: 8 \tLR: 0.0027400000\tLoss: 1.092477\nDev loss: nan",
    "715733": "I have the same issue as well, using Adam",
    "715738": "Maybe  It's because of those learning rate schedulers? ReduceLrOnPlateau didn’t help as well.",
    "715791": "Does the error effect training? Are results hindered?",
    "715830": "What about decreasing the learning rate? I have an experience with NaN loss in the past. Although it was for another problem, in my case, the problem was that the learning rate was too big. After I decreased the learning rate, everything was fine. \n\nIf you google about \"NaN loss\", you might find more suggestions (e.g., https://stackoverflow.com/questions/40050397/deep-learning-nan-loss-reasons). Changing the optimizer might solve the problem in some cases. Some people also suggest to check whether there is NaN value in the training data.",
    "715931": "no nan in our training data,i am using default 1e-3 for AdamW and still getting nan",
    "715932": "error effecting training",
    "716145": "If I remember it correctly, when I got this issue, I used Adam with learning rate 1e-3. Then I changed the learning rate into either 5e-4 or 5e-5.\n\nIn general, NaN is produced when there is a floating-point operation that produces an undefined result, e.g., division by 0, etc (you might want to look more about when NaN is produced here: https://stackoverflow.com/questions/25506281/what-are-all-the-possible-calculations-that-could-cause-a-nan-in-python?answertab=votes#tab-top ). So, I think you might need to find which operation that could cause NaN. In my experience, another reason that I remember was that there was a bug in my code that makes some training batches have size 0. If you suspect that the problem is related to the gradient, maybe you could also try to use gradient clipping (I think it is also a common approach).\n\nIn the end, when I looked into this problem, it seems that there could be many possible reasons for this.",
    "716170": "i just figured out why i got nan but don't know how to solve this issue!\ni just increased kernel size of maxpooling from 2 to 3 and i got validation loss = nan but when i use kernel_size = 2 i don't get this nan error in validation loss! @kagglarsa",
    "716620": "I did not have problems about max-pool. you can check the tensor changes before and after max-pool.\nMaybe some values  changed wrongly.",
    "716978": "I lowered lr and it seems to have fixed the error, an added bonus is my model converges better and faster too.",
    "717119": "Did you also change the stride? Someone reported that: if the stride is larger than the kernel size of the pooling, it may cause NaN. See the link below:\n\nhttps://stackoverflow.com/questions/33962226/common-causes-of-nans-during-training?answertab=votes#tab-top\n\nHowever, theoretically speaking, I still don't really get why this is the case. I have to recall how the derivatives are propagated in max pooling. So far I couldn't think about an example that causes this situation. Probably the reason of NaN is related to the result after pooling, and maybe it is related to the change of the size of the pooling output. Notice that, when you change the max pool kernel size, it will affect the size of the pooling output. Since you make the max pool size bigger, the output dimension should become smaller. Could it be the case that the next layer requires an input with bigger dimension, and when it got an input with smaller dimension, a problem occurred? But this is just a wild guess.",
    "719312": "log(0) gives NaN ."
  },
  "source": "meta"
}