{
  "id": 187773,
  "title": "Exploding Gradient [NaN Output Values]",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/187773",
  "author_name": "Ali Abdin",
  "post_date": "2020-09-30T08:39:27.586000",
  "votes": 18,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I was playing around with the custom negative log likelihood loss from: <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence</a> </p>\n<p>Using a LR of 0.001 and a batch-size of 16 It was relatively easy to get exploding gradients which lead to NaN as output values after 7 iterations. around 1 out of 10 times I was able to get through the training process without exploding gradients, so it was a random effect for me. </p>\n<p>There are several known optimization \"fixes\":</p>\n<ul>\n<li>use lower LR </li>\n<li>use higher batch-size</li>\n<li>add dropout layer</li>\n<li>add batch normalisation layers to your model</li>\n</ul>\n<p>Recommended:</p>\n<ul>\n<li>use gradient clipping (<a href=\"https://towardsdatascience.com/what-is-gradient-clipping-b8e815cdfb48\" target=\"_blank\">https://towardsdatascience.com/what-is-gradient-clipping-b8e815cdfb48</a>)</li>\n</ul>\n<p>Did anyone notice something equal with exploding gradients? How did you fix it?</p>",
  "messages": [
    {
      "id": 1032485,
      "postDate": "2020-09-30T08:39:27.587Z",
      "content": "<p>Hi,</p>\n<p>I was playing around with the custom negative log likelihood loss from: <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence</a> </p>\n<p>Using a LR of 0.001 and a batch-size of 16 It was relatively easy to get exploding gradients which lead to NaN as output values after 7 iterations. around 1 out of 10 times I was able to get through the training process without exploding gradients, so it was a random effect for me. </p>\n<p>There are several known optimization \"fixes\":</p>\n<ul>\n<li>use lower LR </li>\n<li>use higher batch-size</li>\n<li>add dropout layer</li>\n<li>add batch normalisation layers to your model</li>\n</ul>\n<p>Recommended:</p>\n<ul>\n<li>use gradient clipping (<a href=\"https://towardsdatascience.com/what-is-gradient-clipping-b8e815cdfb48\" target=\"_blank\">https://towardsdatascience.com/what-is-gradient-clipping-b8e815cdfb48</a>)</li>\n</ul>\n<p>Did anyone notice something equal with exploding gradients? How did you fix it?</p>",
      "rawMarkdown": "Hi,\n\nI was playing around with the custom negative log likelihood loss from: https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence \n\nUsing a LR of 0.001 and a batch-size of 16 It was relatively easy to get exploding gradients which lead to NaN as output values after 7 iterations. around 1 out of 10 times I was able to get through the training process without exploding gradients, so it was a random effect for me. \n\nThere are several known optimization \"fixes\":\n- use lower LR \n- use higher batch-size\n- add dropout layer\n- add batch normalisation layers to your model\n\nRecommended:\n- use gradient clipping (https://towardsdatascience.com/what-is-gradient-clipping-b8e815cdfb48)\n\nDid anyone notice something equal with exploding gradients? How did you fix it?\n\n",
      "votes": 18
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1032485": "Hi,\n\nI was playing around with the custom negative log likelihood loss from: https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence \n\nUsing a LR of 0.001 and a batch-size of 16 It was relatively easy to get exploding gradients which lead to NaN as output values after 7 iterations. around 1 out of 10 times I was able to get through the training process without exploding gradients, so it was a random effect for me. \n\nThere are several known optimization \"fixes\":\n- use lower LR \n- use higher batch-size\n- add dropout layer\n- add batch normalisation layers to your model\n\nRecommended:\n- use gradient clipping (https://towardsdatascience.com/what-is-gradient-clipping-b8e815cdfb48)\n\nDid anyone notice something equal with exploding gradients? How did you fix it?\n\n"
  }
}