{
  "id": 91820,
  "title": "Vanishing gradient?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91820",
  "author_name": "",
  "post_date": "2019-05-09T11:15:52.566590200Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I heard a lot about vanishing gradient being a problem during model training especially when the neural network is deep, but I've never seen the problem. Has anyone experienced it in practice? What's the effect?</p>",
  "messages": [
    {
      "id": "529185",
      "postDate": "05/09/2019 11:15:52",
      "content": "<p>I heard a lot about vanishing gradient being a problem during model training especially when the neural network is deep, but I've never seen the problem. Has anyone experienced it in practice? What's the effect?</p>",
      "rawMarkdown": "I heard a lot about vanishing gradient being a problem during model training especially when the neural network is deep, but I've never seen the problem. Has anyone experienced it in practice? What's the effect?",
      "votes": null
    },
    {
      "id": "529199",
      "postDate": "05/09/2019 12:04:43",
      "content": "<p>It happened mainly when sigmoid or tanh were the activation functions.  If you use relu, then vanishing gradient becomes a problem only if you use very deep networks and don't initialize weights the right way.  See  for some explanations: <a href=\"https://towardsdatascience.com/weight-initialization-in-neural-networks-a-journey-from-the-basics-to-kaiming-954fb9b47c79\">https://towardsdatascience.com/weight-initialization-in-neural-networks-a-journey-from-the-basics-to-kaiming-954fb9b47c79</a></p>\n\n<p>Note that exploding gradient can now be a problem.  It was not a problem with sigmoid or tanh activations.</p>",
      "rawMarkdown": "It happened mainly when sigmoid or tanh were the activation functions.  If you use relu, then vanishing gradient becomes a problem only if you use very deep networks and don't initialize weights the right way.  See  for some explanations: https://towardsdatascience.com/weight-initialization-in-neural-networks-a-journey-from-the-basics-to-kaiming-954fb9b47c79\n\nNote that exploding gradient can now be a problem.  It was not a problem with sigmoid or tanh activations.",
      "votes": null
    },
    {
      "id": "529218",
      "postDate": "05/09/2019 13:06:08",
      "content": "<p>Thanks for answering! How do you identify during training that vanishing(or exploding) gradient is happening? Is it by seeing that loss remains high and doesn't converge?</p>",
      "rawMarkdown": "Thanks for answering! How do you identify during training that vanishing(or exploding) gradient is happening? Is it by seeing that loss remains high and doesn't converge?",
      "votes": null
    },
    {
      "id": "529256",
      "postDate": "05/09/2019 14:27:19",
      "content": "<p>Hi <a href=\"/vivaroma\">@vivaroma</a>, \nVanishing gradient means that the gradients of earlier layers are very small. Therefore, the corresponding weights barely move after the update of backpropagation. \nYou can identify this problem by looking at the gradients with respect to the weights of earlier layers. If they are very small then you have a vanishing gradient.</p>",
      "rawMarkdown": "Hi @vivaroma, \nVanishing gradient means that the gradients of earlier layers are very small. Therefore, the corresponding weights barely move after the update of backpropagation. \nYou can identify this problem by looking at the gradients with respect to the weights of earlier layers. If they are very small then you have a vanishing gradient.",
      "votes": null
    },
    {
      "id": "529263",
      "postDate": "05/09/2019 14:37:16",
      "content": "<p>As said, it should not happen if you use small number of layers.  Also, recent CNN architectures like resnet pass residual info from layers to layers, eliminating in practice vanishing gradient.  For exploding gradient, you see if if loss suddenly explodes.</p>",
      "rawMarkdown": "As said, it should not happen if you use small number of layers.  Also, recent CNN architectures like resnet pass residual info from layers to layers, eliminating in practice vanishing gradient.  For exploding gradient, you see if if loss suddenly explodes.",
      "votes": null
    },
    {
      "id": "529351",
      "postDate": "05/09/2019 17:39:19",
      "content": "<p><a href=\"/vivaroma\">@vivaroma</a>, a small gradient means that the weights and biases of the initial layers will not be updated effectively with each training session. Since these initial layers are often crucial to recognizing the core elements of the input data, it can lead to overall inaccuracy of the whole network.</p>",
      "rawMarkdown": "vivaroma, a small gradient means that the weights and biases of the initial layers will not be updated effectively with each training session. Since these initial layers are often crucial to recognizing the core elements of the input data, it can lead to overall inaccuracy of the whole network.",
      "votes": null
    },
    {
      "id": "529443",
      "postDate": "05/09/2019 22:50:43",
      "content": "<p>I see...Do you actively check the gradients in practice to avoid vanishing gradient? Sounds like it only happens under certain conditions/architectures, and we might not need to worry about it otherwise.</p>",
      "rawMarkdown": "I see...Do you actively check the gradients in practice to avoid vanishing gradient? Sounds like it only happens under certain conditions/architectures, and we might not need to worry about it otherwise.",
      "votes": null
    },
    {
      "id": "529666",
      "postDate": "05/10/2019 13:14:19",
      "content": "<p>As everyone else said.  I'll also add that vanishing gradient affected specially Vanilla RNN because propagating the error along the time made the gradient to become exponentially small for long time sequences. LSTM and GRU solves this issue teorically but on practice you have to be carefull because with some really long sequences you will most probably have vanishing gradient.</p>\n\n<p>If you are interested more on this topic I'll suggest reading this <a href=\"https://medium.com/datadriveninvestor/how-do-lstm-networks-solve-the-problem-of-vanishing-gradients-a6784971a577\">post</a> which explains everything really well.</p>",
      "rawMarkdown": "As everyone else said.  I'll also add that vanishing gradient affected specially Vanilla RNN because propagating the error along the time made the gradient to become exponentially small for long time sequences. LSTM and GRU solves this issue teorically but on practice you have to be carefull because with some really long sequences you will most probably have vanishing gradient.\n\nIf you are interested more on this topic I'll suggest reading this [post](https://medium.com/datadriveninvestor/how-do-lstm-networks-solve-the-problem-of-vanishing-gradients-a6784971a577) which explains everything really well.",
      "votes": null
    },
    {
      "id": "531785",
      "postDate": "05/15/2019 14:37:09",
      "content": "<p>How do you deal with exploding gradient then?</p>",
      "rawMarkdown": "How do you deal with exploding gradient then?",
      "votes": null
    },
    {
      "id": "531798",
      "postDate": "05/15/2019 15:02:58",
      "content": "<p>lower learning rate, clip gradient, check sample weights, ...</p>",
      "rawMarkdown": "lower learning rate, clip gradient, check sample weights, ...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 529199,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/09/2019 12:04:43",
      "content": "<p>It happened mainly when sigmoid or tanh were the activation functions.  If you use relu, then vanishing gradient becomes a problem only if you use very deep networks and don't initialize weights the right way.  See  for some explanations: <a href=\"https://towardsdatascience.com/weight-initialization-in-neural-networks-a-journey-from-the-basics-to-kaiming-954fb9b47c79\">https://towardsdatascience.com/weight-initialization-in-neural-networks-a-journey-from-the-basics-to-kaiming-954fb9b47c79</a></p>\n\n<p>Note that exploding gradient can now be a problem.  It was not a problem with sigmoid or tanh activations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 529218,
          "author_name": "vivaroma",
          "author_url": "",
          "post_date": "05/09/2019 13:06:08",
          "content": "<p>Thanks for answering! How do you identify during training that vanishing(or exploding) gradient is happening? Is it by seeing that loss remains high and doesn't converge?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 529263,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/09/2019 14:37:16",
          "content": "<p>As said, it should not happen if you use small number of layers.  Also, recent CNN architectures like resnet pass residual info from layers to layers, eliminating in practice vanishing gradient.  For exploding gradient, you see if if loss suddenly explodes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531785,
          "author_name": "vivaroma",
          "author_url": "",
          "post_date": "05/15/2019 14:37:09",
          "content": "<p>How do you deal with exploding gradient then?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531798,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/15/2019 15:02:58",
          "content": "<p>lower learning rate, clip gradient, check sample weights, ...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 529256,
      "author_name": "adnanemiri",
      "author_url": "",
      "post_date": "05/09/2019 14:27:19",
      "content": "<p>Hi <a href=\"/vivaroma\">@vivaroma</a>, \nVanishing gradient means that the gradients of earlier layers are very small. Therefore, the corresponding weights barely move after the update of backpropagation. \nYou can identify this problem by looking at the gradients with respect to the weights of earlier layers. If they are very small then you have a vanishing gradient.</p>",
      "votes": null,
      "replies": [
        {
          "id": 529443,
          "author_name": "vivaroma",
          "author_url": "",
          "post_date": "05/09/2019 22:50:43",
          "content": "<p>I see...Do you actively check the gradients in practice to avoid vanishing gradient? Sounds like it only happens under certain conditions/architectures, and we might not need to worry about it otherwise.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 529351,
      "author_name": "makinghistory",
      "author_url": "",
      "post_date": "05/09/2019 17:39:19",
      "content": "<p><a href=\"/vivaroma\">@vivaroma</a>, a small gradient means that the weights and biases of the initial layers will not be updated effectively with each training session. Since these initial layers are often crucial to recognizing the core elements of the input data, it can lead to overall inaccuracy of the whole network.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529666,
      "author_name": "guillemdelgado",
      "author_url": "",
      "post_date": "05/10/2019 13:14:19",
      "content": "<p>As everyone else said.  I'll also add that vanishing gradient affected specially Vanilla RNN because propagating the error along the time made the gradient to become exponentially small for long time sequences. LSTM and GRU solves this issue teorically but on practice you have to be carefull because with some really long sequences you will most probably have vanishing gradient.</p>\n\n<p>If you are interested more on this topic I'll suggest reading this <a href=\"https://medium.com/datadriveninvestor/how-do-lstm-networks-solve-the-problem-of-vanishing-gradients-a6784971a577\">post</a> which explains everything really well.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "529185": "I heard a lot about vanishing gradient being a problem during model training especially when the neural network is deep, but I've never seen the problem. Has anyone experienced it in practice? What's the effect?",
    "529199": "It happened mainly when sigmoid or tanh were the activation functions.  If you use relu, then vanishing gradient becomes a problem only if you use very deep networks and don't initialize weights the right way.  See  for some explanations: https://towardsdatascience.com/weight-initialization-in-neural-networks-a-journey-from-the-basics-to-kaiming-954fb9b47c79\n\nNote that exploding gradient can now be a problem.  It was not a problem with sigmoid or tanh activations.",
    "529218": "Thanks for answering! How do you identify during training that vanishing(or exploding) gradient is happening? Is it by seeing that loss remains high and doesn't converge?",
    "529256": "Hi @vivaroma, \nVanishing gradient means that the gradients of earlier layers are very small. Therefore, the corresponding weights barely move after the update of backpropagation. \nYou can identify this problem by looking at the gradients with respect to the weights of earlier layers. If they are very small then you have a vanishing gradient.",
    "529263": "As said, it should not happen if you use small number of layers.  Also, recent CNN architectures like resnet pass residual info from layers to layers, eliminating in practice vanishing gradient.  For exploding gradient, you see if if loss suddenly explodes.",
    "529351": "vivaroma, a small gradient means that the weights and biases of the initial layers will not be updated effectively with each training session. Since these initial layers are often crucial to recognizing the core elements of the input data, it can lead to overall inaccuracy of the whole network.",
    "529443": "I see...Do you actively check the gradients in practice to avoid vanishing gradient? Sounds like it only happens under certain conditions/architectures, and we might not need to worry about it otherwise.",
    "529666": "As everyone else said.  I'll also add that vanishing gradient affected specially Vanilla RNN because propagating the error along the time made the gradient to become exponentially small for long time sequences. LSTM and GRU solves this issue teorically but on practice you have to be carefull because with some really long sequences you will most probably have vanishing gradient.\n\nIf you are interested more on this topic I'll suggest reading this [post](https://medium.com/datadriveninvestor/how-do-lstm-networks-solve-the-problem-of-vanishing-gradients-a6784971a577) which explains everything really well.",
    "531785": "How do you deal with exploding gradient then?",
    "531798": "lower learning rate, clip gradient, check sample weights, ..."
  },
  "source": "meta"
}