{
  "id": 201404,
  "title": "Batch Normalization after or before Activation function?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201404",
  "author_name": "",
  "post_date": "2020-12-04T19:05:05.595980Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have read numerous amount of papers that debate whether BN should be applied after activation function or before activation function.</p>\n<p>Has anyone experienced a change in LB here based on the position of BN in fully connected network? </p>",
  "messages": [
    {
      "id": "1102299",
      "postDate": "12/04/2020 19:05:05",
      "content": "<p>I have read numerous amount of papers that debate whether BN should be applied after activation function or before activation function.</p>\n<p>Has anyone experienced a change in LB here based on the position of BN in fully connected network? </p>",
      "rawMarkdown": "I have read numerous amount of papers that debate whether BN should be applied after activation function or before activation function.\n\nHas anyone experienced a change in LB here based on the position of BN in fully connected network?",
      "votes": null
    },
    {
      "id": "1102395",
      "postDate": "12/04/2020 21:57:18",
      "content": "<p>The BN consists of normalizing the output by subtracting the mean and dividing by the variance (and 2 additionals operations ). Therefore, there will be more negative values , so if u are using a ReLU just after, you will eliminate those negative values and you'll get the most of this activation function. So, the best option is to use the BN before the activation function.</p>",
      "rawMarkdown": "The BN consists of normalizing the output by subtracting the mean and dividing by the variance (and 2 additionals operations ). Therefore, there will be more negative values , so if u are using a ReLU just after, you will eliminate those negative values and you'll get the most of this activation function. So, the best option is to use the BN before the activation function.",
      "votes": null
    },
    {
      "id": "1102994",
      "postDate": "12/05/2020 14:47:29",
      "content": "<p>That is right but what in case where my activation function is sigmoid or tanh? Here I will already have value in a particular range.</p>",
      "rawMarkdown": "That is right but what in case where my activation function is sigmoid or tanh? Here I will already have value in a particular range.",
      "votes": null
    },
    {
      "id": "1103024",
      "postDate": "12/05/2020 15:19:16",
      "content": "<p>For sigmoid and tanh, we would want to stay in the \"Linear\" range of this functions ( to avoid the zero gradient near 0 and 1 for sigmoid and -1 and 1 for tanh ). For consequence, not having a very high (or very low ) values before the activation function would make the probability of bumping into that range higher and that's what we want ( cuz the gradient will not saturate in this case ). The BN will do that for us by shrinking the values before the activation functions.<br>\nIn Geoffrey Hinton paper about ReLU, he proved that a ReLU function is equivalent to a stack of sigmoids, so ReLU would be the best option in your hidden layers according to Hinton.</p>",
      "rawMarkdown": "For sigmoid and tanh, we would want to stay in the \"Linear\" range of this functions ( to avoid the zero gradient near 0 and 1 for sigmoid and -1 and 1 for tanh ). For consequence, not having a very high (or very low ) values before the activation function would make the probability of bumping into that range higher and that's what we want ( cuz the gradient will not saturate in this case ). The BN will do that for us by shrinking the values before the activation functions.\nIn Geoffrey Hinton paper about ReLU, he proved that a ReLU function is equivalent to a stack of sigmoids, so ReLU would be the best option in your hidden layers according to Hinton.",
      "votes": null
    },
    {
      "id": "1103591",
      "postDate": "12/06/2020 03:58:07",
      "content": "<p>Those are good arguments, but after the activation function that may leave non-centered data (which is not good for the next layer, as centered data is beneficial for gradient propagation and for) for activations such as sigmoid and ReLU.<br>\nAlso ReLU will be limited to a \"cut those bellow the mean\" operation and would limit the layer bias ability to define a threshold.<br>\nWe usually do normalization to data entering the network, why not do the same before certain layers (which would be the end of the last one, i.e. after activation)</p>",
      "rawMarkdown": "Those are good arguments, but after the activation function that may leave non-centered data (which is not good for the next layer, as centered data is beneficial for gradient propagation and for) for activations such as sigmoid and ReLU.\nAlso ReLU will be limited to a \"cut those bellow the mean\" operation and would limit the layer bias ability to define a threshold.\nWe usually do normalization to data entering the network, why not do the same before certain layers (which would be the end of the last one, i.e. after activation)",
      "votes": null
    },
    {
      "id": "1104014",
      "postDate": "12/06/2020 14:24:52",
      "content": "<p>Thank you Pedro for your argument, it was convincing. I think they have the same pros and cons in each case, and one should test both of them and see what's best for his particular situation.</p>",
      "rawMarkdown": "Thank you Pedro for your argument, it was convincing. I think they have the same pros and cons in each case, and one should test both of them and see what's best for his particular situation.",
      "votes": null
    },
    {
      "id": "1104339",
      "postDate": "12/06/2020 20:38:24",
      "content": "<p>Testing and seeing what works best, there is no argument against it hahaha</p>",
      "rawMarkdown": "Testing and seeing what works best, there is no argument against it hahaha",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1102395,
      "author_name": "harlequeen",
      "author_url": "",
      "post_date": "12/04/2020 21:57:18",
      "content": "<p>The BN consists of normalizing the output by subtracting the mean and dividing by the variance (and 2 additionals operations ). Therefore, there will be more negative values , so if u are using a ReLU just after, you will eliminate those negative values and you'll get the most of this activation function. So, the best option is to use the BN before the activation function.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1102994,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "12/05/2020 14:47:29",
          "content": "<p>That is right but what in case where my activation function is sigmoid or tanh? Here I will already have value in a particular range.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103024,
          "author_name": "harlequeen",
          "author_url": "",
          "post_date": "12/05/2020 15:19:16",
          "content": "<p>For sigmoid and tanh, we would want to stay in the \"Linear\" range of this functions ( to avoid the zero gradient near 0 and 1 for sigmoid and -1 and 1 for tanh ). For consequence, not having a very high (or very low ) values before the activation function would make the probability of bumping into that range higher and that's what we want ( cuz the gradient will not saturate in this case ). The BN will do that for us by shrinking the values before the activation functions.<br>\nIn Geoffrey Hinton paper about ReLU, he proved that a ReLU function is equivalent to a stack of sigmoids, so ReLU would be the best option in your hidden layers according to Hinton.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103591,
          "author_name": "phmonforte",
          "author_url": "",
          "post_date": "12/06/2020 03:58:07",
          "content": "<p>Those are good arguments, but after the activation function that may leave non-centered data (which is not good for the next layer, as centered data is beneficial for gradient propagation and for) for activations such as sigmoid and ReLU.<br>\nAlso ReLU will be limited to a \"cut those bellow the mean\" operation and would limit the layer bias ability to define a threshold.<br>\nWe usually do normalization to data entering the network, why not do the same before certain layers (which would be the end of the last one, i.e. after activation)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104014,
          "author_name": "harlequeen",
          "author_url": "",
          "post_date": "12/06/2020 14:24:52",
          "content": "<p>Thank you Pedro for your argument, it was convincing. I think they have the same pros and cons in each case, and one should test both of them and see what's best for his particular situation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104339,
          "author_name": "phmonforte",
          "author_url": "",
          "post_date": "12/06/2020 20:38:24",
          "content": "<p>Testing and seeing what works best, there is no argument against it hahaha</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1102299": "I have read numerous amount of papers that debate whether BN should be applied after activation function or before activation function.\n\nHas anyone experienced a change in LB here based on the position of BN in fully connected network?",
    "1102395": "The BN consists of normalizing the output by subtracting the mean and dividing by the variance (and 2 additionals operations ). Therefore, there will be more negative values , so if u are using a ReLU just after, you will eliminate those negative values and you'll get the most of this activation function. So, the best option is to use the BN before the activation function.",
    "1102994": "That is right but what in case where my activation function is sigmoid or tanh? Here I will already have value in a particular range.",
    "1103024": "For sigmoid and tanh, we would want to stay in the \"Linear\" range of this functions ( to avoid the zero gradient near 0 and 1 for sigmoid and -1 and 1 for tanh ). For consequence, not having a very high (or very low ) values before the activation function would make the probability of bumping into that range higher and that's what we want ( cuz the gradient will not saturate in this case ). The BN will do that for us by shrinking the values before the activation functions.\nIn Geoffrey Hinton paper about ReLU, he proved that a ReLU function is equivalent to a stack of sigmoids, so ReLU would be the best option in your hidden layers according to Hinton.",
    "1103591": "Those are good arguments, but after the activation function that may leave non-centered data (which is not good for the next layer, as centered data is beneficial for gradient propagation and for) for activations such as sigmoid and ReLU.\nAlso ReLU will be limited to a \"cut those bellow the mean\" operation and would limit the layer bias ability to define a threshold.\nWe usually do normalization to data entering the network, why not do the same before certain layers (which would be the end of the last one, i.e. after activation)",
    "1104014": "Thank you Pedro for your argument, it was convincing. I think they have the same pros and cons in each case, and one should test both of them and see what's best for his particular situation.",
    "1104339": "Testing and seeing what works best, there is no argument against it hahaha"
  },
  "source": "meta"
}