{
  "id": 74449,
  "title": "Batch Normalization & NLP ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74449",
  "author_name": "",
  "post_date": "2018-12-12T10:28:04.896544300Z",
  "votes": 15,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I was discussing the use of Batch Normalization with Puck Wang, the author of the GRU/Capsule kernel ( <a href=\"https://www.kaggle.com/gmhost/gru-capsule\">https://www.kaggle.com/gmhost/gru-capsule</a> ).</p>\n\n<p>Batch Normalization is a technique that has proven to be very effective for deep neural network. However, I did not read a lot about it for shallow architectures such as the ones common in NLP. </p>\n\n<p>It was also question of where to place the BatchNorm, there are two possibilities :</p>\n\n<ul>\n<li>Dense + BatchNorm +Relu</li>\n<li>Dense + ReLu + BatchNorm</li>\n</ul>\n\n<p>The first architecture being the one prescribbed in the original paper ( see  <a href=\"https://arxiv.org/pdf/1502.03167v3.pdf\">https://arxiv.org/pdf/1502.03167v3.pdf</a> )</p>\n\n<p>If you have any thoughts on the topic, I'd like to hear them ! Thanks !</p>",
  "messages": [
    {
      "id": "437699",
      "postDate": "12/12/2018 10:28:04",
      "content": "<p>I was discussing the use of Batch Normalization with Puck Wang, the author of the GRU/Capsule kernel ( <a href=\"https://www.kaggle.com/gmhost/gru-capsule\">https://www.kaggle.com/gmhost/gru-capsule</a> ).</p>\n\n<p>Batch Normalization is a technique that has proven to be very effective for deep neural network. However, I did not read a lot about it for shallow architectures such as the ones common in NLP. </p>\n\n<p>It was also question of where to place the BatchNorm, there are two possibilities :</p>\n\n<ul>\n<li>Dense + BatchNorm +Relu</li>\n<li>Dense + ReLu + BatchNorm</li>\n</ul>\n\n<p>The first architecture being the one prescribbed in the original paper ( see  <a href=\"https://arxiv.org/pdf/1502.03167v3.pdf\">https://arxiv.org/pdf/1502.03167v3.pdf</a> )</p>\n\n<p>If you have any thoughts on the topic, I'd like to hear them ! Thanks !</p>",
      "rawMarkdown": "I was discussing the use of Batch Normalization with Puck Wang, the author of the GRU/Capsule kernel ( https://www.kaggle.com/gmhost/gru-capsule ).\n\nBatch Normalization is a technique that has proven to be very effective for deep neural network. However, I did not read a lot about it for shallow architectures such as the ones common in NLP. \n\nIt was also question of where to place the BatchNorm, there are two possibilities :\n\n - Dense + BatchNorm +Relu\n - Dense + ReLu + BatchNorm\n\nThe first architecture being the one prescribbed in the original paper ( see  https://arxiv.org/pdf/1502.03167v3.pdf )\n\nIf you have any thoughts on the topic, I'd like to hear them ! Thanks !",
      "votes": null
    },
    {
      "id": "437760",
      "postDate": "12/12/2018 12:59:56",
      "content": "<p>For me <code>Dense + ReLu + BatchNorm</code> , Batch Normalization  after ReLU makes much more sense. The weight matrix is more centred to the data. And, deals with largerer x values from (ReLU <code>max(0,x)</code>)</p>\n\n<p>However, <a href=\"https://forums.fast.ai/t/questions-about-batch-normalization/230/5?u=shaz13\">Jeremy Howard thinks</a> -</p>\n\n<blockquote>\n  <p>When I implemented this (BN between linear and non-linear layers) I did some searching and more recent advice seems to be to put it after the non-linearity, based on some experiments.</p>\n</blockquote>\n\n<p>And in some other discussions I saw that BN doesn't help the model at all. So, it really depends on experimenting and see which one works better</p>",
      "rawMarkdown": "For me `Dense + ReLu + BatchNorm` , Batch Normalization  after ReLU makes much more sense. The weight matrix is more centred to the data. And, deals with largerer x values from (ReLU `max(0,x)`)\n\nHowever, [Jeremy Howard thinks][1] -\n&gt; When I implemented this (BN between linear and non-linear layers) I did some searching and more recent advice seems to be to put it after the non-linearity, based on some experiments.\n\nAnd in some other discussions I saw that BN doesn't help the model at all. So, it really depends on experimenting and see which one works better\n\n\n\n[1]: https://forums.fast.ai/t/questions-about-batch-normalization/230/5?u=shaz13",
      "votes": null
    },
    {
      "id": "437796",
      "postDate": "12/12/2018 13:58:38",
      "content": "<p>hah, same with me.</p>",
      "rawMarkdown": "hah, same with me.",
      "votes": null
    },
    {
      "id": "437962",
      "postDate": "12/12/2018 21:07:51",
      "content": "<p>Thanks for your insights. Doing Dense + BN + ReLu did not work very well for me so far, I'll keep on experimenting.</p>",
      "rawMarkdown": "Thanks for your insights. Doing Dense + BN + ReLu did not work very well for me so far, I'll keep on experimenting.",
      "votes": null
    },
    {
      "id": "441219",
      "postDate": "12/18/2018 12:29:45",
      "content": "<p>Dense + ReLu + BN works well for me, both cv and lb improved, thank you for sharing!</p>",
      "rawMarkdown": "Dense + ReLu + BN works well for me, both cv and lb improved, thank you for sharing!",
      "votes": null
    },
    {
      "id": "441230",
      "postDate": "12/18/2018 12:45:24",
      "content": "<p>Thanks for sharing your results :)</p>",
      "rawMarkdown": "Thanks for sharing your results :)",
      "votes": null
    },
    {
      "id": "441337",
      "postDate": "12/18/2018 15:02:48",
      "content": "<p>Hi Theo Viel,\nBelow, Andrew Ng recommended to use the first choice : Dense + BN + Relu :)\n<a href=\"https://www.youtube.com/watch?v=em6dfRxYkYU\">https://www.youtube.com/watch?v=em6dfRxYkYU</a></p>\n\n<p>For NLP related (not directed task), Dense + BN + Relu also used in DeepSpeech2 of Baidu which is state of the art at that time. (See Table 5 of <a href=\"http://proceedings.mlr.press/v48/amodei16.pdf\">http://proceedings.mlr.press/v48/amodei16.pdf</a> )</p>",
      "rawMarkdown": "Hi Theo Viel,\nBelow, Andrew Ng recommended to use the first choice : Dense + BN + Relu :)\nhttps://www.youtube.com/watch?v=em6dfRxYkYU\n\nFor NLP related (not directed task), Dense + BN + Relu also used in DeepSpeech2 of Baidu which is state of the art at that time. (See Table 5 of http://proceedings.mlr.press/v48/amodei16.pdf )",
      "votes": null
    },
    {
      "id": "441539",
      "postDate": "12/18/2018 19:11:46",
      "content": "<p>Thanks, I'll check that out !</p>",
      "rawMarkdown": "Thanks, I'll check that out !",
      "votes": null
    },
    {
      "id": "443089",
      "postDate": "12/21/2018 02:10:45",
      "content": "<p>i remember checking out the batch norm paper and thing which made it so famous that it managed to achieve state of art result at time with almost half training time.I guess batch norm may not help with improving scores as much but it may lower number of epochs you need to train each model therefore perhaps helping you to sqeeze in one more model in your ensemble :D</p>",
      "rawMarkdown": "i remember checking out the batch norm paper and thing which made it so famous that it managed to achieve state of art result at time with almost half training time.I guess batch norm may not help with improving scores as much but it may lower number of epochs you need to train each model therefore perhaps helping you to sqeeze in one more model in your ensemble :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 437760,
      "author_name": "shaz13",
      "author_url": "",
      "post_date": "12/12/2018 12:59:56",
      "content": "<p>For me <code>Dense + ReLu + BatchNorm</code> , Batch Normalization  after ReLU makes much more sense. The weight matrix is more centred to the data. And, deals with largerer x values from (ReLU <code>max(0,x)</code>)</p>\n\n<p>However, <a href=\"https://forums.fast.ai/t/questions-about-batch-normalization/230/5?u=shaz13\">Jeremy Howard thinks</a> -</p>\n\n<blockquote>\n  <p>When I implemented this (BN between linear and non-linear layers) I did some searching and more recent advice seems to be to put it after the non-linearity, based on some experiments.</p>\n</blockquote>\n\n<p>And in some other discussions I saw that BN doesn't help the model at all. So, it really depends on experimenting and see which one works better</p>",
      "votes": null,
      "replies": [
        {
          "id": 437796,
          "author_name": "gmhost",
          "author_url": "",
          "post_date": "12/12/2018 13:58:38",
          "content": "<p>hah, same with me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437962,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "12/12/2018 21:07:51",
          "content": "<p>Thanks for your insights. Doing Dense + BN + ReLu did not work very well for me so far, I'll keep on experimenting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441219,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "12/18/2018 12:29:45",
      "content": "<p>Dense + ReLu + BN works well for me, both cv and lb improved, thank you for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 441230,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "12/18/2018 12:45:24",
          "content": "<p>Thanks for sharing your results :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441337,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "12/18/2018 15:02:48",
      "content": "<p>Hi Theo Viel,\nBelow, Andrew Ng recommended to use the first choice : Dense + BN + Relu :)\n<a href=\"https://www.youtube.com/watch?v=em6dfRxYkYU\">https://www.youtube.com/watch?v=em6dfRxYkYU</a></p>\n\n<p>For NLP related (not directed task), Dense + BN + Relu also used in DeepSpeech2 of Baidu which is state of the art at that time. (See Table 5 of <a href=\"http://proceedings.mlr.press/v48/amodei16.pdf\">http://proceedings.mlr.press/v48/amodei16.pdf</a> )</p>",
      "votes": null,
      "replies": [
        {
          "id": 441539,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "12/18/2018 19:11:46",
          "content": "<p>Thanks, I'll check that out !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443089,
      "author_name": "chansbest93",
      "author_url": "",
      "post_date": "12/21/2018 02:10:45",
      "content": "<p>i remember checking out the batch norm paper and thing which made it so famous that it managed to achieve state of art result at time with almost half training time.I guess batch norm may not help with improving scores as much but it may lower number of epochs you need to train each model therefore perhaps helping you to sqeeze in one more model in your ensemble :D</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "437699": "I was discussing the use of Batch Normalization with Puck Wang, the author of the GRU/Capsule kernel ( https://www.kaggle.com/gmhost/gru-capsule ).\n\nBatch Normalization is a technique that has proven to be very effective for deep neural network. However, I did not read a lot about it for shallow architectures such as the ones common in NLP. \n\nIt was also question of where to place the BatchNorm, there are two possibilities :\n\n - Dense + BatchNorm +Relu\n - Dense + ReLu + BatchNorm\n\nThe first architecture being the one prescribbed in the original paper ( see  https://arxiv.org/pdf/1502.03167v3.pdf )\n\nIf you have any thoughts on the topic, I'd like to hear them ! Thanks !",
    "437760": "For me `Dense + ReLu + BatchNorm` , Batch Normalization  after ReLU makes much more sense. The weight matrix is more centred to the data. And, deals with largerer x values from (ReLU `max(0,x)`)\n\nHowever, [Jeremy Howard thinks][1] -\n&gt; When I implemented this (BN between linear and non-linear layers) I did some searching and more recent advice seems to be to put it after the non-linearity, based on some experiments.\n\nAnd in some other discussions I saw that BN doesn't help the model at all. So, it really depends on experimenting and see which one works better\n\n\n\n[1]: https://forums.fast.ai/t/questions-about-batch-normalization/230/5?u=shaz13",
    "437796": "hah, same with me.",
    "437962": "Thanks for your insights. Doing Dense + BN + ReLu did not work very well for me so far, I'll keep on experimenting.",
    "441219": "Dense + ReLu + BN works well for me, both cv and lb improved, thank you for sharing!",
    "441230": "Thanks for sharing your results :)",
    "441337": "Hi Theo Viel,\nBelow, Andrew Ng recommended to use the first choice : Dense + BN + Relu :)\nhttps://www.youtube.com/watch?v=em6dfRxYkYU\n\nFor NLP related (not directed task), Dense + BN + Relu also used in DeepSpeech2 of Baidu which is state of the art at that time. (See Table 5 of http://proceedings.mlr.press/v48/amodei16.pdf )",
    "441539": "Thanks, I'll check that out !",
    "443089": "i remember checking out the batch norm paper and thing which made it so famous that it managed to achieve state of art result at time with almost half training time.I guess batch norm may not help with improving scores as much but it may lower number of epochs you need to train each model therefore perhaps helping you to sqeeze in one more model in your ensemble :D"
  },
  "source": "meta"
}