{
  "id": 203104,
  "title": "The order of BatchNorm and Residual operation in Attention-based models",
  "url": "/competitions/riiid-test-answer-prediction/discussion/203104",
  "author_name": "",
  "post_date": "2020-12-13T18:03:49.149750700Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In the original batch norm paper, the batch normalization <code>bn()</code> is performed before the residual operation, i.e., the output is <code>x+bn(f(x))</code>. However, some people argued that the batch norm should be placed after the residual operation <code>bn(x+f(x))</code>.</p>\n<p>In the nice SAKT model <a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> put up, it is the latter:<br>\n'''python<br>\nx = fc_layer_1(att_output)<br>\nx = batch_norm_layer(x + att_output)<br>\nx = fc_layer_2(x)<br>\n'''<br>\nSo I tested two orders using <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> 's CV file (cv3 to be specific) in a controlled environment (seed, batch size fixed) with a warm-up phase for both. Here is the result after epoch 1,3,5:</p>\n<table>\n<thead>\n<tr>\n<th>ep/ val_auc</th>\n<th>BN after Residual</th>\n<th>BN before Residual</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.7380</td>\n<td>0.7369</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.7418</td>\n<td>0.7415</td>\n</tr>\n<tr>\n<td>5</td>\n<td>0.7445</td>\n<td>0.7448</td>\n</tr>\n</tbody>\n</table>\n<p>Conclusion: as expected, the switching order in the later layers seems having little effect on the performance…I am testing the effect of this before feeding the attention outputs to an FC net.</p>",
  "messages": [
    {
      "id": "1111458",
      "postDate": "12/13/2020 18:03:49",
      "content": "<p>In the original batch norm paper, the batch normalization <code>bn()</code> is performed before the residual operation, i.e., the output is <code>x+bn(f(x))</code>. However, some people argued that the batch norm should be placed after the residual operation <code>bn(x+f(x))</code>.</p>\n<p>In the nice SAKT model <a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> put up, it is the latter:<br>\n'''python<br>\nx = fc_layer_1(att_output)<br>\nx = batch_norm_layer(x + att_output)<br>\nx = fc_layer_2(x)<br>\n'''<br>\nSo I tested two orders using <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> 's CV file (cv3 to be specific) in a controlled environment (seed, batch size fixed) with a warm-up phase for both. Here is the result after epoch 1,3,5:</p>\n<table>\n<thead>\n<tr>\n<th>ep/ val_auc</th>\n<th>BN after Residual</th>\n<th>BN before Residual</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.7380</td>\n<td>0.7369</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.7418</td>\n<td>0.7415</td>\n</tr>\n<tr>\n<td>5</td>\n<td>0.7445</td>\n<td>0.7448</td>\n</tr>\n</tbody>\n</table>\n<p>Conclusion: as expected, the switching order in the later layers seems having little effect on the performance…I am testing the effect of this before feeding the attention outputs to an FC net.</p>",
      "rawMarkdown": "In the original batch norm paper, the batch normalization `bn()` is performed before the residual operation, i.e., the output is `x+bn(f(x))`. However, some people argued that the batch norm should be placed after the residual operation `bn(x+f(x))`.\n\nIn the nice SAKT model @wangsg put up, it is the latter:\n'''python\nx = fc_layer_1(att_output)\nx = batch_norm_layer(x + att_output)\nx = fc_layer_2(x)\n'''\nSo I tested two orders using @its7171 's CV file (cv3 to be specific) in a controlled environment (seed, batch size fixed) with a warm-up phase for both. Here is the result after epoch 1,3,5:\n|ep/ val_auc | BN after Residual | BN before Residual  | \n| --- | --- | --- |\n| 1 |  0.7380  |0.7369 |\n| 3 | 0.7418 | 0.7415 |\n| 5 | 0.7445 |0.7448 |\n\n\nConclusion: as expected, the switching order in the later layers seems having little effect on the performance...I am testing the effect of this before feeding the attention outputs to an FC net.",
      "votes": null
    },
    {
      "id": "1116251",
      "postDate": "12/17/2020 01:53:29",
      "content": "<p>If your input <code>x</code> is also normalised, then two methods should perform similar.</p>",
      "rawMarkdown": "If your input `x` is also normalised, then two methods should perform similar.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1116251,
      "author_name": "gogo827jz",
      "author_url": "",
      "post_date": "12/17/2020 01:53:29",
      "content": "<p>If your input <code>x</code> is also normalised, then two methods should perform similar.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1111458": "In the original batch norm paper, the batch normalization `bn()` is performed before the residual operation, i.e., the output is `x+bn(f(x))`. However, some people argued that the batch norm should be placed after the residual operation `bn(x+f(x))`.\n\nIn the nice SAKT model @wangsg put up, it is the latter:\n'''python\nx = fc_layer_1(att_output)\nx = batch_norm_layer(x + att_output)\nx = fc_layer_2(x)\n'''\nSo I tested two orders using @its7171 's CV file (cv3 to be specific) in a controlled environment (seed, batch size fixed) with a warm-up phase for both. Here is the result after epoch 1,3,5:\n|ep/ val_auc | BN after Residual | BN before Residual  | \n| --- | --- | --- |\n| 1 |  0.7380  |0.7369 |\n| 3 | 0.7418 | 0.7415 |\n| 5 | 0.7445 |0.7448 |\n\n\nConclusion: as expected, the switching order in the later layers seems having little effect on the performance...I am testing the effect of this before feeding the attention outputs to an FC net.",
    "1116251": "If your input `x` is also normalised, then two methods should perform similar."
  },
  "source": "meta"
}