{
  "id": 120102,
  "title": "Does Layer Norm still need batch size (>=4)?",
  "url": "/competitions/tensorflow2-question-answering/discussion/120102",
  "author_name": "",
  "post_date": "2019-12-03T17:22:57.644823400Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi kagglers,\nI try to use pytorch to reproduce their Bert joint baseline,\nI only have one 11GB memory GPU. \nTheir Bert large setting will need 8GB memory for batch size = 1.\nTherefore, I apply accumulate gradient with batch size = 1.\nAnd it doesn't work.</p>\n\n<p>My question is: \n1. does Layer Norm still need batch size (original paper has batch size &gt; 4) to get some good estimation on layer's statistic? \n2. or other layer in Bert complain batch size = 1 ?</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "686915",
      "postDate": "12/03/2019 17:22:57",
      "content": "<p>Hi kagglers,\nI try to use pytorch to reproduce their Bert joint baseline,\nI only have one 11GB memory GPU. \nTheir Bert large setting will need 8GB memory for batch size = 1.\nTherefore, I apply accumulate gradient with batch size = 1.\nAnd it doesn't work.</p>\n\n<p>My question is: \n1. does Layer Norm still need batch size (original paper has batch size &gt; 4) to get some good estimation on layer's statistic? \n2. or other layer in Bert complain batch size = 1 ?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi kagglers,\nI try to use pytorch to reproduce their Bert joint baseline,\nI only have one 11GB memory GPU. \nTheir Bert large setting will need 8GB memory for batch size = 1.\nTherefore, I apply accumulate gradient with batch size = 1.\nAnd it doesn't work.\n\nMy question is: \n1. does Layer Norm still need batch size (original paper has batch size &gt; 4) to get some good estimation on layer's statistic? \n2. or other layer in Bert complain batch size = 1 ?\n\nThanks",
      "votes": null
    },
    {
      "id": "687076",
      "postDate": "12/03/2019 22:15:23",
      "content": "<p>I think that layer normalization is computing its means and variances over only a single training example.  It's not supposed to depend on the batch size, so gradient accumulation ought to work.  I don't think there's other stuff in Bert that would cause gradient accumulation to fail, but I haven't tried a pytorch Bert implementation.</p>",
      "rawMarkdown": "I think that layer normalization is computing its means and variances over only a single training example.  It's not supposed to depend on the batch size, so gradient accumulation ought to work.  I don't think there's other stuff in Bert that would cause gradient accumulation to fail, but I haven't tried a pytorch Bert implementation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 687076,
      "author_name": "particlebbq",
      "author_url": "",
      "post_date": "12/03/2019 22:15:23",
      "content": "<p>I think that layer normalization is computing its means and variances over only a single training example.  It's not supposed to depend on the batch size, so gradient accumulation ought to work.  I don't think there's other stuff in Bert that would cause gradient accumulation to fail, but I haven't tried a pytorch Bert implementation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "686915": "Hi kagglers,\nI try to use pytorch to reproduce their Bert joint baseline,\nI only have one 11GB memory GPU. \nTheir Bert large setting will need 8GB memory for batch size = 1.\nTherefore, I apply accumulate gradient with batch size = 1.\nAnd it doesn't work.\n\nMy question is: \n1. does Layer Norm still need batch size (original paper has batch size &gt; 4) to get some good estimation on layer's statistic? \n2. or other layer in Bert complain batch size = 1 ?\n\nThanks",
    "687076": "I think that layer normalization is computing its means and variances over only a single training example.  It's not supposed to depend on the batch size, so gradient accumulation ought to work.  I don't think there's other stuff in Bert that would cause gradient accumulation to fail, but I haven't tried a pytorch Bert implementation."
  },
  "source": "meta"
}