{
  "id": 74490,
  "title": "what batch size are you using ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74490",
  "author_name": "",
  "post_date": "2018-12-12T20:38:19.388018300Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I saw a lot of kernels were using a batch size = 512. But with this batch size, my algorithm always stops at one point then it's doesn't no matter how many epochs I give it will move from that point. </p>\n\n<p>But when I just changed the batch size = 4096 ,\n the training is going fine. </p>\n\n<p>Can I know the reason for this </p>",
  "messages": [
    {
      "id": "437950",
      "postDate": "12/12/2018 20:38:19",
      "content": "<p>I saw a lot of kernels were using a batch size = 512. But with this batch size, my algorithm always stops at one point then it's doesn't no matter how many epochs I give it will move from that point. </p>\n\n<p>But when I just changed the batch size = 4096 ,\n the training is going fine. </p>\n\n<p>Can I know the reason for this </p>",
      "rawMarkdown": "I saw a lot of kernels were using a batch size = 512. But with this batch size, my algorithm always stops at one point then it's doesn't no matter how many epochs I give it will move from that point. \n\nBut when I just changed the batch size = 4096 ,\n the training is going fine. \n\nCan I know the reason for this",
      "votes": null
    },
    {
      "id": "437963",
      "postDate": "12/12/2018 21:09:30",
      "content": "<p>Bigger batch size reduces gradient noise. Multiplying batch size by 8 should have same stabilizing effect as dividing learning rate by 8. That goes for vanilla SGD and SGD with momentum, I suppose ADAM, etc. should converge more robustly (and overlearn more easily too).</p>",
      "rawMarkdown": "Bigger batch size reduces gradient noise. Multiplying batch size by 8 should have same stabilizing effect as dividing learning rate by 8. That goes for vanilla SGD and SGD with momentum, I suppose ADAM, etc. should converge more robustly (and overlearn more easily too).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 437963,
      "author_name": "pollux",
      "author_url": "",
      "post_date": "12/12/2018 21:09:30",
      "content": "<p>Bigger batch size reduces gradient noise. Multiplying batch size by 8 should have same stabilizing effect as dividing learning rate by 8. That goes for vanilla SGD and SGD with momentum, I suppose ADAM, etc. should converge more robustly (and overlearn more easily too).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "437950": "I saw a lot of kernels were using a batch size = 512. But with this batch size, my algorithm always stops at one point then it's doesn't no matter how many epochs I give it will move from that point. \n\nBut when I just changed the batch size = 4096 ,\n the training is going fine. \n\nCan I know the reason for this",
    "437963": "Bigger batch size reduces gradient noise. Multiplying batch size by 8 should have same stabilizing effect as dividing learning rate by 8. That goes for vanilla SGD and SGD with momentum, I suppose ADAM, etc. should converge more robustly (and overlearn more easily too)."
  },
  "source": "meta"
}