{
  "id": 44303,
  "title": "Small batch size (8) and decreasing accuracy",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/44303",
  "author_name": "",
  "post_date": "2017-11-26T23:05:21.162241300Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When i  am training  an Xception model with all parameters trainable, my val train/acc   decrease    if i decrease the batch size from 16 to 8 (I need to do that because of my 6G GPU).  Anyone encontered this problem?</p>",
  "messages": [
    {
      "id": "248767",
      "postDate": "11/26/2017 23:05:21",
      "content": "<p>When i  am training  an Xception model with all parameters trainable, my val train/acc   decrease    if i decrease the batch size from 16 to 8 (I need to do that because of my 6G GPU).  Anyone encontered this problem?</p>",
      "rawMarkdown": "When i  am training  an Xception model with all parameters trainable, my val train/acc   decrease    if i decrease the batch size from 16 to 8 (I need to do that because of my 6G GPU).  Anyone encontered this problem?",
      "votes": null
    },
    {
      "id": "248775",
      "postDate": "11/26/2017 23:37:31",
      "content": "<p><a href=\"https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf\">https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf</a></p>\n\n<blockquote>\n  <p>Linear Scaling Rule: When the minibatch size is multiplied by k, multiply the learning rate by k.</p>\n</blockquote>\n\n<p>If the batch size is reduced, the learning rate should be reduced.</p>",
      "rawMarkdown": "https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf\n&gt; Linear Scaling Rule: When the minibatch size is multiplied by k, multiply the learning rate by k.\n\nIf the batch size is reduced, the learning rate should be reduced.",
      "votes": null
    },
    {
      "id": "248801",
      "postDate": "11/27/2017 01:15:49",
      "content": "<p>There's a weird discrepancy between the theoretical amount the learning rate should scale by (sqrt(k) when batch size is multiplied by k) and the empirical best\n<a href=\"https://arxiv.org/pdf/1404.5997.pdf\">https://arxiv.org/pdf/1404.5997.pdf</a></p>",
      "rawMarkdown": "There's a weird discrepancy between the theoretical amount the learning rate should scale by (sqrt(k) when batch size is multiplied by k) and the empirical best\nhttps://arxiv.org/pdf/1404.5997.pdf",
      "votes": null
    },
    {
      "id": "248806",
      "postDate": "11/27/2017 01:43:44",
      "content": "<p>You may want to try batch renormalization layers instead of batch norm layers if you have those types of memory constraints. <a href=\"https://github.com/titu1994/BatchRenormalization\">Here's a link</a>  to an implementation in Keras.</p>",
      "rawMarkdown": "You may want to try batch renormalization layers instead of batch norm layers if you have those types of memory constraints. [Here's a link][1]  to an implementation in Keras.\n\n\n  [1]: https://github.com/titu1994/BatchRenormalization",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 248775,
      "author_name": "mwbyeon",
      "author_url": "",
      "post_date": "11/26/2017 23:37:31",
      "content": "<p><a href=\"https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf\">https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf</a></p>\n\n<blockquote>\n  <p>Linear Scaling Rule: When the minibatch size is multiplied by k, multiply the learning rate by k.</p>\n</blockquote>\n\n<p>If the batch size is reduced, the learning rate should be reduced.</p>",
      "votes": null,
      "replies": [
        {
          "id": 248801,
          "author_name": "eachshadow",
          "author_url": "",
          "post_date": "11/27/2017 01:15:49",
          "content": "<p>There's a weird discrepancy between the theoretical amount the learning rate should scale by (sqrt(k) when batch size is multiplied by k) and the empirical best\n<a href=\"https://arxiv.org/pdf/1404.5997.pdf\">https://arxiv.org/pdf/1404.5997.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 248806,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "11/27/2017 01:43:44",
      "content": "<p>You may want to try batch renormalization layers instead of batch norm layers if you have those types of memory constraints. <a href=\"https://github.com/titu1994/BatchRenormalization\">Here's a link</a>  to an implementation in Keras.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "248767": "When i  am training  an Xception model with all parameters trainable, my val train/acc   decrease    if i decrease the batch size from 16 to 8 (I need to do that because of my 6G GPU).  Anyone encontered this problem?",
    "248775": "https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf\n&gt; Linear Scaling Rule: When the minibatch size is multiplied by k, multiply the learning rate by k.\n\nIf the batch size is reduced, the learning rate should be reduced.",
    "248801": "There's a weird discrepancy between the theoretical amount the learning rate should scale by (sqrt(k) when batch size is multiplied by k) and the empirical best\nhttps://arxiv.org/pdf/1404.5997.pdf",
    "248806": "You may want to try batch renormalization layers instead of batch norm layers if you have those types of memory constraints. [Here's a link][1]  to an implementation in Keras.\n\n\n  [1]: https://github.com/titu1994/BatchRenormalization"
  },
  "source": "meta"
}