{
  "id": 324208,
  "title": "Does larger batch size of training hurt the performance of model?",
  "url": "/competitions/birdclef-2022/discussion/324208",
  "author_name": "Givan",
  "post_date": "2022-05-10T14:20:32.363000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I rerun <a href=\"https://www.kaggle.com/kaerunantoka\" target=\"_blank\">@kaerunantoka</a> ' s  <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">training notebook</a> in my local machine with larger batch size, which is set to 128. In order to keep optimization iteration as the same roughly, I increase the training epoch from 23 to 200 (and lower lr a little). </p>\n<p>The iteration step of public notebook is (18452 / 5 * 4) / 16 * 23 = 17079, gaining <strong>0.71</strong> in the public leaderboard. And mine is (18452 / 5 * 4) / 128 * 200 = 18565, gaining <strong>0.68</strong> instead.</p>\n<p>What are the possible reasons behind it?</p>",
  "messages": [
    {
      "id": 1784223,
      "postDate": "2022-05-11T03:24:01.153Z",
      "content": "<p>I faced a similar problem. Perhaps the cause is the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318999\" target=\"_blank\">optimal threshold</a>.</p>\n<p>The optimal threshold varies widely by epoch, iteration, and lr. Even by 0.01 up or down, the score can vary greatly.<br>\nUnfortunately, the only way to find the optimal threshold is through trial and trial.</p>",
      "rawMarkdown": "I faced a similar problem. Perhaps the cause is the [optimal threshold](https://www.kaggle.com/competitions/birdclef-2022/discussion/318999).\n\nThe optimal threshold varies widely by epoch, iteration, and lr. Even by 0.01 up or down, the score can vary greatly.\nUnfortunately, the only way to find the optimal threshold is through trial and trial.",
      "votes": 4
    },
    {
      "id": 1784857,
      "postDate": "2022-05-11T14:42:04.623Z",
      "content": "<p>The paper [1] discusses the problem of learning difficulty at large batch sizes.<br>\nThis paper focuses on the problem that existing learning algorithms do not scale to large batch sizes, and proposes a learning algorithm called LARS that adapts to large batch sizes.</p>\n<p>Another problem with some learning algorithms is that the optimal learning rate shifts when the batch size is changed [2]. For example, YOLOv5 addresses this problem by adopting a policy of setting the learning rate proportional to the batch size.</p>\n<h1>Reference</h1>\n<p>[1] LARS: <a href=\"https://arxiv.org/abs/1708.03888\" target=\"_blank\">https://arxiv.org/abs/1708.03888</a><br>\n[2] <a href=\"https://arxiv.org/abs/1706.02677\" target=\"_blank\">https://arxiv.org/abs/1706.02677</a></p>",
      "rawMarkdown": "The paper [1] discusses the problem of learning difficulty at large batch sizes.\nThis paper focuses on the problem that existing learning algorithms do not scale to large batch sizes, and proposes a learning algorithm called LARS that adapts to large batch sizes.\n\nAnother problem with some learning algorithms is that the optimal learning rate shifts when the batch size is changed [2]. For example, YOLOv5 addresses this problem by adopting a policy of setting the learning rate proportional to the batch size.\n\n# Reference\n[1] LARS: https://arxiv.org/abs/1708.03888\n[2] https://arxiv.org/abs/1706.02677",
      "votes": 1,
      "replies": [
        {
          "id": 1784877,
          "postDate": "2022-05-11T14:51:13.877Z",
          "content": "<p>There are other learning algorithm addressing large batch size:</p>\n<ul>\n<li>LAMB: <a href=\"https://arxiv.org/abs/1904.00962\" target=\"_blank\">https://arxiv.org/abs/1904.00962</a></li>\n<li>NVLAMB: <a href=\"https://medium.com/nvidia-ai/a-guide-to-optimizer-implementation-for-bert-at-scale-8338cc7f45fd\" target=\"_blank\">https://medium.com/nvidia-ai/a-guide-to-optimizer-implementation-for-bert-at-scale-8338cc7f45fd</a></li>\n<li>NovoGrad: <a href=\"https://arxiv.org/pdf/1905.11286.pdf\" target=\"_blank\">https://arxiv.org/pdf/1905.11286.pdf</a></li>\n</ul>",
          "rawMarkdown": "There are other learning algorithm addressing large batch size:\n- LAMB: https://arxiv.org/abs/1904.00962\n- NVLAMB: https://medium.com/nvidia-ai/a-guide-to-optimizer-implementation-for-bert-at-scale-8338cc7f45fd\n- NovoGrad: https://arxiv.org/pdf/1905.11286.pdf",
          "votes": 2
        },
        {
          "id": 1784883,
          "postDate": "2022-05-11T15:01:11.780Z",
          "content": "<p>pytorch lightning has its own implementation of LARS.<br>\nI haven't use this yet, but this might solve the problem with large batch size.<br>\n<a href=\"https://lightning-flash.readthedocs.io/en/0.5.0/api/generated/flash.core.optimizers.LARS.html\" target=\"_blank\">https://lightning-flash.readthedocs.io/en/0.5.0/api/generated/flash.core.optimizers.LARS.html</a></p>",
          "rawMarkdown": "pytorch lightning has its own implementation of LARS.\nI haven't use this yet, but this might solve the problem with large batch size.\nhttps://lightning-flash.readthedocs.io/en/0.5.0/api/generated/flash.core.optimizers.LARS.html",
          "votes": 1
        },
        {
          "id": 1786042,
          "postDate": "2022-05-12T15:15:11.027Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1783611,
      "postDate": "2022-05-10T14:20:32.363Z",
      "content": "<p>I rerun <a href=\"https://www.kaggle.com/kaerunantoka\" target=\"_blank\">@kaerunantoka</a> ' s  <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">training notebook</a> in my local machine with larger batch size, which is set to 128. In order to keep optimization iteration as the same roughly, I increase the training epoch from 23 to 200 (and lower lr a little). </p>\n<p>The iteration step of public notebook is (18452 / 5 * 4) / 16 * 23 = 17079, gaining <strong>0.71</strong> in the public leaderboard. And mine is (18452 / 5 * 4) / 128 * 200 = 18565, gaining <strong>0.68</strong> instead.</p>\n<p>What are the possible reasons behind it?</p>",
      "rawMarkdown": "I rerun @kaerunantoka ' s  [training notebook](https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0) in my local machine with larger batch size, which is set to 128. In order to keep optimization iteration as the same roughly, I increase the training epoch from 23 to 200 (and lower lr a little). \n\nThe iteration step of public notebook is (18452 / 5 * 4) / 16 * 23 = 17079, gaining **0.71** in the public leaderboard. And mine is (18452 / 5 * 4) / 128 * 200 = 18565, gaining **0.68** instead.\n\nWhat are the possible reasons behind it?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1784223,
      "author_name": "shinmura0",
      "author_url": "",
      "post_date": "2022-05-11T03:24:01.153000",
      "content": "<p>I faced a similar problem. Perhaps the cause is the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318999\" target=\"_blank\">optimal threshold</a>.</p>\n<p>The optimal threshold varies widely by epoch, iteration, and lr. Even by 0.01 up or down, the score can vary greatly.<br>\nUnfortunately, the only way to find the optimal threshold is through trial and trial.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1784857,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-11T14:42:04.623000",
      "content": "<p>The paper [1] discusses the problem of learning difficulty at large batch sizes.<br>\nThis paper focuses on the problem that existing learning algorithms do not scale to large batch sizes, and proposes a learning algorithm called LARS that adapts to large batch sizes.</p>\n<p>Another problem with some learning algorithms is that the optimal learning rate shifts when the batch size is changed [2]. For example, YOLOv5 addresses this problem by adopting a policy of setting the learning rate proportional to the batch size.</p>\n<h1>Reference</h1>\n<p>[1] LARS: <a href=\"https://arxiv.org/abs/1708.03888\" target=\"_blank\">https://arxiv.org/abs/1708.03888</a><br>\n[2] <a href=\"https://arxiv.org/abs/1706.02677\" target=\"_blank\">https://arxiv.org/abs/1706.02677</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1784877,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-05-11T14:51:13.877000",
          "content": "<p>There are other learning algorithm addressing large batch size:</p>\n<ul>\n<li>LAMB: <a href=\"https://arxiv.org/abs/1904.00962\" target=\"_blank\">https://arxiv.org/abs/1904.00962</a></li>\n<li>NVLAMB: <a href=\"https://medium.com/nvidia-ai/a-guide-to-optimizer-implementation-for-bert-at-scale-8338cc7f45fd\" target=\"_blank\">https://medium.com/nvidia-ai/a-guide-to-optimizer-implementation-for-bert-at-scale-8338cc7f45fd</a></li>\n<li>NovoGrad: <a href=\"https://arxiv.org/pdf/1905.11286.pdf\" target=\"_blank\">https://arxiv.org/pdf/1905.11286.pdf</a></li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1784883,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-05-11T15:01:11.780000",
          "content": "<p>pytorch lightning has its own implementation of LARS.<br>\nI haven't use this yet, but this might solve the problem with large batch size.<br>\n<a href=\"https://lightning-flash.readthedocs.io/en/0.5.0/api/generated/flash.core.optimizers.LARS.html\" target=\"_blank\">https://lightning-flash.readthedocs.io/en/0.5.0/api/generated/flash.core.optimizers.LARS.html</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1786042,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-12T15:15:11.027000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1784223": "I faced a similar problem. Perhaps the cause is the [optimal threshold](https://www.kaggle.com/competitions/birdclef-2022/discussion/318999).\n\nThe optimal threshold varies widely by epoch, iteration, and lr. Even by 0.01 up or down, the score can vary greatly.\nUnfortunately, the only way to find the optimal threshold is through trial and trial.",
    "1784857": "The paper [1] discusses the problem of learning difficulty at large batch sizes.\nThis paper focuses on the problem that existing learning algorithms do not scale to large batch sizes, and proposes a learning algorithm called LARS that adapts to large batch sizes.\n\nAnother problem with some learning algorithms is that the optimal learning rate shifts when the batch size is changed [2]. For example, YOLOv5 addresses this problem by adopting a policy of setting the learning rate proportional to the batch size.\n\n# Reference\n[1] LARS: https://arxiv.org/abs/1708.03888\n[2] https://arxiv.org/abs/1706.02677",
    "1783611": "I rerun @kaerunantoka ' s  [training notebook](https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0) in my local machine with larger batch size, which is set to 128. In order to keep optimization iteration as the same roughly, I increase the training epoch from 23 to 200 (and lower lr a little). \n\nThe iteration step of public notebook is (18452 / 5 * 4) / 16 * 23 = 17079, gaining **0.71** in the public leaderboard. And mine is (18452 / 5 * 4) / 128 * 200 = 18565, gaining **0.68** instead.\n\nWhat are the possible reasons behind it?"
  }
}