{
  "id": 217873,
  "title": "Batch size <= 32 (always)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217873",
  "author_name": "",
  "post_date": "2021-02-08T15:38:50.309536100Z",
  "votes": 11,
  "comment_count": 11,
  "views": 0,
  "content": "<p>This is based on Yann LeCun's paper (and tweet). In short, \"Friends dont let friends use minibatches larger than 32.\" (Yann LeCun) <a href=\"https://arxiv.org/abs/1804.07612\" target=\"_blank\">https://arxiv.org/abs/1804.07612</a></p>\n<p>I believe this is still true. Why or why not wouldn't it be always true?</p>\n<p>Twitter Thread from 2018: <a href=\"https://twitter.com/ylecun/status/989610208497360896\" target=\"_blank\">https://twitter.com/ylecun/status/989610208497360896</a></p>",
  "messages": [
    {
      "id": "1191664",
      "postDate": "02/08/2021 15:38:50",
      "content": "<p>This is based on Yann LeCun's paper (and tweet). In short, \"Friends dont let friends use minibatches larger than 32.\" (Yann LeCun) <a href=\"https://arxiv.org/abs/1804.07612\" target=\"_blank\">https://arxiv.org/abs/1804.07612</a></p>\n<p>I believe this is still true. Why or why not wouldn't it be always true?</p>\n<p>Twitter Thread from 2018: <a href=\"https://twitter.com/ylecun/status/989610208497360896\" target=\"_blank\">https://twitter.com/ylecun/status/989610208497360896</a></p>",
      "rawMarkdown": "This is based on Yann LeCun's paper (and tweet). In short, \"Friends dont let friends use minibatches larger than 32.\" (Yann LeCun) https://arxiv.org/abs/1804.07612\n\nI believe this is still true. Why or why not wouldn't it be always true?\n\nTwitter Thread from 2018: https://twitter.com/ylecun/status/989610208497360896",
      "votes": null
    },
    {
      "id": "1192339",
      "postDate": "02/09/2021 04:59:42",
      "content": "<p>Thanks for the reference. My gut feeling is that bigger batch sizes to be useful when training images are very similar to each other, which is not usually the case. Haven't really verified that though.</p>",
      "rawMarkdown": "Thanks for the reference. My gut feeling is that bigger batch sizes to be useful when training images are very similar to each other, which is not usually the case. Haven't really verified that though.",
      "votes": null
    },
    {
      "id": "1192714",
      "postDate": "02/09/2021 09:40:28",
      "content": "<p>More recent paper (probably from FastAI) actually suggest that you should go with the batch size as large as possible to fit your GPU memory at the same time also increase the learning rate to compensate the fact that you will get less updates per epoch.</p>",
      "rawMarkdown": "More recent paper (probably from FastAI) actually suggest that you should go with the batch size as large as possible to fit your GPU memory at the same time also increase the learning rate to compensate the fact that you will get less updates per epoch.",
      "votes": null
    },
    {
      "id": "1192909",
      "postDate": "02/09/2021 10:57:45",
      "content": "<p>It all depends on the objective, compute power and training schedule. </p>\n<p>Imagenet Top LB is heavily dominated by models from Google and all of them were trained with very huge batch bize</p>",
      "rawMarkdown": "It all depends on the objective, compute power and training schedule. \n\nImagenet Top LB is heavily dominated by models from Google and all of them were trained with very huge batch bize",
      "votes": null
    },
    {
      "id": "1193356",
      "postDate": "02/09/2021 15:39:40",
      "content": "<p>Hi Louis - interesting! Do you remember which paper that was? I have empirically tested different batch sizes on MNIST based on what LeCun said and had the impression that smaller batch sizes were more useful for generalisation. Would be happy to try again and document it but if FastAI did the job…why not trust them? :)</p>",
      "rawMarkdown": "Hi Louis - interesting! Do you remember which paper that was? I have empirically tested different batch sizes on MNIST based on what LeCun said and had the impression that smaller batch sizes were more useful for generalisation. Would be happy to try again and document it but if FastAI did the job...why not trust them? :)",
      "votes": null
    },
    {
      "id": "1193684",
      "postDate": "02/09/2021 19:50:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/btrentini\" target=\"_blank\">@btrentini</a> Interresting paper for sure..thanks for sharing that!<br>\nWith I little googling it is interresting to note that are multiple papers where either small or larger batch sizes are preferred. This is also one where they don't use a learning rate scheduler but achieve the same effect by batch size scheduling <a href=\"https://arxiv.org/abs/1711.00489\" target=\"_blank\">https://arxiv.org/abs/1711.00489</a></p>\n<p>From a practical point of view…the effect of the batch size on the achieved result when training a model comes down to doing some experiments with it. <br>\nAllmost all the papers always use the same datasets: MNIST, CIFAR, Imagenet. Some diversity on the datasets might have different effects on the outcome of those papers.</p>\n<p>That said… I will give different batch sizes a try. Always interresting to experiment.</p>",
      "rawMarkdown": "Hi @btrentini Interresting paper for sure..thanks for sharing that!\nWith I little googling it is interresting to note that are multiple papers where either small or larger batch sizes are preferred. This is also one where they don't use a learning rate scheduler but achieve the same effect by batch size scheduling [https://arxiv.org/abs/1711.00489](https://arxiv.org/abs/1711.00489)\n\nFrom a practical point of view...the effect of the batch size on the achieved result when training a model comes down to doing some experiments with it. \nAllmost all the papers always use the same datasets: MNIST, CIFAR, Imagenet. Some diversity on the datasets might have different effects on the outcome of those papers.\n\nThat said... I will give different batch sizes a try. Always interresting to experiment.",
      "votes": null
    },
    {
      "id": "1193740",
      "postDate": "02/09/2021 20:39:02",
      "content": "<p>One reference is from the paper: \"A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay\"<br>\nby Leslie N. Smith<br>\n<a href=\"https://arxiv.org/abs/1803.09820\" target=\"_blank\">https://arxiv.org/abs/1803.09820</a><br>\nHe is the person who is known for cyclical learning rates and super-convergence.<br>\nI guess the point of the paper is probably more about reaching the same accuracy with faster training speed. But I guess if you are talking about generalizability and the problem that is not limited by training time, then small batch size might still give you some small benefit.</p>",
      "rawMarkdown": "One reference is from the paper: \"A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay\"\nby Leslie N. Smith\n[https://arxiv.org/abs/1803.09820](https://arxiv.org/abs/1803.09820)\nHe is the person who is known for cyclical learning rates and super-convergence.\nI guess the point of the paper is probably more about reaching the same accuracy with faster training speed. But I guess if you are talking about generalizability and the problem that is not limited by training time, then small batch size might still give you some small benefit.",
      "votes": null
    },
    {
      "id": "1194066",
      "postDate": "02/10/2021 03:22:30",
      "content": "<p>Thats so cool you made that dope</p>",
      "rawMarkdown": "Thats so cool you made that dope",
      "votes": null
    },
    {
      "id": "1194617",
      "postDate": "02/10/2021 09:25:23",
      "content": "<p>I guess iit's especially helpful with training ViTs as they use batch sizes of 1024 to 4096</p>",
      "rawMarkdown": "I guess iit's especially helpful with training ViTs as they use batch sizes of 1024 to 4096",
      "votes": null
    },
    {
      "id": "1207394",
      "postDate": "02/17/2021 20:26:15",
      "content": "<p>Small batch size affects the batch norm layer's performance unless you use <a href=\"https://arxiv.org/pdf/1803.08494.pdf\" target=\"_blank\">group normalization</a> with small batch sizes.</p>",
      "rawMarkdown": "Small batch size affects the batch norm layer's performance unless you use [group normalization](https://arxiv.org/pdf/1803.08494.pdf) with small batch sizes.",
      "votes": null
    },
    {
      "id": "2750737",
      "postDate": "04/13/2024 20:19:58",
      "content": "<p>3 years later and I changed my opinion long ago! We have more powerful models and parametrization nowadays and agree it's worth probing different batch sizes prior to training </p>",
      "rawMarkdown": "3 years later and I changed my opinion long ago! We have more powerful models and parametrization nowadays and agree it's worth probing different batch sizes prior to training",
      "votes": null
    },
    {
      "id": "2750739",
      "postDate": "04/13/2024 20:20:53",
      "content": "<p>3 years later and I changed my opinion. Nowadays we have more powerful models and parametrization and it's worth probing and considering the impact of different batch sizes relating to memory footprint, throughput, noise-to-signal ratio, etc. prior to training \"at scale\"</p>",
      "rawMarkdown": "3 years later and I changed my opinion. Nowadays we have more powerful models and parametrization and it's worth probing and considering the impact of different batch sizes relating to memory footprint, throughput, noise-to-signal ratio, etc. prior to training \"at scale\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1192339,
      "author_name": "krisho007",
      "author_url": "",
      "post_date": "02/09/2021 04:59:42",
      "content": "<p>Thanks for the reference. My gut feeling is that bigger batch sizes to be useful when training images are very similar to each other, which is not usually the case. Haven't really verified that though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1192714,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "02/09/2021 09:40:28",
      "content": "<p>More recent paper (probably from FastAI) actually suggest that you should go with the batch size as large as possible to fit your GPU memory at the same time also increase the learning rate to compensate the fact that you will get less updates per epoch.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1193356,
          "author_name": "btrentini",
          "author_url": "",
          "post_date": "02/09/2021 15:39:40",
          "content": "<p>Hi Louis - interesting! Do you remember which paper that was? I have empirically tested different batch sizes on MNIST based on what LeCun said and had the impression that smaller batch sizes were more useful for generalisation. Would be happy to try again and document it but if FastAI did the job…why not trust them? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1193740,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "02/09/2021 20:39:02",
          "content": "<p>One reference is from the paper: \"A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay\"<br>\nby Leslie N. Smith<br>\n<a href=\"https://arxiv.org/abs/1803.09820\" target=\"_blank\">https://arxiv.org/abs/1803.09820</a><br>\nHe is the person who is known for cyclical learning rates and super-convergence.<br>\nI guess the point of the paper is probably more about reaching the same accuracy with faster training speed. But I guess if you are talking about generalizability and the problem that is not limited by training time, then small batch size might still give you some small benefit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1192909,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "02/09/2021 10:57:45",
      "content": "<p>It all depends on the objective, compute power and training schedule. </p>\n<p>Imagenet Top LB is heavily dominated by models from Google and all of them were trained with very huge batch bize</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1193684,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "02/09/2021 19:50:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/btrentini\" target=\"_blank\">@btrentini</a> Interresting paper for sure..thanks for sharing that!<br>\nWith I little googling it is interresting to note that are multiple papers where either small or larger batch sizes are preferred. This is also one where they don't use a learning rate scheduler but achieve the same effect by batch size scheduling <a href=\"https://arxiv.org/abs/1711.00489\" target=\"_blank\">https://arxiv.org/abs/1711.00489</a></p>\n<p>From a practical point of view…the effect of the batch size on the achieved result when training a model comes down to doing some experiments with it. <br>\nAllmost all the papers always use the same datasets: MNIST, CIFAR, Imagenet. Some diversity on the datasets might have different effects on the outcome of those papers.</p>\n<p>That said… I will give different batch sizes a try. Always interresting to experiment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2750737,
          "author_name": "btrentini",
          "author_url": "",
          "post_date": "04/13/2024 20:19:58",
          "content": "<p>3 years later and I changed my opinion long ago! We have more powerful models and parametrization nowadays and agree it's worth probing different batch sizes prior to training </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1194066,
      "author_name": "stevejabari",
      "author_url": "",
      "post_date": "02/10/2021 03:22:30",
      "content": "<p>Thats so cool you made that dope</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1194617,
      "author_name": "alexanderriedel",
      "author_url": "",
      "post_date": "02/10/2021 09:25:23",
      "content": "<p>I guess iit's especially helpful with training ViTs as they use batch sizes of 1024 to 4096</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207394,
      "author_name": "mariammohamed",
      "author_url": "",
      "post_date": "02/17/2021 20:26:15",
      "content": "<p>Small batch size affects the batch norm layer's performance unless you use <a href=\"https://arxiv.org/pdf/1803.08494.pdf\" target=\"_blank\">group normalization</a> with small batch sizes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2750739,
      "author_name": "btrentini",
      "author_url": "",
      "post_date": "04/13/2024 20:20:53",
      "content": "<p>3 years later and I changed my opinion. Nowadays we have more powerful models and parametrization and it's worth probing and considering the impact of different batch sizes relating to memory footprint, throughput, noise-to-signal ratio, etc. prior to training \"at scale\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1191664": "This is based on Yann LeCun's paper (and tweet). In short, \"Friends dont let friends use minibatches larger than 32.\" (Yann LeCun) https://arxiv.org/abs/1804.07612\n\nI believe this is still true. Why or why not wouldn't it be always true?\n\nTwitter Thread from 2018: https://twitter.com/ylecun/status/989610208497360896",
    "1192339": "Thanks for the reference. My gut feeling is that bigger batch sizes to be useful when training images are very similar to each other, which is not usually the case. Haven't really verified that though.",
    "1192714": "More recent paper (probably from FastAI) actually suggest that you should go with the batch size as large as possible to fit your GPU memory at the same time also increase the learning rate to compensate the fact that you will get less updates per epoch.",
    "1192909": "It all depends on the objective, compute power and training schedule. \n\nImagenet Top LB is heavily dominated by models from Google and all of them were trained with very huge batch bize",
    "1193356": "Hi Louis - interesting! Do you remember which paper that was? I have empirically tested different batch sizes on MNIST based on what LeCun said and had the impression that smaller batch sizes were more useful for generalisation. Would be happy to try again and document it but if FastAI did the job...why not trust them? :)",
    "1193684": "Hi @btrentini Interresting paper for sure..thanks for sharing that!\nWith I little googling it is interresting to note that are multiple papers where either small or larger batch sizes are preferred. This is also one where they don't use a learning rate scheduler but achieve the same effect by batch size scheduling [https://arxiv.org/abs/1711.00489](https://arxiv.org/abs/1711.00489)\n\nFrom a practical point of view...the effect of the batch size on the achieved result when training a model comes down to doing some experiments with it. \nAllmost all the papers always use the same datasets: MNIST, CIFAR, Imagenet. Some diversity on the datasets might have different effects on the outcome of those papers.\n\nThat said... I will give different batch sizes a try. Always interresting to experiment.",
    "1193740": "One reference is from the paper: \"A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay\"\nby Leslie N. Smith\n[https://arxiv.org/abs/1803.09820](https://arxiv.org/abs/1803.09820)\nHe is the person who is known for cyclical learning rates and super-convergence.\nI guess the point of the paper is probably more about reaching the same accuracy with faster training speed. But I guess if you are talking about generalizability and the problem that is not limited by training time, then small batch size might still give you some small benefit.",
    "1194066": "Thats so cool you made that dope",
    "1194617": "I guess iit's especially helpful with training ViTs as they use batch sizes of 1024 to 4096",
    "1207394": "Small batch size affects the batch norm layer's performance unless you use [group normalization](https://arxiv.org/pdf/1803.08494.pdf) with small batch sizes.",
    "2750737": "3 years later and I changed my opinion long ago! We have more powerful models and parametrization nowadays and agree it's worth probing different batch sizes prior to training",
    "2750739": "3 years later and I changed my opinion. Nowadays we have more powerful models and parametrization and it's worth probing and considering the impact of different batch sizes relating to memory footprint, throughput, noise-to-signal ratio, etc. prior to training \"at scale\""
  },
  "source": "meta"
}