{
  "id": 40780,
  "title": "batch size trick to accelerate training?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40780",
  "author_name": "",
  "post_date": "2017-10-08T05:44:51.785315600Z",
  "votes": 16,
  "comment_count": 26,
  "views": 0,
  "content": "<p>i cannot explain this and i hope this is not a bug. here are my experiment results. other kaggler may want to confirm if this is correct or not.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228871/7535/batch_size1.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>one paper that might be relevant could be:</p>\n\n<p>(1) \"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour\" - Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, Kaiming He</p>",
  "messages": [
    {
      "id": "228871",
      "postDate": "10/08/2017 05:44:51",
      "content": "<p>i cannot explain this and i hope this is not a bug. here are my experiment results. other kaggler may want to confirm if this is correct or not.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228871/7535/batch_size1.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>one paper that might be relevant could be:</p>\n\n<p>(1) \"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour\" - Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, Kaiming He</p>",
      "rawMarkdown": "i cannot explain this and i hope this is not a bug. here are my experiment results. other kaggler may want to confirm if this is correct or not.\n\n ![enter image description here][1]\n\n\none paper that might be relevant could be:\n\n(1) \"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour\" - Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, Kaiming He\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228871/7535/batch_size1.png",
      "votes": null
    },
    {
      "id": "228900",
      "postDate": "10/08/2017 08:34:06",
      "content": "<p>Interesting, I'll have to give this a try. </p>\n\n<p>Remember that when using mini-batches, at each step SGD is really optimizing an approximation of the true loss function, and this approximation is different for each mini-batch. The true loss function is described by <em>all</em> of the training data and our mini-batches are not representative samples of all of the training data. The smaller the batch, the worse the approximation is.</p>\n\n<p>However... the fact that SGD is noisy (because of the approximations) tends to help with getting out of nasty spots on the loss surface like saddlepoints. Changing the batch size might be enough to make SGD jump out of such a saddlepoint where it has become stuck.</p>\n\n<p>I'm not sure how to interpret your chart, though. What does it mean when you say <code>batch == 4x32 to 8x32</code>?</p>",
      "rawMarkdown": "Interesting, I'll have to give this a try. \n\nRemember that when using mini-batches, at each step SGD is really optimizing an approximation of the true loss function, and this approximation is different for each mini-batch. The true loss function is described by _all_ of the training data and our mini-batches are not representative samples of all of the training data. The smaller the batch, the worse the approximation is.\n\nHowever... the fact that SGD is noisy (because of the approximations) tends to help with getting out of nasty spots on the loss surface like saddlepoints. Changing the batch size might be enough to make SGD jump out of such a saddlepoint where it has become stuck.\n\nI'm not sure how to interpret your chart, though. What does it mean when you say `batch == 4x32 to 8x32`?",
      "votes": null
    },
    {
      "id": "228903",
      "postDate": "10/08/2017 08:38:50",
      "content": "<p>\"batch == 4x32 to 8x32?\"</p>\n\n<p>Initially, i am using  4 gradient accumulation of batch=32 each.\nAfter using  8 gradient accumulation of batch=32 each, the \"jump\" occur</p>\n\n<pre><code>    for images, labels, indices in train_loader:\n\n        ... ###  len(images) is 32  here ### ....\n\n        # one iteration update  -------------\n        images  = Variable(images).cuda()\n        labels  = Variable(labels).cuda()\n        logits = net(images)\n        probs  = F.softmax(logits)\n\n        loss = F.cross_entropy(logits, labels)\n        acc  = top_accuracy(probs, labels, top_k=(1,))\n\n        # accumulate gradients\n        loss.backward()\n        if j%iter_accum == 0:  ### iter_accum is changed from 4 to 8 here ### \n            optimizer.step()\n            optimizer.zero_grad()\n\n       ...\n</code></pre>",
      "rawMarkdown": "\"batch == 4x32 to 8x32?\"\n\nInitially, i am using  4 gradient accumulation of batch=32 each.\nAfter using  8 gradient accumulation of batch=32 each, the \"jump\" occur\n\n\n    \n  \n        for images, labels, indices in train_loader:\n  \n            ... ###  len(images) is 32  here ### ....\n    \n            # one iteration update  -------------\n            images  = Variable(images).cuda()\n            labels  = Variable(labels).cuda()\n            logits = net(images)\n            probs  = F.softmax(logits)\n\n            loss = F.cross_entropy(logits, labels)\n            acc  = top_accuracy(probs, labels, top_k=(1,))\n\n            # accumulate gradients\n            loss.backward()\n            if j%iter_accum == 0:  ### iter_accum is changed from 4 to 8 here ### \n                optimizer.step()\n                optimizer.zero_grad()\n\n           ...",
      "votes": null
    },
    {
      "id": "228907",
      "postDate": "10/08/2017 08:46:25",
      "content": "<p>i am thinking the reason could be similar to equation 3 and 4 of the paper. Because the update is \"small\" per batch (due to class imbalance, noise, under fitting, etc), it \"doesn't matter\" if you update \"one by one\" or update \"all at once\". But if you update \"all at once\", the  backward propagation signals are \"boosted\" and signal to noise ratio improves.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228907/7541/Selection_074.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "i am thinking the reason could be similar to equation 3 and 4 of the paper. Because the update is \"small\" per batch (due to class imbalance, noise, under fitting, etc), it \"doesn't matter\" if you update \"one by one\" or update \"all at once\". But if you update \"all at once\", the  backward propagation signals are \"boosted\" and signal to noise ratio improves.\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228907/7541/Selection_074.png",
      "votes": null
    },
    {
      "id": "228908",
      "postDate": "10/08/2017 08:47:18",
      "content": "<p>Ah OK. I'm using Keras and I think it just computes the gradient over each mini-batch (i.e. no accumulation).</p>\n\n<p>So the validation accuracy jumps up when you compute the gradient over a larger number of examples (8x32 samples instead of 4x32 samples). This would make the SGD slightly less stochastic, not more (i.e. it somewhat reduces the noise). After all, the larger the number of examples, the more the gradients approximate the gradients of the true loss function.</p>",
      "rawMarkdown": "Ah OK. I'm using Keras and I think it just computes the gradient over each mini-batch (i.e. no accumulation).\n\nSo the validation accuracy jumps up when you compute the gradient over a larger number of examples (8x32 samples instead of 4x32 samples). This would make the SGD slightly less stochastic, not more (i.e. it somewhat reduces the noise). After all, the larger the number of examples, the more the gradients approximate the gradients of the true loss function.",
      "votes": null
    },
    {
      "id": "228910",
      "postDate": "10/08/2017 08:53:58",
      "content": "<p>facebook paper results for  minibatch size for imagenet 5K (training for 5K classes)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228910/7542/Selection_075.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "facebook paper results for  minibatch size for imagenet 5K (training for 5K classes)\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228910/7542/Selection_075.png",
      "votes": null
    },
    {
      "id": "228914",
      "postDate": "10/08/2017 08:58:20",
      "content": "<p>\" After all, the larger the number of examples, the more the gradients approximate the gradients of the true loss function.\"</p>\n\n<p>i am thinking something similar, but i always though that batch size of 256 to 512 is always enough. Now I am doing experiment like this:</p>\n\n<pre><code>while in training loop:\n\n         do back propagation\n\n         check if loss plateau or not \n         if yes:\n                increase effective batch size by increasing iter_accum\n</code></pre>",
      "rawMarkdown": "\" After all, the larger the number of examples, the more the gradients approximate the gradients of the true loss function.\"\n\ni am thinking something similar, but i always though that batch size of 256 to 512 is always enough. Now I am doing experiment like this:\n\n    while in training loop:\n\n             do back propagation\n            \n             check if loss plateau or not \n             if yes:\n                    increase effective batch size by increasing iter_accum",
      "votes": null
    },
    {
      "id": "228923",
      "postDate": "10/08/2017 09:30:53",
      "content": "<p>I see that you write \"changing learning rate does not seem to help\", but still: If you are keeping the learning the same while increasing the batch size, it can have a similar effect as decreasing the learning rate, because effective per-image learning rate becomes smaller - so I wonder if the jumps can be explained by this?</p>",
      "rawMarkdown": "I see that you write \"changing learning rate does not seem to help\", but still: If you are keeping the learning the same while increasing the batch size, it can have a similar effect as decreasing the learning rate, because effective per-image learning rate becomes smaller - so I wonder if the jumps can be explained by this?",
      "votes": null
    },
    {
      "id": "228924",
      "postDate": "10/08/2017 09:32:41",
      "content": "<p>\" it can have a similar effect as decreasing the learning rate\"</p>\n\n<p>This is correct. But my observation is that \"fixing the learning rate but increasing batch size\" seems to be faster and gives better results.</p>",
      "rawMarkdown": "\" it can have a similar effect as decreasing the learning rate\"\n\nThis is correct. But my observation is that \"fixing the learning rate but increasing batch size\" seems to be faster and gives better results.",
      "votes": null
    },
    {
      "id": "228925",
      "postDate": "10/08/2017 09:34:52",
      "content": "<p>Thanks! Will check if it helps me too :)</p>",
      "rawMarkdown": "Thanks! Will check if it helps me too :)",
      "votes": null
    },
    {
      "id": "229071",
      "postDate": "10/08/2017 17:01:39",
      "content": "<p>convergence curve for imagnet-1K.<a href=\"https://github.com/apache/incubator-mxnet/tree/master/example/image-classification\">https://github.com/apache/incubator-mxnet/tree/master/example/image-classification</a></p>\n\n<p>there is another one at: <a href=\"http://dmlc.ml/mxnet/2015/10/27/training-deep-net-on-14-million-images.html\">http://dmlc.ml/mxnet/2015/10/27/training-deep-net-on-14-million-images.html</a></p>\n\n<p><img src=\"https://raw.githubusercontent.com/dmlc/web-data/master/mxnet/image/dist_converge.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://github.com/tornadomeet/ResNet/blob/master/log/training-curve.png?raw=true\" alt=\"enter image description here\" title=\"\">\n  <a href=\"https://github.com/tornadomeet/ResNet/tree/master/log\">https://github.com/tornadomeet/ResNet/tree/master/log</a></p>",
      "rawMarkdown": "convergence curve for imagnet-1K.https://github.com/apache/incubator-mxnet/tree/master/example/image-classification\n\nthere is another one at: http://dmlc.ml/mxnet/2015/10/27/training-deep-net-on-14-million-images.html\n\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n  https://github.com/tornadomeet/ResNet/tree/master/log\n\n  [1]: https://raw.githubusercontent.com/dmlc/web-data/master/mxnet/image/dist_converge.png\n  [2]: https://github.com/tornadomeet/ResNet/blob/master/log/training-curve.png?raw=true",
      "votes": null
    },
    {
      "id": "229129",
      "postDate": "10/08/2017 22:15:26",
      "content": "<p>Too bad that there is no large scale experimental results in this paper.</p>\n\n<p><a href=\"https://arxiv.org/pdf/1610.05792.pdf\">https://arxiv.org/pdf/1610.05792.pdf</a></p>\n\n<p>\"We analyzed and studied the behavior of alternative SGD methods in which the batch size increases over time.\nUnlike classical SGD methods, in which stochastic gradients quickly become swamped with noise, these “big\nbatch” methods maintain a nearly constant signal to noise ratio of the approximate gradient. As a result,\nbig batch methods are able to adaptively adjust batch sizes without user oversight. The proposed automated\nmethods are shown to be empirically comparable or better performing than other standard methods, but\nwithout requiring an expert user to choose learning rates and decay parameters.\"</p>",
      "rawMarkdown": "Too bad that there is no large scale experimental results in this paper.\n\nhttps://arxiv.org/pdf/1610.05792.pdf\n\n\"We analyzed and studied the behavior of alternative SGD methods in which the batch size increases over time.\nUnlike classical SGD methods, in which stochastic gradients quickly become swamped with noise, these “big\nbatch” methods maintain a nearly constant signal to noise ratio of the approximate gradient. As a result,\nbig batch methods are able to adaptively adjust batch sizes without user oversight. The proposed automated\nmethods are shown to be empirically comparable or better performing than other standard methods, but\nwithout requiring an expert user to choose learning rates and decay parameters.\"",
      "votes": null
    },
    {
      "id": "229141",
      "postDate": "10/09/2017 00:02:50",
      "content": "<p>This is a great idea. Reminiscent of simulated annealing.</p>",
      "rawMarkdown": "This is a great idea. Reminiscent of simulated annealing.",
      "votes": null
    },
    {
      "id": "229367",
      "postDate": "10/09/2017 13:15:26",
      "content": "<p>my 224 se-resnet50 has been completed. Here are the single crop results:</p>\n\n<ul>\n<li><p>se-resnet50 train=224x224 crops from 256 (perturb scale, shift,rotate):</p>\n\n<p>-- test = resize to 224 gives LB 0.66708  </p></li>\n</ul>\n\n<p>.</p>\n\n<ul>\n<li><p>se-resnet50 train=160x160 crops from 180:</p>\n\n<p>-- 160x160 single center crop gives LB 0.63358</p>\n\n<p>-- 180x180 gives LB 0.63839 </p>\n\n<p>-- 180x180 resize to 160x160 gives LB  0.64356</p></li>\n</ul>\n\n<p>.</p>\n\n<p>train method for 224x224  se-resnet50 train:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/229367/7552/batch_size2.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<ul>\n<li><p>use increasing batch size (hand adjusted) method at rate =0.01 to train until validation accuracy hits 0.64. This takes about 2 epoch.</p></li>\n<li><p>then train at rate=0.001 for another 0.5 epoch till validation accuracy hits 0.65.</p></li>\n</ul>\n\n<p>total training takes about 2 days. long iterations at the end are important for good results. (you can probably get better results if you run longer for lr =0.001, and even lr= 0.0001 later). you can acclerate early training.</p>\n\n<p>train software is at: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>",
      "rawMarkdown": "my 224 se-resnet50 has been completed. Here are the single crop results:\n\n-  se-resnet50 train=224x224 crops from 256 (perturb scale, shift,rotate):\n\n     -- test = resize to 224 gives LB 0.66708  \n\n.\n\n-  se-resnet50 train=160x160 crops from 180:\n\n   -- 160x160 single center crop gives LB 0.63358\n\n   -- 180x180 gives LB 0.63839 \n\n   -- 180x180 resize to 160x160 gives LB  0.64356\n\n\n.\n\ntrain method for 224x224  se-resnet50 train:\n\n\n  ![enter image description here][1]\n\n -  use increasing batch size (hand adjusted) method at rate =0.01 to train until validation accuracy hits 0.64. This takes about 2 epoch.\n\n -  then train at rate=0.001 for another 0.5 epoch till validation accuracy hits 0.65.\n\n\ntotal training takes about 2 days. long iterations at the end are important for good results. (you can probably get better results if you run longer for lr =0.001, and even lr= 0.0001 later). you can acclerate early training.\n\ntrain software is at: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/229367/7552/batch_size2.png",
      "votes": null
    },
    {
      "id": "229491",
      "postDate": "10/09/2017 18:34:39",
      "content": "<p>If you are using Adam optimizer, don't gradients get accumulated ?</p>",
      "rawMarkdown": "If you are using Adam optimizer, don't gradients get accumulated ?",
      "votes": null
    },
    {
      "id": "229841",
      "postDate": "10/10/2017 16:22:56",
      "content": "<p>one side note: if you accumulate gradient for large batch size and you are using sdg+momentum, you might want to try experiments with different momentum. Sometimes, you need to lower them. ( i don't know why but my experiment shows sometimes it work better.  maybe large batch size don't need momentum?) you may also want to read this: <a href=\"http://stanford.edu/~imit/tuneyourmomentum/\">http://stanford.edu/~imit/tuneyourmomentum/</a></p>",
      "rawMarkdown": "one side note: if you accumulate gradient for large batch size and you are using sdg+momentum, you might want to try experiments with different momentum. Sometimes, you need to lower them. ( i don't know why but my experiment shows sometimes it work better.  maybe large batch size don't need momentum?) you may also want to read this: http://stanford.edu/~imit/tuneyourmomentum/",
      "votes": null
    },
    {
      "id": "230003",
      "postDate": "10/11/2017 02:10:39",
      "content": "<p>If I'm not wrong, in a similar article by Facebook, they suggest multiplying the learning rate by k. Being k the number of gradients accumulation steps. Doing so will keep the same \"effective\" learning rate</p>",
      "rawMarkdown": "If I'm not wrong, in a similar article by Facebook, they suggest multiplying the learning rate by k. Being k the number of gradients accumulation steps. Doing so will keep the same \"effective\" learning rate",
      "votes": null
    },
    {
      "id": "230631",
      "postDate": "10/12/2017 11:11:12",
      "content": "<p>How Facebook trained ImageNet in 1 hour?</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=eSmst0CUJAo\">https://www.youtube.com/watch?v=eSmst0CUJAo</a></p>\n\n<p><a href=\"https://caffe2.ai/docs/SynchronousSGD.html\">https://caffe2.ai/docs/SynchronousSGD.html</a></p>",
      "rawMarkdown": "How Facebook trained ImageNet in 1 hour?\n\nhttps://www.youtube.com/watch?v=eSmst0CUJAo\n\nhttps://caffe2.ai/docs/SynchronousSGD.html",
      "votes": null
    },
    {
      "id": "234456",
      "postDate": "10/23/2017 11:12:28",
      "content": "<p>Have you compared accumulate gradients with no accumulate gradients on one gpu? I tried accumulate gradients according to your code, but it seemed to perform worse(batch size is 4x32).</p>",
      "rawMarkdown": "Have you compared accumulate gradients with no accumulate gradients on one gpu? I tried accumulate gradients according to your code, but it seemed to perform worse(batch size is 4x32).",
      "votes": null
    },
    {
      "id": "236702",
      "postDate": "10/27/2017 22:09:15",
      "content": "<p><a href=\"https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf\">https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf</a></p>",
      "rawMarkdown": "https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf",
      "votes": null
    },
    {
      "id": "236708",
      "postDate": "10/27/2017 22:33:18",
      "content": "<p>Adding some \"theory\" to Human Analog's statement to why big batches might be dangerous.\n<a href=\"https://arxiv.org/pdf/1609.04836.pdf\">https://arxiv.org/pdf/1609.04836.pdf</a></p>",
      "rawMarkdown": "Adding some \"theory\" to Human Analog's statement to why big batches might be dangerous.\nhttps://arxiv.org/pdf/1609.04836.pdf",
      "votes": null
    },
    {
      "id": "236744",
      "postDate": "10/28/2017 01:15:42",
      "content": "<p>Thanks. I am aware of the paper. Large batch works but have to be careful. if your loss is already decreasing, probably you don't have to use. If the loss is decreasing slowly, you may may to try large batch. Also, the below mention about experiments on alexnet and some check about gradient and gradient update values are necessary.</p>\n\n<p><a href=\"https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf\">https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf</a></p>\n\n<p>Also note that ICLR 2017 paper does not apply rate multiplication as facebook paper,</p>",
      "rawMarkdown": "Thanks. I am aware of the paper. Large batch works but have to be careful. if your loss is already decreasing, probably you don't have to use. If the loss is decreasing slowly, you may may to try large batch. Also, the below mention about experiments on alexnet and some check about gradient and gradient update values are necessary.\n\nhttps://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf\n\nAlso note that ICLR 2017 paper does not apply rate multiplication as facebook paper,",
      "votes": null
    },
    {
      "id": "237500",
      "postDate": "10/30/2017 13:13:20",
      "content": "<p>A new paper about this topic is currently under review at ICLR2018: <a href=\"https://openreview.net/pdf?id=B1Yy1BxCZ\">https://openreview.net/pdf?id=B1Yy1BxCZ</a></p>\n\n<blockquote>\n  <p>It is common practice to decay the learning rate. Here we show one can usually\n  obtain the same learning curve on both training and test sets by instead increasing\n  the batch size during training. </p>\n</blockquote>",
      "rawMarkdown": "A new paper about this topic is currently under review at ICLR2018: https://openreview.net/pdf?id=B1Yy1BxCZ\n\n&gt; It is common practice to decay the learning rate. Here we show one can usually\nobtain the same learning curve on both training and test sets by instead increasing\nthe batch size during training.",
      "votes": null
    },
    {
      "id": "237515",
      "postDate": "10/30/2017 13:39:37",
      "content": "<p>oh! too late for me to write a paper.</p>\n\n<p>but good for me to know the results and apply to this kaggle competition</p>",
      "rawMarkdown": "oh! too late for me to write a paper.\n\nbut good for me to know the results and apply to this kaggle competition",
      "votes": null
    },
    {
      "id": "237607",
      "postDate": "10/30/2017 16:38:47",
      "content": "<p>The main point arising from our results is that, in contrast to previous conception, there is no inherent\ngeneralization problem with training using big mini batches. That is, model training using big\nmini-batches can generalize as good as models trained using small mini-batches. Previous works\nfound that big batch updates reach a sharp minima - which was claimed to provide poor generalization.\nIn this work we show that this is not necessarily the case: provided the needed adjustments, good\ngeneralization can be observed for big batch regimes</p>\n\n<p><a href=\"https://arxiv.org/pdf/1705.08741.pdf\">https://arxiv.org/pdf/1705.08741.pdf</a>\n\"Train longer, generalize better: closing the generalization gap in large batch training of neural networks\"</p>",
      "rawMarkdown": "The main point arising from our results is that, in contrast to previous conception, there is no inherent\ngeneralization problem with training using big mini batches. That is, model training using big\nmini-batches can generalize as good as models trained using small mini-batches. Previous works\nfound that big batch updates reach a sharp minima - which was claimed to provide poor generalization.\nIn this work we show that this is not necessarily the case: provided the needed adjustments, good\ngeneralization can be observed for big batch regimes\n\nhttps://arxiv.org/pdf/1705.08741.pdf\n\"Train longer, generalize better: closing the generalization gap in large batch training of neural networks\"",
      "votes": null
    },
    {
      "id": "239696",
      "postDate": "11/04/2017 07:29:11",
      "content": "<p>DON’T DECAY THE LEARNING RATE,\nINCREASE THE BATCH SIZE</p>\n\n<p><a href=\"https://arxiv.org/pdf/1711.00489.pdf\">https://arxiv.org/pdf/1711.00489.pdf</a></p>\n\n<p>New Google paper on impact of large batch traning</p>",
      "rawMarkdown": "DON’T DECAY THE LEARNING RATE,\nINCREASE THE BATCH SIZE\n\nhttps://arxiv.org/pdf/1711.00489.pdf\n\nNew Google paper on impact of large batch traning",
      "votes": null
    },
    {
      "id": "427598",
      "postDate": "11/25/2018 20:03:37",
      "content": "<p>With his code... If my machine can run 256 batch and I accumulate for 512 batch for example it runs out of memory :(</p>",
      "rawMarkdown": "With his code... If my machine can run 256 batch and I accumulate for 512 batch for example it runs out of memory :(",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 228900,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "10/08/2017 08:34:06",
      "content": "<p>Interesting, I'll have to give this a try. </p>\n\n<p>Remember that when using mini-batches, at each step SGD is really optimizing an approximation of the true loss function, and this approximation is different for each mini-batch. The true loss function is described by <em>all</em> of the training data and our mini-batches are not representative samples of all of the training data. The smaller the batch, the worse the approximation is.</p>\n\n<p>However... the fact that SGD is noisy (because of the approximations) tends to help with getting out of nasty spots on the loss surface like saddlepoints. Changing the batch size might be enough to make SGD jump out of such a saddlepoint where it has become stuck.</p>\n\n<p>I'm not sure how to interpret your chart, though. What does it mean when you say <code>batch == 4x32 to 8x32</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 228903,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/08/2017 08:38:50",
          "content": "<p>\"batch == 4x32 to 8x32?\"</p>\n\n<p>Initially, i am using  4 gradient accumulation of batch=32 each.\nAfter using  8 gradient accumulation of batch=32 each, the \"jump\" occur</p>\n\n<pre><code>    for images, labels, indices in train_loader:\n\n        ... ###  len(images) is 32  here ### ....\n\n        # one iteration update  -------------\n        images  = Variable(images).cuda()\n        labels  = Variable(labels).cuda()\n        logits = net(images)\n        probs  = F.softmax(logits)\n\n        loss = F.cross_entropy(logits, labels)\n        acc  = top_accuracy(probs, labels, top_k=(1,))\n\n        # accumulate gradients\n        loss.backward()\n        if j%iter_accum == 0:  ### iter_accum is changed from 4 to 8 here ### \n            optimizer.step()\n            optimizer.zero_grad()\n\n       ...\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228908,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/08/2017 08:47:18",
          "content": "<p>Ah OK. I'm using Keras and I think it just computes the gradient over each mini-batch (i.e. no accumulation).</p>\n\n<p>So the validation accuracy jumps up when you compute the gradient over a larger number of examples (8x32 samples instead of 4x32 samples). This would make the SGD slightly less stochastic, not more (i.e. it somewhat reduces the noise). After all, the larger the number of examples, the more the gradients approximate the gradients of the true loss function.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228914,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/08/2017 08:58:20",
          "content": "<p>\" After all, the larger the number of examples, the more the gradients approximate the gradients of the true loss function.\"</p>\n\n<p>i am thinking something similar, but i always though that batch size of 256 to 512 is always enough. Now I am doing experiment like this:</p>\n\n<pre><code>while in training loop:\n\n         do back propagation\n\n         check if loss plateau or not \n         if yes:\n                increase effective batch size by increasing iter_accum\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 229491,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "10/09/2017 18:34:39",
          "content": "<p>If you are using Adam optimizer, don't gradients get accumulated ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 236708,
          "author_name": "mihaskalic",
          "author_url": "",
          "post_date": "10/27/2017 22:33:18",
          "content": "<p>Adding some \"theory\" to Human Analog's statement to why big batches might be dangerous.\n<a href=\"https://arxiv.org/pdf/1609.04836.pdf\">https://arxiv.org/pdf/1609.04836.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 236744,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/28/2017 01:15:42",
          "content": "<p>Thanks. I am aware of the paper. Large batch works but have to be careful. if your loss is already decreasing, probably you don't have to use. If the loss is decreasing slowly, you may may to try large batch. Also, the below mention about experiments on alexnet and some check about gradient and gradient update values are necessary.</p>\n\n<p><a href=\"https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf\">https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf</a></p>\n\n<p>Also note that ICLR 2017 paper does not apply rate multiplication as facebook paper,</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237607,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/30/2017 16:38:47",
          "content": "<p>The main point arising from our results is that, in contrast to previous conception, there is no inherent\ngeneralization problem with training using big mini batches. That is, model training using big\nmini-batches can generalize as good as models trained using small mini-batches. Previous works\nfound that big batch updates reach a sharp minima - which was claimed to provide poor generalization.\nIn this work we show that this is not necessarily the case: provided the needed adjustments, good\ngeneralization can be observed for big batch regimes</p>\n\n<p><a href=\"https://arxiv.org/pdf/1705.08741.pdf\">https://arxiv.org/pdf/1705.08741.pdf</a>\n\"Train longer, generalize better: closing the generalization gap in large batch training of neural networks\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 228907,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/08/2017 08:46:25",
      "content": "<p>i am thinking the reason could be similar to equation 3 and 4 of the paper. Because the update is \"small\" per batch (due to class imbalance, noise, under fitting, etc), it \"doesn't matter\" if you update \"one by one\" or update \"all at once\". But if you update \"all at once\", the  backward propagation signals are \"boosted\" and signal to noise ratio improves.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228907/7541/Selection_074.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 228910,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/08/2017 08:53:58",
      "content": "<p>facebook paper results for  minibatch size for imagenet 5K (training for 5K classes)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228910/7542/Selection_075.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 228923,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "10/08/2017 09:30:53",
      "content": "<p>I see that you write \"changing learning rate does not seem to help\", but still: If you are keeping the learning the same while increasing the batch size, it can have a similar effect as decreasing the learning rate, because effective per-image learning rate becomes smaller - so I wonder if the jumps can be explained by this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 228924,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/08/2017 09:32:41",
          "content": "<p>\" it can have a similar effect as decreasing the learning rate\"</p>\n\n<p>This is correct. But my observation is that \"fixing the learning rate but increasing batch size\" seems to be faster and gives better results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228925,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "10/08/2017 09:34:52",
          "content": "<p>Thanks! Will check if it helps me too :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230003,
          "author_name": "felipesens",
          "author_url": "",
          "post_date": "10/11/2017 02:10:39",
          "content": "<p>If I'm not wrong, in a similar article by Facebook, they suggest multiplying the learning rate by k. Being k the number of gradients accumulation steps. Doing so will keep the same \"effective\" learning rate</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 229071,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/08/2017 17:01:39",
      "content": "<p>convergence curve for imagnet-1K.<a href=\"https://github.com/apache/incubator-mxnet/tree/master/example/image-classification\">https://github.com/apache/incubator-mxnet/tree/master/example/image-classification</a></p>\n\n<p>there is another one at: <a href=\"http://dmlc.ml/mxnet/2015/10/27/training-deep-net-on-14-million-images.html\">http://dmlc.ml/mxnet/2015/10/27/training-deep-net-on-14-million-images.html</a></p>\n\n<p><img src=\"https://raw.githubusercontent.com/dmlc/web-data/master/mxnet/image/dist_converge.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://github.com/tornadomeet/ResNet/blob/master/log/training-curve.png?raw=true\" alt=\"enter image description here\" title=\"\">\n  <a href=\"https://github.com/tornadomeet/ResNet/tree/master/log\">https://github.com/tornadomeet/ResNet/tree/master/log</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 229129,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/08/2017 22:15:26",
      "content": "<p>Too bad that there is no large scale experimental results in this paper.</p>\n\n<p><a href=\"https://arxiv.org/pdf/1610.05792.pdf\">https://arxiv.org/pdf/1610.05792.pdf</a></p>\n\n<p>\"We analyzed and studied the behavior of alternative SGD methods in which the batch size increases over time.\nUnlike classical SGD methods, in which stochastic gradients quickly become swamped with noise, these “big\nbatch” methods maintain a nearly constant signal to noise ratio of the approximate gradient. As a result,\nbig batch methods are able to adaptively adjust batch sizes without user oversight. The proposed automated\nmethods are shown to be empirically comparable or better performing than other standard methods, but\nwithout requiring an expert user to choose learning rates and decay parameters.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 229141,
          "author_name": "alexcoventry",
          "author_url": "",
          "post_date": "10/09/2017 00:02:50",
          "content": "<p>This is a great idea. Reminiscent of simulated annealing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 229367,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/09/2017 13:15:26",
      "content": "<p>my 224 se-resnet50 has been completed. Here are the single crop results:</p>\n\n<ul>\n<li><p>se-resnet50 train=224x224 crops from 256 (perturb scale, shift,rotate):</p>\n\n<p>-- test = resize to 224 gives LB 0.66708  </p></li>\n</ul>\n\n<p>.</p>\n\n<ul>\n<li><p>se-resnet50 train=160x160 crops from 180:</p>\n\n<p>-- 160x160 single center crop gives LB 0.63358</p>\n\n<p>-- 180x180 gives LB 0.63839 </p>\n\n<p>-- 180x180 resize to 160x160 gives LB  0.64356</p></li>\n</ul>\n\n<p>.</p>\n\n<p>train method for 224x224  se-resnet50 train:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/229367/7552/batch_size2.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<ul>\n<li><p>use increasing batch size (hand adjusted) method at rate =0.01 to train until validation accuracy hits 0.64. This takes about 2 epoch.</p></li>\n<li><p>then train at rate=0.001 for another 0.5 epoch till validation accuracy hits 0.65.</p></li>\n</ul>\n\n<p>total training takes about 2 days. long iterations at the end are important for good results. (you can probably get better results if you run longer for lr =0.001, and even lr= 0.0001 later). you can acclerate early training.</p>\n\n<p>train software is at: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 229841,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/10/2017 16:22:56",
      "content": "<p>one side note: if you accumulate gradient for large batch size and you are using sdg+momentum, you might want to try experiments with different momentum. Sometimes, you need to lower them. ( i don't know why but my experiment shows sometimes it work better.  maybe large batch size don't need momentum?) you may also want to read this: <a href=\"http://stanford.edu/~imit/tuneyourmomentum/\">http://stanford.edu/~imit/tuneyourmomentum/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230631,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/12/2017 11:11:12",
      "content": "<p>How Facebook trained ImageNet in 1 hour?</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=eSmst0CUJAo\">https://www.youtube.com/watch?v=eSmst0CUJAo</a></p>\n\n<p><a href=\"https://caffe2.ai/docs/SynchronousSGD.html\">https://caffe2.ai/docs/SynchronousSGD.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 234456,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "10/23/2017 11:12:28",
      "content": "<p>Have you compared accumulate gradients with no accumulate gradients on one gpu? I tried accumulate gradients according to your code, but it seemed to perform worse(batch size is 4x32).</p>",
      "votes": null,
      "replies": [
        {
          "id": 427598,
          "author_name": "maparla",
          "author_url": "",
          "post_date": "11/25/2018 20:03:37",
          "content": "<p>With his code... If my machine can run 256 batch and I accumulate for 512 batch for example it runs out of memory :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 236702,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/27/2017 22:09:15",
      "content": "<p><a href=\"https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf\">https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 237500,
      "author_name": "dudemeister",
      "author_url": "",
      "post_date": "10/30/2017 13:13:20",
      "content": "<p>A new paper about this topic is currently under review at ICLR2018: <a href=\"https://openreview.net/pdf?id=B1Yy1BxCZ\">https://openreview.net/pdf?id=B1Yy1BxCZ</a></p>\n\n<blockquote>\n  <p>It is common practice to decay the learning rate. Here we show one can usually\n  obtain the same learning curve on both training and test sets by instead increasing\n  the batch size during training. </p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 237515,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/30/2017 13:39:37",
          "content": "<p>oh! too late for me to write a paper.</p>\n\n<p>but good for me to know the results and apply to this kaggle competition</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 239696,
      "author_name": "vinhnguyen",
      "author_url": "",
      "post_date": "11/04/2017 07:29:11",
      "content": "<p>DON’T DECAY THE LEARNING RATE,\nINCREASE THE BATCH SIZE</p>\n\n<p><a href=\"https://arxiv.org/pdf/1711.00489.pdf\">https://arxiv.org/pdf/1711.00489.pdf</a></p>\n\n<p>New Google paper on impact of large batch traning</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "228871": "i cannot explain this and i hope this is not a bug. here are my experiment results. other kaggler may want to confirm if this is correct or not.\n\n ![enter image description here][1]\n\n\none paper that might be relevant could be:\n\n(1) \"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour\" - Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, Kaiming He\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228871/7535/batch_size1.png",
    "228900": "Interesting, I'll have to give this a try. \n\nRemember that when using mini-batches, at each step SGD is really optimizing an approximation of the true loss function, and this approximation is different for each mini-batch. The true loss function is described by _all_ of the training data and our mini-batches are not representative samples of all of the training data. The smaller the batch, the worse the approximation is.\n\nHowever... the fact that SGD is noisy (because of the approximations) tends to help with getting out of nasty spots on the loss surface like saddlepoints. Changing the batch size might be enough to make SGD jump out of such a saddlepoint where it has become stuck.\n\nI'm not sure how to interpret your chart, though. What does it mean when you say `batch == 4x32 to 8x32`?",
    "228903": "\"batch == 4x32 to 8x32?\"\n\nInitially, i am using  4 gradient accumulation of batch=32 each.\nAfter using  8 gradient accumulation of batch=32 each, the \"jump\" occur\n\n\n    \n  \n        for images, labels, indices in train_loader:\n  \n            ... ###  len(images) is 32  here ### ....\n    \n            # one iteration update  -------------\n            images  = Variable(images).cuda()\n            labels  = Variable(labels).cuda()\n            logits = net(images)\n            probs  = F.softmax(logits)\n\n            loss = F.cross_entropy(logits, labels)\n            acc  = top_accuracy(probs, labels, top_k=(1,))\n\n            # accumulate gradients\n            loss.backward()\n            if j%iter_accum == 0:  ### iter_accum is changed from 4 to 8 here ### \n                optimizer.step()\n                optimizer.zero_grad()\n\n           ...",
    "228907": "i am thinking the reason could be similar to equation 3 and 4 of the paper. Because the update is \"small\" per batch (due to class imbalance, noise, under fitting, etc), it \"doesn't matter\" if you update \"one by one\" or update \"all at once\". But if you update \"all at once\", the  backward propagation signals are \"boosted\" and signal to noise ratio improves.\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228907/7541/Selection_074.png",
    "228908": "Ah OK. I'm using Keras and I think it just computes the gradient over each mini-batch (i.e. no accumulation).\n\nSo the validation accuracy jumps up when you compute the gradient over a larger number of examples (8x32 samples instead of 4x32 samples). This would make the SGD slightly less stochastic, not more (i.e. it somewhat reduces the noise). After all, the larger the number of examples, the more the gradients approximate the gradients of the true loss function.",
    "228910": "facebook paper results for  minibatch size for imagenet 5K (training for 5K classes)\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228910/7542/Selection_075.png",
    "228914": "\" After all, the larger the number of examples, the more the gradients approximate the gradients of the true loss function.\"\n\ni am thinking something similar, but i always though that batch size of 256 to 512 is always enough. Now I am doing experiment like this:\n\n    while in training loop:\n\n             do back propagation\n            \n             check if loss plateau or not \n             if yes:\n                    increase effective batch size by increasing iter_accum",
    "228923": "I see that you write \"changing learning rate does not seem to help\", but still: If you are keeping the learning the same while increasing the batch size, it can have a similar effect as decreasing the learning rate, because effective per-image learning rate becomes smaller - so I wonder if the jumps can be explained by this?",
    "228924": "\" it can have a similar effect as decreasing the learning rate\"\n\nThis is correct. But my observation is that \"fixing the learning rate but increasing batch size\" seems to be faster and gives better results.",
    "228925": "Thanks! Will check if it helps me too :)",
    "229071": "convergence curve for imagnet-1K.https://github.com/apache/incubator-mxnet/tree/master/example/image-classification\n\nthere is another one at: http://dmlc.ml/mxnet/2015/10/27/training-deep-net-on-14-million-images.html\n\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n  https://github.com/tornadomeet/ResNet/tree/master/log\n\n  [1]: https://raw.githubusercontent.com/dmlc/web-data/master/mxnet/image/dist_converge.png\n  [2]: https://github.com/tornadomeet/ResNet/blob/master/log/training-curve.png?raw=true",
    "229129": "Too bad that there is no large scale experimental results in this paper.\n\nhttps://arxiv.org/pdf/1610.05792.pdf\n\n\"We analyzed and studied the behavior of alternative SGD methods in which the batch size increases over time.\nUnlike classical SGD methods, in which stochastic gradients quickly become swamped with noise, these “big\nbatch” methods maintain a nearly constant signal to noise ratio of the approximate gradient. As a result,\nbig batch methods are able to adaptively adjust batch sizes without user oversight. The proposed automated\nmethods are shown to be empirically comparable or better performing than other standard methods, but\nwithout requiring an expert user to choose learning rates and decay parameters.\"",
    "229141": "This is a great idea. Reminiscent of simulated annealing.",
    "229367": "my 224 se-resnet50 has been completed. Here are the single crop results:\n\n-  se-resnet50 train=224x224 crops from 256 (perturb scale, shift,rotate):\n\n     -- test = resize to 224 gives LB 0.66708  \n\n.\n\n-  se-resnet50 train=160x160 crops from 180:\n\n   -- 160x160 single center crop gives LB 0.63358\n\n   -- 180x180 gives LB 0.63839 \n\n   -- 180x180 resize to 160x160 gives LB  0.64356\n\n\n.\n\ntrain method for 224x224  se-resnet50 train:\n\n\n  ![enter image description here][1]\n\n -  use increasing batch size (hand adjusted) method at rate =0.01 to train until validation accuracy hits 0.64. This takes about 2 epoch.\n\n -  then train at rate=0.001 for another 0.5 epoch till validation accuracy hits 0.65.\n\n\ntotal training takes about 2 days. long iterations at the end are important for good results. (you can probably get better results if you run longer for lr =0.001, and even lr= 0.0001 later). you can acclerate early training.\n\ntrain software is at: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/229367/7552/batch_size2.png",
    "229491": "If you are using Adam optimizer, don't gradients get accumulated ?",
    "229841": "one side note: if you accumulate gradient for large batch size and you are using sdg+momentum, you might want to try experiments with different momentum. Sometimes, you need to lower them. ( i don't know why but my experiment shows sometimes it work better.  maybe large batch size don't need momentum?) you may also want to read this: http://stanford.edu/~imit/tuneyourmomentum/",
    "230003": "If I'm not wrong, in a similar article by Facebook, they suggest multiplying the learning rate by k. Being k the number of gradients accumulation steps. Doing so will keep the same \"effective\" learning rate",
    "230631": "How Facebook trained ImageNet in 1 hour?\n\nhttps://www.youtube.com/watch?v=eSmst0CUJAo\n\nhttps://caffe2.ai/docs/SynchronousSGD.html",
    "234456": "Have you compared accumulate gradients with no accumulate gradients on one gpu? I tried accumulate gradients according to your code, but it seemed to perform worse(batch size is 4x32).",
    "236702": "https://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf",
    "236708": "Adding some \"theory\" to Human Analog's statement to why big batches might be dangerous.\nhttps://arxiv.org/pdf/1609.04836.pdf",
    "236744": "Thanks. I am aware of the paper. Large batch works but have to be careful. if your loss is already decreasing, probably you don't have to use. If the loss is decreasing slowly, you may may to try large batch. Also, the below mention about experiments on alexnet and some check about gradient and gradient update values are necessary.\n\nhttps://people.eecs.berkeley.edu/~youyang/publications/batch32k_slide.pdf\n\nAlso note that ICLR 2017 paper does not apply rate multiplication as facebook paper,",
    "237500": "A new paper about this topic is currently under review at ICLR2018: https://openreview.net/pdf?id=B1Yy1BxCZ\n\n&gt; It is common practice to decay the learning rate. Here we show one can usually\nobtain the same learning curve on both training and test sets by instead increasing\nthe batch size during training.",
    "237515": "oh! too late for me to write a paper.\n\nbut good for me to know the results and apply to this kaggle competition",
    "237607": "The main point arising from our results is that, in contrast to previous conception, there is no inherent\ngeneralization problem with training using big mini batches. That is, model training using big\nmini-batches can generalize as good as models trained using small mini-batches. Previous works\nfound that big batch updates reach a sharp minima - which was claimed to provide poor generalization.\nIn this work we show that this is not necessarily the case: provided the needed adjustments, good\ngeneralization can be observed for big batch regimes\n\nhttps://arxiv.org/pdf/1705.08741.pdf\n\"Train longer, generalize better: closing the generalization gap in large batch training of neural networks\"",
    "239696": "DON’T DECAY THE LEARNING RATE,\nINCREASE THE BATCH SIZE\n\nhttps://arxiv.org/pdf/1711.00489.pdf\n\nNew Google paper on impact of large batch traning",
    "427598": "With his code... If my machine can run 256 batch and I accumulate for 512 batch for example it runs out of memory :("
  },
  "source": "meta"
}