{
  "id": 20316,
  "title": "batch size in CNN affecting performance",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20316",
  "author_name": "woshialex",
  "post_date": "2016-04-21T19:24:31.427000",
  "votes": 1,
  "comment_count": 18,
  "views": 5103,
  "content": "<p>I would naively think a larger batch size is good for performance (and quickly saturated after it is large enough). But in this competition, it seems a smaller batch size actually helps (CV score improved). A batch size of 16 is better than 64. Does anyone else see the same?</p>",
  "messages": [
    {
      "id": 116352,
      "postDate": "2016-04-23T15:52:56.230Z",
      "content": "<p>I have done some benchmarking of it on ImageNet. Except too large batch size cases (1024), they are equal, if you adjust your learning rate. The only loosers in this graphs are ones without adjusting.</p>\n\n<p><img src=\"https://github.com/ducha-aiki/caffenet-benchmark/raw/master/logs/batch_size/img/0.png\" alt=\"Batch size benchmark\" title> </p>\n\n<p><a href=\"https://github.com/ducha-aiki/caffenet-benchmark/blob/master/BatchSize.md\">https://github.com/ducha-aiki/caffenet-benchmark/blob/master/BatchSize.md</a></p>",
      "rawMarkdown": "I have done some benchmarking of it on ImageNet. Except too large batch size cases (1024), they are equal, if you adjust your learning rate. The only loosers in this graphs are ones without adjusting.\r\n\r\n![Batch size benchmark][1] \r\n\r\nhttps://github.com/ducha-aiki/caffenet-benchmark/blob/master/BatchSize.md\r\n\r\n  [1]: https://github.com/ducha-aiki/caffenet-benchmark/raw/master/logs/batch_size/img/0.png",
      "votes": 4
    },
    {
      "id": 116358,
      "postDate": "2016-04-23T16:46:11.127Z",
      "content": "<p>@Neil, the reason is scheduled 10x reducing learning rate at 100K (200K, 300K) iterations. References:</p>\n\n<blockquote>\n  <p>We used an equal learning rate for all layers, which we adjusted manually throughout training.\n  The heuristic which we followed was to <strong>divide the learning rate by 10 when the validation error\n  rate stopped improving with the current learning rate</strong>. The learning rate was initialized at 0.01 and\n  reduced three times prior to termination. </p>\n</blockquote>\n\n<p>Page 6, AlexNet paper. </p>\n\n<blockquote>\n  <p>base_lr: 0.01 lr_policy: &quot;step&quot; gamma: 0.1 stepsize: 100000</p>\n</blockquote>\n\n<p>alexnet solver, <a href=\"https://github.com/BVLC/caffe/blob/master/models/bvlc_alexnet/solver.prototxt\">https://github.com/BVLC/caffe/blob/master/models/bvlc_alexnet/solver.prototxt</a></p>",
      "rawMarkdown": "@Neil, the reason is scheduled 10x reducing learning rate at 100K (200K, 300K) iterations. References:\r\n\r\n> We used an equal learning rate for all layers, which we adjusted manually throughout training.\r\nThe heuristic which we followed was to **divide the learning rate by 10 when the validation error\r\nrate stopped improving with the current learning rate**. The learning rate was initialized at 0.01 and\r\nreduced three times prior to termination. \r\n\r\nPage 6, AlexNet paper. \r\n\r\n> base_lr: 0.01 lr_policy: \"step\" gamma: 0.1 stepsize: 100000\r\n\r\nalexnet solver, https://github.com/BVLC/caffe/blob/master/models/bvlc_alexnet/solver.prototxt\r\n\r\n ",
      "votes": 1
    },
    {
      "id": 116357,
      "postDate": "2016-04-23T16:37:50.720Z",
      "content": "<p>@old-ufo: I've been meaning to ask for a while, what is causing the synchronised upward jumps in test accuracy (in your graph at around 22 and 46 epochs)? I see these in a lot of published learning curves for deep networks, but not found the explanation.</p>",
      "rawMarkdown": "@old-ufo: I've been meaning to ask for a while, what is causing the synchronised upward jumps in test accuracy (in your graph at around 22 and 46 epochs)? I see these in a lot of published learning curves for deep networks, but not found the explanation.\r\n",
      "votes": 1
    },
    {
      "id": 116162,
      "postDate": "2016-04-22T10:14:20.257Z",
      "content": "<p>[quote=machines dislike learning;116096]</p>\n\n<p>Is there anyone can explain why theoretically? thx...</p>\n\n<p>[/quote]\nNot much to explain, smaller batch size =&gt; more gradient updates per epoch. <br>\nIncreasing epoch count should be comparable to decreasing batch size in the long run. The issue is that majority of people settle for 5-10 epochs, whereas it usually takes 100+ for convergence.</p>",
      "rawMarkdown": "[quote=machines dislike learning;116096]\r\n\r\nIs there anyone can explain why theoretically? thx...\r\n\r\n[/quote]\r\nNot much to explain, smaller batch size => more gradient updates per epoch.  \r\nIncreasing epoch count should be comparable to decreasing batch size in the long run. The issue is that majority of people settle for 5-10 epochs, whereas it usually takes 100+ for convergence.",
      "votes": 1
    },
    {
      "id": 116147,
      "postDate": "2016-04-22T08:32:27.193Z",
      "content": "<p>[quote=Lu&#237;s Andr&#233; Dutra e Silva;116132]</p>\n\n<p>The only reason I don't use a pure stochastic approach is that training time increases a lot. In this competition I used a batch size of 64 that resulted in about 10% top-1 error and, when I changed the batch size to 16, it dropped to 4%. This is a fact that someone should write about.</p>\n\n<p>[/quote]</p>\n\n<p>Something else, other than the batch size itself, is contributing to this. Perhaps share the two sets of meta-params that did this after the competition if you'd like some analysis (not saying I could do that, just that I am very skeptical of there being a useful effect here.</p>\n\n<p>One thing to beware of - the accuracy percentages in this competition are all over the place compared to the logloss metric. Between one epoch and the next I am seeing big swings in accuracy compared to logloss. Here's an example excerpt from one of my runs:</p>\n\n<pre><code>Epoch 18/20 21700/21700 [===] - 29s - loss: 1.2887 - acc: 0.5801 - val_loss: 1.1078 - val_acc: 0.6906 \nEpoch 19/20 21700/21700 [===] - 29s - loss: 1.3090 - acc: 0.5718 - val_loss: 0.9286 - val_acc: 0.7459\nEpoch 20/20 21700/21700 [===] - 29s - loss: 1.3075 - acc: 0.5748 - val_loss: 1.0084 - val_acc: 0.8052\n</code></pre>\n\n<p>. . . so to convert to your terms, I got a 31% CV error down to 20% CV error just over two otherwise inconsequential epochs (and where the CV logloss - the metric we care about in this competition) varied by less than 2%). Now 10% down to 4% is a lot better, and perhaps there is really something going on here. But so far I am not convinced this is anything other than human pattern-matching.</p>\n\n<p>If you have time, try changing random seed (i.e. something we can all agree has no reliable effect on metrics) a couple of times, see if you get comparable shifts in CV accuracy reported.</p>\n\n<p>Also, if you don't mind saying, which optimiser are you using? As I mentioned before, some optimisers are sensitive to batch size, and it is worth tuning it. But that is not the same as a generic &quot;smaller value is always better&quot;.</p>",
      "rawMarkdown": "[quote=Luís André Dutra e Silva;116132]\r\n\r\nThe only reason I don't use a pure stochastic approach is that training time increases a lot. In this competition I used a batch size of 64 that resulted in about 10% top-1 error and, when I changed the batch size to 16, it dropped to 4%. This is a fact that someone should write about.\r\n\r\n[/quote]\r\n\r\nSomething else, other than the batch size itself, is contributing to this. Perhaps share the two sets of meta-params that did this after the competition if you'd like some analysis (not saying I could do that, just that I am very skeptical of there being a useful effect here.\r\n\r\nOne thing to beware of - the accuracy percentages in this competition are all over the place compared to the logloss metric. Between one epoch and the next I am seeing big swings in accuracy compared to logloss. Here's an example excerpt from one of my runs:\r\n\r\n    Epoch 18/20 21700/21700 [===] - 29s - loss: 1.2887 - acc: 0.5801 - val_loss: 1.1078 - val_acc: 0.6906 \r\n    Epoch 19/20 21700/21700 [===] - 29s - loss: 1.3090 - acc: 0.5718 - val_loss: 0.9286 - val_acc: 0.7459\r\n    Epoch 20/20 21700/21700 [===] - 29s - loss: 1.3075 - acc: 0.5748 - val_loss: 1.0084 - val_acc: 0.8052\r\n\r\n . . . so to convert to your terms, I got a 31% CV error down to 20% CV error just over two otherwise inconsequential epochs (and where the CV logloss - the metric we care about in this competition) varied by less than 2%). Now 10% down to 4% is a lot better, and perhaps there is really something going on here. But so far I am not convinced this is anything other than human pattern-matching.\r\n\r\nIf you have time, try changing random seed (i.e. something we can all agree has no reliable effect on metrics) a couple of times, see if you get comparable shifts in CV accuracy reported.\r\n\r\nAlso, if you don't mind saying, which optimiser are you using? As I mentioned before, some optimisers are sensitive to batch size, and it is worth tuning it. But that is not the same as a generic \"smaller value is always better\".\r\n\r\n\r\n\r\n\r\n",
      "votes": 1
    },
    {
      "id": 116130,
      "postDate": "2016-04-22T07:04:02.437Z",
      "content": "<p>Didn't work for me. I was working with a batch size of 64, changed only that to 16. Slightly worse result.</p>\n\n<p>Bear in mind that changing batch sizes will have some effects similar to changing random seeds because updates will start working from different locations (in param space) after the first step. So 50% of people are likely to report a small improvement, 50% slightly worse, even if there is no real meaningful effect here.</p>\n\n<p>Some optimisers, like RMSProp, are sensitive to batch size, and you may find there is an optimum value.</p>\n\n<p>The take-away here is that there is no simple magic &quot;make it better&quot; param value (for any of the hyper-params in a neural network). If there was, it would be all over the published papers.</p>",
      "rawMarkdown": "Didn't work for me. I was working with a batch size of 64, changed only that to 16. Slightly worse result.\r\n\r\nBear in mind that changing batch sizes will have some effects similar to changing random seeds because updates will start working from different locations (in param space) after the first step. So 50% of people are likely to report a small improvement, 50% slightly worse, even if there is no real meaningful effect here.\r\n\r\nSome optimisers, like RMSProp, are sensitive to batch size, and you may find there is an optimum value.\r\n\r\nThe take-away here is that there is no simple magic \"make it better\" param value (for any of the hyper-params in a neural network). If there was, it would be all over the published papers.\r\n",
      "votes": 1
    },
    {
      "id": 116060,
      "postDate": "2016-04-21T19:24:31.427Z",
      "content": "<p>I would naively think a larger batch size is good for performance (and quickly saturated after it is large enough). But in this competition, it seems a smaller batch size actually helps (CV score improved). A batch size of 16 is better than 64. Does anyone else see the same?</p>",
      "rawMarkdown": "I would naively think a larger batch size is good for performance (and quickly saturated after it is large enough). But in this competition, it seems a smaller batch size actually helps (CV score improved). A batch size of 16 is better than 64. Does anyone else see the same?",
      "votes": 1
    },
    {
      "id": 116172,
      "postDate": "2016-04-22T11:33:32.827Z",
      "content": "<p>[quote=Roman Ring;116162]</p>\n\n<p>[quote=machines dislike learning;116096]</p>\n\n<p>Is there anyone can explain why theoretically? thx...</p>\n\n<p>[/quote]\nNot much to explain, smaller batch size =&gt; more gradient updates per epoch. <br>\nIncreasing epoch count should be comparable to decreasing batch size in the long run. The issue is that majority of people settle for 5-10 epochs, whereas it usually takes 100+ for convergence.</p>\n\n<p>[/quote]</p>\n\n<p>However, there is an offset effect to this. Smaller batch size =&gt; less accurate measure of gradient.</p>\n\n<p>Depending on how noisy the data is, this can make different batch have different efficiency - in terms of  how many epochs overall will be required. There will be a &quot;sweet spot&quot; where number of epochs to get best score out of an architecture (and specific training data) is minimum.</p>\n\n<p>Batched processing of examples is also faster to compute than one-by-one, although this cuts off fairly quickly. So there is possibly a second &quot;sweet spot&quot; where amount of time to get best score out of an architecture is minimum.</p>\n\n<p>I suggest that the observed improvements to score are exploring this, and are not an indication of &quot;smaller batch size returns better results&quot; -  neither in general, nor just for this competition data. However, that doesn't mean it's not worth trying, just that it may not work for everyone - in fact trying a larger batch size may make just as much sense.</p>\n\n<hr>\n\n<p>My own experiments, keeping everything else the same (modified ZFTurbo's script, a relatively shallow network running on CPU - unfortunate for me, not going to win this one - with SGD optimiser for 20 epochs and 26-fold by driver):</p>\n\n<p>Batch size 16  =&gt;  LB score 1.38259</p>\n\n<p>Batch size 48 =&gt; LB score 1.13532</p>\n\n<p>Batch size 64 =&gt; LB score 1.13394</p>\n\n<p>Batch size 96 =&gt; LB score 1.17227</p>\n\n<p>Batch size 256 =&gt; LB score 1.42374</p>\n\n<p>This is not surprising, I have already tuned other params such as learning rate, momentum, number of epochs, so have probably fit them loosely to work with the original batch size.</p>",
      "rawMarkdown": "[quote=Roman Ring;116162]\r\n\r\n[quote=machines dislike learning;116096]\r\n\r\nIs there anyone can explain why theoretically? thx...\r\n\r\n[/quote]\r\nNot much to explain, smaller batch size => more gradient updates per epoch.  \r\nIncreasing epoch count should be comparable to decreasing batch size in the long run. The issue is that majority of people settle for 5-10 epochs, whereas it usually takes 100+ for convergence.\r\n\r\n[/quote]\r\n\r\nHowever, there is an offset effect to this. Smaller batch size => less accurate measure of gradient.\r\n\r\nDepending on how noisy the data is, this can make different batch have different efficiency - in terms of  how many epochs overall will be required. There will be a \"sweet spot\" where number of epochs to get best score out of an architecture (and specific training data) is minimum.\r\n\r\nBatched processing of examples is also faster to compute than one-by-one, although this cuts off fairly quickly. So there is possibly a second \"sweet spot\" where amount of time to get best score out of an architecture is minimum.\r\n\r\nI suggest that the observed improvements to score are exploring this, and are not an indication of \"smaller batch size returns better results\" -  neither in general, nor just for this competition data. However, that doesn't mean it's not worth trying, just that it may not work for everyone - in fact trying a larger batch size may make just as much sense.\r\n\r\n----\r\n\r\nMy own experiments, keeping everything else the same (modified ZFTurbo's script, a relatively shallow network running on CPU - unfortunate for me, not going to win this one - with SGD optimiser for 20 epochs and 26-fold by driver):\r\n\r\nBatch size 16  =>  LB score 1.38259\r\n\r\nBatch size 48 => LB score 1.13532\r\n\r\nBatch size 64 => LB score 1.13394\r\n\r\nBatch size 96 => LB score 1.17227\r\n\r\nBatch size 256 => LB score 1.42374\r\n\r\nThis is not surprising, I have already tuned other params such as learning rate, momentum, number of epochs, so have probably fit them loosely to work with the original batch size.",
      "votes": 2
    },
    {
      "id": 116150,
      "postDate": "2016-04-22T08:43:42.050Z",
      "content": "<p>[quote=Lu&#237;s Andr&#233; Dutra e Silva;116149]</p>\n\n<p>what do you mean by &quot;human pattern-matching&quot; ?</p>\n\n<p>[/quote]</p>\n\n<p>I mean that you have noticed a trend, and extrapolated from a small sample, and not in rigorous scientific or statistically valid way. People are good at finding patterns when there aren't any.</p>\n\n<p>I also used SGD when testing the idea on my meta-params. I made exactly the same change as you reported, 64 batch size down to 16. The effect was an increase in logloss (around 0.25) - i.e. it was worse. Clearly these params are not as good a starting point as yours. But equally clearly, a message of &quot;lower batch size is better&quot; is not true in all cases.</p>\n\n<p>(Edit: I double-checked my submissions, and actually the increase in logloss was quite significantly worse - makes me more likely to believe there is an effect of some sort here, but of course less able to believe the &quot;smaller is better&quot; suggestion).</p>",
      "rawMarkdown": "[quote=Luís André Dutra e Silva;116149]\r\n\r\nwhat do you mean by \"human pattern-matching\" ?\r\n\r\n[/quote]\r\n\r\nI mean that you have noticed a trend, and extrapolated from a small sample, and not in rigorous scientific or statistically valid way. People are good at finding patterns when there aren't any.\r\n\r\nI also used SGD when testing the idea on my meta-params. I made exactly the same change as you reported, 64 batch size down to 16. The effect was an increase in logloss (around 0.25) - i.e. it was worse. Clearly these params are not as good a starting point as yours. But equally clearly, a message of \"lower batch size is better\" is not true in all cases.\r\n\r\n(Edit: I double-checked my submissions, and actually the increase in logloss was quite significantly worse - makes me more likely to believe there is an effect of some sort here, but of course less able to believe the \"smaller is better\" suggestion).",
      "votes": 2
    },
    {
      "id": 116430,
      "postDate": "2016-04-24T07:26:18.377Z",
      "content": "<p>[quote=Mustyy;116404]</p>\n\n<p>@Neil\nHey I agree running it on a CPU is such a long and slow process.\nI also modified ZFTurbo's script and adjusted the epochs to 25 with kfold =5. And batch size of either 48 or 64 between different experiments.</p>\n\n<p>Any insights as to why you have picked 26?\nSecondly, any specific values for nb_pool &amp; nb_conv? And any input you have on how to adjust the learning rate over time?</p>\n\n<p>If I could get a GPU would you be interested in teaming up? :D</p>\n\n<p>[/quote]</p>\n\n<p>26 is the number of drivers, so you get each driver in the CV set on each fold, and get some insight into what the trained network is learning. I've just left it like that for the tests above, because that setup is also my best LB score to date. I guess other values should work, although I'd probably default to 13 for neatness of having always same number of drivers in each CV set (I have no idea if this makes a difference)</p>\n\n<p>I haven't changed nb_pool and nb_conv. I don't have any real insight into other params, mainly I've been trying to force more regularisation into the model (more dropout, trying L2 regularisation and also maxnorm weight constraints), or change the architecture to find something that runs in reasonable time but is a little more accurate. </p>\n\n<p>I've had some limited success, and I don't think my experience offers much use in the different space of deeper networks using the GPU - something I'm saving for another time. I'm not looking for a team, and probably I will drop out of the competition after I've used it to learn a bit more Python/Keras.</p>",
      "rawMarkdown": "[quote=Mustyy;116404]\r\n\r\n@Neil\r\nHey I agree running it on a CPU is such a long and slow process.\r\nI also modified ZFTurbo's script and adjusted the epochs to 25 with kfold =5. And batch size of either 48 or 64 between different experiments.\r\n\r\nAny insights as to why you have picked 26?\r\nSecondly, any specific values for nb_pool & nb_conv? And any input you have on how to adjust the learning rate over time?\r\n\r\nIf I could get a GPU would you be interested in teaming up? :D\r\n\r\n[/quote]\r\n\r\n26 is the number of drivers, so you get each driver in the CV set on each fold, and get some insight into what the trained network is learning. I've just left it like that for the tests above, because that setup is also my best LB score to date. I guess other values should work, although I'd probably default to 13 for neatness of having always same number of drivers in each CV set (I have no idea if this makes a difference)\r\n\r\nI haven't changed nb_pool and nb_conv. I don't have any real insight into other params, mainly I've been trying to force more regularisation into the model (more dropout, trying L2 regularisation and also maxnorm weight constraints), or change the architecture to find something that runs in reasonable time but is a little more accurate. \r\n\r\nI've had some limited success, and I don't think my experience offers much use in the different space of deeper networks using the GPU - something I'm saving for another time. I'm not looking for a team, and probably I will drop out of the competition after I've used it to learn a bit more Python/Keras.\r\n"
    },
    {
      "id": 116417,
      "postDate": "2016-04-24T03:18:36.177Z",
      "content": "<p>I have the same observation, batch size of 16 works better for me. When I try 32 or 64, it is becoming over-fitted somehow.  </p>",
      "rawMarkdown": "I have the same observation, batch size of 16 works better for me. When I try 32 or 64, it is becoming over-fitted somehow.  "
    },
    {
      "id": 116404,
      "postDate": "2016-04-23T23:43:06.353Z",
      "content": "<p>@Neil\nHey I agree running it on a CPU is such a long and slow process.\nI also modified ZFTurbo's script and adjusted the epochs to 25 with kfold =5. And batch size of either 48 or 64 between different experiments.</p>\n\n<p>Any insights as to why you have picked 26?\nSecondly, any specific values for nb_pool &amp; nb_conv? And any input you have on how to adjust the learning rate over time?</p>\n\n<p>If I could get a GPU would you be interested in teaming up? :D</p>",
      "rawMarkdown": "@Neil\r\nHey I agree running it on a CPU is such a long and slow process.\r\nI also modified ZFTurbo's script and adjusted the epochs to 25 with kfold =5. And batch size of either 48 or 64 between different experiments.\r\n\r\nAny insights as to why you have picked 26?\r\nSecondly, any specific values for nb_pool & nb_conv? And any input you have on how to adjust the learning rate over time?\r\n\r\nIf I could get a GPU would you be interested in teaming up? :D"
    },
    {
      "id": 116154,
      "postDate": "2016-04-22T09:07:36.210Z",
      "content": "<p>Another thing to be aware of is how many epochs you are running. The learning curves for different batch sizes will be different, and the amount of overfit will be different at the same epoch number. So it is possible that people reporting an improvement from smaller batch sizes have in essence tuned their early stopping.</p>",
      "rawMarkdown": "Another thing to be aware of is how many epochs you are running. The learning curves for different batch sizes will be different, and the amount of overfit will be different at the same epoch number. So it is possible that people reporting an improvement from smaller batch sizes have in essence tuned their early stopping."
    },
    {
      "id": 116096,
      "postDate": "2016-04-22T01:01:16.620Z",
      "content": "<p>Is there anyone can explain why theoretically? thx...</p>",
      "rawMarkdown": "Is there anyone can explain why theoretically? thx..."
    },
    {
      "id": 116407,
      "postDate": "2016-04-24T00:06:45.790Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 116149,
      "postDate": "2016-04-22T08:41:34.440Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 116132,
      "postDate": "2016-04-22T07:13:44.600Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 116100,
      "postDate": "2016-04-22T01:51:54.050Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 116066,
      "postDate": "2016-04-21T20:46:01.467Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 116352,
      "author_name": "old-ufo",
      "author_url": "",
      "post_date": "2016-04-23T15:52:56.230000",
      "content": "<p>I have done some benchmarking of it on ImageNet. Except too large batch size cases (1024), they are equal, if you adjust your learning rate. The only loosers in this graphs are ones without adjusting.</p>\n\n<p><img src=\"https://github.com/ducha-aiki/caffenet-benchmark/raw/master/logs/batch_size/img/0.png\" alt=\"Batch size benchmark\" title> </p>\n\n<p><a href=\"https://github.com/ducha-aiki/caffenet-benchmark/blob/master/BatchSize.md\">https://github.com/ducha-aiki/caffenet-benchmark/blob/master/BatchSize.md</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 116358,
      "author_name": "old-ufo",
      "author_url": "",
      "post_date": "2016-04-23T16:46:11.127000",
      "content": "<p>@Neil, the reason is scheduled 10x reducing learning rate at 100K (200K, 300K) iterations. References:</p>\n\n<blockquote>\n  <p>We used an equal learning rate for all layers, which we adjusted manually throughout training.\n  The heuristic which we followed was to <strong>divide the learning rate by 10 when the validation error\n  rate stopped improving with the current learning rate</strong>. The learning rate was initialized at 0.01 and\n  reduced three times prior to termination. </p>\n</blockquote>\n\n<p>Page 6, AlexNet paper. </p>\n\n<blockquote>\n  <p>base_lr: 0.01 lr_policy: &quot;step&quot; gamma: 0.1 stepsize: 100000</p>\n</blockquote>\n\n<p>alexnet solver, <a href=\"https://github.com/BVLC/caffe/blob/master/models/bvlc_alexnet/solver.prototxt\">https://github.com/BVLC/caffe/blob/master/models/bvlc_alexnet/solver.prototxt</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116357,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-04-23T16:37:50.720000",
      "content": "<p>@old-ufo: I've been meaning to ask for a while, what is causing the synchronised upward jumps in test accuracy (in your graph at around 22 and 46 epochs)? I see these in a lot of published learning curves for deep networks, but not found the explanation.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116162,
      "author_name": "Roman Ring",
      "author_url": "",
      "post_date": "2016-04-22T10:14:20.257000",
      "content": "<p>[quote=machines dislike learning;116096]</p>\n\n<p>Is there anyone can explain why theoretically? thx...</p>\n\n<p>[/quote]\nNot much to explain, smaller batch size =&gt; more gradient updates per epoch. <br>\nIncreasing epoch count should be comparable to decreasing batch size in the long run. The issue is that majority of people settle for 5-10 epochs, whereas it usually takes 100+ for convergence.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116147,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-04-22T08:32:27.193000",
      "content": "<p>[quote=Lu&#237;s Andr&#233; Dutra e Silva;116132]</p>\n\n<p>The only reason I don't use a pure stochastic approach is that training time increases a lot. In this competition I used a batch size of 64 that resulted in about 10% top-1 error and, when I changed the batch size to 16, it dropped to 4%. This is a fact that someone should write about.</p>\n\n<p>[/quote]</p>\n\n<p>Something else, other than the batch size itself, is contributing to this. Perhaps share the two sets of meta-params that did this after the competition if you'd like some analysis (not saying I could do that, just that I am very skeptical of there being a useful effect here.</p>\n\n<p>One thing to beware of - the accuracy percentages in this competition are all over the place compared to the logloss metric. Between one epoch and the next I am seeing big swings in accuracy compared to logloss. Here's an example excerpt from one of my runs:</p>\n\n<pre><code>Epoch 18/20 21700/21700 [===] - 29s - loss: 1.2887 - acc: 0.5801 - val_loss: 1.1078 - val_acc: 0.6906 \nEpoch 19/20 21700/21700 [===] - 29s - loss: 1.3090 - acc: 0.5718 - val_loss: 0.9286 - val_acc: 0.7459\nEpoch 20/20 21700/21700 [===] - 29s - loss: 1.3075 - acc: 0.5748 - val_loss: 1.0084 - val_acc: 0.8052\n</code></pre>\n\n<p>. . . so to convert to your terms, I got a 31% CV error down to 20% CV error just over two otherwise inconsequential epochs (and where the CV logloss - the metric we care about in this competition) varied by less than 2%). Now 10% down to 4% is a lot better, and perhaps there is really something going on here. But so far I am not convinced this is anything other than human pattern-matching.</p>\n\n<p>If you have time, try changing random seed (i.e. something we can all agree has no reliable effect on metrics) a couple of times, see if you get comparable shifts in CV accuracy reported.</p>\n\n<p>Also, if you don't mind saying, which optimiser are you using? As I mentioned before, some optimisers are sensitive to batch size, and it is worth tuning it. But that is not the same as a generic &quot;smaller value is always better&quot;.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116130,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-04-22T07:04:02.437000",
      "content": "<p>Didn't work for me. I was working with a batch size of 64, changed only that to 16. Slightly worse result.</p>\n\n<p>Bear in mind that changing batch sizes will have some effects similar to changing random seeds because updates will start working from different locations (in param space) after the first step. So 50% of people are likely to report a small improvement, 50% slightly worse, even if there is no real meaningful effect here.</p>\n\n<p>Some optimisers, like RMSProp, are sensitive to batch size, and you may find there is an optimum value.</p>\n\n<p>The take-away here is that there is no simple magic &quot;make it better&quot; param value (for any of the hyper-params in a neural network). If there was, it would be all over the published papers.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116172,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-04-22T11:33:32.827000",
      "content": "<p>[quote=Roman Ring;116162]</p>\n\n<p>[quote=machines dislike learning;116096]</p>\n\n<p>Is there anyone can explain why theoretically? thx...</p>\n\n<p>[/quote]\nNot much to explain, smaller batch size =&gt; more gradient updates per epoch. <br>\nIncreasing epoch count should be comparable to decreasing batch size in the long run. The issue is that majority of people settle for 5-10 epochs, whereas it usually takes 100+ for convergence.</p>\n\n<p>[/quote]</p>\n\n<p>However, there is an offset effect to this. Smaller batch size =&gt; less accurate measure of gradient.</p>\n\n<p>Depending on how noisy the data is, this can make different batch have different efficiency - in terms of  how many epochs overall will be required. There will be a &quot;sweet spot&quot; where number of epochs to get best score out of an architecture (and specific training data) is minimum.</p>\n\n<p>Batched processing of examples is also faster to compute than one-by-one, although this cuts off fairly quickly. So there is possibly a second &quot;sweet spot&quot; where amount of time to get best score out of an architecture is minimum.</p>\n\n<p>I suggest that the observed improvements to score are exploring this, and are not an indication of &quot;smaller batch size returns better results&quot; -  neither in general, nor just for this competition data. However, that doesn't mean it's not worth trying, just that it may not work for everyone - in fact trying a larger batch size may make just as much sense.</p>\n\n<hr>\n\n<p>My own experiments, keeping everything else the same (modified ZFTurbo's script, a relatively shallow network running on CPU - unfortunate for me, not going to win this one - with SGD optimiser for 20 epochs and 26-fold by driver):</p>\n\n<p>Batch size 16  =&gt;  LB score 1.38259</p>\n\n<p>Batch size 48 =&gt; LB score 1.13532</p>\n\n<p>Batch size 64 =&gt; LB score 1.13394</p>\n\n<p>Batch size 96 =&gt; LB score 1.17227</p>\n\n<p>Batch size 256 =&gt; LB score 1.42374</p>\n\n<p>This is not surprising, I have already tuned other params such as learning rate, momentum, number of epochs, so have probably fit them loosely to work with the original batch size.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 116150,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-04-22T08:43:42.050000",
      "content": "<p>[quote=Lu&#237;s Andr&#233; Dutra e Silva;116149]</p>\n\n<p>what do you mean by &quot;human pattern-matching&quot; ?</p>\n\n<p>[/quote]</p>\n\n<p>I mean that you have noticed a trend, and extrapolated from a small sample, and not in rigorous scientific or statistically valid way. People are good at finding patterns when there aren't any.</p>\n\n<p>I also used SGD when testing the idea on my meta-params. I made exactly the same change as you reported, 64 batch size down to 16. The effect was an increase in logloss (around 0.25) - i.e. it was worse. Clearly these params are not as good a starting point as yours. But equally clearly, a message of &quot;lower batch size is better&quot; is not true in all cases.</p>\n\n<p>(Edit: I double-checked my submissions, and actually the increase in logloss was quite significantly worse - makes me more likely to believe there is an effect of some sort here, but of course less able to believe the &quot;smaller is better&quot; suggestion).</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 116430,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-04-24T07:26:18.377000",
      "content": "<p>[quote=Mustyy;116404]</p>\n\n<p>@Neil\nHey I agree running it on a CPU is such a long and slow process.\nI also modified ZFTurbo's script and adjusted the epochs to 25 with kfold =5. And batch size of either 48 or 64 between different experiments.</p>\n\n<p>Any insights as to why you have picked 26?\nSecondly, any specific values for nb_pool &amp; nb_conv? And any input you have on how to adjust the learning rate over time?</p>\n\n<p>If I could get a GPU would you be interested in teaming up? :D</p>\n\n<p>[/quote]</p>\n\n<p>26 is the number of drivers, so you get each driver in the CV set on each fold, and get some insight into what the trained network is learning. I've just left it like that for the tests above, because that setup is also my best LB score to date. I guess other values should work, although I'd probably default to 13 for neatness of having always same number of drivers in each CV set (I have no idea if this makes a difference)</p>\n\n<p>I haven't changed nb_pool and nb_conv. I don't have any real insight into other params, mainly I've been trying to force more regularisation into the model (more dropout, trying L2 regularisation and also maxnorm weight constraints), or change the architecture to find something that runs in reasonable time but is a little more accurate. </p>\n\n<p>I've had some limited success, and I don't think my experience offers much use in the different space of deeper networks using the GPU - something I'm saving for another time. I'm not looking for a team, and probably I will drop out of the competition after I've used it to learn a bit more Python/Keras.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116417,
      "author_name": "june",
      "author_url": "",
      "post_date": "2016-04-24T03:18:36.177000",
      "content": "<p>I have the same observation, batch size of 16 works better for me. When I try 32 or 64, it is becoming over-fitted somehow.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116404,
      "author_name": "Mustyy",
      "author_url": "",
      "post_date": "2016-04-23T23:43:06.353000",
      "content": "<p>@Neil\nHey I agree running it on a CPU is such a long and slow process.\nI also modified ZFTurbo's script and adjusted the epochs to 25 with kfold =5. And batch size of either 48 or 64 between different experiments.</p>\n\n<p>Any insights as to why you have picked 26?\nSecondly, any specific values for nb_pool &amp; nb_conv? And any input you have on how to adjust the learning rate over time?</p>\n\n<p>If I could get a GPU would you be interested in teaming up? :D</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116154,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-04-22T09:07:36.210000",
      "content": "<p>Another thing to be aware of is how many epochs you are running. The learning curves for different batch sizes will be different, and the amount of overfit will be different at the same epoch number. So it is possible that people reporting an improvement from smaller batch sizes have in essence tuned their early stopping.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116096,
      "author_name": "machines dislike learning",
      "author_url": "",
      "post_date": "2016-04-22T01:01:16.620000",
      "content": "<p>Is there anyone can explain why theoretically? thx...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116407,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-24T00:06:45.790000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116149,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-22T08:41:34.440000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116132,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-22T07:13:44.600000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116100,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-22T01:51:54.050000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116066,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-21T20:46:01.467000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116352": "I have done some benchmarking of it on ImageNet. Except too large batch size cases (1024), they are equal, if you adjust your learning rate. The only loosers in this graphs are ones without adjusting.\r\n\r\n![Batch size benchmark][1] \r\n\r\nhttps://github.com/ducha-aiki/caffenet-benchmark/blob/master/BatchSize.md\r\n\r\n  [1]: https://github.com/ducha-aiki/caffenet-benchmark/raw/master/logs/batch_size/img/0.png",
    "116358": "@Neil, the reason is scheduled 10x reducing learning rate at 100K (200K, 300K) iterations. References:\r\n\r\n> We used an equal learning rate for all layers, which we adjusted manually throughout training.\r\nThe heuristic which we followed was to **divide the learning rate by 10 when the validation error\r\nrate stopped improving with the current learning rate**. The learning rate was initialized at 0.01 and\r\nreduced three times prior to termination. \r\n\r\nPage 6, AlexNet paper. \r\n\r\n> base_lr: 0.01 lr_policy: \"step\" gamma: 0.1 stepsize: 100000\r\n\r\nalexnet solver, https://github.com/BVLC/caffe/blob/master/models/bvlc_alexnet/solver.prototxt\r\n\r\n ",
    "116357": "@old-ufo: I've been meaning to ask for a while, what is causing the synchronised upward jumps in test accuracy (in your graph at around 22 and 46 epochs)? I see these in a lot of published learning curves for deep networks, but not found the explanation.\r\n",
    "116162": "[quote=machines dislike learning;116096]\r\n\r\nIs there anyone can explain why theoretically? thx...\r\n\r\n[/quote]\r\nNot much to explain, smaller batch size => more gradient updates per epoch.  \r\nIncreasing epoch count should be comparable to decreasing batch size in the long run. The issue is that majority of people settle for 5-10 epochs, whereas it usually takes 100+ for convergence.",
    "116147": "[quote=Luís André Dutra e Silva;116132]\r\n\r\nThe only reason I don't use a pure stochastic approach is that training time increases a lot. In this competition I used a batch size of 64 that resulted in about 10% top-1 error and, when I changed the batch size to 16, it dropped to 4%. This is a fact that someone should write about.\r\n\r\n[/quote]\r\n\r\nSomething else, other than the batch size itself, is contributing to this. Perhaps share the two sets of meta-params that did this after the competition if you'd like some analysis (not saying I could do that, just that I am very skeptical of there being a useful effect here.\r\n\r\nOne thing to beware of - the accuracy percentages in this competition are all over the place compared to the logloss metric. Between one epoch and the next I am seeing big swings in accuracy compared to logloss. Here's an example excerpt from one of my runs:\r\n\r\n    Epoch 18/20 21700/21700 [===] - 29s - loss: 1.2887 - acc: 0.5801 - val_loss: 1.1078 - val_acc: 0.6906 \r\n    Epoch 19/20 21700/21700 [===] - 29s - loss: 1.3090 - acc: 0.5718 - val_loss: 0.9286 - val_acc: 0.7459\r\n    Epoch 20/20 21700/21700 [===] - 29s - loss: 1.3075 - acc: 0.5748 - val_loss: 1.0084 - val_acc: 0.8052\r\n\r\n . . . so to convert to your terms, I got a 31% CV error down to 20% CV error just over two otherwise inconsequential epochs (and where the CV logloss - the metric we care about in this competition) varied by less than 2%). Now 10% down to 4% is a lot better, and perhaps there is really something going on here. But so far I am not convinced this is anything other than human pattern-matching.\r\n\r\nIf you have time, try changing random seed (i.e. something we can all agree has no reliable effect on metrics) a couple of times, see if you get comparable shifts in CV accuracy reported.\r\n\r\nAlso, if you don't mind saying, which optimiser are you using? As I mentioned before, some optimisers are sensitive to batch size, and it is worth tuning it. But that is not the same as a generic \"smaller value is always better\".\r\n\r\n\r\n\r\n\r\n",
    "116130": "Didn't work for me. I was working with a batch size of 64, changed only that to 16. Slightly worse result.\r\n\r\nBear in mind that changing batch sizes will have some effects similar to changing random seeds because updates will start working from different locations (in param space) after the first step. So 50% of people are likely to report a small improvement, 50% slightly worse, even if there is no real meaningful effect here.\r\n\r\nSome optimisers, like RMSProp, are sensitive to batch size, and you may find there is an optimum value.\r\n\r\nThe take-away here is that there is no simple magic \"make it better\" param value (for any of the hyper-params in a neural network). If there was, it would be all over the published papers.\r\n",
    "116060": "I would naively think a larger batch size is good for performance (and quickly saturated after it is large enough). But in this competition, it seems a smaller batch size actually helps (CV score improved). A batch size of 16 is better than 64. Does anyone else see the same?",
    "116172": "[quote=Roman Ring;116162]\r\n\r\n[quote=machines dislike learning;116096]\r\n\r\nIs there anyone can explain why theoretically? thx...\r\n\r\n[/quote]\r\nNot much to explain, smaller batch size => more gradient updates per epoch.  \r\nIncreasing epoch count should be comparable to decreasing batch size in the long run. The issue is that majority of people settle for 5-10 epochs, whereas it usually takes 100+ for convergence.\r\n\r\n[/quote]\r\n\r\nHowever, there is an offset effect to this. Smaller batch size => less accurate measure of gradient.\r\n\r\nDepending on how noisy the data is, this can make different batch have different efficiency - in terms of  how many epochs overall will be required. There will be a \"sweet spot\" where number of epochs to get best score out of an architecture (and specific training data) is minimum.\r\n\r\nBatched processing of examples is also faster to compute than one-by-one, although this cuts off fairly quickly. So there is possibly a second \"sweet spot\" where amount of time to get best score out of an architecture is minimum.\r\n\r\nI suggest that the observed improvements to score are exploring this, and are not an indication of \"smaller batch size returns better results\" -  neither in general, nor just for this competition data. However, that doesn't mean it's not worth trying, just that it may not work for everyone - in fact trying a larger batch size may make just as much sense.\r\n\r\n----\r\n\r\nMy own experiments, keeping everything else the same (modified ZFTurbo's script, a relatively shallow network running on CPU - unfortunate for me, not going to win this one - with SGD optimiser for 20 epochs and 26-fold by driver):\r\n\r\nBatch size 16  =>  LB score 1.38259\r\n\r\nBatch size 48 => LB score 1.13532\r\n\r\nBatch size 64 => LB score 1.13394\r\n\r\nBatch size 96 => LB score 1.17227\r\n\r\nBatch size 256 => LB score 1.42374\r\n\r\nThis is not surprising, I have already tuned other params such as learning rate, momentum, number of epochs, so have probably fit them loosely to work with the original batch size.",
    "116150": "[quote=Luís André Dutra e Silva;116149]\r\n\r\nwhat do you mean by \"human pattern-matching\" ?\r\n\r\n[/quote]\r\n\r\nI mean that you have noticed a trend, and extrapolated from a small sample, and not in rigorous scientific or statistically valid way. People are good at finding patterns when there aren't any.\r\n\r\nI also used SGD when testing the idea on my meta-params. I made exactly the same change as you reported, 64 batch size down to 16. The effect was an increase in logloss (around 0.25) - i.e. it was worse. Clearly these params are not as good a starting point as yours. But equally clearly, a message of \"lower batch size is better\" is not true in all cases.\r\n\r\n(Edit: I double-checked my submissions, and actually the increase in logloss was quite significantly worse - makes me more likely to believe there is an effect of some sort here, but of course less able to believe the \"smaller is better\" suggestion).",
    "116430": "[quote=Mustyy;116404]\r\n\r\n@Neil\r\nHey I agree running it on a CPU is such a long and slow process.\r\nI also modified ZFTurbo's script and adjusted the epochs to 25 with kfold =5. And batch size of either 48 or 64 between different experiments.\r\n\r\nAny insights as to why you have picked 26?\r\nSecondly, any specific values for nb_pool & nb_conv? And any input you have on how to adjust the learning rate over time?\r\n\r\nIf I could get a GPU would you be interested in teaming up? :D\r\n\r\n[/quote]\r\n\r\n26 is the number of drivers, so you get each driver in the CV set on each fold, and get some insight into what the trained network is learning. I've just left it like that for the tests above, because that setup is also my best LB score to date. I guess other values should work, although I'd probably default to 13 for neatness of having always same number of drivers in each CV set (I have no idea if this makes a difference)\r\n\r\nI haven't changed nb_pool and nb_conv. I don't have any real insight into other params, mainly I've been trying to force more regularisation into the model (more dropout, trying L2 regularisation and also maxnorm weight constraints), or change the architecture to find something that runs in reasonable time but is a little more accurate. \r\n\r\nI've had some limited success, and I don't think my experience offers much use in the different space of deeper networks using the GPU - something I'm saving for another time. I'm not looking for a team, and probably I will drop out of the competition after I've used it to learn a bit more Python/Keras.\r\n",
    "116417": "I have the same observation, batch size of 16 works better for me. When I try 32 or 64, it is becoming over-fitted somehow.  ",
    "116404": "@Neil\r\nHey I agree running it on a CPU is such a long and slow process.\r\nI also modified ZFTurbo's script and adjusted the epochs to 25 with kfold =5. And batch size of either 48 or 64 between different experiments.\r\n\r\nAny insights as to why you have picked 26?\r\nSecondly, any specific values for nb_pool & nb_conv? And any input you have on how to adjust the learning rate over time?\r\n\r\nIf I could get a GPU would you be interested in teaming up? :D",
    "116154": "Another thing to be aware of is how many epochs you are running. The learning curves for different batch sizes will be different, and the amount of overfit will be different at the same epoch number. So it is possible that people reporting an improvement from smaller batch sizes have in essence tuned their early stopping.",
    "116096": "Is there anyone can explain why theoretically? thx...",
    "116407": "",
    "116149": "",
    "116132": "",
    "116100": "",
    "116066": ""
  }
}