{
  "id": 172033,
  "title": "Large models (Batch size Vs Learning rate)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172033",
  "author_name": "",
  "post_date": "2020-08-03T12:21:56.898363400Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>For someone using PyTorch and GPUs like me, it is not an easy thing to shift directly to TF and TPUs which arguably much faster. </p>\n\n<p>Large models ( e.g. EffB5 ) sometimes cannot fit into a single GPU RAM. However one thing I have learned is that using a smaller batch size can allow the model to hold smaller memory in tradeoff the computation time. </p>\n\n<p>To compensate that, using a smaller learning rate should make the optimiser take smaller (carful) steps of the gradients on these smaller batches. Do you think that is accurate ?</p>\n\n<p>Some of my experiments:\n384x384, effb3, bs 32,  starting lr 5e-4 got 0.9390\n384x384, effb5, bs 16, starting lr 1e-4 got 0.9432</p>",
  "messages": [
    {
      "id": "956334",
      "postDate": "08/03/2020 12:21:56",
      "content": "<p>For someone using PyTorch and GPUs like me, it is not an easy thing to shift directly to TF and TPUs which arguably much faster. </p>\n\n<p>Large models ( e.g. EffB5 ) sometimes cannot fit into a single GPU RAM. However one thing I have learned is that using a smaller batch size can allow the model to hold smaller memory in tradeoff the computation time. </p>\n\n<p>To compensate that, using a smaller learning rate should make the optimiser take smaller (carful) steps of the gradients on these smaller batches. Do you think that is accurate ?</p>\n\n<p>Some of my experiments:\n384x384, effb3, bs 32,  starting lr 5e-4 got 0.9390\n384x384, effb5, bs 16, starting lr 1e-4 got 0.9432</p>",
      "rawMarkdown": "For someone using PyTorch and GPUs like me, it is not an easy thing to shift directly to TF and TPUs which arguably much faster. \n\nLarge models ( e.g. EffB5 ) sometimes cannot fit into a single GPU RAM. However one thing I have learned is that using a smaller batch size can allow the model to hold smaller memory in tradeoff the computation time. \n\nTo compensate that, using a smaller learning rate should make the optimiser take smaller (carful) steps of the gradients on these smaller batches. Do you think that is accurate ?\n\nSome of my experiments:\n384x384, effb3, bs 32,  starting lr 5e-4 got 0.9390\n384x384, effb5, bs 16, starting lr 1e-4 got 0.9432",
      "votes": null
    },
    {
      "id": "956682",
      "postDate": "08/03/2020 17:28:14",
      "content": "<p>I can confirm that, using mixed precision, you can fit your moderately large pytorch model into GPU memory. I managed to test effnet B6 with image size=384 and <strong>batch=24</strong>. Adding mixed precision was just a few lines of code.</p>",
      "rawMarkdown": "I can confirm that, using mixed precision, you can fit your moderately large pytorch model into GPU memory. I managed to test effnet B6 with image size=384 and **batch=24**. Adding mixed precision was just a few lines of code.",
      "votes": null
    },
    {
      "id": "956743",
      "postDate": "08/03/2020 18:39:40",
      "content": "<p>Thank you! I'll definitely check that. </p>",
      "rawMarkdown": "Thank you! I'll definitely check that.",
      "votes": null
    },
    {
      "id": "956860",
      "postDate": "08/03/2020 20:58:42",
      "content": "<p>would you mind sharing the image size you used for your 2 experiments?</p>",
      "rawMarkdown": "would you mind sharing the image size you used for your 2 experiments?",
      "votes": null
    },
    {
      "id": "956900",
      "postDate": "08/03/2020 22:14:17",
      "content": "<p>Both using 384x384 (post updated now)</p>",
      "rawMarkdown": "Both using 384x384 (post updated now)",
      "votes": null
    },
    {
      "id": "957178",
      "postDate": "08/04/2020 05:51:20",
      "content": "<p>You can try gradient accumulation to fit larger batch size. :) </p>\n\n<p>Also, IMO, your observation is correct - larger batch size should have larger learning rate and vice-versa. </p>\n\n<p>Refer here <a href=\"https://arxiv.org/abs/1812.01187\">https://arxiv.org/abs/1812.01187</a></p>",
      "rawMarkdown": "You can try gradient accumulation to fit larger batch size. :) \n\nAlso, IMO, your observation is correct - larger batch size should have larger learning rate and vice-versa. \n\nRefer here https://arxiv.org/abs/1812.01187",
      "votes": null
    },
    {
      "id": "957192",
      "postDate": "08/04/2020 06:11:36",
      "content": "<p>A good rule of thumb once you find a good batch size / learning rate combination, is if you double batch size you can double learning rate. And if you halve batch size you can halve learning rate.</p>\n\n<p>Another trick that people do is increase batch size instead of decaying the learning rate. So for example, instead of halving the learning rate every third epoch, you can instead double the batch size every third epoch (while keeping learning rate constant).</p>",
      "rawMarkdown": "A good rule of thumb once you find a good batch size / learning rate combination, is if you double batch size you can double learning rate. And if you halve batch size you can halve learning rate.\n\nAnother trick that people do is increase batch size instead of decaying the learning rate. So for example, instead of halving the learning rate every third epoch, you can instead double the batch size every third epoch (while keeping learning rate constant).",
      "votes": null
    },
    {
      "id": "957202",
      "postDate": "08/04/2020 06:21:25",
      "content": "<p>how would you increase bs every 3rd epoch? fit the model for 3 epochs double the bs and then refit the model repeatedly?</p>",
      "rawMarkdown": "how would you increase bs every 3rd epoch? fit the model for 3 epochs double the bs and then refit the model repeatedly?",
      "votes": null
    },
    {
      "id": "957753",
      "postDate": "08/04/2020 14:39:37",
      "content": "<p>I've heard it's real easy in PyTorch. You can change batch size during training. </p>\n\n<p>In TensorFlow, you can monitor <code>val_loss</code> or <code>val_auc</code> and stop training with early stopping. Or just stop after a fixed number of epochs like say 3. Then you create a new <code>tf.data.Dataset</code> with larger batch size, then start training again and repeat this process. There is probably a more efficient way in TF but I'm not sure what it is.</p>",
      "rawMarkdown": "I've heard it's real easy in PyTorch. You can change batch size during training. \n\nIn TensorFlow, you can monitor `val_loss` or `val_auc` and stop training with early stopping. Or just stop after a fixed number of epochs like say 3. Then you create a new `tf.data.Dataset` with larger batch size, then start training again and repeat this process. There is probably a more efficient way in TF but I'm not sure what it is.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 956682,
      "author_name": "dunklerwald",
      "author_url": "",
      "post_date": "08/03/2020 17:28:14",
      "content": "<p>I can confirm that, using mixed precision, you can fit your moderately large pytorch model into GPU memory. I managed to test effnet B6 with image size=384 and <strong>batch=24</strong>. Adding mixed precision was just a few lines of code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 956743,
          "author_name": "alhasanabdellatif123",
          "author_url": "",
          "post_date": "08/03/2020 18:39:40",
          "content": "<p>Thank you! I'll definitely check that. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 956860,
      "author_name": "hamonk",
      "author_url": "",
      "post_date": "08/03/2020 20:58:42",
      "content": "<p>would you mind sharing the image size you used for your 2 experiments?</p>",
      "votes": null,
      "replies": [
        {
          "id": 956900,
          "author_name": "alhasanabdellatif123",
          "author_url": "",
          "post_date": "08/03/2020 22:14:17",
          "content": "<p>Both using 384x384 (post updated now)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 957178,
      "author_name": "aroraaman",
      "author_url": "",
      "post_date": "08/04/2020 05:51:20",
      "content": "<p>You can try gradient accumulation to fit larger batch size. :) </p>\n\n<p>Also, IMO, your observation is correct - larger batch size should have larger learning rate and vice-versa. </p>\n\n<p>Refer here <a href=\"https://arxiv.org/abs/1812.01187\">https://arxiv.org/abs/1812.01187</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 957192,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/04/2020 06:11:36",
      "content": "<p>A good rule of thumb once you find a good batch size / learning rate combination, is if you double batch size you can double learning rate. And if you halve batch size you can halve learning rate.</p>\n\n<p>Another trick that people do is increase batch size instead of decaying the learning rate. So for example, instead of halving the learning rate every third epoch, you can instead double the batch size every third epoch (while keeping learning rate constant).</p>",
      "votes": null,
      "replies": [
        {
          "id": 957202,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "08/04/2020 06:21:25",
          "content": "<p>how would you increase bs every 3rd epoch? fit the model for 3 epochs double the bs and then refit the model repeatedly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 957753,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/04/2020 14:39:37",
          "content": "<p>I've heard it's real easy in PyTorch. You can change batch size during training. </p>\n\n<p>In TensorFlow, you can monitor <code>val_loss</code> or <code>val_auc</code> and stop training with early stopping. Or just stop after a fixed number of epochs like say 3. Then you create a new <code>tf.data.Dataset</code> with larger batch size, then start training again and repeat this process. There is probably a more efficient way in TF but I'm not sure what it is.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "956334": "For someone using PyTorch and GPUs like me, it is not an easy thing to shift directly to TF and TPUs which arguably much faster. \n\nLarge models ( e.g. EffB5 ) sometimes cannot fit into a single GPU RAM. However one thing I have learned is that using a smaller batch size can allow the model to hold smaller memory in tradeoff the computation time. \n\nTo compensate that, using a smaller learning rate should make the optimiser take smaller (carful) steps of the gradients on these smaller batches. Do you think that is accurate ?\n\nSome of my experiments:\n384x384, effb3, bs 32,  starting lr 5e-4 got 0.9390\n384x384, effb5, bs 16, starting lr 1e-4 got 0.9432",
    "956682": "I can confirm that, using mixed precision, you can fit your moderately large pytorch model into GPU memory. I managed to test effnet B6 with image size=384 and **batch=24**. Adding mixed precision was just a few lines of code.",
    "956743": "Thank you! I'll definitely check that.",
    "956860": "would you mind sharing the image size you used for your 2 experiments?",
    "956900": "Both using 384x384 (post updated now)",
    "957178": "You can try gradient accumulation to fit larger batch size. :) \n\nAlso, IMO, your observation is correct - larger batch size should have larger learning rate and vice-versa. \n\nRefer here https://arxiv.org/abs/1812.01187",
    "957192": "A good rule of thumb once you find a good batch size / learning rate combination, is if you double batch size you can double learning rate. And if you halve batch size you can halve learning rate.\n\nAnother trick that people do is increase batch size instead of decaying the learning rate. So for example, instead of halving the learning rate every third epoch, you can instead double the batch size every third epoch (while keeping learning rate constant).",
    "957202": "how would you increase bs every 3rd epoch? fit the model for 3 epochs double the bs and then refit the model repeatedly?",
    "957753": "I've heard it's real easy in PyTorch. You can change batch size during training. \n\nIn TensorFlow, you can monitor `val_loss` or `val_auc` and stop training with early stopping. Or just stop after a fixed number of epochs like say 3. Then you create a new `tf.data.Dataset` with larger batch size, then start training again and repeat this process. There is probably a more efficient way in TF but I'm not sure what it is."
  },
  "source": "meta"
}