{
  "id": 217088,
  "title": "Batch Size in TPU",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217088",
  "author_name": "",
  "post_date": "2021-02-05T06:53:47.127950900Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I trained my model with batch size 8 in GPU initially, and then I trained with batch size 64 in TPU because there are 8 cores(so 8 batches for each core).   My thoughts were that giving 8 batches to GPU was same as giving 8 batches for each core in TPU(so in total 64 batches). But the result was a bit different. Is it same giving batches to GPU and TPU like what I've mentioned or is it differenet? and if I use batch size 8 in TPU, is it not efficient?</p>",
  "messages": [
    {
      "id": "1187004",
      "postDate": "02/05/2021 06:53:47",
      "content": "<p>I trained my model with batch size 8 in GPU initially, and then I trained with batch size 64 in TPU because there are 8 cores(so 8 batches for each core).   My thoughts were that giving 8 batches to GPU was same as giving 8 batches for each core in TPU(so in total 64 batches). But the result was a bit different. Is it same giving batches to GPU and TPU like what I've mentioned or is it differenet? and if I use batch size 8 in TPU, is it not efficient?</p>",
      "rawMarkdown": "I trained my model with batch size 8 in GPU initially, and then I trained with batch size 64 in TPU because there are 8 cores(so 8 batches for each core).   My thoughts were that giving 8 batches to GPU was same as giving 8 batches for each core in TPU(so in total 64 batches). But the result was a bit different. Is it same giving batches to GPU and TPU like what I've mentioned or is it differenet? and if I use batch size 8 in TPU, is it not efficient?",
      "votes": null
    },
    {
      "id": "1187229",
      "postDate": "02/05/2021 10:00:40",
      "content": "<p>Hello!</p>\n<p>I'm not a TPU expert, but as it is stated in <strong><a href=\"https://www.kaggle.com/docs/tpu\" target=\"_blank\">Kaggle docs</a></strong> batch size of 16 samples per core (i.e. 128 samples total) is most likely to make all the cores of the TPU busy and thus, the most efficient.</p>\n<p>As for the difference in results - that's pretty normal until you do seed everything and manually initialize each layer of the network you are training. Another reason for this difference is the learning rate which, generally speaking, should not be kept the same if you change the batch size: smaller batches mean more stochasticity while larger batches are more stable. Thus switching from GPU to TPU should be accompanied by a linear increase in the learning rate, which can be taken as 1e-4 times the number of cores when you fine-tuning or <code>0.1 * batch_size / 256</code> when you are training from scratch (the last one is taken from <strong><a href=\"https://arxiv.org/abs/1812.01187\" target=\"_blank\">this paper</a></strong>)</p>",
      "rawMarkdown": "Hello!\n\nI'm not a TPU expert, but as it is stated in **[Kaggle docs](https://www.kaggle.com/docs/tpu)** batch size of 16 samples per core (i.e. 128 samples total) is most likely to make all the cores of the TPU busy and thus, the most efficient.\n\nAs for the difference in results - that's pretty normal until you do seed everything and manually initialize each layer of the network you are training. Another reason for this difference is the learning rate which, generally speaking, should not be kept the same if you change the batch size: smaller batches mean more stochasticity while larger batches are more stable. Thus switching from GPU to TPU should be accompanied by a linear increase in the learning rate, which can be taken as 1e-4 times the number of cores when you fine-tuning or `0.1 * batch_size / 256` when you are training from scratch (the last one is taken from **[this paper](https://arxiv.org/abs/1812.01187)**)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1187229,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "02/05/2021 10:00:40",
      "content": "<p>Hello!</p>\n<p>I'm not a TPU expert, but as it is stated in <strong><a href=\"https://www.kaggle.com/docs/tpu\" target=\"_blank\">Kaggle docs</a></strong> batch size of 16 samples per core (i.e. 128 samples total) is most likely to make all the cores of the TPU busy and thus, the most efficient.</p>\n<p>As for the difference in results - that's pretty normal until you do seed everything and manually initialize each layer of the network you are training. Another reason for this difference is the learning rate which, generally speaking, should not be kept the same if you change the batch size: smaller batches mean more stochasticity while larger batches are more stable. Thus switching from GPU to TPU should be accompanied by a linear increase in the learning rate, which can be taken as 1e-4 times the number of cores when you fine-tuning or <code>0.1 * batch_size / 256</code> when you are training from scratch (the last one is taken from <strong><a href=\"https://arxiv.org/abs/1812.01187\" target=\"_blank\">this paper</a></strong>)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1187004": "I trained my model with batch size 8 in GPU initially, and then I trained with batch size 64 in TPU because there are 8 cores(so 8 batches for each core).   My thoughts were that giving 8 batches to GPU was same as giving 8 batches for each core in TPU(so in total 64 batches). But the result was a bit different. Is it same giving batches to GPU and TPU like what I've mentioned or is it differenet? and if I use batch size 8 in TPU, is it not efficient?",
    "1187229": "Hello!\n\nI'm not a TPU expert, but as it is stated in **[Kaggle docs](https://www.kaggle.com/docs/tpu)** batch size of 16 samples per core (i.e. 128 samples total) is most likely to make all the cores of the TPU busy and thus, the most efficient.\n\nAs for the difference in results - that's pretty normal until you do seed everything and manually initialize each layer of the network you are training. Another reason for this difference is the learning rate which, generally speaking, should not be kept the same if you change the batch size: smaller batches mean more stochasticity while larger batches are more stable. Thus switching from GPU to TPU should be accompanied by a linear increase in the learning rate, which can be taken as 1e-4 times the number of cores when you fine-tuning or `0.1 * batch_size / 256` when you are training from scratch (the last one is taken from **[this paper](https://arxiv.org/abs/1812.01187)**)"
  },
  "source": "meta"
}