{
  "id": 173231,
  "title": "Optimal batch size on TPU for training",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173231",
  "author_name": "Janek Idziak",
  "post_date": "2020-08-08T12:04:36.023000",
  "votes": 5,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I trained multiple models using the TPU lately and I wonder if there is any scientific way to determine optimal batch size. <br>\nI was able to train models with following batch sizes using simple EFnetB5 model : </p>\n<ul>\n<li>256x256 - batch_size = 128 * 8(replicas)</li>\n<li>384x384 - batch_size = 64 * 8(replicas)</li>\n<li>512x512 - batch_size = 32* 8(replicas)</li>\n<li>768x768 - batch_size = 16 * 8(replicas)</li>\n<li>1024x1024 - batch_size = 8 * 8 (replicas) </li>\n</ul>\n<p>Then when I do predictions I can evaluate model using <code>model.predict</code> for approx 4 times bigger batch size. </p>\n<p>What is your experience with batch size while training on TPU?</p>",
  "messages": [
    {
      "id": 962749,
      "postDate": "2020-08-08T12:04:36.023Z",
      "content": "<p>I trained multiple models using the TPU lately and I wonder if there is any scientific way to determine optimal batch size. <br>\nI was able to train models with following batch sizes using simple EFnetB5 model : </p>\n<ul>\n<li>256x256 - batch_size = 128 * 8(replicas)</li>\n<li>384x384 - batch_size = 64 * 8(replicas)</li>\n<li>512x512 - batch_size = 32* 8(replicas)</li>\n<li>768x768 - batch_size = 16 * 8(replicas)</li>\n<li>1024x1024 - batch_size = 8 * 8 (replicas) </li>\n</ul>\n<p>Then when I do predictions I can evaluate model using <code>model.predict</code> for approx 4 times bigger batch size. </p>\n<p>What is your experience with batch size while training on TPU?</p>",
      "rawMarkdown": "I trained multiple models using the TPU lately and I wonder if there is any scientific way to determine optimal batch size. \nI was able to train models with following batch sizes using simple EFnetB5 model : \n\n- 256x256 - batch_size = 128 * 8(replicas)\n- 384x384 - batch_size = 64 * 8(replicas)\n- 512x512 - batch_size = 32* 8(replicas)\n- 768x768 - batch_size = 16 * 8(replicas)\n- 1024x1024 - batch_size = 8 * 8 (replicas) \n\nThen when I do predictions I can evaluate model using `model.predict` for approx 4 times bigger batch size. \n\nWhat is your experience with batch size while training on TPU?",
      "votes": 5
    },
    {
      "id": 963069,
      "postDate": "2020-08-08T16:32:09.183Z",
      "content": "<p>I've gotten away with 24 * 8 for 768x768</p>",
      "rawMarkdown": "I've gotten away with 24 * 8 for 768x768",
      "votes": 1,
      "replies": [
        {
          "id": 964632,
          "postDate": "2020-08-10T04:14:51.343Z",
          "content": "<p>What is the model size? </p>",
          "rawMarkdown": "What is the model size? "
        },
        {
          "id": 965884,
          "postDate": "2020-08-10T23:48:30.657Z",
          "content": "<p>B7 using Chris's triple stratified kernel</p>",
          "rawMarkdown": "B7 using Chris's triple stratified kernel"
        },
        {
          "id": 967698,
          "postDate": "2020-08-12T12:44:11.910Z",
          "content": "<p>This is wired, I tried to train B4 B5 noisy-student and B4 B5 image net (4 models at once) and I can fit only batch size of 3 or 2</p>",
          "rawMarkdown": "This is wired, I tried to train B4 B5 noisy-student and B4 B5 image net (4 models at once) and I can fit only batch size of 3 or 2\n"
        }
      ]
    },
    {
      "id": 969458,
      "postDate": "2020-08-13T17:42:58.603Z",
      "content": "<p>Important information here: <br>\n Kaggle's TPU are v3 and Colab are v2, so you may notice that on Kaggle, you may use larger batches and it will also run the epochs faster!</p>",
      "rawMarkdown": "Important information here: \n Kaggle's TPU are v3 and Colab are v2, so you may notice that on Kaggle, you may use larger batches and it will also run the epochs faster!"
    },
    {
      "id": 964173,
      "postDate": "2020-08-09T16:41:43.200Z",
      "content": "<p>Today I tried to train model on <strong>768</strong> b4 and b5 (two of each, one with imagenet and one with noisy student) weights in the model together <strong>with metadata</strong> and ** it failed for batch size 16**. Now trying to re run for batch size of 8. </p>\n<p>(<a href=\"https://www.kaggle.com/janidziak/crazy-incredible-tpu-noisy-and-imgn-192-metadata\" target=\"_blank\">here is sample of notebook that I used for training</a>) </p>",
      "rawMarkdown": "Today I tried to train model on **768** b4 and b5 (two of each, one with imagenet and one with noisy student) weights in the model together **with metadata** and ** it failed for batch size 16**. Now trying to re run for batch size of 8. \n\n ([here is sample of notebook that I used for training](https://www.kaggle.com/janidziak/crazy-incredible-tpu-noisy-and-imgn-192-metadata)) \n"
    },
    {
      "id": 962779,
      "postDate": "2020-08-08T12:40:36.380Z",
      "content": "<p>Are your numbers for Tensorflow or Pytorch? Because as I posted yesterday I had big problems with memory and 384x384 </p>",
      "rawMarkdown": "Are your numbers for Tensorflow or Pytorch? Because as I posted yesterday I had big problems with memory and 384x384 ",
      "replies": [
        {
          "id": 963066,
          "postDate": "2020-08-08T16:30:00.507Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 963096,
          "postDate": "2020-08-08T17:02:34.200Z",
          "content": "<p>I suppose it does, Abhishek has created this notebook: <a href=\"https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\" target=\"_blank\">https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu</a> <br>\nNever tried yet, let me know how it works</p>",
          "rawMarkdown": "I suppose it does, Abhishek has created this notebook: https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu \nNever tried yet, let me know how it works"
        },
        {
          "id": 963099,
          "postDate": "2020-08-08T17:03:12.950Z",
          "content": "<p>I only tried tensorflow for now. What were your problems with 384? </p>",
          "rawMarkdown": "I only tried tensorflow for now. What were your problems with 384? "
        },
        {
          "id": 963254,
          "postDate": "2020-08-08T19:51:29.660Z",
          "content": "<p>he uses 224 not 384</p>",
          "rawMarkdown": "he uses 224 not 384"
        },
        {
          "id": 963571,
          "postDate": "2020-08-09T06:12:16.473Z",
          "content": "<p><a href=\"https://www.kaggle.com/jacekpoplawski\" target=\"_blank\">@jacekpoplawski</a> can you post the link to your discussion? Are you using tensorflow or pytorch yourself? </p>",
          "rawMarkdown": "@jacekpoplawski can you post the link to your discussion? Are you using tensorflow or pytorch yourself? "
        },
        {
          "id": 963818,
          "postDate": "2020-08-09T10:20:23.597Z",
          "content": "<p>I have problem with pytorch, TPU and 384x384 on b3. Will try to debug on colab today.<br>\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173117\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173117</a></p>",
          "rawMarkdown": "I have problem with pytorch, TPU and 384x384 on b3. Will try to debug on colab today.\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173117"
        }
      ]
    },
    {
      "id": 963077,
      "postDate": "2020-08-08T16:42:15.800Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 963069,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-08-08T16:32:09.183000",
      "content": "<p>I've gotten away with 24 * 8 for 768x768</p>",
      "votes": 1,
      "replies": [
        {
          "id": 964632,
          "author_name": "Janek Idziak",
          "author_url": "",
          "post_date": "2020-08-10T04:14:51.343000",
          "content": "<p>What is the model size? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 965884,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-08-10T23:48:30.657000",
          "content": "<p>B7 using Chris's triple stratified kernel</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 967698,
          "author_name": "Janek Idziak",
          "author_url": "",
          "post_date": "2020-08-12T12:44:11.910000",
          "content": "<p>This is wired, I tried to train B4 B5 noisy-student and B4 B5 image net (4 models at once) and I can fit only batch size of 3 or 2</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 969458,
      "author_name": "Janek Idziak",
      "author_url": "",
      "post_date": "2020-08-13T17:42:58.603000",
      "content": "<p>Important information here: <br>\n Kaggle's TPU are v3 and Colab are v2, so you may notice that on Kaggle, you may use larger batches and it will also run the epochs faster!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 964173,
      "author_name": "Janek Idziak",
      "author_url": "",
      "post_date": "2020-08-09T16:41:43.200000",
      "content": "<p>Today I tried to train model on <strong>768</strong> b4 and b5 (two of each, one with imagenet and one with noisy student) weights in the model together <strong>with metadata</strong> and ** it failed for batch size 16**. Now trying to re run for batch size of 8. </p>\n<p>(<a href=\"https://www.kaggle.com/janidziak/crazy-incredible-tpu-noisy-and-imgn-192-metadata\" target=\"_blank\">here is sample of notebook that I used for training</a>) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 962779,
      "author_name": "Jacek Poplawski",
      "author_url": "",
      "post_date": "2020-08-08T12:40:36.380000",
      "content": "<p>Are your numbers for Tensorflow or Pytorch? Because as I posted yesterday I had big problems with memory and 384x384 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 963066,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-08T16:30:00.507000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 963096,
          "author_name": "Janek Idziak",
          "author_url": "",
          "post_date": "2020-08-08T17:02:34.200000",
          "content": "<p>I suppose it does, Abhishek has created this notebook: <a href=\"https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\" target=\"_blank\">https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu</a> <br>\nNever tried yet, let me know how it works</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 963099,
          "author_name": "Janek Idziak",
          "author_url": "",
          "post_date": "2020-08-08T17:03:12.950000",
          "content": "<p>I only tried tensorflow for now. What were your problems with 384? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 963254,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-08-08T19:51:29.660000",
          "content": "<p>he uses 224 not 384</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 963571,
          "author_name": "Janek Idziak",
          "author_url": "",
          "post_date": "2020-08-09T06:12:16.473000",
          "content": "<p><a href=\"https://www.kaggle.com/jacekpoplawski\" target=\"_blank\">@jacekpoplawski</a> can you post the link to your discussion? Are you using tensorflow or pytorch yourself? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 963818,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-08-09T10:20:23.597000",
          "content": "<p>I have problem with pytorch, TPU and 384x384 on b3. Will try to debug on colab today.<br>\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173117\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173117</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 963077,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-08T16:42:15.800000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "962749": "I trained multiple models using the TPU lately and I wonder if there is any scientific way to determine optimal batch size. \nI was able to train models with following batch sizes using simple EFnetB5 model : \n\n- 256x256 - batch_size = 128 * 8(replicas)\n- 384x384 - batch_size = 64 * 8(replicas)\n- 512x512 - batch_size = 32* 8(replicas)\n- 768x768 - batch_size = 16 * 8(replicas)\n- 1024x1024 - batch_size = 8 * 8 (replicas) \n\nThen when I do predictions I can evaluate model using `model.predict` for approx 4 times bigger batch size. \n\nWhat is your experience with batch size while training on TPU?",
    "963069": "I've gotten away with 24 * 8 for 768x768",
    "969458": "Important information here: \n Kaggle's TPU are v3 and Colab are v2, so you may notice that on Kaggle, you may use larger batches and it will also run the epochs faster!",
    "964173": "Today I tried to train model on **768** b4 and b5 (two of each, one with imagenet and one with noisy student) weights in the model together **with metadata** and ** it failed for batch size 16**. Now trying to re run for batch size of 8. \n\n ([here is sample of notebook that I used for training](https://www.kaggle.com/janidziak/crazy-incredible-tpu-noisy-and-imgn-192-metadata)) \n",
    "962779": "Are your numbers for Tensorflow or Pytorch? Because as I posted yesterday I had big problems with memory and 384x384 ",
    "963077": ""
  }
}