{
  "id": 168366,
  "title": "how long is your epoch on TPU?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168366",
  "author_name": "",
  "post_date": "2020-07-20T10:10:48.854465900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am testing different efficientnets on colab on TPU and I am confused.\nEach one between b0 and b3 has same duration, between 260 and 300 seconds.\n(I use 4/5 train and 1/5 valid)\nI wonder maybe I am doing something incorrectly and there is some bottleneck like CPU.\nCan you say your epoch duration on TPU?</p>",
  "messages": [
    {
      "id": "936557",
      "postDate": "07/20/2020 10:10:48",
      "content": "<p>I am testing different efficientnets on colab on TPU and I am confused.\nEach one between b0 and b3 has same duration, between 260 and 300 seconds.\n(I use 4/5 train and 1/5 valid)\nI wonder maybe I am doing something incorrectly and there is some bottleneck like CPU.\nCan you say your epoch duration on TPU?</p>",
      "rawMarkdown": "I am testing different efficientnets on colab on TPU and I am confused.\nEach one between b0 and b3 has same duration, between 260 and 300 seconds.\n(I use 4/5 train and 1/5 valid)\nI wonder maybe I am doing something incorrectly and there is some bottleneck like CPU.\nCan you say your epoch duration on TPU?",
      "votes": null
    },
    {
      "id": "936561",
      "postDate": "07/20/2020 10:13:39",
      "content": "<p>I tested it once, I think it was about 3-4 minutes per epoch with 16 batch size.</p>",
      "rawMarkdown": "I tested it once, I think it was about 3-4 minutes per epoch with 16 batch size.",
      "votes": null
    },
    {
      "id": "936895",
      "postDate": "07/20/2020 15:18:23",
      "content": "<p>Run time of epoch is function of the model, the size of images being used and the batch size.  You did not mention image size or batch so not possible to say if 300 seconds is good or bad :)</p>\n\n<p>If your doing large image sizes than changing the model might not cause much of a time change as the image size is the major component of the speed.</p>\n\n<p>If you have not moved the batch size to the almost maximum, than again changing the model might not cause a noticeable speed change.  </p>\n\n<p>On local PC it's not too painful to play with batch size until OOM and than back off a bit to get more epoch speed out of a model.  On Kaggle it's a tiny bit of pain to increase batch until OOM to max out the speed.  On all of the shared kernels in this competition so far I have found that I can increase the batch on almost all to get a speed improvement.  </p>\n\n<p>Feeding a TPU seems a bit more of a trick than feeding a GPU but it still starts with maximize the batch size for image size and model if your interest is epoch speed.</p>",
      "rawMarkdown": "Run time of epoch is function of the model, the size of images being used and the batch size.  You did not mention image size or batch so not possible to say if 300 seconds is good or bad :)\n\nIf your doing large image sizes than changing the model might not cause much of a time change as the image size is the major component of the speed.\n\nIf you have not moved the batch size to the almost maximum, than again changing the model might not cause a noticeable speed change.  \n\nOn local PC it's not too painful to play with batch size until OOM and than back off a bit to get more epoch speed out of a model.  On Kaggle it's a tiny bit of pain to increase batch until OOM to max out the speed.  On all of the shared kernels in this competition so far I have found that I can increase the batch on almost all to get a speed improvement.  \n\nFeeding a TPU seems a bit more of a trick than feeding a GPU but it still starts with maximize the batch size for image size and model if your interest is epoch speed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 936561,
      "author_name": "amneves",
      "author_url": "",
      "post_date": "07/20/2020 10:13:39",
      "content": "<p>I tested it once, I think it was about 3-4 minutes per epoch with 16 batch size.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 936895,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/20/2020 15:18:23",
      "content": "<p>Run time of epoch is function of the model, the size of images being used and the batch size.  You did not mention image size or batch so not possible to say if 300 seconds is good or bad :)</p>\n\n<p>If your doing large image sizes than changing the model might not cause much of a time change as the image size is the major component of the speed.</p>\n\n<p>If you have not moved the batch size to the almost maximum, than again changing the model might not cause a noticeable speed change.  </p>\n\n<p>On local PC it's not too painful to play with batch size until OOM and than back off a bit to get more epoch speed out of a model.  On Kaggle it's a tiny bit of pain to increase batch until OOM to max out the speed.  On all of the shared kernels in this competition so far I have found that I can increase the batch on almost all to get a speed improvement.  </p>\n\n<p>Feeding a TPU seems a bit more of a trick than feeding a GPU but it still starts with maximize the batch size for image size and model if your interest is epoch speed.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "936557": "I am testing different efficientnets on colab on TPU and I am confused.\nEach one between b0 and b3 has same duration, between 260 and 300 seconds.\n(I use 4/5 train and 1/5 valid)\nI wonder maybe I am doing something incorrectly and there is some bottleneck like CPU.\nCan you say your epoch duration on TPU?",
    "936561": "I tested it once, I think it was about 3-4 minutes per epoch with 16 batch size.",
    "936895": "Run time of epoch is function of the model, the size of images being used and the batch size.  You did not mention image size or batch so not possible to say if 300 seconds is good or bad :)\n\nIf your doing large image sizes than changing the model might not cause much of a time change as the image size is the major component of the speed.\n\nIf you have not moved the batch size to the almost maximum, than again changing the model might not cause a noticeable speed change.  \n\nOn local PC it's not too painful to play with batch size until OOM and than back off a bit to get more epoch speed out of a model.  On Kaggle it's a tiny bit of pain to increase batch until OOM to max out the speed.  On all of the shared kernels in this competition so far I have found that I can increase the batch on almost all to get a speed improvement.  \n\nFeeding a TPU seems a bit more of a trick than feeding a GPU but it still starts with maximize the batch size for image size and model if your interest is epoch speed."
  },
  "source": "meta"
}