{
  "id": 168525,
  "title": "survey of TPU usage",
  "url": "/competitions/alaska2-image-steganalysis/discussion/168525",
  "author_name": "hengck23",
  "post_date": "2020-07-21T02:02:46.109000",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>it becomes a problem for me in this competition as i am using pytorch + local desk PC with multi GPUs. It is almost impossible to train large network like efficientnetb5,6,7.</p>\n\n<p>I would like to know the speed of using TPU on such for large networks, e.g. time per epoch.\n- any time difference for colab, colab pro, kaggle kernel?\n- how does pytorch xla performs? is tpu only for tf/keras?\n- any performance difference between multi-gpu and tpu? (e.g. multi-gpu has the problem of batchnorm synchronization, is this an issue in replicas in TPU?)</p>",
  "messages": [
    {
      "id": 937405,
      "postDate": "2020-07-21T02:02:46.110Z",
      "content": "<p>it becomes a problem for me in this competition as i am using pytorch + local desk PC with multi GPUs. It is almost impossible to train large network like efficientnetb5,6,7.</p>\n\n<p>I would like to know the speed of using TPU on such for large networks, e.g. time per epoch.\n- any time difference for colab, colab pro, kaggle kernel?\n- how does pytorch xla performs? is tpu only for tf/keras?\n- any performance difference between multi-gpu and tpu? (e.g. multi-gpu has the problem of batchnorm synchronization, is this an issue in replicas in TPU?)</p>",
      "rawMarkdown": "it becomes a problem for me in this competition as i am using pytorch + local desk PC with multi GPUs. It is almost impossible to train large network like efficientnetb5,6,7.\n\nI would like to know the speed of using TPU on such for large networks, e.g. time per epoch.\n- any time difference for colab, colab pro, kaggle kernel?\n- how does pytorch xla performs? is tpu only for tf/keras?\n- any performance difference between multi-gpu and tpu? (e.g. multi-gpu has the problem of batchnorm synchronization, is this an issue in replicas in TPU?)",
      "votes": 4
    },
    {
      "id": 941734,
      "postDate": "2020-07-23T11:33:42.847Z",
      "content": "<p>This is w/ 1950x and 3 x 2080 Ti. Pytorch 1.7, fp16 (autocast), and DataDistributed:</p>\n\n<ul>\n<li>B0: 18 mins per epoch</li>\n<li>B4: ~1hr per epoch</li>\n<li>B7: ~3hr per epoch</li>\n</ul>\n\n<p>TPUv3 is powerful but also expensive: $2.4/hr pre-emptible (in my experience it doesn't last more than 12 hrs). Training 48 hours of would be 48*2.4 = $115. </p>\n\n<p>I am surprised B7 is less than 1 hr/epoch on TPU.</p>",
      "rawMarkdown": "This is w/ 1950x and 3 x 2080 Ti. Pytorch 1.7, fp16 (autocast), and DataDistributed:\n\n* B0: 18 mins per epoch\n* B4: ~1hr per epoch\n* B7: ~3hr per epoch\n\nTPUv3 is powerful but also expensive: $2.4/hr pre-emptible (in my experience it doesn't last more than 12 hrs). Training 48 hours of would be 48*2.4 = $115. \n\nI am surprised B7 is less than 1 hr/epoch on TPU."
    },
    {
      "id": 938080,
      "postDate": "2020-07-21T10:34:53.310Z",
      "content": "<p>TPU is very good in TensorFlow/Keras once you set up your tfrecord data and load it into the GCS bucket. You may need a minimum of 2 cores and at least 8 or 16 RAM. Based on my experience in this competition, EfficientNetB7 requires 30-35 minutes of training time with minimum data augmentation. In the case of Pytorch using TPU in image competition, it is highly recommended to use at least 16 cores and higher RAM for the data loader.</p>",
      "rawMarkdown": "TPU is very good in TensorFlow/Keras once you set up your tfrecord data and load it into the GCS bucket. You may need a minimum of 2 cores and at least 8 or 16 RAM. Based on my experience in this competition, EfficientNetB7 requires 30-35 minutes of training time with minimum data augmentation. In the case of Pytorch using TPU in image competition, it is highly recommended to use at least 16 cores and higher RAM for the data loader.",
      "replies": [
        {
          "id": 938090,
          "postDate": "2020-07-21T10:40:23.143Z",
          "content": "<p><a href=\"/projdev\">@projdev</a> \nthanks.</p>\n\n<p>Seems that i need to think of a way to bridge/convert tf model and pytorch model if i want to stick to pytorch.</p>",
          "rawMarkdown": "@projdev \nthanks.\n\nSeems that i need to think of a way to bridge/convert tf model and pytorch model if i want to stick to pytorch."
        }
      ]
    },
    {
      "id": 938038,
      "postDate": "2020-07-21T10:02:09.380Z",
      "content": "<p>seems that there are still many issues, kaggle can crowd fund kagglers to solve kaggler issues like this (kickstarter data science) ....</p>\n\n<p>some kaggler just want to do competitions and willing to pay for setup, framework, etc solution</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F03307bc0d131ba076d3c58860e74e73c%2FSelection_034.png?generation=1595325712980186&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "seems that there are still many issues, kaggle can crowd fund kagglers to solve kaggler issues like this (kickstarter data science) ....\n\nsome kaggler just want to do competitions and willing to pay for setup, framework, etc solution\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F03307bc0d131ba076d3c58860e74e73c%2FSelection_034.png?generation=1595325712980186&amp;alt=media)\n"
    },
    {
      "id": 937971,
      "postDate": "2020-07-21T09:23:00.700Z",
      "content": "<p>any idea if the latest Nvidia Ampere RTX 3080 Ti will be comparable to TPU?</p>\n\n<p>I am currently using my 4x1080Ti development environment for kaggle hobby. Now i need to decide to upgrade my system or go for the TPU cloud ... It is a tough choice ...</p>",
      "rawMarkdown": "any idea if the latest Nvidia Ampere RTX 3080 Ti will be comparable to TPU?\n\nI am currently using my 4x1080Ti development environment for kaggle hobby. Now i need to decide to upgrade my system or go for the TPU cloud ... It is a tough choice ..."
    },
    {
      "id": 937517,
      "postDate": "2020-07-21T04:15:38.277Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> what was your bottleneck for local development? I was using 4x1080Ti + 1950X CPU as my main development setup. It was loud and hot, with B6 training time around 1h30m per epoch. </p>\n\n<p>I've tried to train B7 on TPU in Kaggle. In general I was able to fit batch 32 on single TPU. On v3-8 effective batch size was 128, but the problem was - small amount of host memory (16Gb) and only 2 CPU cores. \nData IO was a clear bottleneck here and most of the time TPU were waiting for the data to become available.\nThat's my experience with TPU so far.</p>",
      "rawMarkdown": "@hengck23 what was your bottleneck for local development? I was using 4x1080Ti + 1950X CPU as my main development setup. It was loud and hot, with B6 training time around 1h30m per epoch. \n\nI've tried to train B7 on TPU in Kaggle. In general I was able to fit batch 32 on single TPU. On v3-8 effective batch size was 128, but the problem was - small amount of host memory (16Gb) and only 2 CPU cores. \nData IO was a clear bottleneck here and most of the time TPU were waiting for the data to become available.\nThat's my experience with TPU so far.\n\n",
      "replies": [
        {
          "id": 938060,
          "postDate": "2020-07-21T10:15:27.327Z",
          "content": "<p><a href=\"/bloodaxe\">@bloodaxe</a> \nThanks for the information</p>\n\n<p>I have a 4x 1080Ti, Intel(R) Core(TM) i7-6850K CPU @ 3.60GHz (6 cores), 128 RAM + 256 SSD too.\nI train single model per GPU and the timing are:\n```</p>\n\n<pre><code>    torch.__version__              = 1.6.0.dev20200516+cu101\n    torch.version.cuda             = 10.1\n    torch.backends.cudnn.version() = 7603\n    os['CUDA_VISIBLE_DEVICES']     = 3\n    torch.cuda.device_count()      = 1\n    torch.cuda.get_device_properties() = (name='GeForce GTX 1080 Ti', major=6, minor=1, total_memory=11178MB, multi_processor_count=28)\n</code></pre>\n\n<p>train_dataset : \n    60000 image ids (i.e. 1 epoch = 60000x4 images (1 cover,3 stego), each is 512x512x3)</p>\n\n<p>efficienetb0 (mixed precision): 1 epoch = 1hr 30 min (batch size = 32) \nefficienetb1 (mixed precision): 1 epoch = 2hr 18 min (batch size = 24) <br>\nefficienetb2 (mixed precision): 1 epoch = 2hr 30 min (batch size = 16) <br>\nefficienetb3 (mixed precision): 1 epoch = 3hr   (batch size = 20) \nefficienetb4 (mixed precision): 1 epoch = 4hr 20 min  (batch size = 10) </p>\n\n<p>gpu utilization  from smi-nvidia is about 90% to 100% (haven't got time optimize my data loader yet)\n``` </p>",
          "rawMarkdown": "@bloodaxe \nThanks for the information\n\nI have a 4x 1080Ti, Intel(R) Core(TM) i7-6850K CPU @ 3.60GHz (6 cores), 128 RAM + 256 SSD too.\nI train single model per GPU and the timing are:\n```\n\n\t\ttorch.__version__              = 1.6.0.dev20200516+cu101\n\t\ttorch.version.cuda             = 10.1\n\t\ttorch.backends.cudnn.version() = 7603\n\t\tos['CUDA_VISIBLE_DEVICES']     = 3\n\t\ttorch.cuda.device_count()      = 1\n\t\ttorch.cuda.get_device_properties() = (name='GeForce GTX 1080 Ti', major=6, minor=1, total_memory=11178MB, multi_processor_count=28)\n\n\n\ntrain_dataset : \n\t60000 image ids (i.e. 1 epoch = 60000x4 images (1 cover,3 stego), each is 512x512x3)\n\t \nefficienetb0 (mixed precision): 1 epoch = 1hr 30 min (batch size = 32) \nefficienetb1 (mixed precision): 1 epoch = 2hr 18 min (batch size = 24)   \nefficienetb2 (mixed precision): 1 epoch = 2hr 30 min (batch size = 16)   \nefficienetb3 (mixed precision): 1 epoch = 3hr   (batch size = 20) \nefficienetb4 (mixed precision): 1 epoch = 4hr 20 min  (batch size = 10) \n\n\ngpu utilization  from smi-nvidia is about 90% to 100% (haven't got time optimize my data loader yet)\n``` \n"
        },
        {
          "id": 938192,
          "postDate": "2020-07-21T11:50:58.547Z",
          "content": "<p>That is clearly some IO bottleneck. My setup is same 4x1080Ti, but I have 1-st gen 1950x threadripper with 32 cores. I believe this CPU is a must-have to multi-gpu setup. Also, I store training data on NVME SSD in order to ensure fastest reading time. I also observe that DDP gives about 10-15% training time reduction compared to DP mode (I was also using fp16 using apex. torch.amp seems to be even faster). </p>",
          "rawMarkdown": "That is clearly some IO bottleneck. My setup is same 4x1080Ti, but I have 1-st gen 1950x threadripper with 32 cores. I believe this CPU is a must-have to multi-gpu setup. Also, I store training data on NVME SSD in order to ensure fastest reading time. I also observe that DDP gives about 10-15% training time reduction compared to DP mode (I was also using fp16 using apex. torch.amp seems to be even faster). ",
          "votes": 1
        },
        {
          "id": 938204,
          "postDate": "2020-07-21T11:56:32.923Z",
          "content": "<p>\"That is clearly some IO bottleneck .... 1950x threadripper with 32 cores.\"</p>\n\n<p>you are correct. it is a mistake to buy a low-core cpu when i builded the system. </p>",
          "rawMarkdown": "\"That is clearly some IO bottleneck .... 1950x threadripper with 32 cores.\"\n\nyou are correct. it is a mistake to buy a low-core cpu when i builded the system. \n"
        }
      ]
    },
    {
      "id": 937505,
      "postDate": "2020-07-21T04:03:23.543Z",
      "content": "<p>One epoch in 15~20 mins for Effnet B0/B1, and 30~45 mins for B6/B7. The detail is <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168537\">here</a>. </p>",
      "rawMarkdown": "One epoch in 15~20 mins for Effnet B0/B1, and 30~45 mins for B6/B7. The detail is [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168537). "
    },
    {
      "id": 937461,
      "postDate": "2020-07-21T03:29:30Z",
      "content": "<p>I trained models entirely on TPU with keras. ef7 takes about 30 min / epoch. Not sure why, but Keras performs so much worse than same model with pytorch from public notebook</p>",
      "rawMarkdown": "I trained models entirely on TPU with keras. ef7 takes about 30 min / epoch. Not sure why, but Keras performs so much worse than same model with pytorch from public notebook"
    },
    {
      "id": 937460,
      "postDate": "2020-07-21T03:29:11.193Z",
      "content": "<p>Hi, TPU is really fast than GPU!<br>\nFor you questions, </p>\n<ol>\n<li>TPU takes 20 min for full alaska2 dataset per epoch, including validation. I make my <a href=\"https://www.kaggle.com/leonshangguan/fullsetfold0final\" target=\"_blank\">kernel</a> and <a href=\"https://www.kaggle.com/leonshangguan/alaska2train95baseline\" target=\"_blank\">kernel</a> public here which you can have a reference. Another benefits is that we can train with a large batch size with TPU.</li>\n<li>For tensorflow, no difference, but if you are using pytorch, kaggle kernel only supports old version pytorch than colab.</li>\n<li>TPU supports both pytorch and tf/keras, but based on my experience, pytorch xla has a lot of bugs, and with only few documents, while tf/keras is is much more user friendly. </li>\n<li>I am not sure if there is any performance difference between multi-gpu and tpu, didn't test it before.</li>\n</ol>\n<p>Together with my teammates we tried xla by modify public kernel first, <a href=\"https://www.kaggle.com/zihezheng/pytorch-tpu-transfer-learning-baseline\" target=\"_blank\">kernel here</a>, but faced with a lot of problems, even with the same code, sometimes we can run successfully but sometimes not(don't know why).</p>",
      "rawMarkdown": "Hi, TPU is really fast than GPU!\nFor you questions, \n1. TPU takes 20 min for full alaska2 dataset per epoch, including validation. I make my [kernel](https://www.kaggle.com/leonshangguan/fullsetfold0final) and [kernel](https://www.kaggle.com/leonshangguan/alaska2train95baseline) public here which you can have a reference. Another benefits is that we can train with a large batch size with TPU.\n2. For tensorflow, no difference, but if you are using pytorch, kaggle kernel only supports old version pytorch than colab.\n3. TPU supports both pytorch and tf/keras, but based on my experience, pytorch xla has a lot of bugs, and with only few documents, while tf/keras is is much more user friendly. \n4. I am not sure if there is any performance difference between multi-gpu and tpu, didn't test it before.\n\nTogether with my teammates we tried xla by modify public kernel first, [kernel here](https://www.kaggle.com/zihezheng/pytorch-tpu-transfer-learning-baseline), but faced with a lot of problems, even with the same code, sometimes we can run successfully but sometimes not(don't know why)."
    },
    {
      "id": 937524,
      "postDate": "2020-07-21T04:19:17.957Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 941734,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2020-07-23T11:33:42.847000",
      "content": "<p>This is w/ 1950x and 3 x 2080 Ti. Pytorch 1.7, fp16 (autocast), and DataDistributed:</p>\n\n<ul>\n<li>B0: 18 mins per epoch</li>\n<li>B4: ~1hr per epoch</li>\n<li>B7: ~3hr per epoch</li>\n</ul>\n\n<p>TPUv3 is powerful but also expensive: $2.4/hr pre-emptible (in my experience it doesn't last more than 12 hrs). Training 48 hours of would be 48*2.4 = $115. </p>\n\n<p>I am surprised B7 is less than 1 hr/epoch on TPU.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 938080,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2020-07-21T10:34:53.310000",
      "content": "<p>TPU is very good in TensorFlow/Keras once you set up your tfrecord data and load it into the GCS bucket. You may need a minimum of 2 cores and at least 8 or 16 RAM. Based on my experience in this competition, EfficientNetB7 requires 30-35 minutes of training time with minimum data augmentation. In the case of Pytorch using TPU in image competition, it is highly recommended to use at least 16 cores and higher RAM for the data loader.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 938090,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-07-21T10:40:23.143000",
          "content": "<p><a href=\"/projdev\">@projdev</a> \nthanks.</p>\n\n<p>Seems that i need to think of a way to bridge/convert tf model and pytorch model if i want to stick to pytorch.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 938038,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-07-21T10:02:09.380000",
      "content": "<p>seems that there are still many issues, kaggle can crowd fund kagglers to solve kaggler issues like this (kickstarter data science) ....</p>\n\n<p>some kaggler just want to do competitions and willing to pay for setup, framework, etc solution</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F03307bc0d131ba076d3c58860e74e73c%2FSelection_034.png?generation=1595325712980186&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 937971,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-07-21T09:23:00.700000",
      "content": "<p>any idea if the latest Nvidia Ampere RTX 3080 Ti will be comparable to TPU?</p>\n\n<p>I am currently using my 4x1080Ti development environment for kaggle hobby. Now i need to decide to upgrade my system or go for the TPU cloud ... It is a tough choice ...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 937517,
      "author_name": "Eugene Khvedchenya",
      "author_url": "",
      "post_date": "2020-07-21T04:15:38.277000",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> what was your bottleneck for local development? I was using 4x1080Ti + 1950X CPU as my main development setup. It was loud and hot, with B6 training time around 1h30m per epoch. </p>\n\n<p>I've tried to train B7 on TPU in Kaggle. In general I was able to fit batch 32 on single TPU. On v3-8 effective batch size was 128, but the problem was - small amount of host memory (16Gb) and only 2 CPU cores. \nData IO was a clear bottleneck here and most of the time TPU were waiting for the data to become available.\nThat's my experience with TPU so far.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 938060,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-07-21T10:15:27.327000",
          "content": "<p><a href=\"/bloodaxe\">@bloodaxe</a> \nThanks for the information</p>\n\n<p>I have a 4x 1080Ti, Intel(R) Core(TM) i7-6850K CPU @ 3.60GHz (6 cores), 128 RAM + 256 SSD too.\nI train single model per GPU and the timing are:\n```</p>\n\n<pre><code>    torch.__version__              = 1.6.0.dev20200516+cu101\n    torch.version.cuda             = 10.1\n    torch.backends.cudnn.version() = 7603\n    os['CUDA_VISIBLE_DEVICES']     = 3\n    torch.cuda.device_count()      = 1\n    torch.cuda.get_device_properties() = (name='GeForce GTX 1080 Ti', major=6, minor=1, total_memory=11178MB, multi_processor_count=28)\n</code></pre>\n\n<p>train_dataset : \n    60000 image ids (i.e. 1 epoch = 60000x4 images (1 cover,3 stego), each is 512x512x3)</p>\n\n<p>efficienetb0 (mixed precision): 1 epoch = 1hr 30 min (batch size = 32) \nefficienetb1 (mixed precision): 1 epoch = 2hr 18 min (batch size = 24) <br>\nefficienetb2 (mixed precision): 1 epoch = 2hr 30 min (batch size = 16) <br>\nefficienetb3 (mixed precision): 1 epoch = 3hr   (batch size = 20) \nefficienetb4 (mixed precision): 1 epoch = 4hr 20 min  (batch size = 10) </p>\n\n<p>gpu utilization  from smi-nvidia is about 90% to 100% (haven't got time optimize my data loader yet)\n``` </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938192,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-07-21T11:50:58.547000",
          "content": "<p>That is clearly some IO bottleneck. My setup is same 4x1080Ti, but I have 1-st gen 1950x threadripper with 32 cores. I believe this CPU is a must-have to multi-gpu setup. Also, I store training data on NVME SSD in order to ensure fastest reading time. I also observe that DDP gives about 10-15% training time reduction compared to DP mode (I was also using fp16 using apex. torch.amp seems to be even faster). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938204,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-07-21T11:56:32.923000",
          "content": "<p>\"That is clearly some IO bottleneck .... 1950x threadripper with 32 cores.\"</p>\n\n<p>you are correct. it is a mistake to buy a low-core cpu when i builded the system. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 937505,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2020-07-21T04:03:23.543000",
      "content": "<p>One epoch in 15~20 mins for Effnet B0/B1, and 30~45 mins for B6/B7. The detail is <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168537\">here</a>. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 937461,
      "author_name": "Gold Retriever",
      "author_url": "",
      "post_date": "2020-07-21T03:29:30",
      "content": "<p>I trained models entirely on TPU with keras. ef7 takes about 30 min / epoch. Not sure why, but Keras performs so much worse than same model with pytorch from public notebook</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 937460,
      "author_name": "Zhongkai Shangguan",
      "author_url": "",
      "post_date": "2020-07-21T03:29:11.193000",
      "content": "<p>Hi, TPU is really fast than GPU!<br>\nFor you questions, </p>\n<ol>\n<li>TPU takes 20 min for full alaska2 dataset per epoch, including validation. I make my <a href=\"https://www.kaggle.com/leonshangguan/fullsetfold0final\" target=\"_blank\">kernel</a> and <a href=\"https://www.kaggle.com/leonshangguan/alaska2train95baseline\" target=\"_blank\">kernel</a> public here which you can have a reference. Another benefits is that we can train with a large batch size with TPU.</li>\n<li>For tensorflow, no difference, but if you are using pytorch, kaggle kernel only supports old version pytorch than colab.</li>\n<li>TPU supports both pytorch and tf/keras, but based on my experience, pytorch xla has a lot of bugs, and with only few documents, while tf/keras is is much more user friendly. </li>\n<li>I am not sure if there is any performance difference between multi-gpu and tpu, didn't test it before.</li>\n</ol>\n<p>Together with my teammates we tried xla by modify public kernel first, <a href=\"https://www.kaggle.com/zihezheng/pytorch-tpu-transfer-learning-baseline\" target=\"_blank\">kernel here</a>, but faced with a lot of problems, even with the same code, sometimes we can run successfully but sometimes not(don't know why).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 937524,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-21T04:19:17.957000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937405": "it becomes a problem for me in this competition as i am using pytorch + local desk PC with multi GPUs. It is almost impossible to train large network like efficientnetb5,6,7.\n\nI would like to know the speed of using TPU on such for large networks, e.g. time per epoch.\n- any time difference for colab, colab pro, kaggle kernel?\n- how does pytorch xla performs? is tpu only for tf/keras?\n- any performance difference between multi-gpu and tpu? (e.g. multi-gpu has the problem of batchnorm synchronization, is this an issue in replicas in TPU?)",
    "941734": "This is w/ 1950x and 3 x 2080 Ti. Pytorch 1.7, fp16 (autocast), and DataDistributed:\n\n* B0: 18 mins per epoch\n* B4: ~1hr per epoch\n* B7: ~3hr per epoch\n\nTPUv3 is powerful but also expensive: $2.4/hr pre-emptible (in my experience it doesn't last more than 12 hrs). Training 48 hours of would be 48*2.4 = $115. \n\nI am surprised B7 is less than 1 hr/epoch on TPU.",
    "938080": "TPU is very good in TensorFlow/Keras once you set up your tfrecord data and load it into the GCS bucket. You may need a minimum of 2 cores and at least 8 or 16 RAM. Based on my experience in this competition, EfficientNetB7 requires 30-35 minutes of training time with minimum data augmentation. In the case of Pytorch using TPU in image competition, it is highly recommended to use at least 16 cores and higher RAM for the data loader.",
    "938038": "seems that there are still many issues, kaggle can crowd fund kagglers to solve kaggler issues like this (kickstarter data science) ....\n\nsome kaggler just want to do competitions and willing to pay for setup, framework, etc solution\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F03307bc0d131ba076d3c58860e74e73c%2FSelection_034.png?generation=1595325712980186&amp;alt=media)\n",
    "937971": "any idea if the latest Nvidia Ampere RTX 3080 Ti will be comparable to TPU?\n\nI am currently using my 4x1080Ti development environment for kaggle hobby. Now i need to decide to upgrade my system or go for the TPU cloud ... It is a tough choice ...",
    "937517": "@hengck23 what was your bottleneck for local development? I was using 4x1080Ti + 1950X CPU as my main development setup. It was loud and hot, with B6 training time around 1h30m per epoch. \n\nI've tried to train B7 on TPU in Kaggle. In general I was able to fit batch 32 on single TPU. On v3-8 effective batch size was 128, but the problem was - small amount of host memory (16Gb) and only 2 CPU cores. \nData IO was a clear bottleneck here and most of the time TPU were waiting for the data to become available.\nThat's my experience with TPU so far.\n\n",
    "937505": "One epoch in 15~20 mins for Effnet B0/B1, and 30~45 mins for B6/B7. The detail is [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168537). ",
    "937461": "I trained models entirely on TPU with keras. ef7 takes about 30 min / epoch. Not sure why, but Keras performs so much worse than same model with pytorch from public notebook",
    "937460": "Hi, TPU is really fast than GPU!\nFor you questions, \n1. TPU takes 20 min for full alaska2 dataset per epoch, including validation. I make my [kernel](https://www.kaggle.com/leonshangguan/fullsetfold0final) and [kernel](https://www.kaggle.com/leonshangguan/alaska2train95baseline) public here which you can have a reference. Another benefits is that we can train with a large batch size with TPU.\n2. For tensorflow, no difference, but if you are using pytorch, kaggle kernel only supports old version pytorch than colab.\n3. TPU supports both pytorch and tf/keras, but based on my experience, pytorch xla has a lot of bugs, and with only few documents, while tf/keras is is much more user friendly. \n4. I am not sure if there is any performance difference between multi-gpu and tpu, didn't test it before.\n\nTogether with my teammates we tried xla by modify public kernel first, [kernel here](https://www.kaggle.com/zihezheng/pytorch-tpu-transfer-learning-baseline), but faced with a lot of problems, even with the same code, sometimes we can run successfully but sometimes not(don't know why).",
    "937524": ""
  }
}