{
  "id": 164970,
  "title": "[solved?]PyTorch TPU is slower than PyTorch GPU",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164970",
  "author_name": "spoon spoon",
  "post_date": "2020-07-08T04:24:20.404000",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi all,<br>\nI am trying to use TPU accelerator for my pytorch code. The code when run with GPU accelerator takes <code>05:33 ± 0:02</code> minutes per epoch. When I use TPU accelerator the first epoch keeps running for hours until it stops after exceeding the time limit of TPU notebook commit! As the epoch progresses the time to process a batch keeps increasing, eg: (in minutes)<code>0:06</code>, <code>0:12</code>, <code>0:34</code> and so on  . <code>MXU</code> shows <code>0%</code> almost all the time. <code>Idle Time</code> shows <code>--</code>.</p>\n<ul>\n<li>I am fine-tuning a pre-trained <strong>EffNet b-4</strong> model on competition data resized to 256*256.</li>\n<li><code>HYPER_PARAMS = {\n'FOLDS' : 3,\n'EPOCHS' : 12,\n'net_LR' : 1e-4,\n'BS' : 32,\n'TTA' : 1,\n'DEVICE' : xm.xla_device()\n}</code></li>\n</ul>\n<p>I have been tinkering with it since last week but no success. I was suspecting the problem is related to transfer of data from CPU to TPU or vice versa being slow, but since the same code is used for GPU and has pretty good speed there, this doesn't seem the case. Any help will be appreciated. Thank you.</p>\n<p>EDIT: A <em>possible solution</em> is using Pytorch Lightning. It automatically manages multiple TPU cores but still needs some hacks and process is not as smooth as Pytorch + GPU.</p>",
  "messages": [
    {
      "id": 922216,
      "postDate": "2020-07-09T23:00:39.510Z",
      "content": "<p>I am just having same issue, trying to run TPU on Colab and I was able to start training process, but it's quite slow, then I realized \"spawn\" call in example kernels, looks like I need to reorganize my code for to work in multi</p>",
      "rawMarkdown": "I am just having same issue, trying to run TPU on Colab and I was able to start training process, but it's quite slow, then I realized \"spawn\" call in example kernels, looks like I need to reorganize my code for to work in multi",
      "votes": 1
    },
    {
      "id": 920816,
      "postDate": "2020-07-08T20:14:53.403Z",
      "content": "<p>How many core  you use! Can you public xmp code</p>",
      "rawMarkdown": "How many core  you use! Can you public xmp code",
      "votes": 1,
      "replies": [
        {
          "id": 921097,
          "postDate": "2020-07-09T04:33:11.527Z",
          "content": "<p>Hi Manh Lab-san\nI am using single core and not using <code>torch_xla.distributed.xla_multiprocessing</code>.  When I tried <code>xmp</code> and ran my code on all 8 cores it went OOM and gave this error: \n<code>\nRuntimeError: Resource exhausted: From /job:tpu_worker/replica:0/task:0:\nRan out of memory in memory space hbm. Used 16.48G of 15.98G hbm. Exceeded hbm capacity by 506.24M.\n</code>\nI think running on single TPU core is slow and using all cores is exceeding memory limits. I need to find a balance in both.\nThank you.</p>",
          "rawMarkdown": "Hi Manh Lab-san\nI am using single core and not using `torch_xla.distributed.xla_multiprocessing`.  When I tried `xmp` and ran my code on all 8 cores it went OOM and gave this error: \n```\nRuntimeError: Resource exhausted: From /job:tpu_worker/replica:0/task:0:\nRan out of memory in memory space hbm. Used 16.48G of 15.98G hbm. Exceeded hbm capacity by 506.24M.\n```\nI think running on single TPU core is slow and using all cores is exceeding memory limits. I need to find a balance in both.\nThank you.",
          "votes": 1
        },
        {
          "id": 921278,
          "postDate": "2020-07-09T07:15:07.427Z",
          "content": "<p>yeah. I know what you are talking about. The problem is xla only access you run 1 or 8 core. You can try colab with 35GB. I can show you how if you need</p>",
          "rawMarkdown": "yeah. I know what you are talking about. The problem is xla only access you run 1 or 8 core. You can try colab with 35GB. I can show you how if you need",
          "votes": 2
        },
        {
          "id": 921872,
          "postDate": "2020-07-09T16:05:39.487Z",
          "content": "<p>Ok I'll try colab. Seems like everyone is moving from Kaggle kernels to colab pro. If I encounter any problem I'll ask you. Thank you Manh Lab-san.</p>\n\n<p>Edit: I tried running all cores of Kaggle's TPU with <code>torch_xla</code>, it is possible, but with very small batch-size but still it takes hours for one epoch.</p>",
          "rawMarkdown": "Ok I'll try colab. Seems like everyone is moving from Kaggle kernels to colab pro. If I encounter any problem I'll ask you. Thank you Manh Lab-san.\n\nEdit: I tried running all cores of Kaggle's TPU with `torch_xla`, it is possible, but with very small batch-size but still it takes hours for one epoch."
        },
        {
          "id": 922083,
          "postDate": "2020-07-09T19:46:15.420Z",
          "content": "<p>Yeah. I try colab. and have some advice for you. Only load tfrecord. Not clone all original dataset because it too large. With colab you can use \n<code>\npip install kaggle\n</code>\n to download dataset</p>",
          "rawMarkdown": "Yeah. I try colab. and have some advice for you. Only load tfrecord. Not clone all original dataset because it too large. With colab you can use \n```\npip install kaggle\n```\n to download dataset",
          "votes": 1
        }
      ]
    },
    {
      "id": 919700,
      "postDate": "2020-07-08T04:24:20.403Z",
      "content": "<p>Hi all,<br>\nI am trying to use TPU accelerator for my pytorch code. The code when run with GPU accelerator takes <code>05:33 ± 0:02</code> minutes per epoch. When I use TPU accelerator the first epoch keeps running for hours until it stops after exceeding the time limit of TPU notebook commit! As the epoch progresses the time to process a batch keeps increasing, eg: (in minutes)<code>0:06</code>, <code>0:12</code>, <code>0:34</code> and so on  . <code>MXU</code> shows <code>0%</code> almost all the time. <code>Idle Time</code> shows <code>--</code>.</p>\n<ul>\n<li>I am fine-tuning a pre-trained <strong>EffNet b-4</strong> model on competition data resized to 256*256.</li>\n<li><code>HYPER_PARAMS = {\n'FOLDS' : 3,\n'EPOCHS' : 12,\n'net_LR' : 1e-4,\n'BS' : 32,\n'TTA' : 1,\n'DEVICE' : xm.xla_device()\n}</code></li>\n</ul>\n<p>I have been tinkering with it since last week but no success. I was suspecting the problem is related to transfer of data from CPU to TPU or vice versa being slow, but since the same code is used for GPU and has pretty good speed there, this doesn't seem the case. Any help will be appreciated. Thank you.</p>\n<p>EDIT: A <em>possible solution</em> is using Pytorch Lightning. It automatically manages multiple TPU cores but still needs some hacks and process is not as smooth as Pytorch + GPU.</p>",
      "rawMarkdown": "Hi all,\nI am trying to use TPU accelerator for my pytorch code. The code when run with GPU accelerator takes `05:33 ± 0:02` minutes per epoch. When I use TPU accelerator the first epoch keeps running for hours until it stops after exceeding the time limit of TPU notebook commit! As the epoch progresses the time to process a batch keeps increasing, eg: (in minutes)`0:06`, `0:12`, `0:34` and so on  . `MXU` shows `0%` almost all the time. `Idle Time` shows `--`.\n- I am fine-tuning a pre-trained **EffNet b-4** model on competition data resized to 256*256.\n- ``` HYPER_PARAMS = {\n    'FOLDS' : 3,\n    'EPOCHS' : 12,\n    'net_LR' : 1e-4,\n    'BS' : 32,\n    'TTA' : 1,\n    'DEVICE' : xm.xla_device()\n}```\n\nI have been tinkering with it since last week but no success. I was suspecting the problem is related to transfer of data from CPU to TPU or vice versa being slow, but since the same code is used for GPU and has pretty good speed there, this doesn't seem the case. Any help will be appreciated. Thank you.\n\nEDIT: A *possible solution* is using Pytorch Lightning. It automatically manages multiple TPU cores but still needs some hacks and process is not as smooth as Pytorch + GPU.",
      "votes": 2
    },
    {
      "id": 925094,
      "postDate": "2020-07-11T19:56:07.950Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 925291,
          "postDate": "2020-07-12T01:48:28.003Z",
          "content": "<p>looks the issue was learning rate, for some reason learning rate must be smaller not bigger when training on multicore TPU</p>",
          "rawMarkdown": "looks the issue was learning rate, for some reason learning rate must be smaller not bigger when training on multicore TPU"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 922216,
      "author_name": "Jacek Poplawski",
      "author_url": "",
      "post_date": "2020-07-09T23:00:39.510000",
      "content": "<p>I am just having same issue, trying to run TPU on Colab and I was able to start training process, but it's quite slow, then I realized \"spawn\" call in example kernels, looks like I need to reorganize my code for to work in multi</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 920816,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-07-08T20:14:53.403000",
      "content": "<p>How many core  you use! Can you public xmp code</p>",
      "votes": 1,
      "replies": [
        {
          "id": 921097,
          "author_name": "spoon spoon",
          "author_url": "",
          "post_date": "2020-07-09T04:33:11.527000",
          "content": "<p>Hi Manh Lab-san\nI am using single core and not using <code>torch_xla.distributed.xla_multiprocessing</code>.  When I tried <code>xmp</code> and ran my code on all 8 cores it went OOM and gave this error: \n<code>\nRuntimeError: Resource exhausted: From /job:tpu_worker/replica:0/task:0:\nRan out of memory in memory space hbm. Used 16.48G of 15.98G hbm. Exceeded hbm capacity by 506.24M.\n</code>\nI think running on single TPU core is slow and using all cores is exceeding memory limits. I need to find a balance in both.\nThank you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 921278,
          "author_name": "Manh Lab",
          "author_url": "",
          "post_date": "2020-07-09T07:15:07.427000",
          "content": "<p>yeah. I know what you are talking about. The problem is xla only access you run 1 or 8 core. You can try colab with 35GB. I can show you how if you need</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 921872,
          "author_name": "spoon spoon",
          "author_url": "",
          "post_date": "2020-07-09T16:05:39.487000",
          "content": "<p>Ok I'll try colab. Seems like everyone is moving from Kaggle kernels to colab pro. If I encounter any problem I'll ask you. Thank you Manh Lab-san.</p>\n\n<p>Edit: I tried running all cores of Kaggle's TPU with <code>torch_xla</code>, it is possible, but with very small batch-size but still it takes hours for one epoch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922083,
          "author_name": "Manh Lab",
          "author_url": "",
          "post_date": "2020-07-09T19:46:15.420000",
          "content": "<p>Yeah. I try colab. and have some advice for you. Only load tfrecord. Not clone all original dataset because it too large. With colab you can use \n<code>\npip install kaggle\n</code>\n to download dataset</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 925094,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T19:56:07.950000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 925291,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-12T01:48:28.003000",
          "content": "<p>looks the issue was learning rate, for some reason learning rate must be smaller not bigger when training on multicore TPU</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "922216": "I am just having same issue, trying to run TPU on Colab and I was able to start training process, but it's quite slow, then I realized \"spawn\" call in example kernels, looks like I need to reorganize my code for to work in multi",
    "920816": "How many core  you use! Can you public xmp code",
    "919700": "Hi all,\nI am trying to use TPU accelerator for my pytorch code. The code when run with GPU accelerator takes `05:33 ± 0:02` minutes per epoch. When I use TPU accelerator the first epoch keeps running for hours until it stops after exceeding the time limit of TPU notebook commit! As the epoch progresses the time to process a batch keeps increasing, eg: (in minutes)`0:06`, `0:12`, `0:34` and so on  . `MXU` shows `0%` almost all the time. `Idle Time` shows `--`.\n- I am fine-tuning a pre-trained **EffNet b-4** model on competition data resized to 256*256.\n- ``` HYPER_PARAMS = {\n    'FOLDS' : 3,\n    'EPOCHS' : 12,\n    'net_LR' : 1e-4,\n    'BS' : 32,\n    'TTA' : 1,\n    'DEVICE' : xm.xla_device()\n}```\n\nI have been tinkering with it since last week but no success. I was suspecting the problem is related to transfer of data from CPU to TPU or vice versa being slow, but since the same code is used for GPU and has pretty good speed there, this doesn't seem the case. Any help will be appreciated. Thank you.\n\nEDIT: A *possible solution* is using Pytorch Lightning. It automatically manages multiple TPU cores but still needs some hacks and process is not as smooth as Pytorch + GPU.",
    "925094": ""
  }
}