{
  "id": 133777,
  "title": "Working Pytorch/XLA with Kaggle TPU",
  "url": "/competitions/flower-classification-with-tpus/discussion/133777",
  "author_name": "Dhananjay Raut",
  "post_date": "2020-03-04T08:10:44.638000",
  "votes": 24,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I was able to successfully train models using PyTorch on provided TPU using multiprocessing to utilize all the 8 core. thanks to comments of <a href=\"/ifigotin\">@ifigotin</a> and <a href=\"/byrachonok\">@byrachonok</a> 's kernel.\nfollowing are the public kernels if anyone is interested.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing\">kernel 1</a>: Most general one using recommend API and similar to examples in the XLA repo.</li>\n<li><a href=\"https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing\">kernel 2</a>: Using entire data as training set skipping the model evaluation at each Epoch.</li>\n<li><a href=\"https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-single-core\">kernel 3</a>: Using only one core of the TPU.</li>\n</ul>\n\n<p>setup : images of dim 512x512</p>\n\n<p>kernel 1 takes 228 sec per epoch with an effective batch size of 128\nkernel 2 takes 92 sec per epoch with an effective batch size of 128\nkernel 3 takes 140 sec per epoch with an effective batch size of 128</p>\n\n<p>some points I noticed : \n- skipping evaluation significantly reduces the time per epoch even though # of backward passes increased per epoch. (If anyone has an explanation for this I would love to know)\n- <a href=\"https://www.kaggle.com/ratan123/densenet201-flower-classification-with-tpus\">this kernel</a> uses TensorFlow with the same setup as my kernel 2 and takes 80sec per epoch with an effective batch size of 128.\n- the scaling from 1 core to 8 core is not even 2x. also surprisingly I was able to use a batch size of 128 with this image size which is not possible with GPU with 16 gigs of VRAM. (maybe different core share memory in NUMA fashion)</p>\n\n<p>if there are any mistakes/bugs in my kernel/testing let me know.</p>",
  "messages": [
    {
      "id": 763181,
      "postDate": "2020-03-04T08:10:44.640Z",
      "content": "<p>I was able to successfully train models using PyTorch on provided TPU using multiprocessing to utilize all the 8 core. thanks to comments of <a href=\"/ifigotin\">@ifigotin</a> and <a href=\"/byrachonok\">@byrachonok</a> 's kernel.\nfollowing are the public kernels if anyone is interested.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing\">kernel 1</a>: Most general one using recommend API and similar to examples in the XLA repo.</li>\n<li><a href=\"https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing\">kernel 2</a>: Using entire data as training set skipping the model evaluation at each Epoch.</li>\n<li><a href=\"https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-single-core\">kernel 3</a>: Using only one core of the TPU.</li>\n</ul>\n\n<p>setup : images of dim 512x512</p>\n\n<p>kernel 1 takes 228 sec per epoch with an effective batch size of 128\nkernel 2 takes 92 sec per epoch with an effective batch size of 128\nkernel 3 takes 140 sec per epoch with an effective batch size of 128</p>\n\n<p>some points I noticed : \n- skipping evaluation significantly reduces the time per epoch even though # of backward passes increased per epoch. (If anyone has an explanation for this I would love to know)\n- <a href=\"https://www.kaggle.com/ratan123/densenet201-flower-classification-with-tpus\">this kernel</a> uses TensorFlow with the same setup as my kernel 2 and takes 80sec per epoch with an effective batch size of 128.\n- the scaling from 1 core to 8 core is not even 2x. also surprisingly I was able to use a batch size of 128 with this image size which is not possible with GPU with 16 gigs of VRAM. (maybe different core share memory in NUMA fashion)</p>\n\n<p>if there are any mistakes/bugs in my kernel/testing let me know.</p>",
      "rawMarkdown": "I was able to successfully train models using PyTorch on provided TPU using multiprocessing to utilize all the 8 core. thanks to comments of @ifigotin and @byrachonok 's kernel.\nfollowing are the public kernels if anyone is interested.\n\n- [kernel 1](https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing): Most general one using recommend API and similar to examples in the XLA repo.\n- [kernel 2](https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing): Using entire data as training set skipping the model evaluation at each Epoch.\n- [kernel 3](https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-single-core): Using only one core of the TPU.\n\nsetup : images of dim 512x512\n\nkernel 1 takes 228 sec per epoch with an effective batch size of 128\nkernel 2 takes 92 sec per epoch with an effective batch size of 128\nkernel 3 takes 140 sec per epoch with an effective batch size of 128\n\nsome points I noticed : \n- skipping evaluation significantly reduces the time per epoch even though # of backward passes increased per epoch. (If anyone has an explanation for this I would love to know)\n- [this kernel](https://www.kaggle.com/ratan123/densenet201-flower-classification-with-tpus) uses TensorFlow with the same setup as my kernel 2 and takes 80sec per epoch with an effective batch size of 128.\n- the scaling from 1 core to 8 core is not even 2x. also surprisingly I was able to use a batch size of 128 with this image size which is not possible with GPU with 16 gigs of VRAM. (maybe different core share memory in NUMA fashion)\n\n\nif there are any mistakes/bugs in my kernel/testing let me know.",
      "votes": 24
    },
    {
      "id": 763697,
      "postDate": "2020-03-04T18:52:42.347Z",
      "content": "<p>Congrats for making this work !\nThe bad news is that the current version of PyTorch on TPUs is not supported on Kaggle. And I don't mean unsupported as in \"try at your own risk\". I mean unsupported as in \"the PyTorch-TPU team told us there were critical bugs in the code and your models will be failing in funny ways\". Unfortunately, although you can pip-install the latest version of PyTorch-XLA on your Kaggle VM, that does not update the code on the TPU.</p>\n\n<p>We are discussing with the PyTorch-TPU team to be able to offer a stable release. The timeframe for that is &gt; 1 month.</p>",
      "rawMarkdown": "Congrats for making this work !\nThe bad news is that the current version of PyTorch on TPUs is not supported on Kaggle. And I don't mean unsupported as in \"try at your own risk\". I mean unsupported as in \"the PyTorch-TPU team told us there were critical bugs in the code and your models will be failing in funny ways\". Unfortunately, although you can pip-install the latest version of PyTorch-XLA on your Kaggle VM, that does not update the code on the TPU.\n\nWe are discussing with the PyTorch-TPU team to be able to offer a stable release. The timeframe for that is &gt; 1 month.",
      "votes": 9,
      "replies": [
        {
          "id": 763715,
          "postDate": "2020-03-04T19:20:54.320Z",
          "content": "<p>Thanks for the update. Do let us know when it's getting supported.</p>\n\n<p>Is google Colabs also having the same issue?</p>",
          "rawMarkdown": "Thanks for the update. Do let us know when it's getting supported.\n\nIs google Colabs also having the same issue?"
        }
      ]
    },
    {
      "id": 793164,
      "postDate": "2020-03-31T20:59:44.297Z",
      "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a>, Could you please use the following PyTorch/XLA setup:</p>\n\n<p>!curl <a href=\"https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py\">https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py</a> -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev</p>\n\n<p>In your setup, you are missing the call to select the runtime version on the TPU. It will work but with bugs.</p>",
      "rawMarkdown": "@dhananjay3, Could you please use the following PyTorch/XLA setup:\n\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n\nIn your setup, you are missing the call to select the runtime version on the TPU. It will work but with bugs.",
      "votes": 3
    },
    {
      "id": 792982,
      "postDate": "2020-03-31T17:41:21.877Z",
      "content": "<p>Any update on this?</p>",
      "rawMarkdown": "Any update on this?",
      "replies": [
        {
          "id": 793155,
          "postDate": "2020-03-31T20:50:51.410Z",
          "content": "<p>Yes, PyTorch/XLA now works on Kaggle but is still considered \"experimental\".\nGreat PyTorch/XLA - TPU tutorial by <a href=\"/abhishek\">@abhishek</a> in <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265\">this discussion thread</a>, along with interesting memory optimization considerations.</p>",
          "rawMarkdown": "Yes, PyTorch/XLA now works on Kaggle but is still considered \"experimental\".\nGreat PyTorch/XLA - TPU tutorial by @abhishek in [this discussion thread](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265), along with interesting memory optimization considerations."
        },
        {
          "id": 892794,
          "postDate": "2020-06-19T06:57:37.100Z",
          "content": "<p>Hi Martin, are there any updates? or is xla still experimental or unsupported?</p>",
          "rawMarkdown": "Hi Martin, are there any updates? or is xla still experimental or unsupported?",
          "votes": 1
        }
      ]
    },
    {
      "id": 763901,
      "postDate": "2020-03-05T00:41:08.477Z",
      "content": "<p>The weirdest fact is we did the same thing and it didn't work way back :(</p>",
      "rawMarkdown": "The weirdest fact is we did the same thing and it didn't work way back :("
    },
    {
      "id": 763600,
      "postDate": "2020-03-04T16:47:29.320Z",
      "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a>, btw the environment variable hack should no longer be necessary. I.e. you should now have <code>XRT_TPU_CONFIG</code> defined already.</p>",
      "rawMarkdown": "@dhananjay3, btw the environment variable hack should no longer be necessary. I.e. you should now have `XRT_TPU_CONFIG` defined already.",
      "replies": [
        {
          "id": 763636,
          "postDate": "2020-03-04T17:20:48.177Z",
          "content": "<p>Though, I just tried your notebook myself, and it didn't work without the env variable hack for some reason.</p>",
          "rawMarkdown": "Though, I just tried your notebook myself, and it didn't work without the env variable hack for some reason."
        },
        {
          "id": 767465,
          "postDate": "2020-03-09T17:09:48.277Z",
          "content": "<p>Ok, <em>now</em> the hack with <code>XRT_TPU_CONFIG</code> environment variable is no longer needed. It should be properly defined automatically.</p>",
          "rawMarkdown": "Ok, *now* the hack with `XRT_TPU_CONFIG` environment variable is no longer needed. It should be properly defined automatically.",
          "votes": 1
        }
      ]
    },
    {
      "id": 852964,
      "postDate": "2020-05-18T20:57:14.367Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 768301,
      "postDate": "2020-03-10T16:11:19.813Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 763697,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-03-04T18:52:42.347000",
      "content": "<p>Congrats for making this work !\nThe bad news is that the current version of PyTorch on TPUs is not supported on Kaggle. And I don't mean unsupported as in \"try at your own risk\". I mean unsupported as in \"the PyTorch-TPU team told us there were critical bugs in the code and your models will be failing in funny ways\". Unfortunately, although you can pip-install the latest version of PyTorch-XLA on your Kaggle VM, that does not update the code on the TPU.</p>\n\n<p>We are discussing with the PyTorch-TPU team to be able to offer a stable release. The timeframe for that is &gt; 1 month.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 763715,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-03-04T19:20:54.320000",
          "content": "<p>Thanks for the update. Do let us know when it's getting supported.</p>\n\n<p>Is google Colabs also having the same issue?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793164,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-03-31T20:59:44.297000",
      "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a>, Could you please use the following PyTorch/XLA setup:</p>\n\n<p>!curl <a href=\"https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py\">https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py</a> -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev</p>\n\n<p>In your setup, you are missing the call to select the runtime version on the TPU. It will work but with bugs.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 792982,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-03-31T17:41:21.877000",
      "content": "<p>Any update on this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 793155,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-03-31T20:50:51.410000",
          "content": "<p>Yes, PyTorch/XLA now works on Kaggle but is still considered \"experimental\".\nGreat PyTorch/XLA - TPU tutorial by <a href=\"/abhishek\">@abhishek</a> in <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265\">this discussion thread</a>, along with interesting memory optimization considerations.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 892794,
          "author_name": "Aashish Ghosh",
          "author_url": "",
          "post_date": "2020-06-19T06:57:37.100000",
          "content": "<p>Hi Martin, are there any updates? or is xla still experimental or unsupported?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 763901,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-03-05T00:41:08.477000",
      "content": "<p>The weirdest fact is we did the same thing and it didn't work way back :(</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 763600,
      "author_name": "Ilya Figotin",
      "author_url": "",
      "post_date": "2020-03-04T16:47:29.320000",
      "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a>, btw the environment variable hack should no longer be necessary. I.e. you should now have <code>XRT_TPU_CONFIG</code> defined already.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 763636,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-03-04T17:20:48.177000",
          "content": "<p>Though, I just tried your notebook myself, and it didn't work without the env variable hack for some reason.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767465,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-03-09T17:09:48.277000",
          "content": "<p>Ok, <em>now</em> the hack with <code>XRT_TPU_CONFIG</code> environment variable is no longer needed. It should be properly defined automatically.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 852964,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-18T20:57:14.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 768301,
      "author_name": "Kranti Kumar",
      "author_url": "",
      "post_date": "2020-03-10T16:11:19.813000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "763181": "I was able to successfully train models using PyTorch on provided TPU using multiprocessing to utilize all the 8 core. thanks to comments of @ifigotin and @byrachonok 's kernel.\nfollowing are the public kernels if anyone is interested.\n\n- [kernel 1](https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing): Most general one using recommend API and similar to examples in the XLA repo.\n- [kernel 2](https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing): Using entire data as training set skipping the model evaluation at each Epoch.\n- [kernel 3](https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-single-core): Using only one core of the TPU.\n\nsetup : images of dim 512x512\n\nkernel 1 takes 228 sec per epoch with an effective batch size of 128\nkernel 2 takes 92 sec per epoch with an effective batch size of 128\nkernel 3 takes 140 sec per epoch with an effective batch size of 128\n\nsome points I noticed : \n- skipping evaluation significantly reduces the time per epoch even though # of backward passes increased per epoch. (If anyone has an explanation for this I would love to know)\n- [this kernel](https://www.kaggle.com/ratan123/densenet201-flower-classification-with-tpus) uses TensorFlow with the same setup as my kernel 2 and takes 80sec per epoch with an effective batch size of 128.\n- the scaling from 1 core to 8 core is not even 2x. also surprisingly I was able to use a batch size of 128 with this image size which is not possible with GPU with 16 gigs of VRAM. (maybe different core share memory in NUMA fashion)\n\n\nif there are any mistakes/bugs in my kernel/testing let me know.",
    "763697": "Congrats for making this work !\nThe bad news is that the current version of PyTorch on TPUs is not supported on Kaggle. And I don't mean unsupported as in \"try at your own risk\". I mean unsupported as in \"the PyTorch-TPU team told us there were critical bugs in the code and your models will be failing in funny ways\". Unfortunately, although you can pip-install the latest version of PyTorch-XLA on your Kaggle VM, that does not update the code on the TPU.\n\nWe are discussing with the PyTorch-TPU team to be able to offer a stable release. The timeframe for that is &gt; 1 month.",
    "793164": "@dhananjay3, Could you please use the following PyTorch/XLA setup:\n\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n\nIn your setup, you are missing the call to select the runtime version on the TPU. It will work but with bugs.",
    "792982": "Any update on this?",
    "763901": "The weirdest fact is we did the same thing and it didn't work way back :(",
    "763600": "@dhananjay3, btw the environment variable hack should no longer be necessary. I.e. you should now have `XRT_TPU_CONFIG` defined already.",
    "852964": "",
    "768301": "Thanks for sharing!"
  }
}