{
  "id": 138137,
  "title": "PyTorch XLA Support?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/138137",
  "author_name": "",
  "post_date": "2020-03-23T22:50:27.135469Z",
  "votes": 13,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Do we have XLA support now on kaggle TPU's  for PyTorch?</p>",
  "messages": [
    {
      "id": "784046",
      "postDate": "03/23/2020 22:50:27",
      "content": "<p>Do we have XLA support now on kaggle TPU's  for PyTorch?</p>",
      "rawMarkdown": "Do we have XLA support now on kaggle TPU's  for PyTorch?",
      "votes": null
    },
    {
      "id": "784053",
      "postDate": "03/23/2020 22:58:28",
      "content": "<p>PyTorch is not yet officially supported, but you can demo our current workaround as follows:</p>\n\n<p><code>\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1314380%2Fbbb90f9a1220e92b3b5b756de7fa3dd0%2FScreen%20Shot%202020-03-23%20at%204.57.57%20PM.png?generation=1585004290331820&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "PyTorch is not yet officially supported, but you can demo our current workaround as follows:\n\n```\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1314380%2Fbbb90f9a1220e92b3b5b756de7fa3dd0%2FScreen%20Shot%202020-03-23%20at%204.57.57%20PM.png?generation=1585004290331820&amp;alt=media)",
      "votes": null
    },
    {
      "id": "784066",
      "postDate": "03/23/2020 23:09:26",
      "content": "<p>Sounds Great!🔥🔥🔥</p>",
      "rawMarkdown": "Sounds Great!🔥🔥🔥",
      "votes": null
    },
    {
      "id": "784072",
      "postDate": "03/23/2020 23:16:10",
      "content": "<p>So we can now train with PyTorch on Kaggle TPUs?</p>",
      "rawMarkdown": "So we can now train with PyTorch on Kaggle TPUs?",
      "votes": null
    },
    {
      "id": "784076",
      "postDate": "03/23/2020 23:24:53",
      "content": "<p>I've seen a handful of successful examples, yes.  Please share a public notebook if you get something up and running!</p>",
      "rawMarkdown": "I've seen a handful of successful examples, yes.  Please share a public notebook if you get something up and running!",
      "votes": null
    },
    {
      "id": "784083",
      "postDate": "03/23/2020 23:31:37",
      "content": "<p>Sounds good. I am looking forward to using PyTorch for TPU training in this competition!</p>",
      "rawMarkdown": "Sounds good. I am looking forward to using PyTorch for TPU training in this competition!",
      "votes": null
    },
    {
      "id": "784084",
      "postDate": "03/23/2020 23:32:49",
      "content": "<p>Yes, but please be aware that:\n- The <a href=\"https://github.com/pytorch/xla\">PyTorch-XLA</a> has not yet released their first stable release. They are working towards that goal.\n- With PyTorch, the Kaggle VM is feeding data to the TPU directly. With a relatively small VM, the TPU can end up being starved of data. However, in an NLP competition like this one, the data is made of numerical tokens (tokenized words) and is therefore small. It should fit in memory and you should not have an issue with bandwidth.</p>",
      "rawMarkdown": "Yes, but please be aware that:\n- The [PyTorch-XLA](https://github.com/pytorch/xla) has not yet released their first stable release. They are working towards that goal.\n- With PyTorch, the Kaggle VM is feeding data to the TPU directly. With a relatively small VM, the TPU can end up being starved of data. However, in an NLP competition like this one, the data is made of numerical tokens (tokenized words) and is therefore small. It should fit in memory and you should not have an issue with bandwidth.",
      "votes": null
    },
    {
      "id": "784093",
      "postDate": "03/23/2020 23:42:22",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for the info. I have used PyTorch XLA successfully in the past so I am not worried about the fact that there's no stable release yet. Instead, I was more worried if Kaggle would support PyTorch XLA, and if there would be any issues with data loading and optimizing TPU usage. But as you mentioned, the data should be small enough that this won't be too much of a problem. </p>",
      "rawMarkdown": "mgornergoogle Thanks for the info. I have used PyTorch XLA successfully in the past so I am not worried about the fact that there's no stable release yet. Instead, I was more worried if Kaggle would support PyTorch XLA, and if there would be any issues with data loading and optimizing TPU usage. But as you mentioned, the data should be small enough that this won't be too much of a problem.",
      "votes": null
    },
    {
      "id": "784099",
      "postDate": "03/23/2020 23:52:42",
      "content": "<p>never worked with <code>TPU</code> but it sounds exciting! Since many people at home have <code>GPU</code>  I was wondering if it easy to train model using <code>GPU</code> and than use <code>TPU</code> for inference ? </p>\n\n<p>asking for a friend who has Pytorch =) </p>",
      "rawMarkdown": "never worked with `TPU` but it sounds exciting! Since many people at home have `GPU`  I was wondering if it easy to train model using `GPU` and than use `TPU` for inference ? \n\nasking for a friend who has Pytorch =)",
      "votes": null
    },
    {
      "id": "784140",
      "postDate": "03/24/2020 00:49:18",
      "content": "<p>30 TPU hours are free on Kaggle. Just curious why you would want to train on GPU if it's equally easy (especially since it is faster on TPU). I wouldn't see why you couldn't use TPU for inference though.</p>",
      "rawMarkdown": "30 TPU hours are free on Kaggle. Just curious why you would want to train on GPU if it's equally easy (especially since it is faster on TPU). I wouldn't see why you couldn't use TPU for inference though.",
      "votes": null
    },
    {
      "id": "784141",
      "postDate": "03/24/2020 00:49:18",
      "content": "<p>Sure that should work. You can also submit with GPU.\nHowever the goal of TPUs is to iterate faster. Using the <a href=\"https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started\">Getting started notebook</a>:\n- TPU: 58 sec per 500 steps (500 steps = 128,000 sentences at batch size 256)\n- GPU: 246 sec per 500 steps (500 steps = 16,000 sentences at batch size 32)</p>",
      "rawMarkdown": "Sure that should work. You can also submit with GPU.\nHowever the goal of TPUs is to iterate faster. Using the [Getting started notebook](https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started):\n- TPU: 58 sec per 500 steps (500 steps = 128,000 sentences at batch size 256)\n- GPU: 246 sec per 500 steps (500 steps = 16,000 sentences at batch size 32)",
      "votes": null
    },
    {
      "id": "784143",
      "postDate": "03/24/2020 00:52:49",
      "content": "<p><code>30 TPU hours are free on Kaggle. Just curious why you would want to train on GPU if it's equally easy (especially since it is faster on TPU). I wouldn't see why you couldn't use TPU for inference though.</code></p>\n\n<p>Yes just to clearly my question, I think we are on the same page.. I want to train locally on GPU (since it easier to do experiments and test out stuff) and do inference using TPU =) </p>",
      "rawMarkdown": "`30 TPU hours are free on Kaggle. Just curious why you would want to train on GPU if it's equally easy (especially since it is faster on TPU). I wouldn't see why you couldn't use TPU for inference though.`\n\nYes just to clearly my question, I think we are on the same page.. I want to train locally on GPU (since it easier to do experiments and test out stuff) and do inference using TPU =)",
      "votes": null
    },
    {
      "id": "784144",
      "postDate": "03/24/2020 00:53:19",
      "content": "<p>Thank you <a href=\"/mgornergoogle\">@mgornergoogle</a> !</p>",
      "rawMarkdown": "Thank you @mgornergoogle !",
      "votes": null
    },
    {
      "id": "784147",
      "postDate": "03/24/2020 00:56:40",
      "content": "<p>Oh ok. I just disagree with the fact that it's \"easier to do experiments and test out stuff\", since TPU is faster. But I guess that depends on your setup.</p>",
      "rawMarkdown": "Oh ok. I just disagree with the fact that it's \"easier to do experiments and test out stuff\", since TPU is faster. But I guess that depends on your setup.",
      "votes": null
    },
    {
      "id": "784276",
      "postDate": "03/24/2020 04:29:21",
      "content": "<p>I have a concern that is not mentioned in the Requirements. Do we allow to use the external dataset including our local pretrained models? Or Training and inference should be made by the kernel?  </p>",
      "rawMarkdown": "I have a concern that is not mentioned in the Requirements. Do we allow to use the external dataset including our local pretrained models? Or Training and inference should be made by the kernel?",
      "votes": null
    },
    {
      "id": "784307",
      "postDate": "03/24/2020 05:12:27",
      "content": "<p><a href=\"/backaggle\">@backaggle</a> External data is permitted. Including offline trained model(s) that are loaded as external data. Standard publicly available pretrained models count as External data that needs to be declared on the official thread. Note also that for external data to be available to run on Kaggle’s TPUs in a submission notebook, that external data must also be made public.</p>\n\n<p>Note that the TPU star prize, however, requires that you use TPUs end to end, including on training. That might not necessarily mean it happen in Kaggle or entirely in one notebook, but is certainly an option with the integration.</p>\n\n<p>The firm code requirement is that the submission.csv be generated out of a notebook. And that submission’s notebook will be constrained to the Code Requirements stated on the overview page. We won’t be re-running your notebook, but the csv submission has to come out of a notebook to be eligible for submission.</p>",
      "rawMarkdown": "backaggle External data is permitted. Including offline trained model(s) that are loaded as external data. Standard publicly available pretrained models count as External data that needs to be declared on the official thread. Note also that for external data to be available to run on Kaggle’s TPUs in a submission notebook, that external data must also be made public.\n\nNote that the TPU star prize, however, requires that you use TPUs end to end, including on training. That might not necessarily mean it happen in Kaggle or entirely in one notebook, but is certainly an option with the integration.\n\nThe firm code requirement is that the submission.csv be generated out of a notebook. And that submission’s notebook will be constrained to the Code Requirements stated on the overview page. We won’t be re-running your notebook, but the csv submission has to come out of a notebook to be eligible for submission.",
      "votes": null
    },
    {
      "id": "784317",
      "postDate": "03/24/2020 05:26:04",
      "content": "<p>But in my little experience with TPU's, they prefer using data from GCP Buckets for better performance, how do we link that here Julia?</p>",
      "rawMarkdown": "But in my little experience with TPU's, they prefer using data from GCP Buckets for better performance, how do we link that here Julia?",
      "votes": null
    },
    {
      "id": "784673",
      "postDate": "03/24/2020 12:30:59",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271</a></p>",
      "rawMarkdown": "See https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271",
      "votes": null
    },
    {
      "id": "784931",
      "postDate": "03/24/2020 16:07:57",
      "content": "<p>Using <code>KaggleDatasets().get_gcs_path()</code> as in the <a href=\"https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started\">Getting Started notebook</a>.\nThe function copies the dataset to a GCS bucket close to the TPU you have been assigned.</p>",
      "rawMarkdown": "Using `KaggleDatasets().get_gcs_path()` as in the [Getting Started notebook](https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started).\nThe function copies the dataset to a GCS bucket close to the TPU you have been assigned.",
      "votes": null
    },
    {
      "id": "784945",
      "postDate": "03/24/2020 16:19:43",
      "content": "<p>This is super cool! :partyparrot:</p>",
      "rawMarkdown": "This is super cool! :partyparrot:",
      "votes": null
    },
    {
      "id": "785319",
      "postDate": "03/25/2020 01:00:20",
      "content": "<p>But I don't think loading from GCS offers a perf advantage for PyTorch.</p>",
      "rawMarkdown": "But I don't think loading from GCS offers a perf advantage for PyTorch.",
      "votes": null
    },
    {
      "id": "785361",
      "postDate": "03/25/2020 02:00:35",
      "content": "<p>I am aware of the them cpmp :); Rather was probably the first one here to play around along with <a href=\"/tanlikesmath\">@tanlikesmath</a> for Flowers TPU Comp :) (wrt PyTroch-XLA)\nThanks for sharing but cpmp!</p>",
      "rawMarkdown": "I am aware of the them cpmp :); Rather was probably the first one here to play around along with @tanlikesmath for Flowers TPU Comp :) (wrt PyTroch-XLA)\nThanks for sharing but cpmp!",
      "votes": null
    },
    {
      "id": "794812",
      "postDate": "04/02/2020 05:36:21",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "859985",
      "postDate": "05/24/2020 23:18:26",
      "content": "<p>Using the above installs has resulted in this error when torch_xla is imported:</p>\n\n<p>ImportError: /opt/conda/lib/python3.7/site-packages/_XLAC.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZN5torch11CppFunctionC1EN3c1014KernelFunctionESt10unique_ptrINS1_14FunctionSchemaESt14default_deleteIS4_EE</p>\n\n<p>Any idea what the issue is? Anyones help would be appreciated </p>",
      "rawMarkdown": "Using the above installs has resulted in this error when torch_xla is imported:\n\nImportError: /opt/conda/lib/python3.7/site-packages/_XLAC.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZN5torch11CppFunctionC1EN3c1014KernelFunctionESt10unique_ptrINS1_14FunctionSchemaESt14default_deleteIS4_EE\n\nAny idea what the issue is? Anyones help would be appreciated",
      "votes": null
    },
    {
      "id": "873240",
      "postDate": "06/04/2020 00:31:40",
      "content": "<p>Installed as requested on the notebook</p>\n\n<p><code>\nVERSION = \"20200515\" #\"20200325\"  #@param [\"1.5\" , \"20200325\", \"nightly\"]\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code></p>\n\n<p>With and without the extra <code>--apt-packages</code> causes this error... I also tried to install with <code>VERSION=\"20200325\"</code> with less luck :) (because it could not find the repos I guess)</p>\n\n<p>```</p>\n\n<p>version 20200515\n  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current\n                                 Dload  Upload   Total   Spent    Left  Speed\n100  4264  100  4264    0     0  18539      0 --:--:-- --:--:-- --:--:-- 18539\n------------------------------------------------------------- <strong><em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>**<em>*</em></strong> -------------------------------------------------------------\nUpdating TPU and VM. This may take around 2 minutes.\nUpdating TPU runtime to pytorch-dev20200515 ...\nFound existing installation: torch 1.6.0a0+bf2bbd9\nUninstalling torch-1.6.0a0+bf2bbd9:\n  Successfully uninstalled torch-1.6.0a0+bf2bbd9\nFound existing installation: torchvision 0.7.0a0+a6073f0\nUninstalling torchvision-0.7.0a0+a6073f0:\n  Successfully uninstalled torchvision-0.7.0a0+a6073f0\nCopying gs://tpu-pytorch/wheels/torch-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n- [1 files][ 91.0 MiB/ 91.0 MiB] <br>\nOperation completed over 1 objects/91.0 MiB. <br>\nCopying gs://tpu-pytorch/wheels/torch_xla-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n\\ [1 files][119.5 MiB/119.5 MiB] <br>\nOperation completed over 1 objects/119.5 MiB. <br>\nCopying gs://tpu-pytorch/wheels/torchvision-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n/ [1 files][  2.3 MiB/  2.3 MiB] <br>\nOperation completed over 1 objects/2.3 MiB. <br>\nProcessing ./torch-nightly+20200515-cp37-cp37m-linux_x86_64.whl\nDone updating TPU runtime: \nRequirement already satisfied: numpy in /opt/conda/lib/python3.7/site-packages (from torch==nightly+20200515) (1.18.1)\nRequirement already satisfied: future in /opt/conda/lib/python3.7/site-packages (from torch==nightly+20200515) (0.18.2)\nERROR: fastai 1.0.61 requires torchvision, which is not installed.\nERROR: kornia 0.3.1 has requirement torch==1.5.0, but you'll have torch 1.6.0a0+bf2bbd9 which is incompatible.\nERROR: allennlp 0.9.0 has requirement spacy&lt;2.2,&gt;=2.1.0, but you'll have spacy 2.2.3 which is incompatible.\nInstalling collected packages: torch\nSuccessfully installed torch-1.6.0a0+bf2bbd9\nWARNING: You are using pip version 20.1; however, version 20.1.1 is available.\nYou should consider upgrading via the '/opt/conda/bin/python3.7 -m pip install --upgrade pip' command.\nProcessing ./torch_xla-nightly+20200515-cp37-cp37m-linux_x86_64.whl\nInstalling collected packages: torch-xla\n  Attempting uninstall: torch-xla\n    Found existing installation: torch-xla 1.6+2b2085a\n    Uninstalling torch-xla-1.6+2b2085a:\n      Successfully uninstalled torch-xla-1.6+2b2085a\nSuccessfully installed torch-xla-1.6+2b2085a\nWARNING: You are using pip version 20.1; however, version 20.1.1 is available.\nYou should consider upgrading via the '/opt/conda/bin/python3.7 -m pip install --upgrade pip' command.\nProcessing ./torchvision-nightly+20200515-cp37-cp37m-linux_x86_64.whl\nRequirement already satisfied: torch in /opt/conda/lib/python3.7/site-packages (from torchvision==nightly+20200515) (1.6.0a0+bf2bbd9)\nRequirement already satisfied: numpy in /opt/conda/lib/python3.7/site-packages (from torchvision==nightly+20200515) (1.18.1)\nRequirement already satisfied: pillow&gt;=4.1.1 in /opt/conda/lib/python3.7/site-packages (from torchvision==nightly+20200515) (5.4.1)\nRequirement already satisfied: future in /opt/conda/lib/python3.7/site-packages (from torch-&gt;torchvision==nightly+20200515) (0.18.2)\nInstalling collected packages: torchvision\nSuccessfully installed torchvision-0.7.0a0+a6073f0\nWARNING: You are using pip version 20.1; however, version 20.1.1 is available.\nYou should consider upgrading via the '/opt/conda/bin/python3.7 -m pip install --upgrade pip' command.\nReading package lists... Done\nBuilding dependency tree <br>\nReading state information... Done\nlibomp5 is already the newest version (5.0.1-1).\nThe following additional packages will be installed:\n  libgfortran4 libopenblas-base\nThe following NEW packages will be installed:\n  libgfortran4 libopenblas-base libopenblas-dev\n0 upgraded, 3 newly installed, 0 to remove and 17 not upgraded.\nNeed to get 8316 kB of archives.\nAfter this operation, 96.8 MB of additional disk space will be used.\nGet:1 <a href=\"http://archive.ubuntu.com/ubuntu\">http://archive.ubuntu.com/ubuntu</a> bionic-updates/main amd64 libgfortran4 amd64 7.5.0-3ubuntu1~18.04 [492 kB]\nGet:2 <a href=\"http://archive.ubuntu.com/ubuntu\">http://archive.ubuntu.com/ubuntu</a> bionic/universe amd64 libopenblas-base amd64 0.2.20+ds-4 [3964 kB]\nGet:3 <a href=\"http://archive.ubuntu.com/ubuntu\">http://archive.ubuntu.com/ubuntu</a> bionic/universe amd64 libopenblas-dev amd64 0.2.20+ds-4 [3860 kB]\nFetched 8316 kB in 0s (39.2 MB/s) <br>\ndebconf: delaying package configuration, since apt-utils is not installed\nSelecting previously unselected package libgfortran4:amd64.\n(Reading database ... 105812 files and directories currently installed.)\nPreparing to unpack .../libgfortran4_7.5.0-3ubuntu1~18.04_amd64.deb ...\nUnpacking libgfortran4:amd64 (7.5.0-3ubuntu1~18.04) ...\nSelecting previously unselected package libopenblas-base:amd64.\nPreparing to unpack .../libopenblas-base_0.2.20+ds-4_amd64.deb ...\nUnpacking libopenblas-base:amd64 (0.2.20+ds-4) ...\nSelecting previously unselected package libopenblas-dev:amd64.\nPreparing to unpack .../libopenblas-dev_0.2.20+ds-4_amd64.deb ...\nUnpacking libopenblas-dev:amd64 (0.2.20+ds-4) ...\nSetting up libgfortran4:amd64 (7.5.0-3ubuntu1~18.04) ...\nSetting up libopenblas-base:amd64 (0.2.20+ds-4) ...\nupdate-alternatives: using /usr/lib/x86_64-linux-gnu/openblas/libblas.so.3 to provide /usr/lib/x86_64-linux-gnu/libblas.so.3 (libblas.so.3-x86_64-linux-gnu) in auto mode\nupdate-alternatives: using /usr/lib/x86_64-linux-gnu/openblas/liblapack.so.3 to provide /usr/lib/x86_64-linux-gnu/liblapack.so.3 (liblapack.so.3-x86_64-linux-gnu) in auto mode\nSetting up libopenblas-dev:amd64 (0.2.20+ds-4) ...\nupdate-alternatives: using /usr/lib/x86_64-linux-gnu/openblas/libblas.so to provide /usr/lib/x86_64-linux-gnu/libblas.so (libblas.so-x86_64-linux-gnu) in auto mode\nupdate-alternatives: using /usr/lib/x86_64-linux-gnu/openblas/liblapack.so to provide /usr/lib/x86_64-linux-gnu/liblapack.so (liblapack.so-x86_64-linux-gnu) in auto mode\nProcessing triggers for libc-bin (2.27-3ubuntu1) ...\n------------------------------------------------------------- <strong><em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>**<em>*</em></strong> -------------------------------------------------------------</p>\n\n<p>```</p>\n\n<p>Thanks for any input on how to run on TPU+pytorch.</p>",
      "rawMarkdown": "Installed as requested on the notebook\n\n```\nVERSION = \"20200515\" #\"20200325\"  #@param [\"1.5\" , \"20200325\", \"nightly\"]\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n```\n\nWith and without the extra `--apt-packages` causes this error... I also tried to install with `VERSION=\"20200325\"` with less luck :) (because it could not find the repos I guess)\n\n```\n\nversion 20200515\n  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current\n                                 Dload  Upload   Total   Spent    Left  Speed\n100  4264  100  4264    0     0  18539      0 --:--:-- --:--:-- --:--:-- 18539\n------------------------------------------------------------- *************************** -------------------------------------------------------------\nUpdating TPU and VM. This may take around 2 minutes.\nUpdating TPU runtime to pytorch-dev20200515 ...\nFound existing installation: torch 1.6.0a0+bf2bbd9\nUninstalling torch-1.6.0a0+bf2bbd9:\n  Successfully uninstalled torch-1.6.0a0+bf2bbd9\nFound existing installation: torchvision 0.7.0a0+a6073f0\nUninstalling torchvision-0.7.0a0+a6073f0:\n  Successfully uninstalled torchvision-0.7.0a0+a6073f0\nCopying gs://tpu-pytorch/wheels/torch-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n- [1 files][ 91.0 MiB/ 91.0 MiB]                                                \nOperation completed over 1 objects/91.0 MiB.                                     \nCopying gs://tpu-pytorch/wheels/torch_xla-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n\\ [1 files][119.5 MiB/119.5 MiB]                                                \nOperation completed over 1 objects/119.5 MiB.                                    \nCopying gs://tpu-pytorch/wheels/torchvision-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n/ [1 files][  2.3 MiB/  2.3 MiB]                                                \nOperation completed over 1 objects/2.3 MiB.                                      \nProcessing ./torch-nightly+20200515-cp37-cp37m-linux_x86_64.whl\nDone updating TPU runtime:",
      "votes": null
    },
    {
      "id": "1327290",
      "postDate": "05/29/2021 06:20:40",
      "content": "<p>Running into the same problem. Did you manage to find a workaround? </p>",
      "rawMarkdown": "Running into the same problem. Did you manage to find a workaround?",
      "votes": null
    },
    {
      "id": "1686049",
      "postDate": "02/11/2022 18:09:26",
      "content": "<p>I always end up getting this error. Any suggestions?</p>\n<p><code>\nImportError: /opt/conda/lib/python3.7/site-packages/_XLAC.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZN5torch11CppFunctionC1EN3c1014KernelFunctionESt10unique_ptrINS1_14FunctionSchemaESt14default_deleteIS4_EE\n</code></p>\n<p>I have ran - </p>\n<p><code>\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code></p>",
      "rawMarkdown": "I always end up getting this error. Any suggestions?\n\n`\nImportError: /opt/conda/lib/python3.7/site-packages/_XLAC.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZN5torch11CppFunctionC1EN3c1014KernelFunctionESt10unique_ptrINS1_14FunctionSchemaESt14default_deleteIS4_EE\n`\n\nI have ran - \n\n`\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1686049,
      "author_name": "jainishsavalia",
      "author_url": "",
      "post_date": "02/11/2022 18:09:26",
      "content": "<p>I always end up getting this error. Any suggestions?</p>\n<p><code>\nImportError: /opt/conda/lib/python3.7/site-packages/_XLAC.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZN5torch11CppFunctionC1EN3c1014KernelFunctionESt10unique_ptrINS1_14FunctionSchemaESt14default_deleteIS4_EE\n</code></p>\n<p>I have ran - </p>\n<p><code>\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 784053,
      "author_name": "paultimothymooney",
      "author_url": "",
      "post_date": "03/23/2020 22:58:28",
      "content": "<p>PyTorch is not yet officially supported, but you can demo our current workaround as follows:</p>\n\n<p><code>\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1314380%2Fbbb90f9a1220e92b3b5b756de7fa3dd0%2FScreen%20Shot%202020-03-23%20at%204.57.57%20PM.png?generation=1585004290331820&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 784066,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/23/2020 23:09:26",
          "content": "<p>Sounds Great!🔥🔥🔥</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784072,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "03/23/2020 23:16:10",
          "content": "<p>So we can now train with PyTorch on Kaggle TPUs?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784076,
          "author_name": "paultimothymooney",
          "author_url": "",
          "post_date": "03/23/2020 23:24:53",
          "content": "<p>I've seen a handful of successful examples, yes.  Please share a public notebook if you get something up and running!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784083,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "03/23/2020 23:31:37",
          "content": "<p>Sounds good. I am looking forward to using PyTorch for TPU training in this competition!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784084,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/23/2020 23:32:49",
          "content": "<p>Yes, but please be aware that:\n- The <a href=\"https://github.com/pytorch/xla\">PyTorch-XLA</a> has not yet released their first stable release. They are working towards that goal.\n- With PyTorch, the Kaggle VM is feeding data to the TPU directly. With a relatively small VM, the TPU can end up being starved of data. However, in an NLP competition like this one, the data is made of numerical tokens (tokenized words) and is therefore small. It should fit in memory and you should not have an issue with bandwidth.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784093,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "03/23/2020 23:42:22",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for the info. I have used PyTorch XLA successfully in the past so I am not worried about the fact that there's no stable release yet. Instead, I was more worried if Kaggle would support PyTorch XLA, and if there would be any issues with data loading and optimizing TPU usage. But as you mentioned, the data should be small enough that this won't be too much of a problem. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 859985,
          "author_name": "jewelltaylor",
          "author_url": "",
          "post_date": "05/24/2020 23:18:26",
          "content": "<p>Using the above installs has resulted in this error when torch_xla is imported:</p>\n\n<p>ImportError: /opt/conda/lib/python3.7/site-packages/_XLAC.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZN5torch11CppFunctionC1EN3c1014KernelFunctionESt10unique_ptrINS1_14FunctionSchemaESt14default_deleteIS4_EE</p>\n\n<p>Any idea what the issue is? Anyones help would be appreciated </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1327290,
          "author_name": "spsayakpaul",
          "author_url": "",
          "post_date": "05/29/2021 06:20:40",
          "content": "<p>Running into the same problem. Did you manage to find a workaround? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 784099,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "03/23/2020 23:52:42",
      "content": "<p>never worked with <code>TPU</code> but it sounds exciting! Since many people at home have <code>GPU</code>  I was wondering if it easy to train model using <code>GPU</code> and than use <code>TPU</code> for inference ? </p>\n\n<p>asking for a friend who has Pytorch =) </p>",
      "votes": null,
      "replies": [
        {
          "id": 784140,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "03/24/2020 00:49:18",
          "content": "<p>30 TPU hours are free on Kaggle. Just curious why you would want to train on GPU if it's equally easy (especially since it is faster on TPU). I wouldn't see why you couldn't use TPU for inference though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784141,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/24/2020 00:49:18",
          "content": "<p>Sure that should work. You can also submit with GPU.\nHowever the goal of TPUs is to iterate faster. Using the <a href=\"https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started\">Getting started notebook</a>:\n- TPU: 58 sec per 500 steps (500 steps = 128,000 sentences at batch size 256)\n- GPU: 246 sec per 500 steps (500 steps = 16,000 sentences at batch size 32)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784143,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/24/2020 00:52:49",
          "content": "<p><code>30 TPU hours are free on Kaggle. Just curious why you would want to train on GPU if it's equally easy (especially since it is faster on TPU). I wouldn't see why you couldn't use TPU for inference though.</code></p>\n\n<p>Yes just to clearly my question, I think we are on the same page.. I want to train locally on GPU (since it easier to do experiments and test out stuff) and do inference using TPU =) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784144,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/24/2020 00:53:19",
          "content": "<p>Thank you <a href=\"/mgornergoogle\">@mgornergoogle</a> !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784147,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "03/24/2020 00:56:40",
          "content": "<p>Oh ok. I just disagree with the fact that it's \"easier to do experiments and test out stuff\", since TPU is faster. But I guess that depends on your setup.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784276,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "03/24/2020 04:29:21",
          "content": "<p>I have a concern that is not mentioned in the Requirements. Do we allow to use the external dataset including our local pretrained models? Or Training and inference should be made by the kernel?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784307,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "03/24/2020 05:12:27",
          "content": "<p><a href=\"/backaggle\">@backaggle</a> External data is permitted. Including offline trained model(s) that are loaded as external data. Standard publicly available pretrained models count as External data that needs to be declared on the official thread. Note also that for external data to be available to run on Kaggle’s TPUs in a submission notebook, that external data must also be made public.</p>\n\n<p>Note that the TPU star prize, however, requires that you use TPUs end to end, including on training. That might not necessarily mean it happen in Kaggle or entirely in one notebook, but is certainly an option with the integration.</p>\n\n<p>The firm code requirement is that the submission.csv be generated out of a notebook. And that submission’s notebook will be constrained to the Code Requirements stated on the overview page. We won’t be re-running your notebook, but the csv submission has to come out of a notebook to be eligible for submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784317,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/24/2020 05:26:04",
          "content": "<p>But in my little experience with TPU's, they prefer using data from GCP Buckets for better performance, how do we link that here Julia?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784931,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/24/2020 16:07:57",
          "content": "<p>Using <code>KaggleDatasets().get_gcs_path()</code> as in the <a href=\"https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started\">Getting Started notebook</a>.\nThe function copies the dataset to a GCS bucket close to the TPU you have been assigned.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784945,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/24/2020 16:19:43",
          "content": "<p>This is super cool! :partyparrot:</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785319,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/25/2020 01:00:20",
          "content": "<p>But I don't think loading from GCS offers a perf advantage for PyTorch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 784673,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/24/2020 12:30:59",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 785361,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/25/2020 02:00:35",
          "content": "<p>I am aware of the them cpmp :); Rather was probably the first one here to play around along with <a href=\"/tanlikesmath\">@tanlikesmath</a> for Flowers TPU Comp :) (wrt PyTroch-XLA)\nThanks for sharing but cpmp!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 794812,
      "author_name": "hemanth007",
      "author_url": "",
      "post_date": "04/02/2020 05:36:21",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 873240,
      "author_name": "tyoc213",
      "author_url": "",
      "post_date": "06/04/2020 00:31:40",
      "content": "<p>Installed as requested on the notebook</p>\n\n<p><code>\nVERSION = \"20200515\" #\"20200325\"  #@param [\"1.5\" , \"20200325\", \"nightly\"]\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code></p>\n\n<p>With and without the extra <code>--apt-packages</code> causes this error... I also tried to install with <code>VERSION=\"20200325\"</code> with less luck :) (because it could not find the repos I guess)</p>\n\n<p>```</p>\n\n<p>version 20200515\n  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current\n                                 Dload  Upload   Total   Spent    Left  Speed\n100  4264  100  4264    0     0  18539      0 --:--:-- --:--:-- --:--:-- 18539\n------------------------------------------------------------- <strong><em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>**<em>*</em></strong> -------------------------------------------------------------\nUpdating TPU and VM. This may take around 2 minutes.\nUpdating TPU runtime to pytorch-dev20200515 ...\nFound existing installation: torch 1.6.0a0+bf2bbd9\nUninstalling torch-1.6.0a0+bf2bbd9:\n  Successfully uninstalled torch-1.6.0a0+bf2bbd9\nFound existing installation: torchvision 0.7.0a0+a6073f0\nUninstalling torchvision-0.7.0a0+a6073f0:\n  Successfully uninstalled torchvision-0.7.0a0+a6073f0\nCopying gs://tpu-pytorch/wheels/torch-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n- [1 files][ 91.0 MiB/ 91.0 MiB] <br>\nOperation completed over 1 objects/91.0 MiB. <br>\nCopying gs://tpu-pytorch/wheels/torch_xla-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n\\ [1 files][119.5 MiB/119.5 MiB] <br>\nOperation completed over 1 objects/119.5 MiB. <br>\nCopying gs://tpu-pytorch/wheels/torchvision-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n/ [1 files][  2.3 MiB/  2.3 MiB] <br>\nOperation completed over 1 objects/2.3 MiB. <br>\nProcessing ./torch-nightly+20200515-cp37-cp37m-linux_x86_64.whl\nDone updating TPU runtime: \nRequirement already satisfied: numpy in /opt/conda/lib/python3.7/site-packages (from torch==nightly+20200515) (1.18.1)\nRequirement already satisfied: future in /opt/conda/lib/python3.7/site-packages (from torch==nightly+20200515) (0.18.2)\nERROR: fastai 1.0.61 requires torchvision, which is not installed.\nERROR: kornia 0.3.1 has requirement torch==1.5.0, but you'll have torch 1.6.0a0+bf2bbd9 which is incompatible.\nERROR: allennlp 0.9.0 has requirement spacy&lt;2.2,&gt;=2.1.0, but you'll have spacy 2.2.3 which is incompatible.\nInstalling collected packages: torch\nSuccessfully installed torch-1.6.0a0+bf2bbd9\nWARNING: You are using pip version 20.1; however, version 20.1.1 is available.\nYou should consider upgrading via the '/opt/conda/bin/python3.7 -m pip install --upgrade pip' command.\nProcessing ./torch_xla-nightly+20200515-cp37-cp37m-linux_x86_64.whl\nInstalling collected packages: torch-xla\n  Attempting uninstall: torch-xla\n    Found existing installation: torch-xla 1.6+2b2085a\n    Uninstalling torch-xla-1.6+2b2085a:\n      Successfully uninstalled torch-xla-1.6+2b2085a\nSuccessfully installed torch-xla-1.6+2b2085a\nWARNING: You are using pip version 20.1; however, version 20.1.1 is available.\nYou should consider upgrading via the '/opt/conda/bin/python3.7 -m pip install --upgrade pip' command.\nProcessing ./torchvision-nightly+20200515-cp37-cp37m-linux_x86_64.whl\nRequirement already satisfied: torch in /opt/conda/lib/python3.7/site-packages (from torchvision==nightly+20200515) (1.6.0a0+bf2bbd9)\nRequirement already satisfied: numpy in /opt/conda/lib/python3.7/site-packages (from torchvision==nightly+20200515) (1.18.1)\nRequirement already satisfied: pillow&gt;=4.1.1 in /opt/conda/lib/python3.7/site-packages (from torchvision==nightly+20200515) (5.4.1)\nRequirement already satisfied: future in /opt/conda/lib/python3.7/site-packages (from torch-&gt;torchvision==nightly+20200515) (0.18.2)\nInstalling collected packages: torchvision\nSuccessfully installed torchvision-0.7.0a0+a6073f0\nWARNING: You are using pip version 20.1; however, version 20.1.1 is available.\nYou should consider upgrading via the '/opt/conda/bin/python3.7 -m pip install --upgrade pip' command.\nReading package lists... Done\nBuilding dependency tree <br>\nReading state information... Done\nlibomp5 is already the newest version (5.0.1-1).\nThe following additional packages will be installed:\n  libgfortran4 libopenblas-base\nThe following NEW packages will be installed:\n  libgfortran4 libopenblas-base libopenblas-dev\n0 upgraded, 3 newly installed, 0 to remove and 17 not upgraded.\nNeed to get 8316 kB of archives.\nAfter this operation, 96.8 MB of additional disk space will be used.\nGet:1 <a href=\"http://archive.ubuntu.com/ubuntu\">http://archive.ubuntu.com/ubuntu</a> bionic-updates/main amd64 libgfortran4 amd64 7.5.0-3ubuntu1~18.04 [492 kB]\nGet:2 <a href=\"http://archive.ubuntu.com/ubuntu\">http://archive.ubuntu.com/ubuntu</a> bionic/universe amd64 libopenblas-base amd64 0.2.20+ds-4 [3964 kB]\nGet:3 <a href=\"http://archive.ubuntu.com/ubuntu\">http://archive.ubuntu.com/ubuntu</a> bionic/universe amd64 libopenblas-dev amd64 0.2.20+ds-4 [3860 kB]\nFetched 8316 kB in 0s (39.2 MB/s) <br>\ndebconf: delaying package configuration, since apt-utils is not installed\nSelecting previously unselected package libgfortran4:amd64.\n(Reading database ... 105812 files and directories currently installed.)\nPreparing to unpack .../libgfortran4_7.5.0-3ubuntu1~18.04_amd64.deb ...\nUnpacking libgfortran4:amd64 (7.5.0-3ubuntu1~18.04) ...\nSelecting previously unselected package libopenblas-base:amd64.\nPreparing to unpack .../libopenblas-base_0.2.20+ds-4_amd64.deb ...\nUnpacking libopenblas-base:amd64 (0.2.20+ds-4) ...\nSelecting previously unselected package libopenblas-dev:amd64.\nPreparing to unpack .../libopenblas-dev_0.2.20+ds-4_amd64.deb ...\nUnpacking libopenblas-dev:amd64 (0.2.20+ds-4) ...\nSetting up libgfortran4:amd64 (7.5.0-3ubuntu1~18.04) ...\nSetting up libopenblas-base:amd64 (0.2.20+ds-4) ...\nupdate-alternatives: using /usr/lib/x86_64-linux-gnu/openblas/libblas.so.3 to provide /usr/lib/x86_64-linux-gnu/libblas.so.3 (libblas.so.3-x86_64-linux-gnu) in auto mode\nupdate-alternatives: using /usr/lib/x86_64-linux-gnu/openblas/liblapack.so.3 to provide /usr/lib/x86_64-linux-gnu/liblapack.so.3 (liblapack.so.3-x86_64-linux-gnu) in auto mode\nSetting up libopenblas-dev:amd64 (0.2.20+ds-4) ...\nupdate-alternatives: using /usr/lib/x86_64-linux-gnu/openblas/libblas.so to provide /usr/lib/x86_64-linux-gnu/libblas.so (libblas.so-x86_64-linux-gnu) in auto mode\nupdate-alternatives: using /usr/lib/x86_64-linux-gnu/openblas/liblapack.so to provide /usr/lib/x86_64-linux-gnu/liblapack.so (liblapack.so-x86_64-linux-gnu) in auto mode\nProcessing triggers for libc-bin (2.27-3ubuntu1) ...\n------------------------------------------------------------- <strong><em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>**<em>*</em></strong> -------------------------------------------------------------</p>\n\n<p>```</p>\n\n<p>Thanks for any input on how to run on TPU+pytorch.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "784046": "Do we have XLA support now on kaggle TPU's  for PyTorch?",
    "784053": "PyTorch is not yet officially supported, but you can demo our current workaround as follows:\n\n```\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1314380%2Fbbb90f9a1220e92b3b5b756de7fa3dd0%2FScreen%20Shot%202020-03-23%20at%204.57.57%20PM.png?generation=1585004290331820&amp;alt=media)",
    "784066": "Sounds Great!🔥🔥🔥",
    "784072": "So we can now train with PyTorch on Kaggle TPUs?",
    "784076": "I've seen a handful of successful examples, yes.  Please share a public notebook if you get something up and running!",
    "784083": "Sounds good. I am looking forward to using PyTorch for TPU training in this competition!",
    "784084": "Yes, but please be aware that:\n- The [PyTorch-XLA](https://github.com/pytorch/xla) has not yet released their first stable release. They are working towards that goal.\n- With PyTorch, the Kaggle VM is feeding data to the TPU directly. With a relatively small VM, the TPU can end up being starved of data. However, in an NLP competition like this one, the data is made of numerical tokens (tokenized words) and is therefore small. It should fit in memory and you should not have an issue with bandwidth.",
    "784093": "mgornergoogle Thanks for the info. I have used PyTorch XLA successfully in the past so I am not worried about the fact that there's no stable release yet. Instead, I was more worried if Kaggle would support PyTorch XLA, and if there would be any issues with data loading and optimizing TPU usage. But as you mentioned, the data should be small enough that this won't be too much of a problem.",
    "784099": "never worked with `TPU` but it sounds exciting! Since many people at home have `GPU`  I was wondering if it easy to train model using `GPU` and than use `TPU` for inference ? \n\nasking for a friend who has Pytorch =)",
    "784140": "30 TPU hours are free on Kaggle. Just curious why you would want to train on GPU if it's equally easy (especially since it is faster on TPU). I wouldn't see why you couldn't use TPU for inference though.",
    "784141": "Sure that should work. You can also submit with GPU.\nHowever the goal of TPUs is to iterate faster. Using the [Getting started notebook](https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started):\n- TPU: 58 sec per 500 steps (500 steps = 128,000 sentences at batch size 256)\n- GPU: 246 sec per 500 steps (500 steps = 16,000 sentences at batch size 32)",
    "784143": "`30 TPU hours are free on Kaggle. Just curious why you would want to train on GPU if it's equally easy (especially since it is faster on TPU). I wouldn't see why you couldn't use TPU for inference though.`\n\nYes just to clearly my question, I think we are on the same page.. I want to train locally on GPU (since it easier to do experiments and test out stuff) and do inference using TPU =)",
    "784144": "Thank you @mgornergoogle !",
    "784147": "Oh ok. I just disagree with the fact that it's \"easier to do experiments and test out stuff\", since TPU is faster. But I guess that depends on your setup.",
    "784276": "I have a concern that is not mentioned in the Requirements. Do we allow to use the external dataset including our local pretrained models? Or Training and inference should be made by the kernel?",
    "784307": "backaggle External data is permitted. Including offline trained model(s) that are loaded as external data. Standard publicly available pretrained models count as External data that needs to be declared on the official thread. Note also that for external data to be available to run on Kaggle’s TPUs in a submission notebook, that external data must also be made public.\n\nNote that the TPU star prize, however, requires that you use TPUs end to end, including on training. That might not necessarily mean it happen in Kaggle or entirely in one notebook, but is certainly an option with the integration.\n\nThe firm code requirement is that the submission.csv be generated out of a notebook. And that submission’s notebook will be constrained to the Code Requirements stated on the overview page. We won’t be re-running your notebook, but the csv submission has to come out of a notebook to be eligible for submission.",
    "784317": "But in my little experience with TPU's, they prefer using data from GCP Buckets for better performance, how do we link that here Julia?",
    "784673": "See https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271",
    "784931": "Using `KaggleDatasets().get_gcs_path()` as in the [Getting Started notebook](https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started).\nThe function copies the dataset to a GCS bucket close to the TPU you have been assigned.",
    "784945": "This is super cool! :partyparrot:",
    "785319": "But I don't think loading from GCS offers a perf advantage for PyTorch.",
    "785361": "I am aware of the them cpmp :); Rather was probably the first one here to play around along with @tanlikesmath for Flowers TPU Comp :) (wrt PyTroch-XLA)\nThanks for sharing but cpmp!",
    "794812": "Thanks for sharing.",
    "859985": "Using the above installs has resulted in this error when torch_xla is imported:\n\nImportError: /opt/conda/lib/python3.7/site-packages/_XLAC.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZN5torch11CppFunctionC1EN3c1014KernelFunctionESt10unique_ptrINS1_14FunctionSchemaESt14default_deleteIS4_EE\n\nAny idea what the issue is? Anyones help would be appreciated",
    "873240": "Installed as requested on the notebook\n\n```\nVERSION = \"20200515\" #\"20200325\"  #@param [\"1.5\" , \"20200325\", \"nightly\"]\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n```\n\nWith and without the extra `--apt-packages` causes this error... I also tried to install with `VERSION=\"20200325\"` with less luck :) (because it could not find the repos I guess)\n\n```\n\nversion 20200515\n  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current\n                                 Dload  Upload   Total   Spent    Left  Speed\n100  4264  100  4264    0     0  18539      0 --:--:-- --:--:-- --:--:-- 18539\n------------------------------------------------------------- *************************** -------------------------------------------------------------\nUpdating TPU and VM. This may take around 2 minutes.\nUpdating TPU runtime to pytorch-dev20200515 ...\nFound existing installation: torch 1.6.0a0+bf2bbd9\nUninstalling torch-1.6.0a0+bf2bbd9:\n  Successfully uninstalled torch-1.6.0a0+bf2bbd9\nFound existing installation: torchvision 0.7.0a0+a6073f0\nUninstalling torchvision-0.7.0a0+a6073f0:\n  Successfully uninstalled torchvision-0.7.0a0+a6073f0\nCopying gs://tpu-pytorch/wheels/torch-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n- [1 files][ 91.0 MiB/ 91.0 MiB]                                                \nOperation completed over 1 objects/91.0 MiB.                                     \nCopying gs://tpu-pytorch/wheels/torch_xla-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n\\ [1 files][119.5 MiB/119.5 MiB]                                                \nOperation completed over 1 objects/119.5 MiB.                                    \nCopying gs://tpu-pytorch/wheels/torchvision-nightly+20200515-cp37-cp37m-linux_x86_64.whl...\n/ [1 files][  2.3 MiB/  2.3 MiB]                                                \nOperation completed over 1 objects/2.3 MiB.                                      \nProcessing ./torch-nightly+20200515-cp37-cp37m-linux_x86_64.whl\nDone updating TPU runtime:",
    "1327290": "Running into the same problem. Did you manage to find a workaround?",
    "1686049": "I always end up getting this error. Any suggestions?\n\n`\nImportError: /opt/conda/lib/python3.7/site-packages/_XLAC.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZN5torch11CppFunctionC1EN3c1014KernelFunctionESt10unique_ptrINS1_14FunctionSchemaESt14default_deleteIS4_EE\n`\n\nI have ran - \n\n`\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n`"
  },
  "source": "meta"
}