{
  "id": 177531,
  "title": "Pytorch XLA",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/177531",
  "author_name": "",
  "post_date": "2020-08-26T08:08:53.106742400Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I cannot install Pytorch XLA TPU but it's not working. Anyone can help?</p>",
  "messages": [
    {
      "id": "986119",
      "postDate": "08/26/2020 08:08:53",
      "content": "<p>I cannot install Pytorch XLA TPU but it's not working. Anyone can help?</p>",
      "rawMarkdown": "I cannot install Pytorch XLA TPU but it's not working. Anyone can help?",
      "votes": null
    },
    {
      "id": "986275",
      "postDate": "08/26/2020 10:35:00",
      "content": "<p>That's very bad if someone downvote because they don't understand the question. Be Contribute!!! plz</p>",
      "rawMarkdown": "That's very bad if someone downvote because they don't understand the question. Be Contribute!!! plz",
      "votes": null
    },
    {
      "id": "986586",
      "postDate": "08/26/2020 15:57:35",
      "content": "<p>Maybe its because internet is not allowed in this comp? I dont really know, just something i thought about! :)</p>",
      "rawMarkdown": "Maybe its because internet is not allowed in this comp? I dont really know, just something i thought about! :)",
      "votes": null
    },
    {
      "id": "987030",
      "postDate": "08/26/2020 23:24:03",
      "content": "<p>You need to run:</p>\n<pre><code>!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code></pre>\n<p>attach the <code>kaggle_l5kit</code> dataset then run:</p>\n<pre><code>import os\n\n## this script transports l5kit and dependencies\nos.system('pip uninstall typing -y')\nos.system('pip install --ignore-installed --target=/kaggle/working l5kit')\n</code></pre>\n<p>It takes about 4-5 minutes to install the kaggle _l5kit using the online approach so be patient. You must run the cell commands in that order - XLA first, then kaggle_l5kit, otherwise it will fail. </p>\n<p>I can get it to train with a TPU but its about 20x slower than running on a CPU. Not sure if there's a bug or perhaps I am implementing the code incorrectly.</p>",
      "rawMarkdown": "You need to run:\n\n```\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n```\n\nattach the `kaggle_l5kit` dataset then run:\n\n```\nimport os\n\n## this script transports l5kit and dependencies\nos.system('pip uninstall typing -y')\nos.system('pip install --ignore-installed --target=/kaggle/working l5kit')\n```\n\nIt takes about 4-5 minutes to install the kaggle _l5kit using the online approach so be patient. You must run the cell commands in that order - XLA first, then kaggle_l5kit, otherwise it will fail. \n\nI can get it to train with a TPU but its about 20x slower than running on a CPU. Not sure if there's a bug or perhaps I am implementing the code incorrectly.",
      "votes": null
    },
    {
      "id": "987953",
      "postDate": "08/27/2020 16:37:23",
      "content": "<p>You probably don't want to use the <code>pytorch_xla</code>master (which you're selecting by using that <code>env_setup.py</code> URL, that specifies specific versions of PyTorch and the TPU runtime). I'd use either:<br>\n<code>https://github.com/pytorch/xla/blob/v1.5.0/contrib/scripts/env-setup.py</code> for PyTorch 1.5 (the version of PyTorch l5kit is tested against)<br>\nOr maybe:<br>\n<code>https://github.com/pytorch/xla/blob/v1.6.0/contrib/scripts/env-setup.py</code> for the PyTorch 1.6 version. That may cause issues with <code>l5kit</code> which specifically doesn't support PyTorch 1.6 as not tested. But it does have some possibly useful <code>pytorch_xla</code> improvements (from the release notes, haven't actually tried it yet). Though actually one of the improvements in the 1.6 version is you no longer need to use <code>env_setup.py</code> on colab/Kaggle as noted <a href=\"https://github.com/pytorch/xla/releases/tag/v1.6.0\" target=\"_blank\">here</a> which links to new install instructions.</p>\n<p>Slowdown may also be due to data loading being a major bottleneck. Locally data loading is taking 50-60% as long as GPU, with a fairly low-end CPU feeding an RTX2070, but still 6 cores not the 2 you get with Kaggle when using TPU. So not sure the 2 cores on Kaggle could hope to keep up with a TPU. The default <code>l5kit</code> data loader config also uses 16 workers so I'd try reducing that given only 2 cores (and another 8 processes to feed the TPU if you're using multiprocessing for TPU as is generally recommended).</p>",
      "rawMarkdown": "You probably don't want to use the `pytorch_xla`master (which you're selecting by using that `env_setup.py` URL, that specifies specific versions of PyTorch and the TPU runtime). I'd use either:\n`https://github.com/pytorch/xla/blob/v1.5.0/contrib/scripts/env-setup.py` for PyTorch 1.5 (the version of PyTorch l5kit is tested against)\nOr maybe:\n`https://github.com/pytorch/xla/blob/v1.6.0/contrib/scripts/env-setup.py` for the PyTorch 1.6 version. That may cause issues with `l5kit` which specifically doesn't support PyTorch 1.6 as not tested. But it does have some possibly useful `pytorch_xla` improvements (from the release notes, haven't actually tried it yet). Though actually one of the improvements in the 1.6 version is you no longer need to use `env_setup.py` on colab/Kaggle as noted [here](https://github.com/pytorch/xla/releases/tag/v1.6.0) which links to new install instructions.\n\nSlowdown may also be due to data loading being a major bottleneck. Locally data loading is taking 50-60% as long as GPU, with a fairly low-end CPU feeding an RTX2070, but still 6 cores not the 2 you get with Kaggle when using TPU. So not sure the 2 cores on Kaggle could hope to keep up with a TPU. The default `l5kit` data loader config also uses 16 workers so I'd try reducing that given only 2 cores (and another 8 processes to feed the TPU if you're using multiprocessing for TPU as is generally recommended).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 986275,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "08/26/2020 10:35:00",
      "content": "<p>That's very bad if someone downvote because they don't understand the question. Be Contribute!!! plz</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 986586,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "08/26/2020 15:57:35",
      "content": "<p>Maybe its because internet is not allowed in this comp? I dont really know, just something i thought about! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 987030,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "08/26/2020 23:24:03",
      "content": "<p>You need to run:</p>\n<pre><code>!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n</code></pre>\n<p>attach the <code>kaggle_l5kit</code> dataset then run:</p>\n<pre><code>import os\n\n## this script transports l5kit and dependencies\nos.system('pip uninstall typing -y')\nos.system('pip install --ignore-installed --target=/kaggle/working l5kit')\n</code></pre>\n<p>It takes about 4-5 minutes to install the kaggle _l5kit using the online approach so be patient. You must run the cell commands in that order - XLA first, then kaggle_l5kit, otherwise it will fail. </p>\n<p>I can get it to train with a TPU but its about 20x slower than running on a CPU. Not sure if there's a bug or perhaps I am implementing the code incorrectly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 987953,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "08/27/2020 16:37:23",
          "content": "<p>You probably don't want to use the <code>pytorch_xla</code>master (which you're selecting by using that <code>env_setup.py</code> URL, that specifies specific versions of PyTorch and the TPU runtime). I'd use either:<br>\n<code>https://github.com/pytorch/xla/blob/v1.5.0/contrib/scripts/env-setup.py</code> for PyTorch 1.5 (the version of PyTorch l5kit is tested against)<br>\nOr maybe:<br>\n<code>https://github.com/pytorch/xla/blob/v1.6.0/contrib/scripts/env-setup.py</code> for the PyTorch 1.6 version. That may cause issues with <code>l5kit</code> which specifically doesn't support PyTorch 1.6 as not tested. But it does have some possibly useful <code>pytorch_xla</code> improvements (from the release notes, haven't actually tried it yet). Though actually one of the improvements in the 1.6 version is you no longer need to use <code>env_setup.py</code> on colab/Kaggle as noted <a href=\"https://github.com/pytorch/xla/releases/tag/v1.6.0\" target=\"_blank\">here</a> which links to new install instructions.</p>\n<p>Slowdown may also be due to data loading being a major bottleneck. Locally data loading is taking 50-60% as long as GPU, with a fairly low-end CPU feeding an RTX2070, but still 6 cores not the 2 you get with Kaggle when using TPU. So not sure the 2 cores on Kaggle could hope to keep up with a TPU. The default <code>l5kit</code> data loader config also uses 16 workers so I'd try reducing that given only 2 cores (and another 8 processes to feed the TPU if you're using multiprocessing for TPU as is generally recommended).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "986119": "I cannot install Pytorch XLA TPU but it's not working. Anyone can help?",
    "986275": "That's very bad if someone downvote because they don't understand the question. Be Contribute!!! plz",
    "986586": "Maybe its because internet is not allowed in this comp? I dont really know, just something i thought about! :)",
    "987030": "You need to run:\n\n```\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --apt-packages libomp5 libopenblas-dev\n```\n\nattach the `kaggle_l5kit` dataset then run:\n\n```\nimport os\n\n## this script transports l5kit and dependencies\nos.system('pip uninstall typing -y')\nos.system('pip install --ignore-installed --target=/kaggle/working l5kit')\n```\n\nIt takes about 4-5 minutes to install the kaggle _l5kit using the online approach so be patient. You must run the cell commands in that order - XLA first, then kaggle_l5kit, otherwise it will fail. \n\nI can get it to train with a TPU but its about 20x slower than running on a CPU. Not sure if there's a bug or perhaps I am implementing the code incorrectly.",
    "987953": "You probably don't want to use the `pytorch_xla`master (which you're selecting by using that `env_setup.py` URL, that specifies specific versions of PyTorch and the TPU runtime). I'd use either:\n`https://github.com/pytorch/xla/blob/v1.5.0/contrib/scripts/env-setup.py` for PyTorch 1.5 (the version of PyTorch l5kit is tested against)\nOr maybe:\n`https://github.com/pytorch/xla/blob/v1.6.0/contrib/scripts/env-setup.py` for the PyTorch 1.6 version. That may cause issues with `l5kit` which specifically doesn't support PyTorch 1.6 as not tested. But it does have some possibly useful `pytorch_xla` improvements (from the release notes, haven't actually tried it yet). Though actually one of the improvements in the 1.6 version is you no longer need to use `env_setup.py` on colab/Kaggle as noted [here](https://github.com/pytorch/xla/releases/tag/v1.6.0) which links to new install instructions.\n\nSlowdown may also be due to data loading being a major bottleneck. Locally data loading is taking 50-60% as long as GPU, with a fairly low-end CPU feeding an RTX2070, but still 6 cores not the 2 you get with Kaggle when using TPU. So not sure the 2 cores on Kaggle could hope to keep up with a TPU. The default `l5kit` data loader config also uses 16 workers so I'd try reducing that given only 2 cores (and another 8 processes to feed the TPU if you're using multiprocessing for TPU as is generally recommended)."
  },
  "source": "meta"
}