{
  "id": 138271,
  "title": "Using Pytorch With TPU",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/138271",
  "author_name": "",
  "post_date": "2020-03-24T11:00:29.152357500Z",
  "votes": 52,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I just shared work I did for Bengali for scoring on TPU.  Given TPU wasn't allowed for scoring I didn't test it.  Let me share code for training, I'll work and share the code for scoring ASAP.</p>\n\n<p>This uses a public dataset that contains what is needed to install Pytorch.</p>\n\n<p>Dataset: <a href=\"https://www.kaggle.com/cpmpml/torchxla\">https://www.kaggle.com/cpmpml/torchxla</a></p>\n\n<p>Test notebook: <a href=\"https://www.kaggle.com/cpmpml/tpu-test\">https://www.kaggle.com/cpmpml/tpu-test</a></p>\n\n<p>A more realistic code: <a href=\"https://www.kaggle.com/cpmpml/tpu-sub-bs256-full-train\">https://www.kaggle.com/cpmpml/tpu-sub-bs256-full-train</a></p>\n\n<p>Just look where <code>xm</code> or <code>device</code> is used to see where you code may need to be modified.</p>\n\n<p>The above uses only one TPU core.  To use the 8  cores you need parallel code.  I'll share an example but you can look at this one already: <a href=\"https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing\">https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing</a></p>\n\n<p>Please upvote the dataset if you use it ;)</p>",
  "messages": [
    {
      "id": "784587",
      "postDate": "03/24/2020 11:00:29",
      "content": "<p>I just shared work I did for Bengali for scoring on TPU.  Given TPU wasn't allowed for scoring I didn't test it.  Let me share code for training, I'll work and share the code for scoring ASAP.</p>\n\n<p>This uses a public dataset that contains what is needed to install Pytorch.</p>\n\n<p>Dataset: <a href=\"https://www.kaggle.com/cpmpml/torchxla\">https://www.kaggle.com/cpmpml/torchxla</a></p>\n\n<p>Test notebook: <a href=\"https://www.kaggle.com/cpmpml/tpu-test\">https://www.kaggle.com/cpmpml/tpu-test</a></p>\n\n<p>A more realistic code: <a href=\"https://www.kaggle.com/cpmpml/tpu-sub-bs256-full-train\">https://www.kaggle.com/cpmpml/tpu-sub-bs256-full-train</a></p>\n\n<p>Just look where <code>xm</code> or <code>device</code> is used to see where you code may need to be modified.</p>\n\n<p>The above uses only one TPU core.  To use the 8  cores you need parallel code.  I'll share an example but you can look at this one already: <a href=\"https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing\">https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing</a></p>\n\n<p>Please upvote the dataset if you use it ;)</p>",
      "rawMarkdown": "I just shared work I did for Bengali for scoring on TPU.  Given TPU wasn't allowed for scoring I didn't test it.  Let me share code for training, I'll work and share the code for scoring ASAP.\n\nThis uses a public dataset that contains what is needed to install Pytorch.\n\nDataset: https://www.kaggle.com/cpmpml/torchxla\n\nTest notebook: https://www.kaggle.com/cpmpml/tpu-test\n\nA more realistic code: https://www.kaggle.com/cpmpml/tpu-sub-bs256-full-train\n\nJust look where `xm` or `device` is used to see where you code may need to be modified.\n\nThe above uses only one TPU core.  To use the 8  cores you need parallel code.  I'll share an example but you can look at this one already: https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing\n\nPlease upvote the dataset if you use it ;)",
      "votes": null
    },
    {
      "id": "784683",
      "postDate": "03/24/2020 12:37:57",
      "content": "<p>Thanks for sharing. I should follow all of your notebooks...</p>",
      "rawMarkdown": "Thanks for sharing. I should follow all of your notebooks...",
      "votes": null
    },
    {
      "id": "784827",
      "postDate": "03/24/2020 14:29:27",
      "content": "<p>Thanks, I just made them public, they were private until today.</p>",
      "rawMarkdown": "Thanks, I just made them public, they were private until today.",
      "votes": null
    },
    {
      "id": "785018",
      "postDate": "03/24/2020 17:52:49",
      "content": "<p>How does training speed on kaggle kernels compare with TF?</p>",
      "rawMarkdown": "How does training speed on kaggle kernels compare with TF?",
      "votes": null
    },
    {
      "id": "785100",
      "postDate": "03/24/2020 19:22:32",
      "content": "<p><a href=\"/abhishek\">@abhishek</a> used a similar approcah to train on TPUs, check it out <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training\">here</a></p>",
      "rawMarkdown": "abhishek used a similar approcah to train on TPUs, check it out [here](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training)",
      "votes": null
    },
    {
      "id": "785301",
      "postDate": "03/25/2020 00:34:26",
      "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> posted a 3x faster version of his notebook here: <a href=\"https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing/notebook\">https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing/notebook</a></p>",
      "rawMarkdown": "dhananjay3 posted a 3x faster version of his notebook here: https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing/notebook",
      "votes": null
    },
    {
      "id": "785310",
      "postDate": "03/25/2020 00:37:30",
      "content": "<p>His kernel downloads pytorch each time it runs. The rest is similar indeed as we all use torch-xla api.</p>",
      "rawMarkdown": "His kernel downloads pytorch each time it runs. The rest is similar indeed as we all use torch-xla api.",
      "votes": null
    },
    {
      "id": "785311",
      "postDate": "03/25/2020 00:37:48",
      "content": "<p>training speed of what?</p>",
      "rawMarkdown": "training speed of what?",
      "votes": null
    },
    {
      "id": "785312",
      "postDate": "03/25/2020 00:38:05",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "785473",
      "postDate": "03/25/2020 05:06:54",
      "content": "<p>Here's the trouble, to use TPU's efficiently wrt PyTorch-XLA, you need more RAM on kernels;</p>",
      "rawMarkdown": "Here's the trouble, to use TPU's efficiently wrt PyTorch-XLA, you need more RAM on kernels;",
      "votes": null
    },
    {
      "id": "785485",
      "postDate": "03/25/2020 05:28:43",
      "content": "<p>There's a simpler way:</p>\n\n<p><code>import tensorflow as torch</code></p>",
      "rawMarkdown": "There's a simpler way:\n\n```import tensorflow as torch```",
      "votes": null
    },
    {
      "id": "785498",
      "postDate": "03/25/2020 05:46:51",
      "content": "<p>Ya sure! That's hell of an import :)</p>",
      "rawMarkdown": "Ya sure! That's hell of an import :)",
      "votes": null
    },
    {
      "id": "786282",
      "postDate": "03/25/2020 18:50:16",
      "content": "<p>Any way to make Pytorch <code>deterministic</code> on TPU? </p>",
      "rawMarkdown": "Any way to make Pytorch `deterministic` on TPU?",
      "votes": null
    },
    {
      "id": "787138",
      "postDate": "03/26/2020 14:47:52",
      "content": "<blockquote>\n  <p>import tensorflow as torch</p>\n</blockquote>\n\n<p>It would be great indeed if it worked with pytorch api.</p>",
      "rawMarkdown": "&gt; import tensorflow as torch\n\nIt would be great indeed if it worked with pytorch api.",
      "votes": null
    },
    {
      "id": "787334",
      "postDate": "03/26/2020 18:04:54",
      "content": "<p>ROFL</p>",
      "rawMarkdown": "ROFL",
      "votes": null
    },
    {
      "id": "794810",
      "postDate": "04/02/2020 05:34:54",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "849830",
      "postDate": "05/16/2020 05:29:21",
      "content": "<p><code>HERE</code> To make <code>TPU</code> training <code>deterministic</code>\n<code>python\ndef seed_everything(seed=396):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n</code></p>",
      "rawMarkdown": "`HERE` To make `TPU` training `deterministic`\n```python\ndef seed_everything(seed=396):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n```",
      "votes": null
    },
    {
      "id": "893754",
      "postDate": "06/19/2020 21:27:03",
      "content": "<p>Thanks for sharing. This seems to be the same code snippet for GPU-Pytorch. Is there anything specific to TPUs? 👀 </p>",
      "rawMarkdown": "Thanks for sharing. This seems to be the same code snippet for GPU-Pytorch. Is there anything specific to TPUs? 👀",
      "votes": null
    },
    {
      "id": "893755",
      "postDate": "06/19/2020 21:28:16",
      "content": "<p>It seems so. I guess being careful with  dataset creation and <code>gc.collect()</code> often are key then.</p>",
      "rawMarkdown": "It seems so. I guess being careful with  dataset creation and `gc.collect()` often are key then.",
      "votes": null
    },
    {
      "id": "934397",
      "postDate": "07/18/2020 12:09:57",
      "content": "<p>No :) same script for <code>GPU</code> and <code>TPU</code></p>",
      "rawMarkdown": "No :) same script for `GPU` and `TPU`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 784683,
      "author_name": "youhanlee",
      "author_url": "",
      "post_date": "03/24/2020 12:37:57",
      "content": "<p>Thanks for sharing. I should follow all of your notebooks...</p>",
      "votes": null,
      "replies": [
        {
          "id": 784827,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/24/2020 14:29:27",
          "content": "<p>Thanks, I just made them public, they were private until today.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 785018,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "03/24/2020 17:52:49",
      "content": "<p>How does training speed on kaggle kernels compare with TF?</p>",
      "votes": null,
      "replies": [
        {
          "id": 785311,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/25/2020 00:37:48",
          "content": "<p>training speed of what?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 785100,
      "author_name": "rimijoker",
      "author_url": "",
      "post_date": "03/24/2020 19:22:32",
      "content": "<p><a href=\"/abhishek\">@abhishek</a> used a similar approcah to train on TPUs, check it out <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 785310,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/25/2020 00:37:30",
          "content": "<p>His kernel downloads pytorch each time it runs. The rest is similar indeed as we all use torch-xla api.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 785301,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/25/2020 00:34:26",
      "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> posted a 3x faster version of his notebook here: <a href=\"https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing/notebook\">https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing/notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 785312,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/25/2020 00:38:05",
          "content": "<p>Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 785473,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "03/25/2020 05:06:54",
      "content": "<p>Here's the trouble, to use TPU's efficiently wrt PyTorch-XLA, you need more RAM on kernels;</p>",
      "votes": null,
      "replies": [
        {
          "id": 893755,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/19/2020 21:28:16",
          "content": "<p>It seems so. I guess being careful with  dataset creation and <code>gc.collect()</code> often are key then.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 785485,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "03/25/2020 05:28:43",
      "content": "<p>There's a simpler way:</p>\n\n<p><code>import tensorflow as torch</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 785498,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/25/2020 05:46:51",
          "content": "<p>Ya sure! That's hell of an import :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787138,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/26/2020 14:47:52",
          "content": "<blockquote>\n  <p>import tensorflow as torch</p>\n</blockquote>\n\n<p>It would be great indeed if it worked with pytorch api.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787334,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/26/2020 18:04:54",
          "content": "<p>ROFL</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 786282,
      "author_name": "coolcoder22",
      "author_url": "",
      "post_date": "03/25/2020 18:50:16",
      "content": "<p>Any way to make Pytorch <code>deterministic</code> on TPU? </p>",
      "votes": null,
      "replies": [
        {
          "id": 849830,
          "author_name": "coolcoder22",
          "author_url": "",
          "post_date": "05/16/2020 05:29:21",
          "content": "<p><code>HERE</code> To make <code>TPU</code> training <code>deterministic</code>\n<code>python\ndef seed_everything(seed=396):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 893754,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/19/2020 21:27:03",
          "content": "<p>Thanks for sharing. This seems to be the same code snippet for GPU-Pytorch. Is there anything specific to TPUs? 👀 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934397,
          "author_name": "coolcoder22",
          "author_url": "",
          "post_date": "07/18/2020 12:09:57",
          "content": "<p>No :) same script for <code>GPU</code> and <code>TPU</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 794810,
      "author_name": "hemanth007",
      "author_url": "",
      "post_date": "04/02/2020 05:34:54",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "784587": "I just shared work I did for Bengali for scoring on TPU.  Given TPU wasn't allowed for scoring I didn't test it.  Let me share code for training, I'll work and share the code for scoring ASAP.\n\nThis uses a public dataset that contains what is needed to install Pytorch.\n\nDataset: https://www.kaggle.com/cpmpml/torchxla\n\nTest notebook: https://www.kaggle.com/cpmpml/tpu-test\n\nA more realistic code: https://www.kaggle.com/cpmpml/tpu-sub-bs256-full-train\n\nJust look where `xm` or `device` is used to see where you code may need to be modified.\n\nThe above uses only one TPU core.  To use the 8  cores you need parallel code.  I'll share an example but you can look at this one already: https://www.kaggle.com/dhananjay3/pytorch-xla-for-tpu-with-multiprocessing\n\nPlease upvote the dataset if you use it ;)",
    "784683": "Thanks for sharing. I should follow all of your notebooks...",
    "784827": "Thanks, I just made them public, they were private until today.",
    "785018": "How does training speed on kaggle kernels compare with TF?",
    "785100": "abhishek used a similar approcah to train on TPUs, check it out [here](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training)",
    "785301": "dhananjay3 posted a 3x faster version of his notebook here: https://www.kaggle.com/dhananjay3/fast-pytorch-xla-for-tpu-with-multiprocessing/notebook",
    "785310": "His kernel downloads pytorch each time it runs. The rest is similar indeed as we all use torch-xla api.",
    "785311": "training speed of what?",
    "785312": "Thanks for sharing.",
    "785473": "Here's the trouble, to use TPU's efficiently wrt PyTorch-XLA, you need more RAM on kernels;",
    "785485": "There's a simpler way:\n\n```import tensorflow as torch```",
    "785498": "Ya sure! That's hell of an import :)",
    "786282": "Any way to make Pytorch `deterministic` on TPU?",
    "787138": "&gt; import tensorflow as torch\n\nIt would be great indeed if it worked with pytorch api.",
    "787334": "ROFL",
    "794810": "Thanks for sharing.",
    "849830": "`HERE` To make `TPU` training `deterministic`\n```python\ndef seed_everything(seed=396):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n```",
    "893754": "Thanks for sharing. This seems to be the same code snippet for GPU-Pytorch. Is there anything specific to TPUs? 👀",
    "893755": "It seems so. I guess being careful with  dataset creation and `gc.collect()` often are key then.",
    "934397": "No :) same script for `GPU` and `TPU`"
  },
  "source": "meta"
}