{
  "id": 210537,
  "title": "PyTorch TPU/XLA starter: training, inference, submission",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210537",
  "author_name": "",
  "post_date": "2021-01-11T07:33:12.326718500Z",
  "votes": 19,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>PyTorch TPU/XLA starter: training, inference, submission</h2>\n<p>Hello all!</p>\n<p>I am a huge fan of using TPU hardware and so I have been interested in using the provided Kaggle TPUs for training models for this competition. <a href=\"https://pytorch.org/xla/\" target=\"_blank\">PyTorch XLA</a> allows us to run PyTorch code on the TPU, so we can easily easily train PyTorch models on the TPU!</p>\n<p>I have written a PyTorch XLA training starter kernel that can train all 5 folds (10 epochs) of a SE-ResNext50 model (from <a href=\"https://rwightman.github.io/pytorch-image-models/\" target=\"_blank\">timm</a>) in &lt;1.5 hr. Unfortunately, submission kernels cannot use TPUs, so I have prepared a submission kernel that uses GPU for inference.</p>\n<p><strong>PyTorch TPU 5-fold training kernel:</strong><br>\n<a href=\"https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training\" target=\"_blank\">https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training</a></p>\n<p><strong>PyTorch TPU 5-fold model inference kernel on GPU</strong>:<br>\n<a href=\"https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-gpu-inference\" target=\"_blank\">https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-gpu-inference</a></p>\n<p>Note: Currently the EfficientNet models from timm do not work as some of the PyTorch operations are not yet fully supported on the TPU. However, check out <a href=\"https://github.com/pytorch/xla/issues/2715\" target=\"_blank\">this issue</a> for updates on the situation.</p>\n<p>Also, If you want a more detailed look at PyTorch XLA and training on TPUS, please check out my <a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\" target=\"_blank\">tutorial kernel</a>.</p>\n<p>Please upvote if you enjoyed these kernels.🙂</p>",
  "messages": [
    {
      "id": "1148520",
      "postDate": "01/11/2021 07:33:12",
      "content": "<h2>PyTorch TPU/XLA starter: training, inference, submission</h2>\n<p>Hello all!</p>\n<p>I am a huge fan of using TPU hardware and so I have been interested in using the provided Kaggle TPUs for training models for this competition. <a href=\"https://pytorch.org/xla/\" target=\"_blank\">PyTorch XLA</a> allows us to run PyTorch code on the TPU, so we can easily easily train PyTorch models on the TPU!</p>\n<p>I have written a PyTorch XLA training starter kernel that can train all 5 folds (10 epochs) of a SE-ResNext50 model (from <a href=\"https://rwightman.github.io/pytorch-image-models/\" target=\"_blank\">timm</a>) in &lt;1.5 hr. Unfortunately, submission kernels cannot use TPUs, so I have prepared a submission kernel that uses GPU for inference.</p>\n<p><strong>PyTorch TPU 5-fold training kernel:</strong><br>\n<a href=\"https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training\" target=\"_blank\">https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training</a></p>\n<p><strong>PyTorch TPU 5-fold model inference kernel on GPU</strong>:<br>\n<a href=\"https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-gpu-inference\" target=\"_blank\">https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-gpu-inference</a></p>\n<p>Note: Currently the EfficientNet models from timm do not work as some of the PyTorch operations are not yet fully supported on the TPU. However, check out <a href=\"https://github.com/pytorch/xla/issues/2715\" target=\"_blank\">this issue</a> for updates on the situation.</p>\n<p>Also, If you want a more detailed look at PyTorch XLA and training on TPUS, please check out my <a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\" target=\"_blank\">tutorial kernel</a>.</p>\n<p>Please upvote if you enjoyed these kernels.🙂</p>",
      "rawMarkdown": "## PyTorch TPU/XLA starter: training, inference, submission\n\nHello all!\n\nI am a huge fan of using TPU hardware and so I have been interested in using the provided Kaggle TPUs for training models for this competition. [PyTorch XLA](https://pytorch.org/xla/) allows us to run PyTorch code on the TPU, so we can easily easily train PyTorch models on the TPU!\n\nI have written a PyTorch XLA training starter kernel that can train all 5 folds (10 epochs) of a SE-ResNext50 model (from [timm](https://rwightman.github.io/pytorch-image-models/)) in <1.5 hr. Unfortunately, submission kernels cannot use TPUs, so I have prepared a submission kernel that uses GPU for inference.\n\n**PyTorch TPU 5-fold training kernel:**\nhttps://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training\n\n**PyTorch TPU 5-fold model inference kernel on GPU**:\nhttps://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-gpu-inference\n\nNote: Currently the EfficientNet models from timm do not work as some of the PyTorch operations are not yet fully supported on the TPU. However, check out [this issue](https://github.com/pytorch/xla/issues/2715) for updates on the situation.\n\nAlso, If you want a more detailed look at PyTorch XLA and training on TPUS, please check out my [tutorial kernel](https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r).\n\nPlease upvote if you enjoyed these kernels.🙂",
      "votes": null
    },
    {
      "id": "1149503",
      "postDate": "01/11/2021 22:14:14",
      "content": "<p>Thanks! I was looking for this exact thing last week.</p>",
      "rawMarkdown": "Thanks! I was looking for this exact thing last week.",
      "votes": null
    },
    {
      "id": "1149751",
      "postDate": "01/12/2021 05:40:12",
      "content": "<p>Fixed a minor mistake and now it runs in less than 1.5 hr instead of more than 2 hrs…</p>",
      "rawMarkdown": "Fixed a minor mistake and now it runs in less than 1.5 hr instead of more than 2 hrs...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1149503,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "01/11/2021 22:14:14",
      "content": "<p>Thanks! I was looking for this exact thing last week.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1149751,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "01/12/2021 05:40:12",
      "content": "<p>Fixed a minor mistake and now it runs in less than 1.5 hr instead of more than 2 hrs…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1148520": "## PyTorch TPU/XLA starter: training, inference, submission\n\nHello all!\n\nI am a huge fan of using TPU hardware and so I have been interested in using the provided Kaggle TPUs for training models for this competition. [PyTorch XLA](https://pytorch.org/xla/) allows us to run PyTorch code on the TPU, so we can easily easily train PyTorch models on the TPU!\n\nI have written a PyTorch XLA training starter kernel that can train all 5 folds (10 epochs) of a SE-ResNext50 model (from [timm](https://rwightman.github.io/pytorch-image-models/)) in <1.5 hr. Unfortunately, submission kernels cannot use TPUs, so I have prepared a submission kernel that uses GPU for inference.\n\n**PyTorch TPU 5-fold training kernel:**\nhttps://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training\n\n**PyTorch TPU 5-fold model inference kernel on GPU**:\nhttps://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-gpu-inference\n\nNote: Currently the EfficientNet models from timm do not work as some of the PyTorch operations are not yet fully supported on the TPU. However, check out [this issue](https://github.com/pytorch/xla/issues/2715) for updates on the situation.\n\nAlso, If you want a more detailed look at PyTorch XLA and training on TPUS, please check out my [tutorial kernel](https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r).\n\nPlease upvote if you enjoyed these kernels.🙂",
    "1149503": "Thanks! I was looking for this exact thing last week.",
    "1149751": "Fixed a minor mistake and now it runs in less than 1.5 hr instead of more than 2 hrs..."
  },
  "source": "meta"
}