{
  "id": 230376,
  "title": "How to write efficient PyTorch code on TPU 🔥🔥",
  "url": "/competitions/bms-molecular-translation/discussion/230376",
  "author_name": "",
  "post_date": "2021-04-03T11:13:49.387827900Z",
  "votes": 16,
  "comment_count": 9,
  "views": 0,
  "content": "<p>As it takes a long time to train and infer models on a single Kaggle GPU (I don't have local machine), I was thinking to reformat my code and run it in parallel on 8 available TPU cores. I searched for PyTorch on TPU tutorials and the best kernels I found were the following:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\" target=\"_blank\">The Ultimate PyTorch+TPU Tutorial (Jigsaw XLM-R)🔥</a> by <a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a> </li>\n<li><a href=\"https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel\" target=\"_blank\">Super-duper fast pytorch tpu kernel… 🔥🔥🔥🔥🔥</a> by <a href=\"https://www.kaggle.com/abhishek\" target=\"_blank\">@abhishek</a> </li>\n<li><a href=\"https://www.kaggle.com/nachiket273/pytorch-tpu-vision-transformer\" target=\"_blank\">pytorch_tpu_vision_transformer</a> by <a href=\"https://www.kaggle.com/nachiket273\" target=\"_blank\">@nachiket273</a> </li>\n</ol>\n<p>These are really great resources and I thank the authors for these awesome material.</p>\n<p>But, as you've noticed, the most common models in the current competition (CNN+RNN, CNN+Transformer, …) are more complicated than a single classification model and writing efficient PyTorch code which runs smoothly on TPU has been really difficult for me and I have failed so far to do so. I'm using CNN + RNN and I still have not been able to run it on 8 TPU cores simultaneously. </p>\n<p>I wanted to ask you that if possible, could you introduce more resources which experiment with more complicated models? Or some important notes that I should keep in mind on which part of the model and data is better to be on CPU and which part on TPU? Any advice in general on PyTorch on TPU?! <br>\nI'm really getting disappointed with this competition because I know I don't have enough computational resources to survive . Using TPU is my last chance!<br>\nThanks in advance!</p>",
  "messages": [
    {
      "id": "1261710",
      "postDate": "04/03/2021 11:13:49",
      "content": "<p>As it takes a long time to train and infer models on a single Kaggle GPU (I don't have local machine), I was thinking to reformat my code and run it in parallel on 8 available TPU cores. I searched for PyTorch on TPU tutorials and the best kernels I found were the following:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\" target=\"_blank\">The Ultimate PyTorch+TPU Tutorial (Jigsaw XLM-R)🔥</a> by <a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a> </li>\n<li><a href=\"https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel\" target=\"_blank\">Super-duper fast pytorch tpu kernel… 🔥🔥🔥🔥🔥</a> by <a href=\"https://www.kaggle.com/abhishek\" target=\"_blank\">@abhishek</a> </li>\n<li><a href=\"https://www.kaggle.com/nachiket273/pytorch-tpu-vision-transformer\" target=\"_blank\">pytorch_tpu_vision_transformer</a> by <a href=\"https://www.kaggle.com/nachiket273\" target=\"_blank\">@nachiket273</a> </li>\n</ol>\n<p>These are really great resources and I thank the authors for these awesome material.</p>\n<p>But, as you've noticed, the most common models in the current competition (CNN+RNN, CNN+Transformer, …) are more complicated than a single classification model and writing efficient PyTorch code which runs smoothly on TPU has been really difficult for me and I have failed so far to do so. I'm using CNN + RNN and I still have not been able to run it on 8 TPU cores simultaneously. </p>\n<p>I wanted to ask you that if possible, could you introduce more resources which experiment with more complicated models? Or some important notes that I should keep in mind on which part of the model and data is better to be on CPU and which part on TPU? Any advice in general on PyTorch on TPU?! <br>\nI'm really getting disappointed with this competition because I know I don't have enough computational resources to survive . Using TPU is my last chance!<br>\nThanks in advance!</p>",
      "rawMarkdown": "As it takes a long time to train and infer models on a single Kaggle GPU (I don't have local machine), I was thinking to reformat my code and run it in parallel on 8 available TPU cores. I searched for PyTorch on TPU tutorials and the best kernels I found were the following:\n\n1. [The Ultimate PyTorch+TPU Tutorial (Jigsaw XLM-R)🔥](https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r) by @tanlikesmath \n2. [Super-duper fast pytorch tpu kernel... 🔥🔥🔥🔥🔥](https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel) by @abhishek \n3. [pytorch_tpu_vision_transformer](https://www.kaggle.com/nachiket273/pytorch-tpu-vision-transformer) by @nachiket273 \n\nThese are really great resources and I thank the authors for these awesome material.\n\n\nBut, as you've noticed, the most common models in the current competition (CNN+RNN, CNN+Transformer, ...) are more complicated than a single classification model and writing efficient PyTorch code which runs smoothly on TPU has been really difficult for me and I have failed so far to do so. I'm using CNN + RNN and I still have not been able to run it on 8 TPU cores simultaneously. \n\nI wanted to ask you that if possible, could you introduce more resources which experiment with more complicated models? Or some important notes that I should keep in mind on which part of the model and data is better to be on CPU and which part on TPU? Any advice in general on PyTorch on TPU?! \nI'm really getting disappointed with this competition because I know I don't have enough computational resources to survive . Using TPU is my last chance!\nThanks in advance!",
      "votes": null
    },
    {
      "id": "1262481",
      "postDate": "04/04/2021 11:05:22",
      "content": "<p>Hey Moein Shariatnia, I also faced problems with GPU. Let build the TPU version. I already start to convert the Y.Nakama<br>\nGPU version of code into TPU but faces so many errors. </p>",
      "rawMarkdown": "Hey Moein Shariatnia, I also faced problems with GPU. Let build the TPU version. I already start to convert the Y.Nakama\nGPU version of code into TPU but faces so many errors.",
      "votes": null
    },
    {
      "id": "1262512",
      "postDate": "04/04/2021 11:52:26",
      "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> Yes I tried that and first I got errors but I handled them. However, even after solving the issues, it runs reallyyyy slowly on TPU and it's obvious that I have not written it efficiently. My problem is that I do not know which parts of the code is better to be on CPU and which parts on TPU. I don't know when to switch devices. </p>",
      "rawMarkdown": "aifahim Yes I tried that and first I got errors but I handled them. However, even after solving the issues, it runs reallyyyy slowly on TPU and it's obvious that I have not written it efficiently. My problem is that I do not know which parts of the code is better to be on CPU and which parts on TPU. I don't know when to switch devices.",
      "votes": null
    },
    {
      "id": "1262614",
      "postDate": "04/04/2021 14:10:24",
      "content": "<p>Have you had a look whether pytorch-lightning can help you? One of their boasts is that you need to hardly change anything to switch from GPU to GPUs to TPU(s) and they've got an example notebook for <a href=\"https://pytorch-lightning.readthedocs.io/en/latest/advanced/tpu.html#kaggle-tpus\" target=\"_blank\">Kaggle TPU usage</a>. I've not tried this much, yet, but easy support for multi-GPU and TPU training always seem like the biggest attraction of lightning to me. There's also various <a href=\"https://pytorch-lightning.readthedocs.io/en/latest/common/fast_training.html\" target=\"_blank\">tips in their documentation</a> for how to make things train fast across devices.</p>",
      "rawMarkdown": "Have you had a look whether pytorch-lightning can help you? One of their boasts is that you need to hardly change anything to switch from GPU to GPUs to TPU(s) and they've got an example notebook for [Kaggle TPU usage](https://pytorch-lightning.readthedocs.io/en/latest/advanced/tpu.html#kaggle-tpus). I've not tried this much, yet, but easy support for multi-GPU and TPU training always seem like the biggest attraction of lightning to me. There's also various [tips in their documentation](https://pytorch-lightning.readthedocs.io/en/latest/common/fast_training.html) for how to make things train fast across devices.",
      "votes": null
    },
    {
      "id": "1262622",
      "postDate": "04/04/2021 14:24:18",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Thanks for the suggestion. Actually it was the same for me; the biggest attraction of PyTorch lightning for me was this TPU training. But unfortunately, like many other examples on such frameworks, they show you how to do it on simple models and dataset (MNIST in this case!) and when you want to apply that knowledge to more complicated projects, you need to manually tune a lot of things that you say to yourself why not doing it all without that framework?!</p>\n<p>This happened for me as well; I converted all my code to PyTorch lightning to use the TPU training, but I got all kinds of errors and couldn't make it work (once the tech errors were gone, I got OOM errors which I couldn't get rid of)! </p>\n<p>I'm not saying that lightning is bad. I just say that for complicated tasks like this competition, their implementation is not efficient as well, of course they cannot think of all the use cases and they're right in this sense.</p>",
      "rawMarkdown": "bjoernholzhauer Thanks for the suggestion. Actually it was the same for me; the biggest attraction of PyTorch lightning for me was this TPU training. But unfortunately, like many other examples on such frameworks, they show you how to do it on simple models and dataset (MNIST in this case!) and when you want to apply that knowledge to more complicated projects, you need to manually tune a lot of things that you say to yourself why not doing it all without that framework?!\n\nThis happened for me as well; I converted all my code to PyTorch lightning to use the TPU training, but I got all kinds of errors and couldn't make it work (once the tech errors were gone, I got OOM errors which I couldn't get rid of)! \n\nI'm not saying that lightning is bad. I just say that for complicated tasks like this competition, their implementation is not efficient as well, of course they cannot think of all the use cases and they're right in this sense.",
      "votes": null
    },
    {
      "id": "1263107",
      "postDate": "04/05/2021 05:08:01",
      "content": "<p>both pytorch lightning and tensorflow light have problems when converting models… somehow big models converting into small compact libraries are still a serious problem… TPU is too early unfortunately… let us know if you find a solution…</p>",
      "rawMarkdown": "both pytorch lightning and tensorflow light have problems when converting models... somehow big models converting into small compact libraries are still a serious problem... TPU is too early unfortunately... let us know if you find a solution...",
      "votes": null
    },
    {
      "id": "1263142",
      "postDate": "04/05/2021 05:53:30",
      "content": "<p>Sure, I'll update this post if I found something useful.</p>",
      "rawMarkdown": "Sure, I'll update this post if I found something useful.",
      "votes": null
    },
    {
      "id": "1267613",
      "postDate": "04/08/2021 16:31:07",
      "content": "<p>To the best of my knowledge, I think it is nearly impossible to get the most out of TPUs using PyTorch at least for now.  The main bottleneck seems to be the input data pipeline. </p>\n<p>Pytorch doesn't seem to have something like the tf.data + TFRecords pipeline in Tensorflow which lets you keep your datasets in GCS buckets which sits closer to the TPU VMs and directly stream your inputs to the TPU VMs without passing through your user VM (Kaggle Kernel / Colab runtime).</p>\n<p>I believe all the toy examples in PyTorch are trying to stream the inputs from the user VM to TPU VM which seems to be not that efficient. </p>\n<p>In summary, I think your best bet now is to use Tensorflow, its tf.data API and keep your dataset in TFRecords.</p>\n<p>Having said that, there IS some work going on with pytorch XLA to overcome this bottleneck with pytorch. you can watch this GitHub issue for details and updates.<br>\n<a href=\"https://github.com/pytorch/xla/issues/1858\" target=\"_blank\">https://github.com/pytorch/xla/issues/1858</a></p>\n<p>PS: I am a big fan of PyTorch too, but stuck with Tensorflow for this sole reason ☹️</p>",
      "rawMarkdown": "To the best of my knowledge, I think it is nearly impossible to get the most out of TPUs using PyTorch at least for now.  The main bottleneck seems to be the input data pipeline. \n\nPytorch doesn't seem to have something like the tf.data + TFRecords pipeline in Tensorflow which lets you keep your datasets in GCS buckets which sits closer to the TPU VMs and directly stream your inputs to the TPU VMs without passing through your user VM (Kaggle Kernel / Colab runtime).\n\nI believe all the toy examples in PyTorch are trying to stream the inputs from the user VM to TPU VM which seems to be not that efficient. \n\nIn summary, I think your best bet now is to use Tensorflow, its tf.data API and keep your dataset in TFRecords.\n\nHaving said that, there IS some work going on with pytorch XLA to overcome this bottleneck with pytorch. you can watch this GitHub issue for details and updates.\nhttps://github.com/pytorch/xla/issues/1858\n\nPS: I am a big fan of PyTorch too, but stuck with Tensorflow for this sole reason ☹️",
      "votes": null
    },
    {
      "id": "1267676",
      "postDate": "04/08/2021 17:37:38",
      "content": "<p><a href=\"https://www.kaggle.com/nisarahamedk\" target=\"_blank\">@nisarahamedk</a> Thanks for your insightful comment! I do not much know about how TPU works and how it will run more efficiently, but the things you said made a lot of sense to me. I didn't know this important point about data pipeline. So I think I should say good for TF users! I'd rather not continuing this competition than leaving PyTorch and re-learning TF :) (joking!)</p>",
      "rawMarkdown": "nisarahamedk Thanks for your insightful comment! I do not much know about how TPU works and how it will run more efficiently, but the things you said made a lot of sense to me. I didn't know this important point about data pipeline. So I think I should say good for TF users! I'd rather not continuing this competition than leaving PyTorch and re-learning TF :) (joking!)",
      "votes": null
    },
    {
      "id": "1296784",
      "postDate": "05/07/2021 13:50:37",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/moeinshariatnia\" target=\"_blank\">@moeinshariatnia</a>,<br>\nCould you publish the tpu kernel if possible. I have been trying to convert the kernel to a tpu version but am facing too many errors.</p>",
      "rawMarkdown": "Hey @moeinshariatnia,\nCould you publish the tpu kernel if possible. I have been trying to convert the kernel to a tpu version but am facing too many errors.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1262481,
      "author_name": "aifahim",
      "author_url": "",
      "post_date": "04/04/2021 11:05:22",
      "content": "<p>Hey Moein Shariatnia, I also faced problems with GPU. Let build the TPU version. I already start to convert the Y.Nakama<br>\nGPU version of code into TPU but faces so many errors. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1262512,
          "author_name": "moeinshariatnia",
          "author_url": "",
          "post_date": "04/04/2021 11:52:26",
          "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> Yes I tried that and first I got errors but I handled them. However, even after solving the issues, it runs reallyyyy slowly on TPU and it's obvious that I have not written it efficiently. My problem is that I do not know which parts of the code is better to be on CPU and which parts on TPU. I don't know when to switch devices. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1296784,
          "author_name": "thejineaswar",
          "author_url": "",
          "post_date": "05/07/2021 13:50:37",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/moeinshariatnia\" target=\"_blank\">@moeinshariatnia</a>,<br>\nCould you publish the tpu kernel if possible. I have been trying to convert the kernel to a tpu version but am facing too many errors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1262614,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "04/04/2021 14:10:24",
      "content": "<p>Have you had a look whether pytorch-lightning can help you? One of their boasts is that you need to hardly change anything to switch from GPU to GPUs to TPU(s) and they've got an example notebook for <a href=\"https://pytorch-lightning.readthedocs.io/en/latest/advanced/tpu.html#kaggle-tpus\" target=\"_blank\">Kaggle TPU usage</a>. I've not tried this much, yet, but easy support for multi-GPU and TPU training always seem like the biggest attraction of lightning to me. There's also various <a href=\"https://pytorch-lightning.readthedocs.io/en/latest/common/fast_training.html\" target=\"_blank\">tips in their documentation</a> for how to make things train fast across devices.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1262622,
          "author_name": "moeinshariatnia",
          "author_url": "",
          "post_date": "04/04/2021 14:24:18",
          "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Thanks for the suggestion. Actually it was the same for me; the biggest attraction of PyTorch lightning for me was this TPU training. But unfortunately, like many other examples on such frameworks, they show you how to do it on simple models and dataset (MNIST in this case!) and when you want to apply that knowledge to more complicated projects, you need to manually tune a lot of things that you say to yourself why not doing it all without that framework?!</p>\n<p>This happened for me as well; I converted all my code to PyTorch lightning to use the TPU training, but I got all kinds of errors and couldn't make it work (once the tech errors were gone, I got OOM errors which I couldn't get rid of)! </p>\n<p>I'm not saying that lightning is bad. I just say that for complicated tasks like this competition, their implementation is not efficient as well, of course they cannot think of all the use cases and they're right in this sense.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263107,
          "author_name": "alincijov",
          "author_url": "",
          "post_date": "04/05/2021 05:08:01",
          "content": "<p>both pytorch lightning and tensorflow light have problems when converting models… somehow big models converting into small compact libraries are still a serious problem… TPU is too early unfortunately… let us know if you find a solution…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263142,
          "author_name": "moeinshariatnia",
          "author_url": "",
          "post_date": "04/05/2021 05:53:30",
          "content": "<p>Sure, I'll update this post if I found something useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1267613,
      "author_name": "nisarahamedk",
      "author_url": "",
      "post_date": "04/08/2021 16:31:07",
      "content": "<p>To the best of my knowledge, I think it is nearly impossible to get the most out of TPUs using PyTorch at least for now.  The main bottleneck seems to be the input data pipeline. </p>\n<p>Pytorch doesn't seem to have something like the tf.data + TFRecords pipeline in Tensorflow which lets you keep your datasets in GCS buckets which sits closer to the TPU VMs and directly stream your inputs to the TPU VMs without passing through your user VM (Kaggle Kernel / Colab runtime).</p>\n<p>I believe all the toy examples in PyTorch are trying to stream the inputs from the user VM to TPU VM which seems to be not that efficient. </p>\n<p>In summary, I think your best bet now is to use Tensorflow, its tf.data API and keep your dataset in TFRecords.</p>\n<p>Having said that, there IS some work going on with pytorch XLA to overcome this bottleneck with pytorch. you can watch this GitHub issue for details and updates.<br>\n<a href=\"https://github.com/pytorch/xla/issues/1858\" target=\"_blank\">https://github.com/pytorch/xla/issues/1858</a></p>\n<p>PS: I am a big fan of PyTorch too, but stuck with Tensorflow for this sole reason ☹️</p>",
      "votes": null,
      "replies": [
        {
          "id": 1267676,
          "author_name": "moeinshariatnia",
          "author_url": "",
          "post_date": "04/08/2021 17:37:38",
          "content": "<p><a href=\"https://www.kaggle.com/nisarahamedk\" target=\"_blank\">@nisarahamedk</a> Thanks for your insightful comment! I do not much know about how TPU works and how it will run more efficiently, but the things you said made a lot of sense to me. I didn't know this important point about data pipeline. So I think I should say good for TF users! I'd rather not continuing this competition than leaving PyTorch and re-learning TF :) (joking!)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1261710": "As it takes a long time to train and infer models on a single Kaggle GPU (I don't have local machine), I was thinking to reformat my code and run it in parallel on 8 available TPU cores. I searched for PyTorch on TPU tutorials and the best kernels I found were the following:\n\n1. [The Ultimate PyTorch+TPU Tutorial (Jigsaw XLM-R)🔥](https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r) by @tanlikesmath \n2. [Super-duper fast pytorch tpu kernel... 🔥🔥🔥🔥🔥](https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel) by @abhishek \n3. [pytorch_tpu_vision_transformer](https://www.kaggle.com/nachiket273/pytorch-tpu-vision-transformer) by @nachiket273 \n\nThese are really great resources and I thank the authors for these awesome material.\n\n\nBut, as you've noticed, the most common models in the current competition (CNN+RNN, CNN+Transformer, ...) are more complicated than a single classification model and writing efficient PyTorch code which runs smoothly on TPU has been really difficult for me and I have failed so far to do so. I'm using CNN + RNN and I still have not been able to run it on 8 TPU cores simultaneously. \n\nI wanted to ask you that if possible, could you introduce more resources which experiment with more complicated models? Or some important notes that I should keep in mind on which part of the model and data is better to be on CPU and which part on TPU? Any advice in general on PyTorch on TPU?! \nI'm really getting disappointed with this competition because I know I don't have enough computational resources to survive . Using TPU is my last chance!\nThanks in advance!",
    "1262481": "Hey Moein Shariatnia, I also faced problems with GPU. Let build the TPU version. I already start to convert the Y.Nakama\nGPU version of code into TPU but faces so many errors.",
    "1262512": "aifahim Yes I tried that and first I got errors but I handled them. However, even after solving the issues, it runs reallyyyy slowly on TPU and it's obvious that I have not written it efficiently. My problem is that I do not know which parts of the code is better to be on CPU and which parts on TPU. I don't know when to switch devices.",
    "1262614": "Have you had a look whether pytorch-lightning can help you? One of their boasts is that you need to hardly change anything to switch from GPU to GPUs to TPU(s) and they've got an example notebook for [Kaggle TPU usage](https://pytorch-lightning.readthedocs.io/en/latest/advanced/tpu.html#kaggle-tpus). I've not tried this much, yet, but easy support for multi-GPU and TPU training always seem like the biggest attraction of lightning to me. There's also various [tips in their documentation](https://pytorch-lightning.readthedocs.io/en/latest/common/fast_training.html) for how to make things train fast across devices.",
    "1262622": "bjoernholzhauer Thanks for the suggestion. Actually it was the same for me; the biggest attraction of PyTorch lightning for me was this TPU training. But unfortunately, like many other examples on such frameworks, they show you how to do it on simple models and dataset (MNIST in this case!) and when you want to apply that knowledge to more complicated projects, you need to manually tune a lot of things that you say to yourself why not doing it all without that framework?!\n\nThis happened for me as well; I converted all my code to PyTorch lightning to use the TPU training, but I got all kinds of errors and couldn't make it work (once the tech errors were gone, I got OOM errors which I couldn't get rid of)! \n\nI'm not saying that lightning is bad. I just say that for complicated tasks like this competition, their implementation is not efficient as well, of course they cannot think of all the use cases and they're right in this sense.",
    "1263107": "both pytorch lightning and tensorflow light have problems when converting models... somehow big models converting into small compact libraries are still a serious problem... TPU is too early unfortunately... let us know if you find a solution...",
    "1263142": "Sure, I'll update this post if I found something useful.",
    "1267613": "To the best of my knowledge, I think it is nearly impossible to get the most out of TPUs using PyTorch at least for now.  The main bottleneck seems to be the input data pipeline. \n\nPytorch doesn't seem to have something like the tf.data + TFRecords pipeline in Tensorflow which lets you keep your datasets in GCS buckets which sits closer to the TPU VMs and directly stream your inputs to the TPU VMs without passing through your user VM (Kaggle Kernel / Colab runtime).\n\nI believe all the toy examples in PyTorch are trying to stream the inputs from the user VM to TPU VM which seems to be not that efficient. \n\nIn summary, I think your best bet now is to use Tensorflow, its tf.data API and keep your dataset in TFRecords.\n\nHaving said that, there IS some work going on with pytorch XLA to overcome this bottleneck with pytorch. you can watch this GitHub issue for details and updates.\nhttps://github.com/pytorch/xla/issues/1858\n\nPS: I am a big fan of PyTorch too, but stuck with Tensorflow for this sole reason ☹️",
    "1267676": "nisarahamedk Thanks for your insightful comment! I do not much know about how TPU works and how it will run more efficiently, but the things you said made a lot of sense to me. I didn't know this important point about data pipeline. So I think I should say good for TF users! I'd rather not continuing this competition than leaving PyTorch and re-learning TF :) (joking!)",
    "1296784": "Hey @moeinshariatnia,\nCould you publish the tpu kernel if possible. I have been trying to convert the kernel to a tpu version but am facing too many errors."
  },
  "source": "meta"
}