{
  "id": 170527,
  "title": "There is nothing more frustrating than trying to use TPUs...",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/170527",
  "author_name": "Carlos Souza",
  "post_date": "2020-07-28T05:14:10.467000",
  "votes": 15,
  "comment_count": 26,
  "views": 0,
  "content": "<p>XLA documentation is very poor. Notebook examples are either so simple they are not useful at all, or so advanced/complex it is impossible to understand what is happening. </p>\n\n<p>The code frequently freezes, and it is impossible to know what is happening in the background...</p>\n\n<p>Is this only with PyTorch, or the user experience with Tensorflow is as bad as this as well?</p>",
  "messages": [
    {
      "id": 948611,
      "postDate": "2020-07-28T05:14:10.467Z",
      "content": "<p>XLA documentation is very poor. Notebook examples are either so simple they are not useful at all, or so advanced/complex it is impossible to understand what is happening. </p>\n\n<p>The code frequently freezes, and it is impossible to know what is happening in the background...</p>\n\n<p>Is this only with PyTorch, or the user experience with Tensorflow is as bad as this as well?</p>",
      "rawMarkdown": "XLA documentation is very poor. Notebook examples are either so simple they are not useful at all, or so advanced/complex it is impossible to understand what is happening. \n\nThe code frequently freezes, and it is impossible to know what is happening in the background...\n\nIs this only with PyTorch, or the user experience with Tensorflow is as bad as this as well?",
      "votes": 14
    },
    {
      "id": 949829,
      "postDate": "2020-07-29T01:37:13.610Z",
      "content": "<p>Hi <a href=\"/carlossouza\">@carlossouza</a> --- I'm really excited that you're trying out TPUs, and I'm really sorry that you're encountering some of the challenges with PyTorch + TPUs. I know that that can be incredibly frustrating.</p>\n\n<p>There are some additional challenges to getting PyTorch up and running on TPUs, but we've pulled together some documentation <strong><a href=\"https://www.kaggle.com/docs/tpu#tpu8\">here</a></strong>. Our community also has some members who have created some excellent resources on PyTorch + TPUs---I'm not sure what you've had a chance to look at, but these are some of my favorites:\n- <strong><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159723\">PyTorch TPU Improvements</a></strong> from Psi\n- <strong><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271\">Using PyTorch with TPU</a></strong> from CPMP\n- <strong><a href=\"https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel\">Super duper fast PyTorch TPU kernel</a></strong> from Abhishek\n- <strong><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005\">PyTorch XLA/TPU training</a></strong> from ilovescience</p>\n\n<p>I'd love to know if any of these are helpful to you!</p>",
      "rawMarkdown": "Hi @carlossouza --- I'm really excited that you're trying out TPUs, and I'm really sorry that you're encountering some of the challenges with PyTorch + TPUs. I know that that can be incredibly frustrating.\n\nThere are some additional challenges to getting PyTorch up and running on TPUs, but we've pulled together some documentation **[here](https://www.kaggle.com/docs/tpu#tpu8)**. Our community also has some members who have created some excellent resources on PyTorch + TPUs---I'm not sure what you've had a chance to look at, but these are some of my favorites:\n- **[PyTorch TPU Improvements](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159723)** from Psi\n- **[Using PyTorch with TPU](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271)** from CPMP\n- **[Super duper fast PyTorch TPU kernel](https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel)** from Abhishek\n- **[PyTorch XLA/TPU training](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005)** from ilovescience\n\nI'd love to know if any of these are helpful to you!",
      "votes": 5,
      "replies": [
        {
          "id": 949850,
          "postDate": "2020-07-29T02:37:48.280Z",
          "content": "<p>Hi <a href=\"/jessemostipak\">@jessemostipak</a> ! Thanks for your prompt response!</p>\n\n<p>Yes, I've read all these tutorials several times, each. And more tutorials.</p>\n\n<p>Then, I adapted my code (which runs perfectly on GPUs), following exactly these tutorials. I specifically tried 3 from the list you provided: I followed their procedures to the letter, changing only the model, and the datasets. It failed miserably, without explicit errors: the code would take forever to train, indicating that something was obviously not right, but it didn't explicitly break.</p>\n\n<p>IMHO, free TPU credits without proper support is generating the exact opposite effect intended. Instead of generating happy users, it is generating frustrated users who'd never want to try it again. That's because of 3 reasons:</p>\n\n<p>First, it is <strong>impossible to debug</strong>. The code fails without generating any hint on what might be wrong. And filling the code with print statements trying to guess where the problem is is very very very frustrating.</p>\n\n<p>Second, this documentation that you provided is <strong>only useful if you are solving these specific problems</strong> (i.e. using the same models/datasets/etc). They do not help whatsoever in adapting the code to solve problems different than what they showcase.</p>\n\n<p>Third, and maybe the most important: there is nothing in the documentation, in forums, in github, etc, that help us <strong>when something goes wrong</strong>. Nothing. All the (very few) tutorials, examples and documents only show working code. They don't <strong>show potential mistakes/problems and how to solve them</strong>, which is exactly what we need to learn how to use TPUs.</p>\n\n<p>I'd really love to learn how to use TPUs. But after unsuccessfully trying to adapt my code, following exactly the examples provided, I don't know how to continue. I'd love a suggestion actually. Without debugging messages/tools, without documentation that either i) shows how to adapt code to different problems and/or ii) shows main mistakes/errors/problems and how to fix them (troubleshooting), I'm not sure how to continue playing with TPUs. :)</p>",
          "rawMarkdown": "Hi @jessemostipak ! Thanks for your prompt response!\n\nYes, I've read all these tutorials several times, each. And more tutorials.\n\nThen, I adapted my code (which runs perfectly on GPUs), following exactly these tutorials. I specifically tried 3 from the list you provided: I followed their procedures to the letter, changing only the model, and the datasets. It failed miserably, without explicit errors: the code would take forever to train, indicating that something was obviously not right, but it didn't explicitly break.\n\nIMHO, free TPU credits without proper support is generating the exact opposite effect intended. Instead of generating happy users, it is generating frustrated users who'd never want to try it again. That's because of 3 reasons:\n\nFirst, it is **impossible to debug**. The code fails without generating any hint on what might be wrong. And filling the code with print statements trying to guess where the problem is is very very very frustrating.\n\nSecond, this documentation that you provided is **only useful if you are solving these specific problems** (i.e. using the same models/datasets/etc). They do not help whatsoever in adapting the code to solve problems different than what they showcase.\n\nThird, and maybe the most important: there is nothing in the documentation, in forums, in github, etc, that help us **when something goes wrong**. Nothing. All the (very few) tutorials, examples and documents only show working code. They don't **show potential mistakes/problems and how to solve them**, which is exactly what we need to learn how to use TPUs.\n\nI'd really love to learn how to use TPUs. But after unsuccessfully trying to adapt my code, following exactly the examples provided, I don't know how to continue. I'd love a suggestion actually. Without debugging messages/tools, without documentation that either i) shows how to adapt code to different problems and/or ii) shows main mistakes/errors/problems and how to fix them (troubleshooting), I'm not sure how to continue playing with TPUs. :)",
          "votes": 3
        },
        {
          "id": 949852,
          "postDate": "2020-07-29T02:43:27.953Z",
          "content": "<p>Btw, here's my code running perfectly on GPUs: <a href=\"https://www.kaggle.com/carlossouza/osic-autoencoder-training\">https://www.kaggle.com/carlossouza/osic-autoencoder-training</a></p>",
          "rawMarkdown": "Btw, here's my code running perfectly on GPUs: https://www.kaggle.com/carlossouza/osic-autoencoder-training"
        },
        {
          "id": 949934,
          "postDate": "2020-07-29T04:49:54.213Z",
          "content": "<p><a href=\"/jessemostipak\">@jessemostipak</a> Thank you for sharing my kernel with others and I am glad you think it's one of community favorites! However, I might recommend that you instead share <a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\">this updated kernel</a> that explains PyTorch XLA in much more detail.</p>\n\n<p><a href=\"/carlossouza\">@carlossouza</a> Please check the above kernel and see if you get a better understanding of PyTorch XLA. It is quite general and hopefully after looking at that kernel, you could then adapt PyTorch XLA for your own problem.</p>",
          "rawMarkdown": "@jessemostipak Thank you for sharing my kernel with others and I am glad you think it's one of community favorites! However, I might recommend that you instead share [this updated kernel](https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r) that explains PyTorch XLA in much more detail.\n\n@carlossouza Please check the above kernel and see if you get a better understanding of PyTorch XLA. It is quite general and hopefully after looking at that kernel, you could then adapt PyTorch XLA for your own problem."
        },
        {
          "id": 949936,
          "postDate": "2020-07-29T04:53:46.150Z",
          "content": "<p>Also <a href=\"/carlossouza\">@carlossouza</a> please share what your specific issue is and I bet the community (including myself) would be glad to help you out.</p>",
          "rawMarkdown": "Also @carlossouza please share what your specific issue is and I bet the community (including myself) would be glad to help you out.",
          "votes": 3
        },
        {
          "id": 950745,
          "postDate": "2020-07-29T15:22:29.357Z",
          "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a> --- thank you for linking the updated kernel! </p>\n\n<p><a href=\"/carlossouza\">@carlossouza</a> --- you've hit on a lot of the pain points with PyTorch and TPUs. one of the cool things about Kaggle is that we can bring really new and exciting technologies (like TPUs!) to our community, but <em>because</em> they're new, documentation is still being developed. We're <em>all</em> learning together and building out the documentation as we go (which is exciting! but can also be frustrating.)</p>\n\n<p>TPUs were designed to work with TensorFlow, so you may encounter fewer issues and more documentation by switching, but that being said, our community is doing a phenomenal job at learning together and creating the documentation to make PyTorch on TPUs as seamless as possible.</p>\n\n<p>I love the idea of documentation that shows different approaches and//or how to deal with different types of errors, and it would be great if you helped create it! Documenting what you're trying, what you're learning, and what works is a great way to help those who in the future find themselves in similar situations.</p>\n\n<p>Our community is working hard on figuring out PyTorch and TPUs, and I strongly recommend leaning on them for insights on what could be going wrong. The forums are a great way to get an open exchange of ideas and collectively learn more.</p>",
          "rawMarkdown": "@tanlikesmath --- thank you for linking the updated kernel! \n\n@carlossouza --- you've hit on a lot of the pain points with PyTorch and TPUs. one of the cool things about Kaggle is that we can bring really new and exciting technologies (like TPUs!) to our community, but _because_ they're new, documentation is still being developed. We're _all_ learning together and building out the documentation as we go (which is exciting! but can also be frustrating.)\n\nTPUs were designed to work with TensorFlow, so you may encounter fewer issues and more documentation by switching, but that being said, our community is doing a phenomenal job at learning together and creating the documentation to make PyTorch on TPUs as seamless as possible.\n\nI love the idea of documentation that shows different approaches and//or how to deal with different types of errors, and it would be great if you helped create it! Documenting what you're trying, what you're learning, and what works is a great way to help those who in the future find themselves in similar situations.\n\nOur community is working hard on figuring out PyTorch and TPUs, and I strongly recommend leaning on them for insights on what could be going wrong. The forums are a great way to get an open exchange of ideas and collectively learn more.",
          "votes": 4
        },
        {
          "id": 950803,
          "postDate": "2020-07-29T16:21:01.627Z",
          "content": "<p>Hi Carlos, we do have this TROUBLESHOOTING (<a href=\"https://github.com/pytorch/xla/blob/master/TROUBLESHOOTING.md\">https://github.com/pytorch/xla/blob/master/TROUBLESHOOTING.md</a>) so please take a look at it. The metrics report is a very useful tool to understand what part of the execution is taking so long (compilation of graph? device step time? data transfer? data retrieval? round trips to CPU?). But thanks for the suggestions, we'll work on improving our documentation and debugging UX. Our team has been stretched and we haven't had enough time to improve that part.</p>\n\n<p>In the meantime if you face problems that the troubleshooting guide isn't able to help with much, open a Github issue on our pytorch/xla repo and we'll try to help get it resolved.</p>",
          "rawMarkdown": "Hi Carlos, we do have this TROUBLESHOOTING (https://github.com/pytorch/xla/blob/master/TROUBLESHOOTING.md) so please take a look at it. The metrics report is a very useful tool to understand what part of the execution is taking so long (compilation of graph? device step time? data transfer? data retrieval? round trips to CPU?). But thanks for the suggestions, we'll work on improving our documentation and debugging UX. Our team has been stretched and we haven't had enough time to improve that part.\n\nIn the meantime if you face problems that the troubleshooting guide isn't able to help with much, open a Github issue on our pytorch/xla repo and we'll try to help get it resolved.",
          "votes": 1
        },
        {
          "id": 951118,
          "postDate": "2020-07-29T22:25:07.780Z",
          "content": "<p><a href=\"/jysohn23\">@jysohn23</a> , <a href=\"/tanlikesmath\">@tanlikesmath</a> , <a href=\"/jessemostipak\">@jessemostipak</a> ,\nThank you all for your answers! You gave me hope, and I tried again :)\nHere's the code (still not working):</p>\n\n<p><a href=\"https://www.kaggle.com/carlossouza/osic-autoencoder-training-on-tpus\">https://www.kaggle.com/carlossouza/osic-autoencoder-training-on-tpus</a></p>\n\n<p>The GPU code works perfectly: <a href=\"https://www.kaggle.com/carlossouza/osic-autoencoder-training\">https://www.kaggle.com/carlossouza/osic-autoencoder-training</a></p>\n\n<p><a href=\"/tanlikesmath\">@tanlikesmath</a> , IMHO I followed your notebook to the letter. Very few changes: data, model, and loss function. The code is running, but it is very very very very slow, and session monitor shows MXU at constant 0%, with -- Idle Time.</p>\n\n<p><strong>Can you help me understand what I am doing wrong and how to fix?</strong></p>\n\n<p>As soon as I understand what's going on, why it is not working, and how to fix, I'll gladly contribute with a tutorial :)</p>",
          "rawMarkdown": "@jysohn23 , @tanlikesmath , @jessemostipak ,\nThank you all for your answers! You gave me hope, and I tried again :)\nHere's the code (still not working):\n\nhttps://www.kaggle.com/carlossouza/osic-autoencoder-training-on-tpus\n\nThe GPU code works perfectly: https://www.kaggle.com/carlossouza/osic-autoencoder-training\n\n@tanlikesmath , IMHO I followed your notebook to the letter. Very few changes: data, model, and loss function. The code is running, but it is very very very very slow, and session monitor shows MXU at constant 0%, with -- Idle Time.\n\n**Can you help me understand what I am doing wrong and how to fix?**\n\nAs soon as I understand what's going on, why it is not working, and how to fix, I'll gladly contribute with a tutorial :)",
          "votes": 2,
          "replies": [
            {
              "id": 951165,
              "postDate": "2020-07-30T00:33:12.980Z",
              "content": "<p>Have you tried a larger Batch Size with the TPU? You have more memory available in the TPU then you do in the GPU. Try doubling your batch size until you run out of memory, then back off a bit. If you are only using 2D images, you might get a batch size like 64. Less if you are using 3D images.</p>",
              "rawMarkdown": "Have you tried a larger Batch Size with the TPU? You have more memory available in the TPU then you do in the GPU. Try doubling your batch size until you run out of memory, then back off a bit. If you are only using 2D images, you might get a batch size like 64. Less if you are using 3D images."
            }
          ]
        },
        {
          "id": 951189,
          "postDate": "2020-07-30T01:19:59.363Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Thanks for sharing a public kernel of your issue. I will look into it tomorrow. In the meantime, I also highly recommend you open an issue in the PyTorch XLA GitHub repository (<a href=\"https://github.com/pytorch/xla/issues\">here</a>). The PyTorch XLA team is quite responsive and very helpful!</p>",
          "rawMarkdown": "@carlossouza Thanks for sharing a public kernel of your issue. I will look into it tomorrow. In the meantime, I also highly recommend you open an issue in the PyTorch XLA GitHub repository ([here](https://github.com/pytorch/xla/issues)). The PyTorch XLA team is quite responsive and very helpful!"
        },
        {
          "id": 951840,
          "postDate": "2020-07-30T12:54:31.963Z",
          "content": "<p>Just asked.. thanks!</p>",
          "rawMarkdown": "Just asked.. thanks!"
        },
        {
          "id": 954361,
          "postDate": "2020-08-01T17:03:30.370Z",
          "content": "<p>Just to give an update on this thread we're continuing our investigation on the issue in <a href=\"https://github.com/pytorch/xla/issues/2383\">https://github.com/pytorch/xla/issues/2383</a></p>\n\n<p>This model uses 3d convolutions, which our PyTorch/XLA stack has never tested out yet (so far we had been focussed on image classification and transformer/bert style models). So thanks Carlos for reporting this bug and we'll continue to expand the set of models/ops that we cover.</p>",
          "rawMarkdown": "Just to give an update on this thread we're continuing our investigation on the issue in https://github.com/pytorch/xla/issues/2383\n\nThis model uses 3d convolutions, which our PyTorch/XLA stack has never tested out yet (so far we had been focussed on image classification and transformer/bert style models). So thanks Carlos for reporting this bug and we'll continue to expand the set of models/ops that we cover.",
          "votes": 1
        },
        {
          "id": 954371,
          "postDate": "2020-08-01T17:26:02.540Z",
          "content": "<p><a href=\"/jysohn23\">@jysohn23</a> Just wanted to thank the contributors for making PyTorch more robust and bug-free for TPUs. I currently use PyTorch with GPUs, but with the rapid development pace of PyTorch-XLA, transitioning to TPUs would be much easier in the future, thanks to efforts put by the community.</p>\n\n<p>I would love to contribute, but I lack experience working with large open-source projects, hopefully I'll make a contribution in the future :)</p>\n\n<p>Huge thanks to the contributors!</p>",
          "rawMarkdown": "@jysohn23 Just wanted to thank the contributors for making PyTorch more robust and bug-free for TPUs. I currently use PyTorch with GPUs, but with the rapid development pace of PyTorch-XLA, transitioning to TPUs would be much easier in the future, thanks to efforts put by the community.\n\nI would love to contribute, but I lack experience working with large open-source projects, hopefully I'll make a contribution in the future :)\n\nHuge thanks to the contributors!"
        },
        {
          "id": 969215,
          "postDate": "2020-08-13T14:51:35.957Z",
          "content": "<p>Hi Jesse, thank you for the links, are really useful. Could you please share few similar one for migrating from using Tensorflow with GPU to using it with TPU?</p>",
          "rawMarkdown": "Hi Jesse, thank you for the links, are really useful. Could you please share few similar one for migrating from using Tensorflow with GPU to using it with TPU?"
        }
      ]
    },
    {
      "id": 948626,
      "postDate": "2020-07-28T05:35:50.993Z",
      "content": "<p>It's literally impossible to debug. I was patient enough to put print statements between every single line in the training code. Now, I know that <code>xm.optimizer_step(optimizer, barrier=True)</code> takes a lot of time, almost freezes. So, what to do now? There's no documentation, nobody to ask for help. Searching gives me no answers... really annoying... It would be great to exchange TPU credits by GPU credits: at least GPUs are usable...</p>",
      "rawMarkdown": "It's literally impossible to debug. I was patient enough to put print statements between every single line in the training code. Now, I know that `xm.optimizer_step(optimizer, barrier=True)` takes a lot of time, almost freezes. So, what to do now? There's no documentation, nobody to ask for help. Searching gives me no answers... really annoying... It would be great to exchange TPU credits by GPU credits: at least GPUs are usable...",
      "votes": 3,
      "replies": [
        {
          "id": 948642,
          "postDate": "2020-07-28T05:55:50.403Z",
          "content": "<p>What I'd like to say is, PyTorch documentation is better than other libraries. The real cause of frustration should be the lack of examples using <code>PyTorch-XLA</code>.</p>\n\n<p>You should try to download the competition data through Kaggle API, upload and run your kernel on Colab if you are really eager to experiment with GPU.</p>",
          "rawMarkdown": "What I'd like to say is, PyTorch documentation is better than other libraries. The real cause of frustration should be the lack of examples using `PyTorch-XLA`.\n\nYou should try to download the competition data through Kaggle API, upload and run your kernel on Colab if you are really eager to experiment with GPU."
        },
        {
          "id": 948839,
          "postDate": "2020-07-28T09:03:31.360Z",
          "content": "<p>From my experience, the first epoch often takes much longer than the rest. I've had to wait for over 10 minutes for the first epoch to finish, then the remaining were all under 30s, and faster than a GPU. This was while using TPU on Kaggle. Surprisingly, the same code (with a few changes to run on Colab) didn't get out of the first epoch on Colab. Waited for well over an hour. All of this was with TensorFlow. </p>",
          "rawMarkdown": "From my experience, the first epoch often takes much longer than the rest. I've had to wait for over 10 minutes for the first epoch to finish, then the remaining were all under 30s, and faster than a GPU. This was while using TPU on Kaggle. Surprisingly, the same code (with a few changes to run on Colab) didn't get out of the first epoch on Colab. Waited for well over an hour. All of this was with TensorFlow. \n"
        }
      ]
    },
    {
      "id": 1009279,
      "postDate": "2020-09-13T20:02:22.177Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>, really sorry to read about your troubles with TPUs. <br>\nMy honest experience: Using Tensorflow makes TPU-usage quite easy and it works mostly out of the box. <br>\nOne great (and big!) TPU-example notebook is: <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">triple-stratified-kfold-with-tfrecords</a> by Chris Deotte, and this really is a worthwhile read!</p>\n<p>I highly admire your Pytorch skills, but sadly it seems that TPUs/Pytorch are currently not as highly compatible as Tensorflow/TPUs.</p>",
      "rawMarkdown": "Dear @carlossouza, really sorry to read about your troubles with TPUs. \nMy honest experience: Using Tensorflow makes TPU-usage quite easy and it works mostly out of the box. \nOne great (and big!) TPU-example notebook is: [triple-stratified-kfold-with-tfrecords](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) by Chris Deotte, and this really is a worthwhile read!\n\nI highly admire your Pytorch skills, but sadly it seems that TPUs/Pytorch are currently not as highly compatible as Tensorflow/TPUs.",
      "votes": 1,
      "replies": [
        {
          "id": 1009351,
          "postDate": "2020-09-13T22:14:50.817Z",
          "content": "<p>Yeah, I'm sure it works great with Tensorflow… I'm not that smart to use Tensorflow, so I have to use simpler tools like PyTorch :)</p>",
          "rawMarkdown": "Yeah, I'm sure it works great with Tensorflow... I'm not that smart to use Tensorflow, so I have to use simpler tools like PyTorch :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 948631,
      "postDate": "2020-07-28T05:39:13.583Z",
      "content": "<p>I'm not sure how the frameworks differ, since I primarily use Tensorflow and Keras. In my experience, the transition from GPU to TPU is pretty seamless, at least for TF on Colab. Kaggle notebooks are very similar to Colab, so the experience should be roughly the same.</p>\n\n<p>I'm not sure about PyTorch. XLA is indeed a nightmare - it was bad enough trying to fix only XLA-GPU's showing up on my local machine, so I don't even want to think about XLA TPU's on a remote one.</p>\n\n<p>Honestly TPU's are pretty overrated. In my experience, the performance difference between TPU's an GPU's are minimal and GPU's actually usually perform faster. Perhaps on certain batchsizes, TPU's can outclass GPU's, but I've yet to see an mainstream case where using TPU's have really made a difference.</p>",
      "rawMarkdown": "I'm not sure how the frameworks differ, since I primarily use Tensorflow and Keras. In my experience, the transition from GPU to TPU is pretty seamless, at least for TF on Colab. Kaggle notebooks are very similar to Colab, so the experience should be roughly the same.\n\nI'm not sure about PyTorch. XLA is indeed a nightmare - it was bad enough trying to fix only XLA-GPU's showing up on my local machine, so I don't even want to think about XLA TPU's on a remote one.\n\nHonestly TPU's are pretty overrated. In my experience, the performance difference between TPU's an GPU's are minimal and GPU's actually usually perform faster. Perhaps on certain batchsizes, TPU's can outclass GPU's, but I've yet to see an mainstream case where using TPU's have really made a difference.",
      "votes": 1
    },
    {
      "id": 970728,
      "postDate": "2020-08-14T18:01:44.213Z",
      "content": "<p>If the industry is really serious about using TPUs, they MUST release home hardware TPUs … the only way to help improve ML frameworks for ML is to get lots of users to use them…. cheerios</p>",
      "rawMarkdown": "If the industry is really serious about using TPUs, they MUST release home hardware TPUs ... the only way to help improve ML frameworks for ML is to get lots of users to use them.... cheerios"
    },
    {
      "id": 956137,
      "postDate": "2020-08-03T09:16:38.167Z",
      "content": "<p>Already turned to TF2, XLA is really unstable, TF2 on TPU is pretty stable and easy to use.</p>",
      "rawMarkdown": "Already turned to TF2, XLA is really unstable, TF2 on TPU is pretty stable and easy to use."
    },
    {
      "id": 951005,
      "postDate": "2020-07-29T19:50:11.663Z",
      "content": "<p>I personally believe that Tensorflow is more TPU-friendly than PyTorch XLA, although it has some troubles as well.</p>",
      "rawMarkdown": "I personally believe that Tensorflow is more TPU-friendly than PyTorch XLA, although it has some troubles as well."
    },
    {
      "id": 949719,
      "postDate": "2020-07-28T20:58:30.143Z",
      "content": "<p>The documentation for TPU is indeed very poor, as is generally the interface to use them IMHO. You also need to optimize so many things it  gets compilcated. Where your data is located, how it is formatted on disk, what size of files is it in, how your patch size divides by 8, how to fit all in memory fast, ... And how to configure it all and make sure you had it right.</p>\n\n<p>I think TPU are great advance for hard-core ML focused work, but needs much more user-friendliness still. Nice of Kaggle to provide more streamlined interfaces for it though, hopefully over time it will get close to ease of the GPU versions.. It's mostly a software problem after all.</p>",
      "rawMarkdown": "The documentation for TPU is indeed very poor, as is generally the interface to use them IMHO. You also need to optimize so many things it  gets compilcated. Where your data is located, how it is formatted on disk, what size of files is it in, how your patch size divides by 8, how to fit all in memory fast, ... And how to configure it all and make sure you had it right.\n\nI think TPU are great advance for hard-core ML focused work, but needs much more user-friendliness still. Nice of Kaggle to provide more streamlined interfaces for it though, hopefully over time it will get close to ease of the GPU versions.. It's mostly a software problem after all.",
      "replies": [
        {
          "id": 949733,
          "postDate": "2020-07-28T21:28:08.313Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 970729,
      "postDate": "2020-08-14T18:01:44.543Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 949829,
      "author_name": "Jesse Mostipak",
      "author_url": "",
      "post_date": "2020-07-29T01:37:13.610000",
      "content": "<p>Hi <a href=\"/carlossouza\">@carlossouza</a> --- I'm really excited that you're trying out TPUs, and I'm really sorry that you're encountering some of the challenges with PyTorch + TPUs. I know that that can be incredibly frustrating.</p>\n\n<p>There are some additional challenges to getting PyTorch up and running on TPUs, but we've pulled together some documentation <strong><a href=\"https://www.kaggle.com/docs/tpu#tpu8\">here</a></strong>. Our community also has some members who have created some excellent resources on PyTorch + TPUs---I'm not sure what you've had a chance to look at, but these are some of my favorites:\n- <strong><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159723\">PyTorch TPU Improvements</a></strong> from Psi\n- <strong><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271\">Using PyTorch with TPU</a></strong> from CPMP\n- <strong><a href=\"https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel\">Super duper fast PyTorch TPU kernel</a></strong> from Abhishek\n- <strong><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005\">PyTorch XLA/TPU training</a></strong> from ilovescience</p>\n\n<p>I'd love to know if any of these are helpful to you!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 949850,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-07-29T02:37:48.280000",
          "content": "<p>Hi <a href=\"/jessemostipak\">@jessemostipak</a> ! Thanks for your prompt response!</p>\n\n<p>Yes, I've read all these tutorials several times, each. And more tutorials.</p>\n\n<p>Then, I adapted my code (which runs perfectly on GPUs), following exactly these tutorials. I specifically tried 3 from the list you provided: I followed their procedures to the letter, changing only the model, and the datasets. It failed miserably, without explicit errors: the code would take forever to train, indicating that something was obviously not right, but it didn't explicitly break.</p>\n\n<p>IMHO, free TPU credits without proper support is generating the exact opposite effect intended. Instead of generating happy users, it is generating frustrated users who'd never want to try it again. That's because of 3 reasons:</p>\n\n<p>First, it is <strong>impossible to debug</strong>. The code fails without generating any hint on what might be wrong. And filling the code with print statements trying to guess where the problem is is very very very frustrating.</p>\n\n<p>Second, this documentation that you provided is <strong>only useful if you are solving these specific problems</strong> (i.e. using the same models/datasets/etc). They do not help whatsoever in adapting the code to solve problems different than what they showcase.</p>\n\n<p>Third, and maybe the most important: there is nothing in the documentation, in forums, in github, etc, that help us <strong>when something goes wrong</strong>. Nothing. All the (very few) tutorials, examples and documents only show working code. They don't <strong>show potential mistakes/problems and how to solve them</strong>, which is exactly what we need to learn how to use TPUs.</p>\n\n<p>I'd really love to learn how to use TPUs. But after unsuccessfully trying to adapt my code, following exactly the examples provided, I don't know how to continue. I'd love a suggestion actually. Without debugging messages/tools, without documentation that either i) shows how to adapt code to different problems and/or ii) shows main mistakes/errors/problems and how to fix them (troubleshooting), I'm not sure how to continue playing with TPUs. :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 949852,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-07-29T02:43:27.953000",
          "content": "<p>Btw, here's my code running perfectly on GPUs: <a href=\"https://www.kaggle.com/carlossouza/osic-autoencoder-training\">https://www.kaggle.com/carlossouza/osic-autoencoder-training</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949934,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-07-29T04:49:54.213000",
          "content": "<p><a href=\"/jessemostipak\">@jessemostipak</a> Thank you for sharing my kernel with others and I am glad you think it's one of community favorites! However, I might recommend that you instead share <a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\">this updated kernel</a> that explains PyTorch XLA in much more detail.</p>\n\n<p><a href=\"/carlossouza\">@carlossouza</a> Please check the above kernel and see if you get a better understanding of PyTorch XLA. It is quite general and hopefully after looking at that kernel, you could then adapt PyTorch XLA for your own problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949936,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-07-29T04:53:46.150000",
          "content": "<p>Also <a href=\"/carlossouza\">@carlossouza</a> please share what your specific issue is and I bet the community (including myself) would be glad to help you out.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 950745,
          "author_name": "Jesse Mostipak",
          "author_url": "",
          "post_date": "2020-07-29T15:22:29.357000",
          "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a> --- thank you for linking the updated kernel! </p>\n\n<p><a href=\"/carlossouza\">@carlossouza</a> --- you've hit on a lot of the pain points with PyTorch and TPUs. one of the cool things about Kaggle is that we can bring really new and exciting technologies (like TPUs!) to our community, but <em>because</em> they're new, documentation is still being developed. We're <em>all</em> learning together and building out the documentation as we go (which is exciting! but can also be frustrating.)</p>\n\n<p>TPUs were designed to work with TensorFlow, so you may encounter fewer issues and more documentation by switching, but that being said, our community is doing a phenomenal job at learning together and creating the documentation to make PyTorch on TPUs as seamless as possible.</p>\n\n<p>I love the idea of documentation that shows different approaches and//or how to deal with different types of errors, and it would be great if you helped create it! Documenting what you're trying, what you're learning, and what works is a great way to help those who in the future find themselves in similar situations.</p>\n\n<p>Our community is working hard on figuring out PyTorch and TPUs, and I strongly recommend leaning on them for insights on what could be going wrong. The forums are a great way to get an open exchange of ideas and collectively learn more.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 950803,
          "author_name": "Jin Young (Daniel) Sohn",
          "author_url": "",
          "post_date": "2020-07-29T16:21:01.627000",
          "content": "<p>Hi Carlos, we do have this TROUBLESHOOTING (<a href=\"https://github.com/pytorch/xla/blob/master/TROUBLESHOOTING.md\">https://github.com/pytorch/xla/blob/master/TROUBLESHOOTING.md</a>) so please take a look at it. The metrics report is a very useful tool to understand what part of the execution is taking so long (compilation of graph? device step time? data transfer? data retrieval? round trips to CPU?). But thanks for the suggestions, we'll work on improving our documentation and debugging UX. Our team has been stretched and we haven't had enough time to improve that part.</p>\n\n<p>In the meantime if you face problems that the troubleshooting guide isn't able to help with much, open a Github issue on our pytorch/xla repo and we'll try to help get it resolved.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 951118,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-07-29T22:25:07.780000",
          "content": "<p><a href=\"/jysohn23\">@jysohn23</a> , <a href=\"/tanlikesmath\">@tanlikesmath</a> , <a href=\"/jessemostipak\">@jessemostipak</a> ,\nThank you all for your answers! You gave me hope, and I tried again :)\nHere's the code (still not working):</p>\n\n<p><a href=\"https://www.kaggle.com/carlossouza/osic-autoencoder-training-on-tpus\">https://www.kaggle.com/carlossouza/osic-autoencoder-training-on-tpus</a></p>\n\n<p>The GPU code works perfectly: <a href=\"https://www.kaggle.com/carlossouza/osic-autoencoder-training\">https://www.kaggle.com/carlossouza/osic-autoencoder-training</a></p>\n\n<p><a href=\"/tanlikesmath\">@tanlikesmath</a> , IMHO I followed your notebook to the letter. Very few changes: data, model, and loss function. The code is running, but it is very very very very slow, and session monitor shows MXU at constant 0%, with -- Idle Time.</p>\n\n<p><strong>Can you help me understand what I am doing wrong and how to fix?</strong></p>\n\n<p>As soon as I understand what's going on, why it is not working, and how to fix, I'll gladly contribute with a tutorial :)</p>",
          "votes": 2,
          "replies": [
            {
              "id": 951165,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-07-30T00:33:12.980000",
              "content": "<p>Have you tried a larger Batch Size with the TPU? You have more memory available in the TPU then you do in the GPU. Try doubling your batch size until you run out of memory, then back off a bit. If you are only using 2D images, you might get a batch size like 64. Less if you are using 3D images.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 951189,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-07-30T01:19:59.363000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Thanks for sharing a public kernel of your issue. I will look into it tomorrow. In the meantime, I also highly recommend you open an issue in the PyTorch XLA GitHub repository (<a href=\"https://github.com/pytorch/xla/issues\">here</a>). The PyTorch XLA team is quite responsive and very helpful!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 951840,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-07-30T12:54:31.963000",
          "content": "<p>Just asked.. thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 954361,
          "author_name": "Jin Young (Daniel) Sohn",
          "author_url": "",
          "post_date": "2020-08-01T17:03:30.370000",
          "content": "<p>Just to give an update on this thread we're continuing our investigation on the issue in <a href=\"https://github.com/pytorch/xla/issues/2383\">https://github.com/pytorch/xla/issues/2383</a></p>\n\n<p>This model uses 3d convolutions, which our PyTorch/XLA stack has never tested out yet (so far we had been focussed on image classification and transformer/bert style models). So thanks Carlos for reporting this bug and we'll continue to expand the set of models/ops that we cover.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954371,
          "author_name": "Aadhav Vignesh",
          "author_url": "",
          "post_date": "2020-08-01T17:26:02.540000",
          "content": "<p><a href=\"/jysohn23\">@jysohn23</a> Just wanted to thank the contributors for making PyTorch more robust and bug-free for TPUs. I currently use PyTorch with GPUs, but with the rapid development pace of PyTorch-XLA, transitioning to TPUs would be much easier in the future, thanks to efforts put by the community.</p>\n\n<p>I would love to contribute, but I lack experience working with large open-source projects, hopefully I'll make a contribution in the future :)</p>\n\n<p>Huge thanks to the contributors!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 969215,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2020-08-13T14:51:35.957000",
          "content": "<p>Hi Jesse, thank you for the links, are really useful. Could you please share few similar one for migrating from using Tensorflow with GPU to using it with TPU?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 948626,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2020-07-28T05:35:50.993000",
      "content": "<p>It's literally impossible to debug. I was patient enough to put print statements between every single line in the training code. Now, I know that <code>xm.optimizer_step(optimizer, barrier=True)</code> takes a lot of time, almost freezes. So, what to do now? There's no documentation, nobody to ask for help. Searching gives me no answers... really annoying... It would be great to exchange TPU credits by GPU credits: at least GPUs are usable...</p>",
      "votes": 3,
      "replies": [
        {
          "id": 948642,
          "author_name": "Aadhav Vignesh",
          "author_url": "",
          "post_date": "2020-07-28T05:55:50.403000",
          "content": "<p>What I'd like to say is, PyTorch documentation is better than other libraries. The real cause of frustration should be the lack of examples using <code>PyTorch-XLA</code>.</p>\n\n<p>You should try to download the competition data through Kaggle API, upload and run your kernel on Colab if you are really eager to experiment with GPU.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 948839,
          "author_name": "NighthawX",
          "author_url": "",
          "post_date": "2020-07-28T09:03:31.360000",
          "content": "<p>From my experience, the first epoch often takes much longer than the rest. I've had to wait for over 10 minutes for the first epoch to finish, then the remaining were all under 30s, and faster than a GPU. This was while using TPU on Kaggle. Surprisingly, the same code (with a few changes to run on Colab) didn't get out of the first epoch on Colab. Waited for well over an hour. All of this was with TensorFlow. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1009279,
      "author_name": "from coffee import *",
      "author_url": "",
      "post_date": "2020-09-13T20:02:22.177000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>, really sorry to read about your troubles with TPUs. <br>\nMy honest experience: Using Tensorflow makes TPU-usage quite easy and it works mostly out of the box. <br>\nOne great (and big!) TPU-example notebook is: <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">triple-stratified-kfold-with-tfrecords</a> by Chris Deotte, and this really is a worthwhile read!</p>\n<p>I highly admire your Pytorch skills, but sadly it seems that TPUs/Pytorch are currently not as highly compatible as Tensorflow/TPUs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1009351,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-09-13T22:14:50.817000",
          "content": "<p>Yeah, I'm sure it works great with Tensorflow… I'm not that smart to use Tensorflow, so I have to use simpler tools like PyTorch :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 948631,
      "author_name": "Henry Zhang",
      "author_url": "",
      "post_date": "2020-07-28T05:39:13.583000",
      "content": "<p>I'm not sure how the frameworks differ, since I primarily use Tensorflow and Keras. In my experience, the transition from GPU to TPU is pretty seamless, at least for TF on Colab. Kaggle notebooks are very similar to Colab, so the experience should be roughly the same.</p>\n\n<p>I'm not sure about PyTorch. XLA is indeed a nightmare - it was bad enough trying to fix only XLA-GPU's showing up on my local machine, so I don't even want to think about XLA TPU's on a remote one.</p>\n\n<p>Honestly TPU's are pretty overrated. In my experience, the performance difference between TPU's an GPU's are minimal and GPU's actually usually perform faster. Perhaps on certain batchsizes, TPU's can outclass GPU's, but I've yet to see an mainstream case where using TPU's have really made a difference.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 970728,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2020-08-14T18:01:44.213000",
      "content": "<p>If the industry is really serious about using TPUs, they MUST release home hardware TPUs … the only way to help improve ML frameworks for ML is to get lots of users to use them…. cheerios</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 956137,
      "author_name": "Anthony Chan",
      "author_url": "",
      "post_date": "2020-08-03T09:16:38.167000",
      "content": "<p>Already turned to TF2, XLA is really unstable, TF2 on TPU is pretty stable and easy to use.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 951005,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2020-07-29T19:50:11.663000",
      "content": "<p>I personally believe that Tensorflow is more TPU-friendly than PyTorch XLA, although it has some troubles as well.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949719,
      "author_name": "averagemn",
      "author_url": "",
      "post_date": "2020-07-28T20:58:30.143000",
      "content": "<p>The documentation for TPU is indeed very poor, as is generally the interface to use them IMHO. You also need to optimize so many things it  gets compilcated. Where your data is located, how it is formatted on disk, what size of files is it in, how your patch size divides by 8, how to fit all in memory fast, ... And how to configure it all and make sure you had it right.</p>\n\n<p>I think TPU are great advance for hard-core ML focused work, but needs much more user-friendliness still. Nice of Kaggle to provide more streamlined interfaces for it though, hopefully over time it will get close to ease of the GPU versions.. It's mostly a software problem after all.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 949733,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-28T21:28:08.313000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 970729,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-14T18:01:44.543000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "948611": "XLA documentation is very poor. Notebook examples are either so simple they are not useful at all, or so advanced/complex it is impossible to understand what is happening. \n\nThe code frequently freezes, and it is impossible to know what is happening in the background...\n\nIs this only with PyTorch, or the user experience with Tensorflow is as bad as this as well?",
    "949829": "Hi @carlossouza --- I'm really excited that you're trying out TPUs, and I'm really sorry that you're encountering some of the challenges with PyTorch + TPUs. I know that that can be incredibly frustrating.\n\nThere are some additional challenges to getting PyTorch up and running on TPUs, but we've pulled together some documentation **[here](https://www.kaggle.com/docs/tpu#tpu8)**. Our community also has some members who have created some excellent resources on PyTorch + TPUs---I'm not sure what you've had a chance to look at, but these are some of my favorites:\n- **[PyTorch TPU Improvements](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159723)** from Psi\n- **[Using PyTorch with TPU](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138271)** from CPMP\n- **[Super duper fast PyTorch TPU kernel](https://www.kaggle.com/abhishek/super-duper-fast-pytorch-tpu-kernel)** from Abhishek\n- **[PyTorch XLA/TPU training](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005)** from ilovescience\n\nI'd love to know if any of these are helpful to you!",
    "948626": "It's literally impossible to debug. I was patient enough to put print statements between every single line in the training code. Now, I know that `xm.optimizer_step(optimizer, barrier=True)` takes a lot of time, almost freezes. So, what to do now? There's no documentation, nobody to ask for help. Searching gives me no answers... really annoying... It would be great to exchange TPU credits by GPU credits: at least GPUs are usable...",
    "1009279": "Dear @carlossouza, really sorry to read about your troubles with TPUs. \nMy honest experience: Using Tensorflow makes TPU-usage quite easy and it works mostly out of the box. \nOne great (and big!) TPU-example notebook is: [triple-stratified-kfold-with-tfrecords](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) by Chris Deotte, and this really is a worthwhile read!\n\nI highly admire your Pytorch skills, but sadly it seems that TPUs/Pytorch are currently not as highly compatible as Tensorflow/TPUs.",
    "948631": "I'm not sure how the frameworks differ, since I primarily use Tensorflow and Keras. In my experience, the transition from GPU to TPU is pretty seamless, at least for TF on Colab. Kaggle notebooks are very similar to Colab, so the experience should be roughly the same.\n\nI'm not sure about PyTorch. XLA is indeed a nightmare - it was bad enough trying to fix only XLA-GPU's showing up on my local machine, so I don't even want to think about XLA TPU's on a remote one.\n\nHonestly TPU's are pretty overrated. In my experience, the performance difference between TPU's an GPU's are minimal and GPU's actually usually perform faster. Perhaps on certain batchsizes, TPU's can outclass GPU's, but I've yet to see an mainstream case where using TPU's have really made a difference.",
    "970728": "If the industry is really serious about using TPUs, they MUST release home hardware TPUs ... the only way to help improve ML frameworks for ML is to get lots of users to use them.... cheerios",
    "956137": "Already turned to TF2, XLA is really unstable, TF2 on TPU is pretty stable and easy to use.",
    "951005": "I personally believe that Tensorflow is more TPU-friendly than PyTorch XLA, although it has some troubles as well.",
    "949719": "The documentation for TPU is indeed very poor, as is generally the interface to use them IMHO. You also need to optimize so many things it  gets compilcated. Where your data is located, how it is formatted on disk, what size of files is it in, how your patch size divides by 8, how to fit all in memory fast, ... And how to configure it all and make sure you had it right.\n\nI think TPU are great advance for hard-core ML focused work, but needs much more user-friendliness still. Nice of Kaggle to provide more streamlined interfaces for it though, hopefully over time it will get close to ease of the GPU versions.. It's mostly a software problem after all.",
    "970729": ""
  }
}