{
  "id": 581134,
  "title": "How to use TPU in pytorch rightly?",
  "url": "/competitions/waveform-inversion/discussion/581134",
  "author_name": "",
  "post_date": "2025-05-28T14:08:14.906718Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello everyone, I'm new come here and I want to know how to use TPU in pytorch, I have looked the torch_xla but have some bad review on it. Can someone give me a advice? Because I found this competition really need much compute resource.</p>",
  "messages": [
    {
      "id": "3211511",
      "postDate": "05/28/2025 14:08:14",
      "content": "<p>Hello everyone, I'm new come here and I want to know how to use TPU in pytorch, I have looked the torch_xla but have some bad review on it. Can someone give me a advice? Because I found this competition really need much compute resource.</p>",
      "rawMarkdown": "Hello everyone, I'm new come here and I want to know how to use TPU in pytorch, I have looked the torch_xla but have some bad review on it. Can someone give me a advice? Because I found this competition really need much compute resource.",
      "votes": null
    },
    {
      "id": "3211519",
      "postDate": "05/28/2025 14:16:15",
      "content": "<p>You can try pytorch lightning <a href=\"https://www.kaggle.com/kurisew\" target=\"_blank\">@kurisew</a> </p>",
      "rawMarkdown": "You can try pytorch lightning @kurisew",
      "votes": null
    },
    {
      "id": "3211522",
      "postDate": "05/28/2025 14:20:24",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "3211524",
      "postDate": "05/28/2025 14:21:14",
      "content": "<p>My advice- forget about pytorch on TPU. You want TPU, learn tensorflow/jax.</p>",
      "rawMarkdown": "My advice- forget about pytorch on TPU. You want TPU, learn tensorflow/jax.",
      "votes": null
    },
    {
      "id": "3211528",
      "postDate": "05/28/2025 14:24:50",
      "content": "<p>I know it will face much bug using TPU in pytorch😭😭</p>",
      "rawMarkdown": "I know it will face much bug using TPU in pytorch😭😭",
      "votes": null
    },
    {
      "id": "3211532",
      "postDate": "05/28/2025 14:26:11",
      "content": "<p>Pytorch-lightning can access TPU, why would one not use this <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a>? <br>\nI am comfortable using TensorFlow as well, but can also use torch on TPU using lightning.</p>",
      "rawMarkdown": "Pytorch-lightning can access TPU, why would one not use this @shlomoron? \nI am comfortable using TensorFlow as well, but can also use torch on TPU using lightning.",
      "votes": null
    },
    {
      "id": "3211539",
      "postDate": "05/28/2025 14:32:44",
      "content": "<p>Working with TPU is never easy. You want to add an extra layer of conplications, sure. Go for it. I don't recommend.  </p>",
      "rawMarkdown": "Working with TPU is never easy. You want to add an extra layer of conplications, sure. Go for it. I don't recommend.",
      "votes": null
    },
    {
      "id": "3211584",
      "postDate": "05/28/2025 15:15:36",
      "content": "<p>but only use one tpu😀</p>",
      "rawMarkdown": "but only use one tpu😀",
      "votes": null
    },
    {
      "id": "3212295",
      "postDate": "05/29/2025 15:43:27",
      "content": "<p>I also tried training with TPU several times!<br>\nHowever, due to the limited quota and unexpectedly poorer convergence compared to GPUs (though I suspect this was my fault - I couldn't identify the exact cause), I no longer use it now. That said, my following code may be incomplete (missing model definitions and other elements), and some parts like the Japanese comments might be somewhat difficult to understand, but it should still be helpful.</p>\n<p><a href=\"https://www.kaggle.com/code/haruiig/unetfloat16-tpu\" target=\"_blank\">https://www.kaggle.com/code/haruiig/unetfloat16-tpu</a></p>",
      "rawMarkdown": "I also tried training with TPU several times!\nHowever, due to the limited quota and unexpectedly poorer convergence compared to GPUs (though I suspect this was my fault - I couldn't identify the exact cause), I no longer use it now. That said, my following code may be incomplete (missing model definitions and other elements), and some parts like the Japanese comments might be somewhat difficult to understand, but it should still be helpful.\n\n\nhttps://www.kaggle.com/code/haruiig/unetfloat16-tpu",
      "votes": null
    },
    {
      "id": "3213350",
      "postDate": "05/29/2025 21:22:38",
      "content": "<p>Like others have said, if you want to use TPU, using Jax and Flax are going to be your friend way more than pytorch. In fact, the Flax team changed their API (recently?) to mimic pytorch. Good time to become semi-dangerous in both. Plus, I believe JAX automatically confirms device, so no using .to(device). Pretty cool.</p>\n<p><a href=\"url\" target=\"_blank\">https://flax.readthedocs.io/en/latest/nnx_basics.html</a></p>",
      "rawMarkdown": "Like others have said, if you want to use TPU, using Jax and Flax are going to be your friend way more than pytorch. In fact, the Flax team changed their API (recently?) to mimic pytorch. Good time to become semi-dangerous in both. Plus, I believe JAX automatically confirms device, so no using .to(device). Pretty cool.\n\n[https://flax.readthedocs.io/en/latest/nnx_basics.html](url)",
      "votes": null
    },
    {
      "id": "3213397",
      "postDate": "05/30/2025 00:28:44",
      "content": "<p>I'll try it! Thanks.</p>",
      "rawMarkdown": "I'll try it! Thanks.",
      "votes": null
    },
    {
      "id": "3213939",
      "postDate": "05/30/2025 17:23:31",
      "content": "<p>Also since you're here, DeepMind put out this guide on how to scale models to multiple TPU's, and I believe the general principles they lay out apply to multi-GPU training as well.</p>\n<p><a href=\"url\" target=\"_blank\">https://jax-ml.github.io/scaling-book/</a></p>",
      "rawMarkdown": "Also since you're here, DeepMind put out this guide on how to scale models to multiple TPU's, and I believe the general principles they lay out apply to multi-GPU training as well.\n\n[https://jax-ml.github.io/scaling-book/](url)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3211519,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/28/2025 14:16:15",
      "content": "<p>You can try pytorch lightning <a href=\"https://www.kaggle.com/kurisew\" target=\"_blank\">@kurisew</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3211522,
          "author_name": "kurisew",
          "author_url": "",
          "post_date": "05/28/2025 14:20:24",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3211524,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "05/28/2025 14:21:14",
      "content": "<p>My advice- forget about pytorch on TPU. You want TPU, learn tensorflow/jax.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3211528,
          "author_name": "kurisew",
          "author_url": "",
          "post_date": "05/28/2025 14:24:50",
          "content": "<p>I know it will face much bug using TPU in pytorch😭😭</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3211532,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "05/28/2025 14:26:11",
          "content": "<p>Pytorch-lightning can access TPU, why would one not use this <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a>? <br>\nI am comfortable using TensorFlow as well, but can also use torch on TPU using lightning.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3211539,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "05/28/2025 14:32:44",
              "content": "<p>Working with TPU is never easy. You want to add an extra layer of conplications, sure. Go for it. I don't recommend.  </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3211584,
              "author_name": "aichangeworld",
              "author_url": "",
              "post_date": "05/28/2025 15:15:36",
              "content": "<p>but only use one tpu😀</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3212295,
      "author_name": "haruiig",
      "author_url": "",
      "post_date": "05/29/2025 15:43:27",
      "content": "<p>I also tried training with TPU several times!<br>\nHowever, due to the limited quota and unexpectedly poorer convergence compared to GPUs (though I suspect this was my fault - I couldn't identify the exact cause), I no longer use it now. That said, my following code may be incomplete (missing model definitions and other elements), and some parts like the Japanese comments might be somewhat difficult to understand, but it should still be helpful.</p>\n<p><a href=\"https://www.kaggle.com/code/haruiig/unetfloat16-tpu\" target=\"_blank\">https://www.kaggle.com/code/haruiig/unetfloat16-tpu</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3213350,
      "author_name": "msthil",
      "author_url": "",
      "post_date": "05/29/2025 21:22:38",
      "content": "<p>Like others have said, if you want to use TPU, using Jax and Flax are going to be your friend way more than pytorch. In fact, the Flax team changed their API (recently?) to mimic pytorch. Good time to become semi-dangerous in both. Plus, I believe JAX automatically confirms device, so no using .to(device). Pretty cool.</p>\n<p><a href=\"url\" target=\"_blank\">https://flax.readthedocs.io/en/latest/nnx_basics.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3213397,
          "author_name": "kurisew",
          "author_url": "",
          "post_date": "05/30/2025 00:28:44",
          "content": "<p>I'll try it! Thanks.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3213939,
              "author_name": "msthil",
              "author_url": "",
              "post_date": "05/30/2025 17:23:31",
              "content": "<p>Also since you're here, DeepMind put out this guide on how to scale models to multiple TPU's, and I believe the general principles they lay out apply to multi-GPU training as well.</p>\n<p><a href=\"url\" target=\"_blank\">https://jax-ml.github.io/scaling-book/</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3211511": "Hello everyone, I'm new come here and I want to know how to use TPU in pytorch, I have looked the torch_xla but have some bad review on it. Can someone give me a advice? Because I found this competition really need much compute resource.",
    "3211519": "You can try pytorch lightning @kurisew",
    "3211522": "Thank you!",
    "3211524": "My advice- forget about pytorch on TPU. You want TPU, learn tensorflow/jax.",
    "3211528": "I know it will face much bug using TPU in pytorch😭😭",
    "3211532": "Pytorch-lightning can access TPU, why would one not use this @shlomoron? \nI am comfortable using TensorFlow as well, but can also use torch on TPU using lightning.",
    "3211539": "Working with TPU is never easy. You want to add an extra layer of conplications, sure. Go for it. I don't recommend.",
    "3211584": "but only use one tpu😀",
    "3212295": "I also tried training with TPU several times!\nHowever, due to the limited quota and unexpectedly poorer convergence compared to GPUs (though I suspect this was my fault - I couldn't identify the exact cause), I no longer use it now. That said, my following code may be incomplete (missing model definitions and other elements), and some parts like the Japanese comments might be somewhat difficult to understand, but it should still be helpful.\n\n\nhttps://www.kaggle.com/code/haruiig/unetfloat16-tpu",
    "3213350": "Like others have said, if you want to use TPU, using Jax and Flax are going to be your friend way more than pytorch. In fact, the Flax team changed their API (recently?) to mimic pytorch. Good time to become semi-dangerous in both. Plus, I believe JAX automatically confirms device, so no using .to(device). Pretty cool.\n\n[https://flax.readthedocs.io/en/latest/nnx_basics.html](url)",
    "3213397": "I'll try it! Thanks.",
    "3213939": "Also since you're here, DeepMind put out this guide on how to scale models to multiple TPU's, and I believe the general principles they lay out apply to multi-GPU training as well.\n\n[https://jax-ml.github.io/scaling-book/](url)"
  },
  "source": "meta"
}