{
  "id": 482135,
  "title": "Why are TPUs almost never used on Kaggle?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/482135",
  "author_name": "",
  "post_date": "2024-03-06T13:39:57.755140600Z",
  "votes": 8,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I have read that TPUs can be up to 10x faster than GPUs but I have never seen anyone use it on Kaggle. There are barely any discussions or notebooks on it and the notebooks who do use them are now broken. When I tried to use them, there was a queue so I am guessing they are somewhat popular, but the TPUs increase RAM to almost 330GB, so maybe only people who are doing EDA with large amounts of data use them.<br>\nIf anyone has any information on why TPUs aren't used for these competitions please let me know. <br>\nThank you.</p>",
  "messages": [
    {
      "id": "2684214",
      "postDate": "03/06/2024 13:39:57",
      "content": "<p>I have read that TPUs can be up to 10x faster than GPUs but I have never seen anyone use it on Kaggle. There are barely any discussions or notebooks on it and the notebooks who do use them are now broken. When I tried to use them, there was a queue so I am guessing they are somewhat popular, but the TPUs increase RAM to almost 330GB, so maybe only people who are doing EDA with large amounts of data use them.<br>\nIf anyone has any information on why TPUs aren't used for these competitions please let me know. <br>\nThank you.</p>",
      "rawMarkdown": "I have read that TPUs can be up to 10x faster than GPUs but I have never seen anyone use it on Kaggle. There are barely any discussions or notebooks on it and the notebooks who do use them are now broken. When I tried to use them, there was a queue so I am guessing they are somewhat popular, but the TPUs increase RAM to almost 330GB, so maybe only people who are doing EDA with large amounts of data use them.\nIf anyone has any information on why TPUs aren't used for these competitions please let me know. \nThank you.",
      "votes": null
    },
    {
      "id": "2684219",
      "postDate": "03/06/2024 13:42:33",
      "content": "<p>its actually like trend we always prefer GPU &gt;&gt;&gt;TPU but sometimes it is fact <a href=\"https://www.kaggle.com/kawaiicoderuwu\" target=\"_blank\">@kawaiicoderuwu</a> </p>",
      "rawMarkdown": "its actually like trend we always prefer GPU >>>TPU but sometimes it is fact @kawaiicoderuwu",
      "votes": null
    },
    {
      "id": "2684323",
      "postDate": "03/06/2024 14:53:11",
      "content": "<p>Look at my top notebooks and competitions summaries, you will find some nice examples of TPU :) I basically train only on TPU tbh. Most of the people probably have a personal modern GPU so they just don't bother with TPU pipelines</p>",
      "rawMarkdown": "Look at my top notebooks and competitions summaries, you will find some nice examples of TPU :) I basically train only on TPU tbh. Most of the people probably have a personal modern GPU so they just don't bother with TPU pipelines",
      "votes": null
    },
    {
      "id": "2684333",
      "postDate": "03/06/2024 14:57:44",
      "content": "<p>Hello, <br>\nYou can't submit with TPU <a href=\"www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/code-requirements\" target=\"_blank\">see competition code requirements</a>. However you can train on TPU and save the model in separate notebook <br>\nIf you are looking for examples you can search for public notebooks <a href=\"https://www.kaggle.com/code?searchQuery=TPU\" target=\"_blank\">https://www.kaggle.com/code?searchQuery=TPU</a> there are many</p>",
      "rawMarkdown": "Hello, \nYou can't submit with TPU [see competition code requirements](www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/code-requirements). However you can train on TPU and save the model in separate notebook \nIf you are looking for examples you can search for public notebooks https://www.kaggle.com/code?searchQuery=TPU there are many",
      "votes": null
    },
    {
      "id": "2684447",
      "postDate": "03/06/2024 16:04:28",
      "content": "<blockquote>\n  <p>If anyone has any information on why TPUs aren't used for these competitions please let me know.</p>\n</blockquote>\n<p>Main reason I don't is because the TPU VMs have so many missing packages it becomes a pain to fix them to match the GPU ones. </p>",
      "rawMarkdown": ">If anyone has any information on why TPUs aren't used for these competitions please let me know.\n\nMain reason I don't is because the TPU VMs have so many missing packages it becomes a pain to fix them to match the GPU ones.",
      "votes": null
    },
    {
      "id": "2684539",
      "postDate": "03/06/2024 17:28:42",
      "content": "<p>People usually train with GPUs in local PCs/ cloud platforms and usually don't use TPU pipelines for the risk of code breaks in corresponding inferencing from the loaded models <a href=\"https://www.kaggle.com/kawaiicoderuwu\" target=\"_blank\">@kawaiicoderuwu</a> </p>",
      "rawMarkdown": "People usually train with GPUs in local PCs/ cloud platforms and usually don't use TPU pipelines for the risk of code breaks in corresponding inferencing from the loaded models @kawaiicoderuwu",
      "votes": null
    },
    {
      "id": "2684540",
      "postDate": "03/06/2024 17:29:32",
      "content": "<p><a href=\"https://www.kaggle.com/kawaiicoderuwu\" target=\"_blank\">@kawaiicoderuwu</a> this is also a problem, missing packages and dependencies is a factor for using GPUs.<br>\nT4 x 2 is a popular choice though!</p>",
      "rawMarkdown": "kawaiicoderuwu this is also a problem, missing packages and dependencies is a factor for using GPUs.\nT4 x 2 is a popular choice though!",
      "votes": null
    },
    {
      "id": "2684574",
      "postDate": "03/06/2024 17:56:51",
      "content": "<p><a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> Do you have a list of packages you'd like to see in the TPU VM image?<br>\n<a href=\"https://github.com/Kaggle/docker-python/blob/main/tpu/Dockerfile\" target=\"_blank\">https://github.com/Kaggle/docker-python/blob/main/tpu/Dockerfile</a></p>\n<p>I had stripped it down because only a handful of libraries actually get the TPU speedup (tensorflow, jax, pytorch) but if you need other libraries to support those just let us know :)</p>",
      "rawMarkdown": "julianmukaj Do you have a list of packages you'd like to see in the TPU VM image?\nhttps://github.com/Kaggle/docker-python/blob/main/tpu/Dockerfile\n\nI had stripped it down because only a handful of libraries actually get the TPU speedup (tensorflow, jax, pytorch) but if you need other libraries to support those just let us know :)",
      "votes": null
    },
    {
      "id": "2684646",
      "postDate": "03/06/2024 18:24:00",
      "content": "<p>As noted by many - getting it to work can be painful.  But when it does, its nice and fast.  In addition you get more quota hours of fast processing.  But only useful for training.</p>",
      "rawMarkdown": "As noted by many - getting it to work can be painful.  But when it does, its nice and fast.  In addition you get more quota hours of fast processing.  But only useful for training.",
      "votes": null
    },
    {
      "id": "2685227",
      "postDate": "03/07/2024 05:00:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> I have a list of some packages in GPU environment.</p>\n<ul>\n<li>timm</li>\n<li>albumentations</li>\n<li>pytorch-lightning</li>\n<li>humanize</li>\n<li>einops</li>\n<li>pyarrow</li>\n<li>fastparquet</li>\n<li>seaborn (add to list 2024/03/09)</li>\n<li>skimage (add to list 2024/03/09)</li>\n</ul>\n<p>I need pyarrow and fastparquet to read parquet files in this competition data.<br>\nTo use albumentations, I need to reinstall opencv.</p>\n<pre><code>!apt-  &amp;&amp; apt- install - -opencv\n!pip install -q opencv-\n</code></pre>\n<p>humanize package is optional. I am OK if it is not installed by default.</p>",
      "rawMarkdown": "Hi @herbison I have a list of some packages in GPU environment.\n- timm\n- albumentations\n- pytorch-lightning\n- humanize\n- einops\n- pyarrow\n- fastparquet\n- seaborn (add to list 2024/03/09)\n- skimage (add to list 2024/03/09)\n\nI need pyarrow and fastparquet to read parquet files in this competition data.\nTo use albumentations, I need to reinstall opencv.\n\n```\n!apt-get update && apt-get install -y python3-opencv\n!pip install -q opencv-python\n```\n\nhumanize package is optional. I am OK if it is not installed by default.",
      "votes": null
    },
    {
      "id": "2685652",
      "postDate": "03/07/2024 10:29:24",
      "content": "<p>TPUs are good if you are using tensorflow (or keras).  But for pytorch, which mostly is used here especially with computer vision, it is nearly impossible to get to work (have tried XLA many many times).  <br>\nUsed to be you had to use TFRecords for everything, also takes time and effort to set up datasets, but not always necessary now.<br>\nAnd as others have mentioned, inference will need to be GPU for code competitions.</p>\n<p>Having said all that, it is always a good idea to keep skills in different things - TPU or GPU, tensorflow or pytorch.  And if using kaggle resources when you use quota for one you can still use the other.     </p>",
      "rawMarkdown": "TPUs are good if you are using tensorflow (or keras).  But for pytorch, which mostly is used here especially with computer vision, it is nearly impossible to get to work (have tried XLA many many times).  \nUsed to be you had to use TFRecords for everything, also takes time and effort to set up datasets, but not always necessary now.\nAnd as others have mentioned, inference will need to be GPU for code competitions.\n\nHaving said all that, it is always a good idea to keep skills in different things - TPU or GPU, tensorflow or pytorch.  And if using kaggle resources when you use quota for one you can still use the other.",
      "votes": null
    },
    {
      "id": "2686169",
      "postDate": "03/07/2024 17:10:40",
      "content": "<p>There are two main reasons not to use TPU:</p>\n<ol>\n<li>Pytorch is almost impossible to run on TPU, I've seen people do it, but it's \"very difficult\" to implement (it's a tambourine dance). By the way, it is almost impossible to win a competition using tensorflow, at least a lot of people use Pytorch.</li>\n<li>Tpu tensorflow is even more buggy than usual, for example, you can get an unknown error on TPU when everything works on GPU.</li>\n</ol>",
      "rawMarkdown": "There are two main reasons not to use TPU:\n1. Pytorch is almost impossible to run on TPU, I've seen people do it, but it's \"very difficult\" to implement (it's a tambourine dance). By the way, it is almost impossible to win a competition using tensorflow, at least a lot of people use Pytorch.\n2. Tpu tensorflow is even more buggy than usual, for example, you can get an unknown error on TPU when everything works on GPU.",
      "votes": null
    },
    {
      "id": "2686258",
      "postDate": "03/07/2024 18:06:16",
      "content": "<p>Hey, I'm Will from the PyTorch/XLA team. Usability (particularly on TPU) is a known issue for us. Would you mind sharing more detail about your issues with PyTorch+TPU? It's really valuable to us to hear outside perspectives on this topic</p>",
      "rawMarkdown": "Hey, I'm Will from the PyTorch/XLA team. Usability (particularly on TPU) is a known issue for us. Would you mind sharing more detail about your issues with PyTorch+TPU? It's really valuable to us to hear outside perspectives on this topic",
      "votes": null
    },
    {
      "id": "2686352",
      "postDate": "03/07/2024 19:04:45",
      "content": "<p>Hello! First of all, thank you for your work! <br>\nI've never been able to run a pytorch model training on a TPU (I try to do it about once a year consistently), the last time I tried it was 3 months ago with no success. <br>\nThen I found a notebook showing how to do it, I can't find it now. <br>\nI even asked the community for it. If you are interested this is my notebook <a href=\"https://www.kaggle.com/code/aikhmelnytskyy/cancer-subtype-tpu\" target=\"_blank\">https://www.kaggle.com/code/aikhmelnytskyy/cancer-subtype-tpu</a>,<br>\nthere is an error BrokenProcessPool: A process in the process pool was terminated abruptly while the future was running or pending.</p>",
      "rawMarkdown": "Hello! First of all, thank you for your work! \nI've never been able to run a pytorch model training on a TPU (I try to do it about once a year consistently), the last time I tried it was 3 months ago with no success. \nThen I found a notebook showing how to do it, I can't find it now. \nI even asked the community for it. If you are interested this is my notebook https://www.kaggle.com/code/aikhmelnytskyy/cancer-subtype-tpu,\nthere is an error BrokenProcessPool: A process in the process pool was terminated abruptly while the future was running or pending.",
      "votes": null
    },
    {
      "id": "2686382",
      "postDate": "03/07/2024 19:25:31",
      "content": "<p>Thanks, I'll take a look and see if I can fix it up.</p>\n<p>A couple of notes from skimming through:</p>\n<ul>\n<li>I completely forgot our <code>env-setup.py</code> script even existed and it no longer works. Prefer the preinstalled <code>torch_xla</code> package on Kaggle where possible since it's been tested. I just sent a PR to remove the script.</li>\n<li>On Kaggle+TPU VM, you have to <code>os.environ.pop('TPU_PROCESS_ADDRESSES')</code> and <code>os.environ.pop('CLOUD_TPU_TASK_ID')</code> as in our <a href=\"https://github.com/pytorch/xla/blob/master/contrib/kaggle/distributed-pytorch-xla-basics-with-pjrt.ipynb\" target=\"_blank\">example notebook</a>. I'll follow up with <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> and see if we can remove that hack permanently now</li>\n</ul>",
      "rawMarkdown": "Thanks, I'll take a look and see if I can fix it up.\n\nA couple of notes from skimming through:\n\n- I completely forgot our `env-setup.py` script even existed and it no longer works. Prefer the preinstalled `torch_xla` package on Kaggle where possible since it's been tested. I just sent a PR to remove the script.\n- On Kaggle+TPU VM, you have to `os.environ.pop('TPU_PROCESS_ADDRESSES')` and `os.environ.pop('CLOUD_TPU_TASK_ID')` as in our [example notebook](https://github.com/pytorch/xla/blob/master/contrib/kaggle/distributed-pytorch-xla-basics-with-pjrt.ipynb). I'll follow up with @herbison and see if we can remove that hack permanently now",
      "votes": null
    },
    {
      "id": "2686588",
      "postDate": "03/07/2024 23:20:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/wcromar\" target=\"_blank\">@wcromar</a> I found that I got BrokenProcessPool error when I define a model outside pytorch lightning module. Please also check ver. 22 and ver. 31 of <a href=\"https://www.kaggle.com/code/tomooinubushi/train-pytorch-lightning-gpu-tpu-w-b-kfolds?scriptVersionId=165818492\" target=\"_blank\">this notebook</a>.<br>\nI am not sure it is the issue of torch or it is fixable, but just let you know.</p>",
      "rawMarkdown": "Hi @wcromar I found that I got BrokenProcessPool error when I define a model outside pytorch lightning module. Please also check ver. 22 and ver. 31 of [this notebook](https://www.kaggle.com/code/tomooinubushi/train-pytorch-lightning-gpu-tpu-w-b-kfolds?scriptVersionId=165818492).\nI am not sure it is the issue of torch or it is fixable, but just let you know.",
      "votes": null
    },
    {
      "id": "2687298",
      "postDate": "03/08/2024 12:38:05",
      "content": "<p>I tried building some TPU pipelines, but I can't even make them work , they are all buggy for me 😭</p>",
      "rawMarkdown": "I tried building some TPU pipelines, but I can't even make them work , they are all buggy for me 😭",
      "votes": null
    },
    {
      "id": "2687311",
      "postDate": "03/08/2024 12:52:16",
      "content": "<p>Well, working with TPU can be delicate. Fork a working pipeline and carefully change it to fit your needs. Work with tensorflow and tf.data pipeline and avoid the common pitfalls of small batch size (will give you NaNs. Usually 32 or larger should be good) and varying/undefined batch sizes (ds = ds.padded_batch(…, drop_remainder=True) will set you for good generally). Other than that, avoid using np. functions and X.shape in the data pipeline (you have to use only tf./tf.math alternatives and tf.shape(X) instead of X.shape) and debug your code on GPU with the flag jit_compile, i.e. model.compile(…,jit_compile = True). This will give you more info on bugs. It's not easy at the beginning, but it becomes much easier after some experience.</p>",
      "rawMarkdown": "Well, working with TPU can be delicate. Fork a working pipeline and carefully change it to fit your needs. Work with tensorflow and tf.data pipeline and avoid the common pitfalls of small batch size (will give you NaNs. Usually 32 or larger should be good) and varying/undefined batch sizes (ds = ds.padded_batch(..., drop_remainder=True) will set you for good generally). Other than that, avoid using np. functions and X.shape in the data pipeline (you have to use only tf./tf.math alternatives and tf.shape(X) instead of X.shape) and debug your code on GPU with the flag jit_compile, i.e. model.compile(...,jit_compile = True). This will give you more info on bugs. It's not easy at the beginning, but it becomes much easier after some experience.",
      "votes": null
    },
    {
      "id": "2687423",
      "postDate": "03/08/2024 15:22:51",
      "content": "<p>Thank you, I will try my best!</p>",
      "rawMarkdown": "Thank you, I will try my best!",
      "votes": null
    },
    {
      "id": "2687456",
      "postDate": "03/08/2024 15:38:28",
      "content": "<p>I use TPU in some competitions. The main problem is that some computing TPU operations do not work or work too slowly. You cannot train some models, especially for some transformers. The same sometimes the same code, the same model can study differently on GPU and TPU.<br>\nAnd I can assure you that you got the wrong impression. If a competition begins on which it is easier to train the model on the TPU, then you can expect your turn for hours. I personally were once in the fourth ten in the queue for tpu</p>",
      "rawMarkdown": "I use TPU in some competitions. The main problem is that some computing TPU operations do not work or work too slowly. You cannot train some models, especially for some transformers. The same sometimes the same code, the same model can study differently on GPU and TPU.\nAnd I can assure you that you got the wrong impression. If a competition begins on which it is easier to train the model on the TPU, then you can expect your turn for hours. I personally were once in the fourth ten in the queue for tpu",
      "votes": null
    },
    {
      "id": "2687939",
      "postDate": "03/08/2024 21:53:59",
      "content": "<p><a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a> I wasn't able to get the notebook all the way functional, but I was able to get rid of the crash at least. I went through and made some edits/comments: <a href=\"https://www.kaggle.com/code/wcromar/corrections-cancer-subtype-tpu?scriptVersionId=166077629\" target=\"_blank\">https://www.kaggle.com/code/wcromar/corrections-cancer-subtype-tpu?scriptVersionId=166077629</a></p>\n<p>The two big things:</p>\n<ul>\n<li>Try to use the preinstalled <code>torch_xla</code> on Kaggle if you can. Since the environment has to be shared between JAX, TF, and PyTorch, manual installation can get tricky.</li>\n<li>For TPUs, we have to spawn child processes for each chip. Lightning does this in the background with <code>Trainer.fit</code>. Unfortunately, the TPU driver only lets you use one chip from one process, which means the parent process cannot also open the TPU. In your case, this happens when you create a dataset that calls <code>xm.world_size</code> (which causes PyTorch/XLA to open the TPU and find the world size). This is a very common error, so I tried to improve the error messaging in our 2.2 release (Kaggle currently is on 2.1). </li>\n</ul>",
      "rawMarkdown": "aikhmelnytskyy I wasn't able to get the notebook all the way functional, but I was able to get rid of the crash at least. I went through and made some edits/comments: https://www.kaggle.com/code/wcromar/corrections-cancer-subtype-tpu?scriptVersionId=166077629\n\nThe two big things:\n\n- Try to use the preinstalled `torch_xla` on Kaggle if you can. Since the environment has to be shared between JAX, TF, and PyTorch, manual installation can get tricky.\n- For TPUs, we have to spawn child processes for each chip. Lightning does this in the background with `Trainer.fit`. Unfortunately, the TPU driver only lets you use one chip from one process, which means the parent process cannot also open the TPU. In your case, this happens when you create a dataset that calls `xm.world_size` (which causes PyTorch/XLA to open the TPU and find the world size). This is a very common error, so I tried to improve the error messaging in our 2.2 release (Kaggle currently is on 2.1).",
      "votes": null
    },
    {
      "id": "2687957",
      "postDate": "03/08/2024 22:08:46",
      "content": "<p>Thanks Will! I'll try running the model training on the TPU next week, I'll let you know the results.</p>",
      "rawMarkdown": "Thanks Will! I'll try running the model training on the TPU next week, I'll let you know the results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2684219,
      "author_name": "",
      "author_url": "",
      "post_date": "03/06/2024 13:42:33",
      "content": "<p>its actually like trend we always prefer GPU &gt;&gt;&gt;TPU but sometimes it is fact <a href=\"https://www.kaggle.com/kawaiicoderuwu\" target=\"_blank\">@kawaiicoderuwu</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2684323,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "03/06/2024 14:53:11",
      "content": "<p>Look at my top notebooks and competitions summaries, you will find some nice examples of TPU :) I basically train only on TPU tbh. Most of the people probably have a personal modern GPU so they just don't bother with TPU pipelines</p>",
      "votes": null,
      "replies": [
        {
          "id": 2687298,
          "author_name": "arunsensei",
          "author_url": "",
          "post_date": "03/08/2024 12:38:05",
          "content": "<p>I tried building some TPU pipelines, but I can't even make them work , they are all buggy for me 😭</p>",
          "votes": null,
          "replies": [
            {
              "id": 2687311,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "03/08/2024 12:52:16",
              "content": "<p>Well, working with TPU can be delicate. Fork a working pipeline and carefully change it to fit your needs. Work with tensorflow and tf.data pipeline and avoid the common pitfalls of small batch size (will give you NaNs. Usually 32 or larger should be good) and varying/undefined batch sizes (ds = ds.padded_batch(…, drop_remainder=True) will set you for good generally). Other than that, avoid using np. functions and X.shape in the data pipeline (you have to use only tf./tf.math alternatives and tf.shape(X) instead of X.shape) and debug your code on GPU with the flag jit_compile, i.e. model.compile(…,jit_compile = True). This will give you more info on bugs. It's not easy at the beginning, but it becomes much easier after some experience.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2687423,
                  "author_name": "arunsensei",
                  "author_url": "",
                  "post_date": "03/08/2024 15:22:51",
                  "content": "<p>Thank you, I will try my best!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2684333,
      "author_name": "waechter",
      "author_url": "",
      "post_date": "03/06/2024 14:57:44",
      "content": "<p>Hello, <br>\nYou can't submit with TPU <a href=\"www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/code-requirements\" target=\"_blank\">see competition code requirements</a>. However you can train on TPU and save the model in separate notebook <br>\nIf you are looking for examples you can search for public notebooks <a href=\"https://www.kaggle.com/code?searchQuery=TPU\" target=\"_blank\">https://www.kaggle.com/code?searchQuery=TPU</a> there are many</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2684447,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "03/06/2024 16:04:28",
      "content": "<blockquote>\n  <p>If anyone has any information on why TPUs aren't used for these competitions please let me know.</p>\n</blockquote>\n<p>Main reason I don't is because the TPU VMs have so many missing packages it becomes a pain to fix them to match the GPU ones. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2684540,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "03/06/2024 17:29:32",
          "content": "<p><a href=\"https://www.kaggle.com/kawaiicoderuwu\" target=\"_blank\">@kawaiicoderuwu</a> this is also a problem, missing packages and dependencies is a factor for using GPUs.<br>\nT4 x 2 is a popular choice though!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2684574,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "03/06/2024 17:56:51",
          "content": "<p><a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> Do you have a list of packages you'd like to see in the TPU VM image?<br>\n<a href=\"https://github.com/Kaggle/docker-python/blob/main/tpu/Dockerfile\" target=\"_blank\">https://github.com/Kaggle/docker-python/blob/main/tpu/Dockerfile</a></p>\n<p>I had stripped it down because only a handful of libraries actually get the TPU speedup (tensorflow, jax, pytorch) but if you need other libraries to support those just let us know :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2685227,
              "author_name": "tomooinubushi",
              "author_url": "",
              "post_date": "03/07/2024 05:00:02",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> I have a list of some packages in GPU environment.</p>\n<ul>\n<li>timm</li>\n<li>albumentations</li>\n<li>pytorch-lightning</li>\n<li>humanize</li>\n<li>einops</li>\n<li>pyarrow</li>\n<li>fastparquet</li>\n<li>seaborn (add to list 2024/03/09)</li>\n<li>skimage (add to list 2024/03/09)</li>\n</ul>\n<p>I need pyarrow and fastparquet to read parquet files in this competition data.<br>\nTo use albumentations, I need to reinstall opencv.</p>\n<pre><code>!apt-  &amp;&amp; apt- install - -opencv\n!pip install -q opencv-\n</code></pre>\n<p>humanize package is optional. I am OK if it is not installed by default.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2684539,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "03/06/2024 17:28:42",
      "content": "<p>People usually train with GPUs in local PCs/ cloud platforms and usually don't use TPU pipelines for the risk of code breaks in corresponding inferencing from the loaded models <a href=\"https://www.kaggle.com/kawaiicoderuwu\" target=\"_blank\">@kawaiicoderuwu</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2684646,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "03/06/2024 18:24:00",
      "content": "<p>As noted by many - getting it to work can be painful.  But when it does, its nice and fast.  In addition you get more quota hours of fast processing.  But only useful for training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2685652,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "03/07/2024 10:29:24",
      "content": "<p>TPUs are good if you are using tensorflow (or keras).  But for pytorch, which mostly is used here especially with computer vision, it is nearly impossible to get to work (have tried XLA many many times).  <br>\nUsed to be you had to use TFRecords for everything, also takes time and effort to set up datasets, but not always necessary now.<br>\nAnd as others have mentioned, inference will need to be GPU for code competitions.</p>\n<p>Having said all that, it is always a good idea to keep skills in different things - TPU or GPU, tensorflow or pytorch.  And if using kaggle resources when you use quota for one you can still use the other.     </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2686169,
      "author_name": "aikhmelnytskyy",
      "author_url": "",
      "post_date": "03/07/2024 17:10:40",
      "content": "<p>There are two main reasons not to use TPU:</p>\n<ol>\n<li>Pytorch is almost impossible to run on TPU, I've seen people do it, but it's \"very difficult\" to implement (it's a tambourine dance). By the way, it is almost impossible to win a competition using tensorflow, at least a lot of people use Pytorch.</li>\n<li>Tpu tensorflow is even more buggy than usual, for example, you can get an unknown error on TPU when everything works on GPU.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2686258,
          "author_name": "wcromar",
          "author_url": "",
          "post_date": "03/07/2024 18:06:16",
          "content": "<p>Hey, I'm Will from the PyTorch/XLA team. Usability (particularly on TPU) is a known issue for us. Would you mind sharing more detail about your issues with PyTorch+TPU? It's really valuable to us to hear outside perspectives on this topic</p>",
          "votes": null,
          "replies": [
            {
              "id": 2686352,
              "author_name": "aikhmelnytskyy",
              "author_url": "",
              "post_date": "03/07/2024 19:04:45",
              "content": "<p>Hello! First of all, thank you for your work! <br>\nI've never been able to run a pytorch model training on a TPU (I try to do it about once a year consistently), the last time I tried it was 3 months ago with no success. <br>\nThen I found a notebook showing how to do it, I can't find it now. <br>\nI even asked the community for it. If you are interested this is my notebook <a href=\"https://www.kaggle.com/code/aikhmelnytskyy/cancer-subtype-tpu\" target=\"_blank\">https://www.kaggle.com/code/aikhmelnytskyy/cancer-subtype-tpu</a>,<br>\nthere is an error BrokenProcessPool: A process in the process pool was terminated abruptly while the future was running or pending.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2686382,
                  "author_name": "wcromar",
                  "author_url": "",
                  "post_date": "03/07/2024 19:25:31",
                  "content": "<p>Thanks, I'll take a look and see if I can fix it up.</p>\n<p>A couple of notes from skimming through:</p>\n<ul>\n<li>I completely forgot our <code>env-setup.py</code> script even existed and it no longer works. Prefer the preinstalled <code>torch_xla</code> package on Kaggle where possible since it's been tested. I just sent a PR to remove the script.</li>\n<li>On Kaggle+TPU VM, you have to <code>os.environ.pop('TPU_PROCESS_ADDRESSES')</code> and <code>os.environ.pop('CLOUD_TPU_TASK_ID')</code> as in our <a href=\"https://github.com/pytorch/xla/blob/master/contrib/kaggle/distributed-pytorch-xla-basics-with-pjrt.ipynb\" target=\"_blank\">example notebook</a>. I'll follow up with <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> and see if we can remove that hack permanently now</li>\n</ul>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2686588,
                      "author_name": "tomooinubushi",
                      "author_url": "",
                      "post_date": "03/07/2024 23:20:24",
                      "content": "<p>Hi <a href=\"https://www.kaggle.com/wcromar\" target=\"_blank\">@wcromar</a> I found that I got BrokenProcessPool error when I define a model outside pytorch lightning module. Please also check ver. 22 and ver. 31 of <a href=\"https://www.kaggle.com/code/tomooinubushi/train-pytorch-lightning-gpu-tpu-w-b-kfolds?scriptVersionId=165818492\" target=\"_blank\">this notebook</a>.<br>\nI am not sure it is the issue of torch or it is fixable, but just let you know.</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2687939,
                      "author_name": "wcromar",
                      "author_url": "",
                      "post_date": "03/08/2024 21:53:59",
                      "content": "<p><a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a> I wasn't able to get the notebook all the way functional, but I was able to get rid of the crash at least. I went through and made some edits/comments: <a href=\"https://www.kaggle.com/code/wcromar/corrections-cancer-subtype-tpu?scriptVersionId=166077629\" target=\"_blank\">https://www.kaggle.com/code/wcromar/corrections-cancer-subtype-tpu?scriptVersionId=166077629</a></p>\n<p>The two big things:</p>\n<ul>\n<li>Try to use the preinstalled <code>torch_xla</code> on Kaggle if you can. Since the environment has to be shared between JAX, TF, and PyTorch, manual installation can get tricky.</li>\n<li>For TPUs, we have to spawn child processes for each chip. Lightning does this in the background with <code>Trainer.fit</code>. Unfortunately, the TPU driver only lets you use one chip from one process, which means the parent process cannot also open the TPU. In your case, this happens when you create a dataset that calls <code>xm.world_size</code> (which causes PyTorch/XLA to open the TPU and find the world size). This is a very common error, so I tried to improve the error messaging in our 2.2 release (Kaggle currently is on 2.1). </li>\n</ul>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2687957,
                          "author_name": "aikhmelnytskyy",
                          "author_url": "",
                          "post_date": "03/08/2024 22:08:46",
                          "content": "<p>Thanks Will! I'll try running the model training on the TPU next week, I'll let you know the results.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2687456,
      "author_name": "zaakciiru",
      "author_url": "",
      "post_date": "03/08/2024 15:38:28",
      "content": "<p>I use TPU in some competitions. The main problem is that some computing TPU operations do not work or work too slowly. You cannot train some models, especially for some transformers. The same sometimes the same code, the same model can study differently on GPU and TPU.<br>\nAnd I can assure you that you got the wrong impression. If a competition begins on which it is easier to train the model on the TPU, then you can expect your turn for hours. I personally were once in the fourth ten in the queue for tpu</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2684214": "I have read that TPUs can be up to 10x faster than GPUs but I have never seen anyone use it on Kaggle. There are barely any discussions or notebooks on it and the notebooks who do use them are now broken. When I tried to use them, there was a queue so I am guessing they are somewhat popular, but the TPUs increase RAM to almost 330GB, so maybe only people who are doing EDA with large amounts of data use them.\nIf anyone has any information on why TPUs aren't used for these competitions please let me know. \nThank you.",
    "2684219": "its actually like trend we always prefer GPU >>>TPU but sometimes it is fact @kawaiicoderuwu",
    "2684323": "Look at my top notebooks and competitions summaries, you will find some nice examples of TPU :) I basically train only on TPU tbh. Most of the people probably have a personal modern GPU so they just don't bother with TPU pipelines",
    "2684333": "Hello, \nYou can't submit with TPU [see competition code requirements](www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/code-requirements). However you can train on TPU and save the model in separate notebook \nIf you are looking for examples you can search for public notebooks https://www.kaggle.com/code?searchQuery=TPU there are many",
    "2684447": ">If anyone has any information on why TPUs aren't used for these competitions please let me know.\n\nMain reason I don't is because the TPU VMs have so many missing packages it becomes a pain to fix them to match the GPU ones.",
    "2684539": "People usually train with GPUs in local PCs/ cloud platforms and usually don't use TPU pipelines for the risk of code breaks in corresponding inferencing from the loaded models @kawaiicoderuwu",
    "2684540": "kawaiicoderuwu this is also a problem, missing packages and dependencies is a factor for using GPUs.\nT4 x 2 is a popular choice though!",
    "2684574": "julianmukaj Do you have a list of packages you'd like to see in the TPU VM image?\nhttps://github.com/Kaggle/docker-python/blob/main/tpu/Dockerfile\n\nI had stripped it down because only a handful of libraries actually get the TPU speedup (tensorflow, jax, pytorch) but if you need other libraries to support those just let us know :)",
    "2684646": "As noted by many - getting it to work can be painful.  But when it does, its nice and fast.  In addition you get more quota hours of fast processing.  But only useful for training.",
    "2685227": "Hi @herbison I have a list of some packages in GPU environment.\n- timm\n- albumentations\n- pytorch-lightning\n- humanize\n- einops\n- pyarrow\n- fastparquet\n- seaborn (add to list 2024/03/09)\n- skimage (add to list 2024/03/09)\n\nI need pyarrow and fastparquet to read parquet files in this competition data.\nTo use albumentations, I need to reinstall opencv.\n\n```\n!apt-get update && apt-get install -y python3-opencv\n!pip install -q opencv-python\n```\n\nhumanize package is optional. I am OK if it is not installed by default.",
    "2685652": "TPUs are good if you are using tensorflow (or keras).  But for pytorch, which mostly is used here especially with computer vision, it is nearly impossible to get to work (have tried XLA many many times).  \nUsed to be you had to use TFRecords for everything, also takes time and effort to set up datasets, but not always necessary now.\nAnd as others have mentioned, inference will need to be GPU for code competitions.\n\nHaving said all that, it is always a good idea to keep skills in different things - TPU or GPU, tensorflow or pytorch.  And if using kaggle resources when you use quota for one you can still use the other.",
    "2686169": "There are two main reasons not to use TPU:\n1. Pytorch is almost impossible to run on TPU, I've seen people do it, but it's \"very difficult\" to implement (it's a tambourine dance). By the way, it is almost impossible to win a competition using tensorflow, at least a lot of people use Pytorch.\n2. Tpu tensorflow is even more buggy than usual, for example, you can get an unknown error on TPU when everything works on GPU.",
    "2686258": "Hey, I'm Will from the PyTorch/XLA team. Usability (particularly on TPU) is a known issue for us. Would you mind sharing more detail about your issues with PyTorch+TPU? It's really valuable to us to hear outside perspectives on this topic",
    "2686352": "Hello! First of all, thank you for your work! \nI've never been able to run a pytorch model training on a TPU (I try to do it about once a year consistently), the last time I tried it was 3 months ago with no success. \nThen I found a notebook showing how to do it, I can't find it now. \nI even asked the community for it. If you are interested this is my notebook https://www.kaggle.com/code/aikhmelnytskyy/cancer-subtype-tpu,\nthere is an error BrokenProcessPool: A process in the process pool was terminated abruptly while the future was running or pending.",
    "2686382": "Thanks, I'll take a look and see if I can fix it up.\n\nA couple of notes from skimming through:\n\n- I completely forgot our `env-setup.py` script even existed and it no longer works. Prefer the preinstalled `torch_xla` package on Kaggle where possible since it's been tested. I just sent a PR to remove the script.\n- On Kaggle+TPU VM, you have to `os.environ.pop('TPU_PROCESS_ADDRESSES')` and `os.environ.pop('CLOUD_TPU_TASK_ID')` as in our [example notebook](https://github.com/pytorch/xla/blob/master/contrib/kaggle/distributed-pytorch-xla-basics-with-pjrt.ipynb). I'll follow up with @herbison and see if we can remove that hack permanently now",
    "2686588": "Hi @wcromar I found that I got BrokenProcessPool error when I define a model outside pytorch lightning module. Please also check ver. 22 and ver. 31 of [this notebook](https://www.kaggle.com/code/tomooinubushi/train-pytorch-lightning-gpu-tpu-w-b-kfolds?scriptVersionId=165818492).\nI am not sure it is the issue of torch or it is fixable, but just let you know.",
    "2687298": "I tried building some TPU pipelines, but I can't even make them work , they are all buggy for me 😭",
    "2687311": "Well, working with TPU can be delicate. Fork a working pipeline and carefully change it to fit your needs. Work with tensorflow and tf.data pipeline and avoid the common pitfalls of small batch size (will give you NaNs. Usually 32 or larger should be good) and varying/undefined batch sizes (ds = ds.padded_batch(..., drop_remainder=True) will set you for good generally). Other than that, avoid using np. functions and X.shape in the data pipeline (you have to use only tf./tf.math alternatives and tf.shape(X) instead of X.shape) and debug your code on GPU with the flag jit_compile, i.e. model.compile(...,jit_compile = True). This will give you more info on bugs. It's not easy at the beginning, but it becomes much easier after some experience.",
    "2687423": "Thank you, I will try my best!",
    "2687456": "I use TPU in some competitions. The main problem is that some computing TPU operations do not work or work too slowly. You cannot train some models, especially for some transformers. The same sometimes the same code, the same model can study differently on GPU and TPU.\nAnd I can assure you that you got the wrong impression. If a competition begins on which it is easier to train the model on the TPU, then you can expect your turn for hours. I personally were once in the fourth ten in the queue for tpu",
    "2687939": "aikhmelnytskyy I wasn't able to get the notebook all the way functional, but I was able to get rid of the crash at least. I went through and made some edits/comments: https://www.kaggle.com/code/wcromar/corrections-cancer-subtype-tpu?scriptVersionId=166077629\n\nThe two big things:\n\n- Try to use the preinstalled `torch_xla` on Kaggle if you can. Since the environment has to be shared between JAX, TF, and PyTorch, manual installation can get tricky.\n- For TPUs, we have to spawn child processes for each chip. Lightning does this in the background with `Trainer.fit`. Unfortunately, the TPU driver only lets you use one chip from one process, which means the parent process cannot also open the TPU. In your case, this happens when you create a dataset that calls `xm.world_size` (which causes PyTorch/XLA to open the TPU and find the world size). This is a very common error, so I tried to improve the error messaging in our 2.2 release (Kaggle currently is on 2.1).",
    "2687957": "Thanks Will! I'll try running the model training on the TPU next week, I'll let you know the results."
  },
  "source": "meta"
}