{
  "id": 164020,
  "title": "pytorch and TPU",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164020",
  "author_name": "",
  "post_date": "2020-07-04T12:13:50.585500600Z",
  "votes": 4,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Question to people using pytorch - are we forced to use this tfrecords format or for pytorch it is possible to use TPU with pytorch on Kaggle with normal images? Is it still as fast as people say about tensorflow TPU? I have seen some pytorch kernels but not investigated them yet (I am working on my local GPU during research phase).</p>",
  "messages": [
    {
      "id": "915006",
      "postDate": "07/04/2020 12:13:50",
      "content": "<p>Question to people using pytorch - are we forced to use this tfrecords format or for pytorch it is possible to use TPU with pytorch on Kaggle with normal images? Is it still as fast as people say about tensorflow TPU? I have seen some pytorch kernels but not investigated them yet (I am working on my local GPU during research phase).</p>",
      "rawMarkdown": "Question to people using pytorch - are we forced to use this tfrecords format or for pytorch it is possible to use TPU with pytorch on Kaggle with normal images? Is it still as fast as people say about tensorflow TPU? I have seen some pytorch kernels but not investigated them yet (I am working on my local GPU during research phase).",
      "votes": null
    },
    {
      "id": "915051",
      "postDate": "07/04/2020 12:57:26",
      "content": "<p>TFRecords speedup image data loading, but are not mandatory.</p>\n\n<p>PyTorch on TPUs uses standard \"from torchvision import datasets\", \"datasets.CIFAR10\" and \"torch.utils.data.DataLoader\"\nSee <a href=\"https://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/resnet18-training.ipynb\">PyTorch example on TPU</a></p>",
      "rawMarkdown": "TFRecords speedup image data loading, but are not mandatory.\n\nPyTorch on TPUs uses standard \"from torchvision import datasets\", \"datasets.CIFAR10\" and \"torch.utils.data.DataLoader\"\nSee [PyTorch example on TPU](https://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/resnet18-training.ipynb)",
      "votes": null
    },
    {
      "id": "915288",
      "postDate": "07/04/2020 16:04:02",
      "content": "<p>do you use pytorch in your kernel for this competition?</p>",
      "rawMarkdown": "do you use pytorch in your kernel for this competition?",
      "votes": null
    },
    {
      "id": "915348",
      "postDate": "07/04/2020 16:57:38",
      "content": "<p>Jacek, I uploaded PyTorch JPEGs <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>. Instead of downloading 110GB of the original data, you only need to download 1GB, 2GB, 3GB, 5GB depending on whether you want 256x256, 384x384, 512x512, 768x768. </p>",
      "rawMarkdown": "Jacek, I uploaded PyTorch JPEGs [here][1]. Instead of downloading 110GB of the original data, you only need to download 1GB, 2GB, 3GB, 5GB depending on whether you want 256x256, 384x384, 512x512, 768x768. \n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
      "votes": null
    },
    {
      "id": "915467",
      "postDate": "07/04/2020 19:07:37",
      "content": "<p>I don't understand, do you mean JPEGs are usable by TPU with pytorch? </p>\n\n<p>I already downloaded all original datasets from 2020 and 2019 but downloading files is for local training, not for TPU on Kaggle.</p>",
      "rawMarkdown": "I don't understand, do you mean JPEGs are usable by TPU with pytorch? \n\nI already downloaded all original datasets from 2020 and 2019 but downloading files is for local training, not for TPU on Kaggle.",
      "votes": null
    },
    {
      "id": "915533",
      "postDate": "07/04/2020 19:58:19",
      "content": "<p>Yes, JPEGs if you wish to use TPU and PyTorch at Kaggle, CoLab, or GCP.</p>\n\n<p>If you use TPU or GPU and PyTorch, it is most convenient to use folders of JPEGs. If you use TPU and TensorFlow is it most convenient to use TFRecords. If you use GPU and TensorFlow it is equally convenient to use folders of JPEGs or TFRecords.</p>",
      "rawMarkdown": "Yes, JPEGs if you wish to use TPU and PyTorch at Kaggle, CoLab, or GCP.\n\nIf you use TPU or GPU and PyTorch, it is most convenient to use folders of JPEGs. If you use TPU and TensorFlow is it most convenient to use TFRecords. If you use GPU and TensorFlow it is equally convenient to use folders of JPEGs or TFRecords.",
      "votes": null
    },
    {
      "id": "915536",
      "postDate": "07/04/2020 20:01:05",
      "content": "<p>There are example notebooks showing how to do all of the above. TPU are very fast with TFRecords. Previously there were problems maximizing the performance of TPU with PyTorch but i believe that has all been fixed now and PyTorch can achieve similar speed as TensorFlow on TPU</p>",
      "rawMarkdown": "There are example notebooks showing how to do all of the above. TPU are very fast with TFRecords. Previously there were problems maximizing the performance of TPU with PyTorch but i believe that has all been fixed now and PyTorch can achieve similar speed as TensorFlow on TPU",
      "votes": null
    },
    {
      "id": "915552",
      "postDate": "07/04/2020 20:22:32",
      "content": "<p>Hm, but tensoflow is most popular in this competition, wonder why.</p>",
      "rawMarkdown": "Hm, but tensoflow is most popular in this competition, wonder why.",
      "votes": null
    },
    {
      "id": "915925",
      "postDate": "07/05/2020 07:55:10",
      "content": "<p>You can use JPEGs for tpu training on Pytorch. The tf_records has the same files just a different format. Apart from that if you still want to use tf_records with pytorch there is a tf_record reader in torch_xla utilities. </p>\n\n<p>Havent tried that but yes it exists. </p>",
      "rawMarkdown": "You can use JPEGs for tpu training on Pytorch. The tf_records has the same files just a different format. Apart from that if you still want to use tf_records with pytorch there is a tf_record reader in torch_xla utilities. \n\nHavent tried that but yes it exists.",
      "votes": null
    },
    {
      "id": "915992",
      "postDate": "07/05/2020 09:10:08",
      "content": "<p>In addition to the regular torch dataset, PyTorch XLA also implemented a TFRecord reader last year: <a href=\"https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py\">https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py</a></p>",
      "rawMarkdown": "In addition to the regular torch dataset, PyTorch XLA also implemented a TFRecord reader last year: https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py",
      "votes": null
    },
    {
      "id": "916082",
      "postDate": "07/05/2020 10:43:52",
      "content": "<p>In case of TPUs, I believe that running PyTorch in a Cloud VM is still slower than TF on the same (VM) hardware.</p>\n\n<p>This is because TF has a dedicated TF server running on the TPU VM. The TF running on a cloud VM does not use much resources because all the input pipelines (including augmentations) actually run on the TPU VMs resources (which are pretty rich in terms on CPU and RAM).</p>\n\n<p>In case of PyTorch the input pipelines run on the cloud VMs, and there is no dedicated PyTorch server on the TPU VM side yet (future enhancement). So, a cloud VM with say 2 CPU cores will become quickly saturated in some cases where the model requires a lot of host computation resources because of a large dataset, custom augmentation pipelines, and smaller model step time. A solution for this is to use higher number of CPU cores, say 32+ cores, and then you will notice that PyTorch performance is now on par with that of TF on TPUs.</p>",
      "rawMarkdown": "In case of TPUs, I believe that running PyTorch in a Cloud VM is still slower than TF on the same (VM) hardware.\n\nThis is because TF has a dedicated TF server running on the TPU VM. The TF running on a cloud VM does not use much resources because all the input pipelines (including augmentations) actually run on the TPU VMs resources (which are pretty rich in terms on CPU and RAM).\n\nIn case of PyTorch the input pipelines run on the cloud VMs, and there is no dedicated PyTorch server on the TPU VM side yet (future enhancement). So, a cloud VM with say 2 CPU cores will become quickly saturated in some cases where the model requires a lot of host computation resources because of a large dataset, custom augmentation pipelines, and smaller model step time. A solution for this is to use higher number of CPU cores, say 32+ cores, and then you will notice that PyTorch performance is now on par with that of TF on TPUs.",
      "votes": null
    },
    {
      "id": "916155",
      "postDate": "07/05/2020 12:25:19",
      "content": "<p>What do you mean by \"locally\"? I thought you can use TPU only in the cloud.</p>",
      "rawMarkdown": "What do you mean by \"locally\"? I thought you can use TPU only in the cloud.",
      "votes": null
    },
    {
      "id": "916201",
      "postDate": "07/05/2020 13:18:02",
      "content": "<p>Updated above reply to use consistent terminology \"cloud VM\", instead of locally/user VM.</p>",
      "rawMarkdown": "Updated above reply to use consistent terminology \"cloud VM\", instead of locally/user VM.",
      "votes": null
    },
    {
      "id": "916221",
      "postDate": "07/05/2020 13:34:17",
      "content": "<p>I totally agree with <a href=\"/sirishks\">@sirishks</a> </p>\n\n<p>That's why i use Pytorch TPU on Colab.  You have much more RAM and CPU cores there. </p>\n\n<p>So, Pytorch TPU is blazingly fast on Colab. </p>",
      "rawMarkdown": "I totally agree with @sirishks \n\nThat's why i use Pytorch TPU on Colab.  You have much more RAM and CPU cores there. \n\nSo, Pytorch TPU is blazingly fast on Colab.",
      "votes": null
    },
    {
      "id": "916229",
      "postDate": "07/05/2020 13:40:26",
      "content": "<p><a href=\"/serigne\">@serigne</a> do you mean you train models for this competition on Google Colab? can you tell more about current state of Colab? how do you use Kaggle data? do you need paid account?</p>\n\n<p>EDIT just tested this on colab:</p>\n\n<p>import torch_xla_py.xla_model as xm</p>\n\n<p>ModuleNotFoundError: No module named 'torch_xla_py'</p>\n\n<p>do I have to install it same way as on Kaggle kernels?</p>",
      "rawMarkdown": "serigne do you mean you train models for this competition on Google Colab? can you tell more about current state of Colab? how do you use Kaggle data? do you need paid account?\n\nEDIT just tested this on colab:\n\nimport torch_xla_py.xla_model as xm\n\nModuleNotFoundError: No module named 'torch_xla_py'\n\ndo I have to install it same way as on Kaggle kernels?",
      "votes": null
    },
    {
      "id": "916325",
      "postDate": "07/05/2020 14:56:35",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  Yes, Serigne is using paid account, which gives approx twice the resources and time limits (24-hrs) compared to Colab.</p>\n\n<p>Please see <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161713#903803\">earlier discussion</a>.</p>",
      "rawMarkdown": "jacekpoplawski  Yes, Serigne is using paid account, which gives approx twice the resources and time limits (24-hrs) compared to Colab.\n\nPlease see [earlier discussion](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161713#903803).",
      "votes": null
    },
    {
      "id": "916365",
      "postDate": "07/05/2020 15:26:09",
      "content": "<p><a href=\"/serigne\">@serigne</a> so what are requirements to run pytorch kernel for this kaggle competition? how to use the data?</p>",
      "rawMarkdown": "serigne so what are requirements to run pytorch kernel for this kaggle competition? how to use the data?",
      "votes": null
    },
    {
      "id": "917876",
      "postDate": "07/06/2020 19:21:39",
      "content": "<p>I don't use kernel . I use Colab only. </p>\n\n<p>For Pytorch the data is uploaded in my Google Drive. For TF, I use GCS bucket. </p>",
      "rawMarkdown": "I don't use kernel . I use Colab only. \n\nFor Pytorch the data is uploaded in my Google Drive. For TF, I use GCS bucket.",
      "votes": null
    },
    {
      "id": "917905",
      "postDate": "07/06/2020 19:42:25",
      "content": "<p><a href=\"/serigne\">@serigne</a> the data for this competition is very big, my whole google account has 17GB space, so that's not the option</p>",
      "rawMarkdown": "serigne the data for this competition is very big, my whole google account has 17GB space, so that's not the option",
      "votes": null
    },
    {
      "id": "917916",
      "postDate": "07/06/2020 19:54:21",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> The highest scoring public notebook <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">here</a> only uses 800MB of data. If you use resized JPEGs, 17GB is more than enough.</p>",
      "rawMarkdown": "jacekpoplawski The highest scoring public notebook [here][1] only uses 800MB of data. If you use resized JPEGs, 17GB is more than enough.\n\n[1]: https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once",
      "votes": null
    },
    {
      "id": "917950",
      "postDate": "07/06/2020 20:34:01",
      "content": "<p>You can follow <a href=\"/cdeotte\">@cdeotte</a>  recommandations . I have uploaded both his 512 and 768 on my Gdrive yesterday.  Even though I have 100 GB Gdrive premium storage. </p>",
      "rawMarkdown": "You can follow @cdeotte  recommandations . I have uploaded both his 512 and 768 on my Gdrive yesterday.  Even though I have 100 GB Gdrive premium storage.",
      "votes": null
    },
    {
      "id": "935600",
      "postDate": "07/19/2020 14:13:33",
      "content": "<p>I tried implementing TPU with PyTorch with JPEG's for my work for the first time, I saw in the metrics section that the idle time sometimes jumps to about 70% while sometimes it is 0%. Also, the MX %age remains on 0.</p>\n\n<p>I am not using any multiprocessing and using only one core of TPU.</p>\n\n<p>Now, I tried the same on two different notebooks, one with less augmentation and one with heavy augmentation, b1 completes one epoch in about 2.5 minutes in case of less aug while 12 minutes in case of heavy aug.</p>\n\n<p>Can this be only due to the augmentation? Or I am doing something wrong while implementing the TPU?</p>\n\n<p>Also, again, this is my first time using it, so please tell me any of your thoughts, it might help me in some way! :)</p>",
      "rawMarkdown": "I tried implementing TPU with PyTorch with JPEG's for my work for the first time, I saw in the metrics section that the idle time sometimes jumps to about 70% while sometimes it is 0%. Also, the MX %age remains on 0.\n\nI am not using any multiprocessing and using only one core of TPU.\n\nNow, I tried the same on two different notebooks, one with less augmentation and one with heavy augmentation, b1 completes one epoch in about 2.5 minutes in case of less aug while 12 minutes in case of heavy aug.\n\nCan this be only due to the augmentation? Or I am doing something wrong while implementing the TPU?\n\nAlso, again, this is my first time using it, so please tell me any of your thoughts, it might help me in some way! :)",
      "votes": null
    },
    {
      "id": "935944",
      "postDate": "07/19/2020 19:48:44",
      "content": "<p>I did some experiments with performance on my local GPU first - focus on batch size and number of workers. Yes, augmentation means CPU is used, so GPU must wait for CPU to finish doing things.</p>",
      "rawMarkdown": "I did some experiments with performance on my local GPU first - focus on batch size and number of workers. Yes, augmentation means CPU is used, so GPU must wait for CPU to finish doing things.",
      "votes": null
    },
    {
      "id": "939298",
      "postDate": "07/22/2020 06:39:06",
      "content": "<p>Hey, I saw your issue on Github regarding TPU. Were you able to implement TPU? I am still trying and I can't even start training due to some error. I checked the value of <code>xm.xrt_world_size()</code> in that course and found its value to be 1 rather than 8. I believe it should be 8 in case I want to implement multi-processing. :)</p>",
      "rawMarkdown": "Hey, I saw your issue on Github regarding TPU. Were you able to implement TPU? I am still trying and I can't even start training due to some error. I checked the value of `xm.xrt_world_size()` in that course and found its value to be 1 rather than 8. I believe it should be 8 in case I want to implement multi-processing. :)",
      "votes": null
    },
    {
      "id": "939452",
      "postDate": "07/22/2020 08:50:04",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> xm.xrt<em>world</em>size() used to work fine (i.e. return 8) but they screwed it up in some recent version<br>\nnevertheless, you still have 8 cores, and can hardcode that number…</p>\n<p>pt lightning also works and gives you the full abstraction e.g.: <a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\" target=\"_blank\">https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp</a><br>\nHowever, something causes it to use massive amounts of RAM, especially in Kaggle. It works much better in Colab, as I mention in the comments, but it still much slower than TF with tfrecords</p>\n<p>For those reasons, I'd recommend using TF for now… :( </p>",
      "rawMarkdown": "sarques xm.xrt_world_size() used to work fine (i.e. return 8) but they screwed it up in some recent version\nnevertheless, you still have 8 cores, and can hardcode that number...\n\npt lightning also works and gives you the full abstraction e.g.: https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\nHowever, something causes it to use massive amounts of RAM, especially in Kaggle. It works much better in Colab, as I mention in the comments, but it still much slower than TF with tfrecords\n\nFor those reasons, I'd recommend using TF for now... :(",
      "votes": null
    },
    {
      "id": "939468",
      "postDate": "07/22/2020 08:56:21",
      "content": "<p>Ohh, that's really disheartening, I did read somewhere that only Kaggle kernels are encountering this problem, while Colab is not having any such problem. I am working on this competition from day 1 and was ready to start building some real models. Maybe I should look for other options if hard coding doesn't work, like TF or pytorch lightning as you said, because as of now, I can only train one model for one fold at a time, with a batch size of only 4 and its not feasible. :/</p>\n<p>Thanks a lot for your advice! :)</p>",
      "rawMarkdown": "Ohh, that's really disheartening, I did read somewhere that only Kaggle kernels are encountering this problem, while Colab is not having any such problem. I am working on this competition from day 1 and was ready to start building some real models. Maybe I should look for other options if hard coding doesn't work, like TF or pytorch lightning as you said, because as of now, I can only train one model for one fold at a time, with a batch size of only 4 and its not feasible. :/\n\nThanks a lot for your advice! :)",
      "votes": null
    },
    {
      "id": "941295",
      "postDate": "07/23/2020 06:53:58",
      "content": "<p>I bumped into the same issue in the last couple of days... frustrating! but if you read over this thread on Github, you will understand the issue better. Pytorch/XLA does not have the TPU VM like Tensorflow/TPU has. Therefore, it relies on the Kernal's RAM to pass dataset to TPU. You would think 16G RAM is enough.... but remember, it's x 8 copies for 8 TPU cores. So runs out of memory very quickly. A solution to this, which is the only solution I came across so far, is to train on Colab Pro as it offers 36G RAM. But again, if you have a huge dataset to pass to TPU, there is still risk of running out of memory!! why Tensorflow has no such issue??? Because Google has countless CPU RAMS to offer for TensorFlow!!!\n<a href=\"https://github.com/pytorch/xla/issues/1870\">https://github.com/pytorch/xla/issues/1870</a></p>",
      "rawMarkdown": "I bumped into the same issue in the last couple of days... frustrating! but if you read over this thread on Github, you will understand the issue better. Pytorch/XLA does not have the TPU VM like Tensorflow/TPU has. Therefore, it relies on the Kernal's RAM to pass dataset to TPU. You would think 16G RAM is enough.... but remember, it's x 8 copies for 8 TPU cores. So runs out of memory very quickly. A solution to this, which is the only solution I came across so far, is to train on Colab Pro as it offers 36G RAM. But again, if you have a huge dataset to pass to TPU, there is still risk of running out of memory!! why Tensorflow has no such issue??? Because Google has countless CPU RAMS to offer for TensorFlow!!!\n[https://github.com/pytorch/xla/issues/1870](https://github.com/pytorch/xla/issues/1870)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 915051,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "07/04/2020 12:57:26",
      "content": "<p>TFRecords speedup image data loading, but are not mandatory.</p>\n\n<p>PyTorch on TPUs uses standard \"from torchvision import datasets\", \"datasets.CIFAR10\" and \"torch.utils.data.DataLoader\"\nSee <a href=\"https://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/resnet18-training.ipynb\">PyTorch example on TPU</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 915288,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/04/2020 16:04:02",
          "content": "<p>do you use pytorch in your kernel for this competition?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 915348,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/04/2020 16:57:38",
      "content": "<p>Jacek, I uploaded PyTorch JPEGs <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>. Instead of downloading 110GB of the original data, you only need to download 1GB, 2GB, 3GB, 5GB depending on whether you want 256x256, 384x384, 512x512, 768x768. </p>",
      "votes": null,
      "replies": [
        {
          "id": 915467,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/04/2020 19:07:37",
          "content": "<p>I don't understand, do you mean JPEGs are usable by TPU with pytorch? </p>\n\n<p>I already downloaded all original datasets from 2020 and 2019 but downloading files is for local training, not for TPU on Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 915533,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/04/2020 19:58:19",
          "content": "<p>Yes, JPEGs if you wish to use TPU and PyTorch at Kaggle, CoLab, or GCP.</p>\n\n<p>If you use TPU or GPU and PyTorch, it is most convenient to use folders of JPEGs. If you use TPU and TensorFlow is it most convenient to use TFRecords. If you use GPU and TensorFlow it is equally convenient to use folders of JPEGs or TFRecords.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 915536,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/04/2020 20:01:05",
          "content": "<p>There are example notebooks showing how to do all of the above. TPU are very fast with TFRecords. Previously there were problems maximizing the performance of TPU with PyTorch but i believe that has all been fixed now and PyTorch can achieve similar speed as TensorFlow on TPU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 915552,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/04/2020 20:22:32",
          "content": "<p>Hm, but tensoflow is most popular in this competition, wonder why.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916082,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/05/2020 10:43:52",
          "content": "<p>In case of TPUs, I believe that running PyTorch in a Cloud VM is still slower than TF on the same (VM) hardware.</p>\n\n<p>This is because TF has a dedicated TF server running on the TPU VM. The TF running on a cloud VM does not use much resources because all the input pipelines (including augmentations) actually run on the TPU VMs resources (which are pretty rich in terms on CPU and RAM).</p>\n\n<p>In case of PyTorch the input pipelines run on the cloud VMs, and there is no dedicated PyTorch server on the TPU VM side yet (future enhancement). So, a cloud VM with say 2 CPU cores will become quickly saturated in some cases where the model requires a lot of host computation resources because of a large dataset, custom augmentation pipelines, and smaller model step time. A solution for this is to use higher number of CPU cores, say 32+ cores, and then you will notice that PyTorch performance is now on par with that of TF on TPUs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916155,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/05/2020 12:25:19",
          "content": "<p>What do you mean by \"locally\"? I thought you can use TPU only in the cloud.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916201,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/05/2020 13:18:02",
          "content": "<p>Updated above reply to use consistent terminology \"cloud VM\", instead of locally/user VM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916221,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "07/05/2020 13:34:17",
          "content": "<p>I totally agree with <a href=\"/sirishks\">@sirishks</a> </p>\n\n<p>That's why i use Pytorch TPU on Colab.  You have much more RAM and CPU cores there. </p>\n\n<p>So, Pytorch TPU is blazingly fast on Colab. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916229,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/05/2020 13:40:26",
          "content": "<p><a href=\"/serigne\">@serigne</a> do you mean you train models for this competition on Google Colab? can you tell more about current state of Colab? how do you use Kaggle data? do you need paid account?</p>\n\n<p>EDIT just tested this on colab:</p>\n\n<p>import torch_xla_py.xla_model as xm</p>\n\n<p>ModuleNotFoundError: No module named 'torch_xla_py'</p>\n\n<p>do I have to install it same way as on Kaggle kernels?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916325,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/05/2020 14:56:35",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  Yes, Serigne is using paid account, which gives approx twice the resources and time limits (24-hrs) compared to Colab.</p>\n\n<p>Please see <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161713#903803\">earlier discussion</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916365,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/05/2020 15:26:09",
          "content": "<p><a href=\"/serigne\">@serigne</a> so what are requirements to run pytorch kernel for this kaggle competition? how to use the data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917876,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "07/06/2020 19:21:39",
          "content": "<p>I don't use kernel . I use Colab only. </p>\n\n<p>For Pytorch the data is uploaded in my Google Drive. For TF, I use GCS bucket. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917905,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/06/2020 19:42:25",
          "content": "<p><a href=\"/serigne\">@serigne</a> the data for this competition is very big, my whole google account has 17GB space, so that's not the option</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917916,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/06/2020 19:54:21",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> The highest scoring public notebook <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">here</a> only uses 800MB of data. If you use resized JPEGs, 17GB is more than enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917950,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "07/06/2020 20:34:01",
          "content": "<p>You can follow <a href=\"/cdeotte\">@cdeotte</a>  recommandations . I have uploaded both his 512 and 768 on my Gdrive yesterday.  Even though I have 100 GB Gdrive premium storage. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 915925,
      "author_name": "rhtsingh",
      "author_url": "",
      "post_date": "07/05/2020 07:55:10",
      "content": "<p>You can use JPEGs for tpu training on Pytorch. The tf_records has the same files just a different format. Apart from that if you still want to use tf_records with pytorch there is a tf_record reader in torch_xla utilities. </p>\n\n<p>Havent tried that but yes it exists. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 915992,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "07/05/2020 09:10:08",
      "content": "<p>In addition to the regular torch dataset, PyTorch XLA also implemented a TFRecord reader last year: <a href=\"https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py\">https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 935600,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "07/19/2020 14:13:33",
      "content": "<p>I tried implementing TPU with PyTorch with JPEG's for my work for the first time, I saw in the metrics section that the idle time sometimes jumps to about 70% while sometimes it is 0%. Also, the MX %age remains on 0.</p>\n\n<p>I am not using any multiprocessing and using only one core of TPU.</p>\n\n<p>Now, I tried the same on two different notebooks, one with less augmentation and one with heavy augmentation, b1 completes one epoch in about 2.5 minutes in case of less aug while 12 minutes in case of heavy aug.</p>\n\n<p>Can this be only due to the augmentation? Or I am doing something wrong while implementing the TPU?</p>\n\n<p>Also, again, this is my first time using it, so please tell me any of your thoughts, it might help me in some way! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 935944,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/19/2020 19:48:44",
          "content": "<p>I did some experiments with performance on my local GPU first - focus on batch size and number of workers. Yes, augmentation means CPU is used, so GPU must wait for CPU to finish doing things.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 939298,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "07/22/2020 06:39:06",
          "content": "<p>Hey, I saw your issue on Github regarding TPU. Were you able to implement TPU? I am still trying and I can't even start training due to some error. I checked the value of <code>xm.xrt_world_size()</code> in that course and found its value to be 1 rather than 8. I believe it should be 8 in case I want to implement multi-processing. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 939452,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/22/2020 08:50:04",
          "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> xm.xrt<em>world</em>size() used to work fine (i.e. return 8) but they screwed it up in some recent version<br>\nnevertheless, you still have 8 cores, and can hardcode that number…</p>\n<p>pt lightning also works and gives you the full abstraction e.g.: <a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\" target=\"_blank\">https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp</a><br>\nHowever, something causes it to use massive amounts of RAM, especially in Kaggle. It works much better in Colab, as I mention in the comments, but it still much slower than TF with tfrecords</p>\n<p>For those reasons, I'd recommend using TF for now… :( </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 939468,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "07/22/2020 08:56:21",
          "content": "<p>Ohh, that's really disheartening, I did read somewhere that only Kaggle kernels are encountering this problem, while Colab is not having any such problem. I am working on this competition from day 1 and was ready to start building some real models. Maybe I should look for other options if hard coding doesn't work, like TF or pytorch lightning as you said, because as of now, I can only train one model for one fold at a time, with a batch size of only 4 and its not feasible. :/</p>\n<p>Thanks a lot for your advice! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 941295,
      "author_name": "sheldonburg",
      "author_url": "",
      "post_date": "07/23/2020 06:53:58",
      "content": "<p>I bumped into the same issue in the last couple of days... frustrating! but if you read over this thread on Github, you will understand the issue better. Pytorch/XLA does not have the TPU VM like Tensorflow/TPU has. Therefore, it relies on the Kernal's RAM to pass dataset to TPU. You would think 16G RAM is enough.... but remember, it's x 8 copies for 8 TPU cores. So runs out of memory very quickly. A solution to this, which is the only solution I came across so far, is to train on Colab Pro as it offers 36G RAM. But again, if you have a huge dataset to pass to TPU, there is still risk of running out of memory!! why Tensorflow has no such issue??? Because Google has countless CPU RAMS to offer for TensorFlow!!!\n<a href=\"https://github.com/pytorch/xla/issues/1870\">https://github.com/pytorch/xla/issues/1870</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "915006": "Question to people using pytorch - are we forced to use this tfrecords format or for pytorch it is possible to use TPU with pytorch on Kaggle with normal images? Is it still as fast as people say about tensorflow TPU? I have seen some pytorch kernels but not investigated them yet (I am working on my local GPU during research phase).",
    "915051": "TFRecords speedup image data loading, but are not mandatory.\n\nPyTorch on TPUs uses standard \"from torchvision import datasets\", \"datasets.CIFAR10\" and \"torch.utils.data.DataLoader\"\nSee [PyTorch example on TPU](https://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/resnet18-training.ipynb)",
    "915288": "do you use pytorch in your kernel for this competition?",
    "915348": "Jacek, I uploaded PyTorch JPEGs [here][1]. Instead of downloading 110GB of the original data, you only need to download 1GB, 2GB, 3GB, 5GB depending on whether you want 256x256, 384x384, 512x512, 768x768. \n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
    "915467": "I don't understand, do you mean JPEGs are usable by TPU with pytorch? \n\nI already downloaded all original datasets from 2020 and 2019 but downloading files is for local training, not for TPU on Kaggle.",
    "915533": "Yes, JPEGs if you wish to use TPU and PyTorch at Kaggle, CoLab, or GCP.\n\nIf you use TPU or GPU and PyTorch, it is most convenient to use folders of JPEGs. If you use TPU and TensorFlow is it most convenient to use TFRecords. If you use GPU and TensorFlow it is equally convenient to use folders of JPEGs or TFRecords.",
    "915536": "There are example notebooks showing how to do all of the above. TPU are very fast with TFRecords. Previously there were problems maximizing the performance of TPU with PyTorch but i believe that has all been fixed now and PyTorch can achieve similar speed as TensorFlow on TPU",
    "915552": "Hm, but tensoflow is most popular in this competition, wonder why.",
    "915925": "You can use JPEGs for tpu training on Pytorch. The tf_records has the same files just a different format. Apart from that if you still want to use tf_records with pytorch there is a tf_record reader in torch_xla utilities. \n\nHavent tried that but yes it exists.",
    "915992": "In addition to the regular torch dataset, PyTorch XLA also implemented a TFRecord reader last year: https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py",
    "916082": "In case of TPUs, I believe that running PyTorch in a Cloud VM is still slower than TF on the same (VM) hardware.\n\nThis is because TF has a dedicated TF server running on the TPU VM. The TF running on a cloud VM does not use much resources because all the input pipelines (including augmentations) actually run on the TPU VMs resources (which are pretty rich in terms on CPU and RAM).\n\nIn case of PyTorch the input pipelines run on the cloud VMs, and there is no dedicated PyTorch server on the TPU VM side yet (future enhancement). So, a cloud VM with say 2 CPU cores will become quickly saturated in some cases where the model requires a lot of host computation resources because of a large dataset, custom augmentation pipelines, and smaller model step time. A solution for this is to use higher number of CPU cores, say 32+ cores, and then you will notice that PyTorch performance is now on par with that of TF on TPUs.",
    "916155": "What do you mean by \"locally\"? I thought you can use TPU only in the cloud.",
    "916201": "Updated above reply to use consistent terminology \"cloud VM\", instead of locally/user VM.",
    "916221": "I totally agree with @sirishks \n\nThat's why i use Pytorch TPU on Colab.  You have much more RAM and CPU cores there. \n\nSo, Pytorch TPU is blazingly fast on Colab.",
    "916229": "serigne do you mean you train models for this competition on Google Colab? can you tell more about current state of Colab? how do you use Kaggle data? do you need paid account?\n\nEDIT just tested this on colab:\n\nimport torch_xla_py.xla_model as xm\n\nModuleNotFoundError: No module named 'torch_xla_py'\n\ndo I have to install it same way as on Kaggle kernels?",
    "916325": "jacekpoplawski  Yes, Serigne is using paid account, which gives approx twice the resources and time limits (24-hrs) compared to Colab.\n\nPlease see [earlier discussion](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161713#903803).",
    "916365": "serigne so what are requirements to run pytorch kernel for this kaggle competition? how to use the data?",
    "917876": "I don't use kernel . I use Colab only. \n\nFor Pytorch the data is uploaded in my Google Drive. For TF, I use GCS bucket.",
    "917905": "serigne the data for this competition is very big, my whole google account has 17GB space, so that's not the option",
    "917916": "jacekpoplawski The highest scoring public notebook [here][1] only uses 800MB of data. If you use resized JPEGs, 17GB is more than enough.\n\n[1]: https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once",
    "917950": "You can follow @cdeotte  recommandations . I have uploaded both his 512 and 768 on my Gdrive yesterday.  Even though I have 100 GB Gdrive premium storage.",
    "935600": "I tried implementing TPU with PyTorch with JPEG's for my work for the first time, I saw in the metrics section that the idle time sometimes jumps to about 70% while sometimes it is 0%. Also, the MX %age remains on 0.\n\nI am not using any multiprocessing and using only one core of TPU.\n\nNow, I tried the same on two different notebooks, one with less augmentation and one with heavy augmentation, b1 completes one epoch in about 2.5 minutes in case of less aug while 12 minutes in case of heavy aug.\n\nCan this be only due to the augmentation? Or I am doing something wrong while implementing the TPU?\n\nAlso, again, this is my first time using it, so please tell me any of your thoughts, it might help me in some way! :)",
    "935944": "I did some experiments with performance on my local GPU first - focus on batch size and number of workers. Yes, augmentation means CPU is used, so GPU must wait for CPU to finish doing things.",
    "939298": "Hey, I saw your issue on Github regarding TPU. Were you able to implement TPU? I am still trying and I can't even start training due to some error. I checked the value of `xm.xrt_world_size()` in that course and found its value to be 1 rather than 8. I believe it should be 8 in case I want to implement multi-processing. :)",
    "939452": "sarques xm.xrt_world_size() used to work fine (i.e. return 8) but they screwed it up in some recent version\nnevertheless, you still have 8 cores, and can hardcode that number...\n\npt lightning also works and gives you the full abstraction e.g.: https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\nHowever, something causes it to use massive amounts of RAM, especially in Kaggle. It works much better in Colab, as I mention in the comments, but it still much slower than TF with tfrecords\n\nFor those reasons, I'd recommend using TF for now... :(",
    "939468": "Ohh, that's really disheartening, I did read somewhere that only Kaggle kernels are encountering this problem, while Colab is not having any such problem. I am working on this competition from day 1 and was ready to start building some real models. Maybe I should look for other options if hard coding doesn't work, like TF or pytorch lightning as you said, because as of now, I can only train one model for one fold at a time, with a batch size of only 4 and its not feasible. :/\n\nThanks a lot for your advice! :)",
    "941295": "I bumped into the same issue in the last couple of days... frustrating! but if you read over this thread on Github, you will understand the issue better. Pytorch/XLA does not have the TPU VM like Tensorflow/TPU has. Therefore, it relies on the Kernal's RAM to pass dataset to TPU. You would think 16G RAM is enough.... but remember, it's x 8 copies for 8 TPU cores. So runs out of memory very quickly. A solution to this, which is the only solution I came across so far, is to train on Colab Pro as it offers 36G RAM. But again, if you have a huge dataset to pass to TPU, there is still risk of running out of memory!! why Tensorflow has no such issue??? Because Google has countless CPU RAMS to offer for TensorFlow!!!\n[https://github.com/pytorch/xla/issues/1870](https://github.com/pytorch/xla/issues/1870)"
  },
  "source": "meta"
}