{
  "id": 207495,
  "title": "Help needed! Pytorch with TPU",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207495",
  "author_name": "",
  "post_date": "2020-12-29T22:54:55.971170600Z",
  "votes": 4,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I have been trying to use TPU with modifications to my Pytorch code. After seeing several notebooks I built a pipeline but when I run it my TPU usage is 0. The entire code is running on CPU</p>\n<p>Please refer to the attachment.<br>\nI have also made my notebook public <a href=\"https://www.kaggle.com/debarshichanda/tpu-training\" target=\"_blank\">here</a><br>\nIf anyone can point out the errors I would be very thankful</p>",
  "messages": [
    {
      "id": "1131717",
      "postDate": "12/29/2020 22:54:55",
      "content": "<p>I have been trying to use TPU with modifications to my Pytorch code. After seeing several notebooks I built a pipeline but when I run it my TPU usage is 0. The entire code is running on CPU</p>\n<p>Please refer to the attachment.<br>\nI have also made my notebook public <a href=\"https://www.kaggle.com/debarshichanda/tpu-training\" target=\"_blank\">here</a><br>\nIf anyone can point out the errors I would be very thankful</p>",
      "rawMarkdown": "I have been trying to use TPU with modifications to my Pytorch code. After seeing several notebooks I built a pipeline but when I run it my TPU usage is 0. The entire code is running on CPU\n\nPlease refer to the attachment.\nI have also made my notebook public [here](https://www.kaggle.com/debarshichanda/tpu-training)\nIf anyone can point out the errors I would be very thankful",
      "votes": null
    },
    {
      "id": "1131766",
      "postDate": "12/30/2020 00:17:41",
      "content": "<p>I am working on an XLA starter and I will share once public.</p>",
      "rawMarkdown": "I am working on an XLA starter and I will share once public.",
      "votes": null
    },
    {
      "id": "1131946",
      "postDate": "12/30/2020 05:00:50",
      "content": "<p>Eagerly waiting for this!</p>",
      "rawMarkdown": "Eagerly waiting for this!",
      "votes": null
    },
    {
      "id": "1132833",
      "postDate": "12/30/2020 17:57:38",
      "content": "<p>I'm facing the same problem. Even after moving model and data to TPU device, everything seems to run on CPU.</p>",
      "rawMarkdown": "I'm facing the same problem. Even after moving model and data to TPU device, everything seems to run on CPU.",
      "votes": null
    },
    {
      "id": "1132898",
      "postDate": "12/30/2020 19:08:38",
      "content": "<p>I have made a working notebook for the melanoma competition before. I just rerun my notebook and it was also just using CPU eventhough it was working very well before. I will try the same code on Colab and report again.</p>\n<p>The notebook for TPU (for melanoma competiton): <a href=\"https://www.kaggle.com/aliabdin1/siim-isic-melanoma-pytorch-xla-tpu-8-cores\" target=\"_blank\">https://www.kaggle.com/aliabdin1/siim-isic-melanoma-pytorch-xla-tpu-8-cores</a></p>\n<p><strong>EDIT:</strong> The same problem occurs on Google Colab, its training on CPU only.</p>",
      "rawMarkdown": "I have made a working notebook for the melanoma competition before. I just rerun my notebook and it was also just using CPU eventhough it was working very well before. I will try the same code on Colab and report again.\n\nThe notebook for TPU (for melanoma competiton): https://www.kaggle.com/aliabdin1/siim-isic-melanoma-pytorch-xla-tpu-8-cores\n\n**EDIT:** The same problem occurs on Google Colab, its training on CPU only.",
      "votes": null
    },
    {
      "id": "1132974",
      "postDate": "12/30/2020 20:47:27",
      "content": "<p>Might be an issue with pytorch_xla🙁</p>",
      "rawMarkdown": "Might be an issue with pytorch_xla🙁",
      "votes": null
    },
    {
      "id": "1132988",
      "postDate": "12/30/2020 20:54:54",
      "content": "<p>It could be, but I have tried code from a recent official medium post by PyTorch: <a href=\"https://medium.com/pytorch/pytorch-xla-is-now-generally-available-on-google-cloud-tpus-f9267f437832\" target=\"_blank\">https://medium.com/pytorch/pytorch-xla-is-now-generally-available-on-google-cloud-tpus-f9267f437832</a> </p>\n<p>and it worked on multiple cores, but it also looks like its a CPU trainable tiny model on mnist (28x28 image-size). It could be that they have changed certain ways how things work e.g:</p>\n<pre><code>WRAPPED_MODEL = xmp.MpModelWrapper(ToyModel())\nmodel = WRAPPED_MODEL.to(device)\n</code></pre>\n<p>I havent seen wrapping of models before. </p>",
      "rawMarkdown": "It could be, but I have tried code from a recent official medium post by PyTorch: https://medium.com/pytorch/pytorch-xla-is-now-generally-available-on-google-cloud-tpus-f9267f437832 \n\nand it worked on multiple cores, but it also looks like its a CPU trainable tiny model on mnist (28x28 image-size). It could be that they have changed certain ways how things work e.g:\n\n```\nWRAPPED_MODEL = xmp.MpModelWrapper(ToyModel())\nmodel = WRAPPED_MODEL.to(device)\n```\n\nI havent seen wrapping of models before.",
      "votes": null
    },
    {
      "id": "1133791",
      "postDate": "12/31/2020 14:28:47",
      "content": "<p>I don't think it's issue with PyTorch XLA. Here is what I found bottleneck: <a href=\"https://github.com/pytorch/xla/issues/2707\" target=\"_blank\">https://github.com/pytorch/xla/issues/2707</a></p>",
      "rawMarkdown": "I don't think it's issue with PyTorch XLA. Here is what I found bottleneck: https://github.com/pytorch/xla/issues/2707",
      "votes": null
    },
    {
      "id": "1138850",
      "postDate": "01/05/2021 02:36:28",
      "content": "<p>My PyTorch TPU starter is <a href=\"https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-starter-training\" target=\"_blank\">here</a>. I have noticed that the EfficientNet models in timm should run but are extremely slow. Hence, in my kernel I have used a SE-ResNext50. I have raised an issue in the PyTorch XLA over <a href=\"https://github.com/pytorch/xla/issues/2715\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "My PyTorch TPU starter is [here](https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-starter-training). I have noticed that the EfficientNet models in timm should run but are extremely slow. Hence, in my kernel I have used a SE-ResNext50. I have raised an issue in the PyTorch XLA over [here](https://github.com/pytorch/xla/issues/2715).",
      "votes": null
    },
    {
      "id": "1139701",
      "postDate": "01/05/2021 15:13:31",
      "content": "<p>Sorry to bother you again, I hope that the bug gets fixed quickly<br>\nCan you tell me the batch size you are using for SE-ResNext50?<br>\nIn your kernel you have used resnext50_32x4d</p>",
      "rawMarkdown": "Sorry to bother you again, I hope that the bug gets fixed quickly\nCan you tell me the batch size you are using for SE-ResNext50?\nIn your kernel you have used resnext50_32x4d",
      "votes": null
    },
    {
      "id": "1140161",
      "postDate": "01/05/2021 20:36:32",
      "content": "<p><a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a> is right, changing the EfficientNet model to any other model started to use the TPU instead of CPU. Thank you for this great finding.</p>",
      "rawMarkdown": "tanlikesmath is right, changing the EfficientNet model to any other model started to use the TPU instead of CPU. Thank you for this great finding.",
      "votes": null
    },
    {
      "id": "1140164",
      "postDate": "01/05/2021 20:38:40",
      "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> Error has been found by <a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a>:</p>\n<p>The problem is EfficientNet from <code>timm</code> which is currently not using the TPU during the training, switching to a different model e.g. Resnet50 uses the TPU again and works how it is supposed to.</p>",
      "rawMarkdown": "debarshichanda Error has been found by @tanlikesmath:\n\nThe problem is EfficientNet from `timm` which is currently not using the TPU during the training, switching to a different model e.g. Resnet50 uses the TPU again and works how it is supposed to.",
      "votes": null
    },
    {
      "id": "1140224",
      "postDate": "01/05/2021 21:37:58",
      "content": "<p>Is the problem only in EfficientNet or the entire <code>timm</code> package?</p>",
      "rawMarkdown": "Is the problem only in EfficientNet or the entire `timm` package?",
      "votes": null
    },
    {
      "id": "1140252",
      "postDate": "01/05/2021 22:01:57",
      "content": "<p>I have tried Resnet50 from the <code>timm</code> package and it worked with TPU, but there could be also different models besides EfficientNet which are also affected. </p>",
      "rawMarkdown": "I have tried Resnet50 from the `timm` package and it worked with TPU, but there could be also different models besides EfficientNet which are also affected.",
      "votes": null
    },
    {
      "id": "1140254",
      "postDate": "01/05/2021 22:06:01",
      "content": "<p>Thank you so much for your help!</p>",
      "rawMarkdown": "Thank you so much for your help!",
      "votes": null
    },
    {
      "id": "1140281",
      "postDate": "01/05/2021 22:48:20",
      "content": "<p>Anytime, I'm glad to help :)</p>",
      "rawMarkdown": "Anytime, I'm glad to help :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1131766,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "12/30/2020 00:17:41",
      "content": "<p>I am working on an XLA starter and I will share once public.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131946,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "12/30/2020 05:00:50",
          "content": "<p>Eagerly waiting for this!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1132833,
      "author_name": "kaushal2896",
      "author_url": "",
      "post_date": "12/30/2020 17:57:38",
      "content": "<p>I'm facing the same problem. Even after moving model and data to TPU device, everything seems to run on CPU.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1132898,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "12/30/2020 19:08:38",
      "content": "<p>I have made a working notebook for the melanoma competition before. I just rerun my notebook and it was also just using CPU eventhough it was working very well before. I will try the same code on Colab and report again.</p>\n<p>The notebook for TPU (for melanoma competiton): <a href=\"https://www.kaggle.com/aliabdin1/siim-isic-melanoma-pytorch-xla-tpu-8-cores\" target=\"_blank\">https://www.kaggle.com/aliabdin1/siim-isic-melanoma-pytorch-xla-tpu-8-cores</a></p>\n<p><strong>EDIT:</strong> The same problem occurs on Google Colab, its training on CPU only.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1132974,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "12/30/2020 20:47:27",
          "content": "<p>Might be an issue with pytorch_xla🙁</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132988,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "12/30/2020 20:54:54",
          "content": "<p>It could be, but I have tried code from a recent official medium post by PyTorch: <a href=\"https://medium.com/pytorch/pytorch-xla-is-now-generally-available-on-google-cloud-tpus-f9267f437832\" target=\"_blank\">https://medium.com/pytorch/pytorch-xla-is-now-generally-available-on-google-cloud-tpus-f9267f437832</a> </p>\n<p>and it worked on multiple cores, but it also looks like its a CPU trainable tiny model on mnist (28x28 image-size). It could be that they have changed certain ways how things work e.g:</p>\n<pre><code>WRAPPED_MODEL = xmp.MpModelWrapper(ToyModel())\nmodel = WRAPPED_MODEL.to(device)\n</code></pre>\n<p>I havent seen wrapping of models before. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140164,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/05/2021 20:38:40",
          "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> Error has been found by <a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a>:</p>\n<p>The problem is EfficientNet from <code>timm</code> which is currently not using the TPU during the training, switching to a different model e.g. Resnet50 uses the TPU again and works how it is supposed to.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140224,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "01/05/2021 21:37:58",
          "content": "<p>Is the problem only in EfficientNet or the entire <code>timm</code> package?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140252,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/05/2021 22:01:57",
          "content": "<p>I have tried Resnet50 from the <code>timm</code> package and it worked with TPU, but there could be also different models besides EfficientNet which are also affected. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140254,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "01/05/2021 22:06:01",
          "content": "<p>Thank you so much for your help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140281,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/05/2021 22:48:20",
          "content": "<p>Anytime, I'm glad to help :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1133791,
      "author_name": "kaushal2896",
      "author_url": "",
      "post_date": "12/31/2020 14:28:47",
      "content": "<p>I don't think it's issue with PyTorch XLA. Here is what I found bottleneck: <a href=\"https://github.com/pytorch/xla/issues/2707\" target=\"_blank\">https://github.com/pytorch/xla/issues/2707</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1138850,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "01/05/2021 02:36:28",
      "content": "<p>My PyTorch TPU starter is <a href=\"https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-starter-training\" target=\"_blank\">here</a>. I have noticed that the EfficientNet models in timm should run but are extremely slow. Hence, in my kernel I have used a SE-ResNext50. I have raised an issue in the PyTorch XLA over <a href=\"https://github.com/pytorch/xla/issues/2715\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1139701,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "01/05/2021 15:13:31",
          "content": "<p>Sorry to bother you again, I hope that the bug gets fixed quickly<br>\nCan you tell me the batch size you are using for SE-ResNext50?<br>\nIn your kernel you have used resnext50_32x4d</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140161,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/05/2021 20:36:32",
          "content": "<p><a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a> is right, changing the EfficientNet model to any other model started to use the TPU instead of CPU. Thank you for this great finding.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1131717": "I have been trying to use TPU with modifications to my Pytorch code. After seeing several notebooks I built a pipeline but when I run it my TPU usage is 0. The entire code is running on CPU\n\nPlease refer to the attachment.\nI have also made my notebook public [here](https://www.kaggle.com/debarshichanda/tpu-training)\nIf anyone can point out the errors I would be very thankful",
    "1131766": "I am working on an XLA starter and I will share once public.",
    "1131946": "Eagerly waiting for this!",
    "1132833": "I'm facing the same problem. Even after moving model and data to TPU device, everything seems to run on CPU.",
    "1132898": "I have made a working notebook for the melanoma competition before. I just rerun my notebook and it was also just using CPU eventhough it was working very well before. I will try the same code on Colab and report again.\n\nThe notebook for TPU (for melanoma competiton): https://www.kaggle.com/aliabdin1/siim-isic-melanoma-pytorch-xla-tpu-8-cores\n\n**EDIT:** The same problem occurs on Google Colab, its training on CPU only.",
    "1132974": "Might be an issue with pytorch_xla🙁",
    "1132988": "It could be, but I have tried code from a recent official medium post by PyTorch: https://medium.com/pytorch/pytorch-xla-is-now-generally-available-on-google-cloud-tpus-f9267f437832 \n\nand it worked on multiple cores, but it also looks like its a CPU trainable tiny model on mnist (28x28 image-size). It could be that they have changed certain ways how things work e.g:\n\n```\nWRAPPED_MODEL = xmp.MpModelWrapper(ToyModel())\nmodel = WRAPPED_MODEL.to(device)\n```\n\nI havent seen wrapping of models before.",
    "1133791": "I don't think it's issue with PyTorch XLA. Here is what I found bottleneck: https://github.com/pytorch/xla/issues/2707",
    "1138850": "My PyTorch TPU starter is [here](https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-starter-training). I have noticed that the EfficientNet models in timm should run but are extremely slow. Hence, in my kernel I have used a SE-ResNext50. I have raised an issue in the PyTorch XLA over [here](https://github.com/pytorch/xla/issues/2715).",
    "1139701": "Sorry to bother you again, I hope that the bug gets fixed quickly\nCan you tell me the batch size you are using for SE-ResNext50?\nIn your kernel you have used resnext50_32x4d",
    "1140161": "tanlikesmath is right, changing the EfficientNet model to any other model started to use the TPU instead of CPU. Thank you for this great finding.",
    "1140164": "debarshichanda Error has been found by @tanlikesmath:\n\nThe problem is EfficientNet from `timm` which is currently not using the TPU during the training, switching to a different model e.g. Resnet50 uses the TPU again and works how it is supposed to.",
    "1140224": "Is the problem only in EfficientNet or the entire `timm` package?",
    "1140252": "I have tried Resnet50 from the `timm` package and it worked with TPU, but there could be also different models besides EfficientNet which are also affected.",
    "1140254": "Thank you so much for your help!",
    "1140281": "Anytime, I'm glad to help :)"
  },
  "source": "meta"
}