{
  "id": 154854,
  "title": "Ask to TPU users",
  "url": "/competitions/alaska2-image-steganalysis/discussion/154854",
  "author_name": "",
  "post_date": "2020-05-30T04:53:20.990070Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Is this normal for TPU? This is my first using TPU and I think it's kinda of slow...or am I missing something?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F3ca1b18e35612da69e39fbd7a4757d7f%2Ftpu_kaggle.bmp?generation=1590814367137454&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "867203",
      "postDate": "05/30/2020 04:53:20",
      "content": "<p>Is this normal for TPU? This is my first using TPU and I think it's kinda of slow...or am I missing something?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F3ca1b18e35612da69e39fbd7a4757d7f%2Ftpu_kaggle.bmp?generation=1590814367137454&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Is this normal for TPU? This is my first using TPU and I think it's kinda of slow...or am I missing something?![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F3ca1b18e35612da69e39fbd7a4757d7f%2Ftpu_kaggle.bmp?generation=1590814367137454&amp;alt=media)",
      "votes": null
    },
    {
      "id": "867407",
      "postDate": "05/30/2020 08:55:55",
      "content": "<p>i wasted a lot of time on pytorch tpu for this competition.\npytorch tpu is using colab's v2 where tf tpu is using v3,,tf tpu is way way more faster than pytorch tpu\nmy personal recommendation : \"don't waste your working hours using pytorch tpu for any computer vision task,instead use tf tpu or pytorch gpu\". xla team confirmed me that they are working on it still,by the end of this year they will hopefully make torch xla much faster and then we probably will get torch tpu which will be close to tf tpu in terms of performance and speed</p>",
      "rawMarkdown": "i wasted a lot of time on pytorch tpu for this competition.\npytorch tpu is using colab's v2 where tf tpu is using v3,,tf tpu is way way more faster than pytorch tpu\nmy personal recommendation : \"don't waste your working hours using pytorch tpu for any computer vision task,instead use tf tpu or pytorch gpu\". xla team confirmed me that they are working on it still,by the end of this year they will hopefully make torch xla much faster and then we probably will get torch tpu which will be close to tf tpu in terms of performance and speed",
      "votes": null
    },
    {
      "id": "867692",
      "postDate": "05/30/2020 14:30:41",
      "content": "<p>Thanks man!! No wonder it was super slow :(</p>",
      "rawMarkdown": "Thanks man!! No wonder it was super slow :(",
      "votes": null
    },
    {
      "id": "872484",
      "postDate": "06/03/2020 08:24:47",
      "content": "<p>To add to that, I think with 16GB of RAM, I find that having <code>num_workers &amp;gt; 0</code> in the <code>DataLoader</code> I quickly run out of memory. I have a feeling that maybe one of the bottlenecks is in the data preparation. The XLA README also says:</p>\n\n<p>&gt; ideally create a VM that has at least 16 cores (n1-standard-16) to not be VM compute/network bound.</p>\n\n<p>Shame as I was looking forward to trying out TPUs for the first time :(</p>",
      "rawMarkdown": "To add to that, I think with 16GB of RAM, I find that having `num_workers &gt; 0` in the `DataLoader` I quickly run out of memory. I have a feeling that maybe one of the bottlenecks is in the data preparation. The XLA README also says:\n\n&gt; ideally create a VM that has at least 16 cores (n1-standard-16) to not be VM compute/network bound.\n\nShame as I was looking forward to trying out TPUs for the first time :(",
      "votes": null
    },
    {
      "id": "872620",
      "postDate": "06/03/2020 11:23:00",
      "content": "<p><a href=\"/bibek777\">@bibek777</a>  and <a href=\"/anjum48\">@anjum48</a> \nI strongly suggest you to  read <a href=\"https://github.com/pytorch/xla/issues/1870\"><strong>XLM-R model OOM (PyTorch XLA limitations vs TF</strong></a> . There are a lot of tips lying around that can help you with not having OOM as such</p>",
      "rawMarkdown": "bibek777  and @anjum48 \nI strongly suggest you to  read [**XLM-R model OOM (PyTorch XLA limitations vs TF**](https://github.com/pytorch/xla/issues/1870) . There are a lot of tips lying around that can help you with not having OOM as such",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 867407,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "05/30/2020 08:55:55",
      "content": "<p>i wasted a lot of time on pytorch tpu for this competition.\npytorch tpu is using colab's v2 where tf tpu is using v3,,tf tpu is way way more faster than pytorch tpu\nmy personal recommendation : \"don't waste your working hours using pytorch tpu for any computer vision task,instead use tf tpu or pytorch gpu\". xla team confirmed me that they are working on it still,by the end of this year they will hopefully make torch xla much faster and then we probably will get torch tpu which will be close to tf tpu in terms of performance and speed</p>",
      "votes": null,
      "replies": [
        {
          "id": 867692,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "05/30/2020 14:30:41",
          "content": "<p>Thanks man!! No wonder it was super slow :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 872484,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "06/03/2020 08:24:47",
          "content": "<p>To add to that, I think with 16GB of RAM, I find that having <code>num_workers &amp;gt; 0</code> in the <code>DataLoader</code> I quickly run out of memory. I have a feeling that maybe one of the bottlenecks is in the data preparation. The XLA README also says:</p>\n\n<p>&gt; ideally create a VM that has at least 16 cores (n1-standard-16) to not be VM compute/network bound.</p>\n\n<p>Shame as I was looking forward to trying out TPUs for the first time :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 872620,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "06/03/2020 11:23:00",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>  and <a href=\"/anjum48\">@anjum48</a> \nI strongly suggest you to  read <a href=\"https://github.com/pytorch/xla/issues/1870\"><strong>XLM-R model OOM (PyTorch XLA limitations vs TF</strong></a> . There are a lot of tips lying around that can help you with not having OOM as such</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "867203": "Is this normal for TPU? This is my first using TPU and I think it's kinda of slow...or am I missing something?![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F3ca1b18e35612da69e39fbd7a4757d7f%2Ftpu_kaggle.bmp?generation=1590814367137454&amp;alt=media)",
    "867407": "i wasted a lot of time on pytorch tpu for this competition.\npytorch tpu is using colab's v2 where tf tpu is using v3,,tf tpu is way way more faster than pytorch tpu\nmy personal recommendation : \"don't waste your working hours using pytorch tpu for any computer vision task,instead use tf tpu or pytorch gpu\". xla team confirmed me that they are working on it still,by the end of this year they will hopefully make torch xla much faster and then we probably will get torch tpu which will be close to tf tpu in terms of performance and speed",
    "867692": "Thanks man!! No wonder it was super slow :(",
    "872484": "To add to that, I think with 16GB of RAM, I find that having `num_workers &gt; 0` in the `DataLoader` I quickly run out of memory. I have a feeling that maybe one of the bottlenecks is in the data preparation. The XLA README also says:\n\n&gt; ideally create a VM that has at least 16 cores (n1-standard-16) to not be VM compute/network bound.\n\nShame as I was looking forward to trying out TPUs for the first time :(",
    "872620": "bibek777  and @anjum48 \nI strongly suggest you to  read [**XLM-R model OOM (PyTorch XLA limitations vs TF**](https://github.com/pytorch/xla/issues/1870) . There are a lot of tips lying around that can help you with not having OOM as such"
  },
  "source": "meta"
}