{
  "id": 182515,
  "title": "Cloud TPU vs Colab TPU? Also, any tricks to speed up TPU training?",
  "url": "/competitions/landmark-recognition-2020/discussion/182515",
  "author_name": "",
  "post_date": "2020-09-13T07:02:00.603443700Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>So far I've been testing my training pipeline and training my models on Google Colab Pro + TPU (and also using <a href=\"https://stackoverflow.com/questions/57113226/how-to-prevent-google-colab-from-disconnecting\" target=\"_blank\">this JS trick to keep the Colab session alive</a>). However, 4.2 hours for EfficientNetB7 and 2.5 hours for ResNet152 is too slow. I'm wondering if Cloud Services TPU (assuming that I'm using the v3-8) can give substancially better performance?</p>\n<blockquote>\n  <p><strong>Question:</strong> can Google Cloud TPU give a substancially better performance in comparison to Colab Pro TPUs?</p>\n</blockquote>\n<p>Also, I would be very thankful if you can share your experience on using TPU. Here is what I've been doing so far:</p>\n<ul>\n<li>I use <code>tf.data.experimental.AUTOTUNE</code> wherever I can in the data pipeline</li>\n<li>I use <code>tf.data.Dataset.prefetch</code> to prepare the next batch of data in parallel</li>\n</ul>",
  "messages": [
    {
      "id": "1008504",
      "postDate": "09/13/2020 07:02:00",
      "content": "<p>So far I've been testing my training pipeline and training my models on Google Colab Pro + TPU (and also using <a href=\"https://stackoverflow.com/questions/57113226/how-to-prevent-google-colab-from-disconnecting\" target=\"_blank\">this JS trick to keep the Colab session alive</a>). However, 4.2 hours for EfficientNetB7 and 2.5 hours for ResNet152 is too slow. I'm wondering if Cloud Services TPU (assuming that I'm using the v3-8) can give substancially better performance?</p>\n<blockquote>\n  <p><strong>Question:</strong> can Google Cloud TPU give a substancially better performance in comparison to Colab Pro TPUs?</p>\n</blockquote>\n<p>Also, I would be very thankful if you can share your experience on using TPU. Here is what I've been doing so far:</p>\n<ul>\n<li>I use <code>tf.data.experimental.AUTOTUNE</code> wherever I can in the data pipeline</li>\n<li>I use <code>tf.data.Dataset.prefetch</code> to prepare the next batch of data in parallel</li>\n</ul>",
      "rawMarkdown": "So far I've been testing my training pipeline and training my models on Google Colab Pro + TPU (and also using [this JS trick to keep the Colab session alive](https://stackoverflow.com/questions/57113226/how-to-prevent-google-colab-from-disconnecting)). However, 4.2 hours for EfficientNetB7 and 2.5 hours for ResNet152 is too slow. I'm wondering if Cloud Services TPU (assuming that I'm using the v3-8) can give substancially better performance?\n\n> **Question:** can Google Cloud TPU give a substancially better performance in comparison to Colab Pro TPUs?\n\nAlso, I would be very thankful if you can share your experience on using TPU. Here is what I've been doing so far:\n* I use `tf.data.experimental.AUTOTUNE` wherever I can in the data pipeline\n* I use `tf.data.Dataset.prefetch` to prepare the next batch of data in parallel",
      "votes": null
    },
    {
      "id": "1008627",
      "postDate": "09/13/2020 09:23:00",
      "content": "<p>Google Cloud TPU uses TPU v3.8 with the lastest version kernel. It's fast like Kaggle TPU kernel. Google Colab use TPU v2.8 is much slower than GCP.</p>",
      "rawMarkdown": "Google Cloud TPU uses TPU v3.8 with the lastest version kernel. It's fast like Kaggle TPU kernel. Google Colab use TPU v2.8 is much slower than GCP.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1008627,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "09/13/2020 09:23:00",
      "content": "<p>Google Cloud TPU uses TPU v3.8 with the lastest version kernel. It's fast like Kaggle TPU kernel. Google Colab use TPU v2.8 is much slower than GCP.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1008504": "So far I've been testing my training pipeline and training my models on Google Colab Pro + TPU (and also using [this JS trick to keep the Colab session alive](https://stackoverflow.com/questions/57113226/how-to-prevent-google-colab-from-disconnecting)). However, 4.2 hours for EfficientNetB7 and 2.5 hours for ResNet152 is too slow. I'm wondering if Cloud Services TPU (assuming that I'm using the v3-8) can give substancially better performance?\n\n> **Question:** can Google Cloud TPU give a substancially better performance in comparison to Colab Pro TPUs?\n\nAlso, I would be very thankful if you can share your experience on using TPU. Here is what I've been doing so far:\n* I use `tf.data.experimental.AUTOTUNE` wherever I can in the data pipeline\n* I use `tf.data.Dataset.prefetch` to prepare the next batch of data in parallel",
    "1008627": "Google Cloud TPU uses TPU v3.8 with the lastest version kernel. It's fast like Kaggle TPU kernel. Google Colab use TPU v2.8 is much slower than GCP."
  },
  "source": "meta"
}