{
  "id": 133447,
  "title": "TPU Idle time is now displayed in the UI",
  "url": "/competitions/flower-classification-with-tpus/discussion/133447",
  "author_name": "",
  "post_date": "2020-03-02T21:46:52.027270100Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Ff70fc3744bd57681ff109f3f135ed323%2FTPU%20idle%20time2.png?generation=1583185609964304&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "761722",
      "postDate": "03/02/2020 21:46:52",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Ff70fc3744bd57681ff109f3f135ed323%2FTPU%20idle%20time2.png?generation=1583185609964304&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Ff70fc3744bd57681ff109f3f135ed323%2FTPU%20idle%20time2.png?generation=1583185609964304&amp;alt=media)",
      "votes": null
    },
    {
      "id": "761724",
      "postDate": "03/02/2020 21:49:20",
      "content": "<p>Idle time is only displayed when the TPU is training.\nThis is the time when the TPU is waiting for data instead of actually running the training (forward and backward pass).\nYou goal is to optimize your data pipeline to bring this as close as possible to zero.</p>\n\n<p>The screenshot above was taken while training the : <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">Getting Started Notebook</a></p>",
      "rawMarkdown": "Idle time is only displayed when the TPU is training.\nThis is the time when the TPU is waiting for data instead of actually running the training (forward and backward pass).\nYou goal is to optimize your data pipeline to bring this as close as possible to zero.\n\nThe screenshot above was taken while training the : [Getting Started Notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/)",
      "votes": null
    },
    {
      "id": "761753",
      "postDate": "03/02/2020 22:31:34",
      "content": "<p>Glad to know.</p>",
      "rawMarkdown": "Glad to know.",
      "votes": null
    },
    {
      "id": "762067",
      "postDate": "03/03/2020 07:10:32",
      "content": "<p>Seems like unfortunate underutilization of such costly resource. due to limited CPU/RAM power to feed data to the TPU. colab gives access to a fairly high specs machine while using TPU. </p>",
      "rawMarkdown": "Seems like unfortunate underutilization of such costly resource. due to limited CPU/RAM power to feed data to the TPU. colab gives access to a fairly high specs machine while using TPU.",
      "votes": null
    },
    {
      "id": "762570",
      "postDate": "03/03/2020 16:06:36",
      "content": "<p>You meant CPU power provided by Kaggle kernels is low?</p>",
      "rawMarkdown": "You meant CPU power provided by Kaggle kernels is low?",
      "votes": null
    },
    {
      "id": "762622",
      "postDate": "03/03/2020 16:51:00",
      "content": "<p>The machine the TPU is attached to (through a PCI link) is designed to have the necessary power never to become a bottleneck. When your VM (Colab or Kaggle) connects to a \"TPU\" it connects to the whole set. While training, the Colab or Kaggle VM do practically nothing, the data pipeline as well as the forward and backward pass of training are executed on the powerful TPU side.</p>",
      "rawMarkdown": "The machine the TPU is attached to (through a PCI link) is designed to have the necessary power never to become a bottleneck. When your VM (Colab or Kaggle) connects to a \"TPU\" it connects to the whole set. While training, the Colab or Kaggle VM do practically nothing, the data pipeline as well as the forward and backward pass of training are executed on the powerful TPU side.",
      "votes": null
    },
    {
      "id": "762777",
      "postDate": "03/03/2020 19:51:04",
      "content": "<p>We added another new feature to help you save precious TPU quota. If a notebook has a TPU accelerator ON but never uses it, we will alert you the next time you open or fork the notebook.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Fb3de410c85e788c2c928258dde6b936d%2FScreen%20Shot%202020-03-03%20at%2011.14.40.png?generation=1583264802357148&amp;alt=media\" alt=\"\"></p>\n\n<p>Please report any cases where you see this message and think it is unhelpful (and please include the session ID in your message, i.e. the number in the URL like <a href=\"https://.../.../edit/run/29598674\">https://.../.../edit/run/29598674</a>)</p>",
      "rawMarkdown": "We added another new feature to help you save precious TPU quota. If a notebook has a TPU accelerator ON but never uses it, we will alert you the next time you open or fork the notebook.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Fb3de410c85e788c2c928258dde6b936d%2FScreen%20Shot%202020-03-03%20at%2011.14.40.png?generation=1583264802357148&amp;alt=media)\n\nPlease report any cases where you see this message and think it is unhelpful (and please include the session ID in your message, i.e. the number in the URL like https://.../.../edit/run/29598674)",
      "votes": null
    },
    {
      "id": "965457",
      "postDate": "08/10/2020 16:47:16",
      "content": "<p>It will be great if there are some guidelines to illustrate how to drive TPU idle time to zero. I consistently have it around 50%. I use PyTorch and I tried increasing the data loaders' num_workers hoping to supply data faster for training, but this did not help. </p>",
      "rawMarkdown": "It will be great if there are some guidelines to illustrate how to drive TPU idle time to zero. I consistently have it around 50%. I use PyTorch and I tried increasing the data loaders' num_workers hoping to supply data faster for training, but this did not help.",
      "votes": null
    },
    {
      "id": "1260153",
      "postDate": "04/01/2021 22:26:18",
      "content": "<p>I have the same question as Krishna.  Would increasing batch size work?  I also have this set up in my get_training_dataset function:   <br>\n<code>AUTO = tf.data.experimental.AUTOTUNE</code><br>\n<code>dataset = dataset.prefetch(AUTO)</code></p>",
      "rawMarkdown": "I have the same question as Krishna.  Would increasing batch size work?  I also have this set up in my get_training_dataset function:   \n`AUTO = tf.data.experimental.AUTOTUNE`\n`dataset = dataset.prefetch(AUTO)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 761724,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/02/2020 21:49:20",
      "content": "<p>Idle time is only displayed when the TPU is training.\nThis is the time when the TPU is waiting for data instead of actually running the training (forward and backward pass).\nYou goal is to optimize your data pipeline to bring this as close as possible to zero.</p>\n\n<p>The screenshot above was taken while training the : <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">Getting Started Notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 965457,
          "author_name": "krisho007",
          "author_url": "",
          "post_date": "08/10/2020 16:47:16",
          "content": "<p>It will be great if there are some guidelines to illustrate how to drive TPU idle time to zero. I consistently have it around 50%. I use PyTorch and I tried increasing the data loaders' num_workers hoping to supply data faster for training, but this did not help. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1260153,
          "author_name": "saukha",
          "author_url": "",
          "post_date": "04/01/2021 22:26:18",
          "content": "<p>I have the same question as Krishna.  Would increasing batch size work?  I also have this set up in my get_training_dataset function:   <br>\n<code>AUTO = tf.data.experimental.AUTOTUNE</code><br>\n<code>dataset = dataset.prefetch(AUTO)</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761753,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "03/02/2020 22:31:34",
      "content": "<p>Glad to know.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 762067,
      "author_name": "dhananjay3",
      "author_url": "",
      "post_date": "03/03/2020 07:10:32",
      "content": "<p>Seems like unfortunate underutilization of such costly resource. due to limited CPU/RAM power to feed data to the TPU. colab gives access to a fairly high specs machine while using TPU. </p>",
      "votes": null,
      "replies": [
        {
          "id": 762570,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "03/03/2020 16:06:36",
          "content": "<p>You meant CPU power provided by Kaggle kernels is low?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 762622,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/03/2020 16:51:00",
          "content": "<p>The machine the TPU is attached to (through a PCI link) is designed to have the necessary power never to become a bottleneck. When your VM (Colab or Kaggle) connects to a \"TPU\" it connects to the whole set. While training, the Colab or Kaggle VM do practically nothing, the data pipeline as well as the forward and backward pass of training are executed on the powerful TPU side.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 762777,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/03/2020 19:51:04",
      "content": "<p>We added another new feature to help you save precious TPU quota. If a notebook has a TPU accelerator ON but never uses it, we will alert you the next time you open or fork the notebook.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Fb3de410c85e788c2c928258dde6b936d%2FScreen%20Shot%202020-03-03%20at%2011.14.40.png?generation=1583264802357148&amp;alt=media\" alt=\"\"></p>\n\n<p>Please report any cases where you see this message and think it is unhelpful (and please include the session ID in your message, i.e. the number in the URL like <a href=\"https://.../.../edit/run/29598674\">https://.../.../edit/run/29598674</a>)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "761722": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Ff70fc3744bd57681ff109f3f135ed323%2FTPU%20idle%20time2.png?generation=1583185609964304&amp;alt=media)",
    "761724": "Idle time is only displayed when the TPU is training.\nThis is the time when the TPU is waiting for data instead of actually running the training (forward and backward pass).\nYou goal is to optimize your data pipeline to bring this as close as possible to zero.\n\nThe screenshot above was taken while training the : [Getting Started Notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/)",
    "761753": "Glad to know.",
    "762067": "Seems like unfortunate underutilization of such costly resource. due to limited CPU/RAM power to feed data to the TPU. colab gives access to a fairly high specs machine while using TPU.",
    "762570": "You meant CPU power provided by Kaggle kernels is low?",
    "762622": "The machine the TPU is attached to (through a PCI link) is designed to have the necessary power never to become a bottleneck. When your VM (Colab or Kaggle) connects to a \"TPU\" it connects to the whole set. While training, the Colab or Kaggle VM do practically nothing, the data pipeline as well as the forward and backward pass of training are executed on the powerful TPU side.",
    "762777": "We added another new feature to help you save precious TPU quota. If a notebook has a TPU accelerator ON but never uses it, we will alert you the next time you open or fork the notebook.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Fb3de410c85e788c2c928258dde6b936d%2FScreen%20Shot%202020-03-03%20at%2011.14.40.png?generation=1583264802357148&amp;alt=media)\n\nPlease report any cases where you see this message and think it is unhelpful (and please include the session ID in your message, i.e. the number in the URL like https://.../.../edit/run/29598674)",
    "965457": "It will be great if there are some guidelines to illustrate how to drive TPU idle time to zero. I consistently have it around 50%. I use PyTorch and I tried increasing the data loaders' num_workers hoping to supply data faster for training, but this did not help.",
    "1260153": "I have the same question as Krishna.  Would increasing batch size work?  I also have this set up in my get_training_dataset function:   \n`AUTO = tf.data.experimental.AUTOTUNE`\n`dataset = dataset.prefetch(AUTO)`"
  },
  "source": "meta"
}