{
  "id": 150412,
  "title": "TPU MXU Utilization and Idle Time",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/150412",
  "author_name": "",
  "post_date": "2020-05-12T07:47:07.170536Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>What is the maximum MXU utilization you are able to get? I can't get more than 20%. However I can keep idle time below 20%. Am I having some bottleneck when transforming the data? Or these figures are just normal?</p>",
  "messages": [
    {
      "id": "843683",
      "postDate": "05/12/2020 07:47:07",
      "content": "<p>What is the maximum MXU utilization you are able to get? I can't get more than 20%. However I can keep idle time below 20%. Am I having some bottleneck when transforming the data? Or these figures are just normal?</p>",
      "rawMarkdown": "What is the maximum MXU utilization you are able to get? I can't get more than 20%. However I can keep idle time below 20%. Am I having some bottleneck when transforming the data? Or these figures are just normal?",
      "votes": null
    },
    {
      "id": "844811",
      "postDate": "05/12/2020 21:46:37",
      "content": "<p>You can get TPU idle time to 0% using a custom training loop with a host loop that runs multiple batches on the TPU at once. Example here: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">TPU: extreme optimizations</a>\nThis will be much easier in TF 2.3 because the host loop functionality has been added directly into <code>model.compile(..., steps_per_execution=XXX)</code>. Right now, it's a bit of work to get there.</p>\n\n<p>As for the MXU utilization, optimizing it has a lot to do with how the model you are using is built and what kind of padding the compiler has to apply to your tensors to get them onto the 128x128 MXU (MatriX multiplication Unit). If you are using pre-made models, there is not much you can do, apart from making sure your batch size is a multiple of 8 or 128.</p>\n\n<p>More in the TPU perf guide here: <a href=\"https://cloud.google.com/tpu/docs/performance-guide\">https://cloud.google.com/tpu/docs/performance-guide</a>\nEven better is to play with the TPU profiler tool on GCP: <a href=\"https://cloud.google.com/tpu/docs/cloud-tpu-tools\">https://cloud.google.com/tpu/docs/cloud-tpu-tools</a></p>",
      "rawMarkdown": "You can get TPU idle time to 0% using a custom training loop with a host loop that runs multiple batches on the TPU at once. Example here: [TPU: extreme optimizations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443)\nThis will be much easier in TF 2.3 because the host loop functionality has been added directly into `model.compile(..., steps_per_execution=XXX)`. Right now, it's a bit of work to get there.\n\nAs for the MXU utilization, optimizing it has a lot to do with how the model you are using is built and what kind of padding the compiler has to apply to your tensors to get them onto the 128x128 MXU (MatriX multiplication Unit). If you are using pre-made models, there is not much you can do, apart from making sure your batch size is a multiple of 8 or 128.\n\nMore in the TPU perf guide here: https://cloud.google.com/tpu/docs/performance-guide\nEven better is to play with the TPU profiler tool on GCP: https://cloud.google.com/tpu/docs/cloud-tpu-tools",
      "votes": null
    },
    {
      "id": "845605",
      "postDate": "05/13/2020 10:11:46",
      "content": "<p>Thx Martin, I'll take a look at those examples</p>",
      "rawMarkdown": "Thx Martin, I'll take a look at those examples",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 844811,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "05/12/2020 21:46:37",
      "content": "<p>You can get TPU idle time to 0% using a custom training loop with a host loop that runs multiple batches on the TPU at once. Example here: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">TPU: extreme optimizations</a>\nThis will be much easier in TF 2.3 because the host loop functionality has been added directly into <code>model.compile(..., steps_per_execution=XXX)</code>. Right now, it's a bit of work to get there.</p>\n\n<p>As for the MXU utilization, optimizing it has a lot to do with how the model you are using is built and what kind of padding the compiler has to apply to your tensors to get them onto the 128x128 MXU (MatriX multiplication Unit). If you are using pre-made models, there is not much you can do, apart from making sure your batch size is a multiple of 8 or 128.</p>\n\n<p>More in the TPU perf guide here: <a href=\"https://cloud.google.com/tpu/docs/performance-guide\">https://cloud.google.com/tpu/docs/performance-guide</a>\nEven better is to play with the TPU profiler tool on GCP: <a href=\"https://cloud.google.com/tpu/docs/cloud-tpu-tools\">https://cloud.google.com/tpu/docs/cloud-tpu-tools</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 845605,
          "author_name": "benayas",
          "author_url": "",
          "post_date": "05/13/2020 10:11:46",
          "content": "<p>Thx Martin, I'll take a look at those examples</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "843683": "What is the maximum MXU utilization you are able to get? I can't get more than 20%. However I can keep idle time below 20%. Am I having some bottleneck when transforming the data? Or these figures are just normal?",
    "844811": "You can get TPU idle time to 0% using a custom training loop with a host loop that runs multiple batches on the TPU at once. Example here: [TPU: extreme optimizations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443)\nThis will be much easier in TF 2.3 because the host loop functionality has been added directly into `model.compile(..., steps_per_execution=XXX)`. Right now, it's a bit of work to get there.\n\nAs for the MXU utilization, optimizing it has a lot to do with how the model you are using is built and what kind of padding the compiler has to apply to your tensors to get them onto the 128x128 MXU (MatriX multiplication Unit). If you are using pre-made models, there is not much you can do, apart from making sure your batch size is a multiple of 8 or 128.\n\nMore in the TPU perf guide here: https://cloud.google.com/tpu/docs/performance-guide\nEven better is to play with the TPU profiler tool on GCP: https://cloud.google.com/tpu/docs/cloud-tpu-tools",
    "845605": "Thx Martin, I'll take a look at those examples"
  },
  "source": "meta"
}