{
  "id": 212227,
  "title": "TPU works obtusely on kaggle kernel",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/212227",
  "author_name": "",
  "post_date": "2021-01-18T05:48:02.312613400Z",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Started  from yesterday, I'm curious why the utilization rate of TPU is so low:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6221164%2F889327a9f8ddcd9bb867bcd6bc361a12%2Fbc10d148e634217dbbe4d9c193652b3.png?generation=1610948869233495&amp;alt=media\" alt=\"\"><br>\nwhich might be a reason for my code (using TPU) to be implemented quite slowly (even the public baseline). Did you feel puzzled about the same case?  </p>",
  "messages": [
    {
      "id": "1157757",
      "postDate": "01/18/2021 05:48:02",
      "content": "<p>Started  from yesterday, I'm curious why the utilization rate of TPU is so low:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6221164%2F889327a9f8ddcd9bb867bcd6bc361a12%2Fbc10d148e634217dbbe4d9c193652b3.png?generation=1610948869233495&amp;alt=media\" alt=\"\"><br>\nwhich might be a reason for my code (using TPU) to be implemented quite slowly (even the public baseline). Did you feel puzzled about the same case?  </p>",
      "rawMarkdown": "Started  from yesterday, I'm curious why the utilization rate of TPU is so low:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6221164%2F889327a9f8ddcd9bb867bcd6bc361a12%2Fbc10d148e634217dbbe4d9c193652b3.png?generation=1610948869233495&alt=media)\nwhich might be a reason for my code (using TPU) to be implemented quite slowly (even the public baseline). Did you feel puzzled about the same case?",
      "votes": null
    },
    {
      "id": "1157761",
      "postDate": "01/18/2021 05:54:59",
      "content": "<p>are you using pytorch xla or tf tpu?</p>",
      "rawMarkdown": "are you using pytorch xla or tf tpu?",
      "votes": null
    },
    {
      "id": "1157767",
      "postDate": "01/18/2021 06:05:50",
      "content": "<p>tf tpu, similar as the public kernel:<br>\n<a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train</a></p>",
      "rawMarkdown": "tf tpu, similar as the public kernel:\nhttps://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train",
      "votes": null
    },
    {
      "id": "1157774",
      "postDate": "01/18/2021 06:15:04",
      "content": "<p>i observed that while using tpu pod,the utilization rate of TPU iseven  much lower than tpu v3-8<br>\n<a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> </p>",
      "rawMarkdown": "i observed that while using tpu pod,the utilization rate of TPU iseven  much lower than tpu v3-8\n@herbison",
      "votes": null
    },
    {
      "id": "1159976",
      "postDate": "01/19/2021 15:30:13",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> Do you mean the utilization rate is less than it was before (like this post seems to imply), or  just that you see a lower utilization rate between regular TPUs and TPU pods (I'll pass this to TPU experts for a better answer, but my initial reaction would be that TPU pods probably have more processing power and therefore might be able to handle similar throughput with less total utilization).</p>",
      "rawMarkdown": "mobassir Do you mean the utilization rate is less than it was before (like this post seems to imply), or  just that you see a lower utilization rate between regular TPUs and TPU pods (I'll pass this to TPU experts for a better answer, but my initial reaction would be that TPU pods probably have more processing power and therefore might be able to handle similar throughput with less total utilization).",
      "votes": null
    },
    {
      "id": "1159985",
      "postDate": "01/19/2021 15:39:23",
      "content": "<p>dear <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> <br>\ni observed lower utilization rate between regular TPUs and TPU pods in tf tpu,i will test in torch tpu soon and will let you know what i get,thank you</p>",
      "rawMarkdown": "dear @herbison \ni observed lower utilization rate between regular TPUs and TPU pods in tf tpu,i will test in torch tpu soon and will let you know what i get,thank you",
      "votes": null
    },
    {
      "id": "1164371",
      "postDate": "01/22/2021 10:30:07",
      "content": "<p>Thank you for your help sir. what did you observe from torch tpu? my TPUs recovered with no reason and can work normally now</p>",
      "rawMarkdown": "Thank you for your help sir. what did you observe from torch tpu? my TPUs recovered with no reason and can work normally now",
      "votes": null
    },
    {
      "id": "1164410",
      "postDate": "01/22/2021 11:02:26",
      "content": "<p><a href=\"https://www.kaggle.com/southsakura\" target=\"_blank\">@southsakura</a> i have lost my all tpu hours few days ago,i will be able to use tpu again from tomorrow for next week :D (will check then) thank you<br>\ni observed in torch tpu if i go for long train then most of the time after 8-10 epochs for single fold i get OOM,,,OOM comes after 8-10 epochs and i don't know how to handle this issue,i lost a lot of tpu hour that way,,,,get OOM within 2 hours and then commit gets stuck and eats 9 hours TPU quota for single experiment run(that too ends with an error), <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> and team working on fixing this issue i guess </p>",
      "rawMarkdown": "southsakura i have lost my all tpu hours few days ago,i will be able to use tpu again from tomorrow for next week :D (will check then) thank you\ni observed in torch tpu if i go for long train then most of the time after 8-10 epochs for single fold i get OOM,,,OOM comes after 8-10 epochs and i don't know how to handle this issue,i lost a lot of tpu hour that way,,,,get OOM within 2 hours and then commit gets stuck and eats 9 hours TPU quota for single experiment run(that too ends with an error), @herbison and team working on fixing this issue i guess",
      "votes": null
    },
    {
      "id": "1164990",
      "postDate": "01/22/2021 16:42:46",
      "content": "<p>Yes, my colleagues are digging into the OOM causing overrun issue.</p>\n<p>TPU utilization is independent of that though.</p>",
      "rawMarkdown": "Yes, my colleagues are digging into the OOM causing overrun issue.\n\nTPU utilization is independent of that though.",
      "votes": null
    },
    {
      "id": "1173636",
      "postDate": "01/28/2021 03:00:28",
      "content": "<p>TF 2.4 introduces a new parameter in tf.keras.Model.compile called steps_per_execution. <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/Model#compile\" target=\"_blank\">Docs here</a>.</p>\n<p>It send multiple batches to the TPU for processing before returning control to Keras. Not only does it shave TPU&lt;&gt;Keras interop overheads, but if multiple batches are processed on the TPU at once, the XLA compiler can unroll the loop and optimize matrix filling. You can do the same by had using a custom training loop if you are not afraid of writing a fair bit of code as I explained a while back in <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\" target=\"_blank\">TPU: extreme optimizations</a>.</p>\n<p>This setting is even more important, from a performance point of view on TPU pods. Stay tuned for TF 2.4.</p>\n<p>Final note: 22% MXU utilization is not necessarily horrible. I am not 100% sure the new parameter will push it higher. Sometimes, the only thing you can do to push MXU utilization higher is to look at individual matrix paddings as reported by <a href=\"https://cloud.google.com/tpu/docs/cloud-tpu-tools\" target=\"_blank\">capture_tpu_profile</a> and then adjust layer sizes to make it better. Not something you can do on Kaggle or when using pre-trained models.</p>",
      "rawMarkdown": "TF 2.4 introduces a new parameter in tf.keras.Model.compile called steps_per_execution. [Docs here](https://www.tensorflow.org/api_docs/python/tf/keras/Model#compile).\n\nIt send multiple batches to the TPU for processing before returning control to Keras. Not only does it shave TPU<>Keras interop overheads, but if multiple batches are processed on the TPU at once, the XLA compiler can unroll the loop and optimize matrix filling. You can do the same by had using a custom training loop if you are not afraid of writing a fair bit of code as I explained a while back in [TPU: extreme optimizations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443).\n\nThis setting is even more important, from a performance point of view on TPU pods. Stay tuned for TF 2.4.\n\nFinal note: 22% MXU utilization is not necessarily horrible. I am not 100% sure the new parameter will push it higher. Sometimes, the only thing you can do to push MXU utilization higher is to look at individual matrix paddings as reported by [capture_tpu_profile](https://cloud.google.com/tpu/docs/cloud-tpu-tools) and then adjust layer sizes to make it better. Not something you can do on Kaggle or when using pre-trained models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1157761,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "01/18/2021 05:54:59",
      "content": "<p>are you using pytorch xla or tf tpu?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1157767,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "01/18/2021 06:05:50",
          "content": "<p>tf tpu, similar as the public kernel:<br>\n<a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1157774,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/18/2021 06:15:04",
          "content": "<p>i observed that while using tpu pod,the utilization rate of TPU iseven  much lower than tpu v3-8<br>\n<a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1159976,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "01/19/2021 15:30:13",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> Do you mean the utilization rate is less than it was before (like this post seems to imply), or  just that you see a lower utilization rate between regular TPUs and TPU pods (I'll pass this to TPU experts for a better answer, but my initial reaction would be that TPU pods probably have more processing power and therefore might be able to handle similar throughput with less total utilization).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1159985,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/19/2021 15:39:23",
          "content": "<p>dear <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> <br>\ni observed lower utilization rate between regular TPUs and TPU pods in tf tpu,i will test in torch tpu soon and will let you know what i get,thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1164371,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "01/22/2021 10:30:07",
          "content": "<p>Thank you for your help sir. what did you observe from torch tpu? my TPUs recovered with no reason and can work normally now</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1164410,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/22/2021 11:02:26",
          "content": "<p><a href=\"https://www.kaggle.com/southsakura\" target=\"_blank\">@southsakura</a> i have lost my all tpu hours few days ago,i will be able to use tpu again from tomorrow for next week :D (will check then) thank you<br>\ni observed in torch tpu if i go for long train then most of the time after 8-10 epochs for single fold i get OOM,,,OOM comes after 8-10 epochs and i don't know how to handle this issue,i lost a lot of tpu hour that way,,,,get OOM within 2 hours and then commit gets stuck and eats 9 hours TPU quota for single experiment run(that too ends with an error), <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> and team working on fixing this issue i guess </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1164990,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "01/22/2021 16:42:46",
          "content": "<p>Yes, my colleagues are digging into the OOM causing overrun issue.</p>\n<p>TPU utilization is independent of that though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1173636,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "01/28/2021 03:00:28",
      "content": "<p>TF 2.4 introduces a new parameter in tf.keras.Model.compile called steps_per_execution. <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/Model#compile\" target=\"_blank\">Docs here</a>.</p>\n<p>It send multiple batches to the TPU for processing before returning control to Keras. Not only does it shave TPU&lt;&gt;Keras interop overheads, but if multiple batches are processed on the TPU at once, the XLA compiler can unroll the loop and optimize matrix filling. You can do the same by had using a custom training loop if you are not afraid of writing a fair bit of code as I explained a while back in <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\" target=\"_blank\">TPU: extreme optimizations</a>.</p>\n<p>This setting is even more important, from a performance point of view on TPU pods. Stay tuned for TF 2.4.</p>\n<p>Final note: 22% MXU utilization is not necessarily horrible. I am not 100% sure the new parameter will push it higher. Sometimes, the only thing you can do to push MXU utilization higher is to look at individual matrix paddings as reported by <a href=\"https://cloud.google.com/tpu/docs/cloud-tpu-tools\" target=\"_blank\">capture_tpu_profile</a> and then adjust layer sizes to make it better. Not something you can do on Kaggle or when using pre-trained models.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1157757": "Started  from yesterday, I'm curious why the utilization rate of TPU is so low:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6221164%2F889327a9f8ddcd9bb867bcd6bc361a12%2Fbc10d148e634217dbbe4d9c193652b3.png?generation=1610948869233495&alt=media)\nwhich might be a reason for my code (using TPU) to be implemented quite slowly (even the public baseline). Did you feel puzzled about the same case?",
    "1157761": "are you using pytorch xla or tf tpu?",
    "1157767": "tf tpu, similar as the public kernel:\nhttps://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train",
    "1157774": "i observed that while using tpu pod,the utilization rate of TPU iseven  much lower than tpu v3-8\n@herbison",
    "1159976": "mobassir Do you mean the utilization rate is less than it was before (like this post seems to imply), or  just that you see a lower utilization rate between regular TPUs and TPU pods (I'll pass this to TPU experts for a better answer, but my initial reaction would be that TPU pods probably have more processing power and therefore might be able to handle similar throughput with less total utilization).",
    "1159985": "dear @herbison \ni observed lower utilization rate between regular TPUs and TPU pods in tf tpu,i will test in torch tpu soon and will let you know what i get,thank you",
    "1164371": "Thank you for your help sir. what did you observe from torch tpu? my TPUs recovered with no reason and can work normally now",
    "1164410": "southsakura i have lost my all tpu hours few days ago,i will be able to use tpu again from tomorrow for next week :D (will check then) thank you\ni observed in torch tpu if i go for long train then most of the time after 8-10 epochs for single fold i get OOM,,,OOM comes after 8-10 epochs and i don't know how to handle this issue,i lost a lot of tpu hour that way,,,,get OOM within 2 hours and then commit gets stuck and eats 9 hours TPU quota for single experiment run(that too ends with an error), @herbison and team working on fixing this issue i guess",
    "1164990": "Yes, my colleagues are digging into the OOM causing overrun issue.\n\nTPU utilization is independent of that though.",
    "1173636": "TF 2.4 introduces a new parameter in tf.keras.Model.compile called steps_per_execution. [Docs here](https://www.tensorflow.org/api_docs/python/tf/keras/Model#compile).\n\nIt send multiple batches to the TPU for processing before returning control to Keras. Not only does it shave TPU<>Keras interop overheads, but if multiple batches are processed on the TPU at once, the XLA compiler can unroll the loop and optimize matrix filling. You can do the same by had using a custom training loop if you are not afraid of writing a fair bit of code as I explained a while back in [TPU: extreme optimizations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443).\n\nThis setting is even more important, from a performance point of view on TPU pods. Stay tuned for TF 2.4.\n\nFinal note: 22% MXU utilization is not necessarily horrible. I am not 100% sure the new parameter will push it higher. Sometimes, the only thing you can do to push MXU utilization higher is to look at individual matrix paddings as reported by [capture_tpu_profile](https://cloud.google.com/tpu/docs/cloud-tpu-tools) and then adjust layer sizes to make it better. Not something you can do on Kaggle or when using pre-trained models."
  },
  "source": "meta"
}