{
  "id": 164633,
  "title": "Problem when using TPU, MXU. ",
  "url": "/competitions/alaska2-image-steganalysis/discussion/164633",
  "author_name": "",
  "post_date": "2020-07-07T02:00:15.963653900Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'm trying to learn TPU but face with a problem. \nThe notebook monitor always says MXU 2.00% or lower as the figure shows. I've tried many methods to improve it but none of them works. Anyone know how to improve the MXU? What's your MXU when using TPU?\nAppriciate it if you can help.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2911977%2F400973179e455805d8249cd1ae9f880c%2F885123393.jpg?generation=1594087637768642&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "918109",
      "postDate": "07/07/2020 02:00:15",
      "content": "<p>I'm trying to learn TPU but face with a problem. \nThe notebook monitor always says MXU 2.00% or lower as the figure shows. I've tried many methods to improve it but none of them works. Anyone know how to improve the MXU? What's your MXU when using TPU?\nAppriciate it if you can help.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2911977%2F400973179e455805d8249cd1ae9f880c%2F885123393.jpg?generation=1594087637768642&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'm trying to learn TPU but face with a problem. \nThe notebook monitor always says MXU 2.00% or lower as the figure shows. I've tried many methods to improve it but none of them works. Anyone know how to improve the MXU? What's your MXU when using TPU?\nAppriciate it if you can help.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2911977%2F400973179e455805d8249cd1ae9f880c%2F885123393.jpg?generation=1594087637768642&amp;alt=media)",
      "votes": null
    },
    {
      "id": "918347",
      "postDate": "07/07/2020 07:41:42",
      "content": "<p>This means that the TPU is processing faster than the CPU is feeding the TPU with data. Try a larger batch size or more workers in the dataloader.</p>",
      "rawMarkdown": "This means that the TPU is processing faster than the CPU is feeding the TPU with data. Try a larger batch size or more workers in the dataloader.",
      "votes": null
    },
    {
      "id": "918494",
      "postDate": "07/07/2020 09:26:21",
      "content": "<p>Here we face a limit of available RAM to save pre-cached batches in memory and CPU bottleneck to prepare those fast enough. \nIt’s a huge imbalance in processing power betweeen TPU and CPU. And this is an issue. I hope kaggle team will address it in future.</p>",
      "rawMarkdown": "Here we face a limit of available RAM to save pre-cached batches in memory and CPU bottleneck to prepare those fast enough. \nIt’s a huge imbalance in processing power betweeen TPU and CPU. And this is an issue. I hope kaggle team will address it in future.",
      "votes": null
    },
    {
      "id": "918562",
      "postDate": "07/07/2020 10:15:30",
      "content": "<p>Hi,</p>\n\n<p>If we use a custom training loop then I think we can run training faster. I saw one such kernel in TPU Flower Classification Challenge. Based on that I created <a href=\"https://www.kaggle.com/urvishp80/tf-tpu-custom-training\">this</a> kernel but I could not implement the prediction loop properly. However, someone can do that by simply using <code>preds = model(test_data, training=False)</code>.</p>\n\n<p>You can modify the kernel to get max MXU and less idle time. </p>",
      "rawMarkdown": "Hi,\n\nIf we use a custom training loop then I think we can run training faster. I saw one such kernel in TPU Flower Classification Challenge. Based on that I created [this](https://www.kaggle.com/urvishp80/tf-tpu-custom-training) kernel but I could not implement the prediction loop properly. However, someone can do that by simply using `preds = model(test_data, training=False)`.\n\nYou can modify the kernel to get max MXU and less idle time.",
      "votes": null
    },
    {
      "id": "919075",
      "postDate": "07/07/2020 17:37:44",
      "content": "<p>I am using pytorch Imao. Upvoted and will try it. Thank you.</p>",
      "rawMarkdown": "I am using pytorch Imao. Upvoted and will try it. Thank you.",
      "votes": null
    },
    {
      "id": "919076",
      "postDate": "07/07/2020 17:38:21",
      "content": "<p>I totally agree with you.</p>",
      "rawMarkdown": "I totally agree with you.",
      "votes": null
    },
    {
      "id": "919226",
      "postDate": "07/07/2020 19:06:12",
      "content": "<p>It is also a huge problem for GPU kernels often times.</p>",
      "rawMarkdown": "It is also a huge problem for GPU kernels often times.",
      "votes": null
    },
    {
      "id": "919680",
      "postDate": "07/08/2020 03:58:02",
      "content": "<p>I will try to publish a new version of the kernel maybe that will help. Thanks for the upvote. </p>",
      "rawMarkdown": "I will try to publish a new version of the kernel maybe that will help. Thanks for the upvote.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 918347,
      "author_name": "njelicic",
      "author_url": "",
      "post_date": "07/07/2020 07:41:42",
      "content": "<p>This means that the TPU is processing faster than the CPU is feeding the TPU with data. Try a larger batch size or more workers in the dataloader.</p>",
      "votes": null,
      "replies": [
        {
          "id": 918494,
          "author_name": "bloodaxe",
          "author_url": "",
          "post_date": "07/07/2020 09:26:21",
          "content": "<p>Here we face a limit of available RAM to save pre-cached batches in memory and CPU bottleneck to prepare those fast enough. \nIt’s a huge imbalance in processing power betweeen TPU and CPU. And this is an issue. I hope kaggle team will address it in future.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 919076,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "07/07/2020 17:38:21",
          "content": "<p>I totally agree with you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 919226,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/07/2020 19:06:12",
          "content": "<p>It is also a huge problem for GPU kernels often times.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 918562,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "07/07/2020 10:15:30",
      "content": "<p>Hi,</p>\n\n<p>If we use a custom training loop then I think we can run training faster. I saw one such kernel in TPU Flower Classification Challenge. Based on that I created <a href=\"https://www.kaggle.com/urvishp80/tf-tpu-custom-training\">this</a> kernel but I could not implement the prediction loop properly. However, someone can do that by simply using <code>preds = model(test_data, training=False)</code>.</p>\n\n<p>You can modify the kernel to get max MXU and less idle time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 919075,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "07/07/2020 17:37:44",
          "content": "<p>I am using pytorch Imao. Upvoted and will try it. Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 919680,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "07/08/2020 03:58:02",
          "content": "<p>I will try to publish a new version of the kernel maybe that will help. Thanks for the upvote. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "918109": "I'm trying to learn TPU but face with a problem. \nThe notebook monitor always says MXU 2.00% or lower as the figure shows. I've tried many methods to improve it but none of them works. Anyone know how to improve the MXU? What's your MXU when using TPU?\nAppriciate it if you can help.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2911977%2F400973179e455805d8249cd1ae9f880c%2F885123393.jpg?generation=1594087637768642&amp;alt=media)",
    "918347": "This means that the TPU is processing faster than the CPU is feeding the TPU with data. Try a larger batch size or more workers in the dataloader.",
    "918494": "Here we face a limit of available RAM to save pre-cached batches in memory and CPU bottleneck to prepare those fast enough. \nIt’s a huge imbalance in processing power betweeen TPU and CPU. And this is an issue. I hope kaggle team will address it in future.",
    "918562": "Hi,\n\nIf we use a custom training loop then I think we can run training faster. I saw one such kernel in TPU Flower Classification Challenge. Based on that I created [this](https://www.kaggle.com/urvishp80/tf-tpu-custom-training) kernel but I could not implement the prediction loop properly. However, someone can do that by simply using `preds = model(test_data, training=False)`.\n\nYou can modify the kernel to get max MXU and less idle time.",
    "919075": "I am using pytorch Imao. Upvoted and will try it. Thank you.",
    "919076": "I totally agree with you.",
    "919226": "It is also a huge problem for GPU kernels often times.",
    "919680": "I will try to publish a new version of the kernel maybe that will help. Thanks for the upvote."
  },
  "source": "meta"
}