{
  "id": 170106,
  "title": "TPU utilizing 0%",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/170106",
  "author_name": "",
  "post_date": "2020-07-26T13:03:37.649957800Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>My model seems to be executing on TPU, however MXU shows 0% and idle time either -- or 30%??</p>\n\n<p>Can anyone help?</p>",
  "messages": [
    {
      "id": "946238",
      "postDate": "07/26/2020 13:03:37",
      "content": "<p>My model seems to be executing on TPU, however MXU shows 0% and idle time either -- or 30%??</p>\n\n<p>Can anyone help?</p>",
      "rawMarkdown": "My model seems to be executing on TPU, however MXU shows 0% and idle time either -- or 30%??\n\nCan anyone help?",
      "votes": null
    },
    {
      "id": "946798",
      "postDate": "07/26/2020 21:07:19",
      "content": "<p>My models also often don't register at all on the TPU, although clearly the TPU is being used (epochs with 30,000 images in 2 minutes). I think it takes a lot to use the TPU enough to show up on the Meter. A 456 x 456 image with EfficientNetB5 and Batchsize 32x8 = 256 get me:</p>\n\n<p>MXU 8%\nIdle Time 11%</p>\n\n<p>But I have run many models that always have 0%,0% on the meter.</p>\n\n<p>There are various guidelines to maximinize the TPU. Some have to do with the size of certain parameters being multiples of certain numbers. The most important is to preprocess your data so that you are not CPU bound.</p>\n\n<p>-Rich</p>",
      "rawMarkdown": "My models also often don't register at all on the TPU, although clearly the TPU is being used (epochs with 30,000 images in 2 minutes). I think it takes a lot to use the TPU enough to show up on the Meter. A 456 x 456 image with EfficientNetB5 and Batchsize 32x8 = 256 get me:\n\nMXU 8%\nIdle Time 11%\n\nBut I have run many models that always have 0%,0% on the meter.\n\nThere are various guidelines to maximinize the TPU. Some have to do with the size of certain parameters being multiples of certain numbers. The most important is to preprocess your data so that you are not CPU bound.\n\n-Rich",
      "votes": null
    },
    {
      "id": "947332",
      "postDate": "07/27/2020 07:31:53",
      "content": "<p>Yes, I have noticed that too. Images are processed lightning fast during training. It's clear that TPU is working in the background! I suppose your assumption of rather late display is correct.</p>",
      "rawMarkdown": "Yes, I have noticed that too. Images are processed lightning fast during training. It's clear that TPU is working in the background! I suppose your assumption of rather late display is correct.",
      "votes": null
    },
    {
      "id": "949209",
      "postDate": "07/28/2020 13:20:05",
      "content": "<p>We only sample the metrics every 10s or so, so it's very possible that we never observe real usage during those samples.</p>\n\n<p>Though I think this also means the TPU isn't working very hard/long to process your data, maybe your pipeline spends more time loading/feeding data then processing it.</p>",
      "rawMarkdown": "We only sample the metrics every 10s or so, so it's very possible that we never observe real usage during those samples.\n\nThough I think this also means the TPU isn't working very hard/long to process your data, maybe your pipeline spends more time loading/feeding data then processing it.",
      "votes": null
    },
    {
      "id": "950156",
      "postDate": "07/29/2020 08:03:59",
      "content": "<p>This might be! 23k images got processed in a few minutes, but there wasn't any change in MXU meter. Maybe it is crunching(that explains the speed) the images...</p>",
      "rawMarkdown": "This might be! 23k images got processed in a few minutes, but there wasn't any change in MXU meter. Maybe it is crunching(that explains the speed) the images...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 946798,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "07/26/2020 21:07:19",
      "content": "<p>My models also often don't register at all on the TPU, although clearly the TPU is being used (epochs with 30,000 images in 2 minutes). I think it takes a lot to use the TPU enough to show up on the Meter. A 456 x 456 image with EfficientNetB5 and Batchsize 32x8 = 256 get me:</p>\n\n<p>MXU 8%\nIdle Time 11%</p>\n\n<p>But I have run many models that always have 0%,0% on the meter.</p>\n\n<p>There are various guidelines to maximinize the TPU. Some have to do with the size of certain parameters being multiples of certain numbers. The most important is to preprocess your data so that you are not CPU bound.</p>\n\n<p>-Rich</p>",
      "votes": null,
      "replies": [
        {
          "id": 947332,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "07/27/2020 07:31:53",
          "content": "<p>Yes, I have noticed that too. Images are processed lightning fast during training. It's clear that TPU is working in the background! I suppose your assumption of rather late display is correct.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 949209,
      "author_name": "herbison",
      "author_url": "",
      "post_date": "07/28/2020 13:20:05",
      "content": "<p>We only sample the metrics every 10s or so, so it's very possible that we never observe real usage during those samples.</p>\n\n<p>Though I think this also means the TPU isn't working very hard/long to process your data, maybe your pipeline spends more time loading/feeding data then processing it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 950156,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "07/29/2020 08:03:59",
          "content": "<p>This might be! 23k images got processed in a few minutes, but there wasn't any change in MXU meter. Maybe it is crunching(that explains the speed) the images...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "946238": "My model seems to be executing on TPU, however MXU shows 0% and idle time either -- or 30%??\n\nCan anyone help?",
    "946798": "My models also often don't register at all on the TPU, although clearly the TPU is being used (epochs with 30,000 images in 2 minutes). I think it takes a lot to use the TPU enough to show up on the Meter. A 456 x 456 image with EfficientNetB5 and Batchsize 32x8 = 256 get me:\n\nMXU 8%\nIdle Time 11%\n\nBut I have run many models that always have 0%,0% on the meter.\n\nThere are various guidelines to maximinize the TPU. Some have to do with the size of certain parameters being multiples of certain numbers. The most important is to preprocess your data so that you are not CPU bound.\n\n-Rich",
    "947332": "Yes, I have noticed that too. Images are processed lightning fast during training. It's clear that TPU is working in the background! I suppose your assumption of rather late display is correct.",
    "949209": "We only sample the metrics every 10s or so, so it's very possible that we never observe real usage during those samples.\n\nThough I think this also means the TPU isn't working very hard/long to process your data, maybe your pipeline spends more time loading/feeding data then processing it.",
    "950156": "This might be! 23k images got processed in a few minutes, but there wasn't any change in MXU meter. Maybe it is crunching(that explains the speed) the images..."
  },
  "source": "meta"
}