{
  "id": 201704,
  "title": "TPU Usage extremely low",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201704",
  "author_name": "Aditya Baurai",
  "post_date": "2020-12-06T10:30:41.783000",
  "votes": 1,
  "comment_count": 12,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F0e1ec5d29fa4381ab2b11783e1d941d5%2Ftpu.png?generation=1607250599764361&amp;alt=media\" alt=\"\"></p>\n<p>What to do? It's a standard pipeline using TFRecords. </p>",
  "messages": [
    {
      "id": 1104116,
      "postDate": "2020-12-06T16:23:57.003Z",
      "content": "<p>This may not be related to your TF pipeline, TPU usage is alo related to model architecture/size and image size.</p>\n<p>In your case, I think you should worry more about the high idle time.</p>",
      "rawMarkdown": "This may not be related to your TF pipeline, TPU usage is alo related to model architecture/size and image size.\n\nIn your case, I think you should worry more about the high idle time.",
      "votes": 1,
      "replies": [
        {
          "id": 1104242,
          "postDate": "2020-12-06T18:31:27.567Z",
          "content": "<p>I'd assume since one epoch is 34 seconds it must be a small model with minimal augmentation. I don't have much TPU experience, but on my GPU I quickly get to a few min per epoch with larger models and augmentation. Even so, GPU usage is usually on average not so high (indicating I'm not being super efficient).</p>",
          "rawMarkdown": "I'd assume since one epoch is 34 seconds it must be a small model with minimal augmentation. I don't have much TPU experience, but on my GPU I quickly get to a few min per epoch with larger models and augmentation. Even so, GPU usage is usually on average not so high (indicating I'm not being super efficient).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1103827,
      "postDate": "2020-12-06T10:30:41.783Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F0e1ec5d29fa4381ab2b11783e1d941d5%2Ftpu.png?generation=1607250599764361&amp;alt=media\" alt=\"\"></p>\n<p>What to do? It's a standard pipeline using TFRecords. </p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F0e1ec5d29fa4381ab2b11783e1d941d5%2Ftpu.png?generation=1607250599764361&alt=media)\n\nWhat to do? It's a standard pipeline using TFRecords. ",
      "votes": 1
    },
    {
      "id": 1107265,
      "postDate": "2020-12-09T14:42:34.580Z",
      "content": "<p>Two things i think in your case are possible, the batch size is smaller or the model itself is smaller, so the computations are somewhat lower which explains the idle time. Maybe you should switch to GPU for such cases.</p>",
      "rawMarkdown": "Two things i think in your case are possible, the batch size is smaller or the model itself is smaller, so the computations are somewhat lower which explains the idle time. Maybe you should switch to GPU for such cases.",
      "replies": [
        {
          "id": 1107311,
          "postDate": "2020-12-09T15:26:54.637Z",
          "content": "<p>I used Resnet50v2. I think it's medium sized… </p>",
          "rawMarkdown": "I used Resnet50v2. I think it's medium sized... ",
          "votes": 1
        },
        {
          "id": 1107899,
          "postDate": "2020-12-10T04:05:55.350Z",
          "content": "<p>Maybe not? idk, You should compare EfficientNets parameters to get exact figures.</p>",
          "rawMarkdown": "Maybe not? idk, You should compare EfficientNets parameters to get exact figures.",
          "votes": 1
        },
        {
          "id": 1108179,
          "postDate": "2020-12-10T11:17:57.047Z",
          "content": "<p>ok, will look into it. Thanks for the assist!</p>",
          "rawMarkdown": "ok, will look into it. Thanks for the assist!"
        }
      ]
    },
    {
      "id": 1106463,
      "postDate": "2020-12-08T21:21:51.867Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1106772,
          "postDate": "2020-12-09T05:39:26.050Z",
          "content": "<p>No no. I am not using image augmentations. I augment the images beforehand in custom TFRec files and directly use those. </p>",
          "rawMarkdown": "No no. I am not using image augmentations. I augment the images beforehand in custom TFRec files and directly use those. "
        },
        {
          "id": 1107232,
          "postDate": "2020-12-09T14:01:11.483Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1107310,
          "postDate": "2020-12-09T15:26:23.773Z",
          "content": "<p>I'll look into it. Thanks</p>",
          "rawMarkdown": "I'll look into it. Thanks"
        },
        {
          "id": 1107386,
          "postDate": "2020-12-09T16:46:45.463Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1108177,
          "postDate": "2020-12-10T11:17:31.197Z",
          "content": "<p>roger that!</p>",
          "rawMarkdown": "roger that!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1104116,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-12-06T16:23:57.003000",
      "content": "<p>This may not be related to your TF pipeline, TPU usage is alo related to model architecture/size and image size.</p>\n<p>In your case, I think you should worry more about the high idle time.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1104242,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-06T18:31:27.567000",
          "content": "<p>I'd assume since one epoch is 34 seconds it must be a small model with minimal augmentation. I don't have much TPU experience, but on my GPU I quickly get to a few min per epoch with larger models and augmentation. Even so, GPU usage is usually on average not so high (indicating I'm not being super efficient).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1107265,
      "author_name": "Ankush kuwar",
      "author_url": "",
      "post_date": "2020-12-09T14:42:34.580000",
      "content": "<p>Two things i think in your case are possible, the batch size is smaller or the model itself is smaller, so the computations are somewhat lower which explains the idle time. Maybe you should switch to GPU for such cases.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1107311,
          "author_name": "Aditya Baurai",
          "author_url": "",
          "post_date": "2020-12-09T15:26:54.637000",
          "content": "<p>I used Resnet50v2. I think it's medium sized… </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107899,
          "author_name": "Ankush kuwar",
          "author_url": "",
          "post_date": "2020-12-10T04:05:55.350000",
          "content": "<p>Maybe not? idk, You should compare EfficientNets parameters to get exact figures.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1108179,
          "author_name": "Aditya Baurai",
          "author_url": "",
          "post_date": "2020-12-10T11:17:57.047000",
          "content": "<p>ok, will look into it. Thanks for the assist!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1106463,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-08T21:21:51.867000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1106772,
          "author_name": "Aditya Baurai",
          "author_url": "",
          "post_date": "2020-12-09T05:39:26.050000",
          "content": "<p>No no. I am not using image augmentations. I augment the images beforehand in custom TFRec files and directly use those. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107232,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-09T14:01:11.483000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107310,
          "author_name": "Aditya Baurai",
          "author_url": "",
          "post_date": "2020-12-09T15:26:23.773000",
          "content": "<p>I'll look into it. Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107386,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-09T16:46:45.463000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1108177,
          "author_name": "Aditya Baurai",
          "author_url": "",
          "post_date": "2020-12-10T11:17:31.197000",
          "content": "<p>roger that!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1104116": "This may not be related to your TF pipeline, TPU usage is alo related to model architecture/size and image size.\n\nIn your case, I think you should worry more about the high idle time.",
    "1103827": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F0e1ec5d29fa4381ab2b11783e1d941d5%2Ftpu.png?generation=1607250599764361&alt=media)\n\nWhat to do? It's a standard pipeline using TFRecords. ",
    "1107265": "Two things i think in your case are possible, the batch size is smaller or the model itself is smaller, so the computations are somewhat lower which explains the idle time. Maybe you should switch to GPU for such cases.",
    "1106463": ""
  }
}