{
  "id": 142748,
  "title": "Best Practices / Troubleshooting TPU Memory Issues",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/142748",
  "author_name": "Jerry Qu",
  "post_date": "2020-04-12T04:08:01.192000",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This competition is my first time using TPUs, and I assume it may be the same for others as well.</p>\n\n<p>I've been running into TPU memory issues, making this GCP resource insanely useful: <a href=\"https://cloud.google.com/tpu/docs/troubleshooting#memory-usage\">https://cloud.google.com/tpu/docs/troubleshooting#memory-usage</a></p>\n\n<p>Some key insights I've used:</p>\n\n<ul>\n<li>Adam uses 8 extra bytes per weight (4GB for me)</li>\n<li><p>Adafactor uses no extra memory</p></li>\n<li><p>tf2.1 supports Mixed Precision (float16 instead of float32)\n<a href=\"https://www.tensorflow.org/guide/keras/mixed_precision\">https://www.tensorflow.org/guide/keras/mixed_precision</a> (memory &amp; speed improvements)</p></li>\n<li><p>The total batch size should be a multiple of 64 (8 per TPU core), and feature dimensions should be a multiple of 128</p></li>\n</ul>\n\n<p>Hope this helps! (&amp; let me know if anything here is incorrect)</p>",
  "messages": [
    {
      "id": 804864,
      "postDate": "2020-04-12T04:08:01.193Z",
      "content": "<p>This competition is my first time using TPUs, and I assume it may be the same for others as well.</p>\n\n<p>I've been running into TPU memory issues, making this GCP resource insanely useful: <a href=\"https://cloud.google.com/tpu/docs/troubleshooting#memory-usage\">https://cloud.google.com/tpu/docs/troubleshooting#memory-usage</a></p>\n\n<p>Some key insights I've used:</p>\n\n<ul>\n<li>Adam uses 8 extra bytes per weight (4GB for me)</li>\n<li><p>Adafactor uses no extra memory</p></li>\n<li><p>tf2.1 supports Mixed Precision (float16 instead of float32)\n<a href=\"https://www.tensorflow.org/guide/keras/mixed_precision\">https://www.tensorflow.org/guide/keras/mixed_precision</a> (memory &amp; speed improvements)</p></li>\n<li><p>The total batch size should be a multiple of 64 (8 per TPU core), and feature dimensions should be a multiple of 128</p></li>\n</ul>\n\n<p>Hope this helps! (&amp; let me know if anything here is incorrect)</p>",
      "rawMarkdown": "This competition is my first time using TPUs, and I assume it may be the same for others as well.\n\nI've been running into TPU memory issues, making this GCP resource insanely useful: https://cloud.google.com/tpu/docs/troubleshooting#memory-usage\n\nSome key insights I've used:\n\n- Adam uses 8 extra bytes per weight (4GB for me)\n- Adafactor uses no extra memory\n\n- tf2.1 supports Mixed Precision (float16 instead of float32)\nhttps://www.tensorflow.org/guide/keras/mixed_precision (memory &amp; speed improvements)\n\n- The total batch size should be a multiple of 64 (8 per TPU core), and feature dimensions should be a multiple of 128\n\nHope this helps! (&amp; let me know if anything here is incorrect)",
      "votes": 15
    },
    {
      "id": 805436,
      "postDate": "2020-04-12T17:18:38Z",
      "content": "<p>Thank you very much <a href=\"/jerryqu\">@jerryqu</a> , it is really infomative. </p>",
      "rawMarkdown": "Thank you very much @jerryqu , it is really infomative. ",
      "votes": 1
    },
    {
      "id": 805317,
      "postDate": "2020-04-12T15:10:16.350Z",
      "content": "<p>NOTE: If you use Mixed Precision, BE SURE to read the TensorFlow guide. You have to manually set certain values to float32 to insure model quality. (ie. final activation should be float32)</p>",
      "rawMarkdown": "NOTE: If you use Mixed Precision, BE SURE to read the TensorFlow guide. You have to manually set certain values to float32 to insure model quality. (ie. final activation should be float32)",
      "votes": 1
    },
    {
      "id": 812750,
      "postDate": "2020-04-19T03:55:01.783Z",
      "content": "<p>Does anyone know how to print TPU memory usage in a kernel? Thanks.</p>",
      "rawMarkdown": "Does anyone know how to print TPU memory usage in a kernel? Thanks."
    },
    {
      "id": 806704,
      "postDate": "2020-04-14T01:12:54.447Z",
      "content": "<p>Thanks, it helps a lot!\nI have a few questions,\n1. Have you actually experienced the benefit of mixed precision on TPU as well?\n2. Why batch size should be multiple of 64, not just 8?</p>",
      "rawMarkdown": "Thanks, it helps a lot!\nI have a few questions,\n1. Have you actually experienced the benefit of mixed precision on TPU as well?\n2. Why batch size should be multiple of 64, not just 8?",
      "replies": [
        {
          "id": 807324,
          "postDate": "2020-04-14T15:14:48.110Z",
          "content": "<ol>\n<li><p>I haven't used Mixed Precision w/ TensorFlow, but I have a working version with PyTorch-XLA. Torch-XLA was running into OOM issues way more than TensorFlow for some reason.</p></li>\n<li><p>I believe it has to do with how Tensors are padded when fed into TPU cores. They mention on the website that 'TPU rounds up the sizes of tensors stored in memory to perform computations more efficiently', which is why you should use the dimensions they recommend. (So, if you use a batch size of 8, I think they'll just pad it up to 64, which is wasting a lot of computation)</p></li>\n</ol>\n\n<p>As to why they must pad the tensors, I'm not so sure. But it may have to do with how computations are actually done on TPUs, with systolic arrays: <a href=\"https://www.youtube.com/watch?v=JC84GCU7zqA\">https://www.youtube.com/watch?v=JC84GCU7zqA</a></p>",
          "rawMarkdown": "1. I haven't used Mixed Precision w/ TensorFlow, but I have a working version with PyTorch-XLA. Torch-XLA was running into OOM issues way more than TensorFlow for some reason.\n\n2. I believe it has to do with how Tensors are padded when fed into TPU cores. They mention on the website that 'TPU rounds up the sizes of tensors stored in memory to perform computations more efficiently', which is why you should use the dimensions they recommend. (So, if you use a batch size of 8, I think they'll just pad it up to 64, which is wasting a lot of computation)\n\nAs to why they must pad the tensors, I'm not so sure. But it may have to do with how computations are actually done on TPUs, with systolic arrays: https://www.youtube.com/watch?v=JC84GCU7zqA"
        },
        {
          "id": 807363,
          "postDate": "2020-04-14T15:50:08.610Z",
          "content": "<p><a href=\"/jerryqu\">@jerryqu</a> Oh, you're using pytorch, I got it. And thank you for the answer! Eventually stopping using Adam was most effective:)</p>",
          "rawMarkdown": "@jerryqu Oh, you're using pytorch, I got it. And thank you for the answer! Eventually stopping using Adam was most effective:)",
          "votes": 1
        }
      ]
    },
    {
      "id": 913372,
      "postDate": "2020-07-03T07:20:17.073Z",
      "content": "<p>When I used TPU Pytorch XLA version, It throws out of memory issue on 2nd epoch. I suspect the cache cannot be cleared timely. Do u know how to tackle this issue? </p>",
      "rawMarkdown": "When I used TPU Pytorch XLA version, It throws out of memory issue on 2nd epoch. I suspect the cache cannot be cleared timely. Do u know how to tackle this issue? ",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 805436,
      "author_name": "Mohammed Deeb",
      "author_url": "",
      "post_date": "2020-04-12T17:18:38",
      "content": "<p>Thank you very much <a href=\"/jerryqu\">@jerryqu</a> , it is really infomative. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 805317,
      "author_name": "Jerry Qu",
      "author_url": "",
      "post_date": "2020-04-12T15:10:16.350000",
      "content": "<p>NOTE: If you use Mixed Precision, BE SURE to read the TensorFlow guide. You have to manually set certain values to float32 to insure model quality. (ie. final activation should be float32)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 812750,
      "author_name": "DHZM",
      "author_url": "",
      "post_date": "2020-04-19T03:55:01.783000",
      "content": "<p>Does anyone know how to print TPU memory usage in a kernel? Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 806704,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-04-14T01:12:54.447000",
      "content": "<p>Thanks, it helps a lot!\nI have a few questions,\n1. Have you actually experienced the benefit of mixed precision on TPU as well?\n2. Why batch size should be multiple of 64, not just 8?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 807324,
          "author_name": "Jerry Qu",
          "author_url": "",
          "post_date": "2020-04-14T15:14:48.110000",
          "content": "<ol>\n<li><p>I haven't used Mixed Precision w/ TensorFlow, but I have a working version with PyTorch-XLA. Torch-XLA was running into OOM issues way more than TensorFlow for some reason.</p></li>\n<li><p>I believe it has to do with how Tensors are padded when fed into TPU cores. They mention on the website that 'TPU rounds up the sizes of tensors stored in memory to perform computations more efficiently', which is why you should use the dimensions they recommend. (So, if you use a batch size of 8, I think they'll just pad it up to 64, which is wasting a lot of computation)</p></li>\n</ol>\n\n<p>As to why they must pad the tensors, I'm not so sure. But it may have to do with how computations are actually done on TPUs, with systolic arrays: <a href=\"https://www.youtube.com/watch?v=JC84GCU7zqA\">https://www.youtube.com/watch?v=JC84GCU7zqA</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 807363,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-04-14T15:50:08.610000",
          "content": "<p><a href=\"/jerryqu\">@jerryqu</a> Oh, you're using pytorch, I got it. And thank you for the answer! Eventually stopping using Adam was most effective:)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 913372,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-03T07:20:17.073000",
      "content": "<p>When I used TPU Pytorch XLA version, It throws out of memory issue on 2nd epoch. I suspect the cache cannot be cleared timely. Do u know how to tackle this issue? </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "804864": "This competition is my first time using TPUs, and I assume it may be the same for others as well.\n\nI've been running into TPU memory issues, making this GCP resource insanely useful: https://cloud.google.com/tpu/docs/troubleshooting#memory-usage\n\nSome key insights I've used:\n\n- Adam uses 8 extra bytes per weight (4GB for me)\n- Adafactor uses no extra memory\n\n- tf2.1 supports Mixed Precision (float16 instead of float32)\nhttps://www.tensorflow.org/guide/keras/mixed_precision (memory &amp; speed improvements)\n\n- The total batch size should be a multiple of 64 (8 per TPU core), and feature dimensions should be a multiple of 128\n\nHope this helps! (&amp; let me know if anything here is incorrect)",
    "805436": "Thank you very much @jerryqu , it is really infomative. ",
    "805317": "NOTE: If you use Mixed Precision, BE SURE to read the TensorFlow guide. You have to manually set certain values to float32 to insure model quality. (ie. final activation should be float32)",
    "812750": "Does anyone know how to print TPU memory usage in a kernel? Thanks.",
    "806704": "Thanks, it helps a lot!\nI have a few questions,\n1. Have you actually experienced the benefit of mixed precision on TPU as well?\n2. Why batch size should be multiple of 64, not just 8?",
    "913372": "When I used TPU Pytorch XLA version, It throws out of memory issue on 2nd epoch. I suspect the cache cannot be cleared timely. Do u know how to tackle this issue? "
  }
}