{
  "id": 437673,
  "title": "What is a Tensor Processing Unit (TPU)?",
  "url": "/competitions/predict-ai-model-runtime/discussion/437673",
  "author_name": "Mangpo Phothilimthana",
  "post_date": "2023-09-07T18:24:27.210000",
  "votes": 24,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Since the dataset contains runtime measured on TPUs, some knowledge about TPUs might be useful. This page has a good explanation about TPUs: <a href=\"https://jax.readthedocs.io/en/latest/pallas/tpu.html#what-is-a-tpu\" target=\"_blank\">https://jax.readthedocs.io/en/latest/pallas/tpu.html#what-is-a-tpu</a>. </p>\n<p>One important thing I would like to point out is that a TPU has 2D registers of size 128 x 8. Essentially, the most minor dimension (as specified by the layout config) will map to the 128-size width of the 2D registers, and the second most minor dimension will map to the 8-length of the 2D registers. Consider an output tensor shape [2, 4, 256] with the layout {0, 1, 2} of a node in a graph (recall that the layout specifies the order of minor-to-major tensor dimensions). In this case, dim of size 2 will go to the most minor dimension, which maps to the 128-size width of the 2D registers. This will end up with lots of padding! In contrast, a layout {2, 1, 0}, which maps the dim of size 256 to the 128-size width of the 2D registers is likely a better choice for this particular node in a graph. When the dim size is greater than the size of 2D registers, the program essentially iterates over the entire tensor multiple times.</p>",
  "messages": [
    {
      "id": 2428250,
      "postDate": "2023-09-07T18:24:27.210Z",
      "content": "<p>Since the dataset contains runtime measured on TPUs, some knowledge about TPUs might be useful. This page has a good explanation about TPUs: <a href=\"https://jax.readthedocs.io/en/latest/pallas/tpu.html#what-is-a-tpu\" target=\"_blank\">https://jax.readthedocs.io/en/latest/pallas/tpu.html#what-is-a-tpu</a>. </p>\n<p>One important thing I would like to point out is that a TPU has 2D registers of size 128 x 8. Essentially, the most minor dimension (as specified by the layout config) will map to the 128-size width of the 2D registers, and the second most minor dimension will map to the 8-length of the 2D registers. Consider an output tensor shape [2, 4, 256] with the layout {0, 1, 2} of a node in a graph (recall that the layout specifies the order of minor-to-major tensor dimensions). In this case, dim of size 2 will go to the most minor dimension, which maps to the 128-size width of the 2D registers. This will end up with lots of padding! In contrast, a layout {2, 1, 0}, which maps the dim of size 256 to the 128-size width of the 2D registers is likely a better choice for this particular node in a graph. When the dim size is greater than the size of 2D registers, the program essentially iterates over the entire tensor multiple times.</p>",
      "rawMarkdown": "Since the dataset contains runtime measured on TPUs, some knowledge about TPUs might be useful. This page has a good explanation about TPUs: https://jax.readthedocs.io/en/latest/pallas/tpu.html#what-is-a-tpu. \n\nOne important thing I would like to point out is that a TPU has 2D registers of size 128 x 8. Essentially, the most minor dimension (as specified by the layout config) will map to the 128-size width of the 2D registers, and the second most minor dimension will map to the 8-length of the 2D registers. Consider an output tensor shape [2, 4, 256] with the layout {0, 1, 2} of a node in a graph (recall that the layout specifies the order of minor-to-major tensor dimensions). In this case, dim of size 2 will go to the most minor dimension, which maps to the 128-size width of the 2D registers. This will end up with lots of padding! In contrast, a layout {2, 1, 0}, which maps the dim of size 256 to the 128-size width of the 2D registers is likely a better choice for this particular node in a graph. When the dim size is greater than the size of 2D registers, the program essentially iterates over the entire tensor multiple times.",
      "votes": 23
    },
    {
      "id": 2466531,
      "postDate": "2023-10-03T23:27:27.323Z",
      "content": "<p>To optimize TPU performance, choosing the right tensor layout is crucial, as mapping larger dimensions to the 2D registers' size can minimize padding and improve efficiency.</p>",
      "rawMarkdown": "To optimize TPU performance, choosing the right tensor layout is crucial, as mapping larger dimensions to the 2D registers' size can minimize padding and improve efficiency.",
      "votes": 1
    },
    {
      "id": 2442480,
      "postDate": "2023-09-17T04:23:20.840Z",
      "content": "<p><a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> Thank you for your sharing. \"Consider an output tensor shape [2, 4, 256] with the layout {0, 1, 2} of a node in a graph(recall that the layout specifies the order of minor-to-major tensor dimensions). In this case, dim of size 2 will go to the most minor dimension, which maps to the 128-size width of the 2D registers. \"  dim of size 4 will map to the 8-size width? What about dim of size 256?</p>",
      "rawMarkdown": "@mangpophothilimthana Thank you for your sharing. \"Consider an output tensor shape [2, 4, 256] with the layout {0, 1, 2} of a node in a graph(recall that the layout specifies the order of minor-to-major tensor dimensions). In this case, dim of size 2 will go to the most minor dimension, which maps to the 128-size width of the 2D registers. \"  dim of size 4 will map to the 8-size width? What about dim of size 256?",
      "votes": 2,
      "replies": [
        {
          "id": 2447400,
          "postDate": "2023-09-20T04:00:34.160Z",
          "content": "<p>Dim of size 256 will be an outer loop, where the program iterates through. Essentially, the entire tensor lives in large HMB memory, and in each iteration a tile of the tensor will be copied into a faster scratchpad memory, and in a sub iteration a smaller portion of a tile will be copied into the 2D registers.</p>",
          "rawMarkdown": "Dim of size 256 will be an outer loop, where the program iterates through. Essentially, the entire tensor lives in large HMB memory, and in each iteration a tile of the tensor will be copied into a faster scratchpad memory, and in a sub iteration a smaller portion of a tile will be copied into the 2D registers.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2428643,
      "postDate": "2023-09-08T04:26:32.543Z",
      "content": "<p>Thanks for pointing this out!<br>\nNice competition by the way)</p>",
      "rawMarkdown": "Thanks for pointing this out!\nNice competition by the way)"
    },
    {
      "id": 2456474,
      "postDate": "2023-09-26T08:19:17.473Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2463043,
      "postDate": "2023-09-30T21:21:01.593Z",
      "content": "<p>Thanks for the resource! </p>",
      "rawMarkdown": "Thanks for the resource! "
    },
    {
      "id": 2453978,
      "postDate": "2023-09-24T13:10:49.623Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing"
    },
    {
      "id": 2442604,
      "postDate": "2023-09-17T06:23:24.080Z",
      "content": "<p>Thank you for your explanation!</p>",
      "rawMarkdown": "Thank you for your explanation!"
    },
    {
      "id": 2433342,
      "postDate": "2023-09-11T13:57:42.297Z",
      "content": "<p>Thank you for your explanation!</p>",
      "rawMarkdown": "Thank you for your explanation!"
    },
    {
      "id": 2429547,
      "postDate": "2023-09-08T16:49:12.647Z",
      "content": "<p>Thank you for your sharing!</p>",
      "rawMarkdown": "Thank you for your sharing!"
    }
  ],
  "comments": [
    {
      "id": 2466531,
      "author_name": "Tasbirul Alom",
      "author_url": "",
      "post_date": "2023-10-03T23:27:27.323000",
      "content": "<p>To optimize TPU performance, choosing the right tensor layout is crucial, as mapping larger dimensions to the 2D registers' size can minimize padding and improve efficiency.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2442480,
      "author_name": "Bruce",
      "author_url": "",
      "post_date": "2023-09-17T04:23:20.840000",
      "content": "<p><a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> Thank you for your sharing. \"Consider an output tensor shape [2, 4, 256] with the layout {0, 1, 2} of a node in a graph(recall that the layout specifies the order of minor-to-major tensor dimensions). In this case, dim of size 2 will go to the most minor dimension, which maps to the 128-size width of the 2D registers. \"  dim of size 4 will map to the 8-size width? What about dim of size 256?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2447400,
          "author_name": "Mangpo Phothilimthana",
          "author_url": "",
          "post_date": "2023-09-20T04:00:34.160000",
          "content": "<p>Dim of size 256 will be an outer loop, where the program iterates through. Essentially, the entire tensor lives in large HMB memory, and in each iteration a tile of the tensor will be copied into a faster scratchpad memory, and in a sub iteration a smaller portion of a tile will be copied into the 2D registers.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2428643,
      "author_name": "Shekhovtsov Aleksandr",
      "author_url": "",
      "post_date": "2023-09-08T04:26:32.543000",
      "content": "<p>Thanks for pointing this out!<br>\nNice competition by the way)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2456474,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-26T08:19:17.473000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2463043,
      "author_name": "Amirhossein hajavi",
      "author_url": "",
      "post_date": "2023-09-30T21:21:01.593000",
      "content": "<p>Thanks for the resource! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2453978,
      "author_name": "AbdelRahman Sousa",
      "author_url": "",
      "post_date": "2023-09-24T13:10:49.623000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2442604,
      "author_name": "Zhaoyan Pan",
      "author_url": "",
      "post_date": "2023-09-17T06:23:24.080000",
      "content": "<p>Thank you for your explanation!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2433342,
      "author_name": "Jeffrey Lim",
      "author_url": "",
      "post_date": "2023-09-11T13:57:42.297000",
      "content": "<p>Thank you for your explanation!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2429547,
      "author_name": "DwyaneLi",
      "author_url": "",
      "post_date": "2023-09-08T16:49:12.647000",
      "content": "<p>Thank you for your sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2428250": "Since the dataset contains runtime measured on TPUs, some knowledge about TPUs might be useful. This page has a good explanation about TPUs: https://jax.readthedocs.io/en/latest/pallas/tpu.html#what-is-a-tpu. \n\nOne important thing I would like to point out is that a TPU has 2D registers of size 128 x 8. Essentially, the most minor dimension (as specified by the layout config) will map to the 128-size width of the 2D registers, and the second most minor dimension will map to the 8-length of the 2D registers. Consider an output tensor shape [2, 4, 256] with the layout {0, 1, 2} of a node in a graph (recall that the layout specifies the order of minor-to-major tensor dimensions). In this case, dim of size 2 will go to the most minor dimension, which maps to the 128-size width of the 2D registers. This will end up with lots of padding! In contrast, a layout {2, 1, 0}, which maps the dim of size 256 to the 128-size width of the 2D registers is likely a better choice for this particular node in a graph. When the dim size is greater than the size of 2D registers, the program essentially iterates over the entire tensor multiple times.",
    "2466531": "To optimize TPU performance, choosing the right tensor layout is crucial, as mapping larger dimensions to the 2D registers' size can minimize padding and improve efficiency.",
    "2442480": "@mangpophothilimthana Thank you for your sharing. \"Consider an output tensor shape [2, 4, 256] with the layout {0, 1, 2} of a node in a graph(recall that the layout specifies the order of minor-to-major tensor dimensions). In this case, dim of size 2 will go to the most minor dimension, which maps to the 128-size width of the 2D registers. \"  dim of size 4 will map to the 8-size width? What about dim of size 256?",
    "2428643": "Thanks for pointing this out!\nNice competition by the way)",
    "2456474": "",
    "2463043": "Thanks for the resource! ",
    "2453978": "Thank you for sharing",
    "2442604": "Thank you for your explanation!",
    "2433342": "Thank you for your explanation!",
    "2429547": "Thank you for your sharing!"
  }
}