{
  "id": 442258,
  "title": "Understanding the Tile Configuration",
  "url": "/competitions/predict-ai-model-runtime/discussion/442258",
  "author_name": "João Víctor Melo",
  "post_date": "2023-09-21T20:39:40.634000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi, I can't compreend or even, I don't have sources enough to interpret correctly the following: what is the exact meaning of the {output: , kernel:  } in the following, and also, what is the meaning of config = {output: [2,8], kernel: [4,16,8]} runtime = 123ms? </p>\n<p>If something has any sources to recommend me, I'd appreciate a lot. (I'm also inviting to my team) <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6693355%2F2ea0142488f7c5068b8270adc32b9d2f%2Fconfig%20TPU%20tile.png?generation=1695328764498916&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2450415,
      "postDate": "2023-09-21T20:39:40.633Z",
      "content": "<p>Hi, I can't compreend or even, I don't have sources enough to interpret correctly the following: what is the exact meaning of the {output: , kernel:  } in the following, and also, what is the meaning of config = {output: [2,8], kernel: [4,16,8]} runtime = 123ms? </p>\n<p>If something has any sources to recommend me, I'd appreciate a lot. (I'm also inviting to my team) <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6693355%2F2ea0142488f7c5068b8270adc32b9d2f%2Fconfig%20TPU%20tile.png?generation=1695328764498916&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi, I can't compreend or even, I don't have sources enough to interpret correctly the following: what is the exact meaning of the {output: <output tile size multipliers>, kernel: <kernel tile size multipliers> } in the following, and also, what is the meaning of config = {output: [2,8], kernel: [4,16,8]} runtime = 123ms? \n\nIf something has any sources to recommend me, I'd appreciate a lot. (I'm also inviting to my team) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6693355%2F2ea0142488f7c5068b8270adc32b9d2f%2Fconfig%20TPU%20tile.png?generation=1695328764498916&alt=media)",
      "votes": 3
    },
    {
      "id": 2453779,
      "postDate": "2023-09-24T09:54:39.593Z",
      "content": "<p>For what regards 'output' (I still don't know about 'kernel'): since the conv output size is [2,8] if we size the output tile size at [2,4] this convolution will be run two times (on some smaller matrices), while instead with output size [2,8] the conv will be run only once but on the whole matrix.</p>\n<ul>\n<li>in the first case you'll have some added overhead due to repeating the operation two times, BUT it's easier that your matrix will fit in the registry/scratchpad memory. Meaning that you reduce cache misses but at the cost of some overhead</li>\n<li>the second case is the opposite: you reduce the overhead (repetitions) but you might incur in cache misses.</li>\n</ul>\n<p>Now this actually depends on the size of the cache of the TPU, which from what I read should be 128x8.</p>\n<p>Pls correct me if I'm wrong</p>",
      "rawMarkdown": "For what regards 'output' (I still don't know about 'kernel'): since the conv output size is [2,8] if we size the output tile size at [2,4] this convolution will be run two times (on some smaller matrices), while instead with output size [2,8] the conv will be run only once but on the whole matrix.\n- in the first case you'll have some added overhead due to repeating the operation two times, BUT it's easier that your matrix will fit in the registry/scratchpad memory. Meaning that you reduce cache misses but at the cost of some overhead\n- the second case is the opposite: you reduce the overhead (repetitions) but you might incur in cache misses.\n\nNow this actually depends on the size of the cache of the TPU, which from what I read should be 128x8.\n\nPls correct me if I'm wrong",
      "replies": [
        {
          "id": 2456007,
          "postDate": "2023-09-25T21:35:50.803Z",
          "content": "<p>First of all, I would like to apologize that the example is over simplified. Please ignore the actual numbers. The tensor size [2,8] is way too small. And the configs don't actually match the graph that well. </p>\n<p>\"output\" in the config refers to tile size of the output tensor. \"kernel\" refers to tile size of the convolution kernel tensor. For example, in this convolution: Y[i,j] = sum_{ki,kj,c} X[i+ki, j+kj, c] * K[ki,kj,c], K is the convolution kernel.</p>\n<p>In fact, the tile config indicates the multipliers of tile size from major-to-minor dimensions (reverse of the layout config). Let take a look at a realistic example of a convolution with the follow shapes, layouts, and tile config.</p>\n<pre><code>    [,,]   {,,}    {,,}\n   [,,]   {,,}    {,,}\n   [,,]   {,,}    {,,}\n</code></pre>\n<p>Consider the kernel tensor, layout specifies minor-to-major dimensions, so the tensor is laid out as [1,3072,6144] in the physical memory (major to minor), and tile config of {1,96,7} translates to tile size of {1, 96 x 8, 7 x 128}, where the last two dimensions are multiplied by the 2D registers shape. The output tensor is laid out as [4,2048,3072] in the physical memory, and tile config of {4,16,6} translates to tile size of {4, 16 x 8, 6 x 128}. This is quite completed, so we didn't provide all the details in the description.</p>",
          "rawMarkdown": "First of all, I would like to apologize that the example is over simplified. Please ignore the actual numbers. The tensor size [2,8] is way too small. And the configs don't actually match the graph that well. \n\n\"output\" in the config refers to tile size of the output tensor. \"kernel\" refers to tile size of the convolution kernel tensor. For example, in this convolution: Y[i,j] = sum_{ki,kj,c} X[i+ki, j+kj, c] * K[ki,kj,c], K is the convolution kernel.\n\nIn fact, the tile config indicates the multipliers of tile size from major-to-minor dimensions (reverse of the layout config). Let take a look at a realistic example of a convolution with the follow shapes, layouts, and tile config.\n\n```\ninput:  shape = [4,2048,6144], layout = {2,1,0}, tile config = {4,16,7}\nkernel: shape = [3072,6144,1], layout = {1,0,2}, tile config = {1,96,7}\noutput: shape = [4,2048,3072], layout = {2,1,0}, tile config = {4,16,6}\n```\n\nConsider the kernel tensor, layout specifies minor-to-major dimensions, so the tensor is laid out as [1,3072,6144] in the physical memory (major to minor), and tile config of {1,96,7} translates to tile size of {1, 96 x 8, 7 x 128}, where the last two dimensions are multiplied by the 2D registers shape. The output tensor is laid out as [4,2048,3072] in the physical memory, and tile config of {4,16,6} translates to tile size of {4, 16 x 8, 6 x 128}. This is quite completed, so we didn't provide all the details in the description.",
          "votes": 5,
          "replies": [
            {
              "id": 2456887,
              "postDate": "2023-09-26T13:41:26.037Z",
              "content": "<p>Thanks for the correction!</p>",
              "rawMarkdown": "Thanks for the correction!\n"
            },
            {
              "id": 2458500,
              "postDate": "2023-09-27T16:18:01.720Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2458501,
              "postDate": "2023-09-27T16:18:21.440Z",
              "content": "<p>From where did we get 8 and 128 in {1, 96 x 8, 7 x 128}?</p>",
              "rawMarkdown": "From where did we get 8 and 128 in {1, 96 x 8, 7 x 128}?"
            },
            {
              "id": 2460766,
              "postDate": "2023-09-29T05:16:37.480Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2460769,
              "postDate": "2023-09-29T05:17:17.660Z",
              "content": "<p>They are the sizes of the 2D registers on the TPU. More details <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/437673\" target=\"_blank\">here</a>.</p>",
              "rawMarkdown": "They are the sizes of the 2D registers on the TPU. More details [here](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/437673).",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2453779,
      "author_name": "Matteo Migliarini",
      "author_url": "",
      "post_date": "2023-09-24T09:54:39.593000",
      "content": "<p>For what regards 'output' (I still don't know about 'kernel'): since the conv output size is [2,8] if we size the output tile size at [2,4] this convolution will be run two times (on some smaller matrices), while instead with output size [2,8] the conv will be run only once but on the whole matrix.</p>\n<ul>\n<li>in the first case you'll have some added overhead due to repeating the operation two times, BUT it's easier that your matrix will fit in the registry/scratchpad memory. Meaning that you reduce cache misses but at the cost of some overhead</li>\n<li>the second case is the opposite: you reduce the overhead (repetitions) but you might incur in cache misses.</li>\n</ul>\n<p>Now this actually depends on the size of the cache of the TPU, which from what I read should be 128x8.</p>\n<p>Pls correct me if I'm wrong</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2456007,
          "author_name": "Mangpo Phothilimthana",
          "author_url": "",
          "post_date": "2023-09-25T21:35:50.803000",
          "content": "<p>First of all, I would like to apologize that the example is over simplified. Please ignore the actual numbers. The tensor size [2,8] is way too small. And the configs don't actually match the graph that well. </p>\n<p>\"output\" in the config refers to tile size of the output tensor. \"kernel\" refers to tile size of the convolution kernel tensor. For example, in this convolution: Y[i,j] = sum_{ki,kj,c} X[i+ki, j+kj, c] * K[ki,kj,c], K is the convolution kernel.</p>\n<p>In fact, the tile config indicates the multipliers of tile size from major-to-minor dimensions (reverse of the layout config). Let take a look at a realistic example of a convolution with the follow shapes, layouts, and tile config.</p>\n<pre><code>    [,,]   {,,}    {,,}\n   [,,]   {,,}    {,,}\n   [,,]   {,,}    {,,}\n</code></pre>\n<p>Consider the kernel tensor, layout specifies minor-to-major dimensions, so the tensor is laid out as [1,3072,6144] in the physical memory (major to minor), and tile config of {1,96,7} translates to tile size of {1, 96 x 8, 7 x 128}, where the last two dimensions are multiplied by the 2D registers shape. The output tensor is laid out as [4,2048,3072] in the physical memory, and tile config of {4,16,6} translates to tile size of {4, 16 x 8, 6 x 128}. This is quite completed, so we didn't provide all the details in the description.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2456887,
              "author_name": "Matteo Migliarini",
              "author_url": "",
              "post_date": "2023-09-26T13:41:26.037000",
              "content": "<p>Thanks for the correction!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2458500,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-27T16:18:01.720000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2458501,
              "author_name": "João Víctor Melo",
              "author_url": "",
              "post_date": "2023-09-27T16:18:21.440000",
              "content": "<p>From where did we get 8 and 128 in {1, 96 x 8, 7 x 128}?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2460766,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-29T05:16:37.480000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2460769,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-09-29T05:17:17.660000",
              "content": "<p>They are the sizes of the 2D registers on the TPU. More details <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/437673\" target=\"_blank\">here</a>.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2450415": "Hi, I can't compreend or even, I don't have sources enough to interpret correctly the following: what is the exact meaning of the {output: <output tile size multipliers>, kernel: <kernel tile size multipliers> } in the following, and also, what is the meaning of config = {output: [2,8], kernel: [4,16,8]} runtime = 123ms? \n\nIf something has any sources to recommend me, I'd appreciate a lot. (I'm also inviting to my team) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6693355%2F2ea0142488f7c5068b8270adc32b9d2f%2Fconfig%20TPU%20tile.png?generation=1695328764498916&alt=media)",
    "2453779": "For what regards 'output' (I still don't know about 'kernel'): since the conv output size is [2,8] if we size the output tile size at [2,4] this convolution will be run two times (on some smaller matrices), while instead with output size [2,8] the conv will be run only once but on the whole matrix.\n- in the first case you'll have some added overhead due to repeating the operation two times, BUT it's easier that your matrix will fit in the registry/scratchpad memory. Meaning that you reduce cache misses but at the cost of some overhead\n- the second case is the opposite: you reduce the overhead (repetitions) but you might incur in cache misses.\n\nNow this actually depends on the size of the cache of the TPU, which from what I read should be 128x8.\n\nPls correct me if I'm wrong"
  }
}