{
  "id": 438074,
  "title": "Understanding the layout config features of a convolutional node",
  "url": "/competitions/predict-ai-model-runtime/discussion/438074",
  "author_name": "",
  "post_date": "2023-09-09T11:15:12.024081700Z",
  "votes": 17,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi!</p>\n<p>I'm doing exploratory data analysis on the data and got stuck on understanding the layout config features for convolutional nodes. To have something specific to discuss around, I am using <code>npz_all/npz/layout/xla/default/train/alexnet_train_batch_32.npz</code> as the development graph. The first configurable node id is <strong>86</strong>, which is a convolution. Below you can see descriptions of relevant features on the form <code>{(feature_index, feature_name): feature_value}</code></p>\n<p>The output tensor features are</p>\n<pre><code>{(, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): }}\n</code></pre>\n<p>the convolution features are</p>\n<pre><code>{(,  ): ,\n (,  ): ,\n (,  ): ,\n (,  ): ,\n (,  ): ,\n (,  ): ,\n (,  ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): }\n</code></pre>\n<p>and the layout config features are</p>\n<pre><code>{(,  ): ,\n (,  ): ,\n (,  ): -,\n (,  ): -,\n (,  ): -,\n (,  ): -,\n (,  ): ,\n (,  ): ,\n (,  ): -,\n (,  ): -,\n (, ): -,\n (, ): -,\n (, ): ,\n (, ): ,\n (, ): -,\n (, ): -,\n (, ): -,\n (, ): -}\n</code></pre>\n<p>Why does the layout config only describes 2 dimensions and not all 4 (seems to be missing channel and batch dimensions).</p>\n<p>Cheers!</p>",
  "messages": [
    {
      "id": "2430558",
      "postDate": "09/09/2023 11:15:12",
      "content": "<p>Hi!</p>\n<p>I'm doing exploratory data analysis on the data and got stuck on understanding the layout config features for convolutional nodes. To have something specific to discuss around, I am using <code>npz_all/npz/layout/xla/default/train/alexnet_train_batch_32.npz</code> as the development graph. The first configurable node id is <strong>86</strong>, which is a convolution. Below you can see descriptions of relevant features on the form <code>{(feature_index, feature_name): feature_value}</code></p>\n<p>The output tensor features are</p>\n<pre><code>{(, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): }}\n</code></pre>\n<p>the convolution features are</p>\n<pre><code>{(,  ): ,\n (,  ): ,\n (,  ): ,\n (,  ): ,\n (,  ): ,\n (,  ): ,\n (,  ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): ,\n (, ): }\n</code></pre>\n<p>and the layout config features are</p>\n<pre><code>{(,  ): ,\n (,  ): ,\n (,  ): -,\n (,  ): -,\n (,  ): -,\n (,  ): -,\n (,  ): ,\n (,  ): ,\n (,  ): -,\n (,  ): -,\n (, ): -,\n (, ): -,\n (, ): ,\n (, ): ,\n (, ): -,\n (, ): -,\n (, ): -,\n (, ): -}\n</code></pre>\n<p>Why does the layout config only describes 2 dimensions and not all 4 (seems to be missing channel and batch dimensions).</p>\n<p>Cheers!</p>",
      "rawMarkdown": "Hi!\n\nI'm doing exploratory data analysis on the data and got stuck on understanding the layout config features for convolutional nodes. To have something specific to discuss around, I am using `npz_all/npz/layout/xla/default/train/alexnet_train_batch_32.npz` as the development graph. The first configurable node id is **86**, which is a convolution. Below you can see descriptions of relevant features on the form `{(feature_index, feature_name): feature_value}`\n\nThe output tensor features are\n\n```python\n{(21, 'shape_dimensions_0'): 32,\n (22, 'shape_dimensions_1'): 54,\n (23, 'shape_dimensions_2'): 54,\n (24, 'shape_dimensions_3'): 64,\n (25, 'shape_dimensions_4'): 0,\n (26, 'shape_dimensions_5'): 0}}\n```\n\nthe convolution features are\n\n```python\n{(93,  'convolution_dim_numbers_input_batch_dim'): 0,\n (94,  'convolution_dim_numbers_input_feature_dim'): 3,\n (95,  'convolution_dim_numbers_input_spatial_dims_0'): 1,\n (96,  'convolution_dim_numbers_input_spatial_dims_1'): 2,\n (97,  'convolution_dim_numbers_input_spatial_dims_2'): 0,\n (98,  'convolution_dim_numbers_input_spatial_dims_3'): 0,\n (99,  'convolution_dim_numbers_kernel_input_feature_dim'): 2,\n (100, 'convolution_dim_numbers_kernel_output_feature_dim'): 3,\n (101, 'convolution_dim_numbers_kernel_spatial_dims_0'): 0,\n (102, 'convolution_dim_numbers_kernel_spatial_dims_1'): 1,\n (103, 'convolution_dim_numbers_kernel_spatial_dims_2'): 0,\n (104, 'convolution_dim_numbers_kernel_spatial_dims_3'): 0,\n (105, 'convolution_dim_numbers_output_batch_dim'): 0,\n (106, 'convolution_dim_numbers_output_feature_dim'): 3}\n```\n\nand the layout config features are\n\n```python\n{(0,  'output_layout_0'): 3,\n (1,  'output_layout_1'): 0,\n (2,  'output_layout_2'): -1,\n (3,  'output_layout_3'): -1,\n (4,  'output_layout_4'): -1,\n (5,  'output_layout_5'): -1,\n (6,  'intput_layout_0'): 0,\n (7,  'intput_layout_1'): 3,\n (8,  'intput_layout_2'): -1,\n (9,  'intput_layout_3'): -1,\n (10, 'intput_layout_4'): -1,\n (11, 'intput_layout_5'): -1,\n (12, 'kernel_layout_0'): 2,\n (13, 'kernel_layout_1'): 3,\n (14, 'kernel_layout_2'): -1,\n (15, 'kernel_layout_3'): -1,\n (16, 'kernel_layout_4'): -1,\n (17, 'kernel_layout_5'): -1}\n```\n\n\nWhy does the layout config only describes 2 dimensions and not all 4 (seems to be missing channel and batch dimensions).\n\nCheers!",
      "votes": null
    },
    {
      "id": "2441545",
      "postDate": "09/16/2023 08:55:40",
      "content": "<p>I made a tool to render the data in the layout configuration datasets using graphviz. Here is a visual representation of the <code>alexnet_train_batch_32 graph</code> from the <strong>random</strong> (note, not <strong>default</strong> as above, minor detail but good to point it out) for the 0:th configuration. The image is huge, I recommend downloading it and viewing locally on your computer. Some notes about the image</p>\n<ol>\n<li>The graph is split according to <code>node_splits</code></li>\n<li>Nodes are labelled by node index and opcode name</li>\n<li>Edges are annotated with tensor dimensions in square brackets and tensor layout in braces</li>\n<li>Configurable nodes are blue</li>\n<li>Configurable nodes are annotated with the input and output layout configuration</li>\n</ol>\n<p>As you can see, all convolutional nodes have underspecified layout configuration (only 2 out of 4 dimensions are configured, rest are set to -1), dot nodes configuration have single non negative configuration dimension. Reshapes seems to be the only node type that fully specifies the layout config for every dimension in the output tensor.<br>\nCan you please describe how to interpret -1 for nodes with convolution and dot opcodes?<br>\nThanks!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7962244%2Fe7f2d1eac645d5f16e515d9b561fdb4e%2Flayout-xla-random-train-alexnet_train_batch_32-config-0.svg?generation=1694854469659730&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I made a tool to render the data in the layout configuration datasets using graphviz. Here is a visual representation of the `alexnet_train_batch_32 graph` from the **random** (note, not **default** as above, minor detail but good to point it out) for the 0:th configuration. The image is huge, I recommend downloading it and viewing locally on your computer. Some notes about the image\n\n1. The graph is split according to `node_splits`\n2. Nodes are labelled by node index and opcode name\n3. Edges are annotated with tensor dimensions in square brackets and tensor layout in braces\n4. Configurable nodes are blue\n5. Configurable nodes are annotated with the input and output layout configuration\n\nAs you can see, all convolutional nodes have underspecified layout configuration (only 2 out of 4 dimensions are configured, rest are set to -1), dot nodes configuration have single non negative configuration dimension. Reshapes seems to be the only node type that fully specifies the layout config for every dimension in the output tensor.\nCan you please describe how to interpret -1 for nodes with convolution and dot opcodes?\nThanks!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7962244%2Fe7f2d1eac645d5f16e515d9b561fdb4e%2Flayout-xla-random-train-alexnet_train_batch_32-config-0.svg?generation=1694854469659730&alt=media)",
      "votes": null
    },
    {
      "id": "2443451",
      "postDate": "09/17/2023 18:05:41",
      "content": "<p>The layout configuration of a convolution/dot node specifies only the first two dimensions because of the 2D nature of registers on a TPU. The other dimensions are essentially outer loops that do not matter to the speed of the program.</p>\n<p>-1 is just a padding value.</p>",
      "rawMarkdown": "The layout configuration of a convolution/dot node specifies only the first two dimensions because of the 2D nature of registers on a TPU. The other dimensions are essentially outer loops that do not matter to the speed of the program.\n\n-1 is just a padding value.",
      "votes": null
    },
    {
      "id": "2444086",
      "postDate": "09/18/2023 06:38:49",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> for you input. I got a bit confused since in the data description it says: <em>For example, the layout of {1, 0, 2, -1, -1, -1} of a 3D tensor indicates that dimension 1 is the most minor (elements of the most minor dimension are consecutive in the physical space), and dimension 2 is the most major.</em>, which mentions configuration of all 3 dimensions. I understand your reasoning, and I will trust that only the 2 inner dimensions have an impact on the results. I listened to one of your talks and I recall you saying that if a input tensor layout does not match the node input layout configuration, then a copy operation has to be inserted. Does that mean that copy operations only have to be inserted when the 2 most inner dimensions are mismatching, and the outer loops will just be rearranged to match?</p>",
      "rawMarkdown": "Thanks @mangpophothilimthana for you input. I got a bit confused since in the data description it says: *For example, the layout of {1, 0, 2, -1, -1, -1} of a 3D tensor indicates that dimension 1 is the most minor (elements of the most minor dimension are consecutive in the physical space), and dimension 2 is the most major.*, which mentions configuration of all 3 dimensions. I understand your reasoning, and I will trust that only the 2 inner dimensions have an impact on the results. I listened to one of your talks and I recall you saying that if a input tensor layout does not match the node input layout configuration, then a copy operation has to be inserted. Does that mean that copy operations only have to be inserted when the 2 most inner dimensions are mismatching, and the outer loops will just be rearranged to match?",
      "votes": null
    },
    {
      "id": "2444095",
      "postDate": "09/18/2023 06:48:28",
      "content": "<p>Another interesting observation I've found after looking at your graph &amp; doing a bit research myself: HLO seems to not be a DAG, which is a bit counter-intuitive because I thought most CV / NLP models are DAGs.</p>",
      "rawMarkdown": "Another interesting observation I've found after looking at your graph & doing a bit research myself: HLO seems to not be a DAG, which is a bit counter-intuitive because I thought most CV / NLP models are DAGs.",
      "votes": null
    },
    {
      "id": "2447387",
      "postDate": "09/20/2023 03:45:30",
      "content": "<p>Clarification: Only the first two dimensions of the layout matters for performance of convolution and dot. All dimensions of layout matters for reshape.</p>\n<blockquote>\n  <p>Does that mean that copy operations only have to be inserted when the 2 most inner dimensions are mismatching, and the outer loops will just be rearranged to match?</p>\n</blockquote>\n<p>Correct</p>",
      "rawMarkdown": "Clarification: Only the first two dimensions of the layout matters for performance of convolution and dot. All dimensions of layout matters for reshape.\n\n>Does that mean that copy operations only have to be inserted when the 2 most inner dimensions are mismatching, and the outer loops will just be rearranged to match?\n\nCorrect",
      "votes": null
    },
    {
      "id": "2452398",
      "postDate": "09/23/2023 08:56:44",
      "content": "<p>Can you see any loops in the alexnet graph i rendered, or do you refer to another network?</p>",
      "rawMarkdown": "Can you see any loops in the alexnet graph i rendered, or do you refer to another network?",
      "votes": null
    },
    {
      "id": "2453015",
      "postDate": "09/23/2023 18:06:53",
      "content": "<p>All networks in this dataset have a loop. In your network, it's around the <code>while</code> I believe.</p>",
      "rawMarkdown": "All networks in this dataset have a loop. In your network, it's around the `while` I believe.",
      "votes": null
    },
    {
      "id": "2454221",
      "postDate": "09/24/2023 16:49:55",
      "content": "<p>I was a confused about that part of the graph as well. AlexNet does not have any loops, I am curious what that OP means. Technically, the while OP does not introduce any loops in the graph, so it's still a DAG. However, as the name suggests, it introduces a loop implicitly. Implicit information would make it really hard to learn to predict the runtime behaviou also…</p>",
      "rawMarkdown": "I was a confused about that part of the graph as well. AlexNet does not have any loops, I am curious what that OP means. Technically, the while OP does not introduce any loops in the graph, so it's still a DAG. However, as the name suggests, it introduces a loop implicitly. Implicit information would make it really hard to learn to predict the runtime behaviou also...",
      "votes": null
    },
    {
      "id": "2455910",
      "postDate": "09/25/2023 19:01:47",
      "content": "<p>Nice observation! A few operations (such as <code>while</code>, <code>call</code>, and <code>custom-call</code>) in an HLO graph are special in a sense that they introduce nested computations (graphs). You can find the semantics of HLO operations <a href=\"https://www.tensorflow.org/xla/operation_semantics\" target=\"_blank\">here</a> (e.g. <a href=\"https://www.tensorflow.org/xla/operation_semantics#while\" target=\"_blank\">while</a>).</p>\n<p>The way we handle these nodes and the nested graphs is to conceptually flattening everything into one graph. For example, a <a href=\"https://www.tensorflow.org/xla/operation_semantics#while\" target=\"_blank\">while</a> operation contains nested computations for condition, body, and init value. We handle a while node by connecting the while node with the  root (output) nodes of the nested computation graphs representing condition, body, and init value. You can see this logic in our code that extracts edges <a href=\"https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/hlo_encoder.cc#L435\" target=\"_blank\">here</a>.</p>\n<p>Now, one could treat these types of edges differently from normal dependency edges. For example, GNNs can handle edges of different types. However, the baseline models treat all edges the same.</p>",
      "rawMarkdown": "Nice observation! A few operations (such as `while`, `call`, and `custom-call`) in an HLO graph are special in a sense that they introduce nested computations (graphs). You can find the semantics of HLO operations [here](https://www.tensorflow.org/xla/operation_semantics) (e.g. [while](https://www.tensorflow.org/xla/operation_semantics#while)).\n\nThe way we handle these nodes and the nested graphs is to conceptually flattening everything into one graph. For example, a [while](https://www.tensorflow.org/xla/operation_semantics#while) operation contains nested computations for condition, body, and init value. We handle a while node by connecting the while node with the  root (output) nodes of the nested computation graphs representing condition, body, and init value. You can see this logic in our code that extracts edges [here](https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/hlo_encoder.cc#L435).\n\nNow, one could treat these types of edges differently from normal dependency edges. For example, GNNs can handle edges of different types. However, the baseline models treat all edges the same.",
      "votes": null
    },
    {
      "id": "2482584",
      "postDate": "10/15/2023 06:43:17",
      "content": "<p>What about the output layout? Since it needs to be propagated (<a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/443265#2473896)\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/443265#2473896)</a>, that must be handled somehow? Current output layout config only gives 2 inner dims for conv nodes. How to \"infer\" the other two for layout propagation purposes? Or is that unnecessary to do?</p>",
      "rawMarkdown": "What about the output layout? Since it needs to be propagated (https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/443265#2473896), that must be handled somehow? Current output layout config only gives 2 inner dims for conv nodes. How to \"infer\" the other two for layout propagation purposes? Or is that unnecessary to do?",
      "votes": null
    },
    {
      "id": "2485137",
      "postDate": "10/17/2023 00:52:48",
      "content": "<p>From what I understand, the other dimensions beyond the two most minor dimensions of the layouts of convolution don't matter in term of performance of the convolution itself because they will all be treated as outer loops. I suppose it would matter if it is eventually connected to a reshape operation. In that case I guess, the layout of the other dimensions could follow the reshape's layout. Discloser: I didn't implement layout propagation in the compiler, so this information is from the best of my knowledge.</p>",
      "rawMarkdown": "From what I understand, the other dimensions beyond the two most minor dimensions of the layouts of convolution don't matter in term of performance of the convolution itself because they will all be treated as outer loops. I suppose it would matter if it is eventually connected to a reshape operation. In that case I guess, the layout of the other dimensions could follow the reshape's layout. Discloser: I didn't implement layout propagation in the compiler, so this information is from the best of my knowledge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2441545,
      "author_name": "carldehlin",
      "author_url": "",
      "post_date": "09/16/2023 08:55:40",
      "content": "<p>I made a tool to render the data in the layout configuration datasets using graphviz. Here is a visual representation of the <code>alexnet_train_batch_32 graph</code> from the <strong>random</strong> (note, not <strong>default</strong> as above, minor detail but good to point it out) for the 0:th configuration. The image is huge, I recommend downloading it and viewing locally on your computer. Some notes about the image</p>\n<ol>\n<li>The graph is split according to <code>node_splits</code></li>\n<li>Nodes are labelled by node index and opcode name</li>\n<li>Edges are annotated with tensor dimensions in square brackets and tensor layout in braces</li>\n<li>Configurable nodes are blue</li>\n<li>Configurable nodes are annotated with the input and output layout configuration</li>\n</ol>\n<p>As you can see, all convolutional nodes have underspecified layout configuration (only 2 out of 4 dimensions are configured, rest are set to -1), dot nodes configuration have single non negative configuration dimension. Reshapes seems to be the only node type that fully specifies the layout config for every dimension in the output tensor.<br>\nCan you please describe how to interpret -1 for nodes with convolution and dot opcodes?<br>\nThanks!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7962244%2Fe7f2d1eac645d5f16e515d9b561fdb4e%2Flayout-xla-random-train-alexnet_train_batch_32-config-0.svg?generation=1694854469659730&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2443451,
      "author_name": "mangpophothilimthana",
      "author_url": "",
      "post_date": "09/17/2023 18:05:41",
      "content": "<p>The layout configuration of a convolution/dot node specifies only the first two dimensions because of the 2D nature of registers on a TPU. The other dimensions are essentially outer loops that do not matter to the speed of the program.</p>\n<p>-1 is just a padding value.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2444086,
          "author_name": "carldehlin",
          "author_url": "",
          "post_date": "09/18/2023 06:38:49",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> for you input. I got a bit confused since in the data description it says: <em>For example, the layout of {1, 0, 2, -1, -1, -1} of a 3D tensor indicates that dimension 1 is the most minor (elements of the most minor dimension are consecutive in the physical space), and dimension 2 is the most major.</em>, which mentions configuration of all 3 dimensions. I understand your reasoning, and I will trust that only the 2 inner dimensions have an impact on the results. I listened to one of your talks and I recall you saying that if a input tensor layout does not match the node input layout configuration, then a copy operation has to be inserted. Does that mean that copy operations only have to be inserted when the 2 most inner dimensions are mismatching, and the outer loops will just be rearranged to match?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2447387,
              "author_name": "mangpophothilimthana",
              "author_url": "",
              "post_date": "09/20/2023 03:45:30",
              "content": "<p>Clarification: Only the first two dimensions of the layout matters for performance of convolution and dot. All dimensions of layout matters for reshape.</p>\n<blockquote>\n  <p>Does that mean that copy operations only have to be inserted when the 2 most inner dimensions are mismatching, and the outer loops will just be rearranged to match?</p>\n</blockquote>\n<p>Correct</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2482584,
                  "author_name": "carldehlin",
                  "author_url": "",
                  "post_date": "10/15/2023 06:43:17",
                  "content": "<p>What about the output layout? Since it needs to be propagated (<a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/443265#2473896)\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/443265#2473896)</a>, that must be handled somehow? Current output layout config only gives 2 inner dims for conv nodes. How to \"infer\" the other two for layout propagation purposes? Or is that unnecessary to do?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2485137,
                      "author_name": "mangpophothilimthana",
                      "author_url": "",
                      "post_date": "10/17/2023 00:52:48",
                      "content": "<p>From what I understand, the other dimensions beyond the two most minor dimensions of the layouts of convolution don't matter in term of performance of the convolution itself because they will all be treated as outer loops. I suppose it would matter if it is eventually connected to a reshape operation. In that case I guess, the layout of the other dimensions could follow the reshape's layout. Discloser: I didn't implement layout propagation in the compiler, so this information is from the best of my knowledge.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2444095,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "09/18/2023 06:48:28",
      "content": "<p>Another interesting observation I've found after looking at your graph &amp; doing a bit research myself: HLO seems to not be a DAG, which is a bit counter-intuitive because I thought most CV / NLP models are DAGs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2452398,
          "author_name": "carldehlin",
          "author_url": "",
          "post_date": "09/23/2023 08:56:44",
          "content": "<p>Can you see any loops in the alexnet graph i rendered, or do you refer to another network?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2453015,
              "author_name": "alexanderliao",
              "author_url": "",
              "post_date": "09/23/2023 18:06:53",
              "content": "<p>All networks in this dataset have a loop. In your network, it's around the <code>while</code> I believe.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2454221,
                  "author_name": "carldehlin",
                  "author_url": "",
                  "post_date": "09/24/2023 16:49:55",
                  "content": "<p>I was a confused about that part of the graph as well. AlexNet does not have any loops, I am curious what that OP means. Technically, the while OP does not introduce any loops in the graph, so it's still a DAG. However, as the name suggests, it introduces a loop implicitly. Implicit information would make it really hard to learn to predict the runtime behaviou also…</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2455910,
                  "author_name": "mangpophothilimthana",
                  "author_url": "",
                  "post_date": "09/25/2023 19:01:47",
                  "content": "<p>Nice observation! A few operations (such as <code>while</code>, <code>call</code>, and <code>custom-call</code>) in an HLO graph are special in a sense that they introduce nested computations (graphs). You can find the semantics of HLO operations <a href=\"https://www.tensorflow.org/xla/operation_semantics\" target=\"_blank\">here</a> (e.g. <a href=\"https://www.tensorflow.org/xla/operation_semantics#while\" target=\"_blank\">while</a>).</p>\n<p>The way we handle these nodes and the nested graphs is to conceptually flattening everything into one graph. For example, a <a href=\"https://www.tensorflow.org/xla/operation_semantics#while\" target=\"_blank\">while</a> operation contains nested computations for condition, body, and init value. We handle a while node by connecting the while node with the  root (output) nodes of the nested computation graphs representing condition, body, and init value. You can see this logic in our code that extracts edges <a href=\"https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/hlo_encoder.cc#L435\" target=\"_blank\">here</a>.</p>\n<p>Now, one could treat these types of edges differently from normal dependency edges. For example, GNNs can handle edges of different types. However, the baseline models treat all edges the same.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2430558": "Hi!\n\nI'm doing exploratory data analysis on the data and got stuck on understanding the layout config features for convolutional nodes. To have something specific to discuss around, I am using `npz_all/npz/layout/xla/default/train/alexnet_train_batch_32.npz` as the development graph. The first configurable node id is **86**, which is a convolution. Below you can see descriptions of relevant features on the form `{(feature_index, feature_name): feature_value}`\n\nThe output tensor features are\n\n```python\n{(21, 'shape_dimensions_0'): 32,\n (22, 'shape_dimensions_1'): 54,\n (23, 'shape_dimensions_2'): 54,\n (24, 'shape_dimensions_3'): 64,\n (25, 'shape_dimensions_4'): 0,\n (26, 'shape_dimensions_5'): 0}}\n```\n\nthe convolution features are\n\n```python\n{(93,  'convolution_dim_numbers_input_batch_dim'): 0,\n (94,  'convolution_dim_numbers_input_feature_dim'): 3,\n (95,  'convolution_dim_numbers_input_spatial_dims_0'): 1,\n (96,  'convolution_dim_numbers_input_spatial_dims_1'): 2,\n (97,  'convolution_dim_numbers_input_spatial_dims_2'): 0,\n (98,  'convolution_dim_numbers_input_spatial_dims_3'): 0,\n (99,  'convolution_dim_numbers_kernel_input_feature_dim'): 2,\n (100, 'convolution_dim_numbers_kernel_output_feature_dim'): 3,\n (101, 'convolution_dim_numbers_kernel_spatial_dims_0'): 0,\n (102, 'convolution_dim_numbers_kernel_spatial_dims_1'): 1,\n (103, 'convolution_dim_numbers_kernel_spatial_dims_2'): 0,\n (104, 'convolution_dim_numbers_kernel_spatial_dims_3'): 0,\n (105, 'convolution_dim_numbers_output_batch_dim'): 0,\n (106, 'convolution_dim_numbers_output_feature_dim'): 3}\n```\n\nand the layout config features are\n\n```python\n{(0,  'output_layout_0'): 3,\n (1,  'output_layout_1'): 0,\n (2,  'output_layout_2'): -1,\n (3,  'output_layout_3'): -1,\n (4,  'output_layout_4'): -1,\n (5,  'output_layout_5'): -1,\n (6,  'intput_layout_0'): 0,\n (7,  'intput_layout_1'): 3,\n (8,  'intput_layout_2'): -1,\n (9,  'intput_layout_3'): -1,\n (10, 'intput_layout_4'): -1,\n (11, 'intput_layout_5'): -1,\n (12, 'kernel_layout_0'): 2,\n (13, 'kernel_layout_1'): 3,\n (14, 'kernel_layout_2'): -1,\n (15, 'kernel_layout_3'): -1,\n (16, 'kernel_layout_4'): -1,\n (17, 'kernel_layout_5'): -1}\n```\n\n\nWhy does the layout config only describes 2 dimensions and not all 4 (seems to be missing channel and batch dimensions).\n\nCheers!",
    "2441545": "I made a tool to render the data in the layout configuration datasets using graphviz. Here is a visual representation of the `alexnet_train_batch_32 graph` from the **random** (note, not **default** as above, minor detail but good to point it out) for the 0:th configuration. The image is huge, I recommend downloading it and viewing locally on your computer. Some notes about the image\n\n1. The graph is split according to `node_splits`\n2. Nodes are labelled by node index and opcode name\n3. Edges are annotated with tensor dimensions in square brackets and tensor layout in braces\n4. Configurable nodes are blue\n5. Configurable nodes are annotated with the input and output layout configuration\n\nAs you can see, all convolutional nodes have underspecified layout configuration (only 2 out of 4 dimensions are configured, rest are set to -1), dot nodes configuration have single non negative configuration dimension. Reshapes seems to be the only node type that fully specifies the layout config for every dimension in the output tensor.\nCan you please describe how to interpret -1 for nodes with convolution and dot opcodes?\nThanks!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7962244%2Fe7f2d1eac645d5f16e515d9b561fdb4e%2Flayout-xla-random-train-alexnet_train_batch_32-config-0.svg?generation=1694854469659730&alt=media)",
    "2443451": "The layout configuration of a convolution/dot node specifies only the first two dimensions because of the 2D nature of registers on a TPU. The other dimensions are essentially outer loops that do not matter to the speed of the program.\n\n-1 is just a padding value.",
    "2444086": "Thanks @mangpophothilimthana for you input. I got a bit confused since in the data description it says: *For example, the layout of {1, 0, 2, -1, -1, -1} of a 3D tensor indicates that dimension 1 is the most minor (elements of the most minor dimension are consecutive in the physical space), and dimension 2 is the most major.*, which mentions configuration of all 3 dimensions. I understand your reasoning, and I will trust that only the 2 inner dimensions have an impact on the results. I listened to one of your talks and I recall you saying that if a input tensor layout does not match the node input layout configuration, then a copy operation has to be inserted. Does that mean that copy operations only have to be inserted when the 2 most inner dimensions are mismatching, and the outer loops will just be rearranged to match?",
    "2444095": "Another interesting observation I've found after looking at your graph & doing a bit research myself: HLO seems to not be a DAG, which is a bit counter-intuitive because I thought most CV / NLP models are DAGs.",
    "2447387": "Clarification: Only the first two dimensions of the layout matters for performance of convolution and dot. All dimensions of layout matters for reshape.\n\n>Does that mean that copy operations only have to be inserted when the 2 most inner dimensions are mismatching, and the outer loops will just be rearranged to match?\n\nCorrect",
    "2452398": "Can you see any loops in the alexnet graph i rendered, or do you refer to another network?",
    "2453015": "All networks in this dataset have a loop. In your network, it's around the `while` I believe.",
    "2454221": "I was a confused about that part of the graph as well. AlexNet does not have any loops, I am curious what that OP means. Technically, the while OP does not introduce any loops in the graph, so it's still a DAG. However, as the name suggests, it introduces a loop implicitly. Implicit information would make it really hard to learn to predict the runtime behaviou also...",
    "2455910": "Nice observation! A few operations (such as `while`, `call`, and `custom-call`) in an HLO graph are special in a sense that they introduce nested computations (graphs). You can find the semantics of HLO operations [here](https://www.tensorflow.org/xla/operation_semantics) (e.g. [while](https://www.tensorflow.org/xla/operation_semantics#while)).\n\nThe way we handle these nodes and the nested graphs is to conceptually flattening everything into one graph. For example, a [while](https://www.tensorflow.org/xla/operation_semantics#while) operation contains nested computations for condition, body, and init value. We handle a while node by connecting the while node with the  root (output) nodes of the nested computation graphs representing condition, body, and init value. You can see this logic in our code that extracts edges [here](https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/hlo_encoder.cc#L435).\n\nNow, one could treat these types of edges differently from normal dependency edges. For example, GNNs can handle edges of different types. However, the baseline models treat all edges the same.",
    "2482584": "What about the output layout? Since it needs to be propagated (https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/443265#2473896), that must be handled somehow? Current output layout config only gives 2 inner dims for conv nodes. How to \"infer\" the other two for layout propagation purposes? Or is that unnecessary to do?",
    "2485137": "From what I understand, the other dimensions beyond the two most minor dimensions of the layouts of convolution don't matter in term of performance of the convolution itself because they will all be treated as outer loops. I suppose it would matter if it is eventually connected to a reshape operation. In that case I guess, the layout of the other dimensions could follow the reshape's layout. Discloser: I didn't implement layout propagation in the compiler, so this information is from the best of my knowledge."
  },
  "source": "meta"
}