{
  "id": 443265,
  "title": "Understanding the \"layout propagation\" technique",
  "url": "/competitions/predict-ai-model-runtime/discussion/443265",
  "author_name": "Shekhovtsov Aleksandr",
  "post_date": "2023-09-26T09:18:22.232000",
  "votes": 8,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I have question about layout configuration. In paper <a href=\"https://arxiv.org/pdf/2308.13490.pdf\" target=\"_blank\">TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs</a>  the authors wrote \"The autotuner tunes the input-output layouts of the most layout-performance-critical nodes and <strong>propagates layouts from these nodes to others</strong>.\" And I heard the same in some videos with <a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> <br>\nDoes anyone understand how exactly \"layout propagation\" works?<br>\nEach (non-configurable) node has 6 features related to layout. Is it a layout for an input or output tensor?<br>\nAnd after we change the layout configurations of the input and output tensors for some configurable node, what kind of propagation is required?</p>\n<p>Thanks for the clarifications in advance</p>",
  "messages": [
    {
      "id": 2456557,
      "postDate": "2023-09-26T09:18:22.233Z",
      "content": "<p>I have question about layout configuration. In paper <a href=\"https://arxiv.org/pdf/2308.13490.pdf\" target=\"_blank\">TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs</a>  the authors wrote \"The autotuner tunes the input-output layouts of the most layout-performance-critical nodes and <strong>propagates layouts from these nodes to others</strong>.\" And I heard the same in some videos with <a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> <br>\nDoes anyone understand how exactly \"layout propagation\" works?<br>\nEach (non-configurable) node has 6 features related to layout. Is it a layout for an input or output tensor?<br>\nAnd after we change the layout configurations of the input and output tensors for some configurable node, what kind of propagation is required?</p>\n<p>Thanks for the clarifications in advance</p>",
      "rawMarkdown": "I have question about layout configuration. In paper [TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs](https://arxiv.org/pdf/2308.13490.pdf)  the authors wrote \"The autotuner tunes the input-output layouts of the most layout-performance-critical nodes and **propagates layouts from these nodes to others**.\" And I heard the same in some videos with @mangpophothilimthana \nDoes anyone understand how exactly \"layout propagation\" works?\nEach (non-configurable) node has 6 features related to layout. Is it a layout for an input or output tensor?\nAnd after we change the layout configurations of the input and output tensors for some configurable node, what kind of propagation is required?\n\nThanks for the clarifications in advance",
      "votes": 7
    },
    {
      "id": 2465300,
      "postDate": "2023-10-02T20:07:31.617Z",
      "content": "<p>The layout propagation algorithm in the compiler is rather complex, and I don't fully understand it myself (I didn't implement it). The high-level idea is that the compiler (1) selects a convolution, dot, or reshape node and assign the best layout for that node, and (2) propagate the layout from that node as far as possible. The compiler integrates between these two steps until all layouts are assigned. There are a bunch of rules on how a layout is propagated between inputs and output of a node. Some can be found <a href=\"https://github.com/openxla/xla/blob/main/xla/service/layout_assignment.cc\" target=\"_blank\">here</a>. For example, an element-wise node will assign the same layout for its inputs and output during propagation.</p>",
      "rawMarkdown": "The layout propagation algorithm in the compiler is rather complex, and I don't fully understand it myself (I didn't implement it). The high-level idea is that the compiler (1) selects a convolution, dot, or reshape node and assign the best layout for that node, and (2) propagate the layout from that node as far as possible. The compiler integrates between these two steps until all layouts are assigned. There are a bunch of rules on how a layout is propagated between inputs and output of a node. Some can be found [here](https://github.com/openxla/xla/blob/main/xla/service/layout_assignment.cc). For example, an element-wise node will assign the same layout for its inputs and output during propagation.",
      "votes": 2,
      "replies": [
        {
          "id": 2473896,
          "postDate": "2023-10-08T16:38:40.957Z",
          "content": "<p>How does this interact with the inserted copy operations? This is how I understand that it works (please correct me where I'm wrong) </p>\n<p><strong>For configurable nodes</strong></p>\n<ol>\n<li>Configurable nodes expect input tensor to be laid out as described by <code>input_layout_&lt;N&gt;</code> fields in the layout config</li>\n<li>A copy operation is inserted when the layout an input tensor does not match the config output layout</li>\n<li>After inserting the copy operation, the layouts match</li>\n<li>The layout of the output tensor of the configurable node is as described by the layout output config</li>\n</ol>\n<p><strong>For non-configurable nodes</strong></p>\n<ol>\n<li>Non-configurable nodes expect input tensors to be laid out as described by <code>layout_minor_to_major_&lt;N&gt;</code> fields</li>\n<li>Non-configurable nodes do dot change the tensor layout (so the output layout is the same as the input)</li>\n<li>If the input tensor does not match the expected input layout, a copy operation is inserted to resolve like for non-configurable nodes</li>\n</ol>\n<p>What I don't understand is how this requires any advanced layout propagation, it should be at most a copy at the input and outputs of each configurable nodes.</p>\n<p>Thankful if you can provide some insights to this!</p>",
          "rawMarkdown": "How does this interact with the inserted copy operations? This is how I understand that it works (please correct me where I'm wrong) \n\n**For configurable nodes**\n1. Configurable nodes expect input tensor to be laid out as described by `input_layout_<N>` fields in the layout config\n1. A copy operation is inserted when the layout an input tensor does not match the config output layout\n2. After inserting the copy operation, the layouts match\n3. The layout of the output tensor of the configurable node is as described by the layout output config\n\n**For non-configurable nodes**\n1. Non-configurable nodes expect input tensors to be laid out as described by `layout_minor_to_major_<N>` fields\n2. Non-configurable nodes do dot change the tensor layout (so the output layout is the same as the input)\n3. If the input tensor does not match the expected input layout, a copy operation is inserted to resolve like for non-configurable nodes\n\nWhat I don't understand is how this requires any advanced layout propagation, it should be at most a copy at the input and outputs of each configurable nodes.\n\nThankful if you can provide some insights to this!",
          "votes": 2,
          "replies": [
            {
              "id": 2474046,
              "postDate": "2023-10-08T21:21:43.403Z",
              "content": "<p>Input and output layout configuration for some node don't have to match. Output layout config from previous node (input node) must be the same as input layout for configurable node, otherwise copy operator is needed.</p>\n<p>Why \"propagation\" is  needed?<br>\nAssume you have two nodes A -&gt; B (A is input for B). And node A is configurable node. Is it important for node B what you choose as output layout config for A  {2,1,-1,-1,-1,-1} or {0,2,-1,-1,-1,-1}? For sure, it is. So, subsequent nodes must to know what choice you made, and that happens via \"layout propagation\"  </p>",
              "rawMarkdown": "Input and output layout configuration for some node don't have to match. Output layout config from previous node (input node) must be the same as input layout for configurable node, otherwise copy operator is needed.\n\nWhy \"propagation\" is  needed?\nAssume you have two nodes A -> B (A is input for B). And node A is configurable node. Is it important for node B what you choose as output layout config for A  {2,1,-1,-1,-1,-1} or {0,2,-1,-1,-1,-1}? For sure, it is. So, subsequent nodes must to know what choice you made, and that happens via \"layout propagation\"  ",
              "votes": 1
            },
            {
              "id": 2479529,
              "postDate": "2023-10-12T17:09:24.773Z",
              "content": "<p>Exactly, I believe you are 100% correct. And if the layout of an output of a configurable node (A) does not match the expected layout at a non-configurable (B), a copy operation is inserted (just like at the input of node A in case of mismatch). I am assuming that the last 6 node features describe both the expected input layout as well as the layout of the output tensor. So as far as I understand it, you only need to consider the immediate neighborhood around each configurable node, and optionally insert copy operations. But I am just speculating, would be great to get some input from organizers / read mode XLA code…</p>",
              "rawMarkdown": "Exactly, I believe you are 100% correct. And if the layout of an output of a configurable node (A) does not match the expected layout at a non-configurable (B), a copy operation is inserted (just like at the input of node A in case of mismatch). I am assuming that the last 6 node features describe both the expected input layout as well as the layout of the output tensor. So as far as I understand it, you only need to consider the immediate neighborhood around each configurable node, and optionally insert copy operations. But I am just speculating, would be great to get some input from organizers / read mode XLA code..."
            },
            {
              "id": 2479993,
              "postDate": "2023-10-13T03:57:42.903Z",
              "content": "<p>This is correct. However, there is no limit on how far a layout propagation will go. For example, if there is a long chain of element-wise operation nodes, then the layout propagation would propagate through the whole chain until it hits a node whose inputs and output don't have the same shape.</p>",
              "rawMarkdown": "This is correct. However, there is no limit on how far a layout propagation will go. For example, if there is a long chain of element-wise operation nodes, then the layout propagation would propagate through the whole chain until it hits a node whose inputs and output don't have the same shape."
            },
            {
              "id": 2480026,
              "postDate": "2023-10-13T05:07:10.660Z",
              "content": "<p>Thanks Mangpo, now I understand how this mechanism works!</p>",
              "rawMarkdown": "Thanks Mangpo, now I understand how this mechanism works!"
            },
            {
              "id": 2498980,
              "postDate": "2023-10-25T16:30:44.393Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2499004,
              "postDate": "2023-10-25T16:55:22.163Z",
              "content": "<p>Sorry to break in, but let me confirm what <code>layout_minor_to_major_&lt;N&gt;</code> is. It seems sometimes configurable nodes have non-zero <code>layout_minor_to_major_&lt;N&gt;</code>. In that case, <code>layout_minor_to_major_&lt;N&gt;</code> is overridden by config <code>input_layout_&lt;N&gt;</code> and <code>output_layout_&lt;N&gt;</code>?</p>",
              "rawMarkdown": "Sorry to break in, but let me confirm what `layout_minor_to_major_<N>` is. It seems sometimes configurable nodes have non-zero `layout_minor_to_major_<N>`. In that case, `layout_minor_to_major_<N>` is overridden by config `input_layout_<N>` and `output_layout_<N>`?",
              "votes": 1
            },
            {
              "id": 2501778,
              "postDate": "2023-10-27T16:43:15.760Z",
              "content": "<p><a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> Could you probably answer this…? I'm wondering why <code>layout_minor_to_major_&lt;N&gt;</code> exists for configurable nodes.</p>",
              "rawMarkdown": "@mangpophothilimthana Could you probably answer this...? I'm wondering why `layout_minor_to_major_<N>` exists for configurable nodes."
            },
            {
              "id": 2505910,
              "postDate": "2023-10-31T00:57:36.210Z",
              "content": "<p>I'm not sure either… I think some compiler passes may have decided the layouts for some nodes before the layout assignment pass. But I don't know which one takes priority to be honest. I would guess that <code>input_layout_&lt;N&gt;</code> and <code>output_layout_&lt;N&gt;</code> take priority.</p>",
              "rawMarkdown": "I'm not sure either... I think some compiler passes may have decided the layouts for some nodes before the layout assignment pass. But I don't know which one takes priority to be honest. I would guess that `input_layout_<N>` and `output_layout_<N>` take priority."
            },
            {
              "id": 2506490,
              "postDate": "2023-10-31T10:39:47.543Z",
              "content": "<p>Okay, thanks for your answer.</p>",
              "rawMarkdown": "Okay, thanks for your answer."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2465300,
      "author_name": "Mangpo Phothilimthana",
      "author_url": "",
      "post_date": "2023-10-02T20:07:31.617000",
      "content": "<p>The layout propagation algorithm in the compiler is rather complex, and I don't fully understand it myself (I didn't implement it). The high-level idea is that the compiler (1) selects a convolution, dot, or reshape node and assign the best layout for that node, and (2) propagate the layout from that node as far as possible. The compiler integrates between these two steps until all layouts are assigned. There are a bunch of rules on how a layout is propagated between inputs and output of a node. Some can be found <a href=\"https://github.com/openxla/xla/blob/main/xla/service/layout_assignment.cc\" target=\"_blank\">here</a>. For example, an element-wise node will assign the same layout for its inputs and output during propagation.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2473896,
          "author_name": "cdeln",
          "author_url": "",
          "post_date": "2023-10-08T16:38:40.957000",
          "content": "<p>How does this interact with the inserted copy operations? This is how I understand that it works (please correct me where I'm wrong) </p>\n<p><strong>For configurable nodes</strong></p>\n<ol>\n<li>Configurable nodes expect input tensor to be laid out as described by <code>input_layout_&lt;N&gt;</code> fields in the layout config</li>\n<li>A copy operation is inserted when the layout an input tensor does not match the config output layout</li>\n<li>After inserting the copy operation, the layouts match</li>\n<li>The layout of the output tensor of the configurable node is as described by the layout output config</li>\n</ol>\n<p><strong>For non-configurable nodes</strong></p>\n<ol>\n<li>Non-configurable nodes expect input tensors to be laid out as described by <code>layout_minor_to_major_&lt;N&gt;</code> fields</li>\n<li>Non-configurable nodes do dot change the tensor layout (so the output layout is the same as the input)</li>\n<li>If the input tensor does not match the expected input layout, a copy operation is inserted to resolve like for non-configurable nodes</li>\n</ol>\n<p>What I don't understand is how this requires any advanced layout propagation, it should be at most a copy at the input and outputs of each configurable nodes.</p>\n<p>Thankful if you can provide some insights to this!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2474046,
              "author_name": "Shekhovtsov Aleksandr",
              "author_url": "",
              "post_date": "2023-10-08T21:21:43.403000",
              "content": "<p>Input and output layout configuration for some node don't have to match. Output layout config from previous node (input node) must be the same as input layout for configurable node, otherwise copy operator is needed.</p>\n<p>Why \"propagation\" is  needed?<br>\nAssume you have two nodes A -&gt; B (A is input for B). And node A is configurable node. Is it important for node B what you choose as output layout config for A  {2,1,-1,-1,-1,-1} or {0,2,-1,-1,-1,-1}? For sure, it is. So, subsequent nodes must to know what choice you made, and that happens via \"layout propagation\"  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2479529,
              "author_name": "cdeln",
              "author_url": "",
              "post_date": "2023-10-12T17:09:24.773000",
              "content": "<p>Exactly, I believe you are 100% correct. And if the layout of an output of a configurable node (A) does not match the expected layout at a non-configurable (B), a copy operation is inserted (just like at the input of node A in case of mismatch). I am assuming that the last 6 node features describe both the expected input layout as well as the layout of the output tensor. So as far as I understand it, you only need to consider the immediate neighborhood around each configurable node, and optionally insert copy operations. But I am just speculating, would be great to get some input from organizers / read mode XLA code…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2479993,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-10-13T03:57:42.903000",
              "content": "<p>This is correct. However, there is no limit on how far a layout propagation will go. For example, if there is a long chain of element-wise operation nodes, then the layout propagation would propagate through the whole chain until it hits a node whose inputs and output don't have the same shape.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2480026,
              "author_name": "cdeln",
              "author_url": "",
              "post_date": "2023-10-13T05:07:10.660000",
              "content": "<p>Thanks Mangpo, now I understand how this mechanism works!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2498980,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-10-25T16:30:44.393000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2499004,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-10-25T16:55:22.163000",
              "content": "<p>Sorry to break in, but let me confirm what <code>layout_minor_to_major_&lt;N&gt;</code> is. It seems sometimes configurable nodes have non-zero <code>layout_minor_to_major_&lt;N&gt;</code>. In that case, <code>layout_minor_to_major_&lt;N&gt;</code> is overridden by config <code>input_layout_&lt;N&gt;</code> and <code>output_layout_&lt;N&gt;</code>?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2501778,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-10-27T16:43:15.760000",
              "content": "<p><a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> Could you probably answer this…? I'm wondering why <code>layout_minor_to_major_&lt;N&gt;</code> exists for configurable nodes.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2505910,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-10-31T00:57:36.210000",
              "content": "<p>I'm not sure either… I think some compiler passes may have decided the layouts for some nodes before the layout assignment pass. But I don't know which one takes priority to be honest. I would guess that <code>input_layout_&lt;N&gt;</code> and <code>output_layout_&lt;N&gt;</code> take priority.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2506490,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-10-31T10:39:47.543000",
              "content": "<p>Okay, thanks for your answer.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2456557": "I have question about layout configuration. In paper [TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs](https://arxiv.org/pdf/2308.13490.pdf)  the authors wrote \"The autotuner tunes the input-output layouts of the most layout-performance-critical nodes and **propagates layouts from these nodes to others**.\" And I heard the same in some videos with @mangpophothilimthana \nDoes anyone understand how exactly \"layout propagation\" works?\nEach (non-configurable) node has 6 features related to layout. Is it a layout for an input or output tensor?\nAnd after we change the layout configurations of the input and output tensors for some configurable node, what kind of propagation is required?\n\nThanks for the clarifications in advance",
    "2465300": "The layout propagation algorithm in the compiler is rather complex, and I don't fully understand it myself (I didn't implement it). The high-level idea is that the compiler (1) selects a convolution, dot, or reshape node and assign the best layout for that node, and (2) propagate the layout from that node as far as possible. The compiler integrates between these two steps until all layouts are assigned. There are a bunch of rules on how a layout is propagated between inputs and output of a node. Some can be found [here](https://github.com/openxla/xla/blob/main/xla/service/layout_assignment.cc). For example, an element-wise node will assign the same layout for its inputs and output during propagation."
  }
}