{
  "id": 437530,
  "title": "Share competition related resources",
  "url": "/competitions/predict-ai-model-runtime/discussion/437530",
  "author_name": "",
  "post_date": "2023-09-07T06:01:19.573349400Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey Guys,</p>\n<p>I am a beginner and this competition seems advance to me. But it's interesting so I'll continue this and gain the knowledge along the way. Please share resources which are prerequisite to this competition. Also explain all the technical terminology. Share whatever the knowledge you have regarding this task.</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "2427233",
      "postDate": "09/07/2023 06:01:19",
      "content": "<p>Hey Guys,</p>\n<p>I am a beginner and this competition seems advance to me. But it's interesting so I'll continue this and gain the knowledge along the way. Please share resources which are prerequisite to this competition. Also explain all the technical terminology. Share whatever the knowledge you have regarding this task.</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hey Guys,\n\nI am a beginner and this competition seems advance to me. But it's interesting so I'll continue this and gain the knowledge along the way. Please share resources which are prerequisite to this competition. Also explain all the technical terminology. Share whatever the knowledge you have regarding this task.\n\nThanks",
      "votes": null
    },
    {
      "id": "2427244",
      "postDate": "09/07/2023 06:05:58",
      "content": "<h2>Here is my notion page regarding this competition:</h2>\n<p>→ Each .npz file contains a graphs with “n” nodes and “m” edges.</p>\n<p>→ Suppose we compile the graph with “c” different configurations.</p>\n<p>→ <strong>XLA (Accelerated Linear Algebra)</strong>: is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes.</p>\n<hr>\n<h3>Tiles:</h3>\n<p><strong>Keys:</strong></p>\n<ol>\n<li>node_feat:<ol>\n<li>float - (n,140)</li></ol></li>\n<li>node_opcode:<ol>\n<li>int32 vector - (n, )</li></ol></li>\n<li>edge_index:<ol>\n<li>int32 - (m, 2)</li></ol></li>\n<li>config_feat:<ol>\n<li>float - (c, 24)</li></ol></li>\n<li>config_runtime:<ol>\n<li>int64 vector of length c</li></ol></li>\n<li>config_runtime_normalizers:<ol>\n<li>int64 vector of length c</li></ol></li>\n</ol>\n<p><strong>Goal:</strong></p>\n<p>For the tile collection, your job is to predict the indices of the best configurations (i.e., ones leading to the smallest&nbsp;<code>d[\"config_runtime\"] / d[\"config_runtime_normalizers\"]</code>).</p>\n<p><strong>Understanding Tiles:</strong></p>\n<p>Tile Configuration:</p>\n<p>For tile configuration, let's first understand tiling. Imagine an operation that needs to be performed on a large matrix (or tensor). Doing it all at once might be inefficient. Instead, what we do is \"tile\" or break the matrix into smaller chunks (tiles) and perform the operation on these smaller chunks individually or simultaneously (depending on the hardware).</p>\n<p>A tile configuration, therefore, determines the size of these tiles. In essence, how big each chunk or tile should be when we break up the tensor.</p>\n<p>Why is this important?</p>\n<p>Choosing the right tile size can significantly impact performance. A tile size that's too small might cause overhead due to the frequent need to load and store data, while a tile size that's too large might not fit well in the cache memory, causing cache misses.</p>\n<p>Key Point:&nbsp;The tile configuration allows Alice to specify the ideal tile size for each fused subgraph in the AI model. This can be crucial for ensuring the efficient use of memory and computational resources.</p>\n<hr>\n<h3>Layout:</h3>\n<p><strong>Keys:</strong></p>\n<ol>\n<li>node_feat:<ol>\n<li>same as above</li></ol></li>\n<li>node_opcode:<ol>\n<li>same as above</li></ol></li>\n<li>edge_index:<ol>\n<li>same as above</li></ol></li>\n<li>node_config_ids:<ol>\n<li>int32 vector - (nc, )</li></ol></li>\n<li>node_config_feat:<ol>\n<li>float32 - (c, nc, 18)</li></ol></li>\n<li>config_runtime:<ol>\n<li>int32 vector - (c, )</li></ol></li>\n</ol>\n<p><strong>Goal:</strong></p>\n<p>Finally, for the layout collections, your job is to predict the order of the indices from best-to-worse configurations (i.e., ones leading to the smallest&nbsp;<code>d[\"config_runtime\"]</code>). We do not have to use runtime normalizers for this task because the runtime variation at the entire program level is very small.</p>\n<p><strong>Understanding Layouts:</strong></p>\n<p>Layout Configuration:</p>\n<p>When we talk about the layout configuration, we are essentially discussing the manner in which tensors are organized in memory. Think of a tensor as a multi-dimensional array (similar to matrices). For instance, for image data in deep learning, we might represent it as a tensor with the dimensions&nbsp;<code>[Batch, Channels, Height, Width]</code>&nbsp;or in shorthand, <strong>BCHW</strong>.</p>\n<p>However, sometimes, due to architectural specificities or optimization purposes, we may want to rearrange this order. For example, a layout might rearrange the tensor dimensions as&nbsp;<code>[Batch, Height, Width, Channels]</code>&nbsp;or <strong>BHWC</strong>.</p>\n<p>Why does this matter?</p>\n<p>The order in which these dimensions are stored can significantly affect the speed of tensor operations. Certain operations might be faster on a BCHW layout, while others on a BHWC. This is due to how these operations access memory and the underlying hardware's capability.</p>\n<p>Key Point:&nbsp;A layout configuration lets Alice specify how the dimensions of a tensor are arranged in memory, which in turn, can optimize specific tensor operations.</p>\n<hr>\n<h3>Learning</h3>\n<ol>\n<li>Computational Graph Compilation</li>\n<li>XLA, HLO</li>\n<li>Competition Papers</li>\n<li>GraphSAGE</li>\n<li>Graph NN</li>\n<li>Domain Knowledge</li>\n<li>Data Understanding</li>\n<li>Related Videos: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436629\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436629</a></li>\n<li>Research Paper 1: <a href=\"https://www.cs.ubc.ca/~kevinlb/papers/2015-IJCAI-EPMs.pdf\" target=\"_blank\">https://www.cs.ubc.ca/~kevinlb/papers/2015-IJCAI-EPMs.pdf</a></li>\n<li>Research Paper 2: <a href=\"https://arxiv.org/pdf/2308.13490v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2308.13490v1.pdf</a></li>\n<li>Research Paper 3: <a href=\"https://sci-hub.se/https://ieeexplore.ieee.org/document/8924184\" target=\"_blank\">https://sci-hub.se/https://ieeexplore.ieee.org/document/8924184</a></li>\n<li>Research Paper 4: <a href=\"https://arxiv.org/pdf/2008.01040.pdf\" target=\"_blank\">https://arxiv.org/pdf/2008.01040.pdf</a></li>\n<li>Slides 1: <a href=\"https://mangpo.net/talks/2021-xla-autotuning-pact.pdf\" target=\"_blank\">https://mangpo.net/talks/2021-xla-autotuning-pact.pdf</a></li>\n<li>YT 1: <a href=\"https://www.youtube.com/watch?v=esD_zvAf49I\" target=\"_blank\">https://www.youtube.com/watch?v=esD_zvAf49I</a></li>\n<li>YT 2: <a href=\"https://www.youtube.com/watch?v=vYQ37pBNAFI\" target=\"_blank\">https://www.youtube.com/watch?v=vYQ37pBNAFI</a></li>\n<li>GNN with PyTorch: <a href=\"https://www.youtube.com/watch?v=-UjytpbqX4A\" target=\"_blank\">https://www.youtube.com/watch?v=-UjytpbqX4A</a></li>\n</ol>\n<hr>\n<h3>IR | LLVM | MLIR</h3>\n<p>Ref: <a href=\"https://www.youtube.com/watch?v=BT2Cv-Tjq7Q\" target=\"_blank\">https://www.youtube.com/watch?v=BT2Cv-Tjq7Q</a></p>\n<p><strong>IR</strong>: Intermediate Representation of Programming Language</p>\n<p><strong>LLVM</strong>: Converts programming language into IR</p>\n<p><strong>MLIR</strong>: Multi-level intermediate representation</p>\n<p>→ LLVM and MLIR are both valuable compiler infrastructure projects, but they have different scopes and focuses. LLVM excels in code generation and optimization for a wide range of target architectures, while MLIR offers flexibility and extensibility to represent and optimize code at various levels of abstraction, making it suitable for a broader set of use cases, including machine learning and domain-specific languages.</p>\n<p>Resource: <a href=\"https://www.youtube.com/playlist?list=PLlONLmJCfHTo9WYfsoQvwjsa5ZB6hjOG5\" target=\"_blank\">https://www.youtube.com/playlist?list=PLlONLmJCfHTo9WYfsoQvwjsa5ZB6hjOG5</a></p>\n<p><strong>Machine Learning Compiler:</strong></p>\n<p>Ref: <a href=\"https://www.youtube.com/watch?v=y7d7CWX_Is8\" target=\"_blank\">https://www.youtube.com/watch?v=y7d7CWX_Is8</a></p>\n<hr>\n<h3>Graph Neural Networks</h3>\n<p>Ref: <a href=\"https://www.youtube.com/playlist?list=PLoROMvodv4rPLKxIpqhjhPgdQy7imNkDn\" target=\"_blank\">https://www.youtube.com/playlist?list=PLoROMvodv4rPLKxIpqhjhPgdQy7imNkDn</a></p>",
      "rawMarkdown": "## Here is my notion page regarding this competition:\n\n\n→ Each .npz file contains a graphs with “n” nodes and “m” edges.\n\n→ Suppose we compile the graph with “c” different configurations.\n\n→ **XLA (Accelerated Linear Algebra)**: is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes.\n\n---\n\n### Tiles:\n\n**Keys:**\n\n1. node_feat:\n    1. float - (n,140)\n2. node_opcode:\n    1. int32 vector - (n, )\n3. edge_index:\n    1. int32 - (m, 2)\n4. config_feat:\n    1. float - (c, 24)\n5. config_runtime:\n    1. int64 vector of length c\n6. config_runtime_normalizers:\n    1. int64 vector of length c\n\n**Goal:**\n\nFor the tile collection, your job is to predict the indices of the best configurations (i.e., ones leading to the smallest `d[\"config_runtime\"] / d[\"config_runtime_normalizers\"]`).\n\n**Understanding Tiles:**\n\nTile Configuration:\n\nFor tile configuration, let's first understand tiling. Imagine an operation that needs to be performed on a large matrix (or tensor). Doing it all at once might be inefficient. Instead, what we do is \"tile\" or break the matrix into smaller chunks (tiles) and perform the operation on these smaller chunks individually or simultaneously (depending on the hardware).\n\nA tile configuration, therefore, determines the size of these tiles. In essence, how big each chunk or tile should be when we break up the tensor.\n\nWhy is this important?\n\nChoosing the right tile size can significantly impact performance. A tile size that's too small might cause overhead due to the frequent need to load and store data, while a tile size that's too large might not fit well in the cache memory, causing cache misses.\n\nKey Point: The tile configuration allows Alice to specify the ideal tile size for each fused subgraph in the AI model. This can be crucial for ensuring the efficient use of memory and computational resources.\n\n---\n\n### Layout:\n\n**Keys:**\n\n1. node_feat:\n    1. same as above\n2. node_opcode:\n    1. same as above\n3. edge_index:\n    1. same as above\n4. node_config_ids:\n    1. int32 vector - (nc, )\n5. node_config_feat:\n    1. float32 - (c, nc, 18)\n6. config_runtime:\n    1. int32 vector - (c, )\n\n**Goal:**\n\nFinally, for the layout collections, your job is to predict the order of the indices from best-to-worse configurations (i.e., ones leading to the smallest `d[\"config_runtime\"]`). We do not have to use runtime normalizers for this task because the runtime variation at the entire program level is very small.\n\n**Understanding Layouts:**\n\nLayout Configuration:\n\nWhen we talk about the layout configuration, we are essentially discussing the manner in which tensors are organized in memory. Think of a tensor as a multi-dimensional array (similar to matrices). For instance, for image data in deep learning, we might represent it as a tensor with the dimensions `[Batch, Channels, Height, Width]` or in shorthand, **BCHW**.\n\nHowever, sometimes, due to architectural specificities or optimization purposes, we may want to rearrange this order. For example, a layout might rearrange the tensor dimensions as `[Batch, Height, Width, Channels]` or **BHWC**.\n\nWhy does this matter?\n\nThe order in which these dimensions are stored can significantly affect the speed of tensor operations. Certain operations might be faster on a BCHW layout, while others on a BHWC. This is due to how these operations access memory and the underlying hardware's capability.\n\nKey Point: A layout configuration lets Alice specify how the dimensions of a tensor are arranged in memory, which in turn, can optimize specific tensor operations.\n\n---\n\n### Learning\n\n1. Computational Graph Compilation\n2. XLA, HLO\n3. Competition Papers\n4. GraphSAGE\n5. Graph NN\n6. Domain Knowledge\n7. Data Understanding\n8. Related Videos: https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436629\n9. Research Paper 1: https://www.cs.ubc.ca/~kevinlb/papers/2015-IJCAI-EPMs.pdf\n10. Research Paper 2: https://arxiv.org/pdf/2308.13490v1.pdf\n11. Research Paper 3: https://sci-hub.se/https://ieeexplore.ieee.org/document/8924184\n12. Research Paper 4: https://arxiv.org/pdf/2008.01040.pdf\n13. Slides 1: https://mangpo.net/talks/2021-xla-autotuning-pact.pdf\n14. YT 1: https://www.youtube.com/watch?v=esD_zvAf49I\n15. YT 2: https://www.youtube.com/watch?v=vYQ37pBNAFI\n16. GNN with PyTorch: https://www.youtube.com/watch?v=-UjytpbqX4A\n\n---\n\n### IR | LLVM | MLIR\n\nRef: https://www.youtube.com/watch?v=BT2Cv-Tjq7Q\n\n**IR**: Intermediate Representation of Programming Language\n\n**LLVM**: Converts programming language into IR\n\n**MLIR**: Multi-level intermediate representation\n\n→ LLVM and MLIR are both valuable compiler infrastructure projects, but they have different scopes and focuses. LLVM excels in code generation and optimization for a wide range of target architectures, while MLIR offers flexibility and extensibility to represent and optimize code at various levels of abstraction, making it suitable for a broader set of use cases, including machine learning and domain-specific languages.\n\nResource: https://www.youtube.com/playlist?list=PLlONLmJCfHTo9WYfsoQvwjsa5ZB6hjOG5\n\n**Machine Learning Compiler:**\n\nRef: https://www.youtube.com/watch?v=y7d7CWX_Is8\n\n---\n\n### Graph Neural Networks\n\nRef: https://www.youtube.com/playlist?list=PLoROMvodv4rPLKxIpqhjhPgdQy7imNkDn",
      "votes": null
    },
    {
      "id": "2427354",
      "postDate": "09/07/2023 07:15:40",
      "content": "<p>Thanks for sharing your resources </p>",
      "rawMarkdown": "Thanks for sharing your resources",
      "votes": null
    },
    {
      "id": "2440863",
      "postDate": "09/15/2023 19:35:27",
      "content": "<p>Is XLA complier knowledge a must for this competition ?</p>",
      "rawMarkdown": "Is XLA complier knowledge a must for this competition ?",
      "votes": null
    },
    {
      "id": "2443457",
      "postDate": "09/17/2023 18:06:45",
      "content": "<p>Not at all.</p>",
      "rawMarkdown": "Not at all.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2427244,
      "author_name": "hussamcheema",
      "author_url": "",
      "post_date": "09/07/2023 06:05:58",
      "content": "<h2>Here is my notion page regarding this competition:</h2>\n<p>→ Each .npz file contains a graphs with “n” nodes and “m” edges.</p>\n<p>→ Suppose we compile the graph with “c” different configurations.</p>\n<p>→ <strong>XLA (Accelerated Linear Algebra)</strong>: is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes.</p>\n<hr>\n<h3>Tiles:</h3>\n<p><strong>Keys:</strong></p>\n<ol>\n<li>node_feat:<ol>\n<li>float - (n,140)</li></ol></li>\n<li>node_opcode:<ol>\n<li>int32 vector - (n, )</li></ol></li>\n<li>edge_index:<ol>\n<li>int32 - (m, 2)</li></ol></li>\n<li>config_feat:<ol>\n<li>float - (c, 24)</li></ol></li>\n<li>config_runtime:<ol>\n<li>int64 vector of length c</li></ol></li>\n<li>config_runtime_normalizers:<ol>\n<li>int64 vector of length c</li></ol></li>\n</ol>\n<p><strong>Goal:</strong></p>\n<p>For the tile collection, your job is to predict the indices of the best configurations (i.e., ones leading to the smallest&nbsp;<code>d[\"config_runtime\"] / d[\"config_runtime_normalizers\"]</code>).</p>\n<p><strong>Understanding Tiles:</strong></p>\n<p>Tile Configuration:</p>\n<p>For tile configuration, let's first understand tiling. Imagine an operation that needs to be performed on a large matrix (or tensor). Doing it all at once might be inefficient. Instead, what we do is \"tile\" or break the matrix into smaller chunks (tiles) and perform the operation on these smaller chunks individually or simultaneously (depending on the hardware).</p>\n<p>A tile configuration, therefore, determines the size of these tiles. In essence, how big each chunk or tile should be when we break up the tensor.</p>\n<p>Why is this important?</p>\n<p>Choosing the right tile size can significantly impact performance. A tile size that's too small might cause overhead due to the frequent need to load and store data, while a tile size that's too large might not fit well in the cache memory, causing cache misses.</p>\n<p>Key Point:&nbsp;The tile configuration allows Alice to specify the ideal tile size for each fused subgraph in the AI model. This can be crucial for ensuring the efficient use of memory and computational resources.</p>\n<hr>\n<h3>Layout:</h3>\n<p><strong>Keys:</strong></p>\n<ol>\n<li>node_feat:<ol>\n<li>same as above</li></ol></li>\n<li>node_opcode:<ol>\n<li>same as above</li></ol></li>\n<li>edge_index:<ol>\n<li>same as above</li></ol></li>\n<li>node_config_ids:<ol>\n<li>int32 vector - (nc, )</li></ol></li>\n<li>node_config_feat:<ol>\n<li>float32 - (c, nc, 18)</li></ol></li>\n<li>config_runtime:<ol>\n<li>int32 vector - (c, )</li></ol></li>\n</ol>\n<p><strong>Goal:</strong></p>\n<p>Finally, for the layout collections, your job is to predict the order of the indices from best-to-worse configurations (i.e., ones leading to the smallest&nbsp;<code>d[\"config_runtime\"]</code>). We do not have to use runtime normalizers for this task because the runtime variation at the entire program level is very small.</p>\n<p><strong>Understanding Layouts:</strong></p>\n<p>Layout Configuration:</p>\n<p>When we talk about the layout configuration, we are essentially discussing the manner in which tensors are organized in memory. Think of a tensor as a multi-dimensional array (similar to matrices). For instance, for image data in deep learning, we might represent it as a tensor with the dimensions&nbsp;<code>[Batch, Channels, Height, Width]</code>&nbsp;or in shorthand, <strong>BCHW</strong>.</p>\n<p>However, sometimes, due to architectural specificities or optimization purposes, we may want to rearrange this order. For example, a layout might rearrange the tensor dimensions as&nbsp;<code>[Batch, Height, Width, Channels]</code>&nbsp;or <strong>BHWC</strong>.</p>\n<p>Why does this matter?</p>\n<p>The order in which these dimensions are stored can significantly affect the speed of tensor operations. Certain operations might be faster on a BCHW layout, while others on a BHWC. This is due to how these operations access memory and the underlying hardware's capability.</p>\n<p>Key Point:&nbsp;A layout configuration lets Alice specify how the dimensions of a tensor are arranged in memory, which in turn, can optimize specific tensor operations.</p>\n<hr>\n<h3>Learning</h3>\n<ol>\n<li>Computational Graph Compilation</li>\n<li>XLA, HLO</li>\n<li>Competition Papers</li>\n<li>GraphSAGE</li>\n<li>Graph NN</li>\n<li>Domain Knowledge</li>\n<li>Data Understanding</li>\n<li>Related Videos: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436629\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436629</a></li>\n<li>Research Paper 1: <a href=\"https://www.cs.ubc.ca/~kevinlb/papers/2015-IJCAI-EPMs.pdf\" target=\"_blank\">https://www.cs.ubc.ca/~kevinlb/papers/2015-IJCAI-EPMs.pdf</a></li>\n<li>Research Paper 2: <a href=\"https://arxiv.org/pdf/2308.13490v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2308.13490v1.pdf</a></li>\n<li>Research Paper 3: <a href=\"https://sci-hub.se/https://ieeexplore.ieee.org/document/8924184\" target=\"_blank\">https://sci-hub.se/https://ieeexplore.ieee.org/document/8924184</a></li>\n<li>Research Paper 4: <a href=\"https://arxiv.org/pdf/2008.01040.pdf\" target=\"_blank\">https://arxiv.org/pdf/2008.01040.pdf</a></li>\n<li>Slides 1: <a href=\"https://mangpo.net/talks/2021-xla-autotuning-pact.pdf\" target=\"_blank\">https://mangpo.net/talks/2021-xla-autotuning-pact.pdf</a></li>\n<li>YT 1: <a href=\"https://www.youtube.com/watch?v=esD_zvAf49I\" target=\"_blank\">https://www.youtube.com/watch?v=esD_zvAf49I</a></li>\n<li>YT 2: <a href=\"https://www.youtube.com/watch?v=vYQ37pBNAFI\" target=\"_blank\">https://www.youtube.com/watch?v=vYQ37pBNAFI</a></li>\n<li>GNN with PyTorch: <a href=\"https://www.youtube.com/watch?v=-UjytpbqX4A\" target=\"_blank\">https://www.youtube.com/watch?v=-UjytpbqX4A</a></li>\n</ol>\n<hr>\n<h3>IR | LLVM | MLIR</h3>\n<p>Ref: <a href=\"https://www.youtube.com/watch?v=BT2Cv-Tjq7Q\" target=\"_blank\">https://www.youtube.com/watch?v=BT2Cv-Tjq7Q</a></p>\n<p><strong>IR</strong>: Intermediate Representation of Programming Language</p>\n<p><strong>LLVM</strong>: Converts programming language into IR</p>\n<p><strong>MLIR</strong>: Multi-level intermediate representation</p>\n<p>→ LLVM and MLIR are both valuable compiler infrastructure projects, but they have different scopes and focuses. LLVM excels in code generation and optimization for a wide range of target architectures, while MLIR offers flexibility and extensibility to represent and optimize code at various levels of abstraction, making it suitable for a broader set of use cases, including machine learning and domain-specific languages.</p>\n<p>Resource: <a href=\"https://www.youtube.com/playlist?list=PLlONLmJCfHTo9WYfsoQvwjsa5ZB6hjOG5\" target=\"_blank\">https://www.youtube.com/playlist?list=PLlONLmJCfHTo9WYfsoQvwjsa5ZB6hjOG5</a></p>\n<p><strong>Machine Learning Compiler:</strong></p>\n<p>Ref: <a href=\"https://www.youtube.com/watch?v=y7d7CWX_Is8\" target=\"_blank\">https://www.youtube.com/watch?v=y7d7CWX_Is8</a></p>\n<hr>\n<h3>Graph Neural Networks</h3>\n<p>Ref: <a href=\"https://www.youtube.com/playlist?list=PLoROMvodv4rPLKxIpqhjhPgdQy7imNkDn\" target=\"_blank\">https://www.youtube.com/playlist?list=PLoROMvodv4rPLKxIpqhjhPgdQy7imNkDn</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2427354,
      "author_name": "faysalmiah1721758",
      "author_url": "",
      "post_date": "09/07/2023 07:15:40",
      "content": "<p>Thanks for sharing your resources </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2440863,
      "author_name": "riteshbhalerao",
      "author_url": "",
      "post_date": "09/15/2023 19:35:27",
      "content": "<p>Is XLA complier knowledge a must for this competition ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2443457,
          "author_name": "mangpophothilimthana",
          "author_url": "",
          "post_date": "09/17/2023 18:06:45",
          "content": "<p>Not at all.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2427233": "Hey Guys,\n\nI am a beginner and this competition seems advance to me. But it's interesting so I'll continue this and gain the knowledge along the way. Please share resources which are prerequisite to this competition. Also explain all the technical terminology. Share whatever the knowledge you have regarding this task.\n\nThanks",
    "2427244": "## Here is my notion page regarding this competition:\n\n\n→ Each .npz file contains a graphs with “n” nodes and “m” edges.\n\n→ Suppose we compile the graph with “c” different configurations.\n\n→ **XLA (Accelerated Linear Algebra)**: is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes.\n\n---\n\n### Tiles:\n\n**Keys:**\n\n1. node_feat:\n    1. float - (n,140)\n2. node_opcode:\n    1. int32 vector - (n, )\n3. edge_index:\n    1. int32 - (m, 2)\n4. config_feat:\n    1. float - (c, 24)\n5. config_runtime:\n    1. int64 vector of length c\n6. config_runtime_normalizers:\n    1. int64 vector of length c\n\n**Goal:**\n\nFor the tile collection, your job is to predict the indices of the best configurations (i.e., ones leading to the smallest `d[\"config_runtime\"] / d[\"config_runtime_normalizers\"]`).\n\n**Understanding Tiles:**\n\nTile Configuration:\n\nFor tile configuration, let's first understand tiling. Imagine an operation that needs to be performed on a large matrix (or tensor). Doing it all at once might be inefficient. Instead, what we do is \"tile\" or break the matrix into smaller chunks (tiles) and perform the operation on these smaller chunks individually or simultaneously (depending on the hardware).\n\nA tile configuration, therefore, determines the size of these tiles. In essence, how big each chunk or tile should be when we break up the tensor.\n\nWhy is this important?\n\nChoosing the right tile size can significantly impact performance. A tile size that's too small might cause overhead due to the frequent need to load and store data, while a tile size that's too large might not fit well in the cache memory, causing cache misses.\n\nKey Point: The tile configuration allows Alice to specify the ideal tile size for each fused subgraph in the AI model. This can be crucial for ensuring the efficient use of memory and computational resources.\n\n---\n\n### Layout:\n\n**Keys:**\n\n1. node_feat:\n    1. same as above\n2. node_opcode:\n    1. same as above\n3. edge_index:\n    1. same as above\n4. node_config_ids:\n    1. int32 vector - (nc, )\n5. node_config_feat:\n    1. float32 - (c, nc, 18)\n6. config_runtime:\n    1. int32 vector - (c, )\n\n**Goal:**\n\nFinally, for the layout collections, your job is to predict the order of the indices from best-to-worse configurations (i.e., ones leading to the smallest `d[\"config_runtime\"]`). We do not have to use runtime normalizers for this task because the runtime variation at the entire program level is very small.\n\n**Understanding Layouts:**\n\nLayout Configuration:\n\nWhen we talk about the layout configuration, we are essentially discussing the manner in which tensors are organized in memory. Think of a tensor as a multi-dimensional array (similar to matrices). For instance, for image data in deep learning, we might represent it as a tensor with the dimensions `[Batch, Channels, Height, Width]` or in shorthand, **BCHW**.\n\nHowever, sometimes, due to architectural specificities or optimization purposes, we may want to rearrange this order. For example, a layout might rearrange the tensor dimensions as `[Batch, Height, Width, Channels]` or **BHWC**.\n\nWhy does this matter?\n\nThe order in which these dimensions are stored can significantly affect the speed of tensor operations. Certain operations might be faster on a BCHW layout, while others on a BHWC. This is due to how these operations access memory and the underlying hardware's capability.\n\nKey Point: A layout configuration lets Alice specify how the dimensions of a tensor are arranged in memory, which in turn, can optimize specific tensor operations.\n\n---\n\n### Learning\n\n1. Computational Graph Compilation\n2. XLA, HLO\n3. Competition Papers\n4. GraphSAGE\n5. Graph NN\n6. Domain Knowledge\n7. Data Understanding\n8. Related Videos: https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436629\n9. Research Paper 1: https://www.cs.ubc.ca/~kevinlb/papers/2015-IJCAI-EPMs.pdf\n10. Research Paper 2: https://arxiv.org/pdf/2308.13490v1.pdf\n11. Research Paper 3: https://sci-hub.se/https://ieeexplore.ieee.org/document/8924184\n12. Research Paper 4: https://arxiv.org/pdf/2008.01040.pdf\n13. Slides 1: https://mangpo.net/talks/2021-xla-autotuning-pact.pdf\n14. YT 1: https://www.youtube.com/watch?v=esD_zvAf49I\n15. YT 2: https://www.youtube.com/watch?v=vYQ37pBNAFI\n16. GNN with PyTorch: https://www.youtube.com/watch?v=-UjytpbqX4A\n\n---\n\n### IR | LLVM | MLIR\n\nRef: https://www.youtube.com/watch?v=BT2Cv-Tjq7Q\n\n**IR**: Intermediate Representation of Programming Language\n\n**LLVM**: Converts programming language into IR\n\n**MLIR**: Multi-level intermediate representation\n\n→ LLVM and MLIR are both valuable compiler infrastructure projects, but they have different scopes and focuses. LLVM excels in code generation and optimization for a wide range of target architectures, while MLIR offers flexibility and extensibility to represent and optimize code at various levels of abstraction, making it suitable for a broader set of use cases, including machine learning and domain-specific languages.\n\nResource: https://www.youtube.com/playlist?list=PLlONLmJCfHTo9WYfsoQvwjsa5ZB6hjOG5\n\n**Machine Learning Compiler:**\n\nRef: https://www.youtube.com/watch?v=y7d7CWX_Is8\n\n---\n\n### Graph Neural Networks\n\nRef: https://www.youtube.com/playlist?list=PLoROMvodv4rPLKxIpqhjhPgdQy7imNkDn",
    "2427354": "Thanks for sharing your resources",
    "2440863": "Is XLA complier knowledge a must for this competition ?",
    "2443457": "Not at all."
  },
  "source": "meta"
}