{
  "id": 436294,
  "title": "Understanding the problem",
  "url": "/competitions/predict-ai-model-runtime/discussion/436294",
  "author_name": "",
  "post_date": "2023-09-01T18:30:27.045654Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>After reading the whole problem, I will just lay out my understanding of the problem. If I misunderstood and missed something please do correct me up. I want to understand the whole data and formulate how to model it.</p>\n<p>We have 2 parts-</p>\n<h2>Layout Configuration</h2>\n<p>In layout, we determine the ordering of tensor dimensions in memory, optimizing for efficient tensor operations.</p>\n<p>We have 7 different types of information in the .npz files for layout.</p>\n<pre><code>[,\n ,\n ,\n ,\n ,\n ,\n ] \n</code></pre>\n<p><code>node_feat</code>: Node Features of fixed dimension 140.<br>\n<code>node_opcode</code>: Opcode for each of the nodes available.<br>\n<code>edge_index</code>: Directed edge from one node to another.<br>\n<code>node_config_feat</code>: Config features of fixed dimension 18.<br>\n<code>node_config_ids</code>: Not all the nodes are configurable, out of all the nodes available it shows the index of the nodes which are configurable (node index 0, 1, 2, .., n)<br>\n<code>config_runtime</code>: Runtime of each configuration in nanoseconds.</p>\n<p>Let's take one file in layout .npz file and its structure-</p>\n<p><code>node_feat</code>: (876, 140)<br>\n<code>node_opcode</code>: (876,)<br>\n<code>edge_index</code>: (1366, 2)<br>\n<code>node_config_feat</code>: (100040, 40, 18)<br>\n<code>node_config_ids</code>: (40,)<br>\n<code>config_runtime</code>: (100040,)</p>\n<p><strong>NOTE:</strong></p>\n<ul>\n<li>I think not all the nodes available are configurable.</li>\n<li>Not all the files have the same <code>config_runtime</code>. Different .npz files have different <code>config_runtime</code>. </li>\n<li><code>node_feat</code> always has 140 features and <code>node_config_feat</code> always has 18 features for every file.</li>\n<li>There are 120 opcodes available.</li>\n</ul>\n<h2>Tile Configuration</h2>\n<p>In tiles, we set the optimal tile sizes for performing operations on tensor subgraphs, optimizing memory usage and computational speed.</p>\n<p>We have 6 different types of information in the .npz files for layout.</p>\n<pre><code>[,\n ,\n ,\n ,\n ,\n ]\n</code></pre>\n<p>Let's take one file in tile .npz file and its structure-</p>\n<p>'node_feat': (31, 140)<br>\n'node_opcode': (31,)<br>\n'edge_index': (33, 2)<br>\n'config_feat': (1537, 24)<br>\n'config_runtime': (1537,)<br>\n'config_runtime_normalizers': (1537,)</p>\n<p><strong>NOTE:</strong></p>\n<ul>\n<li>I think not all the nodes available are configurable.</li>\n<li>Not all the files have the same <code>config_runtime</code>. Different .npz files have different <code>config_runtime</code>. </li>\n<li><code>node_feat</code> always has 140 features same as before and <code>config_feat</code> always has 24 features for every file.</li>\n<li>There are 120 opcodes available.</li>\n</ul>\n<p>I understand the <code>config_runtime</code> is the target for layout and <code>config_runtime/config_runtime_normalizers</code> for tile.<br>\nHow to formulate the data and how to model the data, I mean what approach can we take to prepare the data and modeling techniques? </p>\n<p>Any insights and information are helpful.</p>",
  "messages": [
    {
      "id": "2419170",
      "postDate": "09/01/2023 18:30:27",
      "content": "<p>After reading the whole problem, I will just lay out my understanding of the problem. If I misunderstood and missed something please do correct me up. I want to understand the whole data and formulate how to model it.</p>\n<p>We have 2 parts-</p>\n<h2>Layout Configuration</h2>\n<p>In layout, we determine the ordering of tensor dimensions in memory, optimizing for efficient tensor operations.</p>\n<p>We have 7 different types of information in the .npz files for layout.</p>\n<pre><code>[,\n ,\n ,\n ,\n ,\n ,\n ] \n</code></pre>\n<p><code>node_feat</code>: Node Features of fixed dimension 140.<br>\n<code>node_opcode</code>: Opcode for each of the nodes available.<br>\n<code>edge_index</code>: Directed edge from one node to another.<br>\n<code>node_config_feat</code>: Config features of fixed dimension 18.<br>\n<code>node_config_ids</code>: Not all the nodes are configurable, out of all the nodes available it shows the index of the nodes which are configurable (node index 0, 1, 2, .., n)<br>\n<code>config_runtime</code>: Runtime of each configuration in nanoseconds.</p>\n<p>Let's take one file in layout .npz file and its structure-</p>\n<p><code>node_feat</code>: (876, 140)<br>\n<code>node_opcode</code>: (876,)<br>\n<code>edge_index</code>: (1366, 2)<br>\n<code>node_config_feat</code>: (100040, 40, 18)<br>\n<code>node_config_ids</code>: (40,)<br>\n<code>config_runtime</code>: (100040,)</p>\n<p><strong>NOTE:</strong></p>\n<ul>\n<li>I think not all the nodes available are configurable.</li>\n<li>Not all the files have the same <code>config_runtime</code>. Different .npz files have different <code>config_runtime</code>. </li>\n<li><code>node_feat</code> always has 140 features and <code>node_config_feat</code> always has 18 features for every file.</li>\n<li>There are 120 opcodes available.</li>\n</ul>\n<h2>Tile Configuration</h2>\n<p>In tiles, we set the optimal tile sizes for performing operations on tensor subgraphs, optimizing memory usage and computational speed.</p>\n<p>We have 6 different types of information in the .npz files for layout.</p>\n<pre><code>[,\n ,\n ,\n ,\n ,\n ]\n</code></pre>\n<p>Let's take one file in tile .npz file and its structure-</p>\n<p>'node_feat': (31, 140)<br>\n'node_opcode': (31,)<br>\n'edge_index': (33, 2)<br>\n'config_feat': (1537, 24)<br>\n'config_runtime': (1537,)<br>\n'config_runtime_normalizers': (1537,)</p>\n<p><strong>NOTE:</strong></p>\n<ul>\n<li>I think not all the nodes available are configurable.</li>\n<li>Not all the files have the same <code>config_runtime</code>. Different .npz files have different <code>config_runtime</code>. </li>\n<li><code>node_feat</code> always has 140 features same as before and <code>config_feat</code> always has 24 features for every file.</li>\n<li>There are 120 opcodes available.</li>\n</ul>\n<p>I understand the <code>config_runtime</code> is the target for layout and <code>config_runtime/config_runtime_normalizers</code> for tile.<br>\nHow to formulate the data and how to model the data, I mean what approach can we take to prepare the data and modeling techniques? </p>\n<p>Any insights and information are helpful.</p>",
      "rawMarkdown": "After reading the whole problem, I will just lay out my understanding of the problem. If I misunderstood and missed something please do correct me up. I want to understand the whole data and formulate how to model it.\n\nWe have 2 parts-\n\n## Layout Configuration\n\nIn layout, we determine the ordering of tensor dimensions in memory, optimizing for efficient tensor operations.\n\nWe have 7 different types of information in the .npz files for layout.\n\n```python\n['node_feat',\n 'node_opcode',\n 'edge_index',\n 'node_config_feat',\n 'node_config_ids',\n 'config_runtime',\n 'node_splits'] \n```\n`node_feat`: Node Features of fixed dimension 140.\n`node_opcode`: Opcode for each of the nodes available.\n`edge_index`: Directed edge from one node to another.\n`node_config_feat`: Config features of fixed dimension 18.\n`node_config_ids`: Not all the nodes are configurable, out of all the nodes available it shows the index of the nodes which are configurable (node index 0, 1, 2, .., n)\n`config_runtime`: Runtime of each configuration in nanoseconds.\n\nLet's take one file in layout .npz file and its structure-\n\n`node_feat`: (876, 140)\n`node_opcode`: (876,)\n`edge_index`: (1366, 2)\n`node_config_feat`: (100040, 40, 18)\n`node_config_ids`: (40,)\n`config_runtime`: (100040,)\n\n**NOTE:**\n- I think not all the nodes available are configurable.\n- Not all the files have the same `config_runtime`. Different .npz files have different `config_runtime`. \n- `node_feat` always has 140 features and `node_config_feat` always has 18 features for every file.\n- There are 120 opcodes available.\n\n\n## Tile Configuration\n\nIn tiles, we set the optimal tile sizes for performing operations on tensor subgraphs, optimizing memory usage and computational speed.\n\nWe have 6 different types of information in the .npz files for layout.\n\n```python\n['node_feat',\n 'node_opcode',\n 'edge_index',\n 'config_feat',\n 'config_runtime',\n 'config_runtime_normalizers']\n```\nLet's take one file in tile .npz file and its structure-\n\n'node_feat': (31, 140)\n'node_opcode': (31,)\n'edge_index': (33, 2)\n'config_feat': (1537, 24)\n'config_runtime': (1537,)\n'config_runtime_normalizers': (1537,)\n\n**NOTE:**\n- I think not all the nodes available are configurable.\n- Not all the files have the same `config_runtime`. Different .npz files have different `config_runtime`. \n- `node_feat` always has 140 features same as before and `config_feat` always has 24 features for every file.\n- There are 120 opcodes available.\n\nI understand the `config_runtime` is the target for layout and `config_runtime/config_runtime_normalizers` for tile.\nHow to formulate the data and how to model the data, I mean what approach can we take to prepare the data and modeling techniques? \n\nAny insights and information are helpful.",
      "votes": null
    },
    {
      "id": "2422128",
      "postDate": "09/03/2023 18:09:40",
      "content": "<p>These discussions may help you understanding the data format better:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436214\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436214</a><br>\n<a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436204\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436204</a></p>\n<blockquote>\n  <p>What approach can we take to prepare the data and modeling techniques?<br>\n  We provide baseline models along with a training set up readily for you to get started at <a href=\"https://github.com/google-research-datasets/tpu_graphs\" target=\"_blank\">https://github.com/google-research-datasets/tpu_graphs</a>. You may also find some of the shared notebooks useful: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/code\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/code</a>.</p>\n</blockquote>",
      "rawMarkdown": "These discussions may help you understanding the data format better:\nhttps://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436214\nhttps://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436204\n\n> What approach can we take to prepare the data and modeling techniques?\nWe provide baseline models along with a training set up readily for you to get started at https://github.com/google-research-datasets/tpu_graphs. You may also find some of the shared notebooks useful: https://www.kaggle.com/competitions/predict-ai-model-runtime/code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2422128,
      "author_name": "mangpophothilimthana",
      "author_url": "",
      "post_date": "09/03/2023 18:09:40",
      "content": "<p>These discussions may help you understanding the data format better:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436214\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436214</a><br>\n<a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436204\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436204</a></p>\n<blockquote>\n  <p>What approach can we take to prepare the data and modeling techniques?<br>\n  We provide baseline models along with a training set up readily for you to get started at <a href=\"https://github.com/google-research-datasets/tpu_graphs\" target=\"_blank\">https://github.com/google-research-datasets/tpu_graphs</a>. You may also find some of the shared notebooks useful: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/code\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/code</a>.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2419170": "After reading the whole problem, I will just lay out my understanding of the problem. If I misunderstood and missed something please do correct me up. I want to understand the whole data and formulate how to model it.\n\nWe have 2 parts-\n\n## Layout Configuration\n\nIn layout, we determine the ordering of tensor dimensions in memory, optimizing for efficient tensor operations.\n\nWe have 7 different types of information in the .npz files for layout.\n\n```python\n['node_feat',\n 'node_opcode',\n 'edge_index',\n 'node_config_feat',\n 'node_config_ids',\n 'config_runtime',\n 'node_splits'] \n```\n`node_feat`: Node Features of fixed dimension 140.\n`node_opcode`: Opcode for each of the nodes available.\n`edge_index`: Directed edge from one node to another.\n`node_config_feat`: Config features of fixed dimension 18.\n`node_config_ids`: Not all the nodes are configurable, out of all the nodes available it shows the index of the nodes which are configurable (node index 0, 1, 2, .., n)\n`config_runtime`: Runtime of each configuration in nanoseconds.\n\nLet's take one file in layout .npz file and its structure-\n\n`node_feat`: (876, 140)\n`node_opcode`: (876,)\n`edge_index`: (1366, 2)\n`node_config_feat`: (100040, 40, 18)\n`node_config_ids`: (40,)\n`config_runtime`: (100040,)\n\n**NOTE:**\n- I think not all the nodes available are configurable.\n- Not all the files have the same `config_runtime`. Different .npz files have different `config_runtime`. \n- `node_feat` always has 140 features and `node_config_feat` always has 18 features for every file.\n- There are 120 opcodes available.\n\n\n## Tile Configuration\n\nIn tiles, we set the optimal tile sizes for performing operations on tensor subgraphs, optimizing memory usage and computational speed.\n\nWe have 6 different types of information in the .npz files for layout.\n\n```python\n['node_feat',\n 'node_opcode',\n 'edge_index',\n 'config_feat',\n 'config_runtime',\n 'config_runtime_normalizers']\n```\nLet's take one file in tile .npz file and its structure-\n\n'node_feat': (31, 140)\n'node_opcode': (31,)\n'edge_index': (33, 2)\n'config_feat': (1537, 24)\n'config_runtime': (1537,)\n'config_runtime_normalizers': (1537,)\n\n**NOTE:**\n- I think not all the nodes available are configurable.\n- Not all the files have the same `config_runtime`. Different .npz files have different `config_runtime`. \n- `node_feat` always has 140 features same as before and `config_feat` always has 24 features for every file.\n- There are 120 opcodes available.\n\nI understand the `config_runtime` is the target for layout and `config_runtime/config_runtime_normalizers` for tile.\nHow to formulate the data and how to model the data, I mean what approach can we take to prepare the data and modeling techniques? \n\nAny insights and information are helpful.",
    "2422128": "These discussions may help you understanding the data format better:\nhttps://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436214\nhttps://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/436204\n\n> What approach can we take to prepare the data and modeling techniques?\nWe provide baseline models along with a training set up readily for you to get started at https://github.com/google-research-datasets/tpu_graphs. You may also find some of the shared notebooks useful: https://www.kaggle.com/competitions/predict-ai-model-runtime/code."
  },
  "source": "meta"
}