{
  "id": 436680,
  "title": "Explanations of Data Folder Structure",
  "url": "/competitions/predict-ai-model-runtime/discussion/436680",
  "author_name": "",
  "post_date": "2023-09-03T14:12:02.401184800Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>It's the first time I step into the field of ML compilers. At first glance, it's too much information to digest, so I go through the amazing <a href=\"https://arxiv.org/abs/2308.13490\" target=\"_blank\">paper</a> provided by the host, which clearly explains how the data is generated and arranged. I decide to record and share some of my understandings about the data here.</p>\n<p>In short, the hierarchy is defined as:</p>\n<pre><code>optimization---\n</code></pre>\n<p>That is, collections of data differ in terms of:</p>\n<h3>1. <em>optimization</em> - The compiler optimization</h3>\n<ul>\n<li><code>layout</code>: The tensor computational graph is the input graph to the <strong>layoyt assignment pass</strong>.</li>\n<li><code>tile</code>: The fused subgraph (<strong>kernel</strong>) will goes through <strong>tile size selection</strong>.</li>\n</ul>\n<h3>2. <em>source</em> - The source of graphs (graph collection)<br></h3>\n<p>There are two sources from which the computational graphs are collected:</p>\n<ul>\n<li><code>xla</code>: The combination of the XLA regression benchmark.</li>\n<li><code>nlp</code>: A variety of BERT for training and inference.</li>\n</ul>\n<h3>3. <em>search</em> - The search strategy (configuration generation)<br></h3>\n<p>The XLA autotuner is used to generate the configurations, and there are two strategies to explore the search space:</p>\n<ul>\n<li><code>default</code>: Explore the search space using a <strong>genetic algorithm</strong> starting from the default configuration.</li>\n<li><code>random</code>: Pick random candidates in the search space.</li>\n</ul>\n<h3>4. <em>split</em> - The dataset split</h3>\n<p>As usual, the dataset contains <code>train</code>, <code>valid</code>, and <code>test</code> sets.</p>\n<p>With the information above, let's take a look at the data folder structure. Some observations are summarized as follows:</p>\n<ul>\n<li><code>tile</code> optimization contains graphs only from <code>xla</code> source.</li>\n<li>There's only one search strategy for <code>tile</code> optimization (at the level of the <strong>fused subgraph</strong>).<ul>\n<li><strong>Graph-level</strong> optimization is run beforehand to output a collection of kernels.</li>\n<li>Configurations are generated by <strong>enumerating</strong> all possible tile sizes for the kernel in a random order.  </li></ul></li>\n</ul>\n<pre><code>/kaggle//predict-ai-model-runtime/npz_all/npz\n├── layout\n│   ├── nlp\n│   │   ├── \n│   │   │   ├── test\n│   │   │   ├── train\n│   │   │   └── \n│   │   └── random\n│   │       ├── test\n│   │       ├── train\n│   │       └── \n│   └── xla\n│       ├── \n│       │   ├── test\n│       │   ├── train\n│       │   └── \n│       └── random\n│           ├── test\n│           ├── train\n│           └── \n└── tile\n    └── xla\n        ├── test\n        ├── train\n        └── \n</code></pre>\n<p>For detailed EDA, please see <a href=\"https://www.kaggle.com/code/abaojiang/google-fast-or-slow-detailed-eda\" target=\"_blank\">Google - Fast or Slow? - Detailed EDA</a>.</p>\n<p>Hope this helps, thanks.</p>",
  "messages": [
    {
      "id": "2421774",
      "postDate": "09/03/2023 14:12:02",
      "content": "<p>Hi everyone,</p>\n<p>It's the first time I step into the field of ML compilers. At first glance, it's too much information to digest, so I go through the amazing <a href=\"https://arxiv.org/abs/2308.13490\" target=\"_blank\">paper</a> provided by the host, which clearly explains how the data is generated and arranged. I decide to record and share some of my understandings about the data here.</p>\n<p>In short, the hierarchy is defined as:</p>\n<pre><code>optimization---\n</code></pre>\n<p>That is, collections of data differ in terms of:</p>\n<h3>1. <em>optimization</em> - The compiler optimization</h3>\n<ul>\n<li><code>layout</code>: The tensor computational graph is the input graph to the <strong>layoyt assignment pass</strong>.</li>\n<li><code>tile</code>: The fused subgraph (<strong>kernel</strong>) will goes through <strong>tile size selection</strong>.</li>\n</ul>\n<h3>2. <em>source</em> - The source of graphs (graph collection)<br></h3>\n<p>There are two sources from which the computational graphs are collected:</p>\n<ul>\n<li><code>xla</code>: The combination of the XLA regression benchmark.</li>\n<li><code>nlp</code>: A variety of BERT for training and inference.</li>\n</ul>\n<h3>3. <em>search</em> - The search strategy (configuration generation)<br></h3>\n<p>The XLA autotuner is used to generate the configurations, and there are two strategies to explore the search space:</p>\n<ul>\n<li><code>default</code>: Explore the search space using a <strong>genetic algorithm</strong> starting from the default configuration.</li>\n<li><code>random</code>: Pick random candidates in the search space.</li>\n</ul>\n<h3>4. <em>split</em> - The dataset split</h3>\n<p>As usual, the dataset contains <code>train</code>, <code>valid</code>, and <code>test</code> sets.</p>\n<p>With the information above, let's take a look at the data folder structure. Some observations are summarized as follows:</p>\n<ul>\n<li><code>tile</code> optimization contains graphs only from <code>xla</code> source.</li>\n<li>There's only one search strategy for <code>tile</code> optimization (at the level of the <strong>fused subgraph</strong>).<ul>\n<li><strong>Graph-level</strong> optimization is run beforehand to output a collection of kernels.</li>\n<li>Configurations are generated by <strong>enumerating</strong> all possible tile sizes for the kernel in a random order.  </li></ul></li>\n</ul>\n<pre><code>/kaggle//predict-ai-model-runtime/npz_all/npz\n├── layout\n│   ├── nlp\n│   │   ├── \n│   │   │   ├── test\n│   │   │   ├── train\n│   │   │   └── \n│   │   └── random\n│   │       ├── test\n│   │       ├── train\n│   │       └── \n│   └── xla\n│       ├── \n│       │   ├── test\n│       │   ├── train\n│       │   └── \n│       └── random\n│           ├── test\n│           ├── train\n│           └── \n└── tile\n    └── xla\n        ├── test\n        ├── train\n        └── \n</code></pre>\n<p>For detailed EDA, please see <a href=\"https://www.kaggle.com/code/abaojiang/google-fast-or-slow-detailed-eda\" target=\"_blank\">Google - Fast or Slow? - Detailed EDA</a>.</p>\n<p>Hope this helps, thanks.</p>",
      "rawMarkdown": "Hi everyone,\n\nIt's the first time I step into the field of ML compilers. At first glance, it's too much information to digest, so I go through the amazing [paper](https://arxiv.org/abs/2308.13490) provided by the host, which clearly explains how the data is generated and arranged. I decide to record and share some of my understandings about the data here.\n\nIn short, the hierarchy is defined as:\n```\noptimization-source-search-split\n``` \n\nThat is, collections of data differ in terms of:\n### 1. *optimization* - The compiler optimization\n* `layout`: The tensor computational graph is the input graph to the **layoyt assignment pass**.\n* `tile`: The fused subgraph (**kernel**) will goes through **tile size selection**.\n\n###2. *source* - The source of graphs (graph collection)<br>\nThere are two sources from which the computational graphs are collected:\n* `xla`: The combination of the XLA regression benchmark.\n* `nlp`: A variety of BERT for training and inference.\n\n### 3. *search* - The search strategy (configuration generation)<br>\nThe XLA autotuner is used to generate the configurations, and there are two strategies to explore the search space:\n* `default`: Explore the search space using a **genetic algorithm** starting from the default configuration.\n* `random`: Pick random candidates in the search space.\n\n### 4. *split* - The dataset split\nAs usual, the dataset contains `train`, `valid`, and `test` sets.\n  \nWith the information above, let's take a look at the data folder structure. Some observations are summarized as follows:\n* `tile` optimization contains graphs only from `xla` source.\n* There's only one search strategy for `tile` optimization (at the level of the **fused subgraph**).\n    * **Graph-level** optimization is run beforehand to output a collection of kernels.\n    * Configurations are generated by **enumerating** all possible tile sizes for the kernel in a random order.  \n```\n/kaggle/input/predict-ai-model-runtime/npz_all/npz\n├── layout\n│   ├── nlp\n│   │   ├── default\n│   │   │   ├── test\n│   │   │   ├── train\n│   │   │   └── valid\n│   │   └── random\n│   │       ├── test\n│   │       ├── train\n│   │       └── valid\n│   └── xla\n│       ├── default\n│       │   ├── test\n│       │   ├── train\n│       │   └── valid\n│       └── random\n│           ├── test\n│           ├── train\n│           └── valid\n└── tile\n    └── xla\n        ├── test\n        ├── train\n        └── valid\n```\n\nFor detailed EDA, please see [Google - Fast or Slow? - Detailed EDA](https://www.kaggle.com/code/abaojiang/google-fast-or-slow-detailed-eda).\n\nHope this helps, thanks.",
      "votes": null
    },
    {
      "id": "2422147",
      "postDate": "09/03/2023 18:24:07",
      "content": "<p>Thank you for sharing your insight and the summary!</p>",
      "rawMarkdown": "Thank you for sharing your insight and the summary!",
      "votes": null
    },
    {
      "id": "2478972",
      "postDate": "10/12/2023 10:04:05",
      "content": "<p>Thanks you so much for well explain the file structure</p>",
      "rawMarkdown": "Thanks you so much for well explain the file structure",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2422147,
      "author_name": "mangpophothilimthana",
      "author_url": "",
      "post_date": "09/03/2023 18:24:07",
      "content": "<p>Thank you for sharing your insight and the summary!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2478972,
      "author_name": "phnghiapro",
      "author_url": "",
      "post_date": "10/12/2023 10:04:05",
      "content": "<p>Thanks you so much for well explain the file structure</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2421774": "Hi everyone,\n\nIt's the first time I step into the field of ML compilers. At first glance, it's too much information to digest, so I go through the amazing [paper](https://arxiv.org/abs/2308.13490) provided by the host, which clearly explains how the data is generated and arranged. I decide to record and share some of my understandings about the data here.\n\nIn short, the hierarchy is defined as:\n```\noptimization-source-search-split\n``` \n\nThat is, collections of data differ in terms of:\n### 1. *optimization* - The compiler optimization\n* `layout`: The tensor computational graph is the input graph to the **layoyt assignment pass**.\n* `tile`: The fused subgraph (**kernel**) will goes through **tile size selection**.\n\n###2. *source* - The source of graphs (graph collection)<br>\nThere are two sources from which the computational graphs are collected:\n* `xla`: The combination of the XLA regression benchmark.\n* `nlp`: A variety of BERT for training and inference.\n\n### 3. *search* - The search strategy (configuration generation)<br>\nThe XLA autotuner is used to generate the configurations, and there are two strategies to explore the search space:\n* `default`: Explore the search space using a **genetic algorithm** starting from the default configuration.\n* `random`: Pick random candidates in the search space.\n\n### 4. *split* - The dataset split\nAs usual, the dataset contains `train`, `valid`, and `test` sets.\n  \nWith the information above, let's take a look at the data folder structure. Some observations are summarized as follows:\n* `tile` optimization contains graphs only from `xla` source.\n* There's only one search strategy for `tile` optimization (at the level of the **fused subgraph**).\n    * **Graph-level** optimization is run beforehand to output a collection of kernels.\n    * Configurations are generated by **enumerating** all possible tile sizes for the kernel in a random order.  \n```\n/kaggle/input/predict-ai-model-runtime/npz_all/npz\n├── layout\n│   ├── nlp\n│   │   ├── default\n│   │   │   ├── test\n│   │   │   ├── train\n│   │   │   └── valid\n│   │   └── random\n│   │       ├── test\n│   │       ├── train\n│   │       └── valid\n│   └── xla\n│       ├── default\n│       │   ├── test\n│       │   ├── train\n│       │   └── valid\n│       └── random\n│           ├── test\n│           ├── train\n│           └── valid\n└── tile\n    └── xla\n        ├── test\n        ├── train\n        └── valid\n```\n\nFor detailed EDA, please see [Google - Fast or Slow? - Detailed EDA](https://www.kaggle.com/code/abaojiang/google-fast-or-slow-detailed-eda).\n\nHope this helps, thanks.",
    "2422147": "Thank you for sharing your insight and the summary!",
    "2478972": "Thanks you so much for well explain the file structure"
  },
  "source": "meta"
}