{
  "id": 456579,
  "title": "33rd Solution Writeup and Discussion on GST",
  "url": "/competitions/predict-ai-model-runtime/writeups/abaojiang-33rd-solution-writeup-and-discussion-on-",
  "author_name": "",
  "post_date": "2023-11-20T18:36:35.840Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, thanks Kaggle and the competition host for hosting this exiting competition and congrats to all the winners. I would like to share my solution (though not that good) mainly from the perspective of <a href=\"https://github.com/kaidic/GST\" target=\"_blank\"><strong>G</strong>raph <strong>S</strong>egment <strong>T</strong>raining (<strong>GST</strong>)</a>. The code is released <a href=\"https://github.com/JiangJiaWei1103/Google-Fast-or-Slow\" target=\"_blank\">here</a>.</p>\n<h2>Overview</h2>\n<ul>\n<li>Data cleaning and preprocessing<ul>\n<li>Add new features following instructions <a href=\"https://github.com/google-research-datasets/tpu_graphs/tree/main#graph-feature-extraction\" target=\"_blank\">here</a>.</li>\n<li>Drop constant and quasi-constant features.</li>\n<li>Label encode features related to <code>shape_element_type_is_X</code>.</li>\n<li>Log transform features with max value greater than 20.</li></ul></li>\n<li>Model architecture<ul>\n<li>Train all models in the <strong>early-join</strong> manner (<em>i.e.,</em> fuse node and config features before getting the whole graph context).</li>\n<li>Use <code>SAGEConv</code> as the GNN block.</li></ul></li>\n<li>Training strategy<ul>\n<li>Use <strong>GST</strong> (with history embedding table and stale embedding dropout) to train <em>layout</em> models in a collection-specific manner (<em>i.e.,</em> one model per collection).</li>\n<li>Sample a subset of configurations to train models per iteration.</li></ul></li>\n<li>Experimental Setup<ul>\n<li>Loss criterion: <code>PairWiseHingeLoss</code> for <em>layout</em> and <code>ListMLE</code> for <em>tile</em></li>\n<li>Optimizer: <code>AdamW</code> with base learning rate <code>1e-3</code> (I decrease lr when increasing #epochs)</li>\n<li>Learning rate scheduler: Cosine schedule without warmup</li>\n<li>Checkpoint: Always pick the model at the last epoch</li></ul></li>\n</ul>\n<h2>Data Cleaning and Preprocessing</h2>\n<p>I process the data by a simple four-stage workflow. Firstly, I find out that some of the features are constant among all the datasets, which can be viewed as the redundant dimensions and dropped directly. Also, those with constant ratio like above 0.999 (<em>i.e.,</em> quasi-constant) are thrown away. Then, I label encode the remaining dimensions related to <code>shape_element_type_is_X</code>, which can be represented with a dense embedding. Finally, considering features can span a wide value range (also, some outliers exist), I simply use <code>np.log1p</code> to log transform features with max value greater than 20.<br>\nAfter processing, the node feature dimension drops to 116 and 50 (89 and 33 without new features added) for <em>xla</em> and <em>nlp</em>, respectively.</p>\n<h2>CV Scheme</h2>\n<p>Considering there are only ~4 graphs and 8 graphs for <em>xla</em> and <em>nlp</em> evaluated on public LB, I try to enlarge the validation set by splitting train+val stratified on runtime, which can somewhat balance the intrinsic graph properties (I explore relationship between graph stats and runtime in <a href=\"https://www.kaggle.com/code/abaojiang/google-fast-or-slow-detailed-eda\" target=\"_blank\">this notebook</a>). However, I don't think it's much different from just using the official train-val splitting.</p>\n<h2>Model Architecture</h2>\n<p><a href=\"https://postimg.cc/mtC0KZzX\" target=\"_blank\"><img src=\"https://i.postimg.cc/dtxwkLqY/Screenshot-2023-11-20-at-15-10-31.png\" alt=\"Screenshot-2023-11-20-at-15-10-31.png\"></a><br>\nThe figure above illustrates the overview of model architecture, where \\(d_n \\) and \\(d_c \\) denote the node and config feature dimensions. And, \\(L \\) is the number of graph convolution layers.<br>\nSince my first submission on 22nd, Oct, I use early-join to fuse the node and config features. After experimenting with different GNN blocks (<em>e.g.,</em> <code>GATConv</code>, <code>GATv2Conv</code>, <code>GINConv</code>), <code>SAGEConv</code> always outperforms, so I stick to it till the end. Also \\(L \\) is always set to 3. To be honest, there's no fancy design in my model architecture. Hence, I want to talk more about the training strategy.</p>\n<h2>Training Strategy - <strong>G</strong>raph <strong>S</strong>egment <strong>T</strong>raining (GST)</h2>\n<p>Considering the memory limitation, I quickly decide to choose off-the-shelf <strong>GST</strong> as my training framework. As there exists some unsolved issues in the official implementation of <strong>GST</strong>, I rewrite the pipeline without <a href=\"https://github.com/rampasek/GraphGPS\" target=\"_blank\">GraphGPS</a>.<br>\nThe main concern with <strong>GST</strong> is that the training loss increases as the training process progresses, but validation performance still improves over time. After fixing the \\(\\eta \\), the weight for each graph segment, for final sum pooling, the training loss decreases normally as shown below (special thanks to <a href=\"https://www.kaggle.com/dsfhe49854\" target=\"_blank\">@dsfhe49854</a> 's  analysis <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/448367#2497447\" target=\"_blank\">here</a>),<br>\n<a href=\"https://postimg.cc/jCyB81QT\" target=\"_blank\"><img src=\"https://i.postimg.cc/Vsh1NQJJ/Screenshot-2023-11-20-at-17-50-00-Weights-Biases.png\" alt=\"Screenshot-2023-11-20-at-17-50-00-Weights-Biases.png\"></a><br>\nLet's see how \\(\\eta \\) is derived in the original paper. Let \\(n \\) be the number of segments for one graph and \\(k \\) be the number of segments to be trained per iteration. Also, select \\(p \\) as the dropout ratio of <strong>stale embedding dropout</strong>. Assume we sample only one segment for training per iteration(<em>i.e.,</em> \\(k = 1\\)), as described in the paper. The weight \\(\\alpha \\) of the trained segment can be derived as follows, <br>\n$$<br>\n(n-k)p + k\\alpha = n<br>\n$$<br>\n$$<br>\n\\alpha = (1-p)\\frac{n}{k} + p<br>\n$$</p>\n<p>The logic behind the scene is that the final runtime estimation is the <strong>sum pooling</strong> of runtimes of all segments. Considering some segments are dropped with probability \\(p \\), we need to increase the weight of the trained segment for compensation. However, the problem is that most of the entries in historical embedding table are zeros. Therefore, in early epochs, the objective can be approximated as,<br>\n$$<br>\n\\hat{y} = \\alpha \\hat{y}_{i} ,<br>\n$$</p>\n<p>where \\(\\hat{y} \\) is the predicting runtime of the current graph and \\(\\hat{y}_{i} \\) is the predicting runtime of the segment \\(i \\) of the current graph. What's interesting is that I observe the <strong>unfixed \\(\\eta \\)</strong> always leads to better generalizability compared with the fixed one. Also, if the model is trained with sufficient number of iterations, the training loss actually goes downward (the red line turns the direction at ~100 epochs).</p>\n<h2>Experimental Results</h2>\n<p>Following table shows the local CV scores of my final submission.</p>\n<table>\n<thead>\n<tr>\n<th>Collection</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><em>tile</em></td>\n<td>0.9551</td>\n</tr>\n<tr>\n<td><em>xla-default</em></td>\n<td>0.3188</td>\n</tr>\n<tr>\n<td><em>xla-random</em></td>\n<td>0.5569</td>\n</tr>\n<tr>\n<td><em>nlp-default</em></td>\n<td>0.5053</td>\n</tr>\n<tr>\n<td><em>nlp-random</em></td>\n<td>0.8845</td>\n</tr>\n</tbody>\n</table>\n<h2>What Didn't Work for Me</h2>\n<ul>\n<li>Use other GNN blocks (<em>e.g.,</em> <code>GATConv</code>, <code>GATv2Conv</code>, <code>GINConv</code>)</li>\n<li>Retrain models on the whole dataset</li>\n<li>Finetune <em>default</em> using the pretrained weights from <em>random</em><ul>\n<li>Freezing different parts of network makes no difference.</li></ul></li>\n<li>Segment graphs with other strategies (<em>e.g.,</em> Metis)</li>\n</ul>\n<h2>Conclusion</h2>\n<p>It's not good to stick to only one method (<strong>GST</strong>) for all implementation, I should have explored other potential solutions like all amazing writeups I've digested so far. Though the result isn't that promising this time, I will keep progressing and learning from the top-tiers. This journey never stops! Thanks for your patience!</p>",
  "messages": [
    {
      "id": "2532018",
      "postDate": "11/20/2023 18:07:16",
      "content": "<p>First of all, thanks Kaggle and the competition host for hosting this exiting competition and congrats to all the winners. I would like to share my solution (though not that good) mainly from the perspective of <a href=\"https://github.com/kaidic/GST\" target=\"_blank\"><strong>G</strong>raph <strong>S</strong>egment <strong>T</strong>raining (<strong>GST</strong>)</a>. The code is released <a href=\"https://github.com/JiangJiaWei1103/Google-Fast-or-Slow\" target=\"_blank\">here</a>.</p>\n<h2>Overview</h2>\n<ul>\n<li>Data cleaning and preprocessing<ul>\n<li>Add new features following instructions <a href=\"https://github.com/google-research-datasets/tpu_graphs/tree/main#graph-feature-extraction\" target=\"_blank\">here</a>.</li>\n<li>Drop constant and quasi-constant features.</li>\n<li>Label encode features related to <code>shape_element_type_is_X</code>.</li>\n<li>Log transform features with max value greater than 20.</li></ul></li>\n<li>Model architecture<ul>\n<li>Train all models in the <strong>early-join</strong> manner (<em>i.e.,</em> fuse node and config features before getting the whole graph context).</li>\n<li>Use <code>SAGEConv</code> as the GNN block.</li></ul></li>\n<li>Training strategy<ul>\n<li>Use <strong>GST</strong> (with history embedding table and stale embedding dropout) to train <em>layout</em> models in a collection-specific manner (<em>i.e.,</em> one model per collection).</li>\n<li>Sample a subset of configurations to train models per iteration.</li></ul></li>\n<li>Experimental Setup<ul>\n<li>Loss criterion: <code>PairWiseHingeLoss</code> for <em>layout</em> and <code>ListMLE</code> for <em>tile</em></li>\n<li>Optimizer: <code>AdamW</code> with base learning rate <code>1e-3</code> (I decrease lr when increasing #epochs)</li>\n<li>Learning rate scheduler: Cosine schedule without warmup</li>\n<li>Checkpoint: Always pick the model at the last epoch</li></ul></li>\n</ul>\n<h2>Data Cleaning and Preprocessing</h2>\n<p>I process the data by a simple four-stage workflow. Firstly, I find out that some of the features are constant among all the datasets, which can be viewed as the redundant dimensions and dropped directly. Also, those with constant ratio like above 0.999 (<em>i.e.,</em> quasi-constant) are thrown away. Then, I label encode the remaining dimensions related to <code>shape_element_type_is_X</code>, which can be represented with a dense embedding. Finally, considering features can span a wide value range (also, some outliers exist), I simply use <code>np.log1p</code> to log transform features with max value greater than 20.<br>\nAfter processing, the node feature dimension drops to 116 and 50 (89 and 33 without new features added) for <em>xla</em> and <em>nlp</em>, respectively.</p>\n<h2>CV Scheme</h2>\n<p>Considering there are only ~4 graphs and 8 graphs for <em>xla</em> and <em>nlp</em> evaluated on public LB, I try to enlarge the validation set by splitting train+val stratified on runtime, which can somewhat balance the intrinsic graph properties (I explore relationship between graph stats and runtime in <a href=\"https://www.kaggle.com/code/abaojiang/google-fast-or-slow-detailed-eda\" target=\"_blank\">this notebook</a>). However, I don't think it's much different from just using the official train-val splitting.</p>\n<h2>Model Architecture</h2>\n<p><a href=\"https://postimg.cc/mtC0KZzX\" target=\"_blank\"><img src=\"https://i.postimg.cc/dtxwkLqY/Screenshot-2023-11-20-at-15-10-31.png\" alt=\"Screenshot-2023-11-20-at-15-10-31.png\"></a><br>\nThe figure above illustrates the overview of model architecture, where \\(d_n \\) and \\(d_c \\) denote the node and config feature dimensions. And, \\(L \\) is the number of graph convolution layers.<br>\nSince my first submission on 22nd, Oct, I use early-join to fuse the node and config features. After experimenting with different GNN blocks (<em>e.g.,</em> <code>GATConv</code>, <code>GATv2Conv</code>, <code>GINConv</code>), <code>SAGEConv</code> always outperforms, so I stick to it till the end. Also \\(L \\) is always set to 3. To be honest, there's no fancy design in my model architecture. Hence, I want to talk more about the training strategy.</p>\n<h2>Training Strategy - <strong>G</strong>raph <strong>S</strong>egment <strong>T</strong>raining (GST)</h2>\n<p>Considering the memory limitation, I quickly decide to choose off-the-shelf <strong>GST</strong> as my training framework. As there exists some unsolved issues in the official implementation of <strong>GST</strong>, I rewrite the pipeline without <a href=\"https://github.com/rampasek/GraphGPS\" target=\"_blank\">GraphGPS</a>.<br>\nThe main concern with <strong>GST</strong> is that the training loss increases as the training process progresses, but validation performance still improves over time. After fixing the \\(\\eta \\), the weight for each graph segment, for final sum pooling, the training loss decreases normally as shown below (special thanks to <a href=\"https://www.kaggle.com/dsfhe49854\" target=\"_blank\">@dsfhe49854</a> 's  analysis <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/448367#2497447\" target=\"_blank\">here</a>),<br>\n<a href=\"https://postimg.cc/jCyB81QT\" target=\"_blank\"><img src=\"https://i.postimg.cc/Vsh1NQJJ/Screenshot-2023-11-20-at-17-50-00-Weights-Biases.png\" alt=\"Screenshot-2023-11-20-at-17-50-00-Weights-Biases.png\"></a><br>\nLet's see how \\(\\eta \\) is derived in the original paper. Let \\(n \\) be the number of segments for one graph and \\(k \\) be the number of segments to be trained per iteration. Also, select \\(p \\) as the dropout ratio of <strong>stale embedding dropout</strong>. Assume we sample only one segment for training per iteration(<em>i.e.,</em> \\(k = 1\\)), as described in the paper. The weight \\(\\alpha \\) of the trained segment can be derived as follows, <br>\n$$<br>\n(n-k)p + k\\alpha = n<br>\n$$<br>\n$$<br>\n\\alpha = (1-p)\\frac{n}{k} + p<br>\n$$</p>\n<p>The logic behind the scene is that the final runtime estimation is the <strong>sum pooling</strong> of runtimes of all segments. Considering some segments are dropped with probability \\(p \\), we need to increase the weight of the trained segment for compensation. However, the problem is that most of the entries in historical embedding table are zeros. Therefore, in early epochs, the objective can be approximated as,<br>\n$$<br>\n\\hat{y} = \\alpha \\hat{y}_{i} ,<br>\n$$</p>\n<p>where \\(\\hat{y} \\) is the predicting runtime of the current graph and \\(\\hat{y}_{i} \\) is the predicting runtime of the segment \\(i \\) of the current graph. What's interesting is that I observe the <strong>unfixed \\(\\eta \\)</strong> always leads to better generalizability compared with the fixed one. Also, if the model is trained with sufficient number of iterations, the training loss actually goes downward (the red line turns the direction at ~100 epochs).</p>\n<h2>Experimental Results</h2>\n<p>Following table shows the local CV scores of my final submission.</p>\n<table>\n<thead>\n<tr>\n<th>Collection</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><em>tile</em></td>\n<td>0.9551</td>\n</tr>\n<tr>\n<td><em>xla-default</em></td>\n<td>0.3188</td>\n</tr>\n<tr>\n<td><em>xla-random</em></td>\n<td>0.5569</td>\n</tr>\n<tr>\n<td><em>nlp-default</em></td>\n<td>0.5053</td>\n</tr>\n<tr>\n<td><em>nlp-random</em></td>\n<td>0.8845</td>\n</tr>\n</tbody>\n</table>\n<h2>What Didn't Work for Me</h2>\n<ul>\n<li>Use other GNN blocks (<em>e.g.,</em> <code>GATConv</code>, <code>GATv2Conv</code>, <code>GINConv</code>)</li>\n<li>Retrain models on the whole dataset</li>\n<li>Finetune <em>default</em> using the pretrained weights from <em>random</em><ul>\n<li>Freezing different parts of network makes no difference.</li></ul></li>\n<li>Segment graphs with other strategies (<em>e.g.,</em> Metis)</li>\n</ul>\n<h2>Conclusion</h2>\n<p>It's not good to stick to only one method (<strong>GST</strong>) for all implementation, I should have explored other potential solutions like all amazing writeups I've digested so far. Though the result isn't that promising this time, I will keep progressing and learning from the top-tiers. This journey never stops! Thanks for your patience!</p>",
      "rawMarkdown": "First of all, thanks Kaggle and the competition host for hosting this exiting competition and congrats to all the winners. I would like to share my solution (though not that good) mainly from the perspective of [**G**raph **S**egment **T**raining (**GST**)](https://github.com/kaidic/GST). The code is released [here](https://github.com/JiangJiaWei1103/Google-Fast-or-Slow).\n\n## Overview\n* Data cleaning and preprocessing\n    *  Add new features following instructions [here](https://github.com/google-research-datasets/tpu_graphs/tree/main#graph-feature-extraction).\n    * Drop constant and quasi-constant features.\n    * Label encode features related to `shape_element_type_is_X`.\n    * Log transform features with max value greater than 20.\n* Model architecture\n    * Train all models in the **early-join** manner (*i.e.,* fuse node and config features before getting the whole graph context).\n    * Use `SAGEConv` as the GNN block.\n* Training strategy\n    * Use **GST** (with history embedding table and stale embedding dropout) to train *layout* models in a collection-specific manner (*i.e.,* one model per collection).\n    * Sample a subset of configurations to train models per iteration.\n* Experimental Setup\n    * Loss criterion: `PairWiseHingeLoss` for *layout* and `ListMLE` for *tile*\n    * Optimizer: `AdamW` with base learning rate `1e-3` (I decrease lr when increasing #epochs)\n    * Learning rate scheduler: Cosine schedule without warmup\n    * Checkpoint: Always pick the model at the last epoch\n\n## Data Cleaning and Preprocessing\nI process the data by a simple four-stage workflow. Firstly, I find out that some of the features are constant among all the datasets, which can be viewed as the redundant dimensions and dropped directly. Also, those with constant ratio like above 0.999 (*i.e.,* quasi-constant) are thrown away. Then, I label encode the remaining dimensions related to `shape_element_type_is_X`, which can be represented with a dense embedding. Finally, considering features can span a wide value range (also, some outliers exist), I simply use `np.log1p` to log transform features with max value greater than 20.\nAfter processing, the node feature dimension drops to 116 and 50 (89 and 33 without new features added) for *xla* and *nlp*, respectively.\n\n## CV Scheme\nConsidering there are only ~4 graphs and 8 graphs for *xla* and *nlp* evaluated on public LB, I try to enlarge the validation set by splitting train+val stratified on runtime, which can somewhat balance the intrinsic graph properties (I explore relationship between graph stats and runtime in [this notebook](https://www.kaggle.com/code/abaojiang/google-fast-or-slow-detailed-eda)). However, I don't think it's much different from just using the official train-val splitting.\n\n## Model Architecture\n[![Screenshot-2023-11-20-at-15-10-31.png](https://i.postimg.cc/dtxwkLqY/Screenshot-2023-11-20-at-15-10-31.png)](https://postimg.cc/mtC0KZzX)\nThe figure above illustrates the overview of model architecture, where \\\\(d_n \\\\) and \\\\(d_c \\\\) denote the node and config feature dimensions. And, \\\\(L \\\\) is the number of graph convolution layers.\nSince my first submission on 22nd, Oct, I use early-join to fuse the node and config features. After experimenting with different GNN blocks (*e.g.,* `GATConv`, `GATv2Conv`, `GINConv`), `SAGEConv` always outperforms, so I stick to it till the end. Also \\\\(L \\\\) is always set to 3. To be honest, there's no fancy design in my model architecture. Hence, I want to talk more about the training strategy.\n\n## Training Strategy - **G**raph **S**egment **T**raining (GST)\nConsidering the memory limitation, I quickly decide to choose off-the-shelf **GST** as my training framework. As there exists some unsolved issues in the official implementation of **GST**, I rewrite the pipeline without [GraphGPS](https://github.com/rampasek/GraphGPS).\nThe main concern with **GST** is that the training loss increases as the training process progresses, but validation performance still improves over time. After fixing the \\\\(\\eta \\\\), the weight for each graph segment, for final sum pooling, the training loss decreases normally as shown below (special thanks to @dsfhe49854 's  analysis [here](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/448367#2497447)),\n[![Screenshot-2023-11-20-at-17-50-00-Weights-Biases.png](https://i.postimg.cc/Vsh1NQJJ/Screenshot-2023-11-20-at-17-50-00-Weights-Biases.png)](https://postimg.cc/jCyB81QT)\nLet's see how \\\\(\\eta \\\\) is derived in the original paper. Let \\\\(n \\\\) be the number of segments for one graph and \\\\(k \\\\) be the number of segments to be trained per iteration. Also, select \\\\(p \\\\) as the dropout ratio of **stale embedding dropout**. Assume we sample only one segment for training per iteration(*i.e.,* \\\\(k = 1\\\\)), as described in the paper. The weight \\\\(\\alpha \\\\) of the trained segment can be derived as follows, \n$$\n(n-k)p + k\\alpha = n\n$$\n$$\n\\alpha = (1-p)\\frac{n}{k} + p\n$$\n\nThe logic behind the scene is that the final runtime estimation is the **sum pooling** of runtimes of all segments. Considering some segments are dropped with probability \\\\(p \\\\), we need to increase the weight of the trained segment for compensation. However, the problem is that most of the entries in historical embedding table are zeros. Therefore, in early epochs, the objective can be approximated as,\n$$\n\\hat{y} = \\alpha \\hat{y}_{i} ,\n$$\n\nwhere \\\\(\\hat{y} \\\\) is the predicting runtime of the current graph and \\\\(\\hat{y}_{i} \\\\) is the predicting runtime of the segment \\\\(i \\\\) of the current graph. What's interesting is that I observe the **unfixed \\\\(\\eta \\\\)** always leads to better generalizability compared with the fixed one. Also, if the model is trained with sufficient number of iterations, the training loss actually goes downward (the red line turns the direction at ~100 epochs).\n\n## Experimental Results\nFollowing table shows the local CV scores of my final submission.\n| Collection | CV | \n| --- | --- | \n| *tile* | 0.9551 | \n| *xla-default* | 0.3188 |\n| *xla-random* | 0.5569 |\n| *nlp-default* | 0.5053 |\n| *nlp-random* | 0.8845 |\n\n## What Didn't Work for Me\n* Use other GNN blocks (*e.g.,* `GATConv`, `GATv2Conv`, `GINConv`)\n* Retrain models on the whole dataset\n* Finetune *default* using the pretrained weights from *random*\n    * Freezing different parts of network makes no difference.\n* Segment graphs with other strategies (*e.g.,* Metis)\n\n## Conclusion\nIt's not good to stick to only one method (**GST**) for all implementation, I should have explored other potential solutions like all amazing writeups I've digested so far. Though the result isn't that promising this time, I will keep progressing and learning from the top-tiers. This journey never stops! Thanks for your patience!",
      "votes": null
    },
    {
      "id": "2798881",
      "postDate": "05/07/2024 13:33:20",
      "content": "<p>Great solution and nice work!<br>\nI have a question about the graph partitioning method. It's mentioned that segmenting graphs with other strategies (e.g., Metis) didn't work for your case. Then what algorithm did you use for graph partitioning to use GST?</p>",
      "rawMarkdown": "Great solution and nice work!\nI have a question about the graph partitioning method. It's mentioned that segmenting graphs with other strategies (e.g., Metis) didn't work for your case. Then what algorithm did you use for graph partitioning to use GST?",
      "votes": null
    },
    {
      "id": "2813722",
      "postDate": "05/15/2024 02:35:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/eunseonseong\" target=\"_blank\">@eunseonseong</a>,</p>\n<p>Sorry for the late reply. As shown in the <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">data tab</a>, nodes in a HLO graph are sorted topologically. Hence, I just partition the graph following this order without any change. Also, I use a hyperparameter <code>max_seg_size</code> (1000 by default) to control the size of graph partitions. If you're interested in how it's implemented, you can refer to <a href=\"https://github.com/JiangJiaWei1103/Google-Fast-or-Slow-33rd-Place-Solution/blob/master/data/dataset.py#L127\" target=\"_blank\">my released code</a>. Thanks!</p>",
      "rawMarkdown": "Hi @eunseonseong,\n\nSorry for the late reply. As shown in the [data tab](https://www.kaggle.com/competitions/predict-ai-model-runtime/data), nodes in a HLO graph are sorted topologically. Hence, I just partition the graph following this order without any change. Also, I use a hyperparameter `max_seg_size` (1000 by default) to control the size of graph partitions. If you're interested in how it's implemented, you can refer to [my released code](https://github.com/JiangJiaWei1103/Google-Fast-or-Slow-33rd-Place-Solution/blob/master/data/dataset.py#L127). Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2798881,
      "author_name": "eunseonseong",
      "author_url": "",
      "post_date": "05/07/2024 13:33:20",
      "content": "<p>Great solution and nice work!<br>\nI have a question about the graph partitioning method. It's mentioned that segmenting graphs with other strategies (e.g., Metis) didn't work for your case. Then what algorithm did you use for graph partitioning to use GST?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2813722,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "05/15/2024 02:35:51",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/eunseonseong\" target=\"_blank\">@eunseonseong</a>,</p>\n<p>Sorry for the late reply. As shown in the <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">data tab</a>, nodes in a HLO graph are sorted topologically. Hence, I just partition the graph following this order without any change. Also, I use a hyperparameter <code>max_seg_size</code> (1000 by default) to control the size of graph partitions. If you're interested in how it's implemented, you can refer to <a href=\"https://github.com/JiangJiaWei1103/Google-Fast-or-Slow-33rd-Place-Solution/blob/master/data/dataset.py#L127\" target=\"_blank\">my released code</a>. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2532018": "First of all, thanks Kaggle and the competition host for hosting this exiting competition and congrats to all the winners. I would like to share my solution (though not that good) mainly from the perspective of [**G**raph **S**egment **T**raining (**GST**)](https://github.com/kaidic/GST). The code is released [here](https://github.com/JiangJiaWei1103/Google-Fast-or-Slow).\n\n## Overview\n* Data cleaning and preprocessing\n    *  Add new features following instructions [here](https://github.com/google-research-datasets/tpu_graphs/tree/main#graph-feature-extraction).\n    * Drop constant and quasi-constant features.\n    * Label encode features related to `shape_element_type_is_X`.\n    * Log transform features with max value greater than 20.\n* Model architecture\n    * Train all models in the **early-join** manner (*i.e.,* fuse node and config features before getting the whole graph context).\n    * Use `SAGEConv` as the GNN block.\n* Training strategy\n    * Use **GST** (with history embedding table and stale embedding dropout) to train *layout* models in a collection-specific manner (*i.e.,* one model per collection).\n    * Sample a subset of configurations to train models per iteration.\n* Experimental Setup\n    * Loss criterion: `PairWiseHingeLoss` for *layout* and `ListMLE` for *tile*\n    * Optimizer: `AdamW` with base learning rate `1e-3` (I decrease lr when increasing #epochs)\n    * Learning rate scheduler: Cosine schedule without warmup\n    * Checkpoint: Always pick the model at the last epoch\n\n## Data Cleaning and Preprocessing\nI process the data by a simple four-stage workflow. Firstly, I find out that some of the features are constant among all the datasets, which can be viewed as the redundant dimensions and dropped directly. Also, those with constant ratio like above 0.999 (*i.e.,* quasi-constant) are thrown away. Then, I label encode the remaining dimensions related to `shape_element_type_is_X`, which can be represented with a dense embedding. Finally, considering features can span a wide value range (also, some outliers exist), I simply use `np.log1p` to log transform features with max value greater than 20.\nAfter processing, the node feature dimension drops to 116 and 50 (89 and 33 without new features added) for *xla* and *nlp*, respectively.\n\n## CV Scheme\nConsidering there are only ~4 graphs and 8 graphs for *xla* and *nlp* evaluated on public LB, I try to enlarge the validation set by splitting train+val stratified on runtime, which can somewhat balance the intrinsic graph properties (I explore relationship between graph stats and runtime in [this notebook](https://www.kaggle.com/code/abaojiang/google-fast-or-slow-detailed-eda)). However, I don't think it's much different from just using the official train-val splitting.\n\n## Model Architecture\n[![Screenshot-2023-11-20-at-15-10-31.png](https://i.postimg.cc/dtxwkLqY/Screenshot-2023-11-20-at-15-10-31.png)](https://postimg.cc/mtC0KZzX)\nThe figure above illustrates the overview of model architecture, where \\\\(d_n \\\\) and \\\\(d_c \\\\) denote the node and config feature dimensions. And, \\\\(L \\\\) is the number of graph convolution layers.\nSince my first submission on 22nd, Oct, I use early-join to fuse the node and config features. After experimenting with different GNN blocks (*e.g.,* `GATConv`, `GATv2Conv`, `GINConv`), `SAGEConv` always outperforms, so I stick to it till the end. Also \\\\(L \\\\) is always set to 3. To be honest, there's no fancy design in my model architecture. Hence, I want to talk more about the training strategy.\n\n## Training Strategy - **G**raph **S**egment **T**raining (GST)\nConsidering the memory limitation, I quickly decide to choose off-the-shelf **GST** as my training framework. As there exists some unsolved issues in the official implementation of **GST**, I rewrite the pipeline without [GraphGPS](https://github.com/rampasek/GraphGPS).\nThe main concern with **GST** is that the training loss increases as the training process progresses, but validation performance still improves over time. After fixing the \\\\(\\eta \\\\), the weight for each graph segment, for final sum pooling, the training loss decreases normally as shown below (special thanks to @dsfhe49854 's  analysis [here](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/448367#2497447)),\n[![Screenshot-2023-11-20-at-17-50-00-Weights-Biases.png](https://i.postimg.cc/Vsh1NQJJ/Screenshot-2023-11-20-at-17-50-00-Weights-Biases.png)](https://postimg.cc/jCyB81QT)\nLet's see how \\\\(\\eta \\\\) is derived in the original paper. Let \\\\(n \\\\) be the number of segments for one graph and \\\\(k \\\\) be the number of segments to be trained per iteration. Also, select \\\\(p \\\\) as the dropout ratio of **stale embedding dropout**. Assume we sample only one segment for training per iteration(*i.e.,* \\\\(k = 1\\\\)), as described in the paper. The weight \\\\(\\alpha \\\\) of the trained segment can be derived as follows, \n$$\n(n-k)p + k\\alpha = n\n$$\n$$\n\\alpha = (1-p)\\frac{n}{k} + p\n$$\n\nThe logic behind the scene is that the final runtime estimation is the **sum pooling** of runtimes of all segments. Considering some segments are dropped with probability \\\\(p \\\\), we need to increase the weight of the trained segment for compensation. However, the problem is that most of the entries in historical embedding table are zeros. Therefore, in early epochs, the objective can be approximated as,\n$$\n\\hat{y} = \\alpha \\hat{y}_{i} ,\n$$\n\nwhere \\\\(\\hat{y} \\\\) is the predicting runtime of the current graph and \\\\(\\hat{y}_{i} \\\\) is the predicting runtime of the segment \\\\(i \\\\) of the current graph. What's interesting is that I observe the **unfixed \\\\(\\eta \\\\)** always leads to better generalizability compared with the fixed one. Also, if the model is trained with sufficient number of iterations, the training loss actually goes downward (the red line turns the direction at ~100 epochs).\n\n## Experimental Results\nFollowing table shows the local CV scores of my final submission.\n| Collection | CV | \n| --- | --- | \n| *tile* | 0.9551 | \n| *xla-default* | 0.3188 |\n| *xla-random* | 0.5569 |\n| *nlp-default* | 0.5053 |\n| *nlp-random* | 0.8845 |\n\n## What Didn't Work for Me\n* Use other GNN blocks (*e.g.,* `GATConv`, `GATv2Conv`, `GINConv`)\n* Retrain models on the whole dataset\n* Finetune *default* using the pretrained weights from *random*\n    * Freezing different parts of network makes no difference.\n* Segment graphs with other strategies (*e.g.,* Metis)\n\n## Conclusion\nIt's not good to stick to only one method (**GST**) for all implementation, I should have explored other potential solutions like all amazing writeups I've digested so far. Though the result isn't that promising this time, I will keep progressing and learning from the top-tiers. This journey never stops! Thanks for your patience!",
    "2798881": "Great solution and nice work!\nI have a question about the graph partitioning method. It's mentioned that segmenting graphs with other strategies (e.g., Metis) didn't work for your case. Then what algorithm did you use for graph partitioning to use GST?",
    "2813722": "Hi @eunseonseong,\n\nSorry for the late reply. As shown in the [data tab](https://www.kaggle.com/competitions/predict-ai-model-runtime/data), nodes in a HLO graph are sorted topologically. Hence, I just partition the graph following this order without any change. Also, I use a hyperparameter `max_seg_size` (1000 by default) to control the size of graph partitions. If you're interested in how it's implemented, you can refer to [my released code](https://github.com/JiangJiaWei1103/Google-Fast-or-Slow-33rd-Place-Solution/blob/master/data/dataset.py#L127). Thanks!"
  },
  "source": "meta"
}