{
  "id": 456105,
  "title": "14th Place Solution for the Google - Fast or Slow? Predict AI Model Runtime Competition",
  "url": "/competitions/predict-ai-model-runtime/discussion/456105",
  "author_name": "Matthew Deakos",
  "post_date": "2023-11-18T04:27:58.311000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Context</h1>\n<ul>\n<li>Business Context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/overview</a></li>\n<li>Data Context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/data</a></li>\n</ul>\n<h1>Overview of Approach</h1>\n<h2>Data Preprocessing</h2>\n<ul>\n<li>Developed 2 graph compression techniques to reduce problem complexity (N-hop reduction from config nodes and a config meta-graph).</li>\n<li>Normalized numeric features, dropped features with 0 standard deviation</li>\n<li>One-hot encoded opcodes</li>\n</ul>\n<h2>Feature Engineering</h2>\n<ul>\n<li>Created several node specific features and config specific features</li>\n<li>Created some global features applied to the whole graph</li>\n</ul>\n<h2>Model Design</h2>\n<p>The models all broadly followed the following format:</p>\n<ol>\n<li>Graph/Config/Opcodes concatenated</li>\n<li>Graph representations used to perform some Graph Convolutions (varies slightly between models)</li>\n<li>Global Mean Pooling concatenated with Global Features</li>\n<li>MLP to output layer</li>\n</ol>\n<p>The Tile Dataset result was a single model following this design, with 3 GraphSAGE layers and 3 Linear Layers trained with ListMLE loss. The Layout Dataset results were taken from an ensemble of models with slight variations in their design. All models used GeLu activations, but differed in other respects (detailed below). Output losses used were ListMLE and Pairwise Hinge.</p>\n<h2>Validation</h2>\n<p>We kept the same Train/Val split as provided in the competition dataset.</p>\n<h1>Details of Approach</h1>\n<h2>Graph Reduction</h2>\n<p>Each layout graph was transformed into two distinct graphs:</p>\n<ol>\n<li>A 3-Hop graph (hops from the Configurable Nodes)<br>\nThe configurable nodes are the ones that can differ between graphs, so it makes sense that any graph reduction would try to preserve these nodes. An N-Hop graph transformation will retain configurable nodes, and nodes (and edges) up to N hops away from any configurable node. I landed on a 3-hop graph through val scores, but I think we could have gotten better results with 4 or 5 hop graphs given some time to tune.</li>\n<li>A \"Config Positioning \" graph. This graph removed all non-configurable nodes, but drew edges between all configurable nodes where there was a path from one node to another that did not cross another configurable node. My intuition guiding this was that the relative position of a poorly configured node with respect to downstream configurable nodes might have a meaningful impact on its overall contribution to the runtime. <br>\nAn example of the transformations is shown below (though I just used a 1-hop example).</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2F3ecc17ccdeead41cba76a26a6087d03b%2FGraphReductions3.png?generation=1700281118388593&amp;alt=media\" alt=\"\"></p>\n<h2>The Features</h2>\n<p>Before even talking about new features, it's worth mentioning that normalization is absolutely essential on this problem. If you didn't normalize your data, you didn't score well.</p>\n<h3>Node Features</h3>\n<p>We defined a few extra features. For the nodes, we defined:</p>\n<ul>\n<li>Shape sparsity (shape sum / shape product)</li>\n<li>Dimensionality (count of active shapes)</li>\n<li>Stride Interactions</li>\n<li>Padding Proportions</li>\n<li>Reversal Ratio</li>\n<li>Is configurable (obvious)</li>\n</ul>\n<h3>Configuration Features</h3>\n<p>Additional config features were computed for the output, input and kernel sections. Each feature was replicated for each of those sections:</p>\n<ul>\n<li>is_default (all negative ones)</li>\n<li>active_dims (count of non-negative)</li>\n<li>max order (largest value)</li>\n<li>contiguity rank (count of longest contiguous ordering / active dims)</li>\n<li>section variance</li>\n</ul>\n<p>Additionally, we computed similarity metrics for</p>\n<ul>\n<li>output-input</li>\n<li>output-kernel</li>\n<li>input-kernel</li>\n</ul>\n<h3>Opcodes</h3>\n<p>Opcodes were just one-hot encoded.</p>\n<h3>Global Features</h3>\n<ul>\n<li>Longest Path Length in Graph</li>\n<li>Average shortest path length between connected components</li>\n<li>Number of nodes</li>\n<li>is_default</li>\n</ul>\n<p>The is_default flag was introduced because it <em>seemed</em> like the random vs. default distributions were different enough to warrant having predictive value, since the test set also contained this information. It seemed to provide a small but reliable boost to val scores.</p>\n<h2>The Models</h2>\n<h4>Layout Models</h4>\n<p>For the layout problem, we defined the following GraphBlock, using a GAT with (out channels / 2) channels to process the node features given the 3-Hop Graph, and GAT or GraphSAGE with (out channels / 2) channels to process the features given the Config Positioning graph.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2F893f9ce0ec19a542d2156afc8e951c66%2FGraphBlock2.png?generation=1700278642575220&amp;alt=media\" alt=\"\"></p>\n<p>We then layered the GraphBlocks with residual connections and added some dense feed-forward layers (also with residuals) to which we concatenated the global features. The final result was an ensemble of slightly different versions of this model (varying hidden dims, linear layers, etc.). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2Fc9cfbe2d9197c006df390155fbaab177%2FModelDiagram.png?generation=1700279330883081&amp;alt=media\" alt=\"\"></p>\n<p>The code for of multi-edge block used in the Layout Set is shown below (the Tile Set model is very similar, with fewer complicated pieces. It's less polished but you can view it <a href=\"https://github.com/mattdeak/kaggle-fast-or-slow/blob/master/ml/xla_gcn_v1/model.py\" target=\"_blank\">here</a>. The rest of the code is available <a href=\"https://github.com/mattdeak/kaggle-fast-or-slow/blob/master/readme.md\" target=\"_blank\">here</a></p>\n<pre><code> (nn.Module):\n     ():\n        \n\n        ().__init__()\n\n        output_dim_per_block = output_dim // \n\n         main_block == :\n            self.main_edge_block = GATBlock(\n                input_dim,\n                output_dim_per_block,\n                heads=heads,\n                with_residual=,\n                dropout=dropout,\n            )\n        :\n            self.main_edge_block = SAGEBlock(\n                input_dim,\n                output_dim_per_block,\n                with_residual=,\n                dropout=dropout,\n            )\n\n         alt_block == :\n            self.alternate_edge_block = GATBlock(\n                input_dim,\n                output_dim_per_block,\n                heads=heads,\n                with_residual=,\n                dropout=dropout,\n            )\n\n        :\n            self.alternate_edge_block = SAGEBlock(\n                input_dim,\n                output_dim_per_block,\n                with_residual=,\n                dropout=dropout,\n            )\n\n        self.with_residual = with_residual\n        self.output_dim = output_dim\n\n     ():\n        main_edge_index = data.edge_index\n        alternate_edge_index = data.alt_edge_index\n\n        main_edge_data = Data(\n            x=data.x,\n            edge_index=main_edge_index,\n            batch=data.batch,\n        )\n\n        alternate_edge_data = Data(\n            x=data.x,\n            edge_index=alternate_edge_index,\n            batch=data.batch,\n        )\n\n        main_edge_data = self.main_edge_block(main_edge_data)\n        alternate_edge_data = self.alternate_edge_block(alternate_edge_data)\n\n        f = torch.cat([main_edge_data.x, alternate_edge_data.x], dim=)\n\n         self.with_residual:\n            f += data.x\n\n        new_data = Data(\n            x=f,\n            batch=data.batch,\n        )\n\n         data.update(new_data)\n</code></pre>\n<h4>Tile Model</h4>\n<p>The Model for the Tile Dataset is more or less exactly the same, except it used 3 graph layers, 3 linear layers, and no graph reduction at all (because there were no configurable node).</p>\n<h3>Training Process</h3>\n<p>NLP and XLA were trained separately. I would have loved to play with a unified model more, but I ran out of time and compute, and they seemed to learn well when they were separate.</p>\n<p>All models used:</p>\n<ul>\n<li>AdamW Optimizer</li>\n<li>Batch Size of 16</li>\n<li>GeLu Activations</li>\n<li>LayerNorm in both Graph and MLP</li>\n<li>Global Mean Pooling after the graph blocks</li>\n<li>No Scheduler</li>\n</ul>\n<p>Other parameters are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>XLA-1</th>\n<th>XLA-2</th>\n<th>NLP-1</th>\n<th>NLP-2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Loss</td>\n<td>listMLE</td>\n<td>listMLE</td>\n<td>listMLE</td>\n<td>Rank Margin Loss</td>\n</tr>\n<tr>\n<td>Learning Rate</td>\n<td>0.00028</td>\n<td>0.00028</td>\n<td>0.00028</td>\n<td>0.0001</td>\n</tr>\n<tr>\n<td>Weight Decay</td>\n<td>0.004</td>\n<td>0.004</td>\n<td>0.004</td>\n<td>0.007</td>\n</tr>\n<tr>\n<td>Graph Layers</td>\n<td>4</td>\n<td>4</td>\n<td>4</td>\n<td>4</td>\n</tr>\n<tr>\n<td>Graph Channels</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n</tr>\n<tr>\n<td>FF Layers</td>\n<td>1</td>\n<td>2</td>\n<td>1</td>\n<td>3</td>\n</tr>\n<tr>\n<td>FF Channels</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n<td>64</td>\n</tr>\n<tr>\n<td>3-Hop Graph Conv</td>\n<td>GAT(heads=8)</td>\n<td>GAT(heads=8)</td>\n<td>GAT(heads=8)</td>\n<td>GAT(heads=1)</td>\n</tr>\n<tr>\n<td>Config Graph Conv</td>\n<td>GAT</td>\n<td>GAT</td>\n<td>GAT</td>\n<td>GraphSAGE</td>\n</tr>\n<tr>\n<td>Dropout</td>\n<td>0.15</td>\n<td>0.15</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>Outputs were collected from XLA-1 after 3 epochs, and 2 snapshots of XLA-2 in training (end of epochs 2 and 3) based on their val scores.</p>\n<p>The NLP Models never even finished one epoch, the loss appeared to plateau and I only had so much compute. This indicates to me that I probably could have tuned the LR better or regularized better.</p>\n<h3>Ensembling</h3>\n<p>For a given file id (e.g \"xla:default:abc…\"), we have N (1000 or 1001) predictions per model output. We min-max normalize the predictions so they're all in a zero-to-one range, then we just add the scores elementwise for each model output. We use the summed config scores to derive the rank ordering.</p>\n<p>The normalization is important here, because the models are not guaranteed to be outputting numbers on the same scale if you're using a ranking loss.</p>\n<p>I also tried simple rank averaging and Borda count, both of which worked but not as well as the min-max averaging. This is likely because these methods can't account for things like \"how much better is rank 2 than rank 3\", while the min-max normalized ensemble can.</p>\n<h4>Sources</h4>\n<ul>\n<li>Graph Attention: <a href=\"https://arxiv.org/pdf/1710.10903.pdf\" target=\"_blank\">https://arxiv.org/pdf/1710.10903.pdf</a></li>\n<li>GraphSAGE: <a href=\"https://arxiv.org/pdf/1706.02216.pdf\" target=\"_blank\">https://arxiv.org/pdf/1706.02216.pdf</a></li>\n</ul>",
  "messages": [
    {
      "id": 2529246,
      "postDate": "2023-11-18T04:27:58.310Z",
      "content": "<h1>Context</h1>\n<ul>\n<li>Business Context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/overview</a></li>\n<li>Data Context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/data</a></li>\n</ul>\n<h1>Overview of Approach</h1>\n<h2>Data Preprocessing</h2>\n<ul>\n<li>Developed 2 graph compression techniques to reduce problem complexity (N-hop reduction from config nodes and a config meta-graph).</li>\n<li>Normalized numeric features, dropped features with 0 standard deviation</li>\n<li>One-hot encoded opcodes</li>\n</ul>\n<h2>Feature Engineering</h2>\n<ul>\n<li>Created several node specific features and config specific features</li>\n<li>Created some global features applied to the whole graph</li>\n</ul>\n<h2>Model Design</h2>\n<p>The models all broadly followed the following format:</p>\n<ol>\n<li>Graph/Config/Opcodes concatenated</li>\n<li>Graph representations used to perform some Graph Convolutions (varies slightly between models)</li>\n<li>Global Mean Pooling concatenated with Global Features</li>\n<li>MLP to output layer</li>\n</ol>\n<p>The Tile Dataset result was a single model following this design, with 3 GraphSAGE layers and 3 Linear Layers trained with ListMLE loss. The Layout Dataset results were taken from an ensemble of models with slight variations in their design. All models used GeLu activations, but differed in other respects (detailed below). Output losses used were ListMLE and Pairwise Hinge.</p>\n<h2>Validation</h2>\n<p>We kept the same Train/Val split as provided in the competition dataset.</p>\n<h1>Details of Approach</h1>\n<h2>Graph Reduction</h2>\n<p>Each layout graph was transformed into two distinct graphs:</p>\n<ol>\n<li>A 3-Hop graph (hops from the Configurable Nodes)<br>\nThe configurable nodes are the ones that can differ between graphs, so it makes sense that any graph reduction would try to preserve these nodes. An N-Hop graph transformation will retain configurable nodes, and nodes (and edges) up to N hops away from any configurable node. I landed on a 3-hop graph through val scores, but I think we could have gotten better results with 4 or 5 hop graphs given some time to tune.</li>\n<li>A \"Config Positioning \" graph. This graph removed all non-configurable nodes, but drew edges between all configurable nodes where there was a path from one node to another that did not cross another configurable node. My intuition guiding this was that the relative position of a poorly configured node with respect to downstream configurable nodes might have a meaningful impact on its overall contribution to the runtime. <br>\nAn example of the transformations is shown below (though I just used a 1-hop example).</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2F3ecc17ccdeead41cba76a26a6087d03b%2FGraphReductions3.png?generation=1700281118388593&amp;alt=media\" alt=\"\"></p>\n<h2>The Features</h2>\n<p>Before even talking about new features, it's worth mentioning that normalization is absolutely essential on this problem. If you didn't normalize your data, you didn't score well.</p>\n<h3>Node Features</h3>\n<p>We defined a few extra features. For the nodes, we defined:</p>\n<ul>\n<li>Shape sparsity (shape sum / shape product)</li>\n<li>Dimensionality (count of active shapes)</li>\n<li>Stride Interactions</li>\n<li>Padding Proportions</li>\n<li>Reversal Ratio</li>\n<li>Is configurable (obvious)</li>\n</ul>\n<h3>Configuration Features</h3>\n<p>Additional config features were computed for the output, input and kernel sections. Each feature was replicated for each of those sections:</p>\n<ul>\n<li>is_default (all negative ones)</li>\n<li>active_dims (count of non-negative)</li>\n<li>max order (largest value)</li>\n<li>contiguity rank (count of longest contiguous ordering / active dims)</li>\n<li>section variance</li>\n</ul>\n<p>Additionally, we computed similarity metrics for</p>\n<ul>\n<li>output-input</li>\n<li>output-kernel</li>\n<li>input-kernel</li>\n</ul>\n<h3>Opcodes</h3>\n<p>Opcodes were just one-hot encoded.</p>\n<h3>Global Features</h3>\n<ul>\n<li>Longest Path Length in Graph</li>\n<li>Average shortest path length between connected components</li>\n<li>Number of nodes</li>\n<li>is_default</li>\n</ul>\n<p>The is_default flag was introduced because it <em>seemed</em> like the random vs. default distributions were different enough to warrant having predictive value, since the test set also contained this information. It seemed to provide a small but reliable boost to val scores.</p>\n<h2>The Models</h2>\n<h4>Layout Models</h4>\n<p>For the layout problem, we defined the following GraphBlock, using a GAT with (out channels / 2) channels to process the node features given the 3-Hop Graph, and GAT or GraphSAGE with (out channels / 2) channels to process the features given the Config Positioning graph.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2F893f9ce0ec19a542d2156afc8e951c66%2FGraphBlock2.png?generation=1700278642575220&amp;alt=media\" alt=\"\"></p>\n<p>We then layered the GraphBlocks with residual connections and added some dense feed-forward layers (also with residuals) to which we concatenated the global features. The final result was an ensemble of slightly different versions of this model (varying hidden dims, linear layers, etc.). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2Fc9cfbe2d9197c006df390155fbaab177%2FModelDiagram.png?generation=1700279330883081&amp;alt=media\" alt=\"\"></p>\n<p>The code for of multi-edge block used in the Layout Set is shown below (the Tile Set model is very similar, with fewer complicated pieces. It's less polished but you can view it <a href=\"https://github.com/mattdeak/kaggle-fast-or-slow/blob/master/ml/xla_gcn_v1/model.py\" target=\"_blank\">here</a>. The rest of the code is available <a href=\"https://github.com/mattdeak/kaggle-fast-or-slow/blob/master/readme.md\" target=\"_blank\">here</a></p>\n<pre><code> (nn.Module):\n     ():\n        \n\n        ().__init__()\n\n        output_dim_per_block = output_dim // \n\n         main_block == :\n            self.main_edge_block = GATBlock(\n                input_dim,\n                output_dim_per_block,\n                heads=heads,\n                with_residual=,\n                dropout=dropout,\n            )\n        :\n            self.main_edge_block = SAGEBlock(\n                input_dim,\n                output_dim_per_block,\n                with_residual=,\n                dropout=dropout,\n            )\n\n         alt_block == :\n            self.alternate_edge_block = GATBlock(\n                input_dim,\n                output_dim_per_block,\n                heads=heads,\n                with_residual=,\n                dropout=dropout,\n            )\n\n        :\n            self.alternate_edge_block = SAGEBlock(\n                input_dim,\n                output_dim_per_block,\n                with_residual=,\n                dropout=dropout,\n            )\n\n        self.with_residual = with_residual\n        self.output_dim = output_dim\n\n     ():\n        main_edge_index = data.edge_index\n        alternate_edge_index = data.alt_edge_index\n\n        main_edge_data = Data(\n            x=data.x,\n            edge_index=main_edge_index,\n            batch=data.batch,\n        )\n\n        alternate_edge_data = Data(\n            x=data.x,\n            edge_index=alternate_edge_index,\n            batch=data.batch,\n        )\n\n        main_edge_data = self.main_edge_block(main_edge_data)\n        alternate_edge_data = self.alternate_edge_block(alternate_edge_data)\n\n        f = torch.cat([main_edge_data.x, alternate_edge_data.x], dim=)\n\n         self.with_residual:\n            f += data.x\n\n        new_data = Data(\n            x=f,\n            batch=data.batch,\n        )\n\n         data.update(new_data)\n</code></pre>\n<h4>Tile Model</h4>\n<p>The Model for the Tile Dataset is more or less exactly the same, except it used 3 graph layers, 3 linear layers, and no graph reduction at all (because there were no configurable node).</p>\n<h3>Training Process</h3>\n<p>NLP and XLA were trained separately. I would have loved to play with a unified model more, but I ran out of time and compute, and they seemed to learn well when they were separate.</p>\n<p>All models used:</p>\n<ul>\n<li>AdamW Optimizer</li>\n<li>Batch Size of 16</li>\n<li>GeLu Activations</li>\n<li>LayerNorm in both Graph and MLP</li>\n<li>Global Mean Pooling after the graph blocks</li>\n<li>No Scheduler</li>\n</ul>\n<p>Other parameters are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>XLA-1</th>\n<th>XLA-2</th>\n<th>NLP-1</th>\n<th>NLP-2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Loss</td>\n<td>listMLE</td>\n<td>listMLE</td>\n<td>listMLE</td>\n<td>Rank Margin Loss</td>\n</tr>\n<tr>\n<td>Learning Rate</td>\n<td>0.00028</td>\n<td>0.00028</td>\n<td>0.00028</td>\n<td>0.0001</td>\n</tr>\n<tr>\n<td>Weight Decay</td>\n<td>0.004</td>\n<td>0.004</td>\n<td>0.004</td>\n<td>0.007</td>\n</tr>\n<tr>\n<td>Graph Layers</td>\n<td>4</td>\n<td>4</td>\n<td>4</td>\n<td>4</td>\n</tr>\n<tr>\n<td>Graph Channels</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n</tr>\n<tr>\n<td>FF Layers</td>\n<td>1</td>\n<td>2</td>\n<td>1</td>\n<td>3</td>\n</tr>\n<tr>\n<td>FF Channels</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n<td>64</td>\n</tr>\n<tr>\n<td>3-Hop Graph Conv</td>\n<td>GAT(heads=8)</td>\n<td>GAT(heads=8)</td>\n<td>GAT(heads=8)</td>\n<td>GAT(heads=1)</td>\n</tr>\n<tr>\n<td>Config Graph Conv</td>\n<td>GAT</td>\n<td>GAT</td>\n<td>GAT</td>\n<td>GraphSAGE</td>\n</tr>\n<tr>\n<td>Dropout</td>\n<td>0.15</td>\n<td>0.15</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>Outputs were collected from XLA-1 after 3 epochs, and 2 snapshots of XLA-2 in training (end of epochs 2 and 3) based on their val scores.</p>\n<p>The NLP Models never even finished one epoch, the loss appeared to plateau and I only had so much compute. This indicates to me that I probably could have tuned the LR better or regularized better.</p>\n<h3>Ensembling</h3>\n<p>For a given file id (e.g \"xla:default:abc…\"), we have N (1000 or 1001) predictions per model output. We min-max normalize the predictions so they're all in a zero-to-one range, then we just add the scores elementwise for each model output. We use the summed config scores to derive the rank ordering.</p>\n<p>The normalization is important here, because the models are not guaranteed to be outputting numbers on the same scale if you're using a ranking loss.</p>\n<p>I also tried simple rank averaging and Borda count, both of which worked but not as well as the min-max averaging. This is likely because these methods can't account for things like \"how much better is rank 2 than rank 3\", while the min-max normalized ensemble can.</p>\n<h4>Sources</h4>\n<ul>\n<li>Graph Attention: <a href=\"https://arxiv.org/pdf/1710.10903.pdf\" target=\"_blank\">https://arxiv.org/pdf/1710.10903.pdf</a></li>\n<li>GraphSAGE: <a href=\"https://arxiv.org/pdf/1706.02216.pdf\" target=\"_blank\">https://arxiv.org/pdf/1706.02216.pdf</a></li>\n</ul>",
      "rawMarkdown": "# Context\n* Business Context: https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\n* Data Context: https://www.kaggle.com/competitions/predict-ai-model-runtime/data\n\n# Overview of Approach\n## Data Preprocessing\n* Developed 2 graph compression techniques to reduce problem complexity (N-hop reduction from config nodes and a config meta-graph).\n* Normalized numeric features, dropped features with 0 standard deviation\n* One-hot encoded opcodes\n\n## Feature Engineering\n* Created several node specific features and config specific features\n* Created some global features applied to the whole graph\n\n## Model Design\nThe models all broadly followed the following format:\n1. Graph/Config/Opcodes concatenated\n2. Graph representations used to perform some Graph Convolutions (varies slightly between models)\n3. Global Mean Pooling concatenated with Global Features\n4. MLP to output layer\n\n\nThe Tile Dataset result was a single model following this design, with 3 GraphSAGE layers and 3 Linear Layers trained with ListMLE loss. The Layout Dataset results were taken from an ensemble of models with slight variations in their design. All models used GeLu activations, but differed in other respects (detailed below). Output losses used were ListMLE and Pairwise Hinge.\n\n## Validation\nWe kept the same Train/Val split as provided in the competition dataset.\n\n# Details of Approach \n## Graph Reduction\nEach layout graph was transformed into two distinct graphs:\n\n1. A 3-Hop graph (hops from the Configurable Nodes)\n    The configurable nodes are the ones that can differ between graphs, so it makes sense that any graph reduction would try to preserve these nodes. An N-Hop graph transformation will retain configurable nodes, and nodes (and edges) up to N hops away from any configurable node. I landed on a 3-hop graph through val scores, but I think we could have gotten better results with 4 or 5 hop graphs given some time to tune.\n2. A \"Config Positioning \" graph. This graph removed all non-configurable nodes, but drew edges between all configurable nodes where there was a path from one node to another that did not cross another configurable node. My intuition guiding this was that the relative position of a poorly configured node with respect to downstream configurable nodes might have a meaningful impact on its overall contribution to the runtime. \nAn example of the transformations is shown below (though I just used a 1-hop example).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2F3ecc17ccdeead41cba76a26a6087d03b%2FGraphReductions3.png?generation=1700281118388593&alt=media)\n\n## The Features\nBefore even talking about new features, it's worth mentioning that normalization is absolutely essential on this problem. If you didn't normalize your data, you didn't score well.\n\n### Node Features\nWe defined a few extra features. For the nodes, we defined:\n* Shape sparsity (shape sum / shape product)\n* Dimensionality (count of active shapes)\n* Stride Interactions\n* Padding Proportions\n* Reversal Ratio\n* Is configurable (obvious)\n\n### Configuration Features\nAdditional config features were computed for the output, input and kernel sections. Each feature was replicated for each of those sections:\n* is_default (all negative ones)\n* active_dims (count of non-negative)\n* max order (largest value)\n* contiguity rank (count of longest contiguous ordering / active dims)\n* section variance\n\nAdditionally, we computed similarity metrics for\n* output-input\n* output-kernel\n* input-kernel\n\n### Opcodes\nOpcodes were just one-hot encoded.\n\n### Global Features\n* Longest Path Length in Graph\n* Average shortest path length between connected components\n* Number of nodes\n* is_default\n\nThe is_default flag was introduced because it _seemed_ like the random vs. default distributions were different enough to warrant having predictive value, since the test set also contained this information. It seemed to provide a small but reliable boost to val scores.\n\n## The Models\n#### Layout Models\nFor the layout problem, we defined the following GraphBlock, using a GAT with (out channels / 2) channels to process the node features given the 3-Hop Graph, and GAT or GraphSAGE with (out channels / 2) channels to process the features given the Config Positioning graph.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2F893f9ce0ec19a542d2156afc8e951c66%2FGraphBlock2.png?generation=1700278642575220&alt=media)\n\nWe then layered the GraphBlocks with residual connections and added some dense feed-forward layers (also with residuals) to which we concatenated the global features. The final result was an ensemble of slightly different versions of this model (varying hidden dims, linear layers, etc.). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2Fc9cfbe2d9197c006df390155fbaab177%2FModelDiagram.png?generation=1700279330883081&alt=media)\n\nThe code for of multi-edge block used in the Layout Set is shown below (the Tile Set model is very similar, with fewer complicated pieces. It's less polished but you can view it [here](https://github.com/mattdeak/kaggle-fast-or-slow/blob/master/ml/xla_gcn_v1/model.py). The rest of the code is available [here](https://github.com/mattdeak/kaggle-fast-or-slow/blob/master/readme.md)\n\n\n```\nclass MultiEdgeGATBlock(nn.Module):\n    def __init__(\n        self,\n        *,\n        input_dim: int,\n        output_dim: int,\n        heads: int = 4,\n        with_residual: bool = True,\n        dropout: float = 0.5,\n        main_block: Literal[\"gat\", \"sage\"] = \"gat\",\n        alt_block: Literal[\"gat\", \"sage\"] = \"sage\",\n    ):\n        \"\"\"A block that applies two different edge convolutions to the graph, and then\n        concatenates the results together. Uses an edge mask to determine which edges\n        to apply the main block to, and which to apply the alternate block to.\n        \"\"\"\n\n        super().__init__()\n\n        output_dim_per_block = output_dim // 2\n\n        if main_block == \"gat\":\n            self.main_edge_block = GATBlock(\n                input_dim,\n                output_dim_per_block,\n                heads=heads,\n                with_residual=False,\n                dropout=dropout,\n            )\n        else:\n            self.main_edge_block = SAGEBlock(\n                input_dim,\n                output_dim_per_block,\n                with_residual=False,\n                dropout=dropout,\n            )\n\n        if alt_block == \"gat\":\n            self.alternate_edge_block = GATBlock(\n                input_dim,\n                output_dim_per_block,\n                heads=heads,\n                with_residual=False,\n                dropout=dropout,\n            )\n\n        else:\n            self.alternate_edge_block = SAGEBlock(\n                input_dim,\n                output_dim_per_block,\n                with_residual=False,\n                dropout=dropout,\n            )\n\n        self.with_residual = with_residual\n        self.output_dim = output_dim\n\n    def forward(self, data: Data):\n        main_edge_index = data.edge_index\n        alternate_edge_index = data.alt_edge_index\n\n        main_edge_data = Data(\n            x=data.x,\n            edge_index=main_edge_index,\n            batch=data.batch,\n        )\n\n        alternate_edge_data = Data(\n            x=data.x,\n            edge_index=alternate_edge_index,\n            batch=data.batch,\n        )\n\n        main_edge_data = self.main_edge_block(main_edge_data)\n        alternate_edge_data = self.alternate_edge_block(alternate_edge_data)\n\n        f = torch.cat([main_edge_data.x, alternate_edge_data.x], dim=1)\n\n        if self.with_residual:\n            f += data.x\n\n        new_data = Data(\n            x=f,\n            batch=data.batch,\n        )\n\n        return data.update(new_data)\n```\n\n#### Tile Model\nThe Model for the Tile Dataset is more or less exactly the same, except it used 3 graph layers, 3 linear layers, and no graph reduction at all (because there were no configurable node).\n\n### Training Process\nNLP and XLA were trained separately. I would have loved to play with a unified model more, but I ran out of time and compute, and they seemed to learn well when they were separate.\n\nAll models used:\n* AdamW Optimizer\n* Batch Size of 16\n* GeLu Activations\n* LayerNorm in both Graph and MLP\n* Global Mean Pooling after the graph blocks\n* No Scheduler\n\nOther parameters are as follows:\n\n| Parameter         | XLA-1        | XLA-2        | NLP-1        | NLP-2            |\n|-------------------|--------------|--------------|--------------|------------------|\n| Loss              | listMLE      | listMLE      | listMLE      | Rank Margin Loss |\n| Learning Rate     | 0.00028      | 0.00028      | 0.00028      | 0.0001           |\n| Weight Decay      | 0.004        | 0.004        | 0.004        | 0.007            |\n| Graph Layers      | 4            | 4            | 4            | 4                |\n| Graph Channels    | 128          | 128          | 128          | 128              |\n| FF Layers         | 1            | 2            | 1            | 3                |\n| FF Channels       | 128          | 128          | 128          | 64               |\n| 3-Hop Graph Conv  | GAT(heads=8) | GAT(heads=8) | GAT(heads=8) | GAT(heads=1)     |\n| Config Graph Conv | GAT          | GAT          | GAT          | GraphSAGE        |\n| Dropout           | 0.15         | 0.15         | 0            | 0                |\n\nOutputs were collected from XLA-1 after 3 epochs, and 2 snapshots of XLA-2 in training (end of epochs 2 and 3) based on their val scores.\n\nThe NLP Models never even finished one epoch, the loss appeared to plateau and I only had so much compute. This indicates to me that I probably could have tuned the LR better or regularized better.\n\n### Ensembling\nFor a given file id (e.g \"xla:default:abc...\"), we have N (1000 or 1001) predictions per model output. We min-max normalize the predictions so they're all in a zero-to-one range, then we just add the scores elementwise for each model output. We use the summed config scores to derive the rank ordering.\n\nThe normalization is important here, because the models are not guaranteed to be outputting numbers on the same scale if you're using a ranking loss.\n\nI also tried simple rank averaging and Borda count, both of which worked but not as well as the min-max averaging. This is likely because these methods can't account for things like \"how much better is rank 2 than rank 3\", while the min-max normalized ensemble can.\n\n#### Sources\n* Graph Attention: https://arxiv.org/pdf/1710.10903.pdf\n* GraphSAGE: https://arxiv.org/pdf/1706.02216.pdf",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2529246": "# Context\n* Business Context: https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\n* Data Context: https://www.kaggle.com/competitions/predict-ai-model-runtime/data\n\n# Overview of Approach\n## Data Preprocessing\n* Developed 2 graph compression techniques to reduce problem complexity (N-hop reduction from config nodes and a config meta-graph).\n* Normalized numeric features, dropped features with 0 standard deviation\n* One-hot encoded opcodes\n\n## Feature Engineering\n* Created several node specific features and config specific features\n* Created some global features applied to the whole graph\n\n## Model Design\nThe models all broadly followed the following format:\n1. Graph/Config/Opcodes concatenated\n2. Graph representations used to perform some Graph Convolutions (varies slightly between models)\n3. Global Mean Pooling concatenated with Global Features\n4. MLP to output layer\n\n\nThe Tile Dataset result was a single model following this design, with 3 GraphSAGE layers and 3 Linear Layers trained with ListMLE loss. The Layout Dataset results were taken from an ensemble of models with slight variations in their design. All models used GeLu activations, but differed in other respects (detailed below). Output losses used were ListMLE and Pairwise Hinge.\n\n## Validation\nWe kept the same Train/Val split as provided in the competition dataset.\n\n# Details of Approach \n## Graph Reduction\nEach layout graph was transformed into two distinct graphs:\n\n1. A 3-Hop graph (hops from the Configurable Nodes)\n    The configurable nodes are the ones that can differ between graphs, so it makes sense that any graph reduction would try to preserve these nodes. An N-Hop graph transformation will retain configurable nodes, and nodes (and edges) up to N hops away from any configurable node. I landed on a 3-hop graph through val scores, but I think we could have gotten better results with 4 or 5 hop graphs given some time to tune.\n2. A \"Config Positioning \" graph. This graph removed all non-configurable nodes, but drew edges between all configurable nodes where there was a path from one node to another that did not cross another configurable node. My intuition guiding this was that the relative position of a poorly configured node with respect to downstream configurable nodes might have a meaningful impact on its overall contribution to the runtime. \nAn example of the transformations is shown below (though I just used a 1-hop example).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2F3ecc17ccdeead41cba76a26a6087d03b%2FGraphReductions3.png?generation=1700281118388593&alt=media)\n\n## The Features\nBefore even talking about new features, it's worth mentioning that normalization is absolutely essential on this problem. If you didn't normalize your data, you didn't score well.\n\n### Node Features\nWe defined a few extra features. For the nodes, we defined:\n* Shape sparsity (shape sum / shape product)\n* Dimensionality (count of active shapes)\n* Stride Interactions\n* Padding Proportions\n* Reversal Ratio\n* Is configurable (obvious)\n\n### Configuration Features\nAdditional config features were computed for the output, input and kernel sections. Each feature was replicated for each of those sections:\n* is_default (all negative ones)\n* active_dims (count of non-negative)\n* max order (largest value)\n* contiguity rank (count of longest contiguous ordering / active dims)\n* section variance\n\nAdditionally, we computed similarity metrics for\n* output-input\n* output-kernel\n* input-kernel\n\n### Opcodes\nOpcodes were just one-hot encoded.\n\n### Global Features\n* Longest Path Length in Graph\n* Average shortest path length between connected components\n* Number of nodes\n* is_default\n\nThe is_default flag was introduced because it _seemed_ like the random vs. default distributions were different enough to warrant having predictive value, since the test set also contained this information. It seemed to provide a small but reliable boost to val scores.\n\n## The Models\n#### Layout Models\nFor the layout problem, we defined the following GraphBlock, using a GAT with (out channels / 2) channels to process the node features given the 3-Hop Graph, and GAT or GraphSAGE with (out channels / 2) channels to process the features given the Config Positioning graph.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2F893f9ce0ec19a542d2156afc8e951c66%2FGraphBlock2.png?generation=1700278642575220&alt=media)\n\nWe then layered the GraphBlocks with residual connections and added some dense feed-forward layers (also with residuals) to which we concatenated the global features. The final result was an ensemble of slightly different versions of this model (varying hidden dims, linear layers, etc.). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3626485%2Fc9cfbe2d9197c006df390155fbaab177%2FModelDiagram.png?generation=1700279330883081&alt=media)\n\nThe code for of multi-edge block used in the Layout Set is shown below (the Tile Set model is very similar, with fewer complicated pieces. It's less polished but you can view it [here](https://github.com/mattdeak/kaggle-fast-or-slow/blob/master/ml/xla_gcn_v1/model.py). The rest of the code is available [here](https://github.com/mattdeak/kaggle-fast-or-slow/blob/master/readme.md)\n\n\n```\nclass MultiEdgeGATBlock(nn.Module):\n    def __init__(\n        self,\n        *,\n        input_dim: int,\n        output_dim: int,\n        heads: int = 4,\n        with_residual: bool = True,\n        dropout: float = 0.5,\n        main_block: Literal[\"gat\", \"sage\"] = \"gat\",\n        alt_block: Literal[\"gat\", \"sage\"] = \"sage\",\n    ):\n        \"\"\"A block that applies two different edge convolutions to the graph, and then\n        concatenates the results together. Uses an edge mask to determine which edges\n        to apply the main block to, and which to apply the alternate block to.\n        \"\"\"\n\n        super().__init__()\n\n        output_dim_per_block = output_dim // 2\n\n        if main_block == \"gat\":\n            self.main_edge_block = GATBlock(\n                input_dim,\n                output_dim_per_block,\n                heads=heads,\n                with_residual=False,\n                dropout=dropout,\n            )\n        else:\n            self.main_edge_block = SAGEBlock(\n                input_dim,\n                output_dim_per_block,\n                with_residual=False,\n                dropout=dropout,\n            )\n\n        if alt_block == \"gat\":\n            self.alternate_edge_block = GATBlock(\n                input_dim,\n                output_dim_per_block,\n                heads=heads,\n                with_residual=False,\n                dropout=dropout,\n            )\n\n        else:\n            self.alternate_edge_block = SAGEBlock(\n                input_dim,\n                output_dim_per_block,\n                with_residual=False,\n                dropout=dropout,\n            )\n\n        self.with_residual = with_residual\n        self.output_dim = output_dim\n\n    def forward(self, data: Data):\n        main_edge_index = data.edge_index\n        alternate_edge_index = data.alt_edge_index\n\n        main_edge_data = Data(\n            x=data.x,\n            edge_index=main_edge_index,\n            batch=data.batch,\n        )\n\n        alternate_edge_data = Data(\n            x=data.x,\n            edge_index=alternate_edge_index,\n            batch=data.batch,\n        )\n\n        main_edge_data = self.main_edge_block(main_edge_data)\n        alternate_edge_data = self.alternate_edge_block(alternate_edge_data)\n\n        f = torch.cat([main_edge_data.x, alternate_edge_data.x], dim=1)\n\n        if self.with_residual:\n            f += data.x\n\n        new_data = Data(\n            x=f,\n            batch=data.batch,\n        )\n\n        return data.update(new_data)\n```\n\n#### Tile Model\nThe Model for the Tile Dataset is more or less exactly the same, except it used 3 graph layers, 3 linear layers, and no graph reduction at all (because there were no configurable node).\n\n### Training Process\nNLP and XLA were trained separately. I would have loved to play with a unified model more, but I ran out of time and compute, and they seemed to learn well when they were separate.\n\nAll models used:\n* AdamW Optimizer\n* Batch Size of 16\n* GeLu Activations\n* LayerNorm in both Graph and MLP\n* Global Mean Pooling after the graph blocks\n* No Scheduler\n\nOther parameters are as follows:\n\n| Parameter         | XLA-1        | XLA-2        | NLP-1        | NLP-2            |\n|-------------------|--------------|--------------|--------------|------------------|\n| Loss              | listMLE      | listMLE      | listMLE      | Rank Margin Loss |\n| Learning Rate     | 0.00028      | 0.00028      | 0.00028      | 0.0001           |\n| Weight Decay      | 0.004        | 0.004        | 0.004        | 0.007            |\n| Graph Layers      | 4            | 4            | 4            | 4                |\n| Graph Channels    | 128          | 128          | 128          | 128              |\n| FF Layers         | 1            | 2            | 1            | 3                |\n| FF Channels       | 128          | 128          | 128          | 64               |\n| 3-Hop Graph Conv  | GAT(heads=8) | GAT(heads=8) | GAT(heads=8) | GAT(heads=1)     |\n| Config Graph Conv | GAT          | GAT          | GAT          | GraphSAGE        |\n| Dropout           | 0.15         | 0.15         | 0            | 0                |\n\nOutputs were collected from XLA-1 after 3 epochs, and 2 snapshots of XLA-2 in training (end of epochs 2 and 3) based on their val scores.\n\nThe NLP Models never even finished one epoch, the loss appeared to plateau and I only had so much compute. This indicates to me that I probably could have tuned the LR better or regularized better.\n\n### Ensembling\nFor a given file id (e.g \"xla:default:abc...\"), we have N (1000 or 1001) predictions per model output. We min-max normalize the predictions so they're all in a zero-to-one range, then we just add the scores elementwise for each model output. We use the summed config scores to derive the rank ordering.\n\nThe normalization is important here, because the models are not guaranteed to be outputting numbers on the same scale if you're using a ranking loss.\n\nI also tried simple rank averaging and Borda count, both of which worked but not as well as the min-max averaging. This is likely because these methods can't account for things like \"how much better is rank 2 than rank 3\", while the min-max normalized ensemble can.\n\n#### Sources\n* Graph Attention: https://arxiv.org/pdf/1710.10903.pdf\n* GraphSAGE: https://arxiv.org/pdf/1706.02216.pdf"
  }
}