{
  "id": 456343,
  "title": "1st Place Solution for the Google - Fast or Slow? Predict AI Model Runtime Competition",
  "url": "/competitions/predict-ai-model-runtime/discussion/456343",
  "author_name": "Eduardo Rocha de Andrade",
  "post_date": "2023-11-19T15:44:57.215000",
  "votes": 68,
  "comment_count": 27,
  "views": 0,
  "content": "<p>First of all, thanks Kaggle and Google for hosting this great competition. Also, I'd like to thanks my teammates and friends <a href=\"https://www.kaggle.com/tomirol\" target=\"_blank\">@tomirol</a> and <a href=\"https://www.kaggle.com/thanhhau097a\" target=\"_blank\">@thanhhau097a</a>, which made this journey even more special.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/data</a></li>\n</ul>\n<h1>TLDR:</h1>\n<ul>\n<li>We pruned and compressed the layout graphs in order to increase the efficiency of our experiments</li>\n<li>We removed duplicated configs for layout</li>\n<li>We changed the <code>node_feat</code> to use -1 padding instead of 0</li>\n<li>We used the provided <code>train</code>, <code>val</code>, <code>test</code> splits since we found good correlation with LB</li>\n<li>Data preprocessing:<ul>\n<li>StandardScaler for <code>node_feat[:134]</code></li>\n<li>Shared learned embedding (4 channels) for <code>node_feat[134:]</code> and <code>node_config_feat</code></li>\n<li>Features are concatenated before first linear</li></ul></li>\n<li>Models with the following architecture:<ul>\n<li>Linear on the input features</li>\n<li>2x graph conv with attention blocks <ul>\n<li><code>InstanceNorm</code> -&gt; <code>SAGEConv</code> -&gt; <code>SelfChannelAttetion</code> -&gt; <code>CrossConfigAttetion</code> -&gt; <code>+residual</code> -&gt; <code>GELU</code></li></ul></li>\n<li>Global (graph) mean pooling</li>\n<li>Linear logit layer</li></ul></li>\n<li><code>PairwiseHingeLoss</code> function</li>\n</ul>\n<h1>Data preparation</h1>\n<p>We joined the competition with a bit less than a month to finish, so one of the first problems we tried to tackle was the  low efficiency in our training jobs.</p>\n<h2>Graph pruning</h2>\n<p>For layout, we noticed that only <code>Convolution</code>, <code>Dot</code> and <code>Reshape</code> were configurable nodes. Also, in most cases, the majority of nodes would be identical across the config set. Thus, we opted for a very simple pruning strategy where, for each graph, we would only keep the nodes that were either configurable models themselves or were connected to a configurable node, i.e., input or output to a configurable node. By doing this, we would transform a single big graph into multiple (possibly disconnected) sub-graphs, which was not a problem since the network has a global graph pooling layer in the end that fuses the sub-graphs information. This simple trick reduced 4 times the vRAM usage and sped up training by a factor of 5 in some cases.</p>\n<h2>Deduplication</h2>\n<p>Most of the configuration sets for layout contain a lot of duplication. However, the runtime for the duplicated configs can vary quite a bit and make training less stable. Thus, we opted for removing all the duplicated configs for layout.</p>\n<h2>Compression</h2>\n<p>Even with pruning and de-duplication, the RAM usage to load all configs to memory for NLP collection was super high. We circumvent that issue by compressing <code>node_config_feat</code> beforehand and only decompressing it on-the-fly in the dataloader after config sampling. This enabled us to load all data to memory at the beginning of training, which reduced IO/CPU bottlenecks considerably and allowed us to train faster.</p>\n<p>The idea behind the compression is that each <code>node_config_feat</code> 6-dim vector (input, output and kernel) can only have 7 possible values (-1, 0, 1, 2, 3, 4, 5) and, thus, can be represented by a single integer in base-7 (from 0 to 7^6).</p>\n<h2>Changing pad value in <code>node_feat</code></h2>\n<p>We noted that the features in <code>node_feat</code> were 0 padded. Whilst this is not a problem for most features, for others like <code>layout_minor_to_major_*</code> this can be ambiguous since 0 is a valid axis index. Also, the <code>node_config_feat</code> are -1 padded, which makes it incompatible with <code>layout_minor_to_major_*</code> from <code>node_feat</code>. With that in mind, we re-generated <code>node_feat</code> with -1 padded and this allowed us to use a single embedding matrix for both <code>node_feat[134:]</code> and <code>node_config_feat</code>.</p>\n<h1>Data preprocessing</h1>\n<p>For layout, we split <code>node_feat</code> into <code>node_feat[:134]</code> and <code>node_feat[134:]</code> (<code>layout_minor_to_major_*</code>). The former was simply normalised using a <code>StandardScaler</code>, while the latter, along with <code>node_config_feat</code>, was fed into a learned embedding matrix (4 channels). We found that the normalisation is essential here since <code>node_feat</code> has features like <code>*_sum</code> and <code>*_product</code> that can be very high and, consequently, disrupt the optimisation.</p>\n<p>For <code>node_opcode</code>, we also used a separate embedding layer with 16 channels. The input to the network is the concatenation of all features aforementioned and, for each graph, we sample on-the-fly 64 (default) or 128 (random) configs to form the input batch. For tile, on the other hand, we opt to use late fusion to integrate <code>config_feat</code> into the network.</p>\n<h1>Network architecture</h1>\n<p>Our network architecture was quite simple. We first feed the input features to a Linear block to map it to a 256d embedding vector followed by 2x Conv blocks, global graph mean pooling and a final linear layer.</p>\n<p>As for the graph convolutional layer itself, we tried many types but none was better than <code>SAGEConv</code>. In particular, I had good experience with GAT variants in the past but none worked well in this competition. If I were to guess the reason, I'd say that in the other application the graph itself was quite noisy, so attention helped to \"ignore\" the connections that were not meaningful. However, for TPU graphs, all connections are \"real\" and important so graph attention was not that helpful. Nonetheless, we found two other types of attention that were useful: self-channel attention and cross-config attention.</p>\n<h2>Self-Channel Attention</h2>\n<p>We borrowed the idea from Squeeze-and-Excitation to create a channel-wise attention layer. We first apply a Linear layer to bottleneck the channel dimensions (8x reduction) followed by <code>ReLU</code>. Then, we applied a second linear layer to increase the channels again to the original value followed by sigmoid. We finish by applying element wise multiplication on the obtained feature map and the original input.</p>\n<p>The idea behind this is to capture the correlations between channels and use it to suppress less useful ones while enhancing others.  </p>\n<h2>Cross-Config Attention</h2>\n<p>Another dimension that we can exploit attention is the batch plane (cross-configs). We designed a very simple block that allows the model to explicitly \"compare\" each config against the others throughout the network. We found this to be much better than letting the model infer for each config individually and only compare them implicitly via the loss function (<code>PairwiseHingeLoss</code>). The attention code is as follows:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .temperature = nn.Parameter(torch.tensor())\n\n     ():\n        \n        scores = (x / .temperature).softmax(dim=)\n        x = x * scores\n         x\n</code></pre>\n<p>By applying this simple layer after the self-channel attention at every block of the network, it gave us a huge boost for default collections. For inference, we simply use a reasonably large batch size of 128. However, since the prediction depends on the batch, we can leverage it further by applying TTA to generate <code>N</code> (10) permutations of the configs and average the result after sorting it back to the original order.</p>\n<h2>Linear/Conv blocks design</h2>\n<p>To create our Linear/Conv blocks we followed the good practices in computer vision. We start by using <code>InstanceNorm</code> to normalise the input feature map, followed by <code>Linear</code>/<code>SAGEConv</code> layer, <code>SelfChannelAttetion</code> and  <code>CrossConfigAttetion</code> (we concat the output with its input to preserve the individuality of each sample). Then, we sum the residual connection and finish with <code>GELU</code> and dropout.</p>\n<h1>Ensembling</h1>\n<p>Our best single model prediction scored 0.714 (0.748) on private (public) LB. However, since the number of test samples is quite low for some collections, we opted to use ensembles to improve the results and prevent shaking up. Our best result was 0.736 (0.757) LB by using the simple average of 5-10 models for each collection but we sadly didn't select this sub.</p>\n<h1>Things that didn't work</h1>\n<ul>\n<li>Train together on both random and default data</li>\n<li>2nd level stacking on top of model's embeddings/predictions (worked well locally but not that well on LB)</li>\n<li>Pseudo labelling test set</li>\n<li>Finetunning models for specific graphs in test set (worked well locally but not that well on LB)</li>\n<li>Different ways to sample the configs, e.g., annealing the runtime spacing between configs</li>\n<li>Other loss functions</li>\n<li>Gradient accumulation and training with more than one graph at a time</li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li>Our code can be found <a href=\"https://github.com/thanhhau097/google_fast_or_slow/tree/main\" target=\"_blank\">in this GitHub repo</a></li>\n<li>Our work was based/forked on <a href=\"https://www.kaggle.com/werus23\" target=\"_blank\">@werus23</a> <a href=\"https://www.kaggle.com/code/werus23/tile-xla-end-to-end-train-infer\" target=\"_blank\">public kernel</a> (thank you!)</li>\n<li>The modified layout dataset with -1 padding can be found <a href=\"https://www.kaggle.com/datasets/tomirol/layout-npz-padding\" target=\"_blank\">here</a></li>\n</ul>\n<p>I made a quick diagram of our network, it's a bit crap but should give a good idea 😅</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1648129%2F7cf79a021748ba6f4b38529447419992%2Fwriteup_graph_model.png?generation=1700408602690996&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2530875,
      "postDate": "2023-11-19T15:44:57.217Z",
      "content": "<p>First of all, thanks Kaggle and Google for hosting this great competition. Also, I'd like to thanks my teammates and friends <a href=\"https://www.kaggle.com/tomirol\" target=\"_blank\">@tomirol</a> and <a href=\"https://www.kaggle.com/thanhhau097a\" target=\"_blank\">@thanhhau097a</a>, which made this journey even more special.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/data</a></li>\n</ul>\n<h1>TLDR:</h1>\n<ul>\n<li>We pruned and compressed the layout graphs in order to increase the efficiency of our experiments</li>\n<li>We removed duplicated configs for layout</li>\n<li>We changed the <code>node_feat</code> to use -1 padding instead of 0</li>\n<li>We used the provided <code>train</code>, <code>val</code>, <code>test</code> splits since we found good correlation with LB</li>\n<li>Data preprocessing:<ul>\n<li>StandardScaler for <code>node_feat[:134]</code></li>\n<li>Shared learned embedding (4 channels) for <code>node_feat[134:]</code> and <code>node_config_feat</code></li>\n<li>Features are concatenated before first linear</li></ul></li>\n<li>Models with the following architecture:<ul>\n<li>Linear on the input features</li>\n<li>2x graph conv with attention blocks <ul>\n<li><code>InstanceNorm</code> -&gt; <code>SAGEConv</code> -&gt; <code>SelfChannelAttetion</code> -&gt; <code>CrossConfigAttetion</code> -&gt; <code>+residual</code> -&gt; <code>GELU</code></li></ul></li>\n<li>Global (graph) mean pooling</li>\n<li>Linear logit layer</li></ul></li>\n<li><code>PairwiseHingeLoss</code> function</li>\n</ul>\n<h1>Data preparation</h1>\n<p>We joined the competition with a bit less than a month to finish, so one of the first problems we tried to tackle was the  low efficiency in our training jobs.</p>\n<h2>Graph pruning</h2>\n<p>For layout, we noticed that only <code>Convolution</code>, <code>Dot</code> and <code>Reshape</code> were configurable nodes. Also, in most cases, the majority of nodes would be identical across the config set. Thus, we opted for a very simple pruning strategy where, for each graph, we would only keep the nodes that were either configurable models themselves or were connected to a configurable node, i.e., input or output to a configurable node. By doing this, we would transform a single big graph into multiple (possibly disconnected) sub-graphs, which was not a problem since the network has a global graph pooling layer in the end that fuses the sub-graphs information. This simple trick reduced 4 times the vRAM usage and sped up training by a factor of 5 in some cases.</p>\n<h2>Deduplication</h2>\n<p>Most of the configuration sets for layout contain a lot of duplication. However, the runtime for the duplicated configs can vary quite a bit and make training less stable. Thus, we opted for removing all the duplicated configs for layout.</p>\n<h2>Compression</h2>\n<p>Even with pruning and de-duplication, the RAM usage to load all configs to memory for NLP collection was super high. We circumvent that issue by compressing <code>node_config_feat</code> beforehand and only decompressing it on-the-fly in the dataloader after config sampling. This enabled us to load all data to memory at the beginning of training, which reduced IO/CPU bottlenecks considerably and allowed us to train faster.</p>\n<p>The idea behind the compression is that each <code>node_config_feat</code> 6-dim vector (input, output and kernel) can only have 7 possible values (-1, 0, 1, 2, 3, 4, 5) and, thus, can be represented by a single integer in base-7 (from 0 to 7^6).</p>\n<h2>Changing pad value in <code>node_feat</code></h2>\n<p>We noted that the features in <code>node_feat</code> were 0 padded. Whilst this is not a problem for most features, for others like <code>layout_minor_to_major_*</code> this can be ambiguous since 0 is a valid axis index. Also, the <code>node_config_feat</code> are -1 padded, which makes it incompatible with <code>layout_minor_to_major_*</code> from <code>node_feat</code>. With that in mind, we re-generated <code>node_feat</code> with -1 padded and this allowed us to use a single embedding matrix for both <code>node_feat[134:]</code> and <code>node_config_feat</code>.</p>\n<h1>Data preprocessing</h1>\n<p>For layout, we split <code>node_feat</code> into <code>node_feat[:134]</code> and <code>node_feat[134:]</code> (<code>layout_minor_to_major_*</code>). The former was simply normalised using a <code>StandardScaler</code>, while the latter, along with <code>node_config_feat</code>, was fed into a learned embedding matrix (4 channels). We found that the normalisation is essential here since <code>node_feat</code> has features like <code>*_sum</code> and <code>*_product</code> that can be very high and, consequently, disrupt the optimisation.</p>\n<p>For <code>node_opcode</code>, we also used a separate embedding layer with 16 channels. The input to the network is the concatenation of all features aforementioned and, for each graph, we sample on-the-fly 64 (default) or 128 (random) configs to form the input batch. For tile, on the other hand, we opt to use late fusion to integrate <code>config_feat</code> into the network.</p>\n<h1>Network architecture</h1>\n<p>Our network architecture was quite simple. We first feed the input features to a Linear block to map it to a 256d embedding vector followed by 2x Conv blocks, global graph mean pooling and a final linear layer.</p>\n<p>As for the graph convolutional layer itself, we tried many types but none was better than <code>SAGEConv</code>. In particular, I had good experience with GAT variants in the past but none worked well in this competition. If I were to guess the reason, I'd say that in the other application the graph itself was quite noisy, so attention helped to \"ignore\" the connections that were not meaningful. However, for TPU graphs, all connections are \"real\" and important so graph attention was not that helpful. Nonetheless, we found two other types of attention that were useful: self-channel attention and cross-config attention.</p>\n<h2>Self-Channel Attention</h2>\n<p>We borrowed the idea from Squeeze-and-Excitation to create a channel-wise attention layer. We first apply a Linear layer to bottleneck the channel dimensions (8x reduction) followed by <code>ReLU</code>. Then, we applied a second linear layer to increase the channels again to the original value followed by sigmoid. We finish by applying element wise multiplication on the obtained feature map and the original input.</p>\n<p>The idea behind this is to capture the correlations between channels and use it to suppress less useful ones while enhancing others.  </p>\n<h2>Cross-Config Attention</h2>\n<p>Another dimension that we can exploit attention is the batch plane (cross-configs). We designed a very simple block that allows the model to explicitly \"compare\" each config against the others throughout the network. We found this to be much better than letting the model infer for each config individually and only compare them implicitly via the loss function (<code>PairwiseHingeLoss</code>). The attention code is as follows:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .temperature = nn.Parameter(torch.tensor())\n\n     ():\n        \n        scores = (x / .temperature).softmax(dim=)\n        x = x * scores\n         x\n</code></pre>\n<p>By applying this simple layer after the self-channel attention at every block of the network, it gave us a huge boost for default collections. For inference, we simply use a reasonably large batch size of 128. However, since the prediction depends on the batch, we can leverage it further by applying TTA to generate <code>N</code> (10) permutations of the configs and average the result after sorting it back to the original order.</p>\n<h2>Linear/Conv blocks design</h2>\n<p>To create our Linear/Conv blocks we followed the good practices in computer vision. We start by using <code>InstanceNorm</code> to normalise the input feature map, followed by <code>Linear</code>/<code>SAGEConv</code> layer, <code>SelfChannelAttetion</code> and  <code>CrossConfigAttetion</code> (we concat the output with its input to preserve the individuality of each sample). Then, we sum the residual connection and finish with <code>GELU</code> and dropout.</p>\n<h1>Ensembling</h1>\n<p>Our best single model prediction scored 0.714 (0.748) on private (public) LB. However, since the number of test samples is quite low for some collections, we opted to use ensembles to improve the results and prevent shaking up. Our best result was 0.736 (0.757) LB by using the simple average of 5-10 models for each collection but we sadly didn't select this sub.</p>\n<h1>Things that didn't work</h1>\n<ul>\n<li>Train together on both random and default data</li>\n<li>2nd level stacking on top of model's embeddings/predictions (worked well locally but not that well on LB)</li>\n<li>Pseudo labelling test set</li>\n<li>Finetunning models for specific graphs in test set (worked well locally but not that well on LB)</li>\n<li>Different ways to sample the configs, e.g., annealing the runtime spacing between configs</li>\n<li>Other loss functions</li>\n<li>Gradient accumulation and training with more than one graph at a time</li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li>Our code can be found <a href=\"https://github.com/thanhhau097/google_fast_or_slow/tree/main\" target=\"_blank\">in this GitHub repo</a></li>\n<li>Our work was based/forked on <a href=\"https://www.kaggle.com/werus23\" target=\"_blank\">@werus23</a> <a href=\"https://www.kaggle.com/code/werus23/tile-xla-end-to-end-train-infer\" target=\"_blank\">public kernel</a> (thank you!)</li>\n<li>The modified layout dataset with -1 padding can be found <a href=\"https://www.kaggle.com/datasets/tomirol/layout-npz-padding\" target=\"_blank\">here</a></li>\n</ul>\n<p>I made a quick diagram of our network, it's a bit crap but should give a good idea 😅</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1648129%2F7cf79a021748ba6f4b38529447419992%2Fwriteup_graph_model.png?generation=1700408602690996&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "First of all, thanks Kaggle and Google for hosting this great competition. Also, I'd like to thanks my teammates and friends @tomirol and @thanhhau097a, which made this journey even more special.\n\n# Context\n- Business context: https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\n- Data context: https://www.kaggle.com/competitions/predict-ai-model-runtime/data\n\n# TLDR:\n- We pruned and compressed the layout graphs in order to increase the efficiency of our experiments\n- We removed duplicated configs for layout\n- We changed the `node_feat` to use -1 padding instead of 0\n- We used the provided `train`, `val`, `test` splits since we found good correlation with LB\n- Data preprocessing:\n    - StandardScaler for `node_feat[:134]`\n    - Shared learned embedding (4 channels) for `node_feat[134:]` and `node_config_feat`\n    - Features are concatenated before first linear\n- Models with the following architecture:\n    - Linear on the input features\n    - 2x graph conv with attention blocks \n        - `InstanceNorm` -> `SAGEConv` -> `SelfChannelAttetion` -> `CrossConfigAttetion` -> `+residual` -> `GELU`\n    - Global (graph) mean pooling\n    - Linear logit layer\n- `PairwiseHingeLoss` function\n\n# Data preparation\nWe joined the competition with a bit less than a month to finish, so one of the first problems we tried to tackle was the  low efficiency in our training jobs.\n\n## Graph pruning\nFor layout, we noticed that only `Convolution`, `Dot` and `Reshape` were configurable nodes. Also, in most cases, the majority of nodes would be identical across the config set. Thus, we opted for a very simple pruning strategy where, for each graph, we would only keep the nodes that were either configurable models themselves or were connected to a configurable node, i.e., input or output to a configurable node. By doing this, we would transform a single big graph into multiple (possibly disconnected) sub-graphs, which was not a problem since the network has a global graph pooling layer in the end that fuses the sub-graphs information. This simple trick reduced 4 times the vRAM usage and sped up training by a factor of 5 in some cases.\n\n## Deduplication\nMost of the configuration sets for layout contain a lot of duplication. However, the runtime for the duplicated configs can vary quite a bit and make training less stable. Thus, we opted for removing all the duplicated configs for layout.\n\n## Compression\nEven with pruning and de-duplication, the RAM usage to load all configs to memory for NLP collection was super high. We circumvent that issue by compressing `node_config_feat` beforehand and only decompressing it on-the-fly in the dataloader after config sampling. This enabled us to load all data to memory at the beginning of training, which reduced IO/CPU bottlenecks considerably and allowed us to train faster.\n\nThe idea behind the compression is that each `node_config_feat` 6-dim vector (input, output and kernel) can only have 7 possible values (-1, 0, 1, 2, 3, 4, 5) and, thus, can be represented by a single integer in base-7 (from 0 to 7^6).\n\n## Changing pad value in `node_feat`\nWe noted that the features in `node_feat` were 0 padded. Whilst this is not a problem for most features, for others like `layout_minor_to_major_*` this can be ambiguous since 0 is a valid axis index. Also, the `node_config_feat` are -1 padded, which makes it incompatible with `layout_minor_to_major_*` from `node_feat`. With that in mind, we re-generated `node_feat` with -1 padded and this allowed us to use a single embedding matrix for both `node_feat[134:]` and `node_config_feat`.\n\n# Data preprocessing\nFor layout, we split `node_feat` into `node_feat[:134]` and `node_feat[134:]` (`layout_minor_to_major_*`). The former was simply normalised using a `StandardScaler`, while the latter, along with `node_config_feat`, was fed into a learned embedding matrix (4 channels). We found that the normalisation is essential here since `node_feat` has features like `*_sum` and `*_product` that can be very high and, consequently, disrupt the optimisation.\n\nFor `node_opcode`, we also used a separate embedding layer with 16 channels. The input to the network is the concatenation of all features aforementioned and, for each graph, we sample on-the-fly 64 (default) or 128 (random) configs to form the input batch. For tile, on the other hand, we opt to use late fusion to integrate `config_feat` into the network.\n\n# Network architecture\nOur network architecture was quite simple. We first feed the input features to a Linear block to map it to a 256d embedding vector followed by 2x Conv blocks, global graph mean pooling and a final linear layer.\n\nAs for the graph convolutional layer itself, we tried many types but none was better than `SAGEConv`. In particular, I had good experience with GAT variants in the past but none worked well in this competition. If I were to guess the reason, I'd say that in the other application the graph itself was quite noisy, so attention helped to \"ignore\" the connections that were not meaningful. However, for TPU graphs, all connections are \"real\" and important so graph attention was not that helpful. Nonetheless, we found two other types of attention that were useful: self-channel attention and cross-config attention.\n\n## Self-Channel Attention\nWe borrowed the idea from Squeeze-and-Excitation to create a channel-wise attention layer. We first apply a Linear layer to bottleneck the channel dimensions (8x reduction) followed by `ReLU`. Then, we applied a second linear layer to increase the channels again to the original value followed by sigmoid. We finish by applying element wise multiplication on the obtained feature map and the original input.\n\nThe idea behind this is to capture the correlations between channels and use it to suppress less useful ones while enhancing others.  \n\n## Cross-Config Attention\nAnother dimension that we can exploit attention is the batch plane (cross-configs). We designed a very simple block that allows the model to explicitly \"compare\" each config against the others throughout the network. We found this to be much better than letting the model infer for each config individually and only compare them implicitly via the loss function (`PairwiseHingeLoss`). The attention code is as follows:\n```\nclass CrossConfigAttention(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.temperature = nn.Parameter(torch.tensor(0.5))\n\n    def forward(self, x):\n        # x of shape (nb_configs, nb_nodes, nb_features)\n        scores = (x / self.temperature).softmax(dim=0)\n        x = x * scores\n        return x\n```\n\nBy applying this simple layer after the self-channel attention at every block of the network, it gave us a huge boost for default collections. For inference, we simply use a reasonably large batch size of 128. However, since the prediction depends on the batch, we can leverage it further by applying TTA to generate `N` (10) permutations of the configs and average the result after sorting it back to the original order.\n\n## Linear/Conv blocks design\nTo create our Linear/Conv blocks we followed the good practices in computer vision. We start by using `InstanceNorm` to normalise the input feature map, followed by `Linear`/`SAGEConv` layer, `SelfChannelAttetion` and  `CrossConfigAttetion` (we concat the output with its input to preserve the individuality of each sample). Then, we sum the residual connection and finish with `GELU` and dropout.\n\n# Ensembling\nOur best single model prediction scored 0.714 (0.748) on private (public) LB. However, since the number of test samples is quite low for some collections, we opted to use ensembles to improve the results and prevent shaking up. Our best result was 0.736 (0.757) LB by using the simple average of 5-10 models for each collection but we sadly didn't select this sub.\n\n# Things that didn't work\n- Train together on both random and default data\n- 2nd level stacking on top of model's embeddings/predictions (worked well locally but not that well on LB)\n- Pseudo labelling test set\n- Finetunning models for specific graphs in test set (worked well locally but not that well on LB)\n- Different ways to sample the configs, e.g., annealing the runtime spacing between configs\n- Other loss functions\n- Gradient accumulation and training with more than one graph at a time\n\n# Sources\n- Our code can be found [in this GitHub repo](https://github.com/thanhhau097/google_fast_or_slow/tree/main)\n- Our work was based/forked on @werus23 [public kernel](https://www.kaggle.com/code/werus23/tile-xla-end-to-end-train-infer) (thank you!)\n- The modified layout dataset with -1 padding can be found [here](https://www.kaggle.com/datasets/tomirol/layout-npz-padding)\n\nI made a quick diagram of our network, it's a bit crap but should give a good idea 😅\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1648129%2F7cf79a021748ba6f4b38529447419992%2Fwriteup_graph_model.png?generation=1700408602690996&alt=media)",
      "votes": 68
    },
    {
      "id": 2537036,
      "postDate": "2023-11-24T17:18:51.283Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/arc144\" target=\"_blank\">@arc144</a> and team! Nice to see a brazillian winner 😉<br>\nThanks for sharing.</p>",
      "rawMarkdown": "Congrats @arc144 and team! Nice to see a brazillian winner 😉\nThanks for sharing.",
      "votes": 3,
      "replies": [
        {
          "id": 2537204,
          "postDate": "2023-11-24T22:24:21.030Z",
          "content": "<p>Obrigado Giba! 🙇</p>",
          "rawMarkdown": "Obrigado Giba! 🙇",
          "votes": 2
        }
      ]
    },
    {
      "id": 2531509,
      "postDate": "2023-11-20T09:18:37.973Z",
      "content": "<p>Congrats on the 1st place! I am curious: what tool did you use to make this beautiful block diagram in the writeup?</p>",
      "rawMarkdown": "Congrats on the 1st place! I am curious: what tool did you use to make this beautiful block diagram in the writeup?",
      "votes": 3,
      "replies": [
        {
          "id": 2531680,
          "postDate": "2023-11-20T12:24:44.493Z",
          "content": "<p>Thank you Dmitrii! Congratulation to you and your team as well!<br>\nI used a website called <a href=\"https://app.diagrams.net/\" target=\"_blank\">drawIo</a></p>",
          "rawMarkdown": "Thank you Dmitrii! Congratulation to you and your team as well!\nI used a website called [drawIo](https://app.diagrams.net/)",
          "votes": 3
        }
      ]
    },
    {
      "id": 2530966,
      "postDate": "2023-11-19T17:41:49.437Z",
      "content": "<p>Hi Eduardo,<br>\nI'm so glad to read your write-up. I'm a beginner, so it's very hard to understand it all.</p>\n<p>Congratulations for you and your team (Thanh Hau Nguyen and Tomirol) for the amazing achievement.<br>\nParabéns!</p>",
      "rawMarkdown": "Hi Eduardo,\nI'm so glad to read your write-up. I'm a beginner, so it's very hard to understand it all.\n\nCongratulations for you and your team (Thanh Hau Nguyen and Tomirol) for the amazing achievement.\nParabéns!",
      "votes": 4,
      "replies": [
        {
          "id": 2531165,
          "postDate": "2023-11-20T01:19:41.100Z",
          "content": "<p>Muito obrigado Marília 🙏</p>\n<p>Let me know if there is anything I can help clarifying!</p>",
          "rawMarkdown": "Muito obrigado Marília 🙏\n\nLet me know if there is anything I can help clarifying!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2606250,
      "postDate": "2024-01-17T14:11:39.653Z",
      "content": "<p>Congrats on winning the competition! :) Thanks for implementation details. If you could, kindly share training details, ex: GPU used, GPU training time etc.</p>",
      "rawMarkdown": "Congrats on winning the competition! :) Thanks for implementation details. If you could, kindly share training details, ex: GPU used, GPU training time etc.",
      "votes": 1,
      "replies": [
        {
          "id": 2674796,
          "postDate": "2024-02-29T14:45:53.720Z",
          "content": "<p>Sorry for the late response, was away recently. We used a RTX 4090 and it took around 1-2 hours (depends on the collection) to train the models after the optimisations we described in the thread.</p>",
          "rawMarkdown": "Sorry for the late response, was away recently. We used a RTX 4090 and it took around 1-2 hours (depends on the collection) to train the models after the optimisations we described in the thread.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2561136,
      "postDate": "2023-12-14T09:53:50.293Z",
      "content": "<p>congratulations!</p>",
      "rawMarkdown": "congratulations!",
      "votes": 1
    },
    {
      "id": 2534259,
      "postDate": "2023-11-22T14:07:37.950Z",
      "content": "<p>Great solution. Architecture ideas are quite amazing. Learning embeddings for configs' features is what I was missing :)</p>",
      "rawMarkdown": "Great solution. Architecture ideas are quite amazing. Learning embeddings for configs' features is what I was missing :)",
      "votes": 1
    },
    {
      "id": 2533774,
      "postDate": "2023-11-22T08:10:40.150Z",
      "content": "<p>Congratulations, first place! What tool did you use to create the graph in the last image?</p>",
      "rawMarkdown": "Congratulations, first place! What tool did you use to create the graph in the last image?",
      "votes": 1,
      "replies": [
        {
          "id": 2533970,
          "postDate": "2023-11-22T10:15:30.680Z",
          "content": "<p>Thank you! I use a tool called <a href=\"https://app.diagrams.net/\" target=\"_blank\">drawIo</a></p>",
          "rawMarkdown": "Thank you! I use a tool called [drawIo](https://app.diagrams.net/)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2533350,
      "postDate": "2023-11-21T20:09:01.367Z",
      "content": "<p>congratulations. </p>",
      "rawMarkdown": "congratulations. ",
      "votes": 1
    },
    {
      "id": 2532973,
      "postDate": "2023-11-21T13:49:32.747Z",
      "content": "<p>Congratulations and thanks for sharing! The solution helps me capture all the details.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! The solution helps me capture all the details.",
      "votes": 1
    },
    {
      "id": 2532336,
      "postDate": "2023-11-21T01:15:46.477Z",
      "content": "<p>Congrats for 1st place. Does an ensemble model use the same weight for each model?</p>",
      "rawMarkdown": "Congrats for 1st place. Does an ensemble model use the same weight for each model?",
      "votes": 1,
      "replies": [
        {
          "id": 2532736,
          "postDate": "2023-11-21T09:46:59.803Z",
          "content": "<p>Thank you! For the ensemble, we train the same model architecture many (5 - 10 usually) times with different random seeds (main impact is in the model's weights initialisation). Then, we just load each model/trained weight individually to make the prediction and average them up</p>",
          "rawMarkdown": "Thank you! For the ensemble, we train the same model architecture many (5 - 10 usually) times with different random seeds (main impact is in the model's weights initialisation). Then, we just load each model/trained weight individually to make the prediction and average them up",
          "votes": 1,
          "replies": [
            {
              "id": 2533377,
              "postDate": "2023-11-21T20:51:25.077Z",
              "content": "<p>Using different weights could sometimes improve accuracy. In another competition, I used the Optuna library to experiment with different weight combinations and found the optimal ones, leading to further accuracy gains. Congratulations again on your excellent solutions!</p>",
              "rawMarkdown": "Using different weights could sometimes improve accuracy. In another competition, I used the Optuna library to experiment with different weight combinations and found the optimal ones, leading to further accuracy gains. Congratulations again on your excellent solutions!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2532324,
      "postDate": "2023-11-21T00:34:13.307Z",
      "content": "<p>Congrats brother on achieving 1st place.</p>",
      "rawMarkdown": "Congrats brother on achieving 1st place.",
      "votes": 1
    },
    {
      "id": 2532248,
      "postDate": "2023-11-20T21:28:00.927Z",
      "content": "<p>Thanks for sharing, congrats for winning</p>",
      "rawMarkdown": "Thanks for sharing, congrats for winning",
      "votes": 1
    },
    {
      "id": 2531240,
      "postDate": "2023-11-20T03:53:28.827Z",
      "content": "<p>Congratulations on  topping the leaderboard. Thanks for sharing the details of your solution with flowcharts. </p>",
      "rawMarkdown": "Congratulations on  topping the leaderboard. Thanks for sharing the details of your solution with flowcharts. ",
      "votes": 1
    },
    {
      "id": 2531191,
      "postDate": "2023-11-20T02:28:38.813Z",
      "content": "<p>Congratulations for your team and learn lots from your solution!</p>",
      "rawMarkdown": "Congratulations for your team and learn lots from your solution!",
      "votes": 1
    },
    {
      "id": 2533324,
      "postDate": "2023-11-21T19:25:23.300Z",
      "content": "<p>Great competition solution! Your approach, especially using techniques like graph pruning and deduplication in data preprocessing, is quite interesting for achieving high performance. Additionally, your optimization of network architecture and attention mechanisms is very impressive. The results you obtained with ensembling are noteworthy. Congratulations! 🚀</p>",
      "rawMarkdown": "Great competition solution! Your approach, especially using techniques like graph pruning and deduplication in data preprocessing, is quite interesting for achieving high performance. Additionally, your optimization of network architecture and attention mechanisms is very impressive. The results you obtained with ensembling are noteworthy. Congratulations! 🚀",
      "votes": 2
    },
    {
      "id": 2533525,
      "postDate": "2023-11-22T02:21:27.817Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2532621,
      "postDate": "2023-11-21T07:30:48.087Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2532982,
      "postDate": "2023-11-21T13:59:45.577Z",
      "content": "<p>Good job! Thanks for the sharing!</p>",
      "rawMarkdown": "Good job! Thanks for the sharing!",
      "votes": 1
    },
    {
      "id": 2532390,
      "postDate": "2023-11-21T03:07:21.020Z",
      "content": "<p>congrats! Thanks for sharing</p>",
      "rawMarkdown": "congrats! Thanks for sharing",
      "votes": 1
    },
    {
      "id": 2531335,
      "postDate": "2023-11-20T06:05:09.580Z",
      "content": "<p>Thanks for sharing！</p>",
      "rawMarkdown": "Thanks for sharing！",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2537036,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2023-11-24T17:18:51.283000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/arc144\" target=\"_blank\">@arc144</a> and team! Nice to see a brazillian winner 😉<br>\nThanks for sharing.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2537204,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2023-11-24T22:24:21.030000",
          "content": "<p>Obrigado Giba! 🙇</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2531509,
      "author_name": "Dmitrii Khizbullin",
      "author_url": "",
      "post_date": "2023-11-20T09:18:37.973000",
      "content": "<p>Congrats on the 1st place! I am curious: what tool did you use to make this beautiful block diagram in the writeup?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2531680,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2023-11-20T12:24:44.493000",
          "content": "<p>Thank you Dmitrii! Congratulation to you and your team as well!<br>\nI used a website called <a href=\"https://app.diagrams.net/\" target=\"_blank\">drawIo</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2530966,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2023-11-19T17:41:49.437000",
      "content": "<p>Hi Eduardo,<br>\nI'm so glad to read your write-up. I'm a beginner, so it's very hard to understand it all.</p>\n<p>Congratulations for you and your team (Thanh Hau Nguyen and Tomirol) for the amazing achievement.<br>\nParabéns!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2531165,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2023-11-20T01:19:41.100000",
          "content": "<p>Muito obrigado Marília 🙏</p>\n<p>Let me know if there is anything I can help clarifying!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2606250,
      "author_name": "Kilaru Vasudeva",
      "author_url": "",
      "post_date": "2024-01-17T14:11:39.653000",
      "content": "<p>Congrats on winning the competition! :) Thanks for implementation details. If you could, kindly share training details, ex: GPU used, GPU training time etc.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2674796,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2024-02-29T14:45:53.720000",
          "content": "<p>Sorry for the late response, was away recently. We used a RTX 4090 and it took around 1-2 hours (depends on the collection) to train the models after the optimisations we described in the thread.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2561136,
      "author_name": "Filtered",
      "author_url": "",
      "post_date": "2023-12-14T09:53:50.293000",
      "content": "<p>congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2534259,
      "author_name": "belgraviton",
      "author_url": "",
      "post_date": "2023-11-22T14:07:37.950000",
      "content": "<p>Great solution. Architecture ideas are quite amazing. Learning embeddings for configs' features is what I was missing :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2533774,
      "author_name": "hyohyolim",
      "author_url": "",
      "post_date": "2023-11-22T08:10:40.150000",
      "content": "<p>Congratulations, first place! What tool did you use to create the graph in the last image?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2533970,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2023-11-22T10:15:30.680000",
          "content": "<p>Thank you! I use a tool called <a href=\"https://app.diagrams.net/\" target=\"_blank\">drawIo</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2533350,
      "author_name": "Md. Amanat Ullah",
      "author_url": "",
      "post_date": "2023-11-21T20:09:01.367000",
      "content": "<p>congratulations. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2532973,
      "author_name": "Suma Mallapragada",
      "author_url": "",
      "post_date": "2023-11-21T13:49:32.747000",
      "content": "<p>Congratulations and thanks for sharing! The solution helps me capture all the details.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2532336,
      "author_name": "Min-Hsien Weng",
      "author_url": "",
      "post_date": "2023-11-21T01:15:46.477000",
      "content": "<p>Congrats for 1st place. Does an ensemble model use the same weight for each model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2532736,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2023-11-21T09:46:59.803000",
          "content": "<p>Thank you! For the ensemble, we train the same model architecture many (5 - 10 usually) times with different random seeds (main impact is in the model's weights initialisation). Then, we just load each model/trained weight individually to make the prediction and average them up</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2533377,
              "author_name": "Min-Hsien Weng",
              "author_url": "",
              "post_date": "2023-11-21T20:51:25.077000",
              "content": "<p>Using different weights could sometimes improve accuracy. In another competition, I used the Optuna library to experiment with different weight combinations and found the optimal ones, leading to further accuracy gains. Congratulations again on your excellent solutions!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2532324,
      "author_name": "Indra Sonowal",
      "author_url": "",
      "post_date": "2023-11-21T00:34:13.307000",
      "content": "<p>Congrats brother on achieving 1st place.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2532248,
      "author_name": "Haohan Tsao",
      "author_url": "",
      "post_date": "2023-11-20T21:28:00.927000",
      "content": "<p>Thanks for sharing, congrats for winning</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2531240,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-11-20T03:53:28.827000",
      "content": "<p>Congratulations on  topping the leaderboard. Thanks for sharing the details of your solution with flowcharts. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2531191,
      "author_name": "yy1165",
      "author_url": "",
      "post_date": "2023-11-20T02:28:38.813000",
      "content": "<p>Congratulations for your team and learn lots from your solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2533324,
      "author_name": "Semih Günak",
      "author_url": "",
      "post_date": "2023-11-21T19:25:23.300000",
      "content": "<p>Great competition solution! Your approach, especially using techniques like graph pruning and deduplication in data preprocessing, is quite interesting for achieving high performance. Additionally, your optimization of network architecture and attention mechanisms is very impressive. The results you obtained with ensembling are noteworthy. Congratulations! 🚀</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2533525,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-22T02:21:27.817000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2532621,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-21T07:30:48.087000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2532982,
      "author_name": "Zhuravlev Viktor",
      "author_url": "",
      "post_date": "2023-11-21T13:59:45.577000",
      "content": "<p>Good job! Thanks for the sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2532390,
      "author_name": "Phil",
      "author_url": "",
      "post_date": "2023-11-21T03:07:21.020000",
      "content": "<p>congrats! Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2531335,
      "author_name": "Haoqin Hong",
      "author_url": "",
      "post_date": "2023-11-20T06:05:09.580000",
      "content": "<p>Thanks for sharing！</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2530875": "First of all, thanks Kaggle and Google for hosting this great competition. Also, I'd like to thanks my teammates and friends @tomirol and @thanhhau097a, which made this journey even more special.\n\n# Context\n- Business context: https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\n- Data context: https://www.kaggle.com/competitions/predict-ai-model-runtime/data\n\n# TLDR:\n- We pruned and compressed the layout graphs in order to increase the efficiency of our experiments\n- We removed duplicated configs for layout\n- We changed the `node_feat` to use -1 padding instead of 0\n- We used the provided `train`, `val`, `test` splits since we found good correlation with LB\n- Data preprocessing:\n    - StandardScaler for `node_feat[:134]`\n    - Shared learned embedding (4 channels) for `node_feat[134:]` and `node_config_feat`\n    - Features are concatenated before first linear\n- Models with the following architecture:\n    - Linear on the input features\n    - 2x graph conv with attention blocks \n        - `InstanceNorm` -> `SAGEConv` -> `SelfChannelAttetion` -> `CrossConfigAttetion` -> `+residual` -> `GELU`\n    - Global (graph) mean pooling\n    - Linear logit layer\n- `PairwiseHingeLoss` function\n\n# Data preparation\nWe joined the competition with a bit less than a month to finish, so one of the first problems we tried to tackle was the  low efficiency in our training jobs.\n\n## Graph pruning\nFor layout, we noticed that only `Convolution`, `Dot` and `Reshape` were configurable nodes. Also, in most cases, the majority of nodes would be identical across the config set. Thus, we opted for a very simple pruning strategy where, for each graph, we would only keep the nodes that were either configurable models themselves or were connected to a configurable node, i.e., input or output to a configurable node. By doing this, we would transform a single big graph into multiple (possibly disconnected) sub-graphs, which was not a problem since the network has a global graph pooling layer in the end that fuses the sub-graphs information. This simple trick reduced 4 times the vRAM usage and sped up training by a factor of 5 in some cases.\n\n## Deduplication\nMost of the configuration sets for layout contain a lot of duplication. However, the runtime for the duplicated configs can vary quite a bit and make training less stable. Thus, we opted for removing all the duplicated configs for layout.\n\n## Compression\nEven with pruning and de-duplication, the RAM usage to load all configs to memory for NLP collection was super high. We circumvent that issue by compressing `node_config_feat` beforehand and only decompressing it on-the-fly in the dataloader after config sampling. This enabled us to load all data to memory at the beginning of training, which reduced IO/CPU bottlenecks considerably and allowed us to train faster.\n\nThe idea behind the compression is that each `node_config_feat` 6-dim vector (input, output and kernel) can only have 7 possible values (-1, 0, 1, 2, 3, 4, 5) and, thus, can be represented by a single integer in base-7 (from 0 to 7^6).\n\n## Changing pad value in `node_feat`\nWe noted that the features in `node_feat` were 0 padded. Whilst this is not a problem for most features, for others like `layout_minor_to_major_*` this can be ambiguous since 0 is a valid axis index. Also, the `node_config_feat` are -1 padded, which makes it incompatible with `layout_minor_to_major_*` from `node_feat`. With that in mind, we re-generated `node_feat` with -1 padded and this allowed us to use a single embedding matrix for both `node_feat[134:]` and `node_config_feat`.\n\n# Data preprocessing\nFor layout, we split `node_feat` into `node_feat[:134]` and `node_feat[134:]` (`layout_minor_to_major_*`). The former was simply normalised using a `StandardScaler`, while the latter, along with `node_config_feat`, was fed into a learned embedding matrix (4 channels). We found that the normalisation is essential here since `node_feat` has features like `*_sum` and `*_product` that can be very high and, consequently, disrupt the optimisation.\n\nFor `node_opcode`, we also used a separate embedding layer with 16 channels. The input to the network is the concatenation of all features aforementioned and, for each graph, we sample on-the-fly 64 (default) or 128 (random) configs to form the input batch. For tile, on the other hand, we opt to use late fusion to integrate `config_feat` into the network.\n\n# Network architecture\nOur network architecture was quite simple. We first feed the input features to a Linear block to map it to a 256d embedding vector followed by 2x Conv blocks, global graph mean pooling and a final linear layer.\n\nAs for the graph convolutional layer itself, we tried many types but none was better than `SAGEConv`. In particular, I had good experience with GAT variants in the past but none worked well in this competition. If I were to guess the reason, I'd say that in the other application the graph itself was quite noisy, so attention helped to \"ignore\" the connections that were not meaningful. However, for TPU graphs, all connections are \"real\" and important so graph attention was not that helpful. Nonetheless, we found two other types of attention that were useful: self-channel attention and cross-config attention.\n\n## Self-Channel Attention\nWe borrowed the idea from Squeeze-and-Excitation to create a channel-wise attention layer. We first apply a Linear layer to bottleneck the channel dimensions (8x reduction) followed by `ReLU`. Then, we applied a second linear layer to increase the channels again to the original value followed by sigmoid. We finish by applying element wise multiplication on the obtained feature map and the original input.\n\nThe idea behind this is to capture the correlations between channels and use it to suppress less useful ones while enhancing others.  \n\n## Cross-Config Attention\nAnother dimension that we can exploit attention is the batch plane (cross-configs). We designed a very simple block that allows the model to explicitly \"compare\" each config against the others throughout the network. We found this to be much better than letting the model infer for each config individually and only compare them implicitly via the loss function (`PairwiseHingeLoss`). The attention code is as follows:\n```\nclass CrossConfigAttention(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.temperature = nn.Parameter(torch.tensor(0.5))\n\n    def forward(self, x):\n        # x of shape (nb_configs, nb_nodes, nb_features)\n        scores = (x / self.temperature).softmax(dim=0)\n        x = x * scores\n        return x\n```\n\nBy applying this simple layer after the self-channel attention at every block of the network, it gave us a huge boost for default collections. For inference, we simply use a reasonably large batch size of 128. However, since the prediction depends on the batch, we can leverage it further by applying TTA to generate `N` (10) permutations of the configs and average the result after sorting it back to the original order.\n\n## Linear/Conv blocks design\nTo create our Linear/Conv blocks we followed the good practices in computer vision. We start by using `InstanceNorm` to normalise the input feature map, followed by `Linear`/`SAGEConv` layer, `SelfChannelAttetion` and  `CrossConfigAttetion` (we concat the output with its input to preserve the individuality of each sample). Then, we sum the residual connection and finish with `GELU` and dropout.\n\n# Ensembling\nOur best single model prediction scored 0.714 (0.748) on private (public) LB. However, since the number of test samples is quite low for some collections, we opted to use ensembles to improve the results and prevent shaking up. Our best result was 0.736 (0.757) LB by using the simple average of 5-10 models for each collection but we sadly didn't select this sub.\n\n# Things that didn't work\n- Train together on both random and default data\n- 2nd level stacking on top of model's embeddings/predictions (worked well locally but not that well on LB)\n- Pseudo labelling test set\n- Finetunning models for specific graphs in test set (worked well locally but not that well on LB)\n- Different ways to sample the configs, e.g., annealing the runtime spacing between configs\n- Other loss functions\n- Gradient accumulation and training with more than one graph at a time\n\n# Sources\n- Our code can be found [in this GitHub repo](https://github.com/thanhhau097/google_fast_or_slow/tree/main)\n- Our work was based/forked on @werus23 [public kernel](https://www.kaggle.com/code/werus23/tile-xla-end-to-end-train-infer) (thank you!)\n- The modified layout dataset with -1 padding can be found [here](https://www.kaggle.com/datasets/tomirol/layout-npz-padding)\n\nI made a quick diagram of our network, it's a bit crap but should give a good idea 😅\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1648129%2F7cf79a021748ba6f4b38529447419992%2Fwriteup_graph_model.png?generation=1700408602690996&alt=media)",
    "2537036": "Congrats @arc144 and team! Nice to see a brazillian winner 😉\nThanks for sharing.",
    "2531509": "Congrats on the 1st place! I am curious: what tool did you use to make this beautiful block diagram in the writeup?",
    "2530966": "Hi Eduardo,\nI'm so glad to read your write-up. I'm a beginner, so it's very hard to understand it all.\n\nCongratulations for you and your team (Thanh Hau Nguyen and Tomirol) for the amazing achievement.\nParabéns!",
    "2606250": "Congrats on winning the competition! :) Thanks for implementation details. If you could, kindly share training details, ex: GPU used, GPU training time etc.",
    "2561136": "congratulations!",
    "2534259": "Great solution. Architecture ideas are quite amazing. Learning embeddings for configs' features is what I was missing :)",
    "2533774": "Congratulations, first place! What tool did you use to create the graph in the last image?",
    "2533350": "congratulations. ",
    "2532973": "Congratulations and thanks for sharing! The solution helps me capture all the details.",
    "2532336": "Congrats for 1st place. Does an ensemble model use the same weight for each model?",
    "2532324": "Congrats brother on achieving 1st place.",
    "2532248": "Thanks for sharing, congrats for winning",
    "2531240": "Congratulations on  topping the leaderboard. Thanks for sharing the details of your solution with flowcharts. ",
    "2531191": "Congratulations for your team and learn lots from your solution!",
    "2533324": "Great competition solution! Your approach, especially using techniques like graph pruning and deduplication in data preprocessing, is quite interesting for achieving high performance. Additionally, your optimization of network architecture and attention mechanisms is very impressive. The results you obtained with ensembling are noteworthy. Congratulations! 🚀",
    "2533525": "",
    "2532621": "",
    "2532982": "Good job! Thanks for the sharing!",
    "2532390": "congrats! Thanks for sharing",
    "2531335": "Thanks for sharing！"
  }
}