{
  "id": 456206,
  "title": "9th Place Solution: GNN with Compressed Graphs Using Dijkstra’s Algorithm",
  "url": "/competitions/predict-ai-model-runtime/writeups/preferred-runtimer-9th-place-solution-gnn-with-com",
  "author_name": "",
  "post_date": "2023-11-18T18:35:04.257Z",
  "votes": 21,
  "comment_count": 2,
  "views": 0,
  "content": "<p>We would like to express our sincere gratitude to our Kaggle teammates and to our hosts for providing the opportunity to participate in such an engaging competition. And thanks to my excellent teammate <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a>. </p>\n<p>This was my 7th (and 4th with <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a>) gold medal this year. Wow!</p>\n<p>The problem posed was of a type that we had not tackled much before, which made it extremely fascinating to delve into. Here below we outline our solution.</p>\n<h1>On Tiles</h1>\n<p>We referenced <a href=\"https://github.com/google-research-datasets/tpu_graphs\" target=\"_blank\">GitHub - google-research-datasets/tpu_graphs</a> and public notebooks. <br>\nSince the score from the public notebook was already satisfactory, the successful learning of the layout part was the key to this competition.</p>\n<h1>On Layout</h1>\n<p>We mainly used the implementation of <a href=\"https://github.com/kaidic/GST/tree/main\" target=\"_blank\">GutHub - kaidic/GST</a> as a reference and made modifications that we felt were necessary to improve the score.</p>\n<h2>Graph Compression</h2>\n<p>During the learning of the layout, the configurable data was limited. Therefore, our implementation only extracted nodes influenced by this and their connected components. <br>\nBy applying this, our learning efficiency dramatically improved, significantly contributing to the improvement of the score.</p>\n<p>Specifically, the following procedures were performed for each data set</p>\n<ol>\n<li>An undirected graph was constructed using edge_index, and Dijkstra’s algorithm was applied starting from the node with the largest index <code>s</code> (i.e., <code>s=data[\"node_feat\"].shape[0]-1</code>). We chose <code>s</code> as the starting point because many graphs were trees with <code>s</code> as the parent.</li>\n<li>The shortest path from <code>s</code> to all the nodes in <code>node_config_ids</code> was calculated, and the union set of nodes and edges in the path was considered as a compressed graph.</li>\n</ol>\n<p>This compression method reduced the average number of nodes for each layout data from 13894 to 1736 for xla and from 5711 to 570 for nlp. For features of nodes not included in the compressed graph, we simply ignored them completely.</p>\n<h2>Preprocessing (log-transform)</h2>\n<p>There were times when the learning could not proceed because orders differed depending upon the dimensions of input features. <br>\nTo cope with this, we underwent log transformation for node features, which consequently made learning proceed smoothly.</p>\n<h2>Training Strategy</h2>\n<ul>\n<li>512 configs were sampled for each data per iteration. The bach_size were 2 or 4.</li>\n<li>Trained separate models for the four data types xla-random, xla-default, nlp-random, and nlp-default for 1000 epochs.</li>\n<li>The pairwise hinge loss was used (same as original implementation).</li>\n</ul>\n<h3>CV score</h3>\n<ul>\n<li>xla random: 0.7</li>\n<li>xla default: 0.33</li>\n<li>nlp random: 0.96</li>\n<li>nlp default: 0.51</li>\n</ul>\n<h2>Data Specialized for Specific Model Types</h2>\n<p>Upon examining the data, we found that architectures (such as BERT, U-Net, ResNet…) could be inferred from IDs or node numbers (or op_codes for test data).<br>\n Leveraging this information, we performed learning using data related only to each architecture, which substantially contributed to improving the private score. In addition to the models trained on the entire dataset, we trained models focused on resnet, efficientnet, or bert. Proper EDA is indeed crucial.</p>\n<h2>Multiple GNN Architectures</h2>\n<p>The TransformerConv and GATConv as Graph conv layers had minimal impact but were useful for ensemble purposes.</p>\n<h3>Ensemble</h3>\n<p>Although we didn't have sufficient time for meticulous weight tuning, it certainly contributed to steady score improvement (+0.01 - 0.02).</p>\n<h3>What Didn't Work</h3>\n<ul>\n<li>MSELoss as aux-loss</li>\n<li>Learning using all nodes</li>\n</ul>",
  "messages": [
    {
      "id": "2530018",
      "postDate": "11/18/2023 18:24:13",
      "content": "<p>We would like to express our sincere gratitude to our Kaggle teammates and to our hosts for providing the opportunity to participate in such an engaging competition. And thanks to my excellent teammate <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a>. </p>\n<p>This was my 7th (and 4th with <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a>) gold medal this year. Wow!</p>\n<p>The problem posed was of a type that we had not tackled much before, which made it extremely fascinating to delve into. Here below we outline our solution.</p>\n<h1>On Tiles</h1>\n<p>We referenced <a href=\"https://github.com/google-research-datasets/tpu_graphs\" target=\"_blank\">GitHub - google-research-datasets/tpu_graphs</a> and public notebooks. <br>\nSince the score from the public notebook was already satisfactory, the successful learning of the layout part was the key to this competition.</p>\n<h1>On Layout</h1>\n<p>We mainly used the implementation of <a href=\"https://github.com/kaidic/GST/tree/main\" target=\"_blank\">GutHub - kaidic/GST</a> as a reference and made modifications that we felt were necessary to improve the score.</p>\n<h2>Graph Compression</h2>\n<p>During the learning of the layout, the configurable data was limited. Therefore, our implementation only extracted nodes influenced by this and their connected components. <br>\nBy applying this, our learning efficiency dramatically improved, significantly contributing to the improvement of the score.</p>\n<p>Specifically, the following procedures were performed for each data set</p>\n<ol>\n<li>An undirected graph was constructed using edge_index, and Dijkstra’s algorithm was applied starting from the node with the largest index <code>s</code> (i.e., <code>s=data[\"node_feat\"].shape[0]-1</code>). We chose <code>s</code> as the starting point because many graphs were trees with <code>s</code> as the parent.</li>\n<li>The shortest path from <code>s</code> to all the nodes in <code>node_config_ids</code> was calculated, and the union set of nodes and edges in the path was considered as a compressed graph.</li>\n</ol>\n<p>This compression method reduced the average number of nodes for each layout data from 13894 to 1736 for xla and from 5711 to 570 for nlp. For features of nodes not included in the compressed graph, we simply ignored them completely.</p>\n<h2>Preprocessing (log-transform)</h2>\n<p>There were times when the learning could not proceed because orders differed depending upon the dimensions of input features. <br>\nTo cope with this, we underwent log transformation for node features, which consequently made learning proceed smoothly.</p>\n<h2>Training Strategy</h2>\n<ul>\n<li>512 configs were sampled for each data per iteration. The bach_size were 2 or 4.</li>\n<li>Trained separate models for the four data types xla-random, xla-default, nlp-random, and nlp-default for 1000 epochs.</li>\n<li>The pairwise hinge loss was used (same as original implementation).</li>\n</ul>\n<h3>CV score</h3>\n<ul>\n<li>xla random: 0.7</li>\n<li>xla default: 0.33</li>\n<li>nlp random: 0.96</li>\n<li>nlp default: 0.51</li>\n</ul>\n<h2>Data Specialized for Specific Model Types</h2>\n<p>Upon examining the data, we found that architectures (such as BERT, U-Net, ResNet…) could be inferred from IDs or node numbers (or op_codes for test data).<br>\n Leveraging this information, we performed learning using data related only to each architecture, which substantially contributed to improving the private score. In addition to the models trained on the entire dataset, we trained models focused on resnet, efficientnet, or bert. Proper EDA is indeed crucial.</p>\n<h2>Multiple GNN Architectures</h2>\n<p>The TransformerConv and GATConv as Graph conv layers had minimal impact but were useful for ensemble purposes.</p>\n<h3>Ensemble</h3>\n<p>Although we didn't have sufficient time for meticulous weight tuning, it certainly contributed to steady score improvement (+0.01 - 0.02).</p>\n<h3>What Didn't Work</h3>\n<ul>\n<li>MSELoss as aux-loss</li>\n<li>Learning using all nodes</li>\n</ul>",
      "rawMarkdown": "We would like to express our sincere gratitude to our Kaggle teammates and to our hosts for providing the opportunity to participate in such an engaging competition. And thanks to my excellent teammate @yoichi7yamakawa. \n\nThis was my 7th (and 4th with @yoichi7yamakawa) gold medal this year. Wow!\n\nThe problem posed was of a type that we had not tackled much before, which made it extremely fascinating to delve into. Here below we outline our solution.\n\n# On Tiles\nWe referenced [GitHub - google-research-datasets/tpu_graphs](https://github.com/google-research-datasets/tpu_graphs) and public notebooks. \nSince the score from the public notebook was already satisfactory, the successful learning of the layout part was the key to this competition.\n\n# On Layout\nWe mainly used the implementation of [GutHub - kaidic/GST](https://github.com/kaidic/GST/tree/main) as a reference and made modifications that we felt were necessary to improve the score.\n\n## Graph Compression\nDuring the learning of the layout, the configurable data was limited. Therefore, our implementation only extracted nodes influenced by this and their connected components. \nBy applying this, our learning efficiency dramatically improved, significantly contributing to the improvement of the score.\n\nSpecifically, the following procedures were performed for each data set\n1. An undirected graph was constructed using edge_index, and Dijkstra’s algorithm was applied starting from the node with the largest index `s` (i.e., `s=data[\"node_feat\"].shape[0]-1`). We chose `s` as the starting point because many graphs were trees with `s` as the parent.\n2. The shortest path from `s` to all the nodes in `node_config_ids` was calculated, and the union set of nodes and edges in the path was considered as a compressed graph.\n\nThis compression method reduced the average number of nodes for each layout data from 13894 to 1736 for xla and from 5711 to 570 for nlp. For features of nodes not included in the compressed graph, we simply ignored them completely.\n\n## Preprocessing (log-transform)\nThere were times when the learning could not proceed because orders differed depending upon the dimensions of input features. \nTo cope with this, we underwent log transformation for node features, which consequently made learning proceed smoothly.\n\n## Training Strategy\n- 512 configs were sampled for each data per iteration. The bach_size were 2 or 4.\n- Trained separate models for the four data types xla-random, xla-default, nlp-random, and nlp-default for 1000 epochs.\n- The pairwise hinge loss was used (same as original implementation).\n\n### CV score\n- xla random: 0.7\n- xla default: 0.33\n- nlp random: 0.96\n- nlp default: 0.51\n\n## Data Specialized for Specific Model Types\nUpon examining the data, we found that architectures (such as BERT, U-Net, ResNet...) could be inferred from IDs or node numbers (or op_codes for test data).\n Leveraging this information, we performed learning using data related only to each architecture, which substantially contributed to improving the private score. In addition to the models trained on the entire dataset, we trained models focused on resnet, efficientnet, or bert. Proper EDA is indeed crucial.\n\n## Multiple GNN Architectures\nThe TransformerConv and GATConv as Graph conv layers had minimal impact but were useful for ensemble purposes.\n\n### Ensemble\nAlthough we didn't have sufficient time for meticulous weight tuning, it certainly contributed to steady score improvement (+0.01 - 0.02).\n\n### What Didn't Work\n- MSELoss as aux-loss\n- Learning using all nodes",
      "votes": null
    },
    {
      "id": "2530911",
      "postDate": "11/19/2023 16:27:46",
      "content": "<p>Congratulations!! Really clean solution! </p>\n<p>We also used a similar graph compression algorithm. Did you use the GST on top of the compressed graph?</p>",
      "rawMarkdown": "Congratulations!! Really clean solution! \n\nWe also used a similar graph compression algorithm. Did you use the GST on top of the compressed graph?",
      "votes": null
    },
    {
      "id": "2531536",
      "postDate": "11/20/2023 09:49:17",
      "content": "<p>Thank you!<br>\nYes, we used GST, but the graph is already small enough that it may not be very useful. </p>",
      "rawMarkdown": "Thank you!\nYes, we used GST, but the graph is already small enough that it may not be very useful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2530911,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "11/19/2023 16:27:46",
      "content": "<p>Congratulations!! Really clean solution! </p>\n<p>We also used a similar graph compression algorithm. Did you use the GST on top of the compressed graph?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2531536,
          "author_name": "charmq",
          "author_url": "",
          "post_date": "11/20/2023 09:49:17",
          "content": "<p>Thank you!<br>\nYes, we used GST, but the graph is already small enough that it may not be very useful. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2530018": "We would like to express our sincere gratitude to our Kaggle teammates and to our hosts for providing the opportunity to participate in such an engaging competition. And thanks to my excellent teammate @yoichi7yamakawa. \n\nThis was my 7th (and 4th with @yoichi7yamakawa) gold medal this year. Wow!\n\nThe problem posed was of a type that we had not tackled much before, which made it extremely fascinating to delve into. Here below we outline our solution.\n\n# On Tiles\nWe referenced [GitHub - google-research-datasets/tpu_graphs](https://github.com/google-research-datasets/tpu_graphs) and public notebooks. \nSince the score from the public notebook was already satisfactory, the successful learning of the layout part was the key to this competition.\n\n# On Layout\nWe mainly used the implementation of [GutHub - kaidic/GST](https://github.com/kaidic/GST/tree/main) as a reference and made modifications that we felt were necessary to improve the score.\n\n## Graph Compression\nDuring the learning of the layout, the configurable data was limited. Therefore, our implementation only extracted nodes influenced by this and their connected components. \nBy applying this, our learning efficiency dramatically improved, significantly contributing to the improvement of the score.\n\nSpecifically, the following procedures were performed for each data set\n1. An undirected graph was constructed using edge_index, and Dijkstra’s algorithm was applied starting from the node with the largest index `s` (i.e., `s=data[\"node_feat\"].shape[0]-1`). We chose `s` as the starting point because many graphs were trees with `s` as the parent.\n2. The shortest path from `s` to all the nodes in `node_config_ids` was calculated, and the union set of nodes and edges in the path was considered as a compressed graph.\n\nThis compression method reduced the average number of nodes for each layout data from 13894 to 1736 for xla and from 5711 to 570 for nlp. For features of nodes not included in the compressed graph, we simply ignored them completely.\n\n## Preprocessing (log-transform)\nThere were times when the learning could not proceed because orders differed depending upon the dimensions of input features. \nTo cope with this, we underwent log transformation for node features, which consequently made learning proceed smoothly.\n\n## Training Strategy\n- 512 configs were sampled for each data per iteration. The bach_size were 2 or 4.\n- Trained separate models for the four data types xla-random, xla-default, nlp-random, and nlp-default for 1000 epochs.\n- The pairwise hinge loss was used (same as original implementation).\n\n### CV score\n- xla random: 0.7\n- xla default: 0.33\n- nlp random: 0.96\n- nlp default: 0.51\n\n## Data Specialized for Specific Model Types\nUpon examining the data, we found that architectures (such as BERT, U-Net, ResNet...) could be inferred from IDs or node numbers (or op_codes for test data).\n Leveraging this information, we performed learning using data related only to each architecture, which substantially contributed to improving the private score. In addition to the models trained on the entire dataset, we trained models focused on resnet, efficientnet, or bert. Proper EDA is indeed crucial.\n\n## Multiple GNN Architectures\nThe TransformerConv and GATConv as Graph conv layers had minimal impact but were useful for ensemble purposes.\n\n### Ensemble\nAlthough we didn't have sufficient time for meticulous weight tuning, it certainly contributed to steady score improvement (+0.01 - 0.02).\n\n### What Didn't Work\n- MSELoss as aux-loss\n- Learning using all nodes",
    "2530911": "Congratulations!! Really clean solution! \n\nWe also used a similar graph compression algorithm. Did you use the GST on top of the compressed graph?",
    "2531536": "Thank you!\nYes, we used GST, but the graph is already small enough that it may not be very useful."
  },
  "source": "meta"
}