{
  "id": 456093,
  "title": "5th Place Solution: GNN with Invariant Dimension Features",
  "url": "/competitions/predict-ai-model-runtime/writeups/knshnb-5th-place-solution-gnn-with-invariant-dimen",
  "author_name": "",
  "post_date": "2023-12-16T22:05:02.780Z",
  "votes": 23,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thanks for hosting the interesting competition, and congratulations to the winners!</p>\n<h2>Overview</h2>\n<p>My solution is based on an end-to-end graph neural network (GNN). I implemented a 3-layer GraphSage based on <a href=\"https://pytorch-geometric.readthedocs.io/en/latest/\" target=\"_blank\">PyG</a>. In each layer, I operate graph convolution in both directions of edges by different weights and concatenate the outputs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fa038962047b9285a904e29d787b3d1ff%2FTPUGraphs-GNN.drawio.png?generation=1702755579495636&amp;alt=media\" alt=\"\"><br>\nI trained the model to minimize pairwise hinge loss using the AdamW optimizer using a cosine annealing scheduler.<br>\nFor the loss, I used the average of a pairwise hinge loss among different configurations of the same graph and a pairwise hinge loss among all the samples in a batch (including different graphs). For this reason, I didn't use a subgraph but a whole graph as an input to GNN.</p>\n<h2>Dimension Feature Embed by Transformer</h2>\n<p>Node features include 30 features (including tile and layout configurations) for each of the 6 dimensions. A naive approach to input this to GNN is to simply flatten them (I call this naive model), but I considered the following two disadvantages.</p>\n<ul>\n<li>It drops prior information about feature correspondence across dimensions</li>\n<li>The output should be invariant to the indexing order of dimensions (I'm not sure if this is exactly correct)</li>\n</ul>\n<p>To tackle these issues, I implemented a dimension feature embedding layer using a transformer that handles each dimension as a token. In this layer, I transform (6, 30) input to (6, mid_ch) by a transformer and reduce to (mid_ch) by taking the sum in the token dimension.<br>\nSince most dimension features are exactly the same (padded ones), I could compute this efficiently by calculating embedding for only unique ones in each batch and copying them.</p>\n<h2>Tile Config Dataset</h2>\n<p>I trained the model using only the tile dataset.<br>\nUsing the transformer model, I could easily achieve 0.2 (nearly perfect) in public and private LB. The transformer model was significantly better than the naive approach on the validation Kendall tau score.</p>\n<h2>Layout Config Dataset</h2>\n<p>I trained the model using the whole layout dataset (random and default of xla and nlp). Also, including the tile dataset enhanced the performance a little.</p>\n<p>I could not outperform the naive model by the transformer model in the validation score (due to limited time), but it was comparable. My final submission was an ensemble of naive models and transformer models.</p>\n<h2>Tips</h2>\n<ul>\n<li>use the same opcode embedding for unary operations such as abs, ceil, cosine, etc.</li>\n<li>override layout_minor_to_major by layout config features for configurable nodes</li>\n<li><a href=\"https://arxiv.org/abs/1907.10903\" target=\"_blank\">DropEdge</a></li>\n<li>apply log transformation to input features</li>\n<li>oversampling</li>\n<li>load layout config data by numpy's mmap mode to save RAM</li>\n</ul>\n<h2>What Didn't Work</h2>\n<ul>\n<li>graph pooling</li>\n<li>pretrain on the tile dataset and finetune on the layout dataset</li>\n<li>graph normalization</li>\n<li>dropout node</li>\n<li>GAT, GATv2, GIN</li>\n<li>fp16</li>\n<li>pseudo label</li>\n</ul>\n<h2>Acknowledgement</h2>\n<p>I acknowledge Preferred Networks, Inc. for allowing me to use computational resources.</p>\n<p>Source Code: <a href=\"https://github.com/knshnb/kaggle-tpu-graph-5th-place\" target=\"_blank\">https://github.com/knshnb/kaggle-tpu-graph-5th-place</a></p>",
  "messages": [
    {
      "id": "2529179",
      "postDate": "11/18/2023 02:39:01",
      "content": "<p>Thanks for hosting the interesting competition, and congratulations to the winners!</p>\n<h2>Overview</h2>\n<p>My solution is based on an end-to-end graph neural network (GNN). I implemented a 3-layer GraphSage based on <a href=\"https://pytorch-geometric.readthedocs.io/en/latest/\" target=\"_blank\">PyG</a>. In each layer, I operate graph convolution in both directions of edges by different weights and concatenate the outputs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fa038962047b9285a904e29d787b3d1ff%2FTPUGraphs-GNN.drawio.png?generation=1702755579495636&amp;alt=media\" alt=\"\"><br>\nI trained the model to minimize pairwise hinge loss using the AdamW optimizer using a cosine annealing scheduler.<br>\nFor the loss, I used the average of a pairwise hinge loss among different configurations of the same graph and a pairwise hinge loss among all the samples in a batch (including different graphs). For this reason, I didn't use a subgraph but a whole graph as an input to GNN.</p>\n<h2>Dimension Feature Embed by Transformer</h2>\n<p>Node features include 30 features (including tile and layout configurations) for each of the 6 dimensions. A naive approach to input this to GNN is to simply flatten them (I call this naive model), but I considered the following two disadvantages.</p>\n<ul>\n<li>It drops prior information about feature correspondence across dimensions</li>\n<li>The output should be invariant to the indexing order of dimensions (I'm not sure if this is exactly correct)</li>\n</ul>\n<p>To tackle these issues, I implemented a dimension feature embedding layer using a transformer that handles each dimension as a token. In this layer, I transform (6, 30) input to (6, mid_ch) by a transformer and reduce to (mid_ch) by taking the sum in the token dimension.<br>\nSince most dimension features are exactly the same (padded ones), I could compute this efficiently by calculating embedding for only unique ones in each batch and copying them.</p>\n<h2>Tile Config Dataset</h2>\n<p>I trained the model using only the tile dataset.<br>\nUsing the transformer model, I could easily achieve 0.2 (nearly perfect) in public and private LB. The transformer model was significantly better than the naive approach on the validation Kendall tau score.</p>\n<h2>Layout Config Dataset</h2>\n<p>I trained the model using the whole layout dataset (random and default of xla and nlp). Also, including the tile dataset enhanced the performance a little.</p>\n<p>I could not outperform the naive model by the transformer model in the validation score (due to limited time), but it was comparable. My final submission was an ensemble of naive models and transformer models.</p>\n<h2>Tips</h2>\n<ul>\n<li>use the same opcode embedding for unary operations such as abs, ceil, cosine, etc.</li>\n<li>override layout_minor_to_major by layout config features for configurable nodes</li>\n<li><a href=\"https://arxiv.org/abs/1907.10903\" target=\"_blank\">DropEdge</a></li>\n<li>apply log transformation to input features</li>\n<li>oversampling</li>\n<li>load layout config data by numpy's mmap mode to save RAM</li>\n</ul>\n<h2>What Didn't Work</h2>\n<ul>\n<li>graph pooling</li>\n<li>pretrain on the tile dataset and finetune on the layout dataset</li>\n<li>graph normalization</li>\n<li>dropout node</li>\n<li>GAT, GATv2, GIN</li>\n<li>fp16</li>\n<li>pseudo label</li>\n</ul>\n<h2>Acknowledgement</h2>\n<p>I acknowledge Preferred Networks, Inc. for allowing me to use computational resources.</p>\n<p>Source Code: <a href=\"https://github.com/knshnb/kaggle-tpu-graph-5th-place\" target=\"_blank\">https://github.com/knshnb/kaggle-tpu-graph-5th-place</a></p>",
      "rawMarkdown": "Thanks for hosting the interesting competition, and congratulations to the winners!\n\n## Overview\nMy solution is based on an end-to-end graph neural network (GNN). I implemented a 3-layer GraphSage based on [PyG](https://pytorch-geometric.readthedocs.io/en/latest/). In each layer, I operate graph convolution in both directions of edges by different weights and concatenate the outputs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fa038962047b9285a904e29d787b3d1ff%2FTPUGraphs-GNN.drawio.png?generation=1702755579495636&alt=media)\nI trained the model to minimize pairwise hinge loss using the AdamW optimizer using a cosine annealing scheduler.\nFor the loss, I used the average of a pairwise hinge loss among different configurations of the same graph and a pairwise hinge loss among all the samples in a batch (including different graphs). For this reason, I didn't use a subgraph but a whole graph as an input to GNN.\n\n## Dimension Feature Embed by Transformer\nNode features include 30 features (including tile and layout configurations) for each of the 6 dimensions. A naive approach to input this to GNN is to simply flatten them (I call this naive model), but I considered the following two disadvantages.\n- It drops prior information about feature correspondence across dimensions\n- The output should be invariant to the indexing order of dimensions (I'm not sure if this is exactly correct)\n\nTo tackle these issues, I implemented a dimension feature embedding layer using a transformer that handles each dimension as a token. In this layer, I transform (6, 30) input to (6, mid_ch) by a transformer and reduce to (mid_ch) by taking the sum in the token dimension.\nSince most dimension features are exactly the same (padded ones), I could compute this efficiently by calculating embedding for only unique ones in each batch and copying them.\n\n## Tile Config Dataset\nI trained the model using only the tile dataset.\nUsing the transformer model, I could easily achieve 0.2 (nearly perfect) in public and private LB. The transformer model was significantly better than the naive approach on the validation Kendall tau score.\n\n## Layout Config Dataset\nI trained the model using the whole layout dataset (random and default of xla and nlp). Also, including the tile dataset enhanced the performance a little.\n\nI could not outperform the naive model by the transformer model in the validation score (due to limited time), but it was comparable. My final submission was an ensemble of naive models and transformer models.\n\n## Tips\n- use the same opcode embedding for unary operations such as abs, ceil, cosine, etc.\n- override layout_minor_to_major by layout config features for configurable nodes\n- [DropEdge](https://arxiv.org/abs/1907.10903)\n- apply log transformation to input features\n- oversampling\n- load layout config data by numpy's mmap mode to save RAM\n\n## What Didn't Work\n- graph pooling\n- pretrain on the tile dataset and finetune on the layout dataset\n- graph normalization\n- dropout node\n- GAT, GATv2, GIN\n- fp16\n- pseudo label\n\n## Acknowledgement\nI acknowledge Preferred Networks, Inc. for allowing me to use computational resources.\n\n\nSource Code: https://github.com/knshnb/kaggle-tpu-graph-5th-place",
      "votes": null
    },
    {
      "id": "2529222",
      "postDate": "11/18/2023 04:08:03",
      "content": "<p>How did you get the number 43?<br>\nBTW, congrats for winning!</p>",
      "rawMarkdown": "How did you get the number 43?\nBTW, congrats for winning!",
      "votes": null
    },
    {
      "id": "2529460",
      "postDate": "11/18/2023 09:11:51",
      "content": "<p>Sorry, 43 was a mistake and it was 30 correctly (I was tired from the hard work…). I extracted 10 that end with <code>_{i}</code> from the node features, 14 dim numbers from the node features, 3 from the tile config, and 3 from the layout config.</p>",
      "rawMarkdown": "Sorry, 43 was a mistake and it was 30 correctly (I was tired from the hard work...). I extracted 10 that end with `_{i}` from the node features, 14 dim numbers from the node features, 3 from the tile config, and 3 from the layout config.",
      "votes": null
    },
    {
      "id": "2529666",
      "postDate": "11/18/2023 13:04:44",
      "content": "<p>Thanks for the write-up and thanks for sharing. Would you share your training code at some point?</p>",
      "rawMarkdown": "Thanks for the write-up and thanks for sharing. Would you share your training code at some point?",
      "votes": null
    },
    {
      "id": "2529890",
      "postDate": "11/18/2023 16:15:57",
      "content": "<p>Congratulations on achieving 5th place. Thanks for sharing info about your solution.</p>",
      "rawMarkdown": "Congratulations on achieving 5th place. Thanks for sharing info about your solution.",
      "votes": null
    },
    {
      "id": "2530514",
      "postDate": "11/19/2023 08:04:51",
      "content": "<p>Thanks! Yeah, I'm going to publish the code later.</p>",
      "rawMarkdown": "Thanks! Yeah, I'm going to publish the code later.",
      "votes": null
    },
    {
      "id": "2530752",
      "postDate": "11/19/2023 13:50:05",
      "content": "<p>Update: Due to limited time, I could only train the final best model for 1.2 epochs out of 2 epochs. After the training of whole epochs finished, this single model achieved 1st place score…😢😢😢</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F07b81bbe731009d97fee08622f9997ba%2FF_TAVI6bcAAMHBg.jpeg?generation=1700401008267090&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Update: Due to limited time, I could only train the final best model for 1.2 epochs out of 2 epochs. After the training of whole epochs finished, this single model achieved 1st place score...😢😢😢\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F07b81bbe731009d97fee08622f9997ba%2FF_TAVI6bcAAMHBg.jpeg?generation=1700401008267090&alt=media)",
      "votes": null
    },
    {
      "id": "2530849",
      "postDate": "11/19/2023 15:24:23",
      "content": "<p>Hey man, congratulations!! Really smart how you handled the dimension features!</p>",
      "rawMarkdown": "Hey man, congratulations!! Really smart how you handled the dimension features!",
      "votes": null
    },
    {
      "id": "2531809",
      "postDate": "11/20/2023 14:54:46",
      "content": "<p>Congratulations!</p>\n<p>I have a question about this:</p>\n<blockquote>\n  <p>I operate graph convolution in both directions of edges by different weights and concatenate the outputs.</p>\n</blockquote>\n<p>What were the weights you used? What was the intuition behind this decision?</p>",
      "rawMarkdown": "Congratulations!\n\nI have a question about this:\n> I operate graph convolution in both directions of edges by different weights and concatenate the outputs.\n\nWhat were the weights you used? What was the intuition behind this decision?",
      "votes": null
    },
    {
      "id": "2531918",
      "postDate": "11/20/2023 16:30:44",
      "content": "<p>Thanks, and congratulations on your 1st place!</p>",
      "rawMarkdown": "Thanks, and congratulations on your 1st place!",
      "votes": null
    },
    {
      "id": "2532820",
      "postDate": "11/21/2023 11:08:15",
      "content": "<p>Thanks!<br>\nI mean I concatenated the outputs of two independent graph convolutions. The pseudocode is like this:</p>\n<pre><code>self.gconv = SAGEConv(mid_ch, mid_ch // , ...)\nself.rev_gconv =SAGEConv(mid_ch, mid_ch // , ...)\n\nrev_edge_index = torch.flip(edge_index, (,))\nx = torch.cat([self.gconv(x, edge_index), self.rev_gconv(x, rev_edge_index)], dim=-))\n</code></pre>\n<p>Intuition: If we have an edge (u, v), we want to propagate information not only from u to v but also from v to u. However, the weight should not be equal because it's a directed edge.</p>",
      "rawMarkdown": "Thanks!\nI mean I concatenated the outputs of two independent graph convolutions. The pseudocode is like this:\n```python\nself.gconv = SAGEConv(mid_ch, mid_ch // 2, ...)\nself.rev_gconv =SAGEConv(mid_ch, mid_ch // 2, ...)\n\nrev_edge_index = torch.flip(edge_index, (0,))\nx = torch.cat([self.gconv(x, edge_index), self.rev_gconv(x, rev_edge_index)], dim=-1))\n```\n\nIntuition: If we have an edge (u, v), we want to propagate information not only from u to v but also from v to u. However, the weight should not be equal because it's a directed edge.",
      "votes": null
    },
    {
      "id": "2532933",
      "postDate": "11/21/2023 13:00:13",
      "content": "<p>Awesome! Thanks again. 🥳</p>",
      "rawMarkdown": "Awesome! Thanks again. 🥳",
      "votes": null
    },
    {
      "id": "2551117",
      "postDate": "12/06/2023 14:08:58",
      "content": "<p>I published a source code!<br>\n<a href=\"https://github.com/knshnb/kaggle-tpu-graph-5th-place\" target=\"_blank\">https://github.com/knshnb/kaggle-tpu-graph-5th-place</a></p>",
      "rawMarkdown": "I published a source code!\nhttps://github.com/knshnb/kaggle-tpu-graph-5th-place",
      "votes": null
    },
    {
      "id": "2553751",
      "postDate": "12/08/2023 13:44:22",
      "content": "<p>thanks for the writeup. i was at the winner call video conference with you today. i am impressed with your work!.</p>\n<p>I have a question: any idea why the public score of the post submission (2 epoch trained) is much lower than the private one?<br>\nhow is your validation score? 0.739 and 0.707 is a huge difference</p>",
      "rawMarkdown": "thanks for the writeup. i was at the winner call video conference with you today. i am impressed with your work!.\n\nI have a question: any idea why the public score of the post submission (2 epoch trained) is much lower than the private one?\nhow is your validation score? 0.739 and 0.707 is a huge difference",
      "votes": null
    },
    {
      "id": "2556281",
      "postDate": "12/10/2023 15:18:55",
      "content": "<p>Thanks for the comment. I learned a lot from your presentation!</p>\n<p>I just focused on improving the validation score and did not care about public and private during the competition. According to the LB probing analysis by <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>, it seems public data has some graphs that are far from training graphs, which caused high randomness.<br>\n<a href=\"https://twitter.com/charmq00/status/1725675672222503043\" target=\"_blank\">https://twitter.com/charmq00/status/1725675672222503043</a><br>\n(Sorry, the above tweet is in Japanese…)</p>",
      "rawMarkdown": "Thanks for the comment. I learned a lot from your presentation!\n\nI just focused on improving the validation score and did not care about public and private during the competition. According to the LB probing analysis by @charmq, it seems public data has some graphs that are far from training graphs, which caused high randomness.\nhttps://twitter.com/charmq00/status/1725675672222503043\n(Sorry, the above tweet is in Japanese...)",
      "votes": null
    },
    {
      "id": "2556305",
      "postDate": "12/10/2023 15:31:19",
      "content": "<p>For others, this is the post submission results I mentioned in the winner call.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F7da5830169c4933ab603971a32033ca4%2FScreen%20Shot%202023-12-11%20at%200.22.31.png?generation=1702222223063708&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "For others, this is the post submission results I mentioned in the winner call.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F7da5830169c4933ab603971a32033ca4%2FScreen%20Shot%202023-12-11%20at%200.22.31.png?generation=1702222223063708&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2529222,
      "author_name": "haohantsao",
      "author_url": "",
      "post_date": "11/18/2023 04:08:03",
      "content": "<p>How did you get the number 43?<br>\nBTW, congrats for winning!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2529460,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "11/18/2023 09:11:51",
          "content": "<p>Sorry, 43 was a mistake and it was 30 correctly (I was tired from the hard work…). I extracted 10 that end with <code>_{i}</code> from the node features, 14 dim numbers from the node features, 3 from the tile config, and 3 from the layout config.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2529666,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "11/18/2023 13:04:44",
      "content": "<p>Thanks for the write-up and thanks for sharing. Would you share your training code at some point?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2530514,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "11/19/2023 08:04:51",
          "content": "<p>Thanks! Yeah, I'm going to publish the code later.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2532933,
              "author_name": "yassinealouini",
              "author_url": "",
              "post_date": "11/21/2023 13:00:13",
              "content": "<p>Awesome! Thanks again. 🥳</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2529890,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "11/18/2023 16:15:57",
      "content": "<p>Congratulations on achieving 5th place. Thanks for sharing info about your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2530752,
      "author_name": "knshnb",
      "author_url": "",
      "post_date": "11/19/2023 13:50:05",
      "content": "<p>Update: Due to limited time, I could only train the final best model for 1.2 epochs out of 2 epochs. After the training of whole epochs finished, this single model achieved 1st place score…😢😢😢</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F07b81bbe731009d97fee08622f9997ba%2FF_TAVI6bcAAMHBg.jpeg?generation=1700401008267090&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2553751,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/08/2023 13:44:22",
          "content": "<p>thanks for the writeup. i was at the winner call video conference with you today. i am impressed with your work!.</p>\n<p>I have a question: any idea why the public score of the post submission (2 epoch trained) is much lower than the private one?<br>\nhow is your validation score? 0.739 and 0.707 is a huge difference</p>",
          "votes": null,
          "replies": [
            {
              "id": 2556281,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "12/10/2023 15:18:55",
              "content": "<p>Thanks for the comment. I learned a lot from your presentation!</p>\n<p>I just focused on improving the validation score and did not care about public and private during the competition. According to the LB probing analysis by <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>, it seems public data has some graphs that are far from training graphs, which caused high randomness.<br>\n<a href=\"https://twitter.com/charmq00/status/1725675672222503043\" target=\"_blank\">https://twitter.com/charmq00/status/1725675672222503043</a><br>\n(Sorry, the above tweet is in Japanese…)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2556305,
                  "author_name": "knshnb",
                  "author_url": "",
                  "post_date": "12/10/2023 15:31:19",
                  "content": "<p>For others, this is the post submission results I mentioned in the winner call.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F7da5830169c4933ab603971a32033ca4%2FScreen%20Shot%202023-12-11%20at%200.22.31.png?generation=1702222223063708&amp;alt=media\" alt=\"\"></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2530849,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "11/19/2023 15:24:23",
      "content": "<p>Hey man, congratulations!! Really smart how you handled the dimension features!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2531918,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "11/20/2023 16:30:44",
          "content": "<p>Thanks, and congratulations on your 1st place!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2531809,
      "author_name": "mattdeakos",
      "author_url": "",
      "post_date": "11/20/2023 14:54:46",
      "content": "<p>Congratulations!</p>\n<p>I have a question about this:</p>\n<blockquote>\n  <p>I operate graph convolution in both directions of edges by different weights and concatenate the outputs.</p>\n</blockquote>\n<p>What were the weights you used? What was the intuition behind this decision?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2532820,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "11/21/2023 11:08:15",
          "content": "<p>Thanks!<br>\nI mean I concatenated the outputs of two independent graph convolutions. The pseudocode is like this:</p>\n<pre><code>self.gconv = SAGEConv(mid_ch, mid_ch // , ...)\nself.rev_gconv =SAGEConv(mid_ch, mid_ch // , ...)\n\nrev_edge_index = torch.flip(edge_index, (,))\nx = torch.cat([self.gconv(x, edge_index), self.rev_gconv(x, rev_edge_index)], dim=-))\n</code></pre>\n<p>Intuition: If we have an edge (u, v), we want to propagate information not only from u to v but also from v to u. However, the weight should not be equal because it's a directed edge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2551117,
      "author_name": "knshnb",
      "author_url": "",
      "post_date": "12/06/2023 14:08:58",
      "content": "<p>I published a source code!<br>\n<a href=\"https://github.com/knshnb/kaggle-tpu-graph-5th-place\" target=\"_blank\">https://github.com/knshnb/kaggle-tpu-graph-5th-place</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2529179": "Thanks for hosting the interesting competition, and congratulations to the winners!\n\n## Overview\nMy solution is based on an end-to-end graph neural network (GNN). I implemented a 3-layer GraphSage based on [PyG](https://pytorch-geometric.readthedocs.io/en/latest/). In each layer, I operate graph convolution in both directions of edges by different weights and concatenate the outputs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fa038962047b9285a904e29d787b3d1ff%2FTPUGraphs-GNN.drawio.png?generation=1702755579495636&alt=media)\nI trained the model to minimize pairwise hinge loss using the AdamW optimizer using a cosine annealing scheduler.\nFor the loss, I used the average of a pairwise hinge loss among different configurations of the same graph and a pairwise hinge loss among all the samples in a batch (including different graphs). For this reason, I didn't use a subgraph but a whole graph as an input to GNN.\n\n## Dimension Feature Embed by Transformer\nNode features include 30 features (including tile and layout configurations) for each of the 6 dimensions. A naive approach to input this to GNN is to simply flatten them (I call this naive model), but I considered the following two disadvantages.\n- It drops prior information about feature correspondence across dimensions\n- The output should be invariant to the indexing order of dimensions (I'm not sure if this is exactly correct)\n\nTo tackle these issues, I implemented a dimension feature embedding layer using a transformer that handles each dimension as a token. In this layer, I transform (6, 30) input to (6, mid_ch) by a transformer and reduce to (mid_ch) by taking the sum in the token dimension.\nSince most dimension features are exactly the same (padded ones), I could compute this efficiently by calculating embedding for only unique ones in each batch and copying them.\n\n## Tile Config Dataset\nI trained the model using only the tile dataset.\nUsing the transformer model, I could easily achieve 0.2 (nearly perfect) in public and private LB. The transformer model was significantly better than the naive approach on the validation Kendall tau score.\n\n## Layout Config Dataset\nI trained the model using the whole layout dataset (random and default of xla and nlp). Also, including the tile dataset enhanced the performance a little.\n\nI could not outperform the naive model by the transformer model in the validation score (due to limited time), but it was comparable. My final submission was an ensemble of naive models and transformer models.\n\n## Tips\n- use the same opcode embedding for unary operations such as abs, ceil, cosine, etc.\n- override layout_minor_to_major by layout config features for configurable nodes\n- [DropEdge](https://arxiv.org/abs/1907.10903)\n- apply log transformation to input features\n- oversampling\n- load layout config data by numpy's mmap mode to save RAM\n\n## What Didn't Work\n- graph pooling\n- pretrain on the tile dataset and finetune on the layout dataset\n- graph normalization\n- dropout node\n- GAT, GATv2, GIN\n- fp16\n- pseudo label\n\n## Acknowledgement\nI acknowledge Preferred Networks, Inc. for allowing me to use computational resources.\n\n\nSource Code: https://github.com/knshnb/kaggle-tpu-graph-5th-place",
    "2529222": "How did you get the number 43?\nBTW, congrats for winning!",
    "2529460": "Sorry, 43 was a mistake and it was 30 correctly (I was tired from the hard work...). I extracted 10 that end with `_{i}` from the node features, 14 dim numbers from the node features, 3 from the tile config, and 3 from the layout config.",
    "2529666": "Thanks for the write-up and thanks for sharing. Would you share your training code at some point?",
    "2529890": "Congratulations on achieving 5th place. Thanks for sharing info about your solution.",
    "2530514": "Thanks! Yeah, I'm going to publish the code later.",
    "2530752": "Update: Due to limited time, I could only train the final best model for 1.2 epochs out of 2 epochs. After the training of whole epochs finished, this single model achieved 1st place score...😢😢😢\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F07b81bbe731009d97fee08622f9997ba%2FF_TAVI6bcAAMHBg.jpeg?generation=1700401008267090&alt=media)",
    "2530849": "Hey man, congratulations!! Really smart how you handled the dimension features!",
    "2531809": "Congratulations!\n\nI have a question about this:\n> I operate graph convolution in both directions of edges by different weights and concatenate the outputs.\n\nWhat were the weights you used? What was the intuition behind this decision?",
    "2531918": "Thanks, and congratulations on your 1st place!",
    "2532820": "Thanks!\nI mean I concatenated the outputs of two independent graph convolutions. The pseudocode is like this:\n```python\nself.gconv = SAGEConv(mid_ch, mid_ch // 2, ...)\nself.rev_gconv =SAGEConv(mid_ch, mid_ch // 2, ...)\n\nrev_edge_index = torch.flip(edge_index, (0,))\nx = torch.cat([self.gconv(x, edge_index), self.rev_gconv(x, rev_edge_index)], dim=-1))\n```\n\nIntuition: If we have an edge (u, v), we want to propagate information not only from u to v but also from v to u. However, the weight should not be equal because it's a directed edge.",
    "2532933": "Awesome! Thanks again. 🥳",
    "2551117": "I published a source code!\nhttps://github.com/knshnb/kaggle-tpu-graph-5th-place",
    "2553751": "thanks for the writeup. i was at the winner call video conference with you today. i am impressed with your work!.\n\nI have a question: any idea why the public score of the post submission (2 epoch trained) is much lower than the private one?\nhow is your validation score? 0.739 and 0.707 is a huge difference",
    "2556281": "Thanks for the comment. I learned a lot from your presentation!\n\nI just focused on improving the validation score and did not care about public and private during the competition. According to the LB probing analysis by @charmq, it seems public data has some graphs that are far from training graphs, which caused high randomness.\nhttps://twitter.com/charmq00/status/1725675672222503043\n(Sorry, the above tweet is in Japanese...)",
    "2556305": "For others, this is the post submission results I mentioned in the winner call.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F7da5830169c4933ab603971a32033ca4%2FScreen%20Shot%202023-12-11%20at%200.22.31.png?generation=1702222223063708&alt=media)"
  },
  "source": "meta"
}