{
  "id": 456084,
  "title": "6th solution: Node-level Instance Norm + Residual  SageConv on 5-hop-neighbour Subgraph ",
  "url": "/competitions/predict-ai-model-runtime/writeups/none-6th-solution-node-level-instance-norm-residua",
  "author_name": "",
  "post_date": "2023-12-02T04:00:27.990Z",
  "votes": 28,
  "comment_count": 5,
  "views": 0,
  "content": "<h1><strong>Training and inference code can be downloaded from:</strong></h1>\n<p><a href=\"https://github.com/hengck23/solution-predict-ai-model-runtime/\" target=\"_blank\">https://github.com/hengck23/solution-predict-ai-model-runtime/</a>  <br>\n&nbsp;</p>\n<h2>1. Layout runtime prediction</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb5c1c8af2f166b73061177ae313a2c6b%2FSelection_999(4192).png?generation=1701482437470415&amp;alt=media\" alt=\"\"></p>\n<p>main problem:</p>\n<ul>\n<li>we have very large graph as input. how to design learning model and algorithm that can fit into gpu memory? </li>\n</ul>\n<p>summary of approach :</p>\n<ul>\n<li>instead of using the whole graph, we can reduce it by considering only the 5-hop neighbours from node marked as \"config id\". We call this 5-hop-neighbour subgraph.  We think this is reasonable becuase since we are comparing relative ranking of 2 graphs,  and we just need to input the \"difference nodes\"  (instead of the whole graph) to the neural net.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5e4c66d69dcf5ea590ad619c37dc2179%2FSelection_999(4193).png?generation=1701483096930612&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>using the reduced subgraph, i can sample 32 to 100 configurations using all full subgraphs with a single 48-GB  GPU card at training.</li>\n<li>batch size is not an issue here, bcuase i am using gradient accumulation. We accumuluate over one subgraph at a time when training a batch.</li>\n</ul>\n<pre><code>optimizer.zero_grad()\n\nfor in range(        r = \n        loss = net(r)  \n        \n\n\n</code></pre>\n<ul>\n<li><p>normalisation is important. We use \"graph instance norm\" (over node), see paper[1], which works well with gradient accumulation </p></li>\n<li><p>we use pairwise ranking loss in training loss.</p></li>\n<li><p>We try 2 GNN: </p>\n<ul>\n<li>4-layer SAGE-conv[2] with residual shortcut</li>\n<li>4-layer GIN-conv[3] </li></ul></li>\n</ul>\n<p>SAGE-conv is better than GIN-conv.</p>\n<h2>2. Tile runtime prediction</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc13d2c462b0958161cd4ce09e4249bc7%2FSelection_999(4191).png?generation=1701482450756660&amp;alt=media\" alt=\"\"></p>\n<p>main problem:</p>\n<ul>\n<li>There is no issues here as the graph in the kaggle training data are much smaller. These are actually subgraphs of the much larger original computation graph.</li>\n</ul>\n<p>summary of approach</p>\n<ul>\n<li>We still use \"graph instance norm\" (over node) [1], and smae gradient accumulation apporach, with batch size =64.</li>\n<li>We try both SAGE-conv[2] and GAT-conv[4].  GAT-conv gives better results.</li>\n<li>Since we are interested in top 5 ranks, we find listMLE is a better loss.</li>\n</ul>\n<hr>\n<h2>[Reference]</h2>\n<p>[1] GraphNorm: A Principled Approach to Accelerating Graph Neural Network Training<br>\n<a href=\"https://arxiv.org/abs/2009.03294\" target=\"_blank\">https://arxiv.org/abs/2009.03294</a>  <br>\n[2] Inductive Representation Learning on Large Graphs<br>\n<a href=\"https://arxiv.org/abs/1706.02216\" target=\"_blank\">https://arxiv.org/abs/1706.02216</a>  <br>\n[3] How Powerful are Graph Neural Networks?<br>\n<a href=\"https://arxiv.org/pdf/1810.00826.pdf\" target=\"_blank\">https://arxiv.org/pdf/1810.00826.pdf</a><br>\n[4] Graph Attention Networks <br>\n<a href=\"https://arxiv.org/abs/1710.10903\" target=\"_blank\">https://arxiv.org/abs/1710.10903</a></p>\n<hr>\n<h2>local validation and public/private score</h2>\n<p>The metric are: slowndown (top-5) for tile and kendall tau for layout.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0caab07b1cf25d81b358b1234d459b8c%2FSelection_999(4194).png?generation=1701488206301120&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h2>Acknowledgment</h2>\n<h2><em>\"I would like to express my sincere gratitude to HP for the generous provision of the  Z8-G4 Data Science Workstation that was instrumental in the successful completion of kaggle competition. The two 48GB Nvidia Quadro RTX 8000 GPU cards give me a distinct advantage to easily bulid models with the largest public graph dataset TPUGraphs, with 100 millions graphs of 10 thousands nodes.\"</em></h2>",
  "messages": [
    {
      "id": "2529128",
      "postDate": "11/18/2023 01:02:47",
      "content": "<h1><strong>Training and inference code can be downloaded from:</strong></h1>\n<p><a href=\"https://github.com/hengck23/solution-predict-ai-model-runtime/\" target=\"_blank\">https://github.com/hengck23/solution-predict-ai-model-runtime/</a>  <br>\n&nbsp;</p>\n<h2>1. Layout runtime prediction</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb5c1c8af2f166b73061177ae313a2c6b%2FSelection_999(4192).png?generation=1701482437470415&amp;alt=media\" alt=\"\"></p>\n<p>main problem:</p>\n<ul>\n<li>we have very large graph as input. how to design learning model and algorithm that can fit into gpu memory? </li>\n</ul>\n<p>summary of approach :</p>\n<ul>\n<li>instead of using the whole graph, we can reduce it by considering only the 5-hop neighbours from node marked as \"config id\". We call this 5-hop-neighbour subgraph.  We think this is reasonable becuase since we are comparing relative ranking of 2 graphs,  and we just need to input the \"difference nodes\"  (instead of the whole graph) to the neural net.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5e4c66d69dcf5ea590ad619c37dc2179%2FSelection_999(4193).png?generation=1701483096930612&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>using the reduced subgraph, i can sample 32 to 100 configurations using all full subgraphs with a single 48-GB  GPU card at training.</li>\n<li>batch size is not an issue here, bcuase i am using gradient accumulation. We accumuluate over one subgraph at a time when training a batch.</li>\n</ul>\n<pre><code>optimizer.zero_grad()\n\nfor in range(        r = \n        loss = net(r)  \n        \n\n\n</code></pre>\n<ul>\n<li><p>normalisation is important. We use \"graph instance norm\" (over node), see paper[1], which works well with gradient accumulation </p></li>\n<li><p>we use pairwise ranking loss in training loss.</p></li>\n<li><p>We try 2 GNN: </p>\n<ul>\n<li>4-layer SAGE-conv[2] with residual shortcut</li>\n<li>4-layer GIN-conv[3] </li></ul></li>\n</ul>\n<p>SAGE-conv is better than GIN-conv.</p>\n<h2>2. Tile runtime prediction</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc13d2c462b0958161cd4ce09e4249bc7%2FSelection_999(4191).png?generation=1701482450756660&amp;alt=media\" alt=\"\"></p>\n<p>main problem:</p>\n<ul>\n<li>There is no issues here as the graph in the kaggle training data are much smaller. These are actually subgraphs of the much larger original computation graph.</li>\n</ul>\n<p>summary of approach</p>\n<ul>\n<li>We still use \"graph instance norm\" (over node) [1], and smae gradient accumulation apporach, with batch size =64.</li>\n<li>We try both SAGE-conv[2] and GAT-conv[4].  GAT-conv gives better results.</li>\n<li>Since we are interested in top 5 ranks, we find listMLE is a better loss.</li>\n</ul>\n<hr>\n<h2>[Reference]</h2>\n<p>[1] GraphNorm: A Principled Approach to Accelerating Graph Neural Network Training<br>\n<a href=\"https://arxiv.org/abs/2009.03294\" target=\"_blank\">https://arxiv.org/abs/2009.03294</a>  <br>\n[2] Inductive Representation Learning on Large Graphs<br>\n<a href=\"https://arxiv.org/abs/1706.02216\" target=\"_blank\">https://arxiv.org/abs/1706.02216</a>  <br>\n[3] How Powerful are Graph Neural Networks?<br>\n<a href=\"https://arxiv.org/pdf/1810.00826.pdf\" target=\"_blank\">https://arxiv.org/pdf/1810.00826.pdf</a><br>\n[4] Graph Attention Networks <br>\n<a href=\"https://arxiv.org/abs/1710.10903\" target=\"_blank\">https://arxiv.org/abs/1710.10903</a></p>\n<hr>\n<h2>local validation and public/private score</h2>\n<p>The metric are: slowndown (top-5) for tile and kendall tau for layout.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0caab07b1cf25d81b358b1234d459b8c%2FSelection_999(4194).png?generation=1701488206301120&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h2>Acknowledgment</h2>\n<h2><em>\"I would like to express my sincere gratitude to HP for the generous provision of the  Z8-G4 Data Science Workstation that was instrumental in the successful completion of kaggle competition. The two 48GB Nvidia Quadro RTX 8000 GPU cards give me a distinct advantage to easily bulid models with the largest public graph dataset TPUGraphs, with 100 millions graphs of 10 thousands nodes.\"</em></h2>",
      "rawMarkdown": "# **Training and inference code can be downloaded from:**\nhttps://github.com/hengck23/solution-predict-ai-model-runtime/  \n&nbsp;\n\n\n## 1. Layout runtime prediction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb5c1c8af2f166b73061177ae313a2c6b%2FSelection_999(4192).png?generation=1701482437470415&alt=media)\n\nmain problem:\n- we have very large graph as input. how to design learning model and algorithm that can fit into gpu memory? \n\nsummary of approach :\n- instead of using the whole graph, we can reduce it by considering only the 5-hop neighbours from node marked as \"config id\". We call this 5-hop-neighbour subgraph.  We think this is reasonable becuase since we are comparing relative ranking of 2 graphs,  and we just need to input the \"difference nodes\"  (instead of the whole graph) to the neural net.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5e4c66d69dcf5ea590ad619c37dc2179%2FSelection_999(4193).png?generation=1701483096930612&alt=media)\n\n- using the reduced subgraph, i can sample 32 to 100 configurations using all full subgraphs with a single 48-GB  GPU card at training.\n- batch size is not an issue here, bcuase i am using gradient accumulation. We accumuluate over one subgraph at a time when training a batch.\n\n```\noptimizer.zero_grad()\n\nfor b in range(batch_size):\n        r = batch[r] \n        loss = net(r)  # forward one subgraph \n        scaler.scale(loss).backward() #backward accumuate gradient\n \nscaler.step(optimizer) #update net parameters\nscaler.update()\n```\n- normalisation is important. We use \"graph instance norm\" (over node), see paper[1], which works well with gradient accumulation \n\n- we use pairwise ranking loss in training loss.\n\n- We try 2 GNN: \n   - 4-layer SAGE-conv[2] with residual shortcut\n   - 4-layer GIN-conv[3] \n\nSAGE-conv is better than GIN-conv.\n\n## 2. Tile runtime prediction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc13d2c462b0958161cd4ce09e4249bc7%2FSelection_999(4191).png?generation=1701482450756660&alt=media)\n\nmain problem:\n- There is no issues here as the graph in the kaggle training data are much smaller. These are actually subgraphs of the much larger original computation graph.\n\nsummary of approach\n- We still use \"graph instance norm\" (over node) [1], and smae gradient accumulation apporach, with batch size =64.\n- We try both SAGE-conv[2] and GAT-conv[4].  GAT-conv gives better results.\n- Since we are interested in top 5 ranks, we find listMLE is a better loss.\n\n---\n\n## [Reference]\n\n[1] GraphNorm: A Principled Approach to Accelerating Graph Neural Network Training\nhttps://arxiv.org/abs/2009.03294  \n[2] Inductive Representation Learning on Large Graphs\nhttps://arxiv.org/abs/1706.02216  \n[3] How Powerful are Graph Neural Networks?\nhttps://arxiv.org/pdf/1810.00826.pdf\n[4] Graph Attention Networks \nhttps://arxiv.org/abs/1710.10903\n\n---\n\n## local validation and public/private score\n\nThe metric are: slowndown (top-5) for tile and kendall tau for layout.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0caab07b1cf25d81b358b1234d459b8c%2FSelection_999(4194).png?generation=1701488206301120&alt=media)\n\n---\n\n## Acknowledgment\n## *\"I would like to express my sincere gratitude to HP for the generous provision of the  Z8-G4 Data Science Workstation that was instrumental in the successful completion of kaggle competition. The two 48GB Nvidia Quadro RTX 8000 GPU cards give me a distinct advantage to easily bulid models with the largest public graph dataset TPUGraphs, with 100 millions graphs of 10 thousands nodes.\"*",
      "votes": null
    },
    {
      "id": "2529130",
      "postDate": "11/18/2023 01:09:42",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1ec4d755accef83d3fd01659f985e67e%2FSelection_999(3942).png?generation=1700269712680505&amp;alt=media\" alt=\"\"></p>\n<p>it seems that the trick to winning is to select the correct submission<br>\n(i.e. one should not only look at the average metric of the validation data, but the metric of each validation graph) </p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1ec4d755accef83d3fd01659f985e67e%2FSelection_999(3942).png?generation=1700269712680505&alt=media)\n\nit seems that the trick to winning is to select the correct submission\n(i.e. one should not only look at the average metric of the validation data, but the metric of each validation graph)",
      "votes": null
    },
    {
      "id": "2529135",
      "postDate": "11/18/2023 01:22:51",
      "content": "<p>Thaks for sharing. Congrats!</p>",
      "rawMarkdown": "Thaks for sharing. Congrats!",
      "votes": null
    },
    {
      "id": "2529901",
      "postDate": "11/18/2023 16:24:15",
      "content": "<p>Congratulations on achieving 6th position. Thanks for sharing details of  your  approach. </p>",
      "rawMarkdown": "Congratulations on achieving 6th position. Thanks for sharing details of  your  approach.",
      "votes": null
    },
    {
      "id": "2529904",
      "postDate": "11/18/2023 16:27:52",
      "content": "<p>Private LB of 0.713 would give you 3rd position. I think writing notes while making submissions will be helpful later to decide on selecting final submissions. Please note my note book on Game theory applications for Machine Learning Competitions. I have discussed strategies such as n-version programming for final selections. </p>",
      "rawMarkdown": "Private LB of 0.713 would give you 3rd position. I think writing notes while making submissions will be helpful later to decide on selecting final submissions. Please note my note book on Game theory applications for Machine Learning Competitions. I have discussed strategies such as n-version programming for final selections.",
      "votes": null
    },
    {
      "id": "2530048",
      "postDate": "11/18/2023 18:36:11",
      "content": "<p>You did it in a good way, and it's very good rank.</p>",
      "rawMarkdown": "You did it in a good way, and it's very good rank.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2529130,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/18/2023 01:09:42",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1ec4d755accef83d3fd01659f985e67e%2FSelection_999(3942).png?generation=1700269712680505&amp;alt=media\" alt=\"\"></p>\n<p>it seems that the trick to winning is to select the correct submission<br>\n(i.e. one should not only look at the average metric of the validation data, but the metric of each validation graph) </p>",
      "votes": null,
      "replies": [
        {
          "id": 2529904,
          "author_name": "crsuthikshnkumar",
          "author_url": "",
          "post_date": "11/18/2023 16:27:52",
          "content": "<p>Private LB of 0.713 would give you 3rd position. I think writing notes while making submissions will be helpful later to decide on selecting final submissions. Please note my note book on Game theory applications for Machine Learning Competitions. I have discussed strategies such as n-version programming for final selections. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2529135,
      "author_name": "haohantsao",
      "author_url": "",
      "post_date": "11/18/2023 01:22:51",
      "content": "<p>Thaks for sharing. Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2529901,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "11/18/2023 16:24:15",
      "content": "<p>Congratulations on achieving 6th position. Thanks for sharing details of  your  approach. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2530048,
      "author_name": "vikrampython",
      "author_url": "",
      "post_date": "11/18/2023 18:36:11",
      "content": "<p>You did it in a good way, and it's very good rank.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2529128": "# **Training and inference code can be downloaded from:**\nhttps://github.com/hengck23/solution-predict-ai-model-runtime/  \n&nbsp;\n\n\n## 1. Layout runtime prediction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb5c1c8af2f166b73061177ae313a2c6b%2FSelection_999(4192).png?generation=1701482437470415&alt=media)\n\nmain problem:\n- we have very large graph as input. how to design learning model and algorithm that can fit into gpu memory? \n\nsummary of approach :\n- instead of using the whole graph, we can reduce it by considering only the 5-hop neighbours from node marked as \"config id\". We call this 5-hop-neighbour subgraph.  We think this is reasonable becuase since we are comparing relative ranking of 2 graphs,  and we just need to input the \"difference nodes\"  (instead of the whole graph) to the neural net.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5e4c66d69dcf5ea590ad619c37dc2179%2FSelection_999(4193).png?generation=1701483096930612&alt=media)\n\n- using the reduced subgraph, i can sample 32 to 100 configurations using all full subgraphs with a single 48-GB  GPU card at training.\n- batch size is not an issue here, bcuase i am using gradient accumulation. We accumuluate over one subgraph at a time when training a batch.\n\n```\noptimizer.zero_grad()\n\nfor b in range(batch_size):\n        r = batch[r] \n        loss = net(r)  # forward one subgraph \n        scaler.scale(loss).backward() #backward accumuate gradient\n \nscaler.step(optimizer) #update net parameters\nscaler.update()\n```\n- normalisation is important. We use \"graph instance norm\" (over node), see paper[1], which works well with gradient accumulation \n\n- we use pairwise ranking loss in training loss.\n\n- We try 2 GNN: \n   - 4-layer SAGE-conv[2] with residual shortcut\n   - 4-layer GIN-conv[3] \n\nSAGE-conv is better than GIN-conv.\n\n## 2. Tile runtime prediction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc13d2c462b0958161cd4ce09e4249bc7%2FSelection_999(4191).png?generation=1701482450756660&alt=media)\n\nmain problem:\n- There is no issues here as the graph in the kaggle training data are much smaller. These are actually subgraphs of the much larger original computation graph.\n\nsummary of approach\n- We still use \"graph instance norm\" (over node) [1], and smae gradient accumulation apporach, with batch size =64.\n- We try both SAGE-conv[2] and GAT-conv[4].  GAT-conv gives better results.\n- Since we are interested in top 5 ranks, we find listMLE is a better loss.\n\n---\n\n## [Reference]\n\n[1] GraphNorm: A Principled Approach to Accelerating Graph Neural Network Training\nhttps://arxiv.org/abs/2009.03294  \n[2] Inductive Representation Learning on Large Graphs\nhttps://arxiv.org/abs/1706.02216  \n[3] How Powerful are Graph Neural Networks?\nhttps://arxiv.org/pdf/1810.00826.pdf\n[4] Graph Attention Networks \nhttps://arxiv.org/abs/1710.10903\n\n---\n\n## local validation and public/private score\n\nThe metric are: slowndown (top-5) for tile and kendall tau for layout.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0caab07b1cf25d81b358b1234d459b8c%2FSelection_999(4194).png?generation=1701488206301120&alt=media)\n\n---\n\n## Acknowledgment\n## *\"I would like to express my sincere gratitude to HP for the generous provision of the  Z8-G4 Data Science Workstation that was instrumental in the successful completion of kaggle competition. The two 48GB Nvidia Quadro RTX 8000 GPU cards give me a distinct advantage to easily bulid models with the largest public graph dataset TPUGraphs, with 100 millions graphs of 10 thousands nodes.\"*",
    "2529130": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1ec4d755accef83d3fd01659f985e67e%2FSelection_999(3942).png?generation=1700269712680505&alt=media)\n\nit seems that the trick to winning is to select the correct submission\n(i.e. one should not only look at the average metric of the validation data, but the metric of each validation graph)",
    "2529135": "Thaks for sharing. Congrats!",
    "2529901": "Congratulations on achieving 6th position. Thanks for sharing details of  your  approach.",
    "2529904": "Private LB of 0.713 would give you 3rd position. I think writing notes while making submissions will be helpful later to decide on selecting final submissions. Please note my note book on Game theory applications for Machine Learning Competitions. I have discussed strategies such as n-version programming for final selections.",
    "2530048": "You did it in a good way, and it's very good rank."
  },
  "source": "meta"
}