{
  "id": 456074,
  "title": "19th Place Solution",
  "url": "/competitions/predict-ai-model-runtime/writeups/pistacho-19th-place-solution",
  "author_name": "",
  "post_date": "2023-11-18T00:05:27.230Z",
  "votes": 14,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to Kaggle and Google for organizing this competition! This was my first competition and I really enjoyed it. We didn't know anything about GNN's before starting this competition and we learned a lot throughout this competition. I would also like to thank my great teammate <a href=\"https://www.kaggle.com/roysegalz\" target=\"_blank\">@roysegalz</a> </p>\n<h3>Tile dataset:</h3>\n<p>For the Tile dataset we used an RGCN where the relations are the node_opcode. The model consisted of 4 blocks where a block consisted of RGCN -&gt; LayerNorm -&gt; ReLU</p>\n<ul>\n<li>loss function - ListMLE</li>\n<li>hidden dim = 128</li>\n<li>CosineLRScheduler</li>\n<li>AdamW</li>\n<li>lr = 4e-4</li>\n<li>weight decay = 1e-4</li>\n<li>max aggr for final graph representation</li>\n</ul>\n<p>We reached 0.198 on Tile + sample_submission (random predictions on Layout datasets)</p>\n<h3>Layout datasets:</h3>\n<p>We realized that the best way to differentiate between different graphs is by emphasizing the node_config_feat. We recognized the problem that the number of nodes that contain node configs is relatively small compared to the number of nodes in the graph. Hence we understood that this data might get lost in the model throughout the forward pass. Our way to solve the problem was the following:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6061339%2Fc3795f9d4cb3cd838f865260a2433cd1%2FNew%20node%20config%20feat.png?generation=1700264018548073&amp;alt=media\" alt=\"\"></p>\n<p>First we represent the node_config_feat in a different way (nn.Embedding and one hot vector). After doing that we created this model architecture (which is the most important part of our solution):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6061339%2Fd852425e39c06ed17a4d2c822fda94f4%2FLayout%20Architecture.png?generation=1700264126117504&amp;alt=media\" alt=\"\"></p>\n<p>Concatenating the data again after each block increased our scores dramatically.</p>\n<p>We used the same model architecture for all the different layout datasets.</p>\n<p>The model consisted of 3 blocks where a block consisted of GATv2-&gt; LayerNorm -&gt; ReLU (We decided not to use RGCN since it was really computationally expensive to run on the layout dataset)</p>\n<ul>\n<li>hidden dim = 64</li>\n<li>loss function - nn.MarginRankingLoss(0.5) </li>\n<li>CosineLRScheduler</li>\n<li>AdamW</li>\n<li>lr = 2e-4</li>\n<li>weight decay = 1e-4</li>\n<li>Virtual node for final graph representation</li>\n</ul>\n<p>Our CV Scores:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Kendal-Tau CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XLA default</td>\n<td>~0.3</td>\n</tr>\n<tr>\n<td>NLP default</td>\n<td>~0.5</td>\n</tr>\n<tr>\n<td>XLA random</td>\n<td>~0.62</td>\n</tr>\n<tr>\n<td>NLP random</td>\n<td>~0.94</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2529097",
      "postDate": "11/18/2023 00:02:32",
      "content": "<p>Thanks to Kaggle and Google for organizing this competition! This was my first competition and I really enjoyed it. We didn't know anything about GNN's before starting this competition and we learned a lot throughout this competition. I would also like to thank my great teammate <a href=\"https://www.kaggle.com/roysegalz\" target=\"_blank\">@roysegalz</a> </p>\n<h3>Tile dataset:</h3>\n<p>For the Tile dataset we used an RGCN where the relations are the node_opcode. The model consisted of 4 blocks where a block consisted of RGCN -&gt; LayerNorm -&gt; ReLU</p>\n<ul>\n<li>loss function - ListMLE</li>\n<li>hidden dim = 128</li>\n<li>CosineLRScheduler</li>\n<li>AdamW</li>\n<li>lr = 4e-4</li>\n<li>weight decay = 1e-4</li>\n<li>max aggr for final graph representation</li>\n</ul>\n<p>We reached 0.198 on Tile + sample_submission (random predictions on Layout datasets)</p>\n<h3>Layout datasets:</h3>\n<p>We realized that the best way to differentiate between different graphs is by emphasizing the node_config_feat. We recognized the problem that the number of nodes that contain node configs is relatively small compared to the number of nodes in the graph. Hence we understood that this data might get lost in the model throughout the forward pass. Our way to solve the problem was the following:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6061339%2Fc3795f9d4cb3cd838f865260a2433cd1%2FNew%20node%20config%20feat.png?generation=1700264018548073&amp;alt=media\" alt=\"\"></p>\n<p>First we represent the node_config_feat in a different way (nn.Embedding and one hot vector). After doing that we created this model architecture (which is the most important part of our solution):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6061339%2Fd852425e39c06ed17a4d2c822fda94f4%2FLayout%20Architecture.png?generation=1700264126117504&amp;alt=media\" alt=\"\"></p>\n<p>Concatenating the data again after each block increased our scores dramatically.</p>\n<p>We used the same model architecture for all the different layout datasets.</p>\n<p>The model consisted of 3 blocks where a block consisted of GATv2-&gt; LayerNorm -&gt; ReLU (We decided not to use RGCN since it was really computationally expensive to run on the layout dataset)</p>\n<ul>\n<li>hidden dim = 64</li>\n<li>loss function - nn.MarginRankingLoss(0.5) </li>\n<li>CosineLRScheduler</li>\n<li>AdamW</li>\n<li>lr = 2e-4</li>\n<li>weight decay = 1e-4</li>\n<li>Virtual node for final graph representation</li>\n</ul>\n<p>Our CV Scores:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Kendal-Tau CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XLA default</td>\n<td>~0.3</td>\n</tr>\n<tr>\n<td>NLP default</td>\n<td>~0.5</td>\n</tr>\n<tr>\n<td>XLA random</td>\n<td>~0.62</td>\n</tr>\n<tr>\n<td>NLP random</td>\n<td>~0.94</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Thanks to Kaggle and Google for organizing this competition! This was my first competition and I really enjoyed it. We didn't know anything about GNN's before starting this competition and we learned a lot throughout this competition. I would also like to thank my great teammate @roysegalz \n\n### Tile dataset:\nFor the Tile dataset we used an RGCN where the relations are the node_opcode. The model consisted of 4 blocks where a block consisted of RGCN -> LayerNorm -> ReLU\n* loss function - ListMLE\n* hidden dim = 128\n* CosineLRScheduler\n* AdamW\n* lr = 4e-4\n* weight decay = 1e-4\n* max aggr for final graph representation\n\nWe reached 0.198 on Tile + sample_submission (random predictions on Layout datasets)\n\n### Layout datasets:\nWe realized that the best way to differentiate between different graphs is by emphasizing the node_config_feat. We recognized the problem that the number of nodes that contain node configs is relatively small compared to the number of nodes in the graph. Hence we understood that this data might get lost in the model throughout the forward pass. Our way to solve the problem was the following:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6061339%2Fc3795f9d4cb3cd838f865260a2433cd1%2FNew%20node%20config%20feat.png?generation=1700264018548073&alt=media)\n\nFirst we represent the node_config_feat in a different way (nn.Embedding and one hot vector). After doing that we created this model architecture (which is the most important part of our solution):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6061339%2Fd852425e39c06ed17a4d2c822fda94f4%2FLayout%20Architecture.png?generation=1700264126117504&alt=media)\n\nConcatenating the data again after each block increased our scores dramatically.\n\nWe used the same model architecture for all the different layout datasets.\n\nThe model consisted of 3 blocks where a block consisted of GATv2-> LayerNorm -> ReLU (We decided not to use RGCN since it was really computationally expensive to run on the layout dataset)\n* hidden dim = 64\n* loss function - nn.MarginRankingLoss(0.5) \n* CosineLRScheduler\n* AdamW\n* lr = 2e-4\n* weight decay = 1e-4\n* Virtual node for final graph representation\n\n\nOur CV Scores:\n\n| Dataset | Kendal-Tau CV |\n| --- | --- |\n| XLA default| ~0.3 |\n| NLP default | ~0.5 |\n| XLA random| ~0.62 |\n| NLP random| ~0.94 |",
      "votes": null
    },
    {
      "id": "2529132",
      "postDate": "11/18/2023 01:20:14",
      "content": "<p>thanks for the writeup.</p>\n<p>i have a question, why did you use GATv2?<br>\nmy experiments show that GATv2 is usually worse than sageConv for all layout cases.</p>\n<p>(but GATv2 is indeed better for tile case (only))</p>",
      "rawMarkdown": "thanks for the writeup.\n\ni have a question, why did you use GATv2?\nmy experiments show that GATv2 is usually worse than sageConv for all layout cases.\n\n(but GATv2 is indeed better for tile case (only))",
      "votes": null
    },
    {
      "id": "2529298",
      "postDate": "11/18/2023 05:26:42",
      "content": "<p>For us using gatv2 did better than sage by quite a margin</p>",
      "rawMarkdown": "For us using gatv2 did better than sage by quite a margin",
      "votes": null
    },
    {
      "id": "2529764",
      "postDate": "11/18/2023 14:21:00",
      "content": "<p>Had a great time working with you and learned a lot from you! <a href=\"https://www.kaggle.com/amitaharoni\" target=\"_blank\">@amitaharoni</a> </p>",
      "rawMarkdown": "Had a great time working with you and learned a lot from you! @amitaharoni",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2529132,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/18/2023 01:20:14",
      "content": "<p>thanks for the writeup.</p>\n<p>i have a question, why did you use GATv2?<br>\nmy experiments show that GATv2 is usually worse than sageConv for all layout cases.</p>\n<p>(but GATv2 is indeed better for tile case (only))</p>",
      "votes": null,
      "replies": [
        {
          "id": 2529298,
          "author_name": "amitaharoni",
          "author_url": "",
          "post_date": "11/18/2023 05:26:42",
          "content": "<p>For us using gatv2 did better than sage by quite a margin</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2529764,
      "author_name": "roysegalz",
      "author_url": "",
      "post_date": "11/18/2023 14:21:00",
      "content": "<p>Had a great time working with you and learned a lot from you! <a href=\"https://www.kaggle.com/amitaharoni\" target=\"_blank\">@amitaharoni</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2529097": "Thanks to Kaggle and Google for organizing this competition! This was my first competition and I really enjoyed it. We didn't know anything about GNN's before starting this competition and we learned a lot throughout this competition. I would also like to thank my great teammate @roysegalz \n\n### Tile dataset:\nFor the Tile dataset we used an RGCN where the relations are the node_opcode. The model consisted of 4 blocks where a block consisted of RGCN -> LayerNorm -> ReLU\n* loss function - ListMLE\n* hidden dim = 128\n* CosineLRScheduler\n* AdamW\n* lr = 4e-4\n* weight decay = 1e-4\n* max aggr for final graph representation\n\nWe reached 0.198 on Tile + sample_submission (random predictions on Layout datasets)\n\n### Layout datasets:\nWe realized that the best way to differentiate between different graphs is by emphasizing the node_config_feat. We recognized the problem that the number of nodes that contain node configs is relatively small compared to the number of nodes in the graph. Hence we understood that this data might get lost in the model throughout the forward pass. Our way to solve the problem was the following:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6061339%2Fc3795f9d4cb3cd838f865260a2433cd1%2FNew%20node%20config%20feat.png?generation=1700264018548073&alt=media)\n\nFirst we represent the node_config_feat in a different way (nn.Embedding and one hot vector). After doing that we created this model architecture (which is the most important part of our solution):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6061339%2Fd852425e39c06ed17a4d2c822fda94f4%2FLayout%20Architecture.png?generation=1700264126117504&alt=media)\n\nConcatenating the data again after each block increased our scores dramatically.\n\nWe used the same model architecture for all the different layout datasets.\n\nThe model consisted of 3 blocks where a block consisted of GATv2-> LayerNorm -> ReLU (We decided not to use RGCN since it was really computationally expensive to run on the layout dataset)\n* hidden dim = 64\n* loss function - nn.MarginRankingLoss(0.5) \n* CosineLRScheduler\n* AdamW\n* lr = 2e-4\n* weight decay = 1e-4\n* Virtual node for final graph representation\n\n\nOur CV Scores:\n\n| Dataset | Kendal-Tau CV |\n| --- | --- |\n| XLA default| ~0.3 |\n| NLP default | ~0.5 |\n| XLA random| ~0.62 |\n| NLP random| ~0.94 |",
    "2529132": "thanks for the writeup.\n\ni have a question, why did you use GATv2?\nmy experiments show that GATv2 is usually worse than sageConv for all layout cases.\n\n(but GATv2 is indeed better for tile case (only))",
    "2529298": "For us using gatv2 did better than sage by quite a margin",
    "2529764": "Had a great time working with you and learned a lot from you! @amitaharoni"
  },
  "source": "meta"
}