{
  "id": 456625,
  "title": "97th place solution for Google - Fast or Slow? Predict AI Model Runtime",
  "url": "/competitions/predict-ai-model-runtime/writeups/min-hsien-weng-97th-place-solution-for-google-fast",
  "author_name": "",
  "post_date": "2023-11-30T21:08:05.097Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thank you to Kaggle and Google for hosting this great competition. I learned a lot about training graph data models, especially how to deal with memory efficiency and optimize performance. </p>\n<p>I started by training a modified BERT model by <a href=\"https://www.kaggle.com/KSMCG90\" target=\"_blank\">@KSMCG90</a> using cross-validation and the <a href=\"https://lightning.ai/\" target=\"_blank\">Lightning AI</a> library. But this model was too demanding on memory for the layout dataset. So, I switched to the GNN (Graph Neural Network) model by <a href=\"https://www.kaggle.com/GUSTHEMA\" target=\"_blank\">@GUSTHEMA</a> for the layout task. Since the tile and layout datasets didn't show any significant correlation, I built two separate models: one BERT model for the tile dataset and one GNN model for the layout dataset. </p>\n<p>The solution <a href=\"https://www.kaggle.com/code/minhsienweng/bertlike-tile-model-gnn-layout-model\" target=\"_blank\">notebook</a> focuses on training two models and run them efficiently on limited computing resources (4-core CPUs).</p>\n<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/data</a></p>\n<h1>Overview of the Approach</h1>\n<p>The solution comprised two models: a GNN model for the layout dataset and a modified BERT model for the tile dataset. The inputs were the graph structures generated by an AI model. In this graph, nodes represented tensor operations (e.g., add, max, reshape, convolution) and edges represented tensors.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16611574%2F4548b69d2c2e5cbc11e502a30fece82c%2FPresentation1.jpg?generation=1700528423839754&amp;alt=media\" alt=\"Overview\"></p>\n<h2>TF-GNN Residual Layout model</h2>\n<p>I chose the residual model for layout configurations because it can measure the difference between predicted and actual runtime time of each layout configuration. The model uses TensorFlow's graph neural network library (<a href=\"https://github.com/tensorflow/gnn\" target=\"_blank\">TF-GNN</a>) to train efficiently on TPUs. I trained the layout model on TPUs and saved the best model as a pretrained model for prediction. In this notebook, I built four separate models, one for each type of layout dataset: <code>nlp-default</code>, <code>nlp-random</code>, <code>xla-default</code>, and <code>xla-random</code>.</p>\n<p>The layout model encodes the layout graph into graph embeddings (<code>node sets</code>, <code>edge sets</code>, and <code>context</code>) and includes 5 hidden layers and a global pooling layer. The model is then trained on 10 epochs, with <code>ListMLE</code> ranking loss as the loss function and ordered pair accuracy (OPA) as the performance metric. The optimizer is AdamW with a learning rate of $10^{-3}$. </p>\n<p><em>Things to do</em>: The model can be improved by adjusting the learning rate with cross validation and by pruning irrelevant <code>context</code> to improve memory efficiency. Adding <a href=\"https://github.com/vijaydwivedi75/gnn-lspe\" target=\"_blank\">positional encodings</a> may also help. <a href=\"https://pyg.org/\" target=\"_blank\">PyG</a> library could be used to build and train GNN model. x</p>\n<h2>Modified BERT GAN Tile model</h2>\n<p>Similar to layout dataset, the tile data is encoded into graph embeddings that include 9 data columns (<code>node_feat</code>, <code>node_opcode</code>, <code>edge_index</code>, <code>config_runtime</code>, <code>config_runtime_normalizers</code>, <code>edges_adjecency</code>, <code>node_config_feat</code>, <code>node_config_ids</code> and <code>selected_idxs</code>).</p>\n<p>The tile model is based on the BERT model and uses Graph Attention Networks (<a href=\"https://arxiv.org/abs/1710.10903\" target=\"_blank\">GANs</a>) to give more attention to important parts of the data. This helps improve accuracy by focusing on the most relevant information. The tile model uses a GraphEncoder to encode the tile configuration into graph embeddings and applies a special mask to all layers and heads. This mask is based on the connections between different parts of the graph. The model's loss function compares its predictions to different permutations of the actual results.</p>\n<p>Lastly, the model was wrapped with <a href=\"https://lightning.ai/\" target=\"_blank\">Lightning AI</a> library to perform cross-validation using five folds of training and validation datasets. The best model from each fold was saved and then used to further train and predict the testing dataset.</p>\n<h1>Details of the submission</h1>\n<h2>Training layout model with checkpoints</h2>\n<p>I tried to make the BERT model work with the layout dataset, but it needed too much memory (&gt;300GB). So, I used the TF-GNN-based residual model that uses less memory (&lt;30GB). However, training the TF-GNN model is time-consuming. For instance, training one epoch on the 'xla-random' dataset required approximately one hour on CPUs. Training for 10 epochs would exceed the competition's time constraints (&gt;10 hours). Despite utilizing GPUs to accelerate the training process, the training time has little/no speedups. </p>\n<p>To speed up training, I used a checkpointing technique to save the model's weights after each epoch. This allowed me to resume training, saving time. The best model was saved and used to predict the test datasets.</p>\n<h2>Parallelising the infer loop of tile model</h2>\n<p>The infer process for the tile dataset is time-consuming, for example, the model requires approximately 1.81 hours (6,532 seconds) to predict the testing dataset of 844 tile configurations. This is because the infer loop involves running each fold model separately and then merging the results from the five models to get the final output. I tried to speed up the infer loop by using five separate threads to run the inference for each fold's model concurrently. Using a thread pool of three threads, the total infer time is reduced to 58 minutes (3,448 seconds), achieving approximately a 2x speedup on CPUs.</p>\n<p>This experiment demonstrates that parallelism techniques can be effectively applied to machine learning models, enabling them to run on standard hardware configurations. This eliminates the need for specialized hardware, making machine learning more accessible to students and researchers with limited computing resources.</p>\n<h1>Sources</h1>\n<p>The final solution was inspired by the starter notebook by <a href=\"https://www.kaggle.com/GUSTHEMA\" target=\"_blank\">@GUSTHEMA</a> and BERT-like GAT Tile wLayout Dataset by <a href=\"https://www.kaggle.com/KSMCG90\" target=\"_blank\">@KSMCG90</a></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/gusthema/starter-notebook-fast-or-slow-with-tensorflow-gnn\" target=\"_blank\">https://www.kaggle.com/code/gusthema/starter-notebook-fast-or-slow-with-tensorflow-gnn</a></li>\n<li><a href=\"https://www.kaggle.com/code/ksmcg90/bertlike-gat-tile-wlayout-dataset?scriptVersionId=148717592\" target=\"_blank\">https://www.kaggle.com/code/ksmcg90/bertlike-gat-tile-wlayout-dataset?scriptVersionId=148717592</a></li>\n</ul>",
  "messages": [
    {
      "id": "2532327",
      "postDate": "11/21/2023 00:47:39",
      "content": "<p>Thank you to Kaggle and Google for hosting this great competition. I learned a lot about training graph data models, especially how to deal with memory efficiency and optimize performance. </p>\n<p>I started by training a modified BERT model by <a href=\"https://www.kaggle.com/KSMCG90\" target=\"_blank\">@KSMCG90</a> using cross-validation and the <a href=\"https://lightning.ai/\" target=\"_blank\">Lightning AI</a> library. But this model was too demanding on memory for the layout dataset. So, I switched to the GNN (Graph Neural Network) model by <a href=\"https://www.kaggle.com/GUSTHEMA\" target=\"_blank\">@GUSTHEMA</a> for the layout task. Since the tile and layout datasets didn't show any significant correlation, I built two separate models: one BERT model for the tile dataset and one GNN model for the layout dataset. </p>\n<p>The solution <a href=\"https://www.kaggle.com/code/minhsienweng/bertlike-tile-model-gnn-layout-model\" target=\"_blank\">notebook</a> focuses on training two models and run them efficiently on limited computing resources (4-core CPUs).</p>\n<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/data</a></p>\n<h1>Overview of the Approach</h1>\n<p>The solution comprised two models: a GNN model for the layout dataset and a modified BERT model for the tile dataset. The inputs were the graph structures generated by an AI model. In this graph, nodes represented tensor operations (e.g., add, max, reshape, convolution) and edges represented tensors.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16611574%2F4548b69d2c2e5cbc11e502a30fece82c%2FPresentation1.jpg?generation=1700528423839754&amp;alt=media\" alt=\"Overview\"></p>\n<h2>TF-GNN Residual Layout model</h2>\n<p>I chose the residual model for layout configurations because it can measure the difference between predicted and actual runtime time of each layout configuration. The model uses TensorFlow's graph neural network library (<a href=\"https://github.com/tensorflow/gnn\" target=\"_blank\">TF-GNN</a>) to train efficiently on TPUs. I trained the layout model on TPUs and saved the best model as a pretrained model for prediction. In this notebook, I built four separate models, one for each type of layout dataset: <code>nlp-default</code>, <code>nlp-random</code>, <code>xla-default</code>, and <code>xla-random</code>.</p>\n<p>The layout model encodes the layout graph into graph embeddings (<code>node sets</code>, <code>edge sets</code>, and <code>context</code>) and includes 5 hidden layers and a global pooling layer. The model is then trained on 10 epochs, with <code>ListMLE</code> ranking loss as the loss function and ordered pair accuracy (OPA) as the performance metric. The optimizer is AdamW with a learning rate of $10^{-3}$. </p>\n<p><em>Things to do</em>: The model can be improved by adjusting the learning rate with cross validation and by pruning irrelevant <code>context</code> to improve memory efficiency. Adding <a href=\"https://github.com/vijaydwivedi75/gnn-lspe\" target=\"_blank\">positional encodings</a> may also help. <a href=\"https://pyg.org/\" target=\"_blank\">PyG</a> library could be used to build and train GNN model. x</p>\n<h2>Modified BERT GAN Tile model</h2>\n<p>Similar to layout dataset, the tile data is encoded into graph embeddings that include 9 data columns (<code>node_feat</code>, <code>node_opcode</code>, <code>edge_index</code>, <code>config_runtime</code>, <code>config_runtime_normalizers</code>, <code>edges_adjecency</code>, <code>node_config_feat</code>, <code>node_config_ids</code> and <code>selected_idxs</code>).</p>\n<p>The tile model is based on the BERT model and uses Graph Attention Networks (<a href=\"https://arxiv.org/abs/1710.10903\" target=\"_blank\">GANs</a>) to give more attention to important parts of the data. This helps improve accuracy by focusing on the most relevant information. The tile model uses a GraphEncoder to encode the tile configuration into graph embeddings and applies a special mask to all layers and heads. This mask is based on the connections between different parts of the graph. The model's loss function compares its predictions to different permutations of the actual results.</p>\n<p>Lastly, the model was wrapped with <a href=\"https://lightning.ai/\" target=\"_blank\">Lightning AI</a> library to perform cross-validation using five folds of training and validation datasets. The best model from each fold was saved and then used to further train and predict the testing dataset.</p>\n<h1>Details of the submission</h1>\n<h2>Training layout model with checkpoints</h2>\n<p>I tried to make the BERT model work with the layout dataset, but it needed too much memory (&gt;300GB). So, I used the TF-GNN-based residual model that uses less memory (&lt;30GB). However, training the TF-GNN model is time-consuming. For instance, training one epoch on the 'xla-random' dataset required approximately one hour on CPUs. Training for 10 epochs would exceed the competition's time constraints (&gt;10 hours). Despite utilizing GPUs to accelerate the training process, the training time has little/no speedups. </p>\n<p>To speed up training, I used a checkpointing technique to save the model's weights after each epoch. This allowed me to resume training, saving time. The best model was saved and used to predict the test datasets.</p>\n<h2>Parallelising the infer loop of tile model</h2>\n<p>The infer process for the tile dataset is time-consuming, for example, the model requires approximately 1.81 hours (6,532 seconds) to predict the testing dataset of 844 tile configurations. This is because the infer loop involves running each fold model separately and then merging the results from the five models to get the final output. I tried to speed up the infer loop by using five separate threads to run the inference for each fold's model concurrently. Using a thread pool of three threads, the total infer time is reduced to 58 minutes (3,448 seconds), achieving approximately a 2x speedup on CPUs.</p>\n<p>This experiment demonstrates that parallelism techniques can be effectively applied to machine learning models, enabling them to run on standard hardware configurations. This eliminates the need for specialized hardware, making machine learning more accessible to students and researchers with limited computing resources.</p>\n<h1>Sources</h1>\n<p>The final solution was inspired by the starter notebook by <a href=\"https://www.kaggle.com/GUSTHEMA\" target=\"_blank\">@GUSTHEMA</a> and BERT-like GAT Tile wLayout Dataset by <a href=\"https://www.kaggle.com/KSMCG90\" target=\"_blank\">@KSMCG90</a></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/gusthema/starter-notebook-fast-or-slow-with-tensorflow-gnn\" target=\"_blank\">https://www.kaggle.com/code/gusthema/starter-notebook-fast-or-slow-with-tensorflow-gnn</a></li>\n<li><a href=\"https://www.kaggle.com/code/ksmcg90/bertlike-gat-tile-wlayout-dataset?scriptVersionId=148717592\" target=\"_blank\">https://www.kaggle.com/code/ksmcg90/bertlike-gat-tile-wlayout-dataset?scriptVersionId=148717592</a></li>\n</ul>",
      "rawMarkdown": "Thank you to Kaggle and Google for hosting this great competition. I learned a lot about training graph data models, especially how to deal with memory efficiency and optimize performance. \n\nI started by training a modified BERT model by @KSMCG90 using cross-validation and the [Lightning AI](https://lightning.ai/) library. But this model was too demanding on memory for the layout dataset. So, I switched to the GNN (Graph Neural Network) model by @GUSTHEMA for the layout task. Since the tile and layout datasets didn't show any significant correlation, I built two separate models: one BERT model for the tile dataset and one GNN model for the layout dataset. \n\nThe solution [notebook](https://www.kaggle.com/code/minhsienweng/bertlike-tile-model-gnn-layout-model) focuses on training two models and run them efficiently on limited computing resources (4-core CPUs).\n\n# Context \n\nBusiness context: https://www.kaggle.com/competitions/predict-ai-model-runtime/\nData context: https://www.kaggle.com/competitions/predict-ai-model-runtime/data\n\n# Overview of the Approach\nThe solution comprised two models: a GNN model for the layout dataset and a modified BERT model for the tile dataset. The inputs were the graph structures generated by an AI model. In this graph, nodes represented tensor operations (e.g., add, max, reshape, convolution) and edges represented tensors.\n\n![Overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16611574%2F4548b69d2c2e5cbc11e502a30fece82c%2FPresentation1.jpg?generation=1700528423839754&alt=media)\n \n## TF-GNN Residual Layout model\nI chose the residual model for layout configurations because it can measure the difference between predicted and actual runtime time of each layout configuration. The model uses TensorFlow's graph neural network library ([TF-GNN](https://github.com/tensorflow/gnn)) to train efficiently on TPUs. I trained the layout model on TPUs and saved the best model as a pretrained model for prediction. In this notebook, I built four separate models, one for each type of layout dataset: `nlp-default`, `nlp-random`, `xla-default`, and `xla-random`.\n\nThe layout model encodes the layout graph into graph embeddings (`node sets`, `edge sets`, and `context`) and includes 5 hidden layers and a global pooling layer. The model is then trained on 10 epochs, with `ListMLE` ranking loss as the loss function and ordered pair accuracy (OPA) as the performance metric. The optimizer is AdamW with a learning rate of $10^{-3}$. \n\n*Things to do*: The model can be improved by adjusting the learning rate with cross validation and by pruning irrelevant `context` to improve memory efficiency. Adding [positional encodings](https://github.com/vijaydwivedi75/gnn-lspe) may also help. [PyG](https://pyg.org/) library could be used to build and train GNN model. x\n\n## Modified BERT GAN Tile model\nSimilar to layout dataset, the tile data is encoded into graph embeddings that include 9 data columns (`node_feat`, `node_opcode`, `edge_index`, `config_runtime`, `config_runtime_normalizers`, `edges_adjecency`, `node_config_feat`, `node_config_ids` and `selected_idxs`).\n\nThe tile model is based on the BERT model and uses Graph Attention Networks ([GANs](https://arxiv.org/abs/1710.10903)) to give more attention to important parts of the data. This helps improve accuracy by focusing on the most relevant information. The tile model uses a GraphEncoder to encode the tile configuration into graph embeddings and applies a special mask to all layers and heads. This mask is based on the connections between different parts of the graph. The model's loss function compares its predictions to different permutations of the actual results.\n\nLastly, the model was wrapped with [Lightning AI](https://lightning.ai/) library to perform cross-validation using five folds of training and validation datasets. The best model from each fold was saved and then used to further train and predict the testing dataset.\n\n# Details of the submission\n\n## Training layout model with checkpoints\nI tried to make the BERT model work with the layout dataset, but it needed too much memory (>300GB). So, I used the TF-GNN-based residual model that uses less memory (<30GB). However, training the TF-GNN model is time-consuming. For instance, training one epoch on the 'xla-random' dataset required approximately one hour on CPUs. Training for 10 epochs would exceed the competition's time constraints (>10 hours). Despite utilizing GPUs to accelerate the training process, the training time has little/no speedups. \n\nTo speed up training, I used a checkpointing technique to save the model's weights after each epoch. This allowed me to resume training, saving time. The best model was saved and used to predict the test datasets.\n\n## Parallelising the infer loop of tile model\nThe infer process for the tile dataset is time-consuming, for example, the model requires approximately 1.81 hours (6,532 seconds) to predict the testing dataset of 844 tile configurations. This is because the infer loop involves running each fold model separately and then merging the results from the five models to get the final output. I tried to speed up the infer loop by using five separate threads to run the inference for each fold's model concurrently. Using a thread pool of three threads, the total infer time is reduced to 58 minutes (3,448 seconds), achieving approximately a 2x speedup on CPUs.\n\nThis experiment demonstrates that parallelism techniques can be effectively applied to machine learning models, enabling them to run on standard hardware configurations. This eliminates the need for specialized hardware, making machine learning more accessible to students and researchers with limited computing resources.\n\n\n# Sources\nThe final solution was inspired by the starter notebook by @GUSTHEMA and BERT-like GAT Tile wLayout Dataset by @KSMCG90\n- https://www.kaggle.com/code/gusthema/starter-notebook-fast-or-slow-with-tensorflow-gnn\n- https://www.kaggle.com/code/ksmcg90/bertlike-gat-tile-wlayout-dataset?scriptVersionId=148717592",
      "votes": null
    },
    {
      "id": "2533176",
      "postDate": "11/21/2023 16:37:00",
      "content": "<p>Thank you for sharing your write-up with Modified BERT GAN and infer loop.</p>",
      "rawMarkdown": "Thank you for sharing your write-up with Modified BERT GAN and infer loop.",
      "votes": null
    },
    {
      "id": "2533380",
      "postDate": "11/21/2023 20:53:14",
      "content": "<p>Thanks for liking my solution. Please feel free to reach out if you have any questions.</p>",
      "rawMarkdown": "Thanks for liking my solution. Please feel free to reach out if you have any questions.",
      "votes": null
    },
    {
      "id": "2533445",
      "postDate": "11/21/2023 23:57:57",
      "content": "<p>I'm just a beginner, therefore it's hard to make questions since all is very new to me. I hope that other experienced kagglers could read it and learn from your approach. Thank you M. Weng.</p>",
      "rawMarkdown": "I'm just a beginner, therefore it's hard to make questions since all is very new to me. I hope that other experienced kagglers could read it and learn from your approach. Thank you M. Weng.",
      "votes": null
    },
    {
      "id": "2533459",
      "postDate": "11/22/2023 00:32:18",
      "content": "<p>I am still learning about GNNs, but I am interested to explore their potential.</p>\n<p>Kaggle provides an excellent platform for learning and experimenting with novel ML models. Most importantly,  enjoy the competition and have fun! 😁</p>",
      "rawMarkdown": "I am still learning about GNNs, but I am interested to explore their potential.\n\nKaggle provides an excellent platform for learning and experimenting with novel ML models. Most importantly,  enjoy the competition and have fun! 😁",
      "votes": null
    },
    {
      "id": "2533473",
      "postDate": "11/22/2023 00:44:17",
      "content": "<p>The more you grow (on Kaggle) cause you're already experienced, other users will start to pay attention on your work. Unfortunately, our community is young and shallow. </p>\n<p>So, they don't read much what contributors/experts are writting.  Worst for them, since they lose a precious opportunity to learn from the best professionals. Besides, those that are working, don't have enough time to spend many hours kaggling : )</p>",
      "rawMarkdown": "The more you grow (on Kaggle) cause you're already experienced, other users will start to pay attention on your work. Unfortunately, our community is young and shallow. \n\nSo, they don't read much what contributors/experts are writting.  Worst for them, since they lose a precious opportunity to learn from the best professionals. Besides, those that are working, don't have enough time to spend many hours kaggling : )",
      "votes": null
    },
    {
      "id": "2533502",
      "postDate": "11/22/2023 01:39:18",
      "content": "<p>I wrote the solution partly for myself. If I encounter similar problems, I can leverage the knowledge and experience I gained from this competition.</p>\n<p>Thank you for your encouragement! 😀</p>",
      "rawMarkdown": "I wrote the solution partly for myself. If I encounter similar problems, I can leverage the knowledge and experience I gained from this competition.\n\nThank you for your encouragement! 😀",
      "votes": null
    },
    {
      "id": "2533511",
      "postDate": "11/22/2023 01:54:22",
      "content": "<p>\"Better to write for ourselves (and have few audience), than to write for the others (and have No self)\" Cyril Connoly quote.</p>\n<p>Besides, since everything is on the cloud, anyone/anywhere could read it on the future. Even yourself, can compare how do you think and apply those methods and metrics today and what you could deliver in the future.</p>",
      "rawMarkdown": "\"Better to write for ourselves (and have few audience), than to write for the others (and have No self)\" Cyril Connoly quote.\n\nBesides, since everything is on the cloud, anyone/anywhere could read it on the future. Even yourself, can compare how do you think and apply those methods and metrics today and what you could deliver in the future.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2533176,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "11/21/2023 16:37:00",
      "content": "<p>Thank you for sharing your write-up with Modified BERT GAN and infer loop.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2533380,
          "author_name": "minhsienweng",
          "author_url": "",
          "post_date": "11/21/2023 20:53:14",
          "content": "<p>Thanks for liking my solution. Please feel free to reach out if you have any questions.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2533445,
              "author_name": "mpwolke",
              "author_url": "",
              "post_date": "11/21/2023 23:57:57",
              "content": "<p>I'm just a beginner, therefore it's hard to make questions since all is very new to me. I hope that other experienced kagglers could read it and learn from your approach. Thank you M. Weng.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2533459,
                  "author_name": "minhsienweng",
                  "author_url": "",
                  "post_date": "11/22/2023 00:32:18",
                  "content": "<p>I am still learning about GNNs, but I am interested to explore their potential.</p>\n<p>Kaggle provides an excellent platform for learning and experimenting with novel ML models. Most importantly,  enjoy the competition and have fun! 😁</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2533473,
                      "author_name": "mpwolke",
                      "author_url": "",
                      "post_date": "11/22/2023 00:44:17",
                      "content": "<p>The more you grow (on Kaggle) cause you're already experienced, other users will start to pay attention on your work. Unfortunately, our community is young and shallow. </p>\n<p>So, they don't read much what contributors/experts are writting.  Worst for them, since they lose a precious opportunity to learn from the best professionals. Besides, those that are working, don't have enough time to spend many hours kaggling : )</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2533502,
                          "author_name": "minhsienweng",
                          "author_url": "",
                          "post_date": "11/22/2023 01:39:18",
                          "content": "<p>I wrote the solution partly for myself. If I encounter similar problems, I can leverage the knowledge and experience I gained from this competition.</p>\n<p>Thank you for your encouragement! 😀</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2533511,
                              "author_name": "mpwolke",
                              "author_url": "",
                              "post_date": "11/22/2023 01:54:22",
                              "content": "<p>\"Better to write for ourselves (and have few audience), than to write for the others (and have No self)\" Cyril Connoly quote.</p>\n<p>Besides, since everything is on the cloud, anyone/anywhere could read it on the future. Even yourself, can compare how do you think and apply those methods and metrics today and what you could deliver in the future.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2532327": "Thank you to Kaggle and Google for hosting this great competition. I learned a lot about training graph data models, especially how to deal with memory efficiency and optimize performance. \n\nI started by training a modified BERT model by @KSMCG90 using cross-validation and the [Lightning AI](https://lightning.ai/) library. But this model was too demanding on memory for the layout dataset. So, I switched to the GNN (Graph Neural Network) model by @GUSTHEMA for the layout task. Since the tile and layout datasets didn't show any significant correlation, I built two separate models: one BERT model for the tile dataset and one GNN model for the layout dataset. \n\nThe solution [notebook](https://www.kaggle.com/code/minhsienweng/bertlike-tile-model-gnn-layout-model) focuses on training two models and run them efficiently on limited computing resources (4-core CPUs).\n\n# Context \n\nBusiness context: https://www.kaggle.com/competitions/predict-ai-model-runtime/\nData context: https://www.kaggle.com/competitions/predict-ai-model-runtime/data\n\n# Overview of the Approach\nThe solution comprised two models: a GNN model for the layout dataset and a modified BERT model for the tile dataset. The inputs were the graph structures generated by an AI model. In this graph, nodes represented tensor operations (e.g., add, max, reshape, convolution) and edges represented tensors.\n\n![Overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16611574%2F4548b69d2c2e5cbc11e502a30fece82c%2FPresentation1.jpg?generation=1700528423839754&alt=media)\n \n## TF-GNN Residual Layout model\nI chose the residual model for layout configurations because it can measure the difference between predicted and actual runtime time of each layout configuration. The model uses TensorFlow's graph neural network library ([TF-GNN](https://github.com/tensorflow/gnn)) to train efficiently on TPUs. I trained the layout model on TPUs and saved the best model as a pretrained model for prediction. In this notebook, I built four separate models, one for each type of layout dataset: `nlp-default`, `nlp-random`, `xla-default`, and `xla-random`.\n\nThe layout model encodes the layout graph into graph embeddings (`node sets`, `edge sets`, and `context`) and includes 5 hidden layers and a global pooling layer. The model is then trained on 10 epochs, with `ListMLE` ranking loss as the loss function and ordered pair accuracy (OPA) as the performance metric. The optimizer is AdamW with a learning rate of $10^{-3}$. \n\n*Things to do*: The model can be improved by adjusting the learning rate with cross validation and by pruning irrelevant `context` to improve memory efficiency. Adding [positional encodings](https://github.com/vijaydwivedi75/gnn-lspe) may also help. [PyG](https://pyg.org/) library could be used to build and train GNN model. x\n\n## Modified BERT GAN Tile model\nSimilar to layout dataset, the tile data is encoded into graph embeddings that include 9 data columns (`node_feat`, `node_opcode`, `edge_index`, `config_runtime`, `config_runtime_normalizers`, `edges_adjecency`, `node_config_feat`, `node_config_ids` and `selected_idxs`).\n\nThe tile model is based on the BERT model and uses Graph Attention Networks ([GANs](https://arxiv.org/abs/1710.10903)) to give more attention to important parts of the data. This helps improve accuracy by focusing on the most relevant information. The tile model uses a GraphEncoder to encode the tile configuration into graph embeddings and applies a special mask to all layers and heads. This mask is based on the connections between different parts of the graph. The model's loss function compares its predictions to different permutations of the actual results.\n\nLastly, the model was wrapped with [Lightning AI](https://lightning.ai/) library to perform cross-validation using five folds of training and validation datasets. The best model from each fold was saved and then used to further train and predict the testing dataset.\n\n# Details of the submission\n\n## Training layout model with checkpoints\nI tried to make the BERT model work with the layout dataset, but it needed too much memory (>300GB). So, I used the TF-GNN-based residual model that uses less memory (<30GB). However, training the TF-GNN model is time-consuming. For instance, training one epoch on the 'xla-random' dataset required approximately one hour on CPUs. Training for 10 epochs would exceed the competition's time constraints (>10 hours). Despite utilizing GPUs to accelerate the training process, the training time has little/no speedups. \n\nTo speed up training, I used a checkpointing technique to save the model's weights after each epoch. This allowed me to resume training, saving time. The best model was saved and used to predict the test datasets.\n\n## Parallelising the infer loop of tile model\nThe infer process for the tile dataset is time-consuming, for example, the model requires approximately 1.81 hours (6,532 seconds) to predict the testing dataset of 844 tile configurations. This is because the infer loop involves running each fold model separately and then merging the results from the five models to get the final output. I tried to speed up the infer loop by using five separate threads to run the inference for each fold's model concurrently. Using a thread pool of three threads, the total infer time is reduced to 58 minutes (3,448 seconds), achieving approximately a 2x speedup on CPUs.\n\nThis experiment demonstrates that parallelism techniques can be effectively applied to machine learning models, enabling them to run on standard hardware configurations. This eliminates the need for specialized hardware, making machine learning more accessible to students and researchers with limited computing resources.\n\n\n# Sources\nThe final solution was inspired by the starter notebook by @GUSTHEMA and BERT-like GAT Tile wLayout Dataset by @KSMCG90\n- https://www.kaggle.com/code/gusthema/starter-notebook-fast-or-slow-with-tensorflow-gnn\n- https://www.kaggle.com/code/ksmcg90/bertlike-gat-tile-wlayout-dataset?scriptVersionId=148717592",
    "2533176": "Thank you for sharing your write-up with Modified BERT GAN and infer loop.",
    "2533380": "Thanks for liking my solution. Please feel free to reach out if you have any questions.",
    "2533445": "I'm just a beginner, therefore it's hard to make questions since all is very new to me. I hope that other experienced kagglers could read it and learn from your approach. Thank you M. Weng.",
    "2533459": "I am still learning about GNNs, but I am interested to explore their potential.\n\nKaggle provides an excellent platform for learning and experimenting with novel ML models. Most importantly,  enjoy the competition and have fun! 😁",
    "2533473": "The more you grow (on Kaggle) cause you're already experienced, other users will start to pay attention on your work. Unfortunately, our community is young and shallow. \n\nSo, they don't read much what contributors/experts are writting.  Worst for them, since they lose a precious opportunity to learn from the best professionals. Besides, those that are working, don't have enough time to spend many hours kaggling : )",
    "2533502": "I wrote the solution partly for myself. If I encounter similar problems, I can leverage the knowledge and experience I gained from this competition.\n\nThank you for your encouragement! 😀",
    "2533511": "\"Better to write for ourselves (and have few audience), than to write for the others (and have No self)\" Cyril Connoly quote.\n\nBesides, since everything is on the cloud, anyone/anywhere could read it on the future. Even yourself, can compare how do you think and apply those methods and metrics today and what you could deliver in the future."
  },
  "source": "meta"
}