{
  "id": 457729,
  "title": "Discoveries through the Contest",
  "url": "/competitions/predict-ai-model-runtime/discussion/457729",
  "author_name": "Suma Mallapragada",
  "post_date": "2023-11-26T13:57:15.221000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to the organizers for holding this unique contest! I am currently an undergraduate student in CS and after taking the course on Compiler Design, this contest helped me explore the machine learning driven approaches for compiler optimizations on field to check how AI models run.</p>\n<h1><strong>Background</strong></h1>\n<p>Reference to the paper <a href=\"https://arxiv.org/pdf/2308.13490.pdf\" target=\"_blank\">\"TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs\"</a> gives an overview on the computational graph representation of the programs running on TPUs with a compilation configuration.</p>\n<h1><strong>Approach</strong></h1>\n<p>The detailed description of the node features, opcode for tiles and layout helps deciding the normalization parameters by noting the estimated time taken for the portion of the program to run given their flow representations. The use of <strong>Graph Convolution Network(GCNConv)</strong> to obtain the most probable runtime order is inspired from its ability to capture local and global information with Parameter Sharing and Transferability.</p>\n<h1><strong>Data Preparation</strong></h1>\n<p>Weights are assigned in the range [0.0055-0.01] to each feature vector based on the code against a particular instruction. Runtime per node is obtained as a weighted summation of the feature vectors of each node along with any specified connections through the edge values. The normalized config runtime is calculated as a difference of the found <em>config runtime</em> and the <em>config feature vectors</em>, and finally dividing it by the runtime obtained by summing the runtime of all the nodes. The classes to be predicted by the GCN would be the normalized config runtime values obtained as the order of fastest to slowest arranged from 0,n-1. The GCN is trained from the config features as the input, with edge vectors as the connectives. </p>\n<h1><strong>Interpretation</strong></h1>\n<p>The probabilities of the configuration belonging to a particular class of runtime(0 for the fastest,1,2,etc. specifying the order and n-1 for the slowest) are the outputs obtained as an n*n matrix, where n is the number of configurations. The highest probability obtained for row 'i' at a particular matrix entry pred[i][j] specifies configuration 'i' belongs to class 'j' or the runtime order is the 'j'th fastest. This trained model helps obtaining the predicted runtime classes given nodes of relatable configurations in the test dataset. The nodes bearing the values of the top 5 classes are returned for the tiles dataset. This is similar to the bag-of-words represented as feature vectors used to determine the class of Machine Learning Keywords in the cora dataset.</p>\n<h1><strong>Key Takeaways</strong></h1>\n<ol>\n<li>Reviewing the solution by <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456343\" target=\"_blank\">EDUARDO ROCHA DE ANDRADE</a> helped me understand the need for the use of Cross-Config Attention for making this approach more efficient.</li>\n<li>The approach also helped obtain the runtime most configurations follow based on the most prominent classes obtained.</li>\n</ol>",
  "messages": [
    {
      "id": 2538851,
      "postDate": "2023-11-26T13:57:15.220Z",
      "content": "<p>Thanks to the organizers for holding this unique contest! I am currently an undergraduate student in CS and after taking the course on Compiler Design, this contest helped me explore the machine learning driven approaches for compiler optimizations on field to check how AI models run.</p>\n<h1><strong>Background</strong></h1>\n<p>Reference to the paper <a href=\"https://arxiv.org/pdf/2308.13490.pdf\" target=\"_blank\">\"TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs\"</a> gives an overview on the computational graph representation of the programs running on TPUs with a compilation configuration.</p>\n<h1><strong>Approach</strong></h1>\n<p>The detailed description of the node features, opcode for tiles and layout helps deciding the normalization parameters by noting the estimated time taken for the portion of the program to run given their flow representations. The use of <strong>Graph Convolution Network(GCNConv)</strong> to obtain the most probable runtime order is inspired from its ability to capture local and global information with Parameter Sharing and Transferability.</p>\n<h1><strong>Data Preparation</strong></h1>\n<p>Weights are assigned in the range [0.0055-0.01] to each feature vector based on the code against a particular instruction. Runtime per node is obtained as a weighted summation of the feature vectors of each node along with any specified connections through the edge values. The normalized config runtime is calculated as a difference of the found <em>config runtime</em> and the <em>config feature vectors</em>, and finally dividing it by the runtime obtained by summing the runtime of all the nodes. The classes to be predicted by the GCN would be the normalized config runtime values obtained as the order of fastest to slowest arranged from 0,n-1. The GCN is trained from the config features as the input, with edge vectors as the connectives. </p>\n<h1><strong>Interpretation</strong></h1>\n<p>The probabilities of the configuration belonging to a particular class of runtime(0 for the fastest,1,2,etc. specifying the order and n-1 for the slowest) are the outputs obtained as an n*n matrix, where n is the number of configurations. The highest probability obtained for row 'i' at a particular matrix entry pred[i][j] specifies configuration 'i' belongs to class 'j' or the runtime order is the 'j'th fastest. This trained model helps obtaining the predicted runtime classes given nodes of relatable configurations in the test dataset. The nodes bearing the values of the top 5 classes are returned for the tiles dataset. This is similar to the bag-of-words represented as feature vectors used to determine the class of Machine Learning Keywords in the cora dataset.</p>\n<h1><strong>Key Takeaways</strong></h1>\n<ol>\n<li>Reviewing the solution by <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456343\" target=\"_blank\">EDUARDO ROCHA DE ANDRADE</a> helped me understand the need for the use of Cross-Config Attention for making this approach more efficient.</li>\n<li>The approach also helped obtain the runtime most configurations follow based on the most prominent classes obtained.</li>\n</ol>",
      "rawMarkdown": "Thanks to the organizers for holding this unique contest! I am currently an undergraduate student in CS and after taking the course on Compiler Design, this contest helped me explore the machine learning driven approaches for compiler optimizations on field to check how AI models run.\n\n# **Background**\nReference to the paper [\"TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs\"](https://arxiv.org/pdf/2308.13490.pdf) gives an overview on the computational graph representation of the programs running on TPUs with a compilation configuration.\n\n# **Approach**\nThe detailed description of the node features, opcode for tiles and layout helps deciding the normalization parameters by noting the estimated time taken for the portion of the program to run given their flow representations. The use of **Graph Convolution Network(GCNConv)** to obtain the most probable runtime order is inspired from its ability to capture local and global information with Parameter Sharing and Transferability.\n\n# **Data Preparation**\nWeights are assigned in the range [0.0055-0.01] to each feature vector based on the code against a particular instruction. Runtime per node is obtained as a weighted summation of the feature vectors of each node along with any specified connections through the edge values. The normalized config runtime is calculated as a difference of the found *config runtime* and the *config feature vectors*, and finally dividing it by the runtime obtained by summing the runtime of all the nodes. The classes to be predicted by the GCN would be the normalized config runtime values obtained as the order of fastest to slowest arranged from 0,n-1. The GCN is trained from the config features as the input, with edge vectors as the connectives. \n\n# **Interpretation**\nThe probabilities of the configuration belonging to a particular class of runtime(0 for the fastest,1,2,etc. specifying the order and n-1 for the slowest) are the outputs obtained as an n*n matrix, where n is the number of configurations. The highest probability obtained for row 'i' at a particular matrix entry pred[i][j] specifies configuration 'i' belongs to class 'j' or the runtime order is the 'j'th fastest. This trained model helps obtaining the predicted runtime classes given nodes of relatable configurations in the test dataset. The nodes bearing the values of the top 5 classes are returned for the tiles dataset. This is similar to the bag-of-words represented as feature vectors used to determine the class of Machine Learning Keywords in the cora dataset.\n\n# **Key Takeaways**\n1. Reviewing the solution by [EDUARDO ROCHA DE ANDRADE](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456343) helped me understand the need for the use of Cross-Config Attention for making this approach more efficient.\n2. The approach also helped obtain the runtime most configurations follow based on the most prominent classes obtained."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2538851": "Thanks to the organizers for holding this unique contest! I am currently an undergraduate student in CS and after taking the course on Compiler Design, this contest helped me explore the machine learning driven approaches for compiler optimizations on field to check how AI models run.\n\n# **Background**\nReference to the paper [\"TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs\"](https://arxiv.org/pdf/2308.13490.pdf) gives an overview on the computational graph representation of the programs running on TPUs with a compilation configuration.\n\n# **Approach**\nThe detailed description of the node features, opcode for tiles and layout helps deciding the normalization parameters by noting the estimated time taken for the portion of the program to run given their flow representations. The use of **Graph Convolution Network(GCNConv)** to obtain the most probable runtime order is inspired from its ability to capture local and global information with Parameter Sharing and Transferability.\n\n# **Data Preparation**\nWeights are assigned in the range [0.0055-0.01] to each feature vector based on the code against a particular instruction. Runtime per node is obtained as a weighted summation of the feature vectors of each node along with any specified connections through the edge values. The normalized config runtime is calculated as a difference of the found *config runtime* and the *config feature vectors*, and finally dividing it by the runtime obtained by summing the runtime of all the nodes. The classes to be predicted by the GCN would be the normalized config runtime values obtained as the order of fastest to slowest arranged from 0,n-1. The GCN is trained from the config features as the input, with edge vectors as the connectives. \n\n# **Interpretation**\nThe probabilities of the configuration belonging to a particular class of runtime(0 for the fastest,1,2,etc. specifying the order and n-1 for the slowest) are the outputs obtained as an n*n matrix, where n is the number of configurations. The highest probability obtained for row 'i' at a particular matrix entry pred[i][j] specifies configuration 'i' belongs to class 'j' or the runtime order is the 'j'th fastest. This trained model helps obtaining the predicted runtime classes given nodes of relatable configurations in the test dataset. The nodes bearing the values of the top 5 classes are returned for the tiles dataset. This is similar to the bag-of-words represented as feature vectors used to determine the class of Machine Learning Keywords in the cora dataset.\n\n# **Key Takeaways**\n1. Reviewing the solution by [EDUARDO ROCHA DE ANDRADE](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456343) helped me understand the need for the use of Cross-Config Attention for making this approach more efficient.\n2. The approach also helped obtain the runtime most configurations follow based on the most prominent classes obtained."
  }
}