{
  "id": 458370,
  "title": "13th Place Solution for the Google - Fast or Slow? Predict AI Model Runtime Competition",
  "url": "/competitions/predict-ai-model-runtime/discussion/458370",
  "author_name": "Ignacio Reyes",
  "post_date": "2023-11-29T14:38:24.994000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>13th Place Solution for the Google - Fast or Slow? Predict AI Model Runtime Competition</h1>\n<p>This is my solution for the \"Fast or Slow? Predict AI Model Runtime\" Competition. I hope you find it useful!</p>\n<p>The key principles of my approach were to start with something simple and improve it in many iterations, and to work under the hardware constraints that I had (16 GB RAM, 4 GB VRAM).</p>\n<h2>Context section</h2>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/overview</a></li>\n<li>Data Context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<h3>Layout model</h3>\n<p>The core idea is to extract features from the graph and its nodes, and train a Multi-layer Perceptron with that information. For each configurable node of the computational graph we took some node properties, the layout used, and properties from its \"parents\" and \"siblings\". The \"parent\" nodes are the ones that produce the node inputs. Two nodes are \"siblings\" if they share a parent. For example, in the following figure, nodes 1 and 2 are the parents and nodes 3, 5, and 6 are the siblings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2528048%2F665c76b0456511e708fcdc2859c2ab98%2Fimg_kaggle.jpeg?generation=1701268499581469&amp;alt=media\" alt=\"\"></p>\n<p>For each configurable node in the graph we had a list of features that were processed by 3 fully-connected layers with dropout layers in between. After that, the \"node\" dimension was averaged, so all the node information is represented as a vector. Two extra inputs were concatenated to this vector: a \"graph description\" and the \"subset information\". The \"graph description\" is a vector with how many nodes of each type are present in the graph (normalized to sum 1), with an extra value that is the number of nodes of the graph (with a logarithm) to give the model an idea of the graph's size. The \"subset information\" is a vector that signals if the graph comes from the \"xla:default\", \"xla:random\", \"nlp:default\" or \"nlp:random\" subset, for what we used an \"Embedding\" layer of keras.</p>\n<p>This new vector was processed by 3 additional fully-connected layers (no dropout this time). The Pairwise Hinge loss was used as objective function during training. In order to do that, each training batch of size 128 had examples of 16 different graphs, with 8 configuration examples each. Each batch is built by randomly choosing a subset with equal probability, so the batches are subset-balanced on average. The final submission was an ensemble of 3 independent training runs. </p>\n<h3>Tile model</h3>\n<p>The tile model is a simplified version of the layout model. Each configuration is described with a vector that has the \"config_feat\" information and also a \"graph descriptor\" (the same one from the layout model). This model is a Multi-layer Perceptron with 3 fully-connected layers and one dropout layer after the first layer. The loss function and batch structure is the same as the layout model, but with a larger batch size (600) and a larger number of configurations per graph (20).</p>\n<h3>Validation</h3>\n<p>The \"valid\" folder examples from the dataset were used as the validation set. We performed validation every 10000 training iterations, and we computed the competition metric over that set. In the case of the layout model, the metric was computed for each one of the four subsets and then an average is computed. If the validation metric did not improve after 5 validations, the training is stopped.</p>\n<h2>Details of the submission</h2>\n<h3>Tile model details</h3>\n<p>The \"tile problem\" was relatively easy in comparison to the layout problem, so we did not spend that much effort improving this model. The number of configurations per graph was limited to 160, and we sampled them giving the configurations with a lower runtime a higher probability (using an exponential distribution), as the challenge here was about finding the fastest configuration, not sorting all the configurations.</p>\n<h3>Node features (Layout model)</h3>\n<p>From all the available features in the \"node_feat\" matrices, we selected the ones that we thought were the most important. This helps keeping the required memory low. These features are:</p>\n<ul>\n<li>shape_dimensions (21-26).</li>\n<li>reshape/broadcast dimensions (31-36).</li>\n<li>convolution_dim_numbers_input_spatial_dims (95-98).</li>\n<li>convolution_dim_numbers_kernel_spatial_dims (101-104).</li>\n<li>layout_minor_to_major (134-139)</li>\n</ul>\n<p>Values in parentheses correspond to the selected indices of \"node_feat\". We also gave the node layout information and the opcode to the model (encoded as a vector with a keras Embedding layer). Another important thing is that a re-ordered version of the shapes was given to the network according to the layout information (in addition to the original version). </p>\n<p>For the siblings we took each sibling output shape, its layout and a boolean that compares the node layout with the sibling layout to check if they are the same. As the competition overview mentioned, if the layouts of two siblings are different, an extra copy operation is needed, so that motivated the creation of this variable.</p>\n<p>For the parents we keep their output shapes and physical layouts. We also took the opcodes from the parents and the siblings, and express them as a vector with the help of the Embedding layer of keras.</p>\n<h3>Keep the training stable (Layout model)</h3>\n<p>To facilitate the training process, all the features that took values across many orders of magnitude (e.g. tensor shapes) were passed to a logarithm to avoid very large input values. The features were scaled using a mean / standard deviation normalization, with some clipping in the std estimation to avoid dividing by a very small value.</p>\n<p>A cosine decay schedule was used for the learning rate. After a 10000 iteration linear warm-up in the learning rate, the cosine decay reduced the parameter across 250k iterations until it reached a 5 % of the original value. Adam was used as optimizer, with a clipnorm value of 1.0 to avoid large weight updates.</p>\n<p>Many configurations had the same layout. All of them were replaced by just one instance of that layout, and the runtime replaced by the mean of the runtimes.</p>\n<h3>Things that didn't work that well…</h3>\n<p>We had some instability problems with the List MLE loss, so we chose using the Pairwise hinge loss instead. The problem was that after many iterations, suddenly a NaN value appeared in the model (or loss, idk) and destroyed all the model weights.</p>\n<h3>Keeping the training under the memory budget</h3>\n<p>As we mentioned, we worked with a limited memory budget of 16 GB of RAM and 4 GB of VRAM, so we had to be very careful of not loading too much data at the same time and be conservative with the model size. The first important thing was to process all the npz files and save the necessary information in the tfrecords format of tensorflow (with file compression activated). During training, these files were read from disk, trying to give the model samples from many different graphs instead of seeing just one graph at a time.</p>\n<p>The number of configuration per graph was capped at 7500, and the number of configurable nodes given to the network was capped at 1000. As many different graphs were given to the network and considering that each graph has a different number of configurable nodes, it is necessary to pad and mask tensors to make all the samples the same length across the \"node\" dimension. This can have a heavy memory burden if the number of used nodes is increased too much, so the number 1000 was chosen given this restriction. I also put a limit in the number of parents (2) and siblings (3) for each node.</p>\n<h2>Sources</h2>\n<p>Code: <a href=\"https://github.com/ignacioreyes/kaggle_model_runtime\" target=\"_blank\">https://github.com/ignacioreyes/kaggle_model_runtime</a></p>\n<p>Embedding layer (tf/keras): <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/layers/Embedding\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/layers/Embedding</a></p>\n<h4>Note: I wrote many sections in plural, as is customary in academic papers.</h4>",
  "messages": [
    {
      "id": 2542814,
      "postDate": "2023-11-29T14:38:24.993Z",
      "content": "<h1>13th Place Solution for the Google - Fast or Slow? Predict AI Model Runtime Competition</h1>\n<p>This is my solution for the \"Fast or Slow? Predict AI Model Runtime\" Competition. I hope you find it useful!</p>\n<p>The key principles of my approach were to start with something simple and improve it in many iterations, and to work under the hardware constraints that I had (16 GB RAM, 4 GB VRAM).</p>\n<h2>Context section</h2>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/overview</a></li>\n<li>Data Context: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<h3>Layout model</h3>\n<p>The core idea is to extract features from the graph and its nodes, and train a Multi-layer Perceptron with that information. For each configurable node of the computational graph we took some node properties, the layout used, and properties from its \"parents\" and \"siblings\". The \"parent\" nodes are the ones that produce the node inputs. Two nodes are \"siblings\" if they share a parent. For example, in the following figure, nodes 1 and 2 are the parents and nodes 3, 5, and 6 are the siblings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2528048%2F665c76b0456511e708fcdc2859c2ab98%2Fimg_kaggle.jpeg?generation=1701268499581469&amp;alt=media\" alt=\"\"></p>\n<p>For each configurable node in the graph we had a list of features that were processed by 3 fully-connected layers with dropout layers in between. After that, the \"node\" dimension was averaged, so all the node information is represented as a vector. Two extra inputs were concatenated to this vector: a \"graph description\" and the \"subset information\". The \"graph description\" is a vector with how many nodes of each type are present in the graph (normalized to sum 1), with an extra value that is the number of nodes of the graph (with a logarithm) to give the model an idea of the graph's size. The \"subset information\" is a vector that signals if the graph comes from the \"xla:default\", \"xla:random\", \"nlp:default\" or \"nlp:random\" subset, for what we used an \"Embedding\" layer of keras.</p>\n<p>This new vector was processed by 3 additional fully-connected layers (no dropout this time). The Pairwise Hinge loss was used as objective function during training. In order to do that, each training batch of size 128 had examples of 16 different graphs, with 8 configuration examples each. Each batch is built by randomly choosing a subset with equal probability, so the batches are subset-balanced on average. The final submission was an ensemble of 3 independent training runs. </p>\n<h3>Tile model</h3>\n<p>The tile model is a simplified version of the layout model. Each configuration is described with a vector that has the \"config_feat\" information and also a \"graph descriptor\" (the same one from the layout model). This model is a Multi-layer Perceptron with 3 fully-connected layers and one dropout layer after the first layer. The loss function and batch structure is the same as the layout model, but with a larger batch size (600) and a larger number of configurations per graph (20).</p>\n<h3>Validation</h3>\n<p>The \"valid\" folder examples from the dataset were used as the validation set. We performed validation every 10000 training iterations, and we computed the competition metric over that set. In the case of the layout model, the metric was computed for each one of the four subsets and then an average is computed. If the validation metric did not improve after 5 validations, the training is stopped.</p>\n<h2>Details of the submission</h2>\n<h3>Tile model details</h3>\n<p>The \"tile problem\" was relatively easy in comparison to the layout problem, so we did not spend that much effort improving this model. The number of configurations per graph was limited to 160, and we sampled them giving the configurations with a lower runtime a higher probability (using an exponential distribution), as the challenge here was about finding the fastest configuration, not sorting all the configurations.</p>\n<h3>Node features (Layout model)</h3>\n<p>From all the available features in the \"node_feat\" matrices, we selected the ones that we thought were the most important. This helps keeping the required memory low. These features are:</p>\n<ul>\n<li>shape_dimensions (21-26).</li>\n<li>reshape/broadcast dimensions (31-36).</li>\n<li>convolution_dim_numbers_input_spatial_dims (95-98).</li>\n<li>convolution_dim_numbers_kernel_spatial_dims (101-104).</li>\n<li>layout_minor_to_major (134-139)</li>\n</ul>\n<p>Values in parentheses correspond to the selected indices of \"node_feat\". We also gave the node layout information and the opcode to the model (encoded as a vector with a keras Embedding layer). Another important thing is that a re-ordered version of the shapes was given to the network according to the layout information (in addition to the original version). </p>\n<p>For the siblings we took each sibling output shape, its layout and a boolean that compares the node layout with the sibling layout to check if they are the same. As the competition overview mentioned, if the layouts of two siblings are different, an extra copy operation is needed, so that motivated the creation of this variable.</p>\n<p>For the parents we keep their output shapes and physical layouts. We also took the opcodes from the parents and the siblings, and express them as a vector with the help of the Embedding layer of keras.</p>\n<h3>Keep the training stable (Layout model)</h3>\n<p>To facilitate the training process, all the features that took values across many orders of magnitude (e.g. tensor shapes) were passed to a logarithm to avoid very large input values. The features were scaled using a mean / standard deviation normalization, with some clipping in the std estimation to avoid dividing by a very small value.</p>\n<p>A cosine decay schedule was used for the learning rate. After a 10000 iteration linear warm-up in the learning rate, the cosine decay reduced the parameter across 250k iterations until it reached a 5 % of the original value. Adam was used as optimizer, with a clipnorm value of 1.0 to avoid large weight updates.</p>\n<p>Many configurations had the same layout. All of them were replaced by just one instance of that layout, and the runtime replaced by the mean of the runtimes.</p>\n<h3>Things that didn't work that well…</h3>\n<p>We had some instability problems with the List MLE loss, so we chose using the Pairwise hinge loss instead. The problem was that after many iterations, suddenly a NaN value appeared in the model (or loss, idk) and destroyed all the model weights.</p>\n<h3>Keeping the training under the memory budget</h3>\n<p>As we mentioned, we worked with a limited memory budget of 16 GB of RAM and 4 GB of VRAM, so we had to be very careful of not loading too much data at the same time and be conservative with the model size. The first important thing was to process all the npz files and save the necessary information in the tfrecords format of tensorflow (with file compression activated). During training, these files were read from disk, trying to give the model samples from many different graphs instead of seeing just one graph at a time.</p>\n<p>The number of configuration per graph was capped at 7500, and the number of configurable nodes given to the network was capped at 1000. As many different graphs were given to the network and considering that each graph has a different number of configurable nodes, it is necessary to pad and mask tensors to make all the samples the same length across the \"node\" dimension. This can have a heavy memory burden if the number of used nodes is increased too much, so the number 1000 was chosen given this restriction. I also put a limit in the number of parents (2) and siblings (3) for each node.</p>\n<h2>Sources</h2>\n<p>Code: <a href=\"https://github.com/ignacioreyes/kaggle_model_runtime\" target=\"_blank\">https://github.com/ignacioreyes/kaggle_model_runtime</a></p>\n<p>Embedding layer (tf/keras): <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/layers/Embedding\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/layers/Embedding</a></p>\n<h4>Note: I wrote many sections in plural, as is customary in academic papers.</h4>",
      "rawMarkdown": "# 13th Place Solution for the Google - Fast or Slow? Predict AI Model Runtime Competition\n\nThis is my solution for the \"Fast or Slow? Predict AI Model Runtime\" Competition. I hope you find it useful!\n\nThe key principles of my approach were to start with something simple and improve it in many iterations, and to work under the hardware constraints that I had (16 GB RAM, 4 GB VRAM).\n\n\n## Context section\n\n* Business context: https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\n* Data Context: https://www.kaggle.com/competitions/predict-ai-model-runtime/data\n\n## Overview of the approach\n\n### Layout model\n\nThe core idea is to extract features from the graph and its nodes, and train a Multi-layer Perceptron with that information. For each configurable node of the computational graph we took some node properties, the layout used, and properties from its \"parents\" and \"siblings\". The \"parent\" nodes are the ones that produce the node inputs. Two nodes are \"siblings\" if they share a parent. For example, in the following figure, nodes 1 and 2 are the parents and nodes 3, 5, and 6 are the siblings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2528048%2F665c76b0456511e708fcdc2859c2ab98%2Fimg_kaggle.jpeg?generation=1701268499581469&alt=media)\n\nFor each configurable node in the graph we had a list of features that were processed by 3 fully-connected layers with dropout layers in between. After that, the \"node\" dimension was averaged, so all the node information is represented as a vector. Two extra inputs were concatenated to this vector: a \"graph description\" and the \"subset information\". The \"graph description\" is a vector with how many nodes of each type are present in the graph (normalized to sum 1), with an extra value that is the number of nodes of the graph (with a logarithm) to give the model an idea of the graph's size. The \"subset information\" is a vector that signals if the graph comes from the \"xla:default\", \"xla:random\", \"nlp:default\" or \"nlp:random\" subset, for what we used an \"Embedding\" layer of keras.\n\nThis new vector was processed by 3 additional fully-connected layers (no dropout this time). The Pairwise Hinge loss was used as objective function during training. In order to do that, each training batch of size 128 had examples of 16 different graphs, with 8 configuration examples each. Each batch is built by randomly choosing a subset with equal probability, so the batches are subset-balanced on average. The final submission was an ensemble of 3 independent training runs. \n\n### Tile model\n\nThe tile model is a simplified version of the layout model. Each configuration is described with a vector that has the \"config_feat\" information and also a \"graph descriptor\" (the same one from the layout model). This model is a Multi-layer Perceptron with 3 fully-connected layers and one dropout layer after the first layer. The loss function and batch structure is the same as the layout model, but with a larger batch size (600) and a larger number of configurations per graph (20).\n\n### Validation\n\nThe \"valid\" folder examples from the dataset were used as the validation set. We performed validation every 10000 training iterations, and we computed the competition metric over that set. In the case of the layout model, the metric was computed for each one of the four subsets and then an average is computed. If the validation metric did not improve after 5 validations, the training is stopped.\n\n## Details of the submission\n\n### Tile model details\n\nThe \"tile problem\" was relatively easy in comparison to the layout problem, so we did not spend that much effort improving this model. The number of configurations per graph was limited to 160, and we sampled them giving the configurations with a lower runtime a higher probability (using an exponential distribution), as the challenge here was about finding the fastest configuration, not sorting all the configurations.\n\n### Node features (Layout model)\n\nFrom all the available features in the \"node_feat\" matrices, we selected the ones that we thought were the most important. This helps keeping the required memory low. These features are:\n\n* shape_dimensions (21-26).\n* reshape/broadcast dimensions (31-36).\n* convolution_dim_numbers_input_spatial_dims (95-98).\n* convolution_dim_numbers_kernel_spatial_dims (101-104).\n* layout_minor_to_major (134-139)\n\nValues in parentheses correspond to the selected indices of \"node_feat\". We also gave the node layout information and the opcode to the model (encoded as a vector with a keras Embedding layer). Another important thing is that a re-ordered version of the shapes was given to the network according to the layout information (in addition to the original version). \n\nFor the siblings we took each sibling output shape, its layout and a boolean that compares the node layout with the sibling layout to check if they are the same. As the competition overview mentioned, if the layouts of two siblings are different, an extra copy operation is needed, so that motivated the creation of this variable.\n\nFor the parents we keep their output shapes and physical layouts. We also took the opcodes from the parents and the siblings, and express them as a vector with the help of the Embedding layer of keras.\n\n### Keep the training stable (Layout model)\n\nTo facilitate the training process, all the features that took values across many orders of magnitude (e.g. tensor shapes) were passed to a logarithm to avoid very large input values. The features were scaled using a mean / standard deviation normalization, with some clipping in the std estimation to avoid dividing by a very small value.\n\nA cosine decay schedule was used for the learning rate. After a 10000 iteration linear warm-up in the learning rate, the cosine decay reduced the parameter across 250k iterations until it reached a 5 % of the original value. Adam was used as optimizer, with a clipnorm value of 1.0 to avoid large weight updates.\n\nMany configurations had the same layout. All of them were replaced by just one instance of that layout, and the runtime replaced by the mean of the runtimes.\n\n### Things that didn't work that well...\n\nWe had some instability problems with the List MLE loss, so we chose using the Pairwise hinge loss instead. The problem was that after many iterations, suddenly a NaN value appeared in the model (or loss, idk) and destroyed all the model weights.\n\n### Keeping the training under the memory budget\n\nAs we mentioned, we worked with a limited memory budget of 16 GB of RAM and 4 GB of VRAM, so we had to be very careful of not loading too much data at the same time and be conservative with the model size. The first important thing was to process all the npz files and save the necessary information in the tfrecords format of tensorflow (with file compression activated). During training, these files were read from disk, trying to give the model samples from many different graphs instead of seeing just one graph at a time.\n\nThe number of configuration per graph was capped at 7500, and the number of configurable nodes given to the network was capped at 1000. As many different graphs were given to the network and considering that each graph has a different number of configurable nodes, it is necessary to pad and mask tensors to make all the samples the same length across the \"node\" dimension. This can have a heavy memory burden if the number of used nodes is increased too much, so the number 1000 was chosen given this restriction. I also put a limit in the number of parents (2) and siblings (3) for each node.\n\n## Sources\n\nCode: https://github.com/ignacioreyes/kaggle_model_runtime\n\nEmbedding layer (tf/keras): https://www.tensorflow.org/api_docs/python/tf/keras/layers/Embedding\n\n#### Note: I wrote many sections in plural, as is customary in academic papers."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2542814": "# 13th Place Solution for the Google - Fast or Slow? Predict AI Model Runtime Competition\n\nThis is my solution for the \"Fast or Slow? Predict AI Model Runtime\" Competition. I hope you find it useful!\n\nThe key principles of my approach were to start with something simple and improve it in many iterations, and to work under the hardware constraints that I had (16 GB RAM, 4 GB VRAM).\n\n\n## Context section\n\n* Business context: https://www.kaggle.com/competitions/predict-ai-model-runtime/overview\n* Data Context: https://www.kaggle.com/competitions/predict-ai-model-runtime/data\n\n## Overview of the approach\n\n### Layout model\n\nThe core idea is to extract features from the graph and its nodes, and train a Multi-layer Perceptron with that information. For each configurable node of the computational graph we took some node properties, the layout used, and properties from its \"parents\" and \"siblings\". The \"parent\" nodes are the ones that produce the node inputs. Two nodes are \"siblings\" if they share a parent. For example, in the following figure, nodes 1 and 2 are the parents and nodes 3, 5, and 6 are the siblings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2528048%2F665c76b0456511e708fcdc2859c2ab98%2Fimg_kaggle.jpeg?generation=1701268499581469&alt=media)\n\nFor each configurable node in the graph we had a list of features that were processed by 3 fully-connected layers with dropout layers in between. After that, the \"node\" dimension was averaged, so all the node information is represented as a vector. Two extra inputs were concatenated to this vector: a \"graph description\" and the \"subset information\". The \"graph description\" is a vector with how many nodes of each type are present in the graph (normalized to sum 1), with an extra value that is the number of nodes of the graph (with a logarithm) to give the model an idea of the graph's size. The \"subset information\" is a vector that signals if the graph comes from the \"xla:default\", \"xla:random\", \"nlp:default\" or \"nlp:random\" subset, for what we used an \"Embedding\" layer of keras.\n\nThis new vector was processed by 3 additional fully-connected layers (no dropout this time). The Pairwise Hinge loss was used as objective function during training. In order to do that, each training batch of size 128 had examples of 16 different graphs, with 8 configuration examples each. Each batch is built by randomly choosing a subset with equal probability, so the batches are subset-balanced on average. The final submission was an ensemble of 3 independent training runs. \n\n### Tile model\n\nThe tile model is a simplified version of the layout model. Each configuration is described with a vector that has the \"config_feat\" information and also a \"graph descriptor\" (the same one from the layout model). This model is a Multi-layer Perceptron with 3 fully-connected layers and one dropout layer after the first layer. The loss function and batch structure is the same as the layout model, but with a larger batch size (600) and a larger number of configurations per graph (20).\n\n### Validation\n\nThe \"valid\" folder examples from the dataset were used as the validation set. We performed validation every 10000 training iterations, and we computed the competition metric over that set. In the case of the layout model, the metric was computed for each one of the four subsets and then an average is computed. If the validation metric did not improve after 5 validations, the training is stopped.\n\n## Details of the submission\n\n### Tile model details\n\nThe \"tile problem\" was relatively easy in comparison to the layout problem, so we did not spend that much effort improving this model. The number of configurations per graph was limited to 160, and we sampled them giving the configurations with a lower runtime a higher probability (using an exponential distribution), as the challenge here was about finding the fastest configuration, not sorting all the configurations.\n\n### Node features (Layout model)\n\nFrom all the available features in the \"node_feat\" matrices, we selected the ones that we thought were the most important. This helps keeping the required memory low. These features are:\n\n* shape_dimensions (21-26).\n* reshape/broadcast dimensions (31-36).\n* convolution_dim_numbers_input_spatial_dims (95-98).\n* convolution_dim_numbers_kernel_spatial_dims (101-104).\n* layout_minor_to_major (134-139)\n\nValues in parentheses correspond to the selected indices of \"node_feat\". We also gave the node layout information and the opcode to the model (encoded as a vector with a keras Embedding layer). Another important thing is that a re-ordered version of the shapes was given to the network according to the layout information (in addition to the original version). \n\nFor the siblings we took each sibling output shape, its layout and a boolean that compares the node layout with the sibling layout to check if they are the same. As the competition overview mentioned, if the layouts of two siblings are different, an extra copy operation is needed, so that motivated the creation of this variable.\n\nFor the parents we keep their output shapes and physical layouts. We also took the opcodes from the parents and the siblings, and express them as a vector with the help of the Embedding layer of keras.\n\n### Keep the training stable (Layout model)\n\nTo facilitate the training process, all the features that took values across many orders of magnitude (e.g. tensor shapes) were passed to a logarithm to avoid very large input values. The features were scaled using a mean / standard deviation normalization, with some clipping in the std estimation to avoid dividing by a very small value.\n\nA cosine decay schedule was used for the learning rate. After a 10000 iteration linear warm-up in the learning rate, the cosine decay reduced the parameter across 250k iterations until it reached a 5 % of the original value. Adam was used as optimizer, with a clipnorm value of 1.0 to avoid large weight updates.\n\nMany configurations had the same layout. All of them were replaced by just one instance of that layout, and the runtime replaced by the mean of the runtimes.\n\n### Things that didn't work that well...\n\nWe had some instability problems with the List MLE loss, so we chose using the Pairwise hinge loss instead. The problem was that after many iterations, suddenly a NaN value appeared in the model (or loss, idk) and destroyed all the model weights.\n\n### Keeping the training under the memory budget\n\nAs we mentioned, we worked with a limited memory budget of 16 GB of RAM and 4 GB of VRAM, so we had to be very careful of not loading too much data at the same time and be conservative with the model size. The first important thing was to process all the npz files and save the necessary information in the tfrecords format of tensorflow (with file compression activated). During training, these files were read from disk, trying to give the model samples from many different graphs instead of seeing just one graph at a time.\n\nThe number of configuration per graph was capped at 7500, and the number of configurable nodes given to the network was capped at 1000. As many different graphs were given to the network and considering that each graph has a different number of configurable nodes, it is necessary to pad and mask tensors to make all the samples the same length across the \"node\" dimension. This can have a heavy memory burden if the number of used nodes is increased too much, so the number 1000 was chosen given this restriction. I also put a limit in the number of parents (2) and siblings (3) for each node.\n\n## Sources\n\nCode: https://github.com/ignacioreyes/kaggle_model_runtime\n\nEmbedding layer (tf/keras): https://www.tensorflow.org/api_docs/python/tf/keras/layers/Embedding\n\n#### Note: I wrote many sections in plural, as is customary in academic papers."
  }
}