{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Starter notebook: Fast or Slow with TensorFlow GNN\n\nThis tutorial is designed to walk competitors of [predict-ai-model-runtime](https://kaggle.com/competitions/predict-ai-model-runtime) through the dataset and using [TensorFlow-GNN](https://github.com/tensorflow/gnn).\n\nIn summary, you will:\n- `pip install` libraries\n- imports helper functions from another project, for reading data (`{layout, tile}_data`) and easier programming of GNN models (`implicit`).\n- read batches of graphs from the dataset, prints them on screen and explains them. \n- go through details for writing a GNN model and train it\n- produce an inference `csv` file on the test set.\n","metadata":{}},{"cell_type":"code","source":"!pip install tensorflow_gnn --pre\n!pip install tensorflow_ranking","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-10-09T04:49:32.060403Z","iopub.execute_input":"2023-10-09T04:49:32.060740Z","iopub.status.idle":"2023-10-09T04:50:15.083287Z","shell.execute_reply.started":"2023-10-09T04:49:32.060692Z","shell.execute_reply":"2023-10-09T04:50:15.082157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Install standard modules\n\nimport os\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nimport tensorflow_gnn as tfgnn\nimport tensorflow_ranking as tfr","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:50:15.085062Z","iopub.execute_input":"2023-10-09T04:50:15.085345Z","iopub.status.idle":"2023-10-09T04:50:24.853345Z","shell.execute_reply.started":"2023-10-09T04:50:15.085322Z","shell.execute_reply":"2023-10-09T04:50:24.852202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The utility modules are based on the code on [github](): ","metadata":{}},{"cell_type":"code","source":"# Install utility modules.\n\nimport tpugraphsv1_layout_data_py as layout_data\nimport tpugraphsv1_tile_data_py as tile_data\nimport tpugraphsv1_implicit_py as implicit","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:50:24.854541Z","iopub.execute_input":"2023-10-09T04:50:24.855099Z","iopub.status.idle":"2023-10-09T04:50:47.107361Z","shell.execute_reply.started":"2023-10-09T04:50:24.855076Z","shell.execute_reply":"2023-10-09T04:50:47.106227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training Pipelines\n\nThe following code is organized as:\n\n 1. Helper functions: MLP (`_mlp`) and Embedding layer (`_Opembedding`). The embedding layer amends a feature on the `op` nodes, with name `op_e`, by embedding the integral op IDs.\n 1. Pipeline code for training on the Layout collections.\n 1. Pipeline code for training on the Tile collection.","metadata":{}},{"cell_type":"markdown","source":"## Helper functions, for both Layout and Tile collections.","metadata":{}},{"cell_type":"code","source":"\ndef _mlp(dims, hidden_activation, l2reg=1e-4, use_bias=True):\n  \"\"\"Helper function for multi-layer perceptron (MLP).\"\"\"\n  layers = []\n  for i, dim in enumerate(dims):\n    if i > 0:\n      layers.append(tf.keras.layers.Activation(hidden_activation))\n    layers.append(tf.keras.layers.Dense(\n        dim, kernel_regularizer=tf.keras.regularizers.l2(l2reg),\n        use_bias=use_bias))\n  return tf.keras.Sequential(layers)\n\n\nclass _OpEmbedding(tf.keras.Model):\n  \"\"\"Embeds GraphTensor.node_sets['op']['op'] nodes into feature 'op_e'.\"\"\"\n\n  def __init__(self, num_ops: int, embed_d: int, l2reg: float = 1e-4):\n    super().__init__()\n    self.embedding_layer = tf.keras.layers.Embedding(\n        num_ops, embed_d, activity_regularizer=tf.keras.regularizers.l2(l2reg))\n\n  def call(\n      self, graph: tfgnn.GraphTensor,\n      training: bool = False) -> tfgnn.GraphTensor:\n    op_features = dict(graph.node_sets['op'].features)\n    op_features['op_e'] = self.embedding_layer(\n        tf.cast(graph.node_sets['op']['op'], tf.int32))\n    return graph.replace_features(node_sets={'op': op_features})\n\n","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:50:47.111492Z","iopub.execute_input":"2023-10-09T04:50:47.111849Z","iopub.status.idle":"2023-10-09T04:50:47.122694Z","shell.execute_reply.started":"2023-10-09T04:50:47.111825Z","shell.execute_reply":"2023-10-09T04:50:47.121828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Layout Training Pipeline\n\nWe start by defining constants:\n\n1. Batch sizes = num graphs, num sampled nodes per graph, and num configurations per graph.\n1. Collection to train on: source (`xla` versus `nlp`) and search stragey (`random` versus `default`).\n\nThen, boilerplate code to prepare the datasets.\n\nThen, we dive deeper into the dataset examples (a batch of graphs from the tiles collection).\n\nFinally, details on defining a model.","metadata":{}},{"cell_type":"markdown","source":"## Define constants and choose subcollection\n\nWe load `BATCH_SIZE` graphs per batch. Each will have ","metadata":{}},{"cell_type":"code","source":"LAYOUT_DATA_ROOT = '/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout'\nSOURCE = 'xla'  # Can be \"xla\" or \"nlp\"\nSEARCH = 'random'  # Can be \"random\" or \"default\"\n\n# Batch size information.\nBATCH_SIZE = 16  # Number of graphs per batch.\nCONFIGS_PER_GRAPH = 5  # Number of configurations (features and target values) per graph.\nMAX_KEEP_NODES = 1000  # Useful for dropout.\n# `MAX_KEEP_NODES` is (or, is not) useful for Segment Dropout, if model uses\n# edges \"sampled_config\" and \"sampled_feed\" (or, \"config\" and \"feed\")","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:50:47.124238Z","iopub.execute_input":"2023-10-09T04:50:47.124508Z","iopub.status.idle":"2023-10-09T04:50:47.144731Z","shell.execute_reply.started":"2023-10-09T04:50:47.124487Z","shell.execute_reply":"2023-10-09T04:50:47.143981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare `tf.data.Dataset` instances\nSpecifically, `layout_train_ds` and `layout_valid_ds`.\n\nIt can take ~10 minutes if you are running for the first time, for the caches to be created.","metadata":{}},{"cell_type":"code","source":"\nlayout_data_root_dir = os.path.join(\n      os.path.expanduser(LAYOUT_DATA_ROOT), SOURCE, SEARCH)\n\nlayout_npz_dataset = layout_data.get_npz_dataset(\n    layout_data_root_dir,\n    min_train_configs=CONFIGS_PER_GRAPH,\n    max_train_configs=500,  # If any graph has more than this configurations, it will be filtered [speeds up loading + training]\n    cache_dir='cache'\n)\n\ndef pair_layout_graph_with_label(graph: tfgnn.GraphTensor):\n    \"\"\"Extracts label from graph (`tfgnn.GraphTensor`) and returns a pair of `(graph, label)`\"\"\"\n    # Return runtimes divded over large number: only ranking is required. The\n    # runtimes are in the 100K range\n    label = tf.cast(graph.node_sets['g']['runtimes'], tf.float32) / 1e7\n    return graph, label\n\nlayout_train_ds = (\n      layout_npz_dataset.train.get_graph_tensors_dataset(\n          CONFIGS_PER_GRAPH, max_nodes=MAX_KEEP_NODES)\n      .shuffle(100, reshuffle_each_iteration=True)\n      .batch(BATCH_SIZE, drop_remainder=False)\n      .map(tfgnn.GraphTensor.merge_batch_to_components)\n      .map(pair_layout_graph_with_label))\n\nlayout_valid_ds = (\n      layout_npz_dataset.validation.get_graph_tensors_dataset(\n          CONFIGS_PER_GRAPH)\n      .batch(BATCH_SIZE, drop_remainder=False)\n      .map(tfgnn.GraphTensor.merge_batch_to_components)\n      .map(pair_layout_graph_with_label))","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:50:47.145939Z","iopub.execute_input":"2023-10-09T04:50:47.146178Z","iopub.status.idle":"2023-10-09T04:52:32.019219Z","shell.execute_reply.started":"2023-10-09T04:50:47.146158Z","shell.execute_reply":"2023-10-09T04:52:32.017799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Familiarize yourself with data\n\nLets obtain an example from the dataset `layout_train_ds`, i.e., an instance of `GraphTensor` which encodes a batch\nof graphs. Luckily, using TF-GNN, we can describe our model as-if we are operating on a single graph, and naturally the\nmodel extends to multiple graphs!\n\nLet's take one example (containing a batch) and print it.","metadata":{}},{"cell_type":"code","source":"graph_batch, config_runtimes = next(iter(layout_train_ds.take(1)))\n\nprint('graph_batch = ')\nprint(graph_batch)\nprint('\\n\\n')\nprint('config_runtimes=')\nprint(config_runtimes)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T05:32:31.645759Z","iopub.execute_input":"2023-10-09T05:32:31.646032Z","iopub.status.idle":"2023-10-09T05:32:33.536688Z","shell.execute_reply.started":"2023-10-09T05:32:31.646009Z","shell.execute_reply":"2023-10-09T05:32:33.535126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**< Crash-course on TF-GNN >**\n\nEach `GraphTensor` contains three fields:\n\n1. `node_sets`, can be thought of `dict` from node type (str) in the graph (batch) to feature tensors for that node type.\n1. `edge_sets`, can be thought of `dict` from edge type (str) in the graph (batch) to adjacency, as two integer vectors: source IDs and target IDs -- i.e., all edges are directed, unless explicitly undirected by the model. If edge set `e` connects from node-set `n1` to node-set `n2`, then if `graph.edge_sets[\"e1\"].adjacency.source = [0, 13, ...]` and `graph.edge_sets[\"e1\"].adjacency.target = [1, 14, ...]` (must be of equal length), then node `0` from node-set `n1` points to node `1` from node-set `n2`. The IDs are zero-based, and used to index into the feature tensors at `graph.node_sets[\"n1\"]` and `graph.node_sets[\"n2\"]`.\n1. `context`, contains information per graph in the batch. We do not use this, for the layout collection, as we have singleton nodeset per graph with name `\"g\"` (with features accessible as `graph.node_sets[\"g\"]`)\n\n**</ Crash-course on TF-GNN >**","metadata":{}},{"cell_type":"markdown","source":"Now, lets print the node-sets and the edge-sets of the example `graph_batch`:","metadata":{}},{"cell_type":"code","source":"# The `graph_batch` contains node-sets and edge-sets.\n# There are no context features for layout collection\nprint('graph_batch.context =', graph_batch.context)\n# Note: graph_batch.context.sizes must be equal to BATCH_SIZE.\n# Lets print-out all features for all nodesets.\n\nfor node_set_name in sorted(graph_batch.node_sets.keys()):\n    print(f'\\n\\n #####  NODE SET \"{node_set_name}\" #########')\n    print('** Has sizes: ', graph_batch.node_sets[node_set_name].sizes)\n    for feature_name in graph_batch.node_sets[node_set_name].features.keys():\n        print(f'\\n Feature \"{feature_name}\" has values')\n        print(graph_batch.node_sets[node_set_name][feature_name])\n","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:33.786805Z","iopub.execute_input":"2023-10-09T04:52:33.787145Z","iopub.status.idle":"2023-10-09T04:52:33.796896Z","shell.execute_reply.started":"2023-10-09T04:52:33.787115Z","shell.execute_reply":"2023-10-09T04:52:33.795809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The node set `'g'` corresponds to the \"graph-level\"**. Since `BATCH_SIZE==16`, each tensor in `'g'` should have a leading dimension of `16`. The `graph_id` feature contains model names. Since `CONFIGS_PER_GRAPH_PER_EPOCH=5`, then feature 'runtimes' must be of shape `(16, 5)` with `graph_batch.node_sets['g']['runtimes'][i, j]` indicating the runtime when compiling graph `i` with configuration features `j`. These specific feature values must be found in `nconfig` node-set, as explained next.\n\n**Lets look at nodes per graph**. For instance, node set `op` contains the operation nodes in the tensorflow graph (e.g., element-wise add, matrix multiply, etc). Op-codes are stored in `graph_batch.node_sets['op']['op']`. Since each graph has variable number of nodes, the array `graph_batch.node_sets['op'].sizes` gives the number of `op` nodes per (of the `16`) graphs.\n\nSome nodes are configurable. The (*virtual*) node-set `nconfig` contains features for configurable nodes. The features are in `graph_batch.node_sets['nconfig']['feats']`.\n\n\nThe edge-set `'config'` (next) indicates the correspondence between `nconfig` features and `op` nodes. Specifically, each (*virtual*) `config` node has degree of 1 and each `op` node has degree of 0 or 1 (on edge-set `'config'`).\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"Let's print-out all the edge-sets.\n\n","metadata":{}},{"cell_type":"code","source":"print('\\n config edge set: ', graph_batch.edge_sets['config'])  \nprint('\\n config source nodes: ', graph_batch.edge_sets['config'].adjacency.source)\nprint('\\n config target nodes: ', graph_batch.edge_sets['config'].adjacency.target)\nprint('\\n g_op edge set: ', graph_batch.edge_sets['g_op'])\nprint('\\n g_config edge set: ', graph_batch.edge_sets['g_config'])","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:33.798874Z","iopub.execute_input":"2023-10-09T04:52:33.799238Z","iopub.status.idle":"2023-10-09T04:52:33.819375Z","shell.execute_reply.started":"2023-10-09T04:52:33.799207Z","shell.execute_reply":"2023-10-09T04:52:33.817574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The edge-set `'config'` pairs each `\"nconfig\"` node with one `\"op\"` node. To list the correspondences, you print the `.adjacency.source` and `.adjacency.target`:","metadata":{}},{"cell_type":"code","source":"print(graph_batch.edge_sets['config'])   # Holds directed adjacency as list of pairs of indices: nconfig->op\nprint(graph_batch.edge_sets['config'].adjacency.source)  # Print nconfig indices (should be a range())\nprint(graph_batch.edge_sets['config'].adjacency.target)  # Print corresponding `op` indices.","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:33.822886Z","iopub.execute_input":"2023-10-09T04:52:33.823222Z","iopub.status.idle":"2023-10-09T04:52:33.835130Z","shell.execute_reply.started":"2023-10-09T04:52:33.823199Z","shell.execute_reply":"2023-10-09T04:52:33.834017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Other than `config` edges, the remainder of the edge-sets are:\n```\n'feed', 'g_op', 'g_config', 'sampled_config', 'sampled_feed'\n```\n\nThe first (`feed`) is the actual computation graph! `op` nodes feed into `op` nodes. **Note: The \"transpose\" of this adjacency (implicit) matrix indicates the direction of information flow (models are later in the tutorial).**. The second (`g_op`) and third (`g_config`), group by graph, respectively, `op` nodes and the (virtual) `nconfig` nodes. This edge-set can be helpful for global-pooling operations. \n\n*Segment-level Training*: Finally, to implement some version of **dropout**, `sampled_config` and `sampled_feed` edge-sets contain edges to randomly-sampled `op` nodes. To do full-graph (training or inference), you may use `config` and `feed`. To do training with segment dropout (e.g., a naive version of https://arxiv.org/abs/2308.13490, to appear @ NeurIPS'23), you may use `sampled_config` and `sampled_feed`. You may adjust the number of **keep** nodes by setting `MAX_KEEP_NODES`. An edge only survives in `sampled_feed` only if both of its endpoints survived (segment-level) dropout. In our naive implementation here, nodes with contiguous indices are kept. However, you are welcome to re-implement a better segmentation strategy.\n\n\n*NOTE: When using TF-GNN (models to follow), you dont have to worry about `sizes`: just write your model code as-if you are operating on a single graph, and the code naturally extends to a batch of graphs.*","metadata":{}},{"cell_type":"markdown","source":"## Modeling\n\nBefore we define the full model (`ResModel`), lets run some modeling functions. For example, let's embed the op-codes.\n\nWe have `layout_npz_dataset.num_ops` unique number of op codes, which determines the embedding size.","metadata":{}},{"cell_type":"code","source":"num_ops = layout_npz_dataset.num_ops\nprint('number of ops in the dataset=', num_ops)\n\nembedding_layer = _OpEmbedding(num_ops, 16)  # 16-dimensional embedding, for demonstration.\ngraph_batch_embedded_ops = embedding_layer(graph_batch)\n\nprint('\\n\\n Before embedding, node-set \"op\"=\\n', graph_batch.node_sets['op'])\nprint('\\n\\n After embedding, node-set \"op\"=\\n', graph_batch_embedded_ops.node_sets['op'])","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:33.836182Z","iopub.execute_input":"2023-10-09T04:52:33.836480Z","iopub.status.idle":"2023-10-09T04:52:33.906157Z","shell.execute_reply.started":"2023-10-09T04:52:33.836450Z","shell.execute_reply":"2023-10-09T04:52:33.905182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Note: after embedding, an additional feature `\"op_e\"` shows-up.*\n\nNow, lets concatenate the configuration features with the embedding features.","metadata":{}},{"cell_type":"code","source":"op_e = graph_batch_embedded_ops.node_sets['op']['op_e']\nconfig_features = graph_batch_embedded_ops.node_sets['nconfig']['feats']\n\nprint('op_e.shape ==', op_e.shape)\nprint('config_features.shape ==', config_features.shape)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:33.907124Z","iopub.execute_input":"2023-10-09T04:52:33.907331Z","iopub.status.idle":"2023-10-09T04:52:33.912536Z","shell.execute_reply.started":"2023-10-09T04:52:33.907314Z","shell.execute_reply":"2023-10-09T04:52:33.911517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are two differences in the shapes, yet, we concatenate them.\n\n1. `op_e` has more nodes: every node has an op-code, but not every node is configurable. We first to resize the leading dimension of `config_features` to equal the leading dimension of `op_e`, by filling zeros for nodes that are not configurable.\n1. `config_features` is cuboid. The middle dimension identifies the configuration: there are `CONFIGS_PER_GRAPH` of them.\n\n\nFor the first, we can multiply by the (sparse) \"config\" adjacency matrix -- a binary matrix where every is a one-hot and most rows are zero. If adjacency entry at `[i, j]` is set, then `graph.node_sets[\"nconfig\"][\"feats\"][j]`  contain configuration features for node `i` of `graph.node_sets[\"op\"]`.","metadata":{}},{"cell_type":"code","source":"config_adj = implicit.AdjacencyMultiplier(graph_batch_embedded_ops, 'config')\nprint('config_adj.shape =', config_adj.shape)\nresized_config_features = config_adj @ config_features\nprint('resized_config_features.shape =', resized_config_features.shape)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:33.913514Z","iopub.execute_input":"2023-10-09T04:52:33.913740Z","iopub.status.idle":"2023-10-09T04:52:33.944428Z","shell.execute_reply.started":"2023-10-09T04:52:33.913694Z","shell.execute_reply":"2023-10-09T04:52:33.942901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we want to broadcast the `op_e` feature matrix to a cuboid, by replicating on a (new) inner dimension so that we can finally combine the config features with op-embeddings.","metadata":{}},{"cell_type":"code","source":"broadcasted_op_e = tf.stack([op_e] * CONFIGS_PER_GRAPH, axis=1)\n\ncombined_features = tf.concat([broadcasted_op_e, resized_config_features], axis=-1)\n\nprint('combined_features.shape = ', combined_features.shape)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:33.946031Z","iopub.execute_input":"2023-10-09T04:52:33.946361Z","iopub.status.idle":"2023-10-09T04:52:33.979141Z","shell.execute_reply.started":"2023-10-09T04:52:33.946335Z","shell.execute_reply":"2023-10-09T04:52:33.978530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we want to do graph convolution layer (i.e., message-passing followed by non-linearity) among the `feed` edges. Usually, this can be done by left-multiplying the feature tensor with **some form** of an adjacency matrix. The exact form will determine the pooling (e.g., sum VS average). Let us use the symmetrically-normalized adjacency matrix with self-connections added (by Kipf & Welling, ICLR'17).\n\nWe can compute such a matrix $\\widehat{A}$ as:\n\n$$A_\\textrm{undirected.w.selfconnections} \\leftarrow A + A^\\top + I$$\n\n\n$$D \\leftarrow \\mathbf{1}^\\top A_\\textrm{undirected.w.selfconnections}$$\n\n\n$$\\widehat{A} \\leftarrow D^{-\\frac{1}{2}} (A_\\textrm{undirected.w.selfconnections}) D^{-\\frac{1}{2}} $$\n\nWhich is acheivable by the following code:","metadata":{}},{"cell_type":"code","source":"adj_op_op = implicit.AdjacencyMultiplier(graph_batch_embedded_ops, 'feed')  # op->op\nadj_config = implicit.AdjacencyMultiplier(graph_batch_embedded_ops, 'config')  # nconfig->op\n\nadj_op_op_hat = (adj_op_op + adj_op_op.transpose()).add_eye()\nadj_op_op_hat = adj_op_op_hat.normalize_symmetric()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:33.979972Z","iopub.execute_input":"2023-10-09T04:52:33.980225Z","iopub.status.idle":"2023-10-09T04:52:34.050743Z","shell.execute_reply.started":"2023-10-09T04:52:33.980202Z","shell.execute_reply":"2023-10-09T04:52:34.049594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, the message passing can written as:","metadata":{}},{"cell_type":"code","source":"A_times_X = adj_op_op_hat @ combined_features\nprint('A_times_x.shape =', A_times_X.shape)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:34.051809Z","iopub.execute_input":"2023-10-09T04:52:34.052065Z","iopub.status.idle":"2023-10-09T04:52:37.197989Z","shell.execute_reply.started":"2023-10-09T04:52:34.052044Z","shell.execute_reply":"2023-10-09T04:52:37.196635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we put together everything above to write a model class `ResModel` (next), which has a couple more concepts:\n\n1. Adjacency for `\"g_op\"` and `\"g_config\"`, which is used to pool information from all ops and from configurable ops, to the graph level.\n1. Residual connections.\n1. Segment dropout. a forward-pass is computed on the entire graph (but, with `tf.stop_gradient`). Then, another forward pass is computed using only sampled edge-sets (`edgeset_prefix` is set to `\"sampled_\"` by `forward()`).\n\nWithout further ado, `ResModel`:","metadata":{}},{"cell_type":"code","source":"class ResModel(tf.keras.Model):\n    \"\"\"GNN with residual connections.\"\"\"\n\n    def __init__(\n        self, num_configs: int, num_ops: int, op_embed_dim: int = 32,\n        num_gnns: int = 2, mlp_layers: int = 2,\n        hidden_activation: str = 'leaky_relu',\n        hidden_dim: int = 32, reduction: str = 'sum'):\n        super().__init__()\n        self._num_configs = num_configs\n        self._num_ops = num_ops\n        self._op_embedding = _OpEmbedding(num_ops, op_embed_dim)\n        self._prenet = _mlp([hidden_dim] * mlp_layers, hidden_activation)\n        self._gc_layers = []\n        for _ in range(num_gnns):\n            self._gc_layers.append(_mlp([hidden_dim] * mlp_layers, hidden_activation))\n        self._postnet = _mlp([hidden_dim, 1], hidden_activation, use_bias=False)\n\n    def call(self, graph: tfgnn.GraphTensor, training: bool = False):\n        del training\n        return self.forward(graph, self._num_configs)\n\n    def _node_level_forward(\n        self, node_features: tf.Tensor,\n        config_features: tf.Tensor,\n        graph: tfgnn.GraphTensor, num_configs: int,\n        edgeset_prefix='') -> tf.Tensor:\n        adj_op_op = implicit.AdjacencyMultiplier(\n            graph, edgeset_prefix+'feed')  # op->op\n        adj_config = implicit.AdjacencyMultiplier(\n            graph, edgeset_prefix+'config')  # nconfig->op\n\n        adj_op_op_hat = (adj_op_op + adj_op_op.transpose()).add_eye()\n        adj_op_op_hat = adj_op_op_hat.normalize_symmetric()\n\n        x = node_features\n\n        x = tf.stack([x] * num_configs, axis=1)\n        config_features = 100 * (adj_config @ config_features)\n        x = tf.concat([config_features, x], axis=-1)\n        x = self._prenet(x)\n        x = tf.nn.leaky_relu(x)\n\n        for layer in self._gc_layers:\n            y = x\n            y = tf.concat([config_features, y], axis=-1)\n            y = tf.nn.leaky_relu(layer(adj_op_op_hat @ y))\n            x += y\n        return x\n\n    def forward(\n        self, graph: tfgnn.GraphTensor, num_configs: int,\n        backprop=True) -> tf.Tensor:\n        graph = self._op_embedding(graph)\n\n        config_features = graph.node_sets['nconfig']['feats']\n        node_features = tf.concat([\n            graph.node_sets['op']['feats'],\n            graph.node_sets['op']['op_e']\n        ], axis=-1)\n\n        x_full = self._node_level_forward(\n            node_features=tf.stop_gradient(node_features),\n            config_features=tf.stop_gradient(config_features),\n            graph=graph, num_configs=num_configs)\n\n        if backprop:\n            x_backprop = self._node_level_forward(\n                node_features=node_features,\n                config_features=config_features,\n                graph=graph, num_configs=num_configs, edgeset_prefix='sampled_')\n\n            is_selected = graph.node_sets['op']['selected']\n            # Need to expand twice as `is_selected` is a vector (num_nodes) but\n            # x_{backprop, full} are 3D tensors (num_nodes, num_configs, num_feats).\n            is_selected = tf.expand_dims(is_selected, axis=-1)\n            is_selected = tf.expand_dims(is_selected, axis=-1)\n            x = tf.where(is_selected, x_backprop, x_full)\n        else:\n            x = x_full\n\n        adj_config = implicit.AdjacencyMultiplier(graph, 'config')\n\n        # Features for configurable nodes.\n        config_feats = (adj_config.transpose() @ x)\n\n        # Global pooling\n        adj_pool_op_sum = implicit.AdjacencyMultiplier(graph, 'g_op').transpose()\n        adj_pool_op_mean = adj_pool_op_sum.normalize_right()\n        adj_pool_config_sum = implicit.AdjacencyMultiplier(\n            graph, 'g_config').transpose()\n        x = self._postnet(tf.concat([\n            # (A D^-1) @ Features\n            adj_pool_op_mean @ x,\n            # l2_normalize( A @ Features )\n            tf.nn.l2_normalize(adj_pool_op_sum @ x, axis=-1),\n            # l2_normalize( A @ Features )\n            tf.nn.l2_normalize(adj_pool_config_sum @ config_feats, axis=-1),\n        ], axis=-1))\n\n        x = tf.squeeze(x, -1)\n\n        return x\n\n","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:37.199127Z","iopub.execute_input":"2023-10-09T04:52:37.199388Z","iopub.status.idle":"2023-10-09T04:52:37.214778Z","shell.execute_reply.started":"2023-10-09T04:52:37.199366Z","shell.execute_reply":"2023-10-09T04:52:37.213644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training loop\n\nCreate a model, objective function, and optimizer.","metadata":{}},{"cell_type":"code","source":"model = ResModel(CONFIGS_PER_GRAPH, layout_npz_dataset.num_ops)\n\nloss = tfr.keras.losses.ListMLELoss()  # (temperature=10)\nopt = tf.keras.optimizers.Adam(learning_rate=1e-3, clipnorm=0.5)\n\nmodel.compile(loss=loss, optimizer=opt, metrics=[\n    tfr.keras.metrics.OPAMetric(name='opa_metric'),\n])\n","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:37.216126Z","iopub.execute_input":"2023-10-09T04:52:37.216420Z","iopub.status.idle":"2023-10-09T04:52:37.278641Z","shell.execute_reply.started":"2023-10-09T04:52:37.216398Z","shell.execute_reply":"2023-10-09T04:52:37.277602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train for a few epochs.","metadata":{}},{"cell_type":"code","source":"early_stop = 5  # If validation OPA did not increase in this many epochs, terminate training.\nbest_params = None  # Stores parameters corresponding to best validation OPA, to restore to them after training.\nbest_val_opa = -1  # Tracks best validation OPA\nbest_val_at_epoch = -1  # At which epoch.\nepochs = 1  # Total number of training epochs.\n\nfor i in range(epochs):\n    history = model.fit(\n        layout_train_ds, epochs=1, verbose=1, validation_data=layout_valid_ds,\n        validation_freq=1)\n\n    train_loss = history.history['loss'][-1]\n    train_opa = history.history['opa_metric'][-1]\n    val_loss = history.history['val_loss'][-1]\n    val_opa = history.history['val_opa_metric'][-1]\n    if val_opa > best_val_opa:\n        best_val_opa = val_opa\n        best_val_at_epoch = i\n        best_params = {v.ref: v + 0 for v in model.trainable_variables}\n        print(' * [@%i] Validation (NEW BEST): %s' % (i, str(val_opa)))\n    elif early_stop > 0 and i - best_val_at_epoch >= early_stop:\n      print('[@%i] Best accuracy was attained at epoch %i. Stopping.' % (i, best_val_at_epoch))\n      break\n\n# Restore best parameters.\nprint('Restoring parameters corresponding to the best validation OPA.')\nassert best_params is not None\nfor v in model.trainable_variables:\n    v.assign(best_params[v.ref])","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:52:37.279912Z","iopub.execute_input":"2023-10-09T04:52:37.280362Z","iopub.status.idle":"2023-10-09T04:55:29.270338Z","shell.execute_reply.started":"2023-10-09T04:52:37.280309Z","shell.execute_reply":"2023-10-09T04:55:29.269170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Make Submission CSV file for this task","metadata":{}},{"cell_type":"code","source":"import tqdm\n_INFERENCE_CONFIGS_BATCH_SIZE = 50\n\noutput_csv_filename = f'inference_layout_{SOURCE}_{SEARCH}.csv'\nprint('\\n\\n   Running inference on test set ...\\n\\n')\ntest_rankings = []\n\nassert layout_npz_dataset.test.graph_id is not None\nfor graph in tqdm.tqdm(layout_npz_dataset.test.iter_graph_tensors(),\n                       total=layout_npz_dataset.test.graph_id.shape[-1],\n                       desc='Inference'):\n    num_configs = graph.node_sets['g']['runtimes'].shape[-1]\n    all_scores = []\n    for i in tqdm.tqdm(range(0, num_configs, _INFERENCE_CONFIGS_BATCH_SIZE)):\n        end_i = min(i + _INFERENCE_CONFIGS_BATCH_SIZE, num_configs)\n        # Take a cut of the configs.\n        node_set_g = graph.node_sets['g']\n        subconfigs_graph = tfgnn.GraphTensor.from_pieces(\n            edge_sets=graph.edge_sets,\n            node_sets={\n                'op': graph.node_sets['op'],\n                'nconfig': tfgnn.NodeSet.from_fields(\n                    sizes=graph.node_sets['nconfig'].sizes,\n                    features={\n                        'feats': graph.node_sets['nconfig']['feats'][:, i:end_i],\n                    }),\n                'g': tfgnn.NodeSet.from_fields(\n                    sizes=tf.constant([1]),\n                    features={\n                        'graph_id': node_set_g['graph_id'],\n                        'runtimes': node_set_g['runtimes'][:, i:end_i],\n                        'kept_node_ratio': node_set_g['kept_node_ratio'],\n                    })\n            })\n        h = model.forward(subconfigs_graph, num_configs=(end_i - i),\n                          backprop=False)\n        all_scores.append(h[0])\n    all_scores = tf.concat(all_scores, axis=0)\n    graph_id = graph.node_sets['g']['graph_id'][0].numpy().decode()\n    sorted_indices = tf.strings.join(\n        tf.strings.as_string(tf.argsort(all_scores)), ';').numpy().decode()\n    test_rankings.append((graph_id, sorted_indices))\n\nwith tf.io.gfile.GFile(output_csv_filename, 'w') as fout:\n    fout.write('ID,TopConfigs\\n')\n    for graph_id, ranks in test_rankings:\n        fout.write(f'layout:{SOURCE}:{SEARCH}:{graph_id},{ranks}\\n')\nprint('\\n\\n   ***  Wrote', output_csv_filename, '\\n\\n')\n","metadata":{"execution":{"iopub.status.busy":"2023-10-09T04:55:29.271610Z","iopub.execute_input":"2023-10-09T04:55:29.271885Z","iopub.status.idle":"2023-10-09T05:32:25.670411Z","shell.execute_reply.started":"2023-10-09T04:55:29.271864Z","shell.execute_reply":"2023-10-09T05:32:25.668868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Combine submission CSVs from all layout collections into one CSV file\n\nFinally, after running on all collections, you need to combine the CSV files together (e.g., by concatenation), to prepare the final submission. Specifically, you can modify the constants:\n\n```\nSOURCE = 'xla'  # Can be \"xla\" or \"nlp\"\nSEARCH = 'random'  # Can be \"random\" or \"default\"\n```\n\n(from a few cells ago) and run for all 4 combinations: SOURCE=(\"xla\", \"nlp\") x SEARCH=(\"random\", \"default\"), then combine all inferences into one file:","metadata":{}},{"cell_type":"code","source":"!cat inference_layout_xla_random.csv inference_layout_xla_default.csv inference_layout_nlp_random.csv inference_layout_nlp_default.csv > inference_layout_all.csv","metadata":{"execution":{"iopub.status.busy":"2023-10-09T08:18:02.295968Z","iopub.execute_input":"2023-10-09T08:18:02.296276Z","iopub.status.idle":"2023-10-09T08:18:02.585692Z","shell.execute_reply.started":"2023-10-09T08:18:02.296249Z","shell.execute_reply":"2023-10-09T08:18:02.584032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"producing file `\"inference_layout_all.csv\"` that combines all predictions for all layout subcollections. Finally, this file should be combined with the CSV for the tiles collection, as explained next.","metadata":{}},{"cell_type":"markdown","source":"# Tile Training Pipeline\n\nThis section will be written by end of September. We prioritized getting this notebook out, as soon as possible, as the above Layout section is (1) more tricky and (2) most of the score depends on it.","metadata":{}}]}