{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"tpu1vmV38","dataSources":[{"sourceId":58266,"databundleVersionId":6641124,"sourceType":"competition"},{"sourceId":144043388,"sourceType":"kernelVersion"},{"sourceId":144045966,"sourceType":"kernelVersion"},{"sourceId":144045983,"sourceType":"kernelVersion"}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"!pip install tensorflow_gnn --pre\n!pip install tensorflow_ranking","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:23:51.869684Z","iopub.execute_input":"2025-05-08T10:23:51.869921Z","iopub.status.idle":"2025-05-08T10:26:14.147691Z","shell.execute_reply.started":"2025-05-08T10:23:51.869893Z","shell.execute_reply":"2025-05-08T10:26:14.146858Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Install standard modules\n\nimport os\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nimport tensorflow_gnn as tfgnn\nimport tensorflow_ranking as tfr","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:26:14.149590Z","iopub.execute_input":"2025-05-08T10:26:14.149842Z","iopub.status.idle":"2025-05-08T10:26:26.451124Z","shell.execute_reply.started":"2025-05-08T10:26:14.149816Z","shell.execute_reply":"2025-05-08T10:26:26.450226Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The utility modules are based on the code on [github](): ","metadata":{}},{"cell_type":"code","source":"# Install utility modules.\n\nimport tpugraphsv1_layout_data_py as layout_data\nimport tpugraphsv1_tile_data_py as tile_data\nimport tpugraphsv1_implicit_py as implicit","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:26:26.452115Z","iopub.execute_input":"2025-05-08T10:26:26.452613Z","iopub.status.idle":"2025-05-08T10:26:38.108886Z","shell.execute_reply.started":"2025-05-08T10:26:26.452564Z","shell.execute_reply":"2025-05-08T10:26:38.107921Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Training Pipelines\n\nThe following code is organized as:\n\n 1. Helper functions: MLP (`_mlp`) and Embedding layer (`_Opembedding`). The embedding layer amends a feature on the `op` nodes, with name `op_e`, by embedding the integral op IDs.\n 1. Pipeline code for training on the Layout collections.\n 1. Pipeline code for training on the Tile collection.","metadata":{}},{"cell_type":"markdown","source":"## Helper functions, for both Layout and Tile collections.","metadata":{}},{"cell_type":"code","source":"def _mlp(dims, hidden_activation='relu', l2reg=1e-4, use_bias=True,\n         dropout_rate=0.0, batch_norm=False):\n  \"\"\"Creates a customizable multi-layer perceptron (MLP).\n\n  Args:\n    dims: List[int] - output dimensions of each dense layer.\n    hidden_activation: Activation function between layers.\n    l2reg: L2 regularization factor.\n    use_bias: Whether to use bias in dense layers.\n    dropout_rate: Optional dropout rate after each activation.\n    batch_norm: Whether to apply batch normalization before activations.\n\n  Returns:\n    A tf.keras.Sequential model composed of Dense, Activation, BatchNorm, Dropout layers.\n  \"\"\"\n  layers = []\n  for i, dim in enumerate(dims):\n    if i > 0:\n      if batch_norm:\n        layers.append(tf.keras.layers.BatchNormalization())\n      layers.append(tf.keras.layers.Activation(hidden_activation))\n      if dropout_rate > 0:\n        layers.append(tf.keras.layers.Dropout(dropout_rate))\n    layers.append(tf.keras.layers.Dense(\n        units=dim,\n        use_bias=use_bias,\n        kernel_regularizer=tf.keras.regularizers.l2(l2reg)\n    ))\n  return tf.keras.Sequential(layers)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T10:26:38.110360Z","iopub.execute_input":"2025-05-08T10:26:38.110673Z","iopub.status.idle":"2025-05-08T10:26:38.117864Z","shell.execute_reply.started":"2025-05-08T10:26:38.110642Z","shell.execute_reply":"2025-05-08T10:26:38.117255Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class _OpEmbedding(tf.keras.Model):\n  \"\"\"\n  Embeds GraphTensor.node_sets['op']['op'] nodes into feature 'op_e'.\n\n  Parameters:\n    num_ops (int): Number of unique operations to embed.\n    embed_d (int): Dimension of each embedding vector.\n    l2reg (float, optional): L2 regularization coefficient. Default is 1e-4.\n    input_dtype (tf.DType, optional): Data type of input op IDs. Default is tf.int32.\n    trainable (bool, optional): Whether embedding weights are trainable. Default is True.\n    name (str, optional): Name for the embedding layer. Default is 'op_embedding'.\n  \"\"\"\n\n  def __init__(self,\n               num_ops: int = 100,\n               embed_d: int = 16,\n               l2reg: float = 1e-4,\n               input_dtype: tf.dtypes.DType = tf.int32,\n               trainable: bool = True,\n               name: str = 'op_embedding'):\n    super().__init__(name=name)\n    self.input_dtype = input_dtype\n    self.embedding_layer = tf.keras.layers.Embedding(\n        input_dim=num_ops,\n        output_dim=embed_d,\n        embeddings_initializer='uniform',\n        embeddings_regularizer=tf.keras.regularizers.l2(l2reg),\n        trainable=trainable,\n        name=f\"{name}_layer\"\n    )\n\n  def call(self,\n           graph: tfgnn.GraphTensor,\n           training: bool = False) -> tfgnn.GraphTensor:\n    # Extract current features from the 'op' node set\n    op_features = dict(graph.node_sets['op'].features)\n\n    # Cast op IDs to the correct dtype and generate embeddings\n    op_ids = tf.cast(graph.node_sets['op']['op'], self.input_dtype)\n    op_embeddings = self.embedding_layer(op_ids)\n\n    # Add new embedding feature to the node features\n    op_features['op_e'] = op_embeddings\n\n    # Return updated graph\n    return graph.replace_features(node_sets={'op': op_features})\n","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:26:38.118842Z","iopub.execute_input":"2025-05-08T10:26:38.119083Z","iopub.status.idle":"2025-05-08T10:26:38.131196Z","shell.execute_reply.started":"2025-05-08T10:26:38.119058Z","shell.execute_reply":"2025-05-08T10:26:38.130617Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Layout Training Pipeline\n\nWe start by defining constants:\n\n1. Batch sizes = num graphs, num sampled nodes per graph, and num configurations per graph.\n1. Collection to train on: source (`xla` versus `nlp`) and search stragey (`random` versus `default`).\n\nThen, boilerplate code to prepare the datasets.\n\nThen, we dive deeper into the dataset examples (a batch of graphs from the tiles collection).\n\nFinally, details on defining a model.","metadata":{}},{"cell_type":"markdown","source":"## Define constants and choose subcollection\n\nWe load `BATCH_SIZE` graphs per batch. Each will have ","metadata":{}},{"cell_type":"code","source":"LAYOUT_DATA_ROOT = '/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout'\nSOURCE = 'xla'  # Can be \"xla\" or \"nlp\"\nSEARCH = 'random'  # Can be \"random\" or \"default\"\n\n# Batch size information.\nBATCH_SIZE = 16  # Number of graphs per batch.\nCONFIGS_PER_GRAPH = 5  # Number of configurations (features and target values) per graph.\nMAX_KEEP_NODES = 1000  # Useful for dropout.\n# `MAX_KEEP_NODES` is (or, is not) useful for Segment Dropout, if model uses\n# edges \"sampled_config\" and \"sampled_feed\" (or, \"config\" and \"feed\")","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:26:38.132020Z","iopub.execute_input":"2025-05-08T10:26:38.132255Z","iopub.status.idle":"2025-05-08T10:26:38.144345Z","shell.execute_reply.started":"2025-05-08T10:26:38.132230Z","shell.execute_reply":"2025-05-08T10:26:38.143761Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Prepare `tf.data.Dataset` instances\nSpecifically, `layout_train_ds` and `layout_valid_ds`.\n\nIt can take ~10 minutes if you are running for the first time, for the caches to be created.","metadata":{}},{"cell_type":"code","source":"# Expanded function to generate path for dataset directory\ndef build_layout_data_path(layout_data_root, source, search):\n    \"\"\"\n    Function to construct dataset path dynamically.\n    \n    Args:\n        layout_data_root (str): The root directory of the dataset.\n        source (str): The source configuration (e.g., \"xla\").\n        search (str): The search method (e.g., \"random\").\n\n    Returns:\n        str: The constructed path for dataset.\n    \"\"\"\n    return os.path.join(os.path.expanduser(layout_data_root), source, search)\n\n# Expanded function to load dataset with configurable parameters\ndef load_layout_dataset(layout_data_root_dir, min_train_configs, max_train_configs, cache_dir):\n    \"\"\"\n    Function to load NPZ dataset, applying filters and caching.\n\n    Args:\n        layout_data_root_dir (str): Path to the directory containing the dataset.\n        min_train_configs (int): Minimum configurations per graph.\n        max_train_configs (int): Maximum configurations per graph (filtering).\n        cache_dir (str): Directory to store cached data.\n\n    Returns:\n        layout_data.Dataset: Loaded dataset object.\n    \"\"\"\n    return layout_data.get_npz_dataset(\n        layout_data_root_dir,\n        min_train_configs=min_train_configs,\n        max_train_configs=max_train_configs,\n        cache_dir=cache_dir\n    )\n\n# Updated label extraction to handle multiple graphs\ndef pair_layout_graph_with_label(graph: tfgnn.GraphTensor, scale_factor=1e7):\n    \"\"\"\n    Function to extract label (normalized runtime) from a GraphTensor.\n\n    Args:\n        graph (tfgnn.GraphTensor): The input graph.\n        scale_factor (float, optional): Scaling factor for runtimes, default is 1e7.\n\n    Returns:\n        tuple: A tuple of (graph, normalized_label).\n    \"\"\"\n    label = tf.cast(graph.node_sets['g']['runtimes'], tf.float32) / scale_factor\n    return graph, label\n\n# Dataset preparation function for training and validation\ndef prepare_dataset(dataset, configs_per_graph, batch_size, max_nodes, shuffle=True, reshuffle_each_iteration=True):\n    \"\"\"\n    Function to prepare the dataset for training or validation.\n\n    Args:\n        dataset: The raw dataset to be processed.\n        configs_per_graph (int): Number of configurations per graph.\n        batch_size (int): Batch size.\n        max_nodes (int): Maximum number of nodes to keep.\n        shuffle (bool, optional): Whether to shuffle the dataset, default is True.\n        reshuffle_each_iteration (bool, optional): Whether to reshuffle after each iteration, default is True.\n\n    Returns:\n        tf.data.Dataset: Prepared and processed dataset.\n    \"\"\"\n    dataset = dataset.get_graph_tensors_dataset(configs_per_graph, max_nodes=max_nodes)\n    \n    if shuffle:\n        dataset = dataset.shuffle(100, reshuffle_each_iteration=reshuffle_each_iteration)\n    \n    dataset = (\n        dataset.batch(batch_size, drop_remainder=False)\n        .map(tfgnn.GraphTensor.merge_batch_to_components)\n        .map(pair_layout_graph_with_label)\n    )\n    \n    return dataset\n\n# Construct dataset directory path using the expanded function\nlayout_data_root_dir = build_layout_data_path(LAYOUT_DATA_ROOT, SOURCE, SEARCH)\n\n# Load the dataset using the expanded loading function\nlayout_npz_dataset = load_layout_dataset(layout_data_root_dir, CONFIGS_PER_GRAPH, 500, 'cache')\n\n# Prepare training and validation datasets with the new modular function\nlayout_train_ds = prepare_dataset(layout_npz_dataset.train, CONFIGS_PER_GRAPH, BATCH_SIZE, MAX_KEEP_NODES)\nlayout_valid_ds = prepare_dataset(layout_npz_dataset.validation, CONFIGS_PER_GRAPH, BATCH_SIZE, MAX_KEEP_NODES, shuffle=False)\n\n","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:26:38.146989Z","iopub.execute_input":"2025-05-08T10:26:38.147241Z","iopub.status.idle":"2025-05-08T10:27:58.803420Z","shell.execute_reply.started":"2025-05-08T10:26:38.147215Z","shell.execute_reply":"2025-05-08T10:27:58.802598Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Familiarize yourself with data\n\nLets obtain an example from the dataset `layout_train_ds`, i.e., an instance of `GraphTensor` which encodes a batch\nof graphs. Luckily, using TF-GNN, we can describe our model as-if we are operating on a single graph, and naturally the\nmodel extends to multiple graphs!\n\nLet's take one example (containing a batch) and print it.","metadata":{}},{"cell_type":"code","source":"graph_batch, config_runtimes = next(iter(layout_train_ds.take(1)))\n\nprint('graph_batch = ')\nprint(graph_batch)\nprint('\\n\\n')\nprint('config_runtimes=')\nprint(config_runtimes)","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:27:58.804495Z","iopub.execute_input":"2025-05-08T10:27:58.804764Z","iopub.status.idle":"2025-05-08T10:28:00.991595Z","shell.execute_reply.started":"2025-05-08T10:27:58.804735Z","shell.execute_reply":"2025-05-08T10:28:00.990777Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**< Crash-course on TF-GNN >**\n\nEach `GraphTensor` contains three fields:\n\n1. `node_sets`, can be thought of `dict` from node type (str) in the graph (batch) to feature tensors for that node type.\n1. `edge_sets`, can be thought of `dict` from edge type (str) in the graph (batch) to adjacency, as two integer vectors: source IDs and target IDs -- i.e., all edges are directed, unless explicitly undirected by the model. If edge set `e` connects from node-set `n1` to node-set `n2`, then if `graph.edge_sets[\"e1\"].adjacency.source = [0, 13, ...]` and `graph.edge_sets[\"e1\"].adjacency.target = [1, 14, ...]` (must be of equal length), then node `0` from node-set `n1` points to node `1` from node-set `n2`. The IDs are zero-based, and used to index into the feature tensors at `graph.node_sets[\"n1\"]` and `graph.node_sets[\"n2\"]`.\n1. `context`, contains information per graph in the batch. We do not use this, for the layout collection, as we have singleton nodeset per graph with name `\"g\"` (with features accessible as `graph.node_sets[\"g\"]`)\n\n**</ Crash-course on TF-GNN >**","metadata":{}},{"cell_type":"markdown","source":"Now, lets print the node-sets and the edge-sets of the example `graph_batch`:","metadata":{}},{"cell_type":"code","source":"# The `graph_batch` contains node-sets and edge-sets.\n# There are no context features for layout collection\nprint('graph_batch.context =', graph_batch.context)\n# Note: graph_batch.context.sizes must be equal to BATCH_SIZE.\n# Lets print-out all features for all nodesets.\n\nfor node_set_name in sorted(graph_batch.node_sets.keys()):\n    print(f'\\n\\n #####  NODE SET \"{node_set_name}\" #########')\n    print('** Has sizes: ', graph_batch.node_sets[node_set_name].sizes)\n    for feature_name in graph_batch.node_sets[node_set_name].features.keys():\n        print(f'\\n Feature \"{feature_name}\" has values')\n        print(graph_batch.node_sets[node_set_name][feature_name])\n","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:00.992767Z","iopub.execute_input":"2025-05-08T10:28:00.993080Z","iopub.status.idle":"2025-05-08T10:28:01.002593Z","shell.execute_reply.started":"2025-05-08T10:28:00.993048Z","shell.execute_reply":"2025-05-08T10:28:01.001927Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**The node set `'g'` corresponds to the \"graph-level\"**. Since `BATCH_SIZE==16`, each tensor in `'g'` should have a leading dimension of `16`. The `graph_id` feature contains model names. Since `CONFIGS_PER_GRAPH_PER_EPOCH=5`, then feature 'runtimes' must be of shape `(16, 5)` with `graph_batch.node_sets['g']['runtimes'][i, j]` indicating the runtime when compiling graph `i` with configuration features `j`. These specific feature values must be found in `nconfig` node-set, as explained next.\n\n**Lets look at nodes per graph**. For instance, node set `op` contains the operation nodes in the tensorflow graph (e.g., element-wise add, matrix multiply, etc). Op-codes are stored in `graph_batch.node_sets['op']['op']`. Since each graph has variable number of nodes, the array `graph_batch.node_sets['op'].sizes` gives the number of `op` nodes per (of the `16`) graphs.\n\nSome nodes are configurable. The (*virtual*) node-set `nconfig` contains features for configurable nodes. The features are in `graph_batch.node_sets['nconfig']['feats']`.\n\n\nThe edge-set `'config'` (next) indicates the correspondence between `nconfig` features and `op` nodes. Specifically, each (*virtual*) `config` node has degree of 1 and each `op` node has degree of 0 or 1 (on edge-set `'config'`).\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"Let's print-out all the edge-sets.\n\n","metadata":{}},{"cell_type":"code","source":"print('\\n config edge set: ', graph_batch.edge_sets['config'])  \nprint('\\n config source nodes: ', graph_batch.edge_sets['config'].adjacency.source)\nprint('\\n config target nodes: ', graph_batch.edge_sets['config'].adjacency.target)\nprint('\\n g_op edge set: ', graph_batch.edge_sets['g_op'])\nprint('\\n g_config edge set: ', graph_batch.edge_sets['g_config'])","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:01.003559Z","iopub.execute_input":"2025-05-08T10:28:01.003840Z","iopub.status.idle":"2025-05-08T10:28:01.018321Z","shell.execute_reply.started":"2025-05-08T10:28:01.003810Z","shell.execute_reply":"2025-05-08T10:28:01.017674Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The edge-set `'config'` pairs each `\"nconfig\"` node with one `\"op\"` node. To list the correspondences, you print the `.adjacency.source` and `.adjacency.target`:","metadata":{}},{"cell_type":"code","source":"print(graph_batch.edge_sets['config'])   # Holds directed adjacency as list of pairs of indices: nconfig->op\nprint(graph_batch.edge_sets['config'].adjacency.source)  # Print nconfig indices (should be a range())\nprint(graph_batch.edge_sets['config'].adjacency.target)  # Print corresponding `op` indices.","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:01.019296Z","iopub.execute_input":"2025-05-08T10:28:01.019574Z","iopub.status.idle":"2025-05-08T10:28:01.027345Z","shell.execute_reply.started":"2025-05-08T10:28:01.019547Z","shell.execute_reply":"2025-05-08T10:28:01.026672Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Other than `config` edges, the remainder of the edge-sets are:\n```\n'feed', 'g_op', 'g_config', 'sampled_config', 'sampled_feed'\n```\n\nThe first (`feed`) is the actual computation graph! `op` nodes feed into `op` nodes. **Note: The \"transpose\" of this adjacency (implicit) matrix indicates the direction of information flow (models are later in the tutorial).**. The second (`g_op`) and third (`g_config`), group by graph, respectively, `op` nodes and the (virtual) `nconfig` nodes. This edge-set can be helpful for global-pooling operations. \n\n*Segment-level Training*: Finally, to implement some version of **dropout**, `sampled_config` and `sampled_feed` edge-sets contain edges to randomly-sampled `op` nodes. To do full-graph (training or inference), you may use `config` and `feed`. To do training with segment dropout (e.g., a naive version of https://arxiv.org/abs/2308.13490, to appear @ NeurIPS'23), you may use `sampled_config` and `sampled_feed`. You may adjust the number of **keep** nodes by setting `MAX_KEEP_NODES`. An edge only survives in `sampled_feed` only if both of its endpoints survived (segment-level) dropout. In our naive implementation here, nodes with contiguous indices are kept. However, you are welcome to re-implement a better segmentation strategy.\n\n\n*NOTE: When using TF-GNN (models to follow), you dont have to worry about `sizes`: just write your model code as-if you are operating on a single graph, and the code naturally extends to a batch of graphs.*","metadata":{}},{"cell_type":"markdown","source":"## Modeling\n\nBefore we define the full model (`ResModel`), lets run some modeling functions. For example, let's embed the op-codes.\n\nWe have `layout_npz_dataset.num_ops` unique number of op codes, which determines the embedding size.","metadata":{}},{"cell_type":"code","source":"num_ops = layout_npz_dataset.num_ops\nprint('number of ops in the dataset=', num_ops)\n\nembedding_layer = _OpEmbedding(num_ops, 16)  # 16-dimensional embedding, for demonstration.\ngraph_batch_embedded_ops = embedding_layer(graph_batch)\n\nprint('\\n\\n Before embedding, node-set \"op\"=\\n', graph_batch.node_sets['op'])\nprint('\\n\\n After embedding, node-set \"op\"=\\n', graph_batch_embedded_ops.node_sets['op'])","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:01.028293Z","iopub.execute_input":"2025-05-08T10:28:01.028567Z","iopub.status.idle":"2025-05-08T10:28:01.089566Z","shell.execute_reply.started":"2025-05-08T10:28:01.028540Z","shell.execute_reply":"2025-05-08T10:28:01.088936Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"*Note: after embedding, an additional feature `\"op_e\"` shows-up.*\n\nNow, lets concatenate the configuration features with the embedding features.","metadata":{}},{"cell_type":"code","source":"op_e = graph_batch_embedded_ops.node_sets['op']['op_e']\nconfig_features = graph_batch_embedded_ops.node_sets['nconfig']['feats']\n\nprint('op_e.shape ==', op_e.shape)\nprint('config_features.shape ==', config_features.shape)","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:01.090387Z","iopub.execute_input":"2025-05-08T10:28:01.090619Z","iopub.status.idle":"2025-05-08T10:28:01.094843Z","shell.execute_reply.started":"2025-05-08T10:28:01.090595Z","shell.execute_reply":"2025-05-08T10:28:01.094133Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are two differences in the shapes, yet, we concatenate them.\n\n1. `op_e` has more nodes: every node has an op-code, but not every node is configurable. We first to resize the leading dimension of `config_features` to equal the leading dimension of `op_e`, by filling zeros for nodes that are not configurable.\n1. `config_features` is cuboid. The middle dimension identifies the configuration: there are `CONFIGS_PER_GRAPH` of them.\n\n\nFor the first, we can multiply by the (sparse) \"config\" adjacency matrix -- a binary matrix where every is a one-hot and most rows are zero. If adjacency entry at `[i, j]` is set, then `graph.node_sets[\"nconfig\"][\"feats\"][j]`  contain configuration features for node `i` of `graph.node_sets[\"op\"]`.","metadata":{}},{"cell_type":"code","source":"config_adj = implicit.AdjacencyMultiplier(graph_batch_embedded_ops, 'config')\nprint('config_adj.shape =', config_adj.shape)\nresized_config_features = config_adj @ config_features\nprint('resized_config_features.shape =', resized_config_features.shape)","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:01.095709Z","iopub.execute_input":"2025-05-08T10:28:01.096197Z","iopub.status.idle":"2025-05-08T10:28:01.111802Z","shell.execute_reply.started":"2025-05-08T10:28:01.096170Z","shell.execute_reply":"2025-05-08T10:28:01.111032Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, we want to broadcast the `op_e` feature matrix to a cuboid, by replicating on a (new) inner dimension so that we can finally combine the config features with op-embeddings.","metadata":{}},{"cell_type":"code","source":"broadcasted_op_e = tf.stack([op_e] * CONFIGS_PER_GRAPH, axis=1)\n\ncombined_features = tf.concat([broadcasted_op_e, resized_config_features], axis=-1)\n\nprint('combined_features.shape = ', combined_features.shape)","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:01.112681Z","iopub.execute_input":"2025-05-08T10:28:01.112914Z","iopub.status.idle":"2025-05-08T10:28:01.130250Z","shell.execute_reply.started":"2025-05-08T10:28:01.112890Z","shell.execute_reply":"2025-05-08T10:28:01.129604Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, we want to do graph convolution layer (i.e., message-passing followed by non-linearity) among the `feed` edges. Usually, this can be done by left-multiplying the feature tensor with **some form** of an adjacency matrix. The exact form will determine the pooling (e.g., sum VS average). Let us use the symmetrically-normalized adjacency matrix with self-connections added (by Kipf & Welling, ICLR'17).\n\nWe can compute such a matrix $\\widehat{A}$ as:\n\n$$A_\\textrm{undirected.w.selfconnections} \\leftarrow A + A^\\top + I$$\n\n\n$$D \\leftarrow \\mathbf{1}^\\top A_\\textrm{undirected.w.selfconnections}$$\n\n\n$$\\widehat{A} \\leftarrow D^{-\\frac{1}{2}} (A_\\textrm{undirected.w.selfconnections}) D^{-\\frac{1}{2}} $$\n\nWhich is acheivable by the following code:","metadata":{}},{"cell_type":"code","source":"adj_op_op = implicit.AdjacencyMultiplier(graph_batch_embedded_ops, 'feed')  # op->op\nadj_config = implicit.AdjacencyMultiplier(graph_batch_embedded_ops, 'config')  # nconfig->op\n\nadj_op_op_hat = (adj_op_op + adj_op_op.transpose()).add_eye()\nadj_op_op_hat = adj_op_op_hat.normalize_symmetric()","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:01.131169Z","iopub.execute_input":"2025-05-08T10:28:01.131425Z","iopub.status.idle":"2025-05-08T10:28:01.175001Z","shell.execute_reply.started":"2025-05-08T10:28:01.131393Z","shell.execute_reply":"2025-05-08T10:28:01.174320Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Finally, the message passing can written as:","metadata":{}},{"cell_type":"code","source":"A_times_X = adj_op_op_hat @ combined_features\nprint('A_times_x.shape =', A_times_X.shape)","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:01.175780Z","iopub.execute_input":"2025-05-08T10:28:01.175987Z","iopub.status.idle":"2025-05-08T10:28:03.089293Z","shell.execute_reply.started":"2025-05-08T10:28:01.175965Z","shell.execute_reply":"2025-05-08T10:28:03.088616Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, we put together everything above to write a model class `ResModel` (next), which has a couple more concepts:\n\n1. Adjacency for `\"g_op\"` and `\"g_config\"`, which is used to pool information from all ops and from configurable ops, to the graph level.\n1. Residual connections.\n1. Segment dropout. a forward-pass is computed on the entire graph (but, with `tf.stop_gradient`). Then, another forward pass is computed using only sampled edge-sets (`edgeset_prefix` is set to `\"sampled_\"` by `forward()`).\n\nWithout further ado, `ResModel`:","metadata":{}},{"cell_type":"code","source":"class ResModel(tf.keras.Model):\n    \"\"\"GNN with residual connections.\"\"\"\n\n    def __init__(\n        self, num_configs: int, num_ops: int, op_embed_dim: int = 32,\n        num_gnns: int = 2, mlp_layers: int = 2,\n        hidden_activation: str = 'leaky_relu',\n        hidden_dim: int = 32, reduction: str = 'sum'):\n        super().__init__()\n        self._num_configs = num_configs\n        self._num_ops = num_ops\n        self._op_embedding = _OpEmbedding(num_ops, op_embed_dim)\n        self._prenet = _mlp([hidden_dim] * mlp_layers, hidden_activation)\n        self._gc_layers = []\n        for _ in range(num_gnns):\n            self._gc_layers.append(_mlp([hidden_dim] * mlp_layers, hidden_activation))\n        self._postnet = _mlp([hidden_dim, 1], hidden_activation, use_bias=False)\n\n    def call(self, graph: tfgnn.GraphTensor, training: bool = False):\n        del training\n        return self.forward(graph, self._num_configs)\n\n    def _node_level_forward(\n        self, node_features: tf.Tensor,\n        config_features: tf.Tensor,\n        graph: tfgnn.GraphTensor, num_configs: int,\n        edgeset_prefix='') -> tf.Tensor:\n        adj_op_op = implicit.AdjacencyMultiplier(\n            graph, edgeset_prefix+'feed')  # op->op\n        adj_config = implicit.AdjacencyMultiplier(\n            graph, edgeset_prefix+'config')  # nconfig->op\n\n        adj_op_op_hat = (adj_op_op + adj_op_op.transpose()).add_eye()\n        adj_op_op_hat = adj_op_op_hat.normalize_symmetric()\n\n        x = node_features\n\n        x = tf.stack([x] * num_configs, axis=1)\n        config_features = 100 * (adj_config @ config_features)\n        x = tf.concat([config_features, x], axis=-1)\n        x = self._prenet(x)\n        x = tf.nn.leaky_relu(x)\n\n        for layer in self._gc_layers:\n            y = x\n            y = tf.concat([config_features, y], axis=-1)\n            y = tf.nn.leaky_relu(layer(adj_op_op_hat @ y))\n            x += y\n        return x\n\n    def forward(\n        self, graph: tfgnn.GraphTensor, num_configs: int,\n        backprop=True) -> tf.Tensor:\n        graph = self._op_embedding(graph)\n\n        config_features = graph.node_sets['nconfig']['feats']\n        node_features = tf.concat([\n            graph.node_sets['op']['feats'],\n            graph.node_sets['op']['op_e']\n        ], axis=-1)\n\n        x_full = self._node_level_forward(\n            node_features=tf.stop_gradient(node_features),\n            config_features=tf.stop_gradient(config_features),\n            graph=graph, num_configs=num_configs)\n\n        if backprop:\n            x_backprop = self._node_level_forward(\n                node_features=node_features,\n                config_features=config_features,\n                graph=graph, num_configs=num_configs, edgeset_prefix='sampled_')\n\n            is_selected = graph.node_sets['op']['selected']\n            # Need to expand twice as `is_selected` is a vector (num_nodes) but\n            # x_{backprop, full} are 3D tensors (num_nodes, num_configs, num_feats).\n            is_selected = tf.expand_dims(is_selected, axis=-1)\n            is_selected = tf.expand_dims(is_selected, axis=-1)\n            x = tf.where(is_selected, x_backprop, x_full)\n        else:\n            x = x_full\n\n        adj_config = implicit.AdjacencyMultiplier(graph, 'config')\n\n        # Features for configurable nodes.\n        config_feats = (adj_config.transpose() @ x)\n\n        # Global pooling\n        adj_pool_op_sum = implicit.AdjacencyMultiplier(graph, 'g_op').transpose()\n        adj_pool_op_mean = adj_pool_op_sum.normalize_right()\n        adj_pool_config_sum = implicit.AdjacencyMultiplier(\n            graph, 'g_config').transpose()\n        x = self._postnet(tf.concat([\n            # (A D^-1) @ Features\n            adj_pool_op_mean @ x,\n            # l2_normalize( A @ Features )\n            tf.nn.l2_normalize(adj_pool_op_sum @ x, axis=-1),\n            # l2_normalize( A @ Features )\n            tf.nn.l2_normalize(adj_pool_config_sum @ config_feats, axis=-1),\n        ], axis=-1))\n\n        x = tf.squeeze(x, -1)\n\n        return x\n\n","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:03.090256Z","iopub.execute_input":"2025-05-08T10:28:03.090526Z","iopub.status.idle":"2025-05-08T10:28:03.104743Z","shell.execute_reply.started":"2025-05-08T10:28:03.090501Z","shell.execute_reply":"2025-05-08T10:28:03.104062Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Training loop\n\nCreate a model, objective function, and optimizer.","metadata":{}},{"cell_type":"code","source":"model = ResModel(CONFIGS_PER_GRAPH, layout_npz_dataset.num_ops)\n\nloss = tfr.keras.losses.ListMLELoss()  # (temperature=10)\nopt = tf.keras.optimizers.Adam(learning_rate=1e-3, clipnorm=0.5)\n\nmodel.compile(loss=loss, optimizer=opt, metrics=[\n    tfr.keras.metrics.OPAMetric(name='opa_metric'),\n])\n","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:03.105594Z","iopub.execute_input":"2025-05-08T10:28:03.105804Z","iopub.status.idle":"2025-05-08T10:28:03.159939Z","shell.execute_reply.started":"2025-05-08T10:28:03.105782Z","shell.execute_reply":"2025-05-08T10:28:03.159329Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Train for a few epochs.","metadata":{}},{"cell_type":"code","source":"early_stop = 5  # If validation OPA did not increase in this many epochs, terminate training.\nbest_params = None  # Stores parameters corresponding to best validation OPA, to restore to them after training.\nbest_val_opa = -1  # Tracks best validation OPA\nbest_val_at_epoch = -1  # At which epoch.\nepochs = 1  # Total number of training epochs.\n\nfor i in range(epochs):\n    history = model.fit(\n        layout_train_ds, epochs=1, verbose=1, validation_data=layout_valid_ds,\n        validation_freq=1)\n\n    train_loss = history.history['loss'][-1]\n    train_opa = history.history['opa_metric'][-1]\n    val_loss = history.history['val_loss'][-1]\n    val_opa = history.history['val_opa_metric'][-1]\n    if val_opa > best_val_opa:\n        best_val_opa = val_opa\n        best_val_at_epoch = i\n        best_params = {v.ref: v + 0 for v in model.trainable_variables}\n        print(' * [@%i] Validation (NEW BEST): %s' % (i, str(val_opa)))\n    elif early_stop > 0 and i - best_val_at_epoch >= early_stop:\n      print('[@%i] Best accuracy was attained at epoch %i. Stopping.' % (i, best_val_at_epoch))\n      break\n\n# Restore best parameters.\nprint('Restoring parameters corresponding to the best validation OPA.')\nassert best_params is not None\nfor v in model.trainable_variables:\n    v.assign(best_params[v.ref])","metadata":{"execution":{"iopub.status.busy":"2025-05-08T10:28:03.160728Z","iopub.execute_input":"2025-05-08T10:28:03.160937Z","iopub.status.idle":"2025-05-08T10:29:38.460478Z","shell.execute_reply.started":"2025-05-08T10:28:03.160915Z","shell.execute_reply":"2025-05-08T10:29:38.459741Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Make Submission CSV file for this task","metadata":{}},{"cell_type":"code","source":"import tqdm\n_INFERENCE_CONFIGS_BATCH_SIZE = 50\n\noutput_csv_filename = f'inference_layout_{SOURCE}_{SEARCH}.csv'\nprint('\\n\\n   Running inference on test set ...\\n\\n')\ntest_rankings = []\n\nassert layout_npz_dataset.test.graph_id is not None\nfor graph in tqdm.tqdm(layout_npz_dataset.test.iter_graph_tensors(),\n                       total=layout_npz_dataset.test.graph_id.shape[-1],\n                       desc='Inference'):\n    num_configs = graph.node_sets['g']['runtimes'].shape[-1]\n    all_scores = []\n    for i in tqdm.tqdm(range(0, num_configs, _INFERENCE_CONFIGS_BATCH_SIZE)):\n        end_i = min(i + _INFERENCE_CONFIGS_BATCH_SIZE, num_configs)\n        # Take a cut of the configs.\n        node_set_g = graph.node_sets['g']\n        subconfigs_graph = tfgnn.GraphTensor.from_pieces(\n            edge_sets=graph.edge_sets,\n            node_sets={\n                'op': graph.node_sets['op'],\n                'nconfig': tfgnn.NodeSet.from_fields(\n                    sizes=graph.node_sets['nconfig'].sizes,\n                    features={\n                        'feats': graph.node_sets['nconfig']['feats'][:, i:end_i],\n                    }),\n                'g': tfgnn.NodeSet.from_fields(\n                    sizes=tf.constant([1]),\n                    features={\n                        'graph_id': node_set_g['graph_id'],\n                        'runtimes': node_set_g['runtimes'][:, i:end_i],\n                        'kept_node_ratio': node_set_g['kept_node_ratio'],\n                    })\n            })\n        h = model.forward(subconfigs_graph, num_configs=(end_i - i),\n                          backprop=False)\n        all_scores.append(h[0])\n    all_scores = tf.concat(all_scores, axis=0)\n    graph_id = graph.node_sets['g']['graph_id'][0].numpy().decode()\n    sorted_indices = tf.strings.join(\n        tf.strings.as_string(tf.argsort(all_scores)), ';').numpy().decode()\n    test_rankings.append((graph_id, sorted_indices))\n\nwith tf.io.gfile.GFile(output_csv_filename, 'w') as fout:\n    fout.write('ID,TopConfigs\\n')\n    for graph_id, ranks in test_rankings:\n        fout.write(f'layout:{SOURCE}:{SEARCH}:{graph_id},{ranks}\\n')\nprint('\\n\\n   ***  Wrote', output_csv_filename, '\\n\\n')\n","metadata":{"execution":{"iopub.status.busy":"2025-05-08T11:16:05.354411Z","iopub.execute_input":"2025-05-08T11:16:05.354721Z","iopub.status.idle":"2025-05-08T11:39:24.388439Z","shell.execute_reply.started":"2025-05-08T11:16:05.354693Z","shell.execute_reply":"2025-05-08T11:39:24.387756Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Combine submission CSVs from all layout collections into one CSV file\n\nFinally, after running on all collections, you need to combine the CSV files together (e.g., by concatenation), to prepare the final submission. Specifically, you can modify the constants:\n\n```\nSOURCE = 'xla'  # Can be \"xla\" or \"nlp\"\nSEARCH = 'random'  # Can be \"random\" or \"default\"\n```\n\n(from a few cells ago) and run for all 4 combinations: SOURCE=(\"xla\", \"nlp\") x SEARCH=(\"random\", \"default\"), then combine all inferences into one file:","metadata":{}},{"cell_type":"code","source":"!cat inference_layout_xla_random.csv inference_layout_xla_default.csv inference_layout_nlp_random.csv inference_layout_nlp_default.csv > inference_layout_all.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T11:58:37.001147Z","iopub.execute_input":"2025-05-08T11:58:37.001836Z","iopub.status.idle":"2025-05-08T11:58:37.323896Z","shell.execute_reply.started":"2025-05-08T11:58:37.001801Z","shell.execute_reply":"2025-05-08T11:58:37.322565Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"producing file `\"inference_layout_all.csv\"` that combines all predictions for all layout subcollections. Finally, this file should be combined with the CSV for the tiles collection, as explained next.","metadata":{}},{"cell_type":"markdown","source":"# Tile Training Pipeline\n\nThis section will be written by end of September. We prioritized getting this notebook out, as soon as possible, as the above Layout section is (1) more tricky and (2) most of the score depends on it.","metadata":{}}]}