{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":58266,"databundleVersionId":6641124,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport glob","metadata":{"_uuid":"0a6a8d9a-4685-48e9-ba3e-dd4cac51b669","_cell_guid":"933abec1-4a68-4a00-a894-7cef464df0a2","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:03.647503Z","iopub.execute_input":"2025-04-27T09:09:03.647891Z","iopub.status.idle":"2025-04-27T09:09:04.84163Z","shell.execute_reply.started":"2025-04-27T09:09:03.647855Z","shell.execute_reply":"2025-04-27T09:09:04.84047Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Schematic: Proposed Flow: Start with Layout - Capture Node Features , their opcodes, their edge connectivity  -> concatenate to supply as input to GNN -> get a cumulative runtime for reconfigurable nodes.\n\nWhat Layout has: Node dependencies , Opcode for each node to note its functionality, Node features which are its structures with inpute and output layer configuration, IDs of Nodes which can be configured, features of the nodes which can be experimented with and altered, configuration runtimes to denote the runtime taken for each alteration.\n\nWhat Tile has: Node dependencies , Opcode of each node in a fused subgraph to note their functionality, Rest the same with differences between runtime normalizers as we are observing this at kernel level. Config Features have an extra sum/product differentiators.\n\nHow this data is arranged: Layout has each of alexnet, bert, audio, video data in XLA. Tile has a follow up of these features, different tiles arranged sequentially. To process this, use regex to first optimize tile features. \n\nWill we use GCN or someother architecture?\nGCN would definitely facilitate, provided we make the model predictions variable. Assign data classes to be equal to number of node features. Produce a mathematical equation that computes the runtime, i.e, integrates the multiple dependency runtimes. In channels also should be 140.\n\nApproach: Embed the node opcodes with its features which will be passed as input. Get edge conditions, i.e, node connections and flow.\n\nFigure out: What determines the runtime of each configuration. In other words, output how many times each Opcode's node was used, alongwith info of the connections it has in the value the output node would carry. \nExample: If Addition node was used 10 times, and is connected with max, mul, let it have the value 10+(number of nodes with similar dimension values)*0.5 for max+ (number of nodes with similar dimension values)*2. Having a rough estimate of runtime and a statistial parameter that relates to the actual and predicted value will help.\nDetails on embedding runtime and configuration features: We need an intermediate layer of gcn with 18 input nodes and 1 output node. This embedding should have the opcode of the particular configurable node i.\n\nAlgorithm:\n 1. Determine the runtime of the operation associated with each opcode.\n 2. Multiply 1/100 * opcode * runtime. Name this result oper_run_time.\n 3. Define GRU with 140 input nodes, 18 intermediate nodes and 1 output node.\n 4. Multiply oper_run_time * node_feat to produce consolidated_feat.\n 5. For each connected edge j of node i:\n          set inp_cons = 2 if input is being consumed 1 if it is being given out.\n          consolidated_feat[i] = consolidated_feat[i]+ inp_cons * consolidated_feat[j] \n          i=i+1\n 6. Add all of the consolidated_feat vectors.\n 7. Check in config feat if this node exists and set consolidated_feat[nfi]+=1 else consolidated_feat [nfi] = consolidated_feat[nfi]\n 8. Pass consolidated_feat[nfi] as input to each input gcn node gi = nfi.\n 9. Let intermediate layer 1 have the summation function.\n 10. For each configurable node ci, pass consolidated_feat[ci] * node_config_feat[i] + output of the summation function for each configuration iteration.\n 11. Train the model with runtime/runtime_normalizers of each configuration as output.","metadata":{"_uuid":"32019452-92ae-4a63-b3ef-e863cfde18f5","_cell_guid":"793d5806-ba47-41f5-8240-5b3a675be9d0","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"pip install torch","metadata":{"_uuid":"6b020fe1-fac2-453d-bbd1-4d59d15a6438","_cell_guid":"7839aa80-9ad0-4e20-b677-cfa0b4a4c42e","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:04.842596Z","iopub.execute_input":"2025-04-27T09:09:04.843232Z","iopub.status.idle":"2025-04-27T09:09:10.908874Z","shell.execute_reply.started":"2025-04-27T09:09:04.843188Z","shell.execute_reply":"2025-04-27T09:09:10.907499Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nfrom torch.utils.data import DataLoader","metadata":{"_uuid":"5d687e04-d764-4ab1-b885-cdb97deafafb","_cell_guid":"3b56bba8-f767-48eb-9dad-f337be751f11","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:10.910163Z","iopub.execute_input":"2025-04-27T09:09:10.910499Z","iopub.status.idle":"2025-04-27T09:09:15.578576Z","shell.execute_reply.started":"2025-04-27T09:09:10.910469Z","shell.execute_reply":"2025-04-27T09:09:15.577185Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class GRURuntimePredictor(nn.Module):\n    def __init__(self, input_dim1=140, hidden_dim1=18, gru_input_dim=18, gru_hidden_dim=32, output_dim=1):\n        super(GRURuntimePredictor, self).__init__()\n        self.fc1 = nn.Linear(input_dim1, hidden_dim1)\n        self.gru = nn.GRU(input_size=gru_input_dim, hidden_size=gru_hidden_dim, batch_first=True)\n        self.fc2 = nn.Linear(gru_hidden_dim, output_dim)\n\n    def forward(self,x,y):\n        x_gru_input = torch.relu(x) \n        x_gru_input= x_gru_input.unsqueeze(-1)\n        y_gru_input = torch.relu(y) \n        cons_gru_input= x_gru_input*y_gru_input\n        gru_output, _ = self.gru(cons_gru_input)\n        final_output = self.fc2(gru_output) \n        pred_output = torch.sum(final_output)\n        pred_output= pred_output.unsqueeze(-1)\n        return pred_output","metadata":{"_uuid":"2de70e72-40da-4de8-8fad-715d4c71115f","_cell_guid":"81784995-f10c-40cb-bd0e-b19783a91006","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:15.581314Z","iopub.execute_input":"2025-04-27T09:09:15.581831Z","iopub.status.idle":"2025-04-27T09:09:15.588273Z","shell.execute_reply.started":"2025-04-27T09:09:15.581799Z","shell.execute_reply":"2025-04-27T09:09:15.587089Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def runtime_of_operations_from_opcode(Opcode):\n    runtimes = [ 3,4,3,4, 1, 6, 9, 6, 6, 2, 6,3, 7, 2, 2, 7, 1, 7, 2,3,8,8,\n     8,    2,    5,    5,    6,    1,    9,    2,    4,   14,   21,\n    43,   34,   12,   40,   41,   43,   49,   42,   28,   36,   30,\n    24,   16,   36,   24,   26,   33,   28,   10,   36,   15,   36,\n    11,   10,   17,   15,   31,   16,  305,  219,  160,  243,  214,\n   422,  223,  122,  198,   97,  421,  100,  229,  435,  241,  482,\n   417,  241,  224,  176,  164,  256,   63,  188,  152,  461,  185,\n   297,  167,   60,  782, 4802, 4935, 3594, 3657,  613, 3482, 2442,\n  3008, 2300, 4915, 1195, 1717, 4389, 4474, 4299, 4373,  946, 1662,\n  4832, 2585, 4271, 4712, 3309,  589, 2386, 3385, 2077, 4332]\n    return runtimes[Opcode-1]","metadata":{"_uuid":"8bf64004-c032-4bd4-b439-70e914bd1a78","_cell_guid":"04fe41a5-7af3-44df-b3b6-addf12e35f54","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:15.590237Z","iopub.execute_input":"2025-04-27T09:09:15.590634Z","iopub.status.idle":"2025-04-27T09:09:15.610288Z","shell.execute_reply.started":"2025-04-27T09:09:15.590592Z","shell.execute_reply":"2025-04-27T09:09:15.609249Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"d1= dict(np.load(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/xla/default/train/alexnet_train_batch_32.npz\"))","metadata":{"_uuid":"ffeff9f8-3ec6-40fb-bb42-911e32e2c13d","_cell_guid":"c541c4f0-f402-4f24-8b18-c2b99b219df7","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:15.611354Z","iopub.execute_input":"2025-04-27T09:09:15.611722Z","iopub.status.idle":"2025-04-27T09:09:15.897894Z","shell.execute_reply.started":"2025-04-27T09:09:15.61169Z","shell.execute_reply":"2025-04-27T09:09:15.89685Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def opcode_based_runtime_pred(d_xla,num_nodes):\n    for i in range(num_nodes):\n        oper_run_time = runtime_of_operations_from_opcode(d_xla[\"node_opcode\"][i])\n        d_xla[\"node_feat\"][i] = 0.0000001 * d_xla[\"node_opcode\"][i]* oper_run_time * d_xla[\"node_feat\"][i]","metadata":{"_uuid":"16753950-b0d7-4302-9c5b-18e35b9d64c2","_cell_guid":"e5ad155b-4df7-47c3-971f-9d8d4beb878e","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:15.898781Z","iopub.execute_input":"2025-04-27T09:09:15.899187Z","iopub.status.idle":"2025-04-27T09:09:15.904764Z","shell.execute_reply.started":"2025-04-27T09:09:15.899148Z","shell.execute_reply":"2025-04-27T09:09:15.903285Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def find_row(matrix,target):\n    return [i for i, row in enumerate(matrix) if target in row]","metadata":{"_uuid":"15852a08-aa6c-4a65-9927-42ff8a6a5108","_cell_guid":"0eb8f70f-0c54-46e1-84af-b1c18458b318","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:15.905831Z","iopub.execute_input":"2025-04-27T09:09:15.906203Z","iopub.status.idle":"2025-04-27T09:09:15.936037Z","shell.execute_reply.started":"2025-04-27T09:09:15.906171Z","shell.execute_reply":"2025-04-27T09:09:15.934534Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def edge_dependency_embedding(d_xla,num_nodes):\n    for i in range(num_nodes):\n        occur_node = find_row(d_xla[\"edge_index\"],i)\n        if len(occur_node)>0:\n            for j in range(len(occur_node)):\n                if d_xla[\"edge_index\"][j][0] == i:\n                    d_xla[\"node_feat\"][i] = (d_xla[\"node_feat\"][d_xla[\"edge_index\"][j][1]]*2)+d_xla[\"node_feat\"][i]\n                elif d_xla[\"edge_index\"][j][1] == i:\n                    d_xla[\"node_feat\"][i] = d_xla[\"node_feat\"][d_xla[\"edge_index\"][j][0]]+d_xla[\"node_feat\"][i]","metadata":{"_uuid":"612506c7-a72c-41f8-9e14-d0ceae044f74","_cell_guid":"4473be70-a528-477d-9b05-0d724900a98f","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:15.937358Z","iopub.execute_input":"2025-04-27T09:09:15.937751Z","iopub.status.idle":"2025-04-27T09:09:15.955183Z","shell.execute_reply.started":"2025-04-27T09:09:15.937712Z","shell.execute_reply":"2025-04-27T09:09:15.954061Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def check_if_configurable(d_xla):\n    for i in d_xla[\"node_config_ids\"]:\n        d_xla[\"node_feat\"][i]+=1","metadata":{"_uuid":"1f336270-826f-46a6-aaa2-b96f1b888240","_cell_guid":"946afa78-011d-4faa-aec7-3cd4ed5ce53e","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:15.956672Z","iopub.execute_input":"2025-04-27T09:09:15.9571Z","iopub.status.idle":"2025-04-27T09:09:15.981189Z","shell.execute_reply.started":"2025-04-27T09:09:15.957063Z","shell.execute_reply":"2025-04-27T09:09:15.979731Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def generate_consolidated_node_features(d_xla,num_nodes):\n    consolidated_node_feat= []\n    for i in range(140):\n        cons_feat=0\n        for j in range(num_nodes):\n            cons_feat+=d_xla[\"node_feat\"][j][i]\n        consolidated_node_feat.append(cons_feat)\n    consolidated_node_feat=torch.tensor(consolidated_node_feat)\n    consolidated_node_feat=consolidated_node_feat.float()\n    return consolidated_node_feat","metadata":{"_uuid":"65cfe352-6bb8-4465-a21d-eb48426ed8ca","_cell_guid":"7cec6051-e2ef-4ca3-a697-01755490e3c8","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:15.98255Z","iopub.execute_input":"2025-04-27T09:09:15.98289Z","iopub.status.idle":"2025-04-27T09:09:16.005442Z","shell.execute_reply.started":"2025-04-27T09:09:15.982856Z","shell.execute_reply":"2025-04-27T09:09:16.004018Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def configuration_convolution(d_xla,num_configurable_nodes):\n    node_config_feat = []\n    for i in range(num_configurable_nodes):\n        node_config_cons_feat = []\n        cons_feat=0\n        for j in range(18):\n            for k in range(num_configurable_nodes):\n                cons_feat+=(d_xla[\"node_feat\"][d_xla[\"node_config_ids\"][k]])*(d_xla[\"node_config_feat\"][i][k][j])\n            node_config_cons_feat.append(cons_feat)\n        node_config_feat.append(node_config_cons_feat) \n    return node_config_feat","metadata":{"_uuid":"26da1eed-ea52-4949-81f7-d6071d2ce12a","_cell_guid":"8a1ddb4c-621b-4ebb-9845-e7d87d558b8a","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:16.00719Z","iopub.execute_input":"2025-04-27T09:09:16.007627Z","iopub.status.idle":"2025-04-27T09:09:16.027853Z","shell.execute_reply.started":"2025-04-27T09:09:16.007583Z","shell.execute_reply":"2025-04-27T09:09:16.026556Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def data_transformation(d_xla,num_nodes,num_configurable_nodes):\n    node_config_feat_trans = []\n    config_runtime_trans = []\n    opcode_based_runtime_pred(d_xla,num_nodes)\n    edge_dependency_embedding(d_xla,num_nodes)\n    check_if_configurable(d_xla)\n    cons_node_feat = generate_consolidated_node_features(d_xla,num_nodes)\n    node_config_feature = configuration_convolution(d_xla,num_configurable_nodes)\n    for i in range(num_configurable_nodes):\n        node_config_feat_tensor=torch.tensor(node_config_feature[i])\n        node_config_feat_tensor=torch.t(node_config_feat_tensor)\n        config_runtime_tensor=torch.tensor(0.0001* d_xla[\"config_runtime\"][i]).float()\n        config_runtime_tensor=config_runtime_tensor.unsqueeze(-1)\n        node_config_feat_trans.append(node_config_feat_tensor)\n        config_runtime_trans.append(config_runtime_tensor)\n    return cons_node_feat,node_config_feat_trans,config_runtime_trans","metadata":{"_uuid":"608c4097-47b4-4185-870d-0396eb2c601d","_cell_guid":"4fe1cea1-6603-4e57-9fd6-279ea0f8f9a7","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:16.030844Z","iopub.execute_input":"2025-04-27T09:09:16.031204Z","iopub.status.idle":"2025-04-27T09:09:16.05466Z","shell.execute_reply.started":"2025-04-27T09:09:16.031171Z","shell.execute_reply":"2025-04-27T09:09:16.052932Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def model_training(d_xla,num_nodes,num_configurable_nodes,model):\n    optimizer = torch.optim.Adam(model.parameters(),lr=0.8)\n    criterion = torch.nn.MSELoss()\n    cons_node_feat,node_config_feat_tens,config_runtime_tens = data_transformation(d_xla,num_nodes,num_configurable_nodes)\n    cons_loss = 1\n    for i in range(num_configurable_nodes):\n        pred1 = model(cons_node_feat,node_config_feat_tens[i])\n        loss1 = criterion(pred1,config_runtime_tens[i]) * 0.00000001\n        optimizer.zero_grad()\n        loss1.backward()\n        optimizer.step()\n        cons_loss+=loss1\n    return cons_loss","metadata":{"_uuid":"58b9a7ca-1595-4b74-82ef-37248c0d9d49","_cell_guid":"d531d192-2db6-40d4-94b2-0cc7ed54db06","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:16.0562Z","iopub.execute_input":"2025-04-27T09:09:16.056882Z","iopub.status.idle":"2025-04-27T09:09:16.084683Z","shell.execute_reply.started":"2025-04-27T09:09:16.056832Z","shell.execute_reply":"2025-04-27T09:09:16.083248Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = GRURuntimePredictor()\nxla_dataset = glob.glob(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/xla/default/train/*.npz\")","metadata":{"_uuid":"f7e05c09-af9c-4a56-b915-9483986b7a52","_cell_guid":"b9d464b8-ea1a-44e9-908d-9b75dd520aab","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:16.086086Z","iopub.execute_input":"2025-04-27T09:09:16.086471Z","iopub.status.idle":"2025-04-27T09:09:16.150837Z","shell.execute_reply.started":"2025-04-27T09:09:16.086423Z","shell.execute_reply":"2025-04-27T09:09:16.149759Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len_dataset = len(xla_dataset)\ncons_train_loss = 0\ntrain_loss_per_iter = 100","metadata":{"_uuid":"194effaa-b1af-4880-928d-3b40ddf46cc4","_cell_guid":"b44eef7c-144d-4ea8-acca-ad8c7775b4ea","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:16.151755Z","iopub.execute_input":"2025-04-27T09:09:16.152118Z","iopub.status.idle":"2025-04-27T09:09:16.157216Z","shell.execute_reply.started":"2025-04-27T09:09:16.152088Z","shell.execute_reply":"2025-04-27T09:09:16.155701Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for i in range(len_dataset):\n    xla_dataset_dict = dict(np.load(xla_dataset[i]))\n    node_count = len(xla_dataset_dict[\"node_opcode\"])\n    config_count = len(xla_dataset_dict[\"node_config_ids\"])\n    loss_train = model_training(xla_dataset_dict,node_count,config_count,model)\n    print(\"loss train\"+ str(loss_train))","metadata":{"_uuid":"987b4cbe-4717-402f-9eb9-75eba2680a17","_cell_guid":"de66ce75-7097-41c8-81b2-121a11325513","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-27T09:09:16.158366Z","iopub.execute_input":"2025-04-27T09:09:16.15877Z","iopub.status.idle":"2025-04-27T11:17:19.608088Z","shell.execute_reply.started":"2025-04-27T09:09:16.15873Z","shell.execute_reply":"2025-04-27T11:17:19.606427Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null}]}