{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Hi all!\n\nFirst of all, I wish you all the best of luck for this competition, it seems the guys from Google have gifted us a brand new very interesting problem to look at.\n\nThis is a WIP for getting some insights about the data for the new competition [*Google - Fast or Slow? Predict AI Model Runtime*](https://www.kaggle.com/competitions/predict-ai-model-runtime/) in Kaggle. Hopefully after going through this notebook you will understand the competition and the dataset much better than before.","metadata":{}},{"cell_type":"markdown","source":"# Loading and exploring the data","metadata":{}},{"cell_type":"markdown","source":"We can see from a first glance that the data in this competition comes presented in a more complicated way than usual. To get some information about it we can go through the `README` file of [the dataset GitHub repo](https://github.com/google-research-datasets/tpu_graphs).\n\nJust as an introduction, we can keep in mind the following extract from that `README` file:\n> The dataset consists of two compiler optimization collections: layout and tile. Layout configurations control how tensors are laid out in the physical memory, by specifying the dimension order of each input and output of an operation node. A tile configuration controls the tile size of each fused subgraph.","metadata":{}},{"cell_type":"code","source":"# Importing the necessary libraries\n\nimport numpy as np\nimport pandas as pd\nimport os","metadata":{"execution":{"iopub.status.busy":"2023-08-30T18:17:47.928769Z","iopub.execute_input":"2023-08-30T18:17:47.929296Z","iopub.status.idle":"2023-08-30T18:17:47.963335Z","shell.execute_reply.started":"2023-08-30T18:17:47.929250Z","shell.execute_reply":"2023-08-30T18:17:47.962198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/input'):\n    if len(filenames) != 0:\n        if filenames[0] != \"sample_submission.csv\":\n            avg = np.array([os.path.getsize(os.path.join(dirname, filename)) for filename in filenames]).mean()\n            print(dirname, len(os.listdir(dirname)))\n            print(\"Size: {:.3f} KB\".format(avg/1024))\n            ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-08-30T18:17:47.965847Z","iopub.execute_input":"2023-08-30T18:17:47.967269Z","iopub.status.idle":"2023-08-30T18:17:50.015261Z","shell.execute_reply.started":"2023-08-30T18:17:47.967213Z","shell.execute_reply":"2023-08-30T18:17:50.013359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When looking at the data, we can see the following details:\n- The tile collection does not have any subdivisions inside it except the ones required for working on the competition (train, validation and test).\n- On the other hand, the layout collection is split into nlp and xla (which are further split between *random* and *default*).\n- The tile files are, on average, much smaller in size than the layout files.\n- The layout file sizes are not uniform inside the same categories, resulting in pathologies like the test set of the *nlp/default* subdivission having drastically smaller files on average than both the train and validation sets. This might have to be taken into account during the competition (the train set might not be representing new observations faithfully).\n- Some of the train-val-test splits in the layout collection result in very small subsets, making these subcategories very prone to imbalances in the existing data.\n\nLet's try to get some information about why this could be the case, first by looking at the different characteristics between the tile and the layout collections:","metadata":{}},{"cell_type":"code","source":"# We take a tile .npz file and check its structure\nexample = np.load(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/tile/xla/train/retinanet.4x4.fp32_-431a58cc30e72ec6.npz\")\nexample.files","metadata":{"execution":{"iopub.status.busy":"2023-08-30T18:17:50.016915Z","iopub.execute_input":"2023-08-30T18:17:50.017290Z","iopub.status.idle":"2023-08-30T18:17:50.032796Z","shell.execute_reply.started":"2023-08-30T18:17:50.017262Z","shell.execute_reply":"2023-08-30T18:17:50.031152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are the files that each .npz file from the tile colection contains (as mentioned [in the repo](https://github.com/google-research-datasets/tpu_graphs#tiles-collection-npz-files)).\n\nFrom the repo we get that:\n\n> Suppose a `.npz` file stores a graph (representing a kernel) with `n` nodes and `m` edges. In addition, suppose we compile the graph with `c` different configurations, and run each on a TPU. Crucially, the configuration is at the graph-level. Then, the `.npz` file stores the following dictionary (can be loaded with `d = dict(np.load(\"npz/tile/xla/train/<pick 1>.npz\"))`):\n>   - Key `node_feat`: contains `float32` matrix with shape `(n, 140)`. The `u`-th row contains the feature vector for node `u` < `n` (please see Subsection \"Node Features\", below). Nodes are ordered topologically.\n>   - Key `node_opcode` contains `int32` vector with shape `(n, )`. The `u`-th entry stores the op-code for node u (please see the mapping of opcode to instruction name here).\n>   - Key `edge_index` contains `int32` matrix with shape `(m, 2)`. If entry `i` is = `[u, v]` (where `0 <= u, v < n`), then there is a directed edge from node `u` to node `v`, where `u` consumes the output of `v`.\n>   - Key `config_feat` contains `float32` matrix with shape `(c, 24)` with row `j` containing the (graph-level) configuration feature vector (please see Subsection \"Tile Config Features\").\n>   - Keys `config_runtime` and `config_runtime_normalizers`: both are `int64` vectors of length `c`. Entry `j` stores the runtime (in nanoseconds) of the given graph compiled with configuration `j` and a default configuration, respectively. Samples from the same graph may have slightly different `config_runtime_normalizers` because they are measured from different runs on multiple machines.\n> Finally, for the tile collection, your job is to predict the indices of the best configurations (i.e., ones leading to the smallest `d[\"config_runtime\"] / d[\"config_runtime_normalizers\"]`).","metadata":{}},{"cell_type":"markdown","source":"We can even check if every file has the same structure:","metadata":{}},{"cell_type":"code","source":"basic_structure = ['node_feat',\n 'node_opcode',\n 'edge_index',\n 'config_feat',\n 'config_runtime',\n 'config_runtime_normalizers']\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    flag = False\n\n    for filename in filenames:\n        if filename != \"sample_submission.csv\":\n            if np.load(os.path.join(dirname, filename)).files != basic_structure and not flag:\n                print(dirname)\n                print(np.load(os.path.join(dirname, filename)).files)\n                flag = True","metadata":{"execution":{"iopub.status.busy":"2023-08-30T18:17:50.035560Z","iopub.execute_input":"2023-08-30T18:17:50.036011Z","iopub.status.idle":"2023-08-30T18:18:24.634278Z","shell.execute_reply.started":"2023-08-30T18:17:50.035972Z","shell.execute_reply":"2023-08-30T18:18:24.632658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see that .npz files inside the layout collection contain a different set of files, as we could expect from reading [the `README` inside the repo](https://github.com/google-research-datasets/tpu_graphs#layout-collections-npz-files)).\n\nFrom the repo itself we get:\n\n> Suppose a .npz file stores a graph (representing the entire program) with `n` nodes and `m` edges. In addition, suppose we compile the graph with `c` different configurations, and run each on a TPU. Crucially, the configuration is at the node-level. Suppose that `nc` of the `n` nodes are configurable. Then, the .npz file stores the following dictionary (can be loaded with, e.g., `d = dict(np.load(\"npz/layout/xla/random/train/unet3d.npz\")))`:\n>   - Keys `node_feat`, `node_opcode`, `edge_index`, are like above.\n>   - Key `node_config_ids` contains `int32` vector with shape `(nc, )` and every entry is in `{0, 1, ..., n - 1}` i.e. indicating the indices of the configurable nodes. For these nodes, they can have an additional feature vector that instructs the compiler (described next).\n>   - Key `node_config_feat` contains `float32` tensor with shape `(c, nc, 18)`. Entry `[j, k]` gives an 18-dimensional vector describing the configuration features for node `d[\"node_config_ids\"][k]` for the `j`-th run (please see Subsection \"Layout Config Features\", below).\n>   - Key `config_runtime` contains `int32` vector with shape `(c, )` where the `j`-th entry contains the runtime of the `j`-th run (i.e., when nodes are configured with `d[\"node_config_feat\"][j]`).","metadata":{}},{"cell_type":"markdown","source":"We can now load all the data in dataframes to make working with it easier","metadata":{}},{"cell_type":"code","source":"def load_df(directory):\n    splits = [\"train\", \"valid\", \"test\"]\n    dfs = dict()\n    \n    for split in splits:\n        path = os.path.join(directory, split)\n        files = os.listdir(path)\n        list_df = []\n        \n        for file in files:\n            list_df.append(dict(np.load(os.path.join(path,file))))\n        dfs[split] = pd.DataFrame.from_dict(list_df)\n    return dfs","metadata":{"execution":{"iopub.status.busy":"2023-08-30T18:18:24.636670Z","iopub.execute_input":"2023-08-30T18:18:24.637253Z","iopub.status.idle":"2023-08-30T18:18:24.645894Z","shell.execute_reply.started":"2023-08-30T18:18:24.637191Z","shell.execute_reply":"2023-08-30T18:18:24.644684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If you try to run the following cell completely uncommented the Kaggle kernel will run out of memory and crash, so we will have to study the datasets individually","metadata":{}},{"cell_type":"code","source":"\n\ntile_xla = load_df(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/tile/xla/\")\n\n#layout_nlp_random = load_df(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/nlp/random/\")\n#layout_nlp_default = load_df(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/nlp/default/\")\n#layout_xla_random = load_df(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/xla/random/\")\n#layout_xla_random = load_df(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/xla/default/\")","metadata":{"execution":{"iopub.status.busy":"2023-08-30T18:18:24.647325Z","iopub.execute_input":"2023-08-30T18:18:24.648314Z","iopub.status.idle":"2023-08-30T18:18:57.199551Z","shell.execute_reply.started":"2023-08-30T18:18:24.648230Z","shell.execute_reply":"2023-08-30T18:18:57.197859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"markdown","source":"## Tile collection","metadata":{}},{"cell_type":"code","source":"tile_xla[\"train\"]","metadata":{"execution":{"iopub.status.busy":"2023-08-30T18:19:11.705281Z","iopub.execute_input":"2023-08-30T18:19:11.705825Z","iopub.status.idle":"2023-08-30T18:19:13.172458Z","shell.execute_reply.started":"2023-08-30T18:19:11.705778Z","shell.execute_reply":"2023-08-30T18:19:13.170955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}