{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 1px; color:#000000; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #000000\"> Google - Fast or Slow? Predict AI Model Runtime</p>","metadata":{}},{"cell_type":"markdown","source":"## <b>Hey everyone! </b> <br>\nI hope to create a walkthrough type notebook that could simplify things a bit as I myself am learning and working along. This would be an explanatory work for building up my understadning. Hope you can also benefit by going through it and gain better understanding of the competition.\n<br>\nI would recommend everyone to check out the github link as it contains a lot of details about our dataset. <b>https://github.com/google-research-datasets/tpu_graphs</b>","metadata":{}},{"cell_type":"markdown","source":"## <b><span style='color:#2779d2'> 1</span> | INTRODUCTION</b>\n\nThe dataset consists of two compiler optimization collections: layout and tile. Layout configurations control how tensors are laid out in the physical memory, by specifying the dimension order of each input and output of an operation node. A tile configuration controls the tile size of each fused subgraph.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-08-30T19:36:02.507600Z","iopub.execute_input":"2023-08-30T19:36:02.508095Z","iopub.status.idle":"2023-08-30T19:36:02.516793Z","shell.execute_reply.started":"2023-08-30T19:36:02.508059Z","shell.execute_reply":"2023-08-30T19:36:02.515529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a lot of input files to list using the standard input code given by kaggle so we should begin by navigating the directory. <br>\nThe directeries are stored on linux type environment(servers) so for those familiar and unfamiliar with them we can use `!ls {path}` to list the files and navigate through as demonstarted below.","metadata":{}},{"cell_type":"code","source":"!ls ../input/predict-ai-model-runtime # The sample submission and and our dataset basically\n!ls ../input/predict-ai-model-runtime/npz_all # These are the npz files mentioned in the data description of the competition\n!ls ../input/predict-ai-model-runtime/npz_all/npz # Here they are further seperated into layout and tile","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:58:00.687193Z","iopub.execute_input":"2023-08-30T19:58:00.687733Z","iopub.status.idle":"2023-08-30T19:58:03.992795Z","shell.execute_reply.started":"2023-08-30T19:58:00.687684Z","shell.execute_reply":"2023-08-30T19:58:03.991124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dataset consists of two compiler optimization collections: layout and tile. Layout configurations control how tensors are laid out in the physical memory, by specifying the dimension order of each input and output of an operation node. A tile configuration controls the tile size of each fused subgraph.","metadata":{}},{"cell_type":"code","source":"!ls ../input/predict-ai-model-runtime/npz_all/npz/layout\n\n!ls ../input/predict-ai-model-runtime/npz_all/npz/layout/nlp\n!ls ../input/predict-ai-model-runtime/npz_all/npz/layout/nlp/default\n!ls ../input/predict-ai-model-runtime/npz_all/npz/layout/nlp/default/test # the npz files for testing\n\n#!ls ../input/predict-ai-model-runtime/npz_all/npz/layout/xla\n# The xla follows the same distribution as nlp","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:59:38.375197Z","iopub.execute_input":"2023-08-30T19:59:38.376060Z","iopub.status.idle":"2023-08-30T19:59:42.758304Z","shell.execute_reply.started":"2023-08-30T19:59:38.376011Z","shell.execute_reply":"2023-08-30T19:59:42.756777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" Each layout data file is named as followed : <br>\n{source}: xla or nlp <br>\n{search}: default or random <br>\n{split}: train, valid, or test ","metadata":{}},{"cell_type":"code","source":"!ls ../input/predict-ai-model-runtime/npz_all/npz/tile/xla","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:00:42.676448Z","iopub.execute_input":"2023-08-30T20:00:42.676931Z","iopub.status.idle":"2023-08-30T20:00:43.777272Z","shell.execute_reply.started":"2023-08-30T20:00:42.676889Z","shell.execute_reply":"2023-08-30T20:00:43.775556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The tile folder contain basically the train-test-split files for xla. They further contain the npz files.","metadata":{}},{"cell_type":"markdown","source":"## <b><span style='color:#2779d2'> 2</span> | Exploring the data</b>","metadata":{}},{"cell_type":"code","source":"os.listdir('/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/nlp/default/train')[0]","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:33:11.992172Z","iopub.execute_input":"2023-08-30T20:33:11.992595Z","iopub.status.idle":"2023-08-30T20:33:12.000116Z","shell.execute_reply.started":"2023-08-30T20:33:11.992562Z","shell.execute_reply":"2023-08-30T20:33:11.999012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One of the models in the train set for nlp, lets take a look at its file structure.","metadata":{}},{"cell_type":"code","source":"data = '/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/nlp/default/train/small_bert_bert_en_uncased_L-2_H-768_A-12_batch_size_16_test.npz'\nnpz_file = np.load(data) #standard np.load() to read the data.\nnpz_file.files # This is the basic structure of a file","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:37:31.464672Z","iopub.execute_input":"2023-08-30T20:37:31.465760Z","iopub.status.idle":"2023-08-30T20:37:31.479426Z","shell.execute_reply.started":"2023-08-30T20:37:31.465720Z","shell.execute_reply":"2023-08-30T20:37:31.478503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Further lets view the contents of a npz file.","metadata":{}},{"cell_type":"code","source":"for x in npz_file.files :\n    print('\\033[1m' + x + \"\\033[0m\" + '\\n')\n    print(npz_file[x])\n    print('-'*20)   ","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:42:18.742617Z","iopub.execute_input":"2023-08-30T20:42:18.742974Z","iopub.status.idle":"2023-08-30T20:42:19.289906Z","shell.execute_reply.started":"2023-08-30T20:42:18.742947Z","shell.execute_reply":"2023-08-30T20:42:19.288834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <b><span style='color:#16C2D5'></span>THANK YOU !!</b>\n\n<br>\n    \nPlease comment your thoughts and I'd greatly appreciate any upvotes!\n\n<br>\n    \nThis is for todays update and I will keep improving the notebook.     ＼( °▾° )／\n<br>\n\n<img src = \"https://media.makeameme.org/created/please-upvote-and.jpg\" width=\"400\" height=\"100\">\n","metadata":{}}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}