{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport glob\nimport matplotlib.pyplot as plt\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-09-28T02:19:46.163759Z","iopub.execute_input":"2023-09-28T02:19:46.164273Z","iopub.status.idle":"2023-09-28T02:19:46.170376Z","shell.execute_reply.started":"2023-09-28T02:19:46.164215Z","shell.execute_reply":"2023-09-28T02:19:46.169262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Getting the hang of the keys in the files\n\nThere are 2 different types of tracks: \n- layout\n- tile\n\nIn layout folder there are 2 formats:\n- **nlp**: based on the previous knowledge that xla meanst TPU I am guessing that nlp would mean normal GPU setup? \n- **xla**: since xla usually refers to the TPU setup I am guessing this means TPU\n\nthese is only xla hardware for tile track.\n\nWe start with nlp format in layout tracks and list all of them. \n\nThen we see what does it look like inside the first file. ","metadata":{}},{"cell_type":"code","source":"files = list(glob.glob(\"/kaggle/input/predict-ai-model-runtime/npz_all/npz/layout/nlp/random/train/*.npz\"))\n\ndata = np.load(file=files[0])\nfor key in data.keys():\n    print(key, data[key].shape)","metadata":{"execution":{"iopub.status.busy":"2023-09-28T02:19:46.172439Z","iopub.execute_input":"2023-09-28T02:19:46.172805Z","iopub.status.idle":"2023-09-28T02:19:46.839766Z","shell.execute_reply.started":"2023-09-28T02:19:46.172775Z","shell.execute_reply":"2023-09-28T02:19:46.838518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Getting some stats from the target labels\n\nSince some of the files are very large I define a function to load only 1 file at a time and generate the stats I need and close it\n\nthe runtime_plotter as the name suggests is for plotting reasons","metadata":{}},{"cell_type":"code","source":"def get_runtime_stats(file: dict):\n    data = np.load(file)\n    runtime = data['config_runtime']\n    return {\n            \"min\":np.min(runtime), \n            \"max\":np.max(runtime), \n            \"avg\":np.mean(runtime),\n            \"range\":  (np.max(runtime) - np.min(runtime)) / np.min(runtime)\n    }\n\nruntime_plotter = pd.DataFrame(map(get_runtime_stats, files))","metadata":{"execution":{"iopub.status.busy":"2023-09-28T02:19:46.841449Z","iopub.execute_input":"2023-09-28T02:19:46.842493Z","iopub.status.idle":"2023-09-28T02:19:47.857738Z","shell.execute_reply.started":"2023-09-28T02:19:46.842450Z","shell.execute_reply":"2023-09-28T02:19:47.856643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"figure, (ax1, ax2, ax3, ax4) = plt.subplots(1,4)\n\nfigure.set_size_inches(15, 5)\nax1.hist(runtime_plotter['avg'])\nax2.hist(runtime_plotter['min'])\nax3.hist(runtime_plotter['max'])\nax4.hist(runtime_plotter['range'])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-28T02:19:47.860618Z","iopub.execute_input":"2023-09-28T02:19:47.861083Z","iopub.status.idle":"2023-09-28T02:19:48.653398Z","shell.execute_reply.started":"2023-09-28T02:19:47.861038Z","shell.execute_reply":"2023-09-28T02:19:48.652195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By looking at the plots it seems that the runtimes of the models themselves are close to each other (looking at the range which says that in most cases the max is not more than 2x of min)\n\nHowever the values themselves are a little bit diverse and very high! We need to come up with a ratio that keeps the values close to each other and preferrebly on the low side. \n\nHow about we normalize the values by an L2 norm? Since only the rank of the values is important to us at the end it should be ok to do so right? ","metadata":{}},{"cell_type":"code","source":"def get_runtime_normalized(file: dict):\n    data = np.load(file)\n    runtime = data['config_runtime'].astype(np.float32)\n    \n    runtime /= np.linalg.norm(runtime)\n        \n    return {\n            \"min\":np.min(runtime), \n            \"max\":np.max(runtime), \n            \"avg\":np.mean(runtime),\n            \"range\":  (np.max(runtime) - np.min(runtime)) / np.min(runtime)\n    }\n\nruntime_plotter = pd.DataFrame(map(get_runtime_normalized, files))","metadata":{"execution":{"iopub.status.busy":"2023-09-28T02:19:48.654971Z","iopub.execute_input":"2023-09-28T02:19:48.655331Z","iopub.status.idle":"2023-09-28T02:19:49.666213Z","shell.execute_reply.started":"2023-09-28T02:19:48.655299Z","shell.execute_reply":"2023-09-28T02:19:49.665000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"figure, (ax1, ax2, ax3, ax4) = plt.subplots(1,4)\n\nfigure.set_size_inches(15, 5)\nax1.hist(runtime_plotter['avg'])\nax2.hist(runtime_plotter['min'])\nax3.hist(runtime_plotter['max'])\nax4.hist(runtime_plotter['range'])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-28T02:19:49.667712Z","iopub.execute_input":"2023-09-28T02:19:49.668092Z","iopub.status.idle":"2023-09-28T02:19:50.392390Z","shell.execute_reply.started":"2023-09-28T02:19:49.668057Z","shell.execute_reply":"2023-09-28T02:19:50.391276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now it looks a little bit more coherent and less sparse. This normalization may help the models to not get stuck thinking of valid values as outliers. ","metadata":{}}]}