{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Competition Understanding...\n\nAt first glance this competition has a lot of moving parts, making it hard to understand how to approach. In this notebook I'm going through some of my thought processes to understand the competition and it's data.\n\nInitial thoughts:\n- **What are we predicting?** - Predict the runtime length of ML graphs and configurations.\n- **What does the data look like?** - Graph configurations in npz format.\n    - Two \"collection types\": `tile` and `layout` collections\n- **What does the target look like?** - \"Finally, for the layout collections, your job is to predict the order of the indices from best-to-worse configurations (i.e., ones leading to the smallest `d[\"config_runtime\"]`)\"\n- **How are we evaluated?** two evaluation metrics, described below.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"!python3 -m pip install -q jupyter-black jupyter","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-08-30T15:22:30.447944Z","iopub.execute_input":"2023-08-30T15:22:30.448473Z","iopub.status.idle":"2023-08-30T15:22:45.499699Z","shell.execute_reply.started":"2023-08-30T15:22:30.448428Z","shell.execute_reply":"2023-08-30T15:22:45.498010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom glob import glob\nfrom pathlib import Path\nfrom tqdm import tqdm\n\nimport jupyter_black\n\njupyter_black.load()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:50:03.866018Z","iopub.execute_input":"2023-08-30T15:50:03.866547Z","iopub.status.idle":"2023-08-30T15:50:03.882393Z","shell.execute_reply.started":"2023-08-30T15:50:03.866503Z","shell.execute_reply":"2023-08-30T15:50:03.881027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluation Metric\n\nThis one is a bit confusing..\n\n1. For collection `tile:xla`\n    - We use the (1-slowdown) incurred of the top-K predictions to reflect how much slower the top-K configurations predicted by the model is from the actual fastest configuration\n2. For collection `layout:*`\n    - We use the Kendal Tau Correlation (a ranking metric: how well does your model-predicted ranking, correspond to the real ranking of runtimes).","metadata":{}},{"cell_type":"markdown","source":"# Submission Format:\nDepending on collection type:\n1.  For the `tile:xla` collection, only the *first 5 entries will be considered* and the rest will be ignored.\n2. For the `layout:*` collections, *all entries will be considered* (you should output a permutation of the number of configurations!).","metadata":{}},{"cell_type":"code","source":"ss = pd.read_csv(\"../input/predict-ai-model-runtime/sample_submission.csv\")\nss.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:25:50.273084Z","iopub.execute_input":"2023-08-30T15:25:50.273639Z","iopub.status.idle":"2023-08-30T15:25:50.304456Z","shell.execute_reply.started":"2023-08-30T15:25:50.273599Z","shell.execute_reply":"2023-08-30T15:25:50.303080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is why we see two different types of `TopConfigs` in the sample submission document.\n- The examples with only 5 configs are the `tile:xla`\n- Examples with more are `layout:*` collections.","metadata":{}},{"cell_type":"code","source":"print(\"Unique sample TopConfigs\", ss[\"TopConfigs\"].nunique())\n\nss[\"TopConfigs\"].str[:20].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:25:50.899446Z","iopub.execute_input":"2023-08-30T15:25:50.900385Z","iopub.status.idle":"2023-08-30T15:25:50.921342Z","shell.execute_reply.started":"2023-08-30T15:25:50.900332Z","shell.execute_reply":"2023-08-30T15:25:50.920236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Structure\n- `npz_all` contains the data.\n    - Two collection types: `layout` and `tile` folders.\n        - `layout` has `nlp` and `xla`\n        - `tile` has `xla` only\n    - Each collection type has default/random\n    - The raw data is split `train/test/valid`","metadata":{}},{"cell_type":"code","source":"!tree -I *.npz /kaggle/input/predict-ai-model-runtime/","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:27:07.047474Z","iopub.execute_input":"2023-08-30T15:27:07.047985Z","iopub.status.idle":"2023-08-30T15:27:08.257075Z","shell.execute_reply.started":"2023-08-30T15:27:07.047943Z","shell.execute_reply":"2023-08-30T15:27:08.255725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating our Own Metadata CSV\n- We aren't given a metadata csv but we can create one on our own. This will help with further EDA\n\nUse their provided description to create metadata columns:\n```\nEach layout data file is named as followed: http://download.tensorflow.org/data/tpu_graphs/v0/npz_layout_{source}_{search}_{split}.tar\n- {source}: xla or nlp\n- {search}: default or random\n- {split}: train, valid, or test\n```","metadata":{}},{"cell_type":"code","source":"BASE_PATH = Path(\"/kaggle/input/predict-ai-model-runtime/npz_all/\")\n\nnpz_files = [str(f) for f in BASE_PATH.glob(\"**/*.npz\")]\n\nmetadata = pd.DataFrame(npz_files, columns=[\"file_path\"])\nmetadata[\"file_depth\"] = metadata[\"file_path\"].str.split(\"/\").str.len()\nmetadata[\"file_name\"] = metadata[\"file_path\"].str.split(\"/\").str[-1]\nmetadata[\"base_collection\"] = metadata[\"file_path\"].str.split(\"/\").str[6]\nmetadata[\"split\"] = metadata[\"file_path\"].str.split(\"/\").str[-2]\nmetadata[\"search\"] = (\n    metadata.loc[metadata[\"base_collection\"] == \"layout\", \"file_path\"]\n    .str.split(\"/\")\n    .str[-3]\n)\nmetadata[\"source\"] = metadata[\"file_path\"].str.split(\"/\").str[7]","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:37:34.856594Z","iopub.execute_input":"2023-08-30T15:37:34.857227Z","iopub.status.idle":"2023-08-30T15:37:35.205870Z","shell.execute_reply.started":"2023-08-30T15:37:34.857175Z","shell.execute_reply":"2023-08-30T15:37:35.204654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata[\n    [\n        \"base_collection\",\n        \"source\",\n        \"search\",\n        \"split\",\n    ]\n].value_counts().reset_index().rename(columns={0: \"Count\"})","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:39:21.137223Z","iopub.execute_input":"2023-08-30T15:39:21.137704Z","iopub.status.idle":"2023-08-30T15:39:21.180141Z","shell.execute_reply.started":"2023-08-30T15:39:21.137667Z","shell.execute_reply":"2023-08-30T15:39:21.179061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(ss), metadata.query('split == \"test\"').shape","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:40:07.534020Z","iopub.execute_input":"2023-08-30T15:40:07.535223Z","iopub.status.idle":"2023-08-30T15:40:07.552092Z","shell.execute_reply.started":"2023-08-30T15:40:07.535150Z","shell.execute_reply":"2023-08-30T15:40:07.550941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore `npz` structure and add to metadata","metadata":{}},{"cell_type":"code","source":"example_tile_npz = metadata.query('base_collection == \"tile\"')[\"file_path\"].values[0]\nraw = np.load(example_tile_npz)\n\nprint(\"Keys inside of an example TILE npz file:\")\nlist(raw)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:45:51.993513Z","iopub.execute_input":"2023-08-30T15:45:51.993994Z","iopub.status.idle":"2023-08-30T15:45:52.029629Z","shell.execute_reply.started":"2023-08-30T15:45:51.993950Z","shell.execute_reply":"2023-08-30T15:45:52.028247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example_tile_npz = metadata.query('base_collection == \"layout\"')[\"file_path\"].values[0]\nraw = np.load(example_tile_npz)\n\nprint(\"Keys inside of an example LAYOUT npz file:\")\nlist(raw)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:46:30.010730Z","iopub.execute_input":"2023-08-30T15:46:30.011415Z","iopub.status.idle":"2023-08-30T15:46:30.090198Z","shell.execute_reply.started":"2023-08-30T15:46:30.011374Z","shell.execute_reply":"2023-08-30T15:46:30.089226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Extract the target variables for tile and layout\n\n**Tile**\n\n- For the tile collection, your job is to predict the indices of the best configurations (i.e., ones leading to the smallest `d[\"config_runtime\"] / d[\"config_runtime_normalizers\"]`).\n\n**Layout**\n\n- For the layout collections, your job is to predict the order of the indices from best-to-worse configurations (i.e., ones leading to the smallest `d[\"config_runtime\"]`). We do not have to use runtime normalizers for this task because the runtime variation at the entire program level is very small.","metadata":{}},{"cell_type":"code","source":"def extract_npz_metadata(raw):\n    out = {}\n    return {}","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:51:11.384230Z","iopub.execute_input":"2023-08-30T15:51:11.384730Z","iopub.status.idle":"2023-08-30T15:51:11.394790Z","shell.execute_reply.started":"2023-08-30T15:51:11.384690Z","shell.execute_reply":"2023-08-30T15:51:11.393229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print the data shapes for examples\nprint(\"=\" * 50)\nprint(\"Tile Example Data Size\")\nprint(\"=\" * 50)\n\nexample_tile_npz = metadata.query('base_collection == \"tile\"')[\"file_path\"].values[0]\nraw = np.load(example_tile_npz)\nfor key, value in raw.items():\n    print(key, value.shape)\n\nprint(\"=\" * 50)\nprint(\"Layout Example Data Size\")\nprint(\"=\" * 50)\n\nexample_tile_npz = metadata.query('base_collection == \"layout\"')[\"file_path\"].values[0]\nraw = np.load(example_tile_npz)\nfor key, value in raw.items():\n    print(key, value.shape)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T15:55:46.695122Z","iopub.execute_input":"2023-08-30T15:55:46.695698Z","iopub.status.idle":"2023-08-30T15:55:47.702510Z","shell.execute_reply.started":"2023-08-30T15:55:46.695634Z","shell.execute_reply":"2023-08-30T15:55:47.700674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Parse a tile example","metadata":{}},{"cell_type":"code","source":"TEST_RUN = False","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:11:34.182980Z","iopub.execute_input":"2023-08-30T16:11:34.183501Z","iopub.status.idle":"2023-08-30T16:11:34.191731Z","shell.execute_reply.started":"2023-08-30T16:11:34.183461Z","shell.execute_reply":"2023-08-30T16:11:34.190241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tile_npz_files = metadata.query('base_collection == \"tile\"')[\"file_path\"].values\nif TEST_RUN:\n    tile_npz_files = tile_npz_files[:50]\nexample_dfs = []\nfor example_tile_npz in tqdm(tile_npz_files):\n    raw = np.load(example_tile_npz)\n    config_ids = range(len(raw[\"config_runtime\"]))\n    config_runtimes = raw[\"config_runtime\"]\n    config_runtime_normalizers = raw[\"config_runtime_normalizers\"]\n    assert len(config_ids) == len(config_runtimes) == len(config_runtime_normalizers)\n    my_id = example_tile_npz.split(\"/\")[-1][:-4]\n    base = example_tile_npz.split(\"/\")[-4]\n    source = example_tile_npz.split(\"/\")[-3]\n    ID = f\"{base}:{source}:{my_id}\"\n    example_df = pd.DataFrame(\n        [config_ids, config_runtimes, config_runtime_normalizers]\n    ).T\n    example_df.columns = [\n        \"config_id\",\n        \"config_runtime\",\n        \"config_runtime_normalizers\",\n    ]\n    example_df[\"file_id\"] = my_id\n    example_df[\"ID\"] = ID\n    example_df = example_df[\n        [\"ID\", \"config_id\", \"config_runtime\", \"config_runtime_normalizers\"]\n    ]\n    example_dfs.append(example_df)\ntile_metadata = pd.concat(example_dfs).reset_index(drop=True)\n# For tile -->  d[\"config_runtime\"] / d[\"config_runtime_normalizers\"]\ntile_metadata[\"target\"] = (\n    tile_metadata[\"config_runtime\"] / tile_metadata[\"config_runtime_normalizers\"]\n)\ntile_metadata[\"target_rank\"] = (\n    tile_metadata.groupby(\"ID\")[\"target\"].rank(method=\"dense\").astype(\"int\")\n)\ntile_metadata.to_parquet(\"tile_metadata.parquet\")","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:23:22.673601Z","iopub.execute_input":"2023-08-30T16:23:22.674733Z","iopub.status.idle":"2023-08-30T16:23:26.954903Z","shell.execute_reply.started":"2023-08-30T16:23:22.674684Z","shell.execute_reply":"2023-08-30T16:23:26.953677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Parse a layout examples","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:22:58.820695Z","iopub.execute_input":"2023-08-30T16:22:58.821245Z","iopub.status.idle":"2023-08-30T16:22:58.874949Z","shell.execute_reply.started":"2023-08-30T16:22:58.821193Z","shell.execute_reply":"2023-08-30T16:22:58.873402Z"}}},{"cell_type":"code","source":"tile_npz_files = metadata.query('base_collection == \"layout\"')[\"file_path\"].values\nif TEST_RUN:\n    tile_npz_files = tile_npz_files[:50]\nexample_dfs = []\nfor example_tile_npz in tqdm(tile_npz_files):\n    raw = np.load(example_tile_npz)\n    config_ids = range(len(raw[\"config_runtime\"]))\n    config_runtimes = raw[\"config_runtime\"]\n    assert len(config_ids) == len(config_runtimes)\n    my_id = example_tile_npz.split(\"/\")[-1][:-4]\n    base = example_tile_npz.split(\"/\")[-4]\n    source = example_tile_npz.split(\"/\")[-3]\n    ID = f\"{base}:{source}:{my_id}\"\n    example_df = pd.DataFrame(\n        [config_ids, config_runtimes]\n    ).T\n    example_df.columns = [\n        \"config_id\",\n        \"config_runtime\"\n    ]\n    example_df[\"file_id\"] = my_id\n    example_df[\"ID\"] = ID\n    example_df = example_df[\n        [\"ID\", \"config_id\", \"config_runtime\"]\n    ]\n    example_dfs.append(example_df)\nlayout_metadata = pd.concat(example_dfs).reset_index(drop=True)\n# For tile -->  d[\"config_runtime\"] / d[\"config_runtime_normalizers\"]\nlayout_metadata[\"target\"] = layout_metadata[\"config_runtime\"]\nlayout_metadata[\"target_rank\"] = (\n    layout_metadata.groupby(\"ID\")[\"target\"].rank(method=\"dense\").astype(\"int\")\n)\nlayout_metadata.to_parquet(\"layout_metadata.parquet\")","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:26:27.656861Z","iopub.execute_input":"2023-08-30T16:26:27.657402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata.to_csv('metadata.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Get the Training Data Targets and Compute a Baseline Submission and Score\n- Can we get the ground truth for some of the training data?\n- Can we use the metric above to grade a baseline submission?\n- We would expect the result to be similar to the current leaderboard score of `0.119`","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploring the Baseline Code Base:\n- Link: https://github.com/google-research-datasets/tpu_graphs/tree/main\n\nThis contains:\n","metadata":{}},{"cell_type":"code","source":"!git clone https://github.com/google-research-datasets/tpu_graphs.git","metadata":{"execution":{"iopub.status.busy":"2023-08-30T14:24:54.354321Z","iopub.execute_input":"2023-08-30T14:24:54.354876Z","iopub.status.idle":"2023-08-30T14:24:55.462752Z","shell.execute_reply.started":"2023-08-30T14:24:54.354826Z","shell.execute_reply":"2023-08-30T14:24:55.460821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Display the Repo Code Structure","metadata":{}},{"cell_type":"code","source":"!ls tpu_graphs/ -GFlash --color","metadata":{"execution":{"iopub.status.busy":"2023-08-30T14:25:27.283568Z","iopub.execute_input":"2023-08-30T14:25:27.284795Z","iopub.status.idle":"2023-08-30T14:25:28.391856Z","shell.execute_reply.started":"2023-08-30T14:25:27.284739Z","shell.execute_reply":"2023-08-30T14:25:28.390167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!tree tpu_graphs","metadata":{"execution":{"iopub.status.busy":"2023-08-30T14:25:28.729363Z","iopub.execute_input":"2023-08-30T14:25:28.729874Z","iopub.status.idle":"2023-08-30T14:25:29.872683Z","shell.execute_reply.started":"2023-08-30T14:25:28.729833Z","shell.execute_reply":"2023-08-30T14:25:29.870770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check it out","metadata":{"execution":{"iopub.status.busy":"2023-08-30T14:25:22.504502Z","iopub.execute_input":"2023-08-30T14:25:22.505766Z","iopub.status.idle":"2023-08-30T14:25:23.619410Z","shell.execute_reply.started":"2023-08-30T14:25:22.505695Z","shell.execute_reply":"2023-08-30T14:25:23.618105Z"}}}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}