{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Abstract\n\nIn this notebook I compared loading time of pytorch DataLoader in two file formats: JPEG and numpy array (`*.npy`). As a result, numpy format is about 2-10 times faster than JPEG format. Note that using numpy format as a data source, pytorch DataLoader loads even faster after it loads file second time. I don't find any information source, but it might caches files when once loaded.","metadata":{}},{"cell_type":"code","source":"!pip install nb-black","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-09T00:44:12.527978Z","iopub.execute_input":"2023-03-09T00:44:12.528476Z","iopub.status.idle":"2023-03-09T00:44:29.122791Z","shell.execute_reply.started":"2023-03-09T00:44:12.528422Z","shell.execute_reply":"2023-03-09T00:44:29.121064Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%load_ext lab_black\n%load_ext autoreload\n%autoreload 2\n\nfrom functools import lru_cache\n\nfrom pathlib import Path\nimport numpy as np\nimport cv2\nfrom tqdm import tqdm\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\n\n\nimport matplotlib.pyplot as plt\nimport polars as pl\n\nDATA_DIR = Path(\"/kaggle/jpg\")\n\nif not DATA_DIR.exists():\n    DATA_DIR.mkdir(parents=True)","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:44:29.128437Z","iopub.execute_input":"2023-03-09T00:44:29.128918Z","iopub.status.idle":"2023-03-09T00:44:32.736717Z","shell.execute_reply.started":"2023-03-09T00:44:29.128863Z","shell.execute_reply":"2023-03-09T00:44:32.735198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"te_helmet = pl.read_csv(\n    \"/kaggle/input/nfl-player-contact-detection/test_baseline_helmets.csv\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:44:32.738672Z","iopub.execute_input":"2023-03-09T00:44:32.739381Z","iopub.status.idle":"2023-03-09T00:44:33.431048Z","shell.execute_reply.started":"2023-03-09T00:44:32.739333Z","shell.execute_reply":"2023-03-09T00:44:33.429893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for video_file in tqdm(te_helmet[\"video\"].unique().to_list()):\n    video_file = Path(video_file)\n    output_path = DATA_DIR / video_file.stem\n    !ffmpeg -i /kaggle/input/nfl-player-contact-detection/test/{video_file} -q:v 2 -f image2 {output_path}_%06d.jpg -hide_banner -loglevel error","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:44:33.437099Z","iopub.execute_input":"2023-03-09T00:44:33.438109Z","iopub.status.idle":"2023-03-09T00:45:05.591377Z","shell.execute_reply.started":"2023-03-09T00:44:33.438059Z","shell.execute_reply":"2023-03-09T00:45:05.589828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"jpg_files = list(Path(\"/kaggle/jpg\").glob(\"*.jpg\"))\nout_dir = Path(\"/kaggle/npy\")\nout_dir.mkdir(exist_ok=True)\n\n\nfor jpg_file in tqdm(jpg_files):\n    img = cv2.imread(jpg_file.as_posix())\n    np.save(out_dir / f\"{jpg_file.stem}.npy\", img)","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:45:05.593172Z","iopub.execute_input":"2023-03-09T00:45:05.593566Z","iopub.status.idle":"2023-03-09T00:46:21.452287Z","shell.execute_reply.started":"2023-03-09T00:45:05.593522Z","shell.execute_reply":"2023-03-09T00:46:21.449858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class NPYDataset(Dataset):\n    def __init__(self, path: Path):\n        self.files = list(path.glob(\"*.npy\"))\n\n    def __len__(self):\n        return len(self.files)\n\n    def __getitem__(self, idx):\n        img = np.load(self.files[idx].as_posix())\n        return img\n\n\nclass JPEGDataset(Dataset):\n    def __init__(self, path: Path):\n        self.files = list(path.glob(\"*.jpg\"))\n\n    def __len__(self):\n        return len(self.files)\n\n    def __getitem__(self, idx):\n        img = cv2.imread(self.files[idx].as_posix())\n        return img","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:46:21.455400Z","iopub.execute_input":"2023-03-09T00:46:21.457689Z","iopub.status.idle":"2023-03-09T00:46:21.547739Z","shell.execute_reply.started":"2023-03-09T00:46:21.457620Z","shell.execute_reply":"2023-03-09T00:46:21.546059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def test_jpg_loader(path: Path, epochs=3):\n    dataset = JPEGDataset(path)\n    loader = DataLoader(dataset, pin_memory=True)\n\n    for epoch in range(epochs):\n        for img in tqdm(loader):\n            img.detach().cpu().numpy()\n    del dataset\n    del loader\n\n\ndef test_npy_loader(path: Path, epochs=3):\n    dataset = NPYDataset(path)\n    loader = DataLoader(dataset, pin_memory=True)\n\n    for epoch in range(epochs):\n        for img in tqdm(loader):\n            img.detach().cpu().numpy()\n    del dataset\n    del loader","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:50:34.342386Z","iopub.execute_input":"2023-03-09T00:50:34.342950Z","iopub.status.idle":"2023-03-09T00:50:34.408526Z","shell.execute_reply.started":"2023-03-09T00:50:34.342893Z","shell.execute_reply":"2023-03-09T00:50:34.407215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_jpg_loader(Path(\"/kaggle/jpg\"))","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:46:23.142945Z","iopub.execute_input":"2023-03-09T00:46:23.144312Z","iopub.status.idle":"2023-03-09T00:48:59.526785Z","shell.execute_reply.started":"2023-03-09T00:46:23.144260Z","shell.execute_reply":"2023-03-09T00:48:59.525354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_npy_loader(Path(\"/kaggle/npy\"))","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:48:59.529270Z","iopub.execute_input":"2023-03-09T00:48:59.529900Z","iopub.status.idle":"2023-03-09T00:49:39.385601Z","shell.execute_reply.started":"2023-03-09T00:48:59.529843Z","shell.execute_reply":"2023-03-09T00:49:39.384136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Discussion \n\nLoading from `*.npy` format took 2-10x faster than loading from JPEG format. Note that in the first epoch, it loads only 2x faster, but after 2nd epoch, it loads 10x faster. I don't find any information source, but `np.load()` might cache file when once loaded.","metadata":{}},{"cell_type":"markdown","source":"## Filesize\n\nAlthough numpy data format (`*.npy`) is faster to load than JPEG format, it uses more disk spaces. Compared to ffmpeg's default compression rate in JPEG file, numpy data format uses about 20x disk spaces.","metadata":{}},{"cell_type":"code","source":"!du -sh /kaggle/jpg","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:49:39.389854Z","iopub.execute_input":"2023-03-09T00:49:39.390291Z","iopub.status.idle":"2023-03-09T00:49:40.489492Z","shell.execute_reply.started":"2023-03-09T00:49:39.390246Z","shell.execute_reply":"2023-03-09T00:49:40.487887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!du -sh /kaggle/npy","metadata":{"execution":{"iopub.status.busy":"2023-03-09T00:49:40.492089Z","iopub.execute_input":"2023-03-09T00:49:40.492542Z","iopub.status.idle":"2023-03-09T00:49:41.610181Z","shell.execute_reply.started":"2023-03-09T00:49:40.492492Z","shell.execute_reply":"2023-03-09T00:49:41.608668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion\n\nLoading from ndarray (`*.npy`) is about 2-10x faster than loading from JPEG file.\nHowever, file size in `*.npy` format is about 20x larger than JPEG format.","metadata":{}}]}