{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Video reading speed test\n\nSee also [this discussion topic](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122328).\n\nI currently have a face detector that takes about 30 ms for a batch of 30 images, but loading that batch from the video takes 330 ms (on my local computer). So video loading is 10x slower than doing face detection. Because we're dealing with a LOT of data in this competition, it's worthwhile to reduce any overhead where possible.\n\n> Note that loading speeds appear to be slower on Kaggle than on my local machine (1.2 seconds instead of 330 ms for the same video). Not sure why that is. On my local machine the videos are stored on HDD, not SSD."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Pick a video at random."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train_dir = \"/kaggle/input/deepfake-detection-challenge/train_sample_videos\"\nvideo_path = os.path.join(train_dir, np.random.choice(os.listdir(train_dir)))\nvideo_path","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This version uses OpenCV. It looks at all the frames in the video but only decodes the ones we're interested in:"},{"metadata":{"trusted":true},"cell_type":"code","source":"def grab_frames_from_video(path, num_frames=10):\n    capture = cv2.VideoCapture(path)\n    frame_count = int(capture.get(cv2.CAP_PROP_FRAME_COUNT))\n    frame_idxs = np.linspace(0, frame_count, num_frames, endpoint=False, dtype=np.int)\n\n    i = 0\n    for frame_idx in range(int(frame_count)):\n        # Get the next frame, but don't decode if we're not using it.\n        ret = capture.grab()\n        if not ret: \n            print(\"Error grabbing frame %d from movie %s\" % (frame_idx, path))\n\n        # Need to look at this frame?\n        if frame_idx >= frame_idxs[i]:\n            ret, frame = capture.retrieve()\n            if not ret or frame is None:\n                print(\"Error retrieving frame %d from movie %s\" % (frame_idx, path))\n            else:\n                frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n                # Do something with `frame`\n\n            i += 1\n            if i >= len(frame_idxs):\n                break\n\n    capture.release()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%time grab_frames_from_video(video_path, num_frames=10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"1.2 seconds for reading a single video. Ouch. With 4000 videos that is 1.5 hours to read the entire test set."},{"metadata":{"trusted":true},"cell_type":"code","source":"%time grab_frames_from_video(video_path, num_frames=50)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"At least reading more frames doesn't make the time much worse..."},{"metadata":{},"cell_type":"markdown","source":"The next version jumps directly to the frame you want to read. You might expect this to be faster but it actually isn't. My guess is that a \"streaming\" approach is more efficient than a \"random access\" approach because, unless you happen to grab a keyframe, the decoder still needs to read all the previous frames in order to reconstruct the one you're asking for."},{"metadata":{"trusted":true},"cell_type":"code","source":"# This version.\n\ndef grab_frames_from_video(path, num_frames=10):\n    capture = cv2.VideoCapture(path)\n    frame_count = int(capture.get(cv2.CAP_PROP_FRAME_COUNT))\n    frame_idxs = np.linspace(0, frame_count, num_frames, endpoint=False, dtype=np.int)\n\n    for i, frame_idx in enumerate(frame_idxs):\n        capture.set(cv2.CAP_PROP_POS_FRAMES, frame_idx)\n        ret, frame = capture.read()\n        if not ret or frame is None:\n            print(\"Error retrieving frame %d from movie %s\" % (frame_idx, path))\n        else:\n            frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n\n    capture.release()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%time grab_frames_from_video(video_path, num_frames=10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Yeah that's 4x slower than the other method."},{"metadata":{"trusted":true},"cell_type":"code","source":"%time grab_frames_from_video(video_path, num_frames=50)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And unlike the previous method, it gets way worse the more frames you want to look at."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}