{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Examine MP4 files\n\nI thought it would be fun to write some Python code that digs into the MP4 files, without using any external dependencies such as `ffprobe`.\n\nMore information about the MP4 file format at the following links:\n\n- http://xhelmboyx.tripod.com/formats/mp4-layout.txt\n- https://github.com/OpenAnsible/rust-mp4/raw/master/docs/ISO_IEC_14496-14_2003-11-15.pdf\n- https://developer.apple.com/library/archive/documentation/QuickTime/QTFF/QTFFPreface/qtffPreface.html"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os, sys\nimport struct\nimport numpy as np","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"An MP4 file is a container that consists of several sections, also known as boxes, chunks, or atoms. Boxes can be nested. Getting data out of an MP4 file consists of finding the box(es) you're interested in, then digging into those."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"def find_boxes(f, start_offset=0, end_offset=float(\"inf\")):\n    \"\"\"Returns a dictionary of all the data boxes and their absolute starting\n    and ending offsets inside the mp4 file.\n\n    Specify a start_offset and end_offset to read sub-boxes.\n    \"\"\"\n    s = struct.Struct(\"> I 4s\") \n    boxes = {}\n    offset = start_offset\n    f.seek(offset, 0)\n    while offset < end_offset:\n        data = f.read(8)               # read box header\n        if data == b\"\": break          # EOF\n        length, text = s.unpack(data)\n        f.seek(length - 8, 1)          # skip to next box\n        boxes[text] = (offset, offset + length)\n        offset += length\n    return boxes","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"An interesting box is **mvhd**, the \"movie header\", which is inside the **moov** box. Among other things, the movie header contains the duration of the video."},{"metadata":{"trusted":true},"cell_type":"code","source":"def scan_mvhd(f, offset):\n    f.seek(offset, 0)\n    f.seek(8, 1)            # skip box header\n\n    data = f.read(1)        # read version number\n    version = int.from_bytes(data, \"big\")\n    word_size = 8 if version == 1 else 4\n\n    f.seek(3, 1)            # skip flags\n    f.seek(word_size*2, 1)  # skip dates\n\n    timescale = int.from_bytes(f.read(4), \"big\")\n    if timescale == 0: timescale = 600\n\n    duration = int.from_bytes(f.read(word_size), \"big\")\n\n    print(\"Duration (sec):\", duration / timescale)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The main function for examining an MP4 file. Right now it just prints out the file offsets of some of the more interesting boxes, such as **udta** which contains metadata and **trak** that describes the tracks in the movie file."},{"metadata":{"trusted":true},"cell_type":"code","source":"def examine_mp4(filename):\n    print(\"Examining:\", filename)\n    \n    with open(filename, \"rb\") as f:\n        boxes = find_boxes(f)\n        print(boxes)\n\n        # Sanity check that this really is a movie file.\n        assert(boxes[b\"ftyp\"][0] == 0)\n\n        moov_boxes = find_boxes(f, boxes[b\"moov\"][0] + 8, boxes[b\"moov\"][1])\n        print(moov_boxes)\n\n        trak_boxes = find_boxes(f, moov_boxes[b\"trak\"][0] + 8, moov_boxes[b\"trak\"][1])\n        print(trak_boxes)\n\n        udta_boxes = find_boxes(f, moov_boxes[b\"udta\"][0] + 8, moov_boxes[b\"udta\"][1])\n        print(udta_boxes)\n\n        scan_mvhd(f, moov_boxes[b\"mvhd\"][0])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's try it on some of the videos:"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_dir = \"/kaggle/input/deepfake-detection-challenge/test_videos/\"\ntest_files = [x for x in os.listdir(test_dir) if x[-4:] == \".mp4\"]\n\ntrain_dir = \"/kaggle/input/deepfake-detection-challenge/train_sample_videos/\"\ntrain_files = [x for x in os.listdir(train_dir) if x[-4:] == \".mp4\"]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"examine_mp4(os.path.join(train_dir, np.random.choice(train_files)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"examine_mp4(os.path.join(test_dir, np.random.choice(test_files)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I originally wrote this code to quickly get the durations from the mp4 files but it turns out they're all about the same length (10 seconds).\n\nI also used this to see if there was any leakage inside the mp4 files themselves, such as in the metadata, but so far I haven't found anything. ;-)\n\nAnyway, I just wanted to show that you don't necessarily need external tools to read information from mp4 files. :-)"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}