{
  "id": 122328,
  "title": "Video loading optimization tips",
  "url": "/competitions/deepfake-detection-challenge/discussion/122328",
  "author_name": "Human Analog",
  "post_date": "2019-12-19T10:48:58.895000",
  "votes": 115,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Since loading videos is an important part of this challenge, I thought a topic on how to do this as fast as possible would be useful. :-)</p>\n\n<p>I'm not an expert on handling video in Python (yet -- one of the reasons why this competition is interesting to me) but I did notice the following:</p>\n\n<p>If you're not interested in using all the frames from a video, for example only every 10th frame, don't do the following:</p>\n\n<p><code>\ncapture = cv.VideoCapture(movie_path)\nfor i in range(0, num_frames):\n    ret, frame = capture.read()\n    if i % 10 == 0:\n        # do something with frame\n</code></p>\n\n<p>Instead, it's faster to call <code>grab()</code>, which reads the next frame but doesn't decode it, and only call <code>retrieve()</code> on the frames you're interested in:</p>\n\n<p><code>\ncapture = cv.VideoCapture(movie_path)\nfor i in range(0, num_frames):\n    ret = capture.grab()\n    if i % 10 == 0:\n        ret, frame = capture.retrieve()\n        # do something with frame\ncapture.release()\n</code></p>\n\n<p>This will already make quite a bit of difference in processing time.</p>\n\n<p>If you've got any more tips, please reply below. 😄 </p>",
  "messages": [
    {
      "id": 698522,
      "postDate": "2019-12-19T10:48:58.897Z",
      "content": "<p>Since loading videos is an important part of this challenge, I thought a topic on how to do this as fast as possible would be useful. :-)</p>\n\n<p>I'm not an expert on handling video in Python (yet -- one of the reasons why this competition is interesting to me) but I did notice the following:</p>\n\n<p>If you're not interested in using all the frames from a video, for example only every 10th frame, don't do the following:</p>\n\n<p><code>\ncapture = cv.VideoCapture(movie_path)\nfor i in range(0, num_frames):\n    ret, frame = capture.read()\n    if i % 10 == 0:\n        # do something with frame\n</code></p>\n\n<p>Instead, it's faster to call <code>grab()</code>, which reads the next frame but doesn't decode it, and only call <code>retrieve()</code> on the frames you're interested in:</p>\n\n<p><code>\ncapture = cv.VideoCapture(movie_path)\nfor i in range(0, num_frames):\n    ret = capture.grab()\n    if i % 10 == 0:\n        ret, frame = capture.retrieve()\n        # do something with frame\ncapture.release()\n</code></p>\n\n<p>This will already make quite a bit of difference in processing time.</p>\n\n<p>If you've got any more tips, please reply below. 😄 </p>",
      "rawMarkdown": "Since loading videos is an important part of this challenge, I thought a topic on how to do this as fast as possible would be useful. :-)\n\nI'm not an expert on handling video in Python (yet -- one of the reasons why this competition is interesting to me) but I did notice the following:\n\nIf you're not interested in using all the frames from a video, for example only every 10th frame, don't do the following:\n\n```\ncapture = cv.VideoCapture(movie_path)\nfor i in range(0, num_frames):\n    ret, frame = capture.read()\n    if i % 10 == 0:\n        # do something with frame\n```\n\nInstead, it's faster to call `grab()`, which reads the next frame but doesn't decode it, and only call `retrieve()` on the frames you're interested in:\n\n```\ncapture = cv.VideoCapture(movie_path)\nfor i in range(0, num_frames):\n    ret = capture.grab()\n    if i % 10 == 0:\n        ret, frame = capture.retrieve()\n        # do something with frame\ncapture.release()\n```\n\nThis will already make quite a bit of difference in processing time.\n\nIf you've got any more tips, please reply below. 😄 ",
      "votes": 114
    },
    {
      "id": 701751,
      "postDate": "2019-12-23T21:46:33.513Z",
      "content": "<p>For me, I just take 10 frames. So I used\n<code>cap.set(cv2.CAP_PROP_POS_FRAMES, framenum)</code></p>\n\n<p>For example, this is the code for getting frame number 33: </p>\n\n<p><code>cap.set(cv2.CAPPROPPOSFRAMES, 33)\nret, framenumber33 = cap.read()</code></p>\n\n<p>According to this <a href=\"https://subscription.packtpub.com/book/application_development/9781788474443/1/ch01lvl1sec24/jumping-between-frames-in-video-files\">blog</a></p>",
      "rawMarkdown": "For me, I just take 10 frames. So I used\n`cap.set(cv2.CAP_PROP_POS_FRAMES, framenum)`\n\n For example, this is the code for getting frame number 33: \n\n`cap.set(cv2.CAPPROPPOSFRAMES, 33)\nret, framenumber33 = cap.read()`\n\n\nAccording to this [blog](https://subscription.packtpub.com/book/application_development/9781788474443/1/ch01lvl1sec24/jumping-between-frames-in-video-files)",
      "votes": 7,
      "replies": [
        {
          "id": 715408,
          "postDate": "2020-01-10T13:52:41.717Z",
          "content": "<p>I don't know why but For me it is almost 4 times slower then retrieve-grab method.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1830962%2F4673fb4666d059e68a265ab22a816baf%2FScreenshot%20from%202020-01-10%2019-21-42.png?generation=1578664342203845&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I don't know why but For me it is almost 4 times slower then retrieve-grab method.![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1830962%2F4673fb4666d059e68a265ab22a816baf%2FScreenshot%20from%202020-01-10%2019-21-42.png?generation=1578664342203845&amp;alt=media)\n"
        },
        {
          "id": 716569,
          "postDate": "2020-01-11T22:11:02.570Z",
          "content": "<p>That's most likely because videos don't store the full data for each frame, except in so-called keyframes. So if the frame you're asking for is not a keyframe, OpenCV will have to go back to read the keyframe and then stream until the frame you're asking for anyway.</p>",
          "rawMarkdown": "That's most likely because videos don't store the full data for each frame, except in so-called keyframes. So if the frame you're asking for is not a keyframe, OpenCV will have to go back to read the keyframe and then stream until the frame you're asking for anyway.",
          "votes": 2
        },
        {
          "id": 716570,
          "postDate": "2020-01-11T22:13:41.293Z",
          "content": "<p>This works for me because I just need a single frame. This doesn't work for you because you will need quite a few frames.</p>",
          "rawMarkdown": "This works for me because I just need a single frame. This doesn't work for you because you will need quite a few frames.",
          "votes": 1
        }
      ]
    },
    {
      "id": 701062,
      "postDate": "2019-12-23T03:20:08.593Z",
      "content": "<p>For training, which you do on your own machine, not in a kernel, why not just find all the faces and save them to disk. Then use these for all training runs.</p>",
      "rawMarkdown": "For training, which you do on your own machine, not in a kernel, why not just find all the faces and save them to disk. Then use these for all training runs.",
      "votes": 3
    },
    {
      "id": 699606,
      "postDate": "2019-12-20T16:20:45.670Z",
      "content": "<p>Here is a benchmark of scikit-video and opencv using python .\nto load a single frame\ntime taken by scikit-video : 0.005018000000000189\ntime taken by cv2 : 0.0228900000000003</p>\n\n<p>time taken to load 5637 frames sequentially</p>\n\n<p>time taken by scikit-video 8.149045999999998\ntime taken by cv2 : 55.439571</p>\n\n<p><a href=\"https://medium.com/@amkr/week-2-gsoc19-ccextractor-development-e4ffb6d54376\">blog link</a></p>",
      "rawMarkdown": "Here is a benchmark of scikit-video and opencv using python .\nto load a single frame\ntime taken by scikit-video : 0.005018000000000189\ntime taken by cv2 : 0.0228900000000003\n\ntime taken to load 5637 frames sequentially\n\ntime taken by scikit-video 8.149045999999998\ntime taken by cv2 : 55.439571\n\n[blog link](https://medium.com/@amkr/week-2-gsoc19-ccextractor-development-e4ffb6d54376)",
      "votes": 1,
      "replies": [
        {
          "id": 699609,
          "postDate": "2019-12-20T16:23:29.310Z",
          "content": "<p>Multithreading can also be used to increase the efficiency of frames processing.</p>",
          "rawMarkdown": "Multithreading can also be used to increase the efficiency of frames processing."
        },
        {
          "id": 699667,
          "postDate": "2019-12-20T18:18:28.160Z",
          "content": "<p>How to install scikit-video without internet? Can I turn on internet pip install and turn off internet ?</p>",
          "rawMarkdown": "How to install scikit-video without internet? Can I turn on internet pip install and turn off internet ?"
        },
        {
          "id": 699842,
          "postDate": "2019-12-21T02:39:54.270Z",
          "content": "<p>You will need internet to install it.\n<code>pip install scikit-video</code></p>",
          "rawMarkdown": "You will need internet to install it.\n`pip install scikit-video`"
        },
        {
          "id": 699992,
          "postDate": "2019-12-21T08:19:53.363Z",
          "content": "<p>but is it legal in this competition to use pip install ? internet forbidden :)</p>",
          "rawMarkdown": "but is it legal in this competition to use pip install ? internet forbidden :)\n"
        },
        {
          "id": 700293,
          "postDate": "2019-12-21T18:42:14.490Z",
          "content": "<p>Dang this would be a great package to use but yeah, Jonas is right. It won't be possible to use this since the internet and custom packages aren't allowed </p>",
          "rawMarkdown": "Dang this would be a great package to use but yeah, Jonas is right. It won't be possible to use this since the internet and custom packages aren't allowed "
        },
        {
          "id": 700700,
          "postDate": "2019-12-22T13:15:45.260Z",
          "content": "<p>Where does it say that in the rules? </p>",
          "rawMarkdown": "Where does it say that in the rules? "
        },
        {
          "id": 701002,
          "postDate": "2019-12-23T00:49:57.297Z",
          "content": "<p>We're allowed to install packages by uploading them as a dataset and then instialling with pip. Nothing different than just copying and pasting some code in so no reason why it would be excluded. When they say no custom packages they mean you can't use the old custom packages option kaggle used to have which would essentially be like a requirements.txt that was preinstalled before your kernel was initialized. Those aren't allowed because it would call the internet potentially and download stuff. </p>\n\n<p>That interface isn't even available on kaggle anymore. </p>\n\n<p>But also, tried skvideo locally and wasn't able to reproduce their speed-up. With skvideo I found it was actually slower.  opencv with grab and retrieve has been the fastest I have been able to find so far unfortunately. </p>",
          "rawMarkdown": "We're allowed to install packages by uploading them as a dataset and then instialling with pip. Nothing different than just copying and pasting some code in so no reason why it would be excluded. When they say no custom packages they mean you can't use the old custom packages option kaggle used to have which would essentially be like a requirements.txt that was preinstalled before your kernel was initialized. Those aren't allowed because it would call the internet potentially and download stuff. \n\nThat interface isn't even available on kaggle anymore. \n\nBut also, tried skvideo locally and wasn't able to reproduce their speed-up. With skvideo I found it was actually slower.  opencv with grab and retrieve has been the fastest I have been able to find so far unfortunately. ",
          "votes": 9
        },
        {
          "id": 701019,
          "postDate": "2019-12-23T01:59:26.600Z",
          "content": "<p>I also get no speedup from skvideo.  opencv seems to be the fastest.</p>",
          "rawMarkdown": "I also get no speedup from skvideo.  opencv seems to be the fastest.",
          "votes": 2
        },
        {
          "id": 701308,
          "postDate": "2019-12-23T10:44:02.293Z",
          "content": "<p>In addition to uploading the packages and installing the with pip, for python only packages you can just put the code in directly and add the dataset to the path to load it.</p>",
          "rawMarkdown": "In addition to uploading the packages and installing the with pip, for python only packages you can just put the code in directly and add the dataset to the path to load it."
        },
        {
          "id": 710263,
          "postDate": "2020-01-04T14:14:20.857Z",
          "content": "<p>I'm getting 68 seconds for skvideo &amp; 69 seconds for OpenCV doing the same task (face detection using dlib on each frame). Not much between them!</p>",
          "rawMarkdown": "I'm getting 68 seconds for skvideo &amp; 69 seconds for OpenCV doing the same task (face detection using dlib on each frame). Not much between them!"
        }
      ]
    },
    {
      "id": 707559,
      "postDate": "2020-01-01T03:58:42.143Z",
      "content": "<p>Followed your suggestion and tweaked my Video loading function. My data generator is significantly faster now. \nMy Keras data generator reads the input video (1 in 10 frames), uses <code>MTCNN</code> to detect faces and predict on them. \nBelow are the times taken to predict on 400 videos before and after the tweak.</p>\n\n<p>Before,\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1288210%2F852d6a3ff125c9cffd12865dfda5c417%2FScreenshot%202020-01-01%20at%209.25.21%20AM.png?generation=1577850978156060&amp;alt=media\" alt=\"\"></p>\n\n<p>After,\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1288210%2F770f446b0e72b834d3cee2d7c71f1251%2FScreenshot%202020-01-01%20at%209.25.35%20AM.png?generation=1577851018436037&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Followed your suggestion and tweaked my Video loading function. My data generator is significantly faster now. \nMy Keras data generator reads the input video (1 in 10 frames), uses `MTCNN` to detect faces and predict on them. \nBelow are the times taken to predict on 400 videos before and after the tweak.\n\nBefore,\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1288210%2F852d6a3ff125c9cffd12865dfda5c417%2FScreenshot%202020-01-01%20at%209.25.21%20AM.png?generation=1577850978156060&amp;alt=media)\n\nAfter,\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1288210%2F770f446b0e72b834d3cee2d7c71f1251%2FScreenshot%202020-01-01%20at%209.25.35%20AM.png?generation=1577851018436037&amp;alt=media)\n",
      "votes": 1
    },
    {
      "id": 698729,
      "postDate": "2019-12-19T16:29:33.783Z",
      "content": "<p>Thanks for the tip. Does yield some speed-up for me. What is the timing difference you have seen in your experience?</p>",
      "rawMarkdown": "Thanks for the tip. Does yield some speed-up for me. What is the timing difference you have seen in your experience?",
      "votes": 1,
      "replies": [
        {
          "id": 698854,
          "postDate": "2019-12-19T19:21:18.433Z",
          "content": "<p>I think it's about twice as fast, but obviously this matters more if you're reading more of the video.</p>\n\n<p>Reading the video is obviously not the bottleneck in this competition (running the model is), but there are a lot of videos so any time savings still add up. (If you save 1 second per video, that saves over an hour on the 4000 test set videos.)</p>\n\n<p>I think imageio can read any frame using random access, which would be even faster but I haven't really looked into that in detail yet.</p>",
          "rawMarkdown": "I think it's about twice as fast, but obviously this matters more if you're reading more of the video.\n\nReading the video is obviously not the bottleneck in this competition (running the model is), but there are a lot of videos so any time savings still add up. (If you save 1 second per video, that saves over an hour on the 4000 test set videos.)\n\nI think imageio can read any frame using random access, which would be even faster but I haven't really looked into that in detail yet.",
          "votes": 2
        },
        {
          "id": 698928,
          "postDate": "2019-12-19T22:38:58.207Z",
          "content": "<p>Good idea. Still the bottleneck is MTCNN for a decent face recognition.... Thanks! </p>",
          "rawMarkdown": "Good idea. Still the bottleneck is MTCNN for a decent face recognition.... Thanks! ",
          "votes": 2
        },
        {
          "id": 699124,
          "postDate": "2019-12-20T04:13:08.793Z",
          "content": "<p>I am only ran MTCNN on every 10th frame of the original videos.  The averaged over all 300 frames, and reused the face bounding boxes from the original for all of the fakes rather than running MTCNN again</p>\n\n<p>Also I was able to fit all 30 frames in a singly GPU batch, that combined with pytorch pin memory gave me good performance running mtcnn</p>",
          "rawMarkdown": "I am only ran MTCNN on every 10th frame of the original videos.  The averaged over all 300 frames, and reused the face bounding boxes from the original for all of the fakes rather than running MTCNN again\n\nAlso I was able to fit all 30 frames in a singly GPU batch, that combined with pytorch pin memory gave me good performance running mtcnn",
          "votes": 5
        },
        {
          "id": 700716,
          "postDate": "2019-12-22T13:49:37.850Z",
          "content": "<p><a href=\"/nikperi\">@nikperi</a> that won't work in test sets where you don't know the original video, \nhere are my time findings:</p>\n\n<p>it took 0.6 sec to skip through 200 frames\nand it takes MTCNN around 0.5 second to detect a face each frame. \naround 3 seconds to process a 10 sec video with GPU decoding 3 frames only\nit took me 30 minutes to analyze 400 videos (using the 2 threads kaggle gives us, but we  don't know the specs of the private test run) .. \nand the public test set is 4000 videos so I expect 6 hours just to go through videos in my case ( assuming they are 10 secs as well)</p>",
          "rawMarkdown": "@nikperi that won't work in test sets where you don't know the original video, \nhere are my time findings:\n\nit took 0.6 sec to skip through 200 frames\nand it takes MTCNN around 0.5 second to detect a face each frame. \naround 3 seconds to process a 10 sec video with GPU decoding 3 frames only\nit took me 30 minutes to analyze 400 videos (using the 2 threads kaggle gives us, but we  don't know the specs of the private test run) .. \nand the public test set is 4000 videos so I expect 6 hours just to go through videos in my case ( assuming they are 10 secs as well)\n\n"
        }
      ]
    },
    {
      "id": 701312,
      "postDate": "2019-12-23T10:54:41.053Z",
      "content": "<p>Thanks for the tip, I've update my face capturing function to include the optimization:</p>\n\n<pre>def get_n_faces(file_path,n=100,face_size=(96,96), verbose=0, show_faces=False, detector=None):\n    if verbose &gt; 0: print(file_path)\n    target_frames = n\n    frame_num = 0\n    reader = cv2.VideoCapture(file_path)\n    frame_count = int(reader.get(cv2.CAP_PROP_FRAME_COUNT))\n\n    mod = int(frame_count/target_frames)\n    mod = max(1,mod)\n    faces_counted = 0\n    tries = 0\n    faces = []\n    detections = []\n    if verbose &gt; 0: pbar = tqdm_notebook(total = target_frames)\n    while reader.isOpened() and len(faces)  0:\n                faces_counted += 1\n                new_box = [int(point / (scale_percent / 100)) for point in detections[0]['box']]\n                # get coordinates\n                x1, y1, width, height = new_box\n                x1 -= 20\n                y1 -= 100\n                x1 = max(0,x1)\n                y1 = max(0,y1)\n                width += 40\n                height += 200\n                x2, y2 = x1 + width, y1 + height\n                # extract face\n                try:\n                    face = image[y1:y2, x1:x2][...,::-1]\n                    face = ((face * (255.0/np.max(face)))+1).astype('uint8') #brighten face\n                    face = cv2.resize(face, face_size, interpolation = cv2.INTER_NEAREST)\n                    faces.append(face)\n                    if show_faces:\n                        plot.imshow(face)\n                        plot.show()\n                except:\n                    error = True # but we dont care!\n            if verbose &gt; 0: pbar.update(1)\n        frame_num += 1\n    # Finish reader\n    if verbose &gt; 0: pbar.close()\n    reader.release()\n    if verbose &gt; 0: print('counted:',faces_counted)\n    return_val = np.zeros([n, face_size[0], face_size[1], 3],dtype='uint8')\n    if len(faces) &gt; 0:\n        return_val[:len(faces)] = faces\n    if verbose &gt; 0: print(return_val.shape)\n    return return_val\n\n# for row in balanced_df[balanced_df['label']=='FAKE'].head()['file']:\n#     file_path = TRAIN_VIDEOS +'/'+row\n#     get_n_faces(file_path,100,show_faces=False)\n</pre>",
      "rawMarkdown": "Thanks for the tip, I've update my face capturing function to include the optimization:\n<pre>def get_n_faces(file_path,n=100,face_size=(96,96), verbose=0, show_faces=False, detector=None):\n    if verbose &gt; 0: print(file_path)\n    target_frames = n\n    frame_num = 0\n    reader = cv2.VideoCapture(file_path)\n    frame_count = int(reader.get(cv2.CAP_PROP_FRAME_COUNT))\n\n    mod = int(frame_count/target_frames)\n    mod = max(1,mod)\n    faces_counted = 0\n    tries = 0\n    faces = []\n    detections = []\n    if verbose &gt; 0: pbar = tqdm_notebook(total = target_frames)\n    while reader.isOpened() and len(faces) &lt; n and tries &lt; n*2:\n        tries +=1\n        _ = reader.grab()\n        if frame_num % mod == 0 or len(detections) == 0:\n            _, image = reader.retrieve()\n            scale_percent = 20 # percent of original size\n            width = int(image.shape[1] * scale_percent / 100)\n            height = int(image.shape[0] * scale_percent / 100)\n            dim = (width, height)\n            # resize image\n            resized = cv2.resize(image, dim, interpolation = cv2.INTER_NEAREST)\n            resized = (resized * (255.0/np.max(resized))).astype('uint8') #brighten image\n            img = resized[...,::-1]\n            detections = []\n            try:\n                detections = detector.detect_faces(img)\n            except:\n                print(\"detect error frame:\",frame_num,\"path:\",file_path)\n                #raise # if you want a headache, otherwise just try on another frame\n            if len(detections) &gt; 0:\n                faces_counted += 1\n                new_box = [int(point / (scale_percent / 100)) for point in detections[0]['box']]\n                # get coordinates\n                x1, y1, width, height = new_box\n                x1 -= 20\n                y1 -= 100\n                x1 = max(0,x1)\n                y1 = max(0,y1)\n                width += 40\n                height += 200\n                x2, y2 = x1 + width, y1 + height\n                # extract face\n                try:\n                    face = image[y1:y2, x1:x2][...,::-1]\n                    face = ((face * (255.0/np.max(face)))+1).astype('uint8') #brighten face\n                    face = cv2.resize(face, face_size, interpolation = cv2.INTER_NEAREST)\n                    faces.append(face)\n                    if show_faces:\n                        plot.imshow(face)\n                        plot.show()\n                except:\n                    error = True # but we dont care!\n            if verbose &gt; 0: pbar.update(1)\n        frame_num += 1\n    # Finish reader\n    if verbose &gt; 0: pbar.close()\n    reader.release()\n    if verbose &gt; 0: print('counted:',faces_counted)\n    return_val = np.zeros([n, face_size[0], face_size[1], 3],dtype='uint8')\n    if len(faces) &gt; 0:\n        return_val[:len(faces)] = faces\n    if verbose &gt; 0: print(return_val.shape)\n    return return_val\n    \n# for row in balanced_df[balanced_df['label']=='FAKE'].head()['file']:\n#     file_path = TRAIN_VIDEOS +'/'+row\n#     get_n_faces(file_path,100,show_faces=False)\n</pre>",
      "votes": 2,
      "replies": [
        {
          "id": 701772,
          "postDate": "2019-12-23T22:29:05.060Z",
          "content": "<p>Note that it's a good idea to check the return value from <code>reader.grab()</code> and <code>reader.retrieve()</code> as well, instead of simply ignoring it and assuming every single frame will read without errors. </p>",
          "rawMarkdown": "Note that it's a good idea to check the return value from `reader.grab()` and `reader.retrieve()` as well, instead of simply ignoring it and assuming every single frame will read without errors. ",
          "votes": 3
        },
        {
          "id": 701965,
          "postDate": "2019-12-24T05:31:41.147Z",
          "content": "<p>Thank you</p>",
          "rawMarkdown": "Thank you"
        }
      ]
    },
    {
      "id": 2007831,
      "postDate": "2022-10-28T14:21:36.337Z",
      "content": "<p>Great post.</p>",
      "rawMarkdown": "Great post."
    },
    {
      "id": 749164,
      "postDate": "2020-02-18T11:38:28.383Z",
      "content": "<p>Great post. Significant gains! Thanks man</p>",
      "rawMarkdown": "Great post. Significant gains! Thanks man"
    },
    {
      "id": 699979,
      "postDate": "2019-12-21T08:05:08.733Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice"
    },
    {
      "id": 748305,
      "postDate": "2020-02-17T11:15:41.790Z",
      "content": "<p>Thanks for sharing!! This really helps a lot!! \nHow do you think imageio_ffmpeg?\nI tried to compare it with v_cap.read() and I found the imageio_ffmpeg is a bit faster for 2 reasons. 1) it doesn’t need to do color space conversion (cv2.cvtColor(frame, cv2.COLOR_BRG2RGB)). 2) it can read desired frame directly, needless to do for loop. </p>",
      "rawMarkdown": "Thanks for sharing!! This really helps a lot!! \nHow do you think imageio_ffmpeg?\nI tried to compare it with v\\_cap.read() and I found the imageio\\_ffmpeg is a bit faster for 2 reasons. 1) it doesn’t need to do color space conversion (cv2.cvtColor(frame, cv2.COLOR\\_BRG2RGB)). 2) it can read desired frame directly, needless to do for loop. \n",
      "isDeleted": true
    },
    {
      "id": 699439,
      "postDate": "2019-12-20T13:08:59.257Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 699568,
          "postDate": "2019-12-20T15:44:06.217Z",
          "content": "<p>Won't that just read the first 10% of the whole video? <code>i</code> and <code>capture</code> are not linked so you're just calling read <code>num_frames/10</code> times</p>",
          "rawMarkdown": "Won't that just read the first 10% of the whole video? `i` and `capture` are not linked so you're just calling read `num_frames/10` times\n\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 701751,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2019-12-23T21:46:33.513000",
      "content": "<p>For me, I just take 10 frames. So I used\n<code>cap.set(cv2.CAP_PROP_POS_FRAMES, framenum)</code></p>\n\n<p>For example, this is the code for getting frame number 33: </p>\n\n<p><code>cap.set(cv2.CAPPROPPOSFRAMES, 33)\nret, framenumber33 = cap.read()</code></p>\n\n<p>According to this <a href=\"https://subscription.packtpub.com/book/application_development/9781788474443/1/ch01lvl1sec24/jumping-between-frames-in-video-files\">blog</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 715408,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-01-10T13:52:41.717000",
          "content": "<p>I don't know why but For me it is almost 4 times slower then retrieve-grab method.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1830962%2F4673fb4666d059e68a265ab22a816baf%2FScreenshot%20from%202020-01-10%2019-21-42.png?generation=1578664342203845&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716569,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-01-11T22:11:02.570000",
          "content": "<p>That's most likely because videos don't store the full data for each frame, except in so-called keyframes. So if the frame you're asking for is not a keyframe, OpenCV will have to go back to read the keyframe and then stream until the frame you're asking for anyway.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 716570,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-11T22:13:41.293000",
          "content": "<p>This works for me because I just need a single frame. This doesn't work for you because you will need quite a few frames.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 701062,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2019-12-23T03:20:08.593000",
      "content": "<p>For training, which you do on your own machine, not in a kernel, why not just find all the faces and save them to disk. Then use these for all training runs.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 699606,
      "author_name": "Amit Kumar",
      "author_url": "",
      "post_date": "2019-12-20T16:20:45.670000",
      "content": "<p>Here is a benchmark of scikit-video and opencv using python .\nto load a single frame\ntime taken by scikit-video : 0.005018000000000189\ntime taken by cv2 : 0.0228900000000003</p>\n\n<p>time taken to load 5637 frames sequentially</p>\n\n<p>time taken by scikit-video 8.149045999999998\ntime taken by cv2 : 55.439571</p>\n\n<p><a href=\"https://medium.com/@amkr/week-2-gsoc19-ccextractor-development-e4ffb6d54376\">blog link</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 699609,
          "author_name": "Amit Kumar",
          "author_url": "",
          "post_date": "2019-12-20T16:23:29.310000",
          "content": "<p>Multithreading can also be used to increase the efficiency of frames processing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 699667,
          "author_name": "Jonas Matuzas",
          "author_url": "",
          "post_date": "2019-12-20T18:18:28.160000",
          "content": "<p>How to install scikit-video without internet? Can I turn on internet pip install and turn off internet ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 699842,
          "author_name": "Amit Kumar",
          "author_url": "",
          "post_date": "2019-12-21T02:39:54.270000",
          "content": "<p>You will need internet to install it.\n<code>pip install scikit-video</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 699992,
          "author_name": "Jonas Matuzas",
          "author_url": "",
          "post_date": "2019-12-21T08:19:53.363000",
          "content": "<p>but is it legal in this competition to use pip install ? internet forbidden :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 700293,
          "author_name": "Danny",
          "author_url": "",
          "post_date": "2019-12-21T18:42:14.490000",
          "content": "<p>Dang this would be a great package to use but yeah, Jonas is right. It won't be possible to use this since the internet and custom packages aren't allowed </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 700700,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2019-12-22T13:15:45.260000",
          "content": "<p>Where does it say that in the rules? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 701002,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-12-23T00:49:57.297000",
          "content": "<p>We're allowed to install packages by uploading them as a dataset and then instialling with pip. Nothing different than just copying and pasting some code in so no reason why it would be excluded. When they say no custom packages they mean you can't use the old custom packages option kaggle used to have which would essentially be like a requirements.txt that was preinstalled before your kernel was initialized. Those aren't allowed because it would call the internet potentially and download stuff. </p>\n\n<p>That interface isn't even available on kaggle anymore. </p>\n\n<p>But also, tried skvideo locally and wasn't able to reproduce their speed-up. With skvideo I found it was actually slower.  opencv with grab and retrieve has been the fastest I have been able to find so far unfortunately. </p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 701019,
          "author_name": "David Austin",
          "author_url": "",
          "post_date": "2019-12-23T01:59:26.600000",
          "content": "<p>I also get no speedup from skvideo.  opencv seems to be the fastest.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 701308,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2019-12-23T10:44:02.293000",
          "content": "<p>In addition to uploading the packages and installing the with pip, for python only packages you can just put the code in directly and add the dataset to the path to load it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 710263,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2020-01-04T14:14:20.857000",
          "content": "<p>I'm getting 68 seconds for skvideo &amp; 69 seconds for OpenCV doing the same task (face detection using dlib on each frame). Not much between them!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 707559,
      "author_name": "Manideep",
      "author_url": "",
      "post_date": "2020-01-01T03:58:42.143000",
      "content": "<p>Followed your suggestion and tweaked my Video loading function. My data generator is significantly faster now. \nMy Keras data generator reads the input video (1 in 10 frames), uses <code>MTCNN</code> to detect faces and predict on them. \nBelow are the times taken to predict on 400 videos before and after the tweak.</p>\n\n<p>Before,\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1288210%2F852d6a3ff125c9cffd12865dfda5c417%2FScreenshot%202020-01-01%20at%209.25.21%20AM.png?generation=1577850978156060&amp;alt=media\" alt=\"\"></p>\n\n<p>After,\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1288210%2F770f446b0e72b834d3cee2d7c71f1251%2FScreenshot%202020-01-01%20at%209.25.35%20AM.png?generation=1577851018436037&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 698729,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-12-19T16:29:33.783000",
      "content": "<p>Thanks for the tip. Does yield some speed-up for me. What is the timing difference you have seen in your experience?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 698854,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2019-12-19T19:21:18.433000",
          "content": "<p>I think it's about twice as fast, but obviously this matters more if you're reading more of the video.</p>\n\n<p>Reading the video is obviously not the bottleneck in this competition (running the model is), but there are a lot of videos so any time savings still add up. (If you save 1 second per video, that saves over an hour on the 4000 test set videos.)</p>\n\n<p>I think imageio can read any frame using random access, which would be even faster but I haven't really looked into that in detail yet.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 698928,
          "author_name": "Simon Caby",
          "author_url": "",
          "post_date": "2019-12-19T22:38:58.207000",
          "content": "<p>Good idea. Still the bottleneck is MTCNN for a decent face recognition.... Thanks! </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 699124,
          "author_name": "Nikhil Peri",
          "author_url": "",
          "post_date": "2019-12-20T04:13:08.793000",
          "content": "<p>I am only ran MTCNN on every 10th frame of the original videos.  The averaged over all 300 frames, and reused the face bounding boxes from the original for all of the fakes rather than running MTCNN again</p>\n\n<p>Also I was able to fit all 30 frames in a singly GPU batch, that combined with pytorch pin memory gave me good performance running mtcnn</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 700716,
          "author_name": "B. Allabadi",
          "author_url": "",
          "post_date": "2019-12-22T13:49:37.850000",
          "content": "<p><a href=\"/nikperi\">@nikperi</a> that won't work in test sets where you don't know the original video, \nhere are my time findings:</p>\n\n<p>it took 0.6 sec to skip through 200 frames\nand it takes MTCNN around 0.5 second to detect a face each frame. \naround 3 seconds to process a 10 sec video with GPU decoding 3 frames only\nit took me 30 minutes to analyze 400 videos (using the 2 threads kaggle gives us, but we  don't know the specs of the private test run) .. \nand the public test set is 4000 videos so I expect 6 hours just to go through videos in my case ( assuming they are 10 secs as well)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 701312,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2019-12-23T10:54:41.053000",
      "content": "<p>Thanks for the tip, I've update my face capturing function to include the optimization:</p>\n\n<pre>def get_n_faces(file_path,n=100,face_size=(96,96), verbose=0, show_faces=False, detector=None):\n    if verbose &gt; 0: print(file_path)\n    target_frames = n\n    frame_num = 0\n    reader = cv2.VideoCapture(file_path)\n    frame_count = int(reader.get(cv2.CAP_PROP_FRAME_COUNT))\n\n    mod = int(frame_count/target_frames)\n    mod = max(1,mod)\n    faces_counted = 0\n    tries = 0\n    faces = []\n    detections = []\n    if verbose &gt; 0: pbar = tqdm_notebook(total = target_frames)\n    while reader.isOpened() and len(faces)  0:\n                faces_counted += 1\n                new_box = [int(point / (scale_percent / 100)) for point in detections[0]['box']]\n                # get coordinates\n                x1, y1, width, height = new_box\n                x1 -= 20\n                y1 -= 100\n                x1 = max(0,x1)\n                y1 = max(0,y1)\n                width += 40\n                height += 200\n                x2, y2 = x1 + width, y1 + height\n                # extract face\n                try:\n                    face = image[y1:y2, x1:x2][...,::-1]\n                    face = ((face * (255.0/np.max(face)))+1).astype('uint8') #brighten face\n                    face = cv2.resize(face, face_size, interpolation = cv2.INTER_NEAREST)\n                    faces.append(face)\n                    if show_faces:\n                        plot.imshow(face)\n                        plot.show()\n                except:\n                    error = True # but we dont care!\n            if verbose &gt; 0: pbar.update(1)\n        frame_num += 1\n    # Finish reader\n    if verbose &gt; 0: pbar.close()\n    reader.release()\n    if verbose &gt; 0: print('counted:',faces_counted)\n    return_val = np.zeros([n, face_size[0], face_size[1], 3],dtype='uint8')\n    if len(faces) &gt; 0:\n        return_val[:len(faces)] = faces\n    if verbose &gt; 0: print(return_val.shape)\n    return return_val\n\n# for row in balanced_df[balanced_df['label']=='FAKE'].head()['file']:\n#     file_path = TRAIN_VIDEOS +'/'+row\n#     get_n_faces(file_path,100,show_faces=False)\n</pre>",
      "votes": 2,
      "replies": [
        {
          "id": 701772,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2019-12-23T22:29:05.060000",
          "content": "<p>Note that it's a good idea to check the return value from <code>reader.grab()</code> and <code>reader.retrieve()</code> as well, instead of simply ignoring it and assuming every single frame will read without errors. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 701965,
          "author_name": "Sefa Kurtipek",
          "author_url": "",
          "post_date": "2019-12-24T05:31:41.147000",
          "content": "<p>Thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2007831,
      "author_name": "Reda Ghanem",
      "author_url": "",
      "post_date": "2022-10-28T14:21:36.337000",
      "content": "<p>Great post.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 749164,
      "author_name": "Akash",
      "author_url": "",
      "post_date": "2020-02-18T11:38:28.383000",
      "content": "<p>Great post. Significant gains! Thanks man</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 699979,
      "author_name": "Gaurav Shahi",
      "author_url": "",
      "post_date": "2019-12-21T08:05:08.733000",
      "content": "<p>nice</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 748305,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-17T11:15:41.790000",
      "content": "<p>Thanks for sharing!! This really helps a lot!! \nHow do you think imageio_ffmpeg?\nI tried to compare it with v_cap.read() and I found the imageio_ffmpeg is a bit faster for 2 reasons. 1) it doesn’t need to do color space conversion (cv2.cvtColor(frame, cv2.COLOR_BRG2RGB)). 2) it can read desired frame directly, needless to do for loop. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 699439,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-20T13:08:59.257000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 699568,
          "author_name": "Josh Asprey",
          "author_url": "",
          "post_date": "2019-12-20T15:44:06.217000",
          "content": "<p>Won't that just read the first 10% of the whole video? <code>i</code> and <code>capture</code> are not linked so you're just calling read <code>num_frames/10</code> times</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "698522": "Since loading videos is an important part of this challenge, I thought a topic on how to do this as fast as possible would be useful. :-)\n\nI'm not an expert on handling video in Python (yet -- one of the reasons why this competition is interesting to me) but I did notice the following:\n\nIf you're not interested in using all the frames from a video, for example only every 10th frame, don't do the following:\n\n```\ncapture = cv.VideoCapture(movie_path)\nfor i in range(0, num_frames):\n    ret, frame = capture.read()\n    if i % 10 == 0:\n        # do something with frame\n```\n\nInstead, it's faster to call `grab()`, which reads the next frame but doesn't decode it, and only call `retrieve()` on the frames you're interested in:\n\n```\ncapture = cv.VideoCapture(movie_path)\nfor i in range(0, num_frames):\n    ret = capture.grab()\n    if i % 10 == 0:\n        ret, frame = capture.retrieve()\n        # do something with frame\ncapture.release()\n```\n\nThis will already make quite a bit of difference in processing time.\n\nIf you've got any more tips, please reply below. 😄 ",
    "701751": "For me, I just take 10 frames. So I used\n`cap.set(cv2.CAP_PROP_POS_FRAMES, framenum)`\n\n For example, this is the code for getting frame number 33: \n\n`cap.set(cv2.CAPPROPPOSFRAMES, 33)\nret, framenumber33 = cap.read()`\n\n\nAccording to this [blog](https://subscription.packtpub.com/book/application_development/9781788474443/1/ch01lvl1sec24/jumping-between-frames-in-video-files)",
    "701062": "For training, which you do on your own machine, not in a kernel, why not just find all the faces and save them to disk. Then use these for all training runs.",
    "699606": "Here is a benchmark of scikit-video and opencv using python .\nto load a single frame\ntime taken by scikit-video : 0.005018000000000189\ntime taken by cv2 : 0.0228900000000003\n\ntime taken to load 5637 frames sequentially\n\ntime taken by scikit-video 8.149045999999998\ntime taken by cv2 : 55.439571\n\n[blog link](https://medium.com/@amkr/week-2-gsoc19-ccextractor-development-e4ffb6d54376)",
    "707559": "Followed your suggestion and tweaked my Video loading function. My data generator is significantly faster now. \nMy Keras data generator reads the input video (1 in 10 frames), uses `MTCNN` to detect faces and predict on them. \nBelow are the times taken to predict on 400 videos before and after the tweak.\n\nBefore,\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1288210%2F852d6a3ff125c9cffd12865dfda5c417%2FScreenshot%202020-01-01%20at%209.25.21%20AM.png?generation=1577850978156060&amp;alt=media)\n\nAfter,\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1288210%2F770f446b0e72b834d3cee2d7c71f1251%2FScreenshot%202020-01-01%20at%209.25.35%20AM.png?generation=1577851018436037&amp;alt=media)\n",
    "698729": "Thanks for the tip. Does yield some speed-up for me. What is the timing difference you have seen in your experience?",
    "701312": "Thanks for the tip, I've update my face capturing function to include the optimization:\n<pre>def get_n_faces(file_path,n=100,face_size=(96,96), verbose=0, show_faces=False, detector=None):\n    if verbose &gt; 0: print(file_path)\n    target_frames = n\n    frame_num = 0\n    reader = cv2.VideoCapture(file_path)\n    frame_count = int(reader.get(cv2.CAP_PROP_FRAME_COUNT))\n\n    mod = int(frame_count/target_frames)\n    mod = max(1,mod)\n    faces_counted = 0\n    tries = 0\n    faces = []\n    detections = []\n    if verbose &gt; 0: pbar = tqdm_notebook(total = target_frames)\n    while reader.isOpened() and len(faces) &lt; n and tries &lt; n*2:\n        tries +=1\n        _ = reader.grab()\n        if frame_num % mod == 0 or len(detections) == 0:\n            _, image = reader.retrieve()\n            scale_percent = 20 # percent of original size\n            width = int(image.shape[1] * scale_percent / 100)\n            height = int(image.shape[0] * scale_percent / 100)\n            dim = (width, height)\n            # resize image\n            resized = cv2.resize(image, dim, interpolation = cv2.INTER_NEAREST)\n            resized = (resized * (255.0/np.max(resized))).astype('uint8') #brighten image\n            img = resized[...,::-1]\n            detections = []\n            try:\n                detections = detector.detect_faces(img)\n            except:\n                print(\"detect error frame:\",frame_num,\"path:\",file_path)\n                #raise # if you want a headache, otherwise just try on another frame\n            if len(detections) &gt; 0:\n                faces_counted += 1\n                new_box = [int(point / (scale_percent / 100)) for point in detections[0]['box']]\n                # get coordinates\n                x1, y1, width, height = new_box\n                x1 -= 20\n                y1 -= 100\n                x1 = max(0,x1)\n                y1 = max(0,y1)\n                width += 40\n                height += 200\n                x2, y2 = x1 + width, y1 + height\n                # extract face\n                try:\n                    face = image[y1:y2, x1:x2][...,::-1]\n                    face = ((face * (255.0/np.max(face)))+1).astype('uint8') #brighten face\n                    face = cv2.resize(face, face_size, interpolation = cv2.INTER_NEAREST)\n                    faces.append(face)\n                    if show_faces:\n                        plot.imshow(face)\n                        plot.show()\n                except:\n                    error = True # but we dont care!\n            if verbose &gt; 0: pbar.update(1)\n        frame_num += 1\n    # Finish reader\n    if verbose &gt; 0: pbar.close()\n    reader.release()\n    if verbose &gt; 0: print('counted:',faces_counted)\n    return_val = np.zeros([n, face_size[0], face_size[1], 3],dtype='uint8')\n    if len(faces) &gt; 0:\n        return_val[:len(faces)] = faces\n    if verbose &gt; 0: print(return_val.shape)\n    return return_val\n    \n# for row in balanced_df[balanced_df['label']=='FAKE'].head()['file']:\n#     file_path = TRAIN_VIDEOS +'/'+row\n#     get_n_faces(file_path,100,show_faces=False)\n</pre>",
    "2007831": "Great post.",
    "749164": "Great post. Significant gains! Thanks man",
    "699979": "nice",
    "748305": "Thanks for sharing!! This really helps a lot!! \nHow do you think imageio_ffmpeg?\nI tried to compare it with v\\_cap.read() and I found the imageio\\_ffmpeg is a bit faster for 2 reasons. 1) it doesn’t need to do color space conversion (cv2.cvtColor(frame, cv2.COLOR\\_BRG2RGB)). 2) it can read desired frame directly, needless to do for loop. \n",
    "699439": ""
  }
}