{
  "id": 121716,
  "title": "FaceForensics++ experiment",
  "url": "/competitions/deepfake-detection-challenge/discussion/121716",
  "author_name": "Carlos Souza",
  "post_date": "2019-12-15T03:22:38.753000",
  "votes": 39,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Just quick read all shared papers, and tested FaceForensics++ pre-trained model. Results are pretty cool, here's some samples:</p>\n\n<p><strong>Real sample</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F17e9bed55f9da71196ec53bb39b89dc9%2Fabarnvbtwb.gif?generation=1576380120161560&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Fake sample</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fe1e531517bc51c6889098cf9d6bd2de8%2Faagfhgtpmv.gif?generation=1576380143703413&amp;alt=media\" alt=\"\"></p>\n\n<p>Cheers!</p>",
  "messages": [
    {
      "id": 695377,
      "postDate": "2019-12-15T03:22:38.753Z",
      "content": "<p>Just quick read all shared papers, and tested FaceForensics++ pre-trained model. Results are pretty cool, here's some samples:</p>\n\n<p><strong>Real sample</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F17e9bed55f9da71196ec53bb39b89dc9%2Fabarnvbtwb.gif?generation=1576380120161560&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Fake sample</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fe1e531517bc51c6889098cf9d6bd2de8%2Faagfhgtpmv.gif?generation=1576380143703413&amp;alt=media\" alt=\"\"></p>\n\n<p>Cheers!</p>",
      "rawMarkdown": "Just quick read all shared papers, and tested FaceForensics++ pre-trained model. Results are pretty cool, here's some samples:\n\n**Real sample**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F17e9bed55f9da71196ec53bb39b89dc9%2Fabarnvbtwb.gif?generation=1576380120161560&amp;alt=media)\n\n**Fake sample**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fe1e531517bc51c6889098cf9d6bd2de8%2Faagfhgtpmv.gif?generation=1576380143703413&amp;alt=media)\n\nCheers!",
      "votes": 39
    },
    {
      "id": 700613,
      "postDate": "2019-12-22T10:03:40.730Z",
      "content": "<p>I have managed to achieve ~65 fps face detection on 8xCPU + 1xGPU (1070Ti) + MTCNN at full resolution (1920x1080). Keys:</p>\n\n<ul>\n<li>Opencv for video reading/processing</li>\n<li>Use minsize=60px to keep the pyramidal downscaling under a limit</li>\n<li>No batching (process one frame at a time)</li>\n<li>Use as many processes as CPUs</li>\n</ul>\n\n<p>Overall 8 videos are processed every ~35 or 40s. </p>\n\n<p>Bottleneck is apparently in the CPU (100%) while the GPU stays at a steady ~85%</p>\n\n<p>This approach may benefit from using scikit for video processing as suggested in another thread.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3076212%2Fa0451a8611604d17a819a6cb1872ff27%2FCaptura%20de%20pantalla%202019-12-22%20a%20las%2010.59.50.png?generation=1577008892533969&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I have managed to achieve ~65 fps face detection on 8xCPU + 1xGPU (1070Ti) + MTCNN at full resolution (1920x1080). Keys:\n\n- Opencv for video reading/processing\n- Use minsize=60px to keep the pyramidal downscaling under a limit\n- No batching (process one frame at a time)\n- Use as many processes as CPUs\n\nOverall 8 videos are processed every ~35 or 40s. \n\nBottleneck is apparently in the CPU (100%) while the GPU stays at a steady ~85%\n\nThis approach may benefit from using scikit for video processing as suggested in another thread.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3076212%2Fa0451a8611604d17a819a6cb1872ff27%2FCaptura%20de%20pantalla%202019-12-22%20a%20las%2010.59.50.png?generation=1577008892533969&amp;alt=media)\n",
      "votes": 5
    },
    {
      "id": 695787,
      "postDate": "2019-12-15T15:12:01.923Z",
      "content": "<p>Btw, I scanned through all discussion posts and downloaded all papers. All can be <a href=\"https://www.dropbox.com/sh/ivio1l39chxalkl/AAD1_CXOMrYb8ZdAho0OqB4na?dl=0\">found here</a>.</p>",
      "rawMarkdown": "Btw, I scanned through all discussion posts and downloaded all papers. All can be [found here](https://www.dropbox.com/sh/ivio1l39chxalkl/AAD1_CXOMrYb8ZdAho0OqB4na?dl=0).",
      "votes": 6
    },
    {
      "id": 697292,
      "postDate": "2019-12-17T17:24:00.517Z",
      "content": "<p>I created a notebook that is still a work in progress, but uses the pretrained models from this paper:\n<a href=\"https://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet\">https://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet</a></p>",
      "rawMarkdown": "I created a notebook that is still a work in progress, but uses the pretrained models from this paper:\nhttps://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet",
      "votes": 1,
      "replies": [
        {
          "id": 697307,
          "postDate": "2019-12-17T17:35:19.633Z",
          "content": "<p>nevermind, just figured it out! Great job, thx for sharing!</p>",
          "rawMarkdown": "nevermind, just figured it out! Great job, thx for sharing!\n",
          "votes": 1
        },
        {
          "id": 698095,
          "postDate": "2019-12-18T19:48:42.683Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> - I'm curious if you figured out how to speed up face detection using the GPU. Currently the bottleneck for me is face detection using dlib is pretty slow. I haven't tried compiling with GPU support - did you have any luck with that?</p>",
          "rawMarkdown": "@carlossouza - I'm curious if you figured out how to speed up face detection using the GPU. Currently the bottleneck for me is face detection using dlib is pretty slow. I haven't tried compiling with GPU support - did you have any luck with that?"
        },
        {
          "id": 698201,
          "postDate": "2019-12-18T23:28:49.447Z",
          "content": "<p>Hey <a href=\"/robikscube\">@robikscube</a> ,</p>\n\n<p>Just finished doing several experiments with different VM setups and code. Writing code to squeeze all processing power is very tricky, installing all GPU drivers and compiling all packages to use them is a real hassle, but I managed to get it working.</p>\n\n<p>The best I got so far in face detection is <strong>8.3 fps</strong>. This means ~4 hours to face detect all 120,000 frames from the 400 test videos. If instead of using all frames in a given video, we use 10 initial frames (approach used in FaceForensics++ paper), this drops to 8min.</p>\n\n<p>The problem is in training. There are 119,146 video files. Face detecting all frames from all these videos would take ~50 days of non-stop processing, or ~US$ 1,400 with the setup I'm using. That's out of my budget, and probably an overkill - there must be a more efficient way. If we chose to face detect only 10 frames in training, that would mean ~40 hours of non-stop processing, or US$ 48. That's more doable.</p>\n\n<p>Now, the questions are: \n- What's the optimal # of frames to use in each video? Is 10/video enough?\n- If 10 is enough, how to get them? First 10? Randomly select within the video?\n- Is there a better/dynamic way to set this #? (E.g. if after reading first 10 frames, the algo is confident enough, classify as fake or real; if not, read more 10 frames)\n- Is there another more efficient way to do face detection, that yields results at better rate than 8.3 fps without compromising accuracy?</p>\n\n<p>Cheers</p>",
          "rawMarkdown": "Hey @robikscube ,\n\nJust finished doing several experiments with different VM setups and code. Writing code to squeeze all processing power is very tricky, installing all GPU drivers and compiling all packages to use them is a real hassle, but I managed to get it working.\n\nThe best I got so far in face detection is **8.3 fps**. This means ~4 hours to face detect all 120,000 frames from the 400 test videos. If instead of using all frames in a given video, we use 10 initial frames (approach used in FaceForensics++ paper), this drops to 8min.\n\nThe problem is in training. There are 119,146 video files. Face detecting all frames from all these videos would take ~50 days of non-stop processing, or ~US$ 1,400 with the setup I'm using. That's out of my budget, and probably an overkill - there must be a more efficient way. If we chose to face detect only 10 frames in training, that would mean ~40 hours of non-stop processing, or US$ 48. That's more doable.\n\nNow, the questions are: \n- What's the optimal # of frames to use in each video? Is 10/video enough?\n- If 10 is enough, how to get them? First 10? Randomly select within the video?\n- Is there a better/dynamic way to set this #? (E.g. if after reading first 10 frames, the algo is confident enough, classify as fake or real; if not, read more 10 frames)\n- Is there another more efficient way to do face detection, that yields results at better rate than 8.3 fps without compromising accuracy?\n\nCheers",
          "votes": 3
        },
        {
          "id": 698203,
          "postDate": "2019-12-18T23:34:00.353Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> - looks like we're taking similar approaches, but watch out. You said it takes ~4 hours for the 400 test videos. The test set is actually much larger than 400 videos- I didn't realize this at first, but when you submit your results in the kernel it runs on a different, and much bigger dataset (I'm guessing 4x to 6x larger). I got my pipeline working yesterday except I ran into this issue when doing the final submission, it's frustrating.</p>",
          "rawMarkdown": "@carlossouza - looks like we're taking similar approaches, but watch out. You said it takes ~4 hours for the 400 test videos. The test set is actually much larger than 400 videos- I didn't realize this at first, but when you submit your results in the kernel it runs on a different, and much bigger dataset (I'm guessing 4x to 6x larger). I got my pipeline working yesterday except I ran into this issue when doing the final submission, it's frustrating."
        },
        {
          "id": 698207,
          "postDate": "2019-12-18T23:47:32.690Z",
          "content": "<p>Humm... thanks, I imagined that. So, if we use 10 frames per video, and private test dataset is 6x larger than the public one, this would mean 8 x 6 = 48min to face detect them. Also, trying to face detect all 300 frames per video wouldn't be possible in the 9 hours limit.</p>\n\n<p>My next step now is to check the accuracy of the pretrained model in the training data, using 10 frames per video, do a 1st submission with these parameters, and compare. Then, start experimenting with it.</p>\n\n<p>Cheers</p>",
          "rawMarkdown": "Humm... thanks, I imagined that. So, if we use 10 frames per video, and private test dataset is 6x larger than the public one, this would mean 8 x 6 = 48min to face detect them. Also, trying to face detect all 300 frames per video wouldn't be possible in the 9 hours limit.\n\nMy next step now is to check the accuracy of the pretrained model in the training data, using 10 frames per video, do a 1st submission with these parameters, and compare. Then, start experimenting with it.\n\nCheers",
          "votes": 1
        },
        {
          "id": 698504,
          "postDate": "2019-12-19T10:21:11.280Z",
          "content": "<blockquote>\n  <p>The best I got so far in face detection is 8.3 fps.</p>\n</blockquote>\n\n<p>What face detector are you using? I'm putting some effort into getting a fast face detector (such as BlazeFace) to work, since this is going to be a big bottleneck for this competition.</p>",
          "rawMarkdown": "&gt; The best I got so far in face detection is 8.3 fps.\n\nWhat face detector are you using? I'm putting some effort into getting a fast face detector (such as BlazeFace) to work, since this is going to be a big bottleneck for this competition."
        },
        {
          "id": 698647,
          "postDate": "2019-12-19T14:11:08.803Z",
          "content": "<p>Using Dlib. There’s several options, and some benchmark studies. Here’s one:\n<a href=\"https://medium.com/nodeflux/performance-showdown-of-publicly-available-face-detection-model-7c725747094a\">https://medium.com/nodeflux/performance-showdown-of-publicly-available-face-detection-model-7c725747094a</a>\nI’ve seen algorithms claiming up to 1500 fps speed in face detection, which seems obviously too good to be true. Of course it all boils down to our choice of hardware and how much accuracy we are willing to sacrifice. It’s a trade-off..</p>",
          "rawMarkdown": "Using Dlib. There’s several options, and some benchmark studies. Here’s one:\nhttps://medium.com/nodeflux/performance-showdown-of-publicly-available-face-detection-model-7c725747094a\nI’ve seen algorithms claiming up to 1500 fps speed in face detection, which seems obviously too good to be true. Of course it all boils down to our choice of hardware and how much accuracy we are willing to sacrifice. It’s a trade-off.."
        },
        {
          "id": 698674,
          "postDate": "2019-12-19T14:58:20.557Z",
          "content": "<p>I implementing this which claims to speed up fps processing using <code>FileVideoStream</code> but I didn't see a noticeable speedup:</p>\n\n<p><a href=\"https://www.pyimagesearch.com/2017/02/06/faster-video-file-fps-with-cv2-videocapture-and-opencv/\">https://www.pyimagesearch.com/2017/02/06/faster-video-file-fps-with-cv2-videocapture-and-opencv/</a>\n<a href=\"https://github.com/jrosebr1/imutils/blob/master/imutils/video/filevideostream.py\">https://github.com/jrosebr1/imutils/blob/master/imutils/video/filevideostream.py</a></p>",
          "rawMarkdown": "I implementing this which claims to speed up fps processing using `FileVideoStream` but I didn't see a noticeable speedup:\n\nhttps://www.pyimagesearch.com/2017/02/06/faster-video-file-fps-with-cv2-videocapture-and-opencv/\nhttps://github.com/jrosebr1/imutils/blob/master/imutils/video/filevideostream.py"
        },
        {
          "id": 698681,
          "postDate": "2019-12-19T15:10:10.057Z",
          "content": "<p>That's because the bottleneck is not in reading frames from videos, but in doing face detection in the frames. For perspective, a naive cv2 videocapture single-thread implementation reads 300 frames from a video in 2-2.5s - pretty fast. Btw, are you on Slack (kagglenoobs)?</p>",
          "rawMarkdown": "That's because the bottleneck is not in reading frames from videos, but in doing face detection in the frames. For perspective, a naive cv2 videocapture single-thread implementation reads 300 frames from a video in 2-2.5s - pretty fast. Btw, are you on Slack (kagglenoobs)?"
        },
        {
          "id": 705081,
          "postDate": "2019-12-28T11:41:10.793Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Have You thought about interpolating the face position between frames within a single video? I haven't done any research yet, but according to what You've written, face-detecting say every 10th frame should provide a pretty good guess of where the face was located in the meantime, as people on the videos are moving relatively slowly, in fact they are rarely moving at all. This may possibly be a good means for boosting performance as the bottleneck is in doing face detection, right? This would be a way of using the information that the frames form a consistent video, should we not forget about it! </p>",
          "rawMarkdown": "@carlossouza Have You thought about interpolating the face position between frames within a single video? I haven't done any research yet, but according to what You've written, face-detecting say every 10th frame should provide a pretty good guess of where the face was located in the meantime, as people on the videos are moving relatively slowly, in fact they are rarely moving at all. This may possibly be a good means for boosting performance as the bottleneck is in doing face detection, right? This would be a way of using the information that the frames form a consistent video, should we not forget about it! ",
          "votes": 1
        },
        {
          "id": 705112,
          "postDate": "2019-12-28T12:55:28.403Z",
          "content": "<p>Yes <a href=\"/snufkin77\">@snufkin77</a> ! That’s in my roadmap!\nI actually implemented first 2 other ideas: multiprocessing and resizing images progressively (i.e. start resizing frames to a small square, and try predicting faces... if not detected after N frames, resize to a slightly larger square and try again). These 2 ideas helped me a lot!\nBefore I implement your suggestion, I have to finish dealing with class imbalance. In my pov, interpolating face boxes every N frames will boil down to a tradeoff between speed (low N) and accuracy/score (high N). Right now my model has a high bias to predict videos as fake, generating a lot of false positives and few true negatives (real videos). If I don’t finish solving this first, all ideas in my roadmap that must be evaluated with scores rather than just speed would be compromised. Does it make sense?\nAnyway, it’s in my bucketlist, will report progress soon! Thanks!</p>",
          "rawMarkdown": "Yes @snufkin77 ! That’s in my roadmap!\nI actually implemented first 2 other ideas: multiprocessing and resizing images progressively (i.e. start resizing frames to a small square, and try predicting faces... if not detected after N frames, resize to a slightly larger square and try again). These 2 ideas helped me a lot!\nBefore I implement your suggestion, I have to finish dealing with class imbalance. In my pov, interpolating face boxes every N frames will boil down to a tradeoff between speed (low N) and accuracy/score (high N). Right now my model has a high bias to predict videos as fake, generating a lot of false positives and few true negatives (real videos). If I don’t finish solving this first, all ideas in my roadmap that must be evaluated with scores rather than just speed would be compromised. Does it make sense?\nAnyway, it’s in my bucketlist, will report progress soon! Thanks!"
        },
        {
          "id": 705128,
          "postDate": "2019-12-28T13:39:29.373Z",
          "content": "<p>Sure! Thanks for all observations and results You already shared. If I manage to figure out anything myself, I'll share too</p>",
          "rawMarkdown": "Sure! Thanks for all observations and results You already shared. If I manage to figure out anything myself, I'll share too"
        },
        {
          "id": 705983,
          "postDate": "2019-12-29T19:21:48.613Z",
          "content": "<p>If you are only selecting 10 frames. I would pick the first and last ones. Which leaves us with picking another 8. I will assume for ease every video is 30 seconds long. We don't want to pick frames to close to each other but to far means we run out of video. If you select the other 8 frames somewhere between 2-4 seconds. This should give you the last 8 with enough interval to train properly. But this is a total guess without testing. Someway of value testing the 10 frames would be extremely useful. </p>",
          "rawMarkdown": "If you are only selecting 10 frames. I would pick the first and last ones. Which leaves us with picking another 8. I will assume for ease every video is 30 seconds long. We don't want to pick frames to close to each other but to far means we run out of video. If you select the other 8 frames somewhere between 2-4 seconds. This should give you the last 8 with enough interval to train properly. But this is a total guess without testing. Someway of value testing the 10 frames would be extremely useful. "
        },
        {
          "id": 706047,
          "postDate": "2019-12-29T20:56:27.077Z",
          "content": "<p>Most videos have 10 seconds, with 300 frames.\nBut you have a valid point: if not using all frames but a selection, intuitively it makes sense not to select adjacent frames. The only way to know for sure, though, is testing :)</p>",
          "rawMarkdown": "Most videos have 10 seconds, with 300 frames.\nBut you have a valid point: if not using all frames but a selection, intuitively it makes sense not to select adjacent frames. The only way to know for sure, though, is testing :)"
        },
        {
          "id": 706071,
          "postDate": "2019-12-29T22:24:54.497Z",
          "content": "<p>Yeah I didn't look at the actual data set. I just assumed in a 30 second video there would be atleast 1 frame per second </p>",
          "rawMarkdown": "Yeah I didn't look at the actual data set. I just assumed in a 30 second video there would be atleast 1 frame per second "
        }
      ]
    },
    {
      "id": 697007,
      "postDate": "2019-12-17T11:07:35.897Z",
      "content": "<p>Nice! 💯 </p>",
      "rawMarkdown": "Nice! 💯 ",
      "votes": -1
    },
    {
      "id": 2638119,
      "postDate": "2024-02-06T05:07:15.620Z",
      "content": "<p>I am not able to run the faceforensics++ dataset script. can anyone help me  access this dataset</p>",
      "rawMarkdown": "I am not able to run the faceforensics++ dataset script. can anyone help me  access this dataset"
    },
    {
      "id": 704598,
      "postDate": "2019-12-27T17:17:27.290Z",
      "content": "<p><a href=\"/bacterio\">@bacterio</a> \nI struggled to replicate your results... learned a lot in the process.\nHowever, the max speed I got to detect faces in 8 videos at full resolution, in parallel, was:</p>\n\n<p>| GPU        | FPS |\n|------------|-----|\n| Tesla K80  | 22  |\n| Tesla P4   | 28  |\n| Tesla T4   | 31  |\n| Tesla V100 | 28  |</p>\n\n<p>All experiments were done in Google Cloud, 8 vCPUs / 30 GB memory, and the same key 4 constraints you used (OpenCV, min face size = 60px, no batching, using a pool with 8 processes).</p>\n\n<p>I'm using PyTorch, and this <a href=\"https://github.com/timesler/facenet-pytorch\">facenet-pytorch</a> MTCNN implementation. </p>\n\n<p>The only explanation I can think of for the brutal difference in performance vs. what you achieved is the MTCNN implementation... I will try <a href=\"https://github.com/ipazc/mtcnn\">this TensorFlow implementation</a>. Which one did you use?</p>\n\n<p>Cheers!</p>",
      "rawMarkdown": "@bacterio \nI struggled to replicate your results... learned a lot in the process.\nHowever, the max speed I got to detect faces in 8 videos at full resolution, in parallel, was:\n\n| GPU        | FPS |\n|------------|-----|\n| Tesla K80  | 22  |\n| Tesla P4   | 28  |\n| Tesla T4   | 31  |\n| Tesla V100 | 28  |\n\nAll experiments were done in Google Cloud, 8 vCPUs / 30 GB memory, and the same key 4 constraints you used (OpenCV, min face size = 60px, no batching, using a pool with 8 processes).\n\nI'm using PyTorch, and this [facenet-pytorch](https://github.com/timesler/facenet-pytorch) MTCNN implementation. \n\nThe only explanation I can think of for the brutal difference in performance vs. what you achieved is the MTCNN implementation... I will try [this TensorFlow implementation](https://github.com/ipazc/mtcnn). Which one did you use?\n\nCheers!",
      "replies": [
        {
          "id": 704709,
          "postDate": "2019-12-27T21:54:02.340Z",
          "content": "<p>I used this one: <a href=\"https://github.com/faciallab/FaceDetector\">https://github.com/faciallab/FaceDetector</a></p>\n\n<p>The FPS I mantioned are with minimal pre-processing, ie. just BGR -&gt; RGB and basic histogram equalization.</p>\n\n<p>Also I realized a few videos are &lt;1920x1080 so minsize should depend on the actual video resolution.</p>",
          "rawMarkdown": "I used this one: https://github.com/faciallab/FaceDetector\n\nThe FPS I mantioned are with minimal pre-processing, ie. just BGR -&gt; RGB and basic histogram equalization.\n\nAlso I realized a few videos are &lt;1920x1080 so minsize should depend on the actual video resolution.",
          "votes": 1
        }
      ]
    },
    {
      "id": 700527,
      "postDate": "2019-12-22T06:38:48.410Z",
      "content": "<p>This is great.</p>",
      "rawMarkdown": "This is great."
    },
    {
      "id": 700058,
      "postDate": "2019-12-21T11:35:20.360Z",
      "content": "<p>I found a GitHub repo questioning the ability of FaceForensics++ to detect all kinds of deepfakes, you'll probably be interesting in this repo! See the discussion here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122603\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122603</a></p>",
      "rawMarkdown": "I found a GitHub repo questioning the ability of FaceForensics++ to detect all kinds of deepfakes, you'll probably be interesting in this repo! See the discussion here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122603"
    },
    {
      "id": 697089,
      "postDate": "2019-12-17T13:22:53.663Z",
      "content": "<p>Wow this is really cool. In fact it is a new revolutionary technology</p>",
      "rawMarkdown": "Wow this is really cool. In fact it is a new revolutionary technology"
    },
    {
      "id": 696783,
      "postDate": "2019-12-17T03:58:02.877Z",
      "content": "<p>Cool!</p>",
      "rawMarkdown": "Cool!"
    },
    {
      "id": 695789,
      "postDate": "2019-12-15T15:16:57.960Z",
      "content": "<p>I'm curious to see what this would score on the leaderboard. 😄 </p>",
      "rawMarkdown": "I'm curious to see what this would score on the leaderboard. 😄 ",
      "replies": [
        {
          "id": 695796,
          "postDate": "2019-12-15T15:26:27.837Z",
          "content": "<p>Me too, working on that ;)</p>",
          "rawMarkdown": "Me too, working on that ;)",
          "votes": 2
        },
        {
          "id": 696814,
          "postDate": "2019-12-17T05:18:28.833Z",
          "content": "<p>Hi <a href=\"/carlossouza\">@carlossouza</a> ,\nI'm currently studying the paper. Kindly share your findings when you're done with experiments.</p>",
          "rawMarkdown": "Hi @carlossouza ,\nI'm currently studying the paper. Kindly share your findings when you're done with experiments."
        },
        {
          "id": 697023,
          "postDate": "2019-12-17T11:32:41.643Z",
          "content": "<p>This <a href=\"http://openaccess.thecvf.com/content_CVPRW_2019/papers/Media%20Forensics/Sabir_Recurrent_Convolutional_Strategies_for_Face_Manipulation_Detection_in_Videos_CVPRW_2019_paper.pdf\">paper</a> seems interesting. Might be worth looking into.</p>",
          "rawMarkdown": "This [paper](http://openaccess.thecvf.com/content_CVPRW_2019/papers/Media%20Forensics/Sabir_Recurrent_Convolutional_Strategies_for_Face_Manipulation_Detection_in_Videos_CVPRW_2019_paper.pdf) seems interesting. Might be worth looking into."
        },
        {
          "id": 697309,
          "postDate": "2019-12-17T17:38:09.477Z",
          "content": "<p><a href=\"/abyaadrafid\">@abyaadrafid</a> Currently I’m struggling to make sure dlib is fully utilizing GPUs/CUDA. That’s not easy/out-of-the-box behavior... as soon as I figure out how to do it, I’ll share :)</p>",
          "rawMarkdown": "@abyaadrafid Currently I’m struggling to make sure dlib is fully utilizing GPUs/CUDA. That’s not easy/out-of-the-box behavior... as soon as I figure out how to do it, I’ll share :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 706097,
      "postDate": "2019-12-29T23:12:16.153Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 705296,
      "postDate": "2019-12-28T18:24:18.693Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 700613,
      "author_name": "Javier Martín",
      "author_url": "",
      "post_date": "2019-12-22T10:03:40.730000",
      "content": "<p>I have managed to achieve ~65 fps face detection on 8xCPU + 1xGPU (1070Ti) + MTCNN at full resolution (1920x1080). Keys:</p>\n\n<ul>\n<li>Opencv for video reading/processing</li>\n<li>Use minsize=60px to keep the pyramidal downscaling under a limit</li>\n<li>No batching (process one frame at a time)</li>\n<li>Use as many processes as CPUs</li>\n</ul>\n\n<p>Overall 8 videos are processed every ~35 or 40s. </p>\n\n<p>Bottleneck is apparently in the CPU (100%) while the GPU stays at a steady ~85%</p>\n\n<p>This approach may benefit from using scikit for video processing as suggested in another thread.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3076212%2Fa0451a8611604d17a819a6cb1872ff27%2FCaptura%20de%20pantalla%202019-12-22%20a%20las%2010.59.50.png?generation=1577008892533969&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 695787,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2019-12-15T15:12:01.923000",
      "content": "<p>Btw, I scanned through all discussion posts and downloaded all papers. All can be <a href=\"https://www.dropbox.com/sh/ivio1l39chxalkl/AAD1_CXOMrYb8ZdAho0OqB4na?dl=0\">found here</a>.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 697292,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2019-12-17T17:24:00.517000",
      "content": "<p>I created a notebook that is still a work in progress, but uses the pretrained models from this paper:\n<a href=\"https://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet\">https://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 697307,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-17T17:35:19.633000",
          "content": "<p>nevermind, just figured it out! Great job, thx for sharing!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 698095,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2019-12-18T19:48:42.683000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> - I'm curious if you figured out how to speed up face detection using the GPU. Currently the bottleneck for me is face detection using dlib is pretty slow. I haven't tried compiling with GPU support - did you have any luck with that?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 698201,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-18T23:28:49.447000",
          "content": "<p>Hey <a href=\"/robikscube\">@robikscube</a> ,</p>\n\n<p>Just finished doing several experiments with different VM setups and code. Writing code to squeeze all processing power is very tricky, installing all GPU drivers and compiling all packages to use them is a real hassle, but I managed to get it working.</p>\n\n<p>The best I got so far in face detection is <strong>8.3 fps</strong>. This means ~4 hours to face detect all 120,000 frames from the 400 test videos. If instead of using all frames in a given video, we use 10 initial frames (approach used in FaceForensics++ paper), this drops to 8min.</p>\n\n<p>The problem is in training. There are 119,146 video files. Face detecting all frames from all these videos would take ~50 days of non-stop processing, or ~US$ 1,400 with the setup I'm using. That's out of my budget, and probably an overkill - there must be a more efficient way. If we chose to face detect only 10 frames in training, that would mean ~40 hours of non-stop processing, or US$ 48. That's more doable.</p>\n\n<p>Now, the questions are: \n- What's the optimal # of frames to use in each video? Is 10/video enough?\n- If 10 is enough, how to get them? First 10? Randomly select within the video?\n- Is there a better/dynamic way to set this #? (E.g. if after reading first 10 frames, the algo is confident enough, classify as fake or real; if not, read more 10 frames)\n- Is there another more efficient way to do face detection, that yields results at better rate than 8.3 fps without compromising accuracy?</p>\n\n<p>Cheers</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 698203,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2019-12-18T23:34:00.353000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> - looks like we're taking similar approaches, but watch out. You said it takes ~4 hours for the 400 test videos. The test set is actually much larger than 400 videos- I didn't realize this at first, but when you submit your results in the kernel it runs on a different, and much bigger dataset (I'm guessing 4x to 6x larger). I got my pipeline working yesterday except I ran into this issue when doing the final submission, it's frustrating.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 698207,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-18T23:47:32.690000",
          "content": "<p>Humm... thanks, I imagined that. So, if we use 10 frames per video, and private test dataset is 6x larger than the public one, this would mean 8 x 6 = 48min to face detect them. Also, trying to face detect all 300 frames per video wouldn't be possible in the 9 hours limit.</p>\n\n<p>My next step now is to check the accuracy of the pretrained model in the training data, using 10 frames per video, do a 1st submission with these parameters, and compare. Then, start experimenting with it.</p>\n\n<p>Cheers</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 698504,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2019-12-19T10:21:11.280000",
          "content": "<blockquote>\n  <p>The best I got so far in face detection is 8.3 fps.</p>\n</blockquote>\n\n<p>What face detector are you using? I'm putting some effort into getting a fast face detector (such as BlazeFace) to work, since this is going to be a big bottleneck for this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 698647,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-19T14:11:08.803000",
          "content": "<p>Using Dlib. There’s several options, and some benchmark studies. Here’s one:\n<a href=\"https://medium.com/nodeflux/performance-showdown-of-publicly-available-face-detection-model-7c725747094a\">https://medium.com/nodeflux/performance-showdown-of-publicly-available-face-detection-model-7c725747094a</a>\nI’ve seen algorithms claiming up to 1500 fps speed in face detection, which seems obviously too good to be true. Of course it all boils down to our choice of hardware and how much accuracy we are willing to sacrifice. It’s a trade-off..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 698674,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2019-12-19T14:58:20.557000",
          "content": "<p>I implementing this which claims to speed up fps processing using <code>FileVideoStream</code> but I didn't see a noticeable speedup:</p>\n\n<p><a href=\"https://www.pyimagesearch.com/2017/02/06/faster-video-file-fps-with-cv2-videocapture-and-opencv/\">https://www.pyimagesearch.com/2017/02/06/faster-video-file-fps-with-cv2-videocapture-and-opencv/</a>\n<a href=\"https://github.com/jrosebr1/imutils/blob/master/imutils/video/filevideostream.py\">https://github.com/jrosebr1/imutils/blob/master/imutils/video/filevideostream.py</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 698681,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-19T15:10:10.057000",
          "content": "<p>That's because the bottleneck is not in reading frames from videos, but in doing face detection in the frames. For perspective, a naive cv2 videocapture single-thread implementation reads 300 frames from a video in 2-2.5s - pretty fast. Btw, are you on Slack (kagglenoobs)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 705081,
          "author_name": "Filip Strzałka",
          "author_url": "",
          "post_date": "2019-12-28T11:41:10.793000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Have You thought about interpolating the face position between frames within a single video? I haven't done any research yet, but according to what You've written, face-detecting say every 10th frame should provide a pretty good guess of where the face was located in the meantime, as people on the videos are moving relatively slowly, in fact they are rarely moving at all. This may possibly be a good means for boosting performance as the bottleneck is in doing face detection, right? This would be a way of using the information that the frames form a consistent video, should we not forget about it! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 705112,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-28T12:55:28.403000",
          "content": "<p>Yes <a href=\"/snufkin77\">@snufkin77</a> ! That’s in my roadmap!\nI actually implemented first 2 other ideas: multiprocessing and resizing images progressively (i.e. start resizing frames to a small square, and try predicting faces... if not detected after N frames, resize to a slightly larger square and try again). These 2 ideas helped me a lot!\nBefore I implement your suggestion, I have to finish dealing with class imbalance. In my pov, interpolating face boxes every N frames will boil down to a tradeoff between speed (low N) and accuracy/score (high N). Right now my model has a high bias to predict videos as fake, generating a lot of false positives and few true negatives (real videos). If I don’t finish solving this first, all ideas in my roadmap that must be evaluated with scores rather than just speed would be compromised. Does it make sense?\nAnyway, it’s in my bucketlist, will report progress soon! Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 705128,
          "author_name": "Filip Strzałka",
          "author_url": "",
          "post_date": "2019-12-28T13:39:29.373000",
          "content": "<p>Sure! Thanks for all observations and results You already shared. If I manage to figure out anything myself, I'll share too</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 705983,
          "author_name": "M00100100100w",
          "author_url": "",
          "post_date": "2019-12-29T19:21:48.613000",
          "content": "<p>If you are only selecting 10 frames. I would pick the first and last ones. Which leaves us with picking another 8. I will assume for ease every video is 30 seconds long. We don't want to pick frames to close to each other but to far means we run out of video. If you select the other 8 frames somewhere between 2-4 seconds. This should give you the last 8 with enough interval to train properly. But this is a total guess without testing. Someway of value testing the 10 frames would be extremely useful. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706047,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-29T20:56:27.077000",
          "content": "<p>Most videos have 10 seconds, with 300 frames.\nBut you have a valid point: if not using all frames but a selection, intuitively it makes sense not to select adjacent frames. The only way to know for sure, though, is testing :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706071,
          "author_name": "M00100100100w",
          "author_url": "",
          "post_date": "2019-12-29T22:24:54.497000",
          "content": "<p>Yeah I didn't look at the actual data set. I just assumed in a 30 second video there would be atleast 1 frame per second </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 697007,
      "author_name": "dasmehdixtr",
      "author_url": "",
      "post_date": "2019-12-17T11:07:35.897000",
      "content": "<p>Nice! 💯 </p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2638119,
      "author_name": "MUKESH",
      "author_url": "",
      "post_date": "2024-02-06T05:07:15.620000",
      "content": "<p>I am not able to run the faceforensics++ dataset script. can anyone help me  access this dataset</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 704598,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2019-12-27T17:17:27.290000",
      "content": "<p><a href=\"/bacterio\">@bacterio</a> \nI struggled to replicate your results... learned a lot in the process.\nHowever, the max speed I got to detect faces in 8 videos at full resolution, in parallel, was:</p>\n\n<p>| GPU        | FPS |\n|------------|-----|\n| Tesla K80  | 22  |\n| Tesla P4   | 28  |\n| Tesla T4   | 31  |\n| Tesla V100 | 28  |</p>\n\n<p>All experiments were done in Google Cloud, 8 vCPUs / 30 GB memory, and the same key 4 constraints you used (OpenCV, min face size = 60px, no batching, using a pool with 8 processes).</p>\n\n<p>I'm using PyTorch, and this <a href=\"https://github.com/timesler/facenet-pytorch\">facenet-pytorch</a> MTCNN implementation. </p>\n\n<p>The only explanation I can think of for the brutal difference in performance vs. what you achieved is the MTCNN implementation... I will try <a href=\"https://github.com/ipazc/mtcnn\">this TensorFlow implementation</a>. Which one did you use?</p>\n\n<p>Cheers!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 704709,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2019-12-27T21:54:02.340000",
          "content": "<p>I used this one: <a href=\"https://github.com/faciallab/FaceDetector\">https://github.com/faciallab/FaceDetector</a></p>\n\n<p>The FPS I mantioned are with minimal pre-processing, ie. just BGR -&gt; RGB and basic histogram equalization.</p>\n\n<p>Also I realized a few videos are &lt;1920x1080 so minsize should depend on the actual video resolution.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 700527,
      "author_name": "Shubham Prateek",
      "author_url": "",
      "post_date": "2019-12-22T06:38:48.410000",
      "content": "<p>This is great.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 700058,
      "author_name": "Michaël Karpe",
      "author_url": "",
      "post_date": "2019-12-21T11:35:20.360000",
      "content": "<p>I found a GitHub repo questioning the ability of FaceForensics++ to detect all kinds of deepfakes, you'll probably be interesting in this repo! See the discussion here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122603\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122603</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 697089,
      "author_name": "David Yee",
      "author_url": "",
      "post_date": "2019-12-17T13:22:53.663000",
      "content": "<p>Wow this is really cool. In fact it is a new revolutionary technology</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 696783,
      "author_name": "Ilya Khristoforov",
      "author_url": "",
      "post_date": "2019-12-17T03:58:02.877000",
      "content": "<p>Cool!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 695789,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2019-12-15T15:16:57.960000",
      "content": "<p>I'm curious to see what this would score on the leaderboard. 😄 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 695796,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-15T15:26:27.837000",
          "content": "<p>Me too, working on that ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 696814,
          "author_name": "Rafid Abyaad",
          "author_url": "",
          "post_date": "2019-12-17T05:18:28.833000",
          "content": "<p>Hi <a href=\"/carlossouza\">@carlossouza</a> ,\nI'm currently studying the paper. Kindly share your findings when you're done with experiments.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 697023,
          "author_name": "Rafid Abyaad",
          "author_url": "",
          "post_date": "2019-12-17T11:32:41.643000",
          "content": "<p>This <a href=\"http://openaccess.thecvf.com/content_CVPRW_2019/papers/Media%20Forensics/Sabir_Recurrent_Convolutional_Strategies_for_Face_Manipulation_Detection_in_Videos_CVPRW_2019_paper.pdf\">paper</a> seems interesting. Might be worth looking into.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 697309,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-17T17:38:09.477000",
          "content": "<p><a href=\"/abyaadrafid\">@abyaadrafid</a> Currently I’m struggling to make sure dlib is fully utilizing GPUs/CUDA. That’s not easy/out-of-the-box behavior... as soon as I figure out how to do it, I’ll share :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 706097,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-29T23:12:16.153000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 705296,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-28T18:24:18.693000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "695377": "Just quick read all shared papers, and tested FaceForensics++ pre-trained model. Results are pretty cool, here's some samples:\n\n**Real sample**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F17e9bed55f9da71196ec53bb39b89dc9%2Fabarnvbtwb.gif?generation=1576380120161560&amp;alt=media)\n\n**Fake sample**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fe1e531517bc51c6889098cf9d6bd2de8%2Faagfhgtpmv.gif?generation=1576380143703413&amp;alt=media)\n\nCheers!",
    "700613": "I have managed to achieve ~65 fps face detection on 8xCPU + 1xGPU (1070Ti) + MTCNN at full resolution (1920x1080). Keys:\n\n- Opencv for video reading/processing\n- Use minsize=60px to keep the pyramidal downscaling under a limit\n- No batching (process one frame at a time)\n- Use as many processes as CPUs\n\nOverall 8 videos are processed every ~35 or 40s. \n\nBottleneck is apparently in the CPU (100%) while the GPU stays at a steady ~85%\n\nThis approach may benefit from using scikit for video processing as suggested in another thread.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3076212%2Fa0451a8611604d17a819a6cb1872ff27%2FCaptura%20de%20pantalla%202019-12-22%20a%20las%2010.59.50.png?generation=1577008892533969&amp;alt=media)\n",
    "695787": "Btw, I scanned through all discussion posts and downloaded all papers. All can be [found here](https://www.dropbox.com/sh/ivio1l39chxalkl/AAD1_CXOMrYb8ZdAho0OqB4na?dl=0).",
    "697292": "I created a notebook that is still a work in progress, but uses the pretrained models from this paper:\nhttps://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet",
    "697007": "Nice! 💯 ",
    "2638119": "I am not able to run the faceforensics++ dataset script. can anyone help me  access this dataset",
    "704598": "@bacterio \nI struggled to replicate your results... learned a lot in the process.\nHowever, the max speed I got to detect faces in 8 videos at full resolution, in parallel, was:\n\n| GPU        | FPS |\n|------------|-----|\n| Tesla K80  | 22  |\n| Tesla P4   | 28  |\n| Tesla T4   | 31  |\n| Tesla V100 | 28  |\n\nAll experiments were done in Google Cloud, 8 vCPUs / 30 GB memory, and the same key 4 constraints you used (OpenCV, min face size = 60px, no batching, using a pool with 8 processes).\n\nI'm using PyTorch, and this [facenet-pytorch](https://github.com/timesler/facenet-pytorch) MTCNN implementation. \n\nThe only explanation I can think of for the brutal difference in performance vs. what you achieved is the MTCNN implementation... I will try [this TensorFlow implementation](https://github.com/ipazc/mtcnn). Which one did you use?\n\nCheers!",
    "700527": "This is great.",
    "700058": "I found a GitHub repo questioning the ability of FaceForensics++ to detect all kinds of deepfakes, you'll probably be interesting in this repo! See the discussion here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122603",
    "697089": "Wow this is really cool. In fact it is a new revolutionary technology",
    "696783": "Cool!",
    "695789": "I'm curious to see what this would score on the leaderboard. 😄 ",
    "706097": "",
    "705296": ""
  }
}