{
  "id": 158167,
  "title": "92th place solution",
  "url": "/competitions/deepfake-detection-challenge/discussion/158167",
  "author_name": "Maxime Churin",
  "post_date": "2020-06-13T10:01:56.309000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>With the publication of the final results, we will share in this post the details of the machine learning pipeline that got us into the 92nd place (top 5% of all contestants).</p>\n\n<p>We want to share insight into what we found worked well and what did not for this challenge.</p>\n\n<h2>ML Pipeline</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2541996%2F559ac0932600a1a49a16c482b09fafc1%2Fpipeline_DFDC.png?generation=1592040705943795&amp;alt=media\" alt=\"\">\n<em>Our ML pipeline, including randomly sampling frames from videos, cropping around the face, inference with a trained model, and averaging the predictions to make a prediction for the video.</em></p>\n\n<p>The first choice we had to make involved what to process: should we focus on the entire videos, or on individual frames sampled from the videos? This choice would affect the type of machine learning model we used, the computation time, and probably the final performance. After observing that many fake videos contained obvious facial manipulations and could be identified from still images, we decided early on to focus on detecting a deepfake from individual frames.</p>\n\n<p>We first created a balanced training set by sampling 10 frames from real videos and 2 frames per fake video. We then cropped these images to focus on the faces, used transfer learning to train a pretrained model on this face dataset, and finally for inference we used ensembling with several models and video frames.</p>\n\n<p>We reached the top 5% of participants thanks to a combination of choices we made throughout the ML pipeline, focusing on:</p>\n\n<ul>\n<li>face detection</li>\n<li>model selection</li>\n<li>generalization methods</li>\n<li>frame and model ensembling</li>\n</ul>\n\n<h2>Face Detection</h2>\n\n<p>Since the video manipulations focused on people’s faces, which were only a small part of the videos, it was essential to develop a good face detection system. Specifically, we sought a face detection system that would maximize the detection probability and minimize the false positive rate.</p>\n\n<p>We benchmarked and compared several different open source face detectors, and ended up testing two different approaches:</p>\n\n<ol>\n<li>Face detection with no false positives. We hypothesized that noise in the data created by frames wrongly identified as faces could adversely affect training. For this approach, we used a combination of two different face detectors: faced [1] and dlib [2]. By requiring that a face must be detected by both detectors, we reached 0.1% false positives, which came at a cost of a rather low detection rate of 60%.</li>\n<li>Balanced false positive rate and detection probability. For this approach, we used the BlazeFace [3] detector and immediately reached 97% detection rate and 7% false positives. However, we noticed that this detector performed poorly when the faces were far from the camera. By adding an additional step in which we created horizontal, vertical, and matrix sliding windows with overlapping cells, and then taking those windows to perform face detection, we reached a 99% detection rate with 8% false positives.</li>\n</ol>\n\n<p>We found that the second approach yielded better results on the Kaggle leaderboard, and then adopted it as our baseline for face detection. We parallelized this approach to make it faster and we were extracting frames from videos, detecting faces, and performing model inference at 25 frames per second in the Kaggle environment.</p>\n\n<h2>Choice of model</h2>\n\n<p>By choosing to focus only on frames as opposed to entire videos, we had reduced the challenge to an image classification problem and could rely on pre-existing methods. Specifically, we used transfer learning with Convolutional Neural Networks (CNNs) trained on the ImageNet classification task [4]. Although the ImageNet dataset and the Deepfake dataset are very different, all natural images share common features, and we expect that starting with a pretrained model allows it to focus on the specifics of the problem at hand.</p>\n\n<p>As we tested several models, we found a correlation between the model’s results on ImageNet and the validation loss on our dataset. As a result, we selected ResNext [5] and DenseNet [6] models.</p>\n\n<h2>Generalization methods</h2>\n\n<p>The biggest problem we had to solve in this DeepFake challenge was that of generalization, as the training data provided and the test data used for the leaderboard came from different distributions. Consequently, we found that there was often a significant gap between our validation score and the leaderboard score.</p>\n\n<p>We combined three different approaches to improve generalization:</p>\n\n<ul>\n<li>Regularization: we tried standard methods such as dropout, weight decay, and label smoothing. However, we found that these methods only yielded a small improvement.</li>\n<li>Adding more data: to increase the diversity of images in our training set, we added data from a different source. Specifically, we used the FaceForensics dataset [7], released in 2018, which is also focused on deepfake detection.</li>\n<li>Data augmentation: We synthetically created more data by randomly introducing slight variations to the existing images using translation, rotation, brightness variation, compression, and resolution reduction.</li>\n</ul>\n\n<p>We found that the data augmentation techniques led to the highest improvement in our score. On our final submissions, we only used data augmentation and regularization.</p>\n\n<h2>Ensembling</h2>\n\n<p>To make predictions on videos, we performed two types of ensembling:</p>\n\n<ul>\n<li>Frame ensembling: we randomly sampled frames from a video, performed face detection, made individual predictions for each face detected, and then used the average over all the predictions as our prediction for the video. We started out doing this averaging with 5 frames per video, but quickly realized the more frames we used, the better. In our final submission, we were taking the average over 100 frames per video.</li>\n<li>Model ensembling: since different models tend to learn slightly different patterns in the data, to achieve better generalization, we averaged the predictions made by several different ResNext and DenseNet models.</li>\n</ul>\n\n<p>The use of both types of ensembling significantly improved our score and made us shoot up about 100 places in the ranking.</p>\n\n<h2>What we did and did not try</h2>\n\n<p>There were a few other things we tried that ended up not making the cut in our final pipeline:</p>\n\n<ul>\n<li>We hypothesized that we may be missing out on important information captured in the temporal dimension of the video. We thus trained a Recurrent Neural Network (RNN) on top of the pre-trained Convolutional Neural Network (CNN) with sequences of frames as input. However, we found that our results were not very different from training just a CNN and averaging several predictions. We thus concluded that most information about the real or fake nature of the data could be found in individual frames.</li>\n<li>We hypothesized that simply averaging predictions made by a CNN over several frames of a video was suboptimal and that better results could be reached by training a separate network to aggregate the predictions. However, despite having a huge impact on our validation score, we found that this procedure led to overfitting to our train dataset and worse generalization to the test dataset.</li>\n</ul>\n\n<p>Some things we did not try but could have improved our score are:</p>\n\n<ul>\n<li>We knew from the start that around 5% of the videos were fake because of manipulations not of the image, but of the audio. Although we decided to focus on the visual fakes, extra gains could have been made, either by sticking with our approach and cleaning our train dataset to remove the fake audios or by building a more comprehensive model that could deal with both fake audios and fake visuals.</li>\n<li>Training a classifier with an auxiliary (but relevant) objective is known to often lead to improvement. For instance, we could have trained our model to not only classify the video frames into fake and real categories but also predict which pixels were most likely to have been altered.\n<ul><li>We expect that more data augmentations and fine tuning of our models and pipeline would have led to improved generalization. For example, we found after the competition had officially ended that a wider crop around the faces we detected would have led to a better score.</li></ul></li>\n</ul>\n\n<h2>Our results</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2541996%2F6c5a267127b0fe3e6ac6f75e452342ec%2Fscore_evolution.png?generation=1592041933312850&amp;alt=media\" alt=\"\">\n<em>Evolution of our score (yellow), ranking in the leaderboard (blue) over time, and main technology driver behind score improvement (green banner)</em></p>\n\n<p>The plot above shows the evolution of our leaderboard score (in yellow, smaller is better) and of our ranking among competitors (blue, higher is better). Our score consistently improved over time, as we incrementally tested different approaches and made changes to our pipeline.</p>\n\n<h2>Acknowledgments</h2>\n\n<p>Thank you to our team members: <a href=\"/berlitos\">@berlitos</a> <a href=\"/wrclements\">@wrclements</a> <a href=\"/bastien7898\">@bastien7898</a> <a href=\"/bennox\">@bennox</a>.\nThank you to the organizers for hosting this competition.</p>\n\n<h2>References</h2>\n\n<p>[1] <a href=\"https://github.com/iitzco/faced\">iitzco/faced: 🚀 😏 Near Real Time CPU Face detection using deep learning</a>\n[2] <a href=\"http://dlib.net/\">dlib C++ Library</a>\n[3] <a href=\"https://github.com/hollance/BlazeFace-PyTorch\">hollance/BlazeFace-PyTorch: The BlazeFace face detector model implemented in PyTorch</a>\n[4] <a href=\"https://paperswithcode.com/sota/image-classification-on-imagenet\">ImageNet Leaderboard | Papers with Code</a>\n[5] ResNext: <a href=\"https://arxiv.org/abs/1611.05431\">https://arxiv.org/abs/1611.05431</a>\n[6] DenseNet: <a href=\"https://arxiv.org/abs/1608.06993\">https://arxiv.org/abs/1608.06993</a>\n[7] <a href=\"https://github.com/ondyari/FaceForensics\">ondyari/FaceForensics: Github of the FaceForensics dataset</a></p>",
  "messages": [
    {
      "id": 884357,
      "postDate": "2020-06-13T10:01:56.310Z",
      "content": "<p>With the publication of the final results, we will share in this post the details of the machine learning pipeline that got us into the 92nd place (top 5% of all contestants).</p>\n\n<p>We want to share insight into what we found worked well and what did not for this challenge.</p>\n\n<h2>ML Pipeline</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2541996%2F559ac0932600a1a49a16c482b09fafc1%2Fpipeline_DFDC.png?generation=1592040705943795&amp;alt=media\" alt=\"\">\n<em>Our ML pipeline, including randomly sampling frames from videos, cropping around the face, inference with a trained model, and averaging the predictions to make a prediction for the video.</em></p>\n\n<p>The first choice we had to make involved what to process: should we focus on the entire videos, or on individual frames sampled from the videos? This choice would affect the type of machine learning model we used, the computation time, and probably the final performance. After observing that many fake videos contained obvious facial manipulations and could be identified from still images, we decided early on to focus on detecting a deepfake from individual frames.</p>\n\n<p>We first created a balanced training set by sampling 10 frames from real videos and 2 frames per fake video. We then cropped these images to focus on the faces, used transfer learning to train a pretrained model on this face dataset, and finally for inference we used ensembling with several models and video frames.</p>\n\n<p>We reached the top 5% of participants thanks to a combination of choices we made throughout the ML pipeline, focusing on:</p>\n\n<ul>\n<li>face detection</li>\n<li>model selection</li>\n<li>generalization methods</li>\n<li>frame and model ensembling</li>\n</ul>\n\n<h2>Face Detection</h2>\n\n<p>Since the video manipulations focused on people’s faces, which were only a small part of the videos, it was essential to develop a good face detection system. Specifically, we sought a face detection system that would maximize the detection probability and minimize the false positive rate.</p>\n\n<p>We benchmarked and compared several different open source face detectors, and ended up testing two different approaches:</p>\n\n<ol>\n<li>Face detection with no false positives. We hypothesized that noise in the data created by frames wrongly identified as faces could adversely affect training. For this approach, we used a combination of two different face detectors: faced [1] and dlib [2]. By requiring that a face must be detected by both detectors, we reached 0.1% false positives, which came at a cost of a rather low detection rate of 60%.</li>\n<li>Balanced false positive rate and detection probability. For this approach, we used the BlazeFace [3] detector and immediately reached 97% detection rate and 7% false positives. However, we noticed that this detector performed poorly when the faces were far from the camera. By adding an additional step in which we created horizontal, vertical, and matrix sliding windows with overlapping cells, and then taking those windows to perform face detection, we reached a 99% detection rate with 8% false positives.</li>\n</ol>\n\n<p>We found that the second approach yielded better results on the Kaggle leaderboard, and then adopted it as our baseline for face detection. We parallelized this approach to make it faster and we were extracting frames from videos, detecting faces, and performing model inference at 25 frames per second in the Kaggle environment.</p>\n\n<h2>Choice of model</h2>\n\n<p>By choosing to focus only on frames as opposed to entire videos, we had reduced the challenge to an image classification problem and could rely on pre-existing methods. Specifically, we used transfer learning with Convolutional Neural Networks (CNNs) trained on the ImageNet classification task [4]. Although the ImageNet dataset and the Deepfake dataset are very different, all natural images share common features, and we expect that starting with a pretrained model allows it to focus on the specifics of the problem at hand.</p>\n\n<p>As we tested several models, we found a correlation between the model’s results on ImageNet and the validation loss on our dataset. As a result, we selected ResNext [5] and DenseNet [6] models.</p>\n\n<h2>Generalization methods</h2>\n\n<p>The biggest problem we had to solve in this DeepFake challenge was that of generalization, as the training data provided and the test data used for the leaderboard came from different distributions. Consequently, we found that there was often a significant gap between our validation score and the leaderboard score.</p>\n\n<p>We combined three different approaches to improve generalization:</p>\n\n<ul>\n<li>Regularization: we tried standard methods such as dropout, weight decay, and label smoothing. However, we found that these methods only yielded a small improvement.</li>\n<li>Adding more data: to increase the diversity of images in our training set, we added data from a different source. Specifically, we used the FaceForensics dataset [7], released in 2018, which is also focused on deepfake detection.</li>\n<li>Data augmentation: We synthetically created more data by randomly introducing slight variations to the existing images using translation, rotation, brightness variation, compression, and resolution reduction.</li>\n</ul>\n\n<p>We found that the data augmentation techniques led to the highest improvement in our score. On our final submissions, we only used data augmentation and regularization.</p>\n\n<h2>Ensembling</h2>\n\n<p>To make predictions on videos, we performed two types of ensembling:</p>\n\n<ul>\n<li>Frame ensembling: we randomly sampled frames from a video, performed face detection, made individual predictions for each face detected, and then used the average over all the predictions as our prediction for the video. We started out doing this averaging with 5 frames per video, but quickly realized the more frames we used, the better. In our final submission, we were taking the average over 100 frames per video.</li>\n<li>Model ensembling: since different models tend to learn slightly different patterns in the data, to achieve better generalization, we averaged the predictions made by several different ResNext and DenseNet models.</li>\n</ul>\n\n<p>The use of both types of ensembling significantly improved our score and made us shoot up about 100 places in the ranking.</p>\n\n<h2>What we did and did not try</h2>\n\n<p>There were a few other things we tried that ended up not making the cut in our final pipeline:</p>\n\n<ul>\n<li>We hypothesized that we may be missing out on important information captured in the temporal dimension of the video. We thus trained a Recurrent Neural Network (RNN) on top of the pre-trained Convolutional Neural Network (CNN) with sequences of frames as input. However, we found that our results were not very different from training just a CNN and averaging several predictions. We thus concluded that most information about the real or fake nature of the data could be found in individual frames.</li>\n<li>We hypothesized that simply averaging predictions made by a CNN over several frames of a video was suboptimal and that better results could be reached by training a separate network to aggregate the predictions. However, despite having a huge impact on our validation score, we found that this procedure led to overfitting to our train dataset and worse generalization to the test dataset.</li>\n</ul>\n\n<p>Some things we did not try but could have improved our score are:</p>\n\n<ul>\n<li>We knew from the start that around 5% of the videos were fake because of manipulations not of the image, but of the audio. Although we decided to focus on the visual fakes, extra gains could have been made, either by sticking with our approach and cleaning our train dataset to remove the fake audios or by building a more comprehensive model that could deal with both fake audios and fake visuals.</li>\n<li>Training a classifier with an auxiliary (but relevant) objective is known to often lead to improvement. For instance, we could have trained our model to not only classify the video frames into fake and real categories but also predict which pixels were most likely to have been altered.\n<ul><li>We expect that more data augmentations and fine tuning of our models and pipeline would have led to improved generalization. For example, we found after the competition had officially ended that a wider crop around the faces we detected would have led to a better score.</li></ul></li>\n</ul>\n\n<h2>Our results</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2541996%2F6c5a267127b0fe3e6ac6f75e452342ec%2Fscore_evolution.png?generation=1592041933312850&amp;alt=media\" alt=\"\">\n<em>Evolution of our score (yellow), ranking in the leaderboard (blue) over time, and main technology driver behind score improvement (green banner)</em></p>\n\n<p>The plot above shows the evolution of our leaderboard score (in yellow, smaller is better) and of our ranking among competitors (blue, higher is better). Our score consistently improved over time, as we incrementally tested different approaches and made changes to our pipeline.</p>\n\n<h2>Acknowledgments</h2>\n\n<p>Thank you to our team members: <a href=\"/berlitos\">@berlitos</a> <a href=\"/wrclements\">@wrclements</a> <a href=\"/bastien7898\">@bastien7898</a> <a href=\"/bennox\">@bennox</a>.\nThank you to the organizers for hosting this competition.</p>\n\n<h2>References</h2>\n\n<p>[1] <a href=\"https://github.com/iitzco/faced\">iitzco/faced: 🚀 😏 Near Real Time CPU Face detection using deep learning</a>\n[2] <a href=\"http://dlib.net/\">dlib C++ Library</a>\n[3] <a href=\"https://github.com/hollance/BlazeFace-PyTorch\">hollance/BlazeFace-PyTorch: The BlazeFace face detector model implemented in PyTorch</a>\n[4] <a href=\"https://paperswithcode.com/sota/image-classification-on-imagenet\">ImageNet Leaderboard | Papers with Code</a>\n[5] ResNext: <a href=\"https://arxiv.org/abs/1611.05431\">https://arxiv.org/abs/1611.05431</a>\n[6] DenseNet: <a href=\"https://arxiv.org/abs/1608.06993\">https://arxiv.org/abs/1608.06993</a>\n[7] <a href=\"https://github.com/ondyari/FaceForensics\">ondyari/FaceForensics: Github of the FaceForensics dataset</a></p>",
      "rawMarkdown": "With the publication of the final results, we will share in this post the details of the machine learning pipeline that got us into the 92nd place (top 5% of all contestants).\n\nWe want to share insight into what we found worked well and what did not for this challenge.\n\n## ML Pipeline\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2541996%2F559ac0932600a1a49a16c482b09fafc1%2Fpipeline_DFDC.png?generation=1592040705943795&amp;alt=media)\n*Our ML pipeline, including randomly sampling frames from videos, cropping around the face, inference with a trained model, and averaging the predictions to make a prediction for the video.*\n\nThe first choice we had to make involved what to process: should we focus on the entire videos, or on individual frames sampled from the videos? This choice would affect the type of machine learning model we used, the computation time, and probably the final performance. After observing that many fake videos contained obvious facial manipulations and could be identified from still images, we decided early on to focus on detecting a deepfake from individual frames.\n\nWe first created a balanced training set by sampling 10 frames from real videos and 2 frames per fake video. We then cropped these images to focus on the faces, used transfer learning to train a pretrained model on this face dataset, and finally for inference we used ensembling with several models and video frames.\n\nWe reached the top 5% of participants thanks to a combination of choices we made throughout the ML pipeline, focusing on:\n\n- face detection\n- model selection\n- generalization methods\n- frame and model ensembling\n\n## Face Detection\n\nSince the video manipulations focused on people’s faces, which were only a small part of the videos, it was essential to develop a good face detection system. Specifically, we sought a face detection system that would maximize the detection probability and minimize the false positive rate.\n\nWe benchmarked and compared several different open source face detectors, and ended up testing two different approaches:\n\n1. Face detection with no false positives. We hypothesized that noise in the data created by frames wrongly identified as faces could adversely affect training. For this approach, we used a combination of two different face detectors: faced [1] and dlib [2]. By requiring that a face must be detected by both detectors, we reached 0.1% false positives, which came at a cost of a rather low detection rate of 60%.\n2. Balanced false positive rate and detection probability. For this approach, we used the BlazeFace [3] detector and immediately reached 97% detection rate and 7% false positives. However, we noticed that this detector performed poorly when the faces were far from the camera. By adding an additional step in which we created horizontal, vertical, and matrix sliding windows with overlapping cells, and then taking those windows to perform face detection, we reached a 99% detection rate with 8% false positives.\n\nWe found that the second approach yielded better results on the Kaggle leaderboard, and then adopted it as our baseline for face detection. We parallelized this approach to make it faster and we were extracting frames from videos, detecting faces, and performing model inference at 25 frames per second in the Kaggle environment.\n\n## Choice of model\n\nBy choosing to focus only on frames as opposed to entire videos, we had reduced the challenge to an image classification problem and could rely on pre-existing methods. Specifically, we used transfer learning with Convolutional Neural Networks (CNNs) trained on the ImageNet classification task [4]. Although the ImageNet dataset and the Deepfake dataset are very different, all natural images share common features, and we expect that starting with a pretrained model allows it to focus on the specifics of the problem at hand.\n\nAs we tested several models, we found a correlation between the model’s results on ImageNet and the validation loss on our dataset. As a result, we selected ResNext [5] and DenseNet [6] models.\n\n\n## Generalization methods\n\nThe biggest problem we had to solve in this DeepFake challenge was that of generalization, as the training data provided and the test data used for the leaderboard came from different distributions. Consequently, we found that there was often a significant gap between our validation score and the leaderboard score.\n\nWe combined three different approaches to improve generalization:\n\n- Regularization: we tried standard methods such as dropout, weight decay, and label smoothing. However, we found that these methods only yielded a small improvement.\n- Adding more data: to increase the diversity of images in our training set, we added data from a different source. Specifically, we used the FaceForensics dataset [7], released in 2018, which is also focused on deepfake detection.\n- Data augmentation: We synthetically created more data by randomly introducing slight variations to the existing images using translation, rotation, brightness variation, compression, and resolution reduction.\n\nWe found that the data augmentation techniques led to the highest improvement in our score. On our final submissions, we only used data augmentation and regularization.\n\n## Ensembling\n\nTo make predictions on videos, we performed two types of ensembling:\n\n- Frame ensembling: we randomly sampled frames from a video, performed face detection, made individual predictions for each face detected, and then used the average over all the predictions as our prediction for the video. We started out doing this averaging with 5 frames per video, but quickly realized the more frames we used, the better. In our final submission, we were taking the average over 100 frames per video.\n- Model ensembling: since different models tend to learn slightly different patterns in the data, to achieve better generalization, we averaged the predictions made by several different ResNext and DenseNet models.\n\nThe use of both types of ensembling significantly improved our score and made us shoot up about 100 places in the ranking.\n\n## What we did and did not try\n\nThere were a few other things we tried that ended up not making the cut in our final pipeline:\n\n- We hypothesized that we may be missing out on important information captured in the temporal dimension of the video. We thus trained a Recurrent Neural Network (RNN) on top of the pre-trained Convolutional Neural Network (CNN) with sequences of frames as input. However, we found that our results were not very different from training just a CNN and averaging several predictions. We thus concluded that most information about the real or fake nature of the data could be found in individual frames.\n- We hypothesized that simply averaging predictions made by a CNN over several frames of a video was suboptimal and that better results could be reached by training a separate network to aggregate the predictions. However, despite having a huge impact on our validation score, we found that this procedure led to overfitting to our train dataset and worse generalization to the test dataset.\n\nSome things we did not try but could have improved our score are:\n\n- We knew from the start that around 5% of the videos were fake because of manipulations not of the image, but of the audio. Although we decided to focus on the visual fakes, extra gains could have been made, either by sticking with our approach and cleaning our train dataset to remove the fake audios or by building a more comprehensive model that could deal with both fake audios and fake visuals.\n- Training a classifier with an auxiliary (but relevant) objective is known to often lead to improvement. For instance, we could have trained our model to not only classify the video frames into fake and real categories but also predict which pixels were most likely to have been altered.\n    - We expect that more data augmentations and fine tuning of our models and pipeline would have led to improved generalization. For example, we found after the competition had officially ended that a wider crop around the faces we detected would have led to a better score.\n\n## Our results\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2541996%2F6c5a267127b0fe3e6ac6f75e452342ec%2Fscore_evolution.png?generation=1592041933312850&amp;alt=media)\n*Evolution of our score (yellow), ranking in the leaderboard (blue) over time, and main technology driver behind score improvement (green banner)*\n\n\nThe plot above shows the evolution of our leaderboard score (in yellow, smaller is better) and of our ranking among competitors (blue, higher is better). Our score consistently improved over time, as we incrementally tested different approaches and made changes to our pipeline.\n\n## Acknowledgments\n\nThank you to our team members: @berlitos @wrclements @bastien7898 @bennox.\nThank you to the organizers for hosting this competition.\n\n## References\n\n[1] [iitzco/faced: 🚀 😏 Near Real Time CPU Face detection using deep learning](https://github.com/iitzco/faced)\n[2] [dlib C++ Library](http://dlib.net/)\n[3] [hollance/BlazeFace-PyTorch: The BlazeFace face detector model implemented in PyTorch](https://github.com/hollance/BlazeFace-PyTorch)\n[4] [ImageNet Leaderboard | Papers with Code](https://paperswithcode.com/sota/image-classification-on-imagenet)\n[5] ResNext: [https://arxiv.org/abs/1611.05431](https://arxiv.org/abs/1611.05431)\n[6] DenseNet: [https://arxiv.org/abs/1608.06993](https://arxiv.org/abs/1608.06993)\n[7] [ondyari/FaceForensics: Github of the FaceForensics dataset](https://github.com/ondyari/FaceForensics)",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "884357": "With the publication of the final results, we will share in this post the details of the machine learning pipeline that got us into the 92nd place (top 5% of all contestants).\n\nWe want to share insight into what we found worked well and what did not for this challenge.\n\n## ML Pipeline\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2541996%2F559ac0932600a1a49a16c482b09fafc1%2Fpipeline_DFDC.png?generation=1592040705943795&amp;alt=media)\n*Our ML pipeline, including randomly sampling frames from videos, cropping around the face, inference with a trained model, and averaging the predictions to make a prediction for the video.*\n\nThe first choice we had to make involved what to process: should we focus on the entire videos, or on individual frames sampled from the videos? This choice would affect the type of machine learning model we used, the computation time, and probably the final performance. After observing that many fake videos contained obvious facial manipulations and could be identified from still images, we decided early on to focus on detecting a deepfake from individual frames.\n\nWe first created a balanced training set by sampling 10 frames from real videos and 2 frames per fake video. We then cropped these images to focus on the faces, used transfer learning to train a pretrained model on this face dataset, and finally for inference we used ensembling with several models and video frames.\n\nWe reached the top 5% of participants thanks to a combination of choices we made throughout the ML pipeline, focusing on:\n\n- face detection\n- model selection\n- generalization methods\n- frame and model ensembling\n\n## Face Detection\n\nSince the video manipulations focused on people’s faces, which were only a small part of the videos, it was essential to develop a good face detection system. Specifically, we sought a face detection system that would maximize the detection probability and minimize the false positive rate.\n\nWe benchmarked and compared several different open source face detectors, and ended up testing two different approaches:\n\n1. Face detection with no false positives. We hypothesized that noise in the data created by frames wrongly identified as faces could adversely affect training. For this approach, we used a combination of two different face detectors: faced [1] and dlib [2]. By requiring that a face must be detected by both detectors, we reached 0.1% false positives, which came at a cost of a rather low detection rate of 60%.\n2. Balanced false positive rate and detection probability. For this approach, we used the BlazeFace [3] detector and immediately reached 97% detection rate and 7% false positives. However, we noticed that this detector performed poorly when the faces were far from the camera. By adding an additional step in which we created horizontal, vertical, and matrix sliding windows with overlapping cells, and then taking those windows to perform face detection, we reached a 99% detection rate with 8% false positives.\n\nWe found that the second approach yielded better results on the Kaggle leaderboard, and then adopted it as our baseline for face detection. We parallelized this approach to make it faster and we were extracting frames from videos, detecting faces, and performing model inference at 25 frames per second in the Kaggle environment.\n\n## Choice of model\n\nBy choosing to focus only on frames as opposed to entire videos, we had reduced the challenge to an image classification problem and could rely on pre-existing methods. Specifically, we used transfer learning with Convolutional Neural Networks (CNNs) trained on the ImageNet classification task [4]. Although the ImageNet dataset and the Deepfake dataset are very different, all natural images share common features, and we expect that starting with a pretrained model allows it to focus on the specifics of the problem at hand.\n\nAs we tested several models, we found a correlation between the model’s results on ImageNet and the validation loss on our dataset. As a result, we selected ResNext [5] and DenseNet [6] models.\n\n\n## Generalization methods\n\nThe biggest problem we had to solve in this DeepFake challenge was that of generalization, as the training data provided and the test data used for the leaderboard came from different distributions. Consequently, we found that there was often a significant gap between our validation score and the leaderboard score.\n\nWe combined three different approaches to improve generalization:\n\n- Regularization: we tried standard methods such as dropout, weight decay, and label smoothing. However, we found that these methods only yielded a small improvement.\n- Adding more data: to increase the diversity of images in our training set, we added data from a different source. Specifically, we used the FaceForensics dataset [7], released in 2018, which is also focused on deepfake detection.\n- Data augmentation: We synthetically created more data by randomly introducing slight variations to the existing images using translation, rotation, brightness variation, compression, and resolution reduction.\n\nWe found that the data augmentation techniques led to the highest improvement in our score. On our final submissions, we only used data augmentation and regularization.\n\n## Ensembling\n\nTo make predictions on videos, we performed two types of ensembling:\n\n- Frame ensembling: we randomly sampled frames from a video, performed face detection, made individual predictions for each face detected, and then used the average over all the predictions as our prediction for the video. We started out doing this averaging with 5 frames per video, but quickly realized the more frames we used, the better. In our final submission, we were taking the average over 100 frames per video.\n- Model ensembling: since different models tend to learn slightly different patterns in the data, to achieve better generalization, we averaged the predictions made by several different ResNext and DenseNet models.\n\nThe use of both types of ensembling significantly improved our score and made us shoot up about 100 places in the ranking.\n\n## What we did and did not try\n\nThere were a few other things we tried that ended up not making the cut in our final pipeline:\n\n- We hypothesized that we may be missing out on important information captured in the temporal dimension of the video. We thus trained a Recurrent Neural Network (RNN) on top of the pre-trained Convolutional Neural Network (CNN) with sequences of frames as input. However, we found that our results were not very different from training just a CNN and averaging several predictions. We thus concluded that most information about the real or fake nature of the data could be found in individual frames.\n- We hypothesized that simply averaging predictions made by a CNN over several frames of a video was suboptimal and that better results could be reached by training a separate network to aggregate the predictions. However, despite having a huge impact on our validation score, we found that this procedure led to overfitting to our train dataset and worse generalization to the test dataset.\n\nSome things we did not try but could have improved our score are:\n\n- We knew from the start that around 5% of the videos were fake because of manipulations not of the image, but of the audio. Although we decided to focus on the visual fakes, extra gains could have been made, either by sticking with our approach and cleaning our train dataset to remove the fake audios or by building a more comprehensive model that could deal with both fake audios and fake visuals.\n- Training a classifier with an auxiliary (but relevant) objective is known to often lead to improvement. For instance, we could have trained our model to not only classify the video frames into fake and real categories but also predict which pixels were most likely to have been altered.\n    - We expect that more data augmentations and fine tuning of our models and pipeline would have led to improved generalization. For example, we found after the competition had officially ended that a wider crop around the faces we detected would have led to a better score.\n\n## Our results\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2541996%2F6c5a267127b0fe3e6ac6f75e452342ec%2Fscore_evolution.png?generation=1592041933312850&amp;alt=media)\n*Evolution of our score (yellow), ranking in the leaderboard (blue) over time, and main technology driver behind score improvement (green banner)*\n\n\nThe plot above shows the evolution of our leaderboard score (in yellow, smaller is better) and of our ranking among competitors (blue, higher is better). Our score consistently improved over time, as we incrementally tested different approaches and made changes to our pipeline.\n\n## Acknowledgments\n\nThank you to our team members: @berlitos @wrclements @bastien7898 @bennox.\nThank you to the organizers for hosting this competition.\n\n## References\n\n[1] [iitzco/faced: 🚀 😏 Near Real Time CPU Face detection using deep learning](https://github.com/iitzco/faced)\n[2] [dlib C++ Library](http://dlib.net/)\n[3] [hollance/BlazeFace-PyTorch: The BlazeFace face detector model implemented in PyTorch](https://github.com/hollance/BlazeFace-PyTorch)\n[4] [ImageNet Leaderboard | Papers with Code](https://paperswithcode.com/sota/image-classification-on-imagenet)\n[5] ResNext: [https://arxiv.org/abs/1611.05431](https://arxiv.org/abs/1611.05431)\n[6] DenseNet: [https://arxiv.org/abs/1608.06993](https://arxiv.org/abs/1608.06993)\n[7] [ondyari/FaceForensics: Github of the FaceForensics dataset](https://github.com/ondyari/FaceForensics)"
  }
}