{
  "id": 140625,
  "title": "Journey to 0.3. Public 64th Place Solution",
  "url": "/competitions/deepfake-detection-challenge/discussion/140625",
  "author_name": "Vladislav Ostankovich",
  "post_date": "2020-04-02T17:02:21.076000",
  "votes": 20,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi to all the competitors and organizers of DFDC, this was my first Kaggle competition and here I want to share the methods I've experimented with during the competition and share the details of how I've finally crossed the 0.3 border. </p>\n\n<p>Firstly I would like to write a list of methods that didn't work for me (or they worked but not as good as my final solution).</p>\n\n<h1>Classical methods</h1>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1811.00661\">Exposing Deep Fakes Using Inconsistent Head Poses</a></li>\n</ul>\n\n<p>Code: <a href=\"https://bitbucket.org/ericyang3721/headpose_forensic/src/master/\">https://bitbucket.org/ericyang3721/headpose_forensic/src/master/</a>\nIn few words this method uses difference between face and head poses. Firstly I've implemented it myself, then found the code from authors on bitbucket. Approach didn't work for the majority of cases. I guess it only works good for FaceSwap faces.</p>\n\n<p>Didn't test it against Public LB</p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1911.00686\">Unmasking DeepFakes with simple Features</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/cc-hpc-itwm/DeepFakeDetection\">https://github.com/cc-hpc-itwm/DeepFakeDetection</a>\nBased on analysis of high frequency components in fourier spectra. I've tested the method on a part of dfdc dataset, found out that it works, yet not as good to beat single CNN model.</p>\n\n<p>Didn't test it against Public LB</p>\n\n<h1>Deep Learning-based methods</h1>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1811.00656\">Exposing DeepFake Videos By Detecting Face Warping Artifacts</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/danmohaha/CVPRW2019_Face_Artifacts\">https://github.com/danmohaha/CVPRW2019_Face_Artifacts</a>\nThis was my first submission to the competition. I've used pretrained model provided by authors.</p>\n\n<p><strong>Public Score: 1.16260</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/2001.07444\">Detecting Face2Face Facial Reenactment in Videos</a></li>\n</ul>\n\n<p>I've liked this approach and it even improved my score at some point. I've used ResNet-18 models as stated in the paper and compared it to single ResNet-18 model that was trained on the same data. The improvement in public score was about 0.045. This is either because the approach really works and paying attention to different face regions really matters or because this is just a kind of ensembling (5 similar models, but trained on different image regions). Or maybe both, I don't know 😃. But this approach didn't work with EfficientNet models, which I've used later. </p>\n\n<p><strong>Public Score: 0.55683 (for single ResNet-18)</strong>\n*<em>Public Score: 0.51062 (for ResNets-18)</em>*\n<strong>Public Score: 0.38760 (for EFNets-4)</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1809.00888\">MesoNet: a Compact Facial Video Forgery Detection Network</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/DariusAf/MesoNet\">https://github.com/DariusAf/MesoNet</a>\nI've tested only pretrained models. Didn't train it on dfdc dataset.</p>\n\n<p><strong>Public Score: 1.54918</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1912.11035\">CNN-generated images are surprisingly easy to spot... for now</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/peterwang512/CNNDetection\">https://github.com/peterwang512/CNNDetection</a>\nI've already gave a thought about this approach here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134295#766094\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134295#766094</a></p>\n\n<p>Didn't test it against Public LB</p>\n\n<ul>\n<li><a href=\"https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf\">Deepfake Video Detection Using Recurrent Neural Networks</a></li>\n</ul>\n\n<p>The method which took the majority of my time, for some reason I've thought it has to outperform single CNN model. I've spend several weeks experimenting with Conv-LSTM models. </p>\n\n<p>I've firstly tried to use CNN model as feature extractor (removing last prediction layer) and train LSTM model on those features.</p>\n\n<p>I've tried CNN-LSTM model (similar to <a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference\">https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference</a>), experimented with dropouts, num of parameters, training generators, undersampling, class imbalance, custom losses, augmentations, etc.</p>\n\n<p>I've experimented with number of frames (16 and 32) and with changing the gap between them (sequential frames from different parts of video, every third frame and 2-3 frames per second).</p>\n\n<p>I've sad that it didn't perform even close to single CNN model, I don't know the reason why. Maybe I needed to experiment more or did some mistakes in architecture or training progress.</p>\n\n<p><strong>Public Score: 0.44083 (The best one across many experimental tries. Based on EFNet-4)</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1910.12467\">Use of a Capsule Network to Detect Fake Images and Videos</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/nii-yamagishilab/Capsule-Forensics-v2\">https://github.com/nii-yamagishilab/Capsule-Forensics-v2</a></p>\n\n<p>Trained it for few epochs on subset of DFDC dataset (1:1 real/fake). I suggest it could perform much better with more experiments, but I didn't have time for it. This is the method which I could regret I didn't pay much attention to.</p>\n\n<p><strong>Public Score: 0.60641</strong></p>\n\n<ul>\n<li><a href=\"https://hal.inria.fr/hal-02140558/document\">MARS: Motion-Augmented RGB Stream for Action Recognition</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/craston/MARS\">https://github.com/craston/MARS</a></p>\n\n<p>This is the only method I didn't have time to fully experiment with. The problem is in preparation of optical flows for training. Even with CUDA compiled OpenCV, it takes around 100 hours to extract x and y direction optical flows from 32 frames (the method requires sequences) on my PC. I so much wanted to train/test this approach, but due to the time constraints I've decided to stop it. And this is the second and the last method I could regret I didn't test.</p>\n\n<h1>My Solution</h1>\n\n<p><strong>Public score: 0.29477</strong></p>\n\n<p>The notebook is available here: <a href=\"https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15\">https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15</a></p>\n\n<p>Now I will describe step by step how I've been improving single CNN model and ended up with ensemble. I'm kind of disappointed that after experimenting with other approaches not any of those came closer to single CNN model. But okay, let's continue.</p>\n\n<h3>Data preparation</h3>\n\n<p>Video reader: <a href=\"https://www.kaggle.com/humananalog/deepfakes-inference-demo\">https://www.kaggle.com/humananalog/deepfakes-inference-demo</a> (Big thanks to <a href=\"/humananalog\">@humananalog</a>)\nFace extractor: <a href=\"https://github.com/1adrianb/face-alignment\">https://github.com/1adrianb/face-alignment</a></p>\n\n<p>I've used this specific face extractor, because I wanted to train my model on face only, not the whole head. Since the lib gives landmarks, I could crop face region easily. After some time I've realised that landmarks extraction takes much time (even though it fits to extract 16-20 frames and process them within 8 secs) and decided to keep only detection part, removing all the stuff related to landmarks. The detector used in the lib is S3FD. I've additionally used two more face detection models (MTCNN and BlazeFace) for better confidence. Using additional detectors helped me to remove some FPs and save some FNs. This is especially true for very dark images. Finally I've came up with extracting 32 frames for each person from each video. The reason for that is because I had many experiments with Conv-LSTM models. After all I had around 4M .png images of 224x224 size, but I didn't use so big number of faces for training my final single CNN model. I've randomly picked half of those (so I had 16 faces per person).</p>\n\n<h3>Training</h3>\n\n<p>Training EFNet-4 with setting Keras <code>class_weights</code> during training worked well for me. But I've found out better and what's more important much faster way to train my model. The idea is to balance fake and real data on each epoch. Since the training data has around 5:1 imbalance, each epoch we take all real examples but only 0.2 part of fake examples. In such approach 5 epochs is enough for model to see all examples. But i've experimentally found out that 6 epochs gave the best results. During the 6th epoch I've randomly picked fake examples. </p>\n\n<p>I was afraid that my model would overfit to real faces, but seems like it didn't. Augmentations helped to solve this issue. Couple of words about augmentations. I've used different albumentations functions, but during training I've checked if the image is too dark and applied more brightness with larger aug probability.</p>\n\n<p>If someone needs the generator, here is the code: (I'm too lazy now to clean up my codes and release on github). Note that here <code>fake_images</code> list is 5 times larger than <code>real_images</code>. Amount of images for each epoch was around 750K (350K real and 350K fake).</p>\n\n<pre><code>def generator(real_images, fake_images, preprocess_input_fn, batch_size=1, input_shape=(224,224,3), do_aug=False):\n\n    i = 0\n    k = 0\n\n    images = real_images + fake_images[k*len(real_images):(k+1)*len(real_images)]\n    random.shuffle(images)\n\n    while True:\n\n        x_batch = np.zeros((batch_size, input_shape[0], input_shape[1], 3))\n        y_batch = np.zeros((batch_size), dtype=np.single)\n\n        for b in range(batch_size):\n\n            if i == len(images):\n                i = 0\n                k += 1\n                if k &amp;gt; (len(fake_images)//len(real_images) - 1):\n                    images = real_images + [str(i) for i in np.random.choice(fake_images, size = len(real_images), replace = False)]\n                else:\n                    images = real_images + fake_images[k*len(real_images):(k+1)*len(real_images)]\n                random.shuffle(images)\n\n            x = Image.open(images[i])\n            x = np.array(x) \n\n            if do_aug:\n                x = augment(x)\n\n            y = 1. if '/fake/' in images[i] else 0.\n\n            x_batch[b] = x\n            y_batch[b] = y             \n\n            i += 1\n\n        x_batch = preprocess_input_fn(x_batch)\n\n        yield (x_batch, y_batch)\n</code></pre>\n\n<p>After I've trained several EFNet-4 models, I've found out that all models perform best after 6 epochs. Then I've started to train models for ensemble without validation. I've used RAdam optimizer with initial learning rate of 0.001. This learning rate was changed for 4th and 5th epochs to 0.0005 and for the last 6th epoch I've used 0.0001. Loss is binary crossentropy.</p>\n\n<h3>Inference</h3>\n\n<p>In the last two or three days I decided to switch from my EFNet-4 ensemble of 13 models to EFNet-6 ensemble of 5 models and trained them. Boost was around 0.005. Inference could be completely viewed in my released notebook here: <a href=\"https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15\">https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15</a>.</p>\n\n<h3>Tricks that worked</h3>\n\n<ul>\n<li>Using <code>read_frames_at_indices()</code> along with <code>capture.get(cv2.CAP_PROP_FRAME_COUNT)</code> instead of <code>read_frames()</code>. I don't know the reason why, but it works a little bit faster.</li>\n<li>Resizing video for the sake of saving time for detector.</li>\n<li>Removing small detected objects. These are mainly false positives and don't contain faces. I've used the threshold of <code>min(video.shape)//20</code>.</li>\n<li>Zooming out of face. Didn't find a reason why, but it works better.</li>\n<li>Median averaging across frames. I was surpised, but it gave a boost of more than 0.01 compared to mean averaging.</li>\n</ul>\n\n<h3>Tricks that didn't work</h3>\n\n<ul>\n<li><strong>Clipping</strong>. It worked for single models, but for ensembles of 5+ models it always made worse. Even though for safety I've submitted two versions, one with 0.01-0.99 clip and the second one without it (well there was a clip but very small one of 1e-15).</li>\n<li><strong>TTA</strong>. I've tried with complex TTAs as well as with simple Horizontal Flip. But it turned out that it is better to process more frames without TTA, then apply TTA to less frames.</li>\n<li><strong>Judge by one person</strong>. I've tested the approach of giving FAKE prediction if one of the persons in video is FAKE. I've sorted the multiple faces by x-coordinates, and then gave prediction to each person, finally giving FAKE prediction if one is FAKE. It performs slightly worse than judge all persons equally.</li>\n</ul>\n\n<p><strong>Public score: 0.34 (for single EFNet-4 model)</strong>\n*<em>Public score: 0.30 (for ensemble of 13 EFNet-4 models)</em>*\n<strong>Public score: 0.295 (for ensemble of 5 EFNet-6 models)</strong></p>",
  "messages": [
    {
      "id": 795384,
      "postDate": "2020-04-02T17:02:21.077Z",
      "content": "<p>Hi to all the competitors and organizers of DFDC, this was my first Kaggle competition and here I want to share the methods I've experimented with during the competition and share the details of how I've finally crossed the 0.3 border. </p>\n\n<p>Firstly I would like to write a list of methods that didn't work for me (or they worked but not as good as my final solution).</p>\n\n<h1>Classical methods</h1>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1811.00661\">Exposing Deep Fakes Using Inconsistent Head Poses</a></li>\n</ul>\n\n<p>Code: <a href=\"https://bitbucket.org/ericyang3721/headpose_forensic/src/master/\">https://bitbucket.org/ericyang3721/headpose_forensic/src/master/</a>\nIn few words this method uses difference between face and head poses. Firstly I've implemented it myself, then found the code from authors on bitbucket. Approach didn't work for the majority of cases. I guess it only works good for FaceSwap faces.</p>\n\n<p>Didn't test it against Public LB</p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1911.00686\">Unmasking DeepFakes with simple Features</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/cc-hpc-itwm/DeepFakeDetection\">https://github.com/cc-hpc-itwm/DeepFakeDetection</a>\nBased on analysis of high frequency components in fourier spectra. I've tested the method on a part of dfdc dataset, found out that it works, yet not as good to beat single CNN model.</p>\n\n<p>Didn't test it against Public LB</p>\n\n<h1>Deep Learning-based methods</h1>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1811.00656\">Exposing DeepFake Videos By Detecting Face Warping Artifacts</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/danmohaha/CVPRW2019_Face_Artifacts\">https://github.com/danmohaha/CVPRW2019_Face_Artifacts</a>\nThis was my first submission to the competition. I've used pretrained model provided by authors.</p>\n\n<p><strong>Public Score: 1.16260</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/2001.07444\">Detecting Face2Face Facial Reenactment in Videos</a></li>\n</ul>\n\n<p>I've liked this approach and it even improved my score at some point. I've used ResNet-18 models as stated in the paper and compared it to single ResNet-18 model that was trained on the same data. The improvement in public score was about 0.045. This is either because the approach really works and paying attention to different face regions really matters or because this is just a kind of ensembling (5 similar models, but trained on different image regions). Or maybe both, I don't know 😃. But this approach didn't work with EfficientNet models, which I've used later. </p>\n\n<p><strong>Public Score: 0.55683 (for single ResNet-18)</strong>\n*<em>Public Score: 0.51062 (for ResNets-18)</em>*\n<strong>Public Score: 0.38760 (for EFNets-4)</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1809.00888\">MesoNet: a Compact Facial Video Forgery Detection Network</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/DariusAf/MesoNet\">https://github.com/DariusAf/MesoNet</a>\nI've tested only pretrained models. Didn't train it on dfdc dataset.</p>\n\n<p><strong>Public Score: 1.54918</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1912.11035\">CNN-generated images are surprisingly easy to spot... for now</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/peterwang512/CNNDetection\">https://github.com/peterwang512/CNNDetection</a>\nI've already gave a thought about this approach here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134295#766094\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134295#766094</a></p>\n\n<p>Didn't test it against Public LB</p>\n\n<ul>\n<li><a href=\"https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf\">Deepfake Video Detection Using Recurrent Neural Networks</a></li>\n</ul>\n\n<p>The method which took the majority of my time, for some reason I've thought it has to outperform single CNN model. I've spend several weeks experimenting with Conv-LSTM models. </p>\n\n<p>I've firstly tried to use CNN model as feature extractor (removing last prediction layer) and train LSTM model on those features.</p>\n\n<p>I've tried CNN-LSTM model (similar to <a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference\">https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference</a>), experimented with dropouts, num of parameters, training generators, undersampling, class imbalance, custom losses, augmentations, etc.</p>\n\n<p>I've experimented with number of frames (16 and 32) and with changing the gap between them (sequential frames from different parts of video, every third frame and 2-3 frames per second).</p>\n\n<p>I've sad that it didn't perform even close to single CNN model, I don't know the reason why. Maybe I needed to experiment more or did some mistakes in architecture or training progress.</p>\n\n<p><strong>Public Score: 0.44083 (The best one across many experimental tries. Based on EFNet-4)</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1910.12467\">Use of a Capsule Network to Detect Fake Images and Videos</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/nii-yamagishilab/Capsule-Forensics-v2\">https://github.com/nii-yamagishilab/Capsule-Forensics-v2</a></p>\n\n<p>Trained it for few epochs on subset of DFDC dataset (1:1 real/fake). I suggest it could perform much better with more experiments, but I didn't have time for it. This is the method which I could regret I didn't pay much attention to.</p>\n\n<p><strong>Public Score: 0.60641</strong></p>\n\n<ul>\n<li><a href=\"https://hal.inria.fr/hal-02140558/document\">MARS: Motion-Augmented RGB Stream for Action Recognition</a></li>\n</ul>\n\n<p>Code: <a href=\"https://github.com/craston/MARS\">https://github.com/craston/MARS</a></p>\n\n<p>This is the only method I didn't have time to fully experiment with. The problem is in preparation of optical flows for training. Even with CUDA compiled OpenCV, it takes around 100 hours to extract x and y direction optical flows from 32 frames (the method requires sequences) on my PC. I so much wanted to train/test this approach, but due to the time constraints I've decided to stop it. And this is the second and the last method I could regret I didn't test.</p>\n\n<h1>My Solution</h1>\n\n<p><strong>Public score: 0.29477</strong></p>\n\n<p>The notebook is available here: <a href=\"https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15\">https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15</a></p>\n\n<p>Now I will describe step by step how I've been improving single CNN model and ended up with ensemble. I'm kind of disappointed that after experimenting with other approaches not any of those came closer to single CNN model. But okay, let's continue.</p>\n\n<h3>Data preparation</h3>\n\n<p>Video reader: <a href=\"https://www.kaggle.com/humananalog/deepfakes-inference-demo\">https://www.kaggle.com/humananalog/deepfakes-inference-demo</a> (Big thanks to <a href=\"/humananalog\">@humananalog</a>)\nFace extractor: <a href=\"https://github.com/1adrianb/face-alignment\">https://github.com/1adrianb/face-alignment</a></p>\n\n<p>I've used this specific face extractor, because I wanted to train my model on face only, not the whole head. Since the lib gives landmarks, I could crop face region easily. After some time I've realised that landmarks extraction takes much time (even though it fits to extract 16-20 frames and process them within 8 secs) and decided to keep only detection part, removing all the stuff related to landmarks. The detector used in the lib is S3FD. I've additionally used two more face detection models (MTCNN and BlazeFace) for better confidence. Using additional detectors helped me to remove some FPs and save some FNs. This is especially true for very dark images. Finally I've came up with extracting 32 frames for each person from each video. The reason for that is because I had many experiments with Conv-LSTM models. After all I had around 4M .png images of 224x224 size, but I didn't use so big number of faces for training my final single CNN model. I've randomly picked half of those (so I had 16 faces per person).</p>\n\n<h3>Training</h3>\n\n<p>Training EFNet-4 with setting Keras <code>class_weights</code> during training worked well for me. But I've found out better and what's more important much faster way to train my model. The idea is to balance fake and real data on each epoch. Since the training data has around 5:1 imbalance, each epoch we take all real examples but only 0.2 part of fake examples. In such approach 5 epochs is enough for model to see all examples. But i've experimentally found out that 6 epochs gave the best results. During the 6th epoch I've randomly picked fake examples. </p>\n\n<p>I was afraid that my model would overfit to real faces, but seems like it didn't. Augmentations helped to solve this issue. Couple of words about augmentations. I've used different albumentations functions, but during training I've checked if the image is too dark and applied more brightness with larger aug probability.</p>\n\n<p>If someone needs the generator, here is the code: (I'm too lazy now to clean up my codes and release on github). Note that here <code>fake_images</code> list is 5 times larger than <code>real_images</code>. Amount of images for each epoch was around 750K (350K real and 350K fake).</p>\n\n<pre><code>def generator(real_images, fake_images, preprocess_input_fn, batch_size=1, input_shape=(224,224,3), do_aug=False):\n\n    i = 0\n    k = 0\n\n    images = real_images + fake_images[k*len(real_images):(k+1)*len(real_images)]\n    random.shuffle(images)\n\n    while True:\n\n        x_batch = np.zeros((batch_size, input_shape[0], input_shape[1], 3))\n        y_batch = np.zeros((batch_size), dtype=np.single)\n\n        for b in range(batch_size):\n\n            if i == len(images):\n                i = 0\n                k += 1\n                if k &amp;gt; (len(fake_images)//len(real_images) - 1):\n                    images = real_images + [str(i) for i in np.random.choice(fake_images, size = len(real_images), replace = False)]\n                else:\n                    images = real_images + fake_images[k*len(real_images):(k+1)*len(real_images)]\n                random.shuffle(images)\n\n            x = Image.open(images[i])\n            x = np.array(x) \n\n            if do_aug:\n                x = augment(x)\n\n            y = 1. if '/fake/' in images[i] else 0.\n\n            x_batch[b] = x\n            y_batch[b] = y             \n\n            i += 1\n\n        x_batch = preprocess_input_fn(x_batch)\n\n        yield (x_batch, y_batch)\n</code></pre>\n\n<p>After I've trained several EFNet-4 models, I've found out that all models perform best after 6 epochs. Then I've started to train models for ensemble without validation. I've used RAdam optimizer with initial learning rate of 0.001. This learning rate was changed for 4th and 5th epochs to 0.0005 and for the last 6th epoch I've used 0.0001. Loss is binary crossentropy.</p>\n\n<h3>Inference</h3>\n\n<p>In the last two or three days I decided to switch from my EFNet-4 ensemble of 13 models to EFNet-6 ensemble of 5 models and trained them. Boost was around 0.005. Inference could be completely viewed in my released notebook here: <a href=\"https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15\">https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15</a>.</p>\n\n<h3>Tricks that worked</h3>\n\n<ul>\n<li>Using <code>read_frames_at_indices()</code> along with <code>capture.get(cv2.CAP_PROP_FRAME_COUNT)</code> instead of <code>read_frames()</code>. I don't know the reason why, but it works a little bit faster.</li>\n<li>Resizing video for the sake of saving time for detector.</li>\n<li>Removing small detected objects. These are mainly false positives and don't contain faces. I've used the threshold of <code>min(video.shape)//20</code>.</li>\n<li>Zooming out of face. Didn't find a reason why, but it works better.</li>\n<li>Median averaging across frames. I was surpised, but it gave a boost of more than 0.01 compared to mean averaging.</li>\n</ul>\n\n<h3>Tricks that didn't work</h3>\n\n<ul>\n<li><strong>Clipping</strong>. It worked for single models, but for ensembles of 5+ models it always made worse. Even though for safety I've submitted two versions, one with 0.01-0.99 clip and the second one without it (well there was a clip but very small one of 1e-15).</li>\n<li><strong>TTA</strong>. I've tried with complex TTAs as well as with simple Horizontal Flip. But it turned out that it is better to process more frames without TTA, then apply TTA to less frames.</li>\n<li><strong>Judge by one person</strong>. I've tested the approach of giving FAKE prediction if one of the persons in video is FAKE. I've sorted the multiple faces by x-coordinates, and then gave prediction to each person, finally giving FAKE prediction if one is FAKE. It performs slightly worse than judge all persons equally.</li>\n</ul>\n\n<p><strong>Public score: 0.34 (for single EFNet-4 model)</strong>\n*<em>Public score: 0.30 (for ensemble of 13 EFNet-4 models)</em>*\n<strong>Public score: 0.295 (for ensemble of 5 EFNet-6 models)</strong></p>",
      "rawMarkdown": "Hi to all the competitors and organizers of DFDC, this was my first Kaggle competition and here I want to share the methods I've experimented with during the competition and share the details of how I've finally crossed the 0.3 border. \n\nFirstly I would like to write a list of methods that didn't work for me (or they worked but not as good as my final solution).\n\n# Classical methods\n- [Exposing Deep Fakes Using Inconsistent Head Poses](https://arxiv.org/abs/1811.00661)\n\nCode: https://bitbucket.org/ericyang3721/headpose_forensic/src/master/\nIn few words this method uses difference between face and head poses. Firstly I've implemented it myself, then found the code from authors on bitbucket. Approach didn't work for the majority of cases. I guess it only works good for FaceSwap faces.\n\nDidn't test it against Public LB\n\n- [Unmasking DeepFakes with simple Features](https://arxiv.org/abs/1911.00686)\n\nCode: https://github.com/cc-hpc-itwm/DeepFakeDetection\nBased on analysis of high frequency components in fourier spectra. I've tested the method on a part of dfdc dataset, found out that it works, yet not as good to beat single CNN model.\n\nDidn't test it against Public LB\n\n# Deep Learning-based methods\n- [Exposing DeepFake Videos By Detecting Face Warping Artifacts](https://arxiv.org/abs/1811.00656)\n\nCode: https://github.com/danmohaha/CVPRW2019_Face_Artifacts\nThis was my first submission to the competition. I've used pretrained model provided by authors.\n\n**Public Score: 1.16260**\n\n- [Detecting Face2Face Facial Reenactment in Videos](https://arxiv.org/abs/2001.07444)\n\nI've liked this approach and it even improved my score at some point. I've used ResNet-18 models as stated in the paper and compared it to single ResNet-18 model that was trained on the same data. The improvement in public score was about 0.045. This is either because the approach really works and paying attention to different face regions really matters or because this is just a kind of ensembling (5 similar models, but trained on different image regions). Or maybe both, I don't know 😃. But this approach didn't work with EfficientNet models, which I've used later. \n\n**Public Score: 0.55683 (for single ResNet-18)**\n**Public Score: 0.51062 (for ResNets-18)**\n**Public Score: 0.38760 (for EFNets-4)**\n\n- [MesoNet: a Compact Facial Video Forgery Detection Network](https://arxiv.org/abs/1809.00888)\n\nCode: https://github.com/DariusAf/MesoNet\nI've tested only pretrained models. Didn't train it on dfdc dataset.\n\n**Public Score: 1.54918**\n\n- [CNN-generated images are surprisingly easy to spot... for now](https://arxiv.org/abs/1912.11035)\n\nCode: https://github.com/peterwang512/CNNDetection\nI've already gave a thought about this approach here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134295#766094\n\nDidn't test it against Public LB\n\n- [Deepfake Video Detection Using Recurrent Neural Networks](https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf)\n\nThe method which took the majority of my time, for some reason I've thought it has to outperform single CNN model. I've spend several weeks experimenting with Conv-LSTM models. \n\nI've firstly tried to use CNN model as feature extractor (removing last prediction layer) and train LSTM model on those features.\n\nI've tried CNN-LSTM model (similar to https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference), experimented with dropouts, num of parameters, training generators, undersampling, class imbalance, custom losses, augmentations, etc.\n\nI've experimented with number of frames (16 and 32) and with changing the gap between them (sequential frames from different parts of video, every third frame and 2-3 frames per second).\n\nI've sad that it didn't perform even close to single CNN model, I don't know the reason why. Maybe I needed to experiment more or did some mistakes in architecture or training progress.\n\n**Public Score: 0.44083 (The best one across many experimental tries. Based on EFNet-4)**\n\n- [Use of a Capsule Network to Detect Fake Images and Videos](https://arxiv.org/abs/1910.12467)\n\nCode: https://github.com/nii-yamagishilab/Capsule-Forensics-v2\n\nTrained it for few epochs on subset of DFDC dataset (1:1 real/fake). I suggest it could perform much better with more experiments, but I didn't have time for it. This is the method which I could regret I didn't pay much attention to.\n\n**Public Score: 0.60641**\n\n- [MARS: Motion-Augmented RGB Stream for Action Recognition](https://hal.inria.fr/hal-02140558/document)\n\nCode: https://github.com/craston/MARS\n\nThis is the only method I didn't have time to fully experiment with. The problem is in preparation of optical flows for training. Even with CUDA compiled OpenCV, it takes around 100 hours to extract x and y direction optical flows from 32 frames (the method requires sequences) on my PC. I so much wanted to train/test this approach, but due to the time constraints I've decided to stop it. And this is the second and the last method I could regret I didn't test.\n\n# My Solution\n\n**Public score: 0.29477**\n\nThe notebook is available here: https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15\n\nNow I will describe step by step how I've been improving single CNN model and ended up with ensemble. I'm kind of disappointed that after experimenting with other approaches not any of those came closer to single CNN model. But okay, let's continue.\n\n### Data preparation\nVideo reader: https://www.kaggle.com/humananalog/deepfakes-inference-demo (Big thanks to @humananalog)\nFace extractor: https://github.com/1adrianb/face-alignment\n\nI've used this specific face extractor, because I wanted to train my model on face only, not the whole head. Since the lib gives landmarks, I could crop face region easily. After some time I've realised that landmarks extraction takes much time (even though it fits to extract 16-20 frames and process them within 8 secs) and decided to keep only detection part, removing all the stuff related to landmarks. The detector used in the lib is S3FD. I've additionally used two more face detection models (MTCNN and BlazeFace) for better confidence. Using additional detectors helped me to remove some FPs and save some FNs. This is especially true for very dark images. Finally I've came up with extracting 32 frames for each person from each video. The reason for that is because I had many experiments with Conv-LSTM models. After all I had around 4M .png images of 224x224 size, but I didn't use so big number of faces for training my final single CNN model. I've randomly picked half of those (so I had 16 faces per person).\n\n### Training\nTraining EFNet-4 with setting Keras `class_weights` during training worked well for me. But I've found out better and what's more important much faster way to train my model. The idea is to balance fake and real data on each epoch. Since the training data has around 5:1 imbalance, each epoch we take all real examples but only 0.2 part of fake examples. In such approach 5 epochs is enough for model to see all examples. But i've experimentally found out that 6 epochs gave the best results. During the 6th epoch I've randomly picked fake examples. \n\nI was afraid that my model would overfit to real faces, but seems like it didn't. Augmentations helped to solve this issue. Couple of words about augmentations. I've used different albumentations functions, but during training I've checked if the image is too dark and applied more brightness with larger aug probability.\n\nIf someone needs the generator, here is the code: (I'm too lazy now to clean up my codes and release on github). Note that here `fake_images` list is 5 times larger than `real_images`. Amount of images for each epoch was around 750K (350K real and 350K fake).\n\n    def generator(real_images, fake_images, preprocess_input_fn, batch_size=1, input_shape=(224,224,3), do_aug=False):\n    \n        i = 0\n        k = 0\n    \n        images = real_images + fake_images[k*len(real_images):(k+1)*len(real_images)]\n        random.shuffle(images)\n\n        while True:\n        \n            x_batch = np.zeros((batch_size, input_shape[0], input_shape[1], 3))\n            y_batch = np.zeros((batch_size), dtype=np.single)\n        \n            for b in range(batch_size):\n            \n                if i == len(images):\n                    i = 0\n                    k += 1\n                    if k &gt; (len(fake_images)//len(real_images) - 1):\n                        images = real_images + [str(i) for i in np.random.choice(fake_images, size = len(real_images), replace = False)]\n                    else:\n                        images = real_images + fake_images[k*len(real_images):(k+1)*len(real_images)]\n                    random.shuffle(images)\n                \n                x = Image.open(images[i])\n                x = np.array(x) \n            \n                if do_aug:\n                    x = augment(x)\n\n                y = 1. if '/fake/' in images[i] else 0.\n            \n                x_batch[b] = x\n                y_batch[b] = y             \n                \n                i += 1\n            \n            x_batch = preprocess_input_fn(x_batch)\n    \n            yield (x_batch, y_batch)\n\nAfter I've trained several EFNet-4 models, I've found out that all models perform best after 6 epochs. Then I've started to train models for ensemble without validation. I've used RAdam optimizer with initial learning rate of 0.001. This learning rate was changed for 4th and 5th epochs to 0.0005 and for the last 6th epoch I've used 0.0001. Loss is binary crossentropy.\n\n### Inference\nIn the last two or three days I decided to switch from my EFNet-4 ensemble of 13 models to EFNet-6 ensemble of 5 models and trained them. Boost was around 0.005. Inference could be completely viewed in my released notebook here: https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15.\n\n### Tricks that worked\n- Using `read_frames_at_indices()` along with `capture.get(cv2.CAP_PROP_FRAME_COUNT)` instead of `read_frames()`. I don't know the reason why, but it works a little bit faster.\n- Resizing video for the sake of saving time for detector.\n- Removing small detected objects. These are mainly false positives and don't contain faces. I've used the threshold of `min(video.shape)//20`.\n- Zooming out of face. Didn't find a reason why, but it works better.\n- Median averaging across frames. I was surpised, but it gave a boost of more than 0.01 compared to mean averaging.\n\n### Tricks that didn't work\n- **Clipping**. It worked for single models, but for ensembles of 5+ models it always made worse. Even though for safety I've submitted two versions, one with 0.01-0.99 clip and the second one without it (well there was a clip but very small one of 1e-15).\n- **TTA**. I've tried with complex TTAs as well as with simple Horizontal Flip. But it turned out that it is better to process more frames without TTA, then apply TTA to less frames.\n- **Judge by one person**. I've tested the approach of giving FAKE prediction if one of the persons in video is FAKE. I've sorted the multiple faces by x-coordinates, and then gave prediction to each person, finally giving FAKE prediction if one is FAKE. It performs slightly worse than judge all persons equally.\n\n**Public score: 0.34 (for single EFNet-4 model)**\n**Public score: 0.30 (for ensemble of 13 EFNet-4 models)**\n**Public score: 0.295 (for ensemble of 5 EFNet-6 models)**",
      "votes": 20
    },
    {
      "id": 795401,
      "postDate": "2020-04-02T17:24:58.413Z",
      "content": "<p>Interesting to see what did not work for you. I started the competition with capsule networks and I got the same conclusions as you, my results on around half data (folder 0 to 25):\n- VGG19/Capsulesx10: LB=0.62 (CV4 with Augmentation, 32 first frames, 256x256)\n- ResNeXt/Capsulesx10: LB=0.63 (CV4, 23 frames, 256x256)</p>\n\n<p>Many papers claim they've the best approach and give comparisons to other approaches but none have the truth. It always depends on data, quality, hyper-parameters ...</p>",
      "rawMarkdown": "Interesting to see what did not work for you. I started the competition with capsule networks and I got the same conclusions as you, my results on around half data (folder 0 to 25):\n- VGG19/Capsulesx10: LB=0.62 (CV4 with Augmentation, 32 first frames, 256x256)\n- ResNeXt/Capsulesx10: LB=0.63 (CV4, 23 frames, 256x256)\n\nMany papers claim they've the best approach and give comparisons to other approaches but none have the truth. It always depends on data, quality, hyper-parameters ...",
      "votes": 2,
      "replies": [
        {
          "id": 795543,
          "postDate": "2020-04-02T20:05:46.160Z",
          "content": "<p>A lot of papers also use accuracy as the metric, but logloss is much less forgiving, so the Kaggle competition is much harder than the problem these papers are solving.</p>",
          "rawMarkdown": "A lot of papers also use accuracy as the metric, but logloss is much less forgiving, so the Kaggle competition is much harder than the problem these papers are solving.",
          "votes": 2
        },
        {
          "id": 795569,
          "postDate": "2020-04-02T20:44:19.660Z",
          "content": "<p>Agree</p>",
          "rawMarkdown": "Agree"
        },
        {
          "id": 795622,
          "postDate": "2020-04-02T21:57:37.930Z",
          "content": "<p>True, but even high accuracy is difficult to reach.</p>",
          "rawMarkdown": "True, but even high accuracy is difficult to reach.",
          "votes": 1
        },
        {
          "id": 796079,
          "postDate": "2020-04-03T09:18:02.463Z",
          "content": "<p>you may have 90% accuracy with infinite logloss :/</p>",
          "rawMarkdown": "you may have 90% accuracy with infinite logloss :/"
        }
      ]
    },
    {
      "id": 1325436,
      "postDate": "2021-05-27T18:41:07.753Z",
      "content": "<p>Hi, thank you for sharing your detailed study!<br>\nCan you please share the notebook or give an idea of what preprocessing you did for this paper \"Unmasking DeepFakes with simple Features\". Somehow I can't reproduce the results that they say should appear and end up getting overlapping  plots.</p>",
      "rawMarkdown": "Hi, thank you for sharing your detailed study!\nCan you please share the notebook or give an idea of what preprocessing you did for this paper \"Unmasking DeepFakes with simple Features\". Somehow I can't reproduce the results that they say should appear and end up getting overlapping  plots.\n"
    },
    {
      "id": 795857,
      "postDate": "2020-04-03T05:06:06.517Z",
      "content": "<p>Interestingly, I used median averging too. Gave me a boost of 0.03LB in one of the models that I used in my final ensemble! 0.34LB to 0.31LB.</p>",
      "rawMarkdown": "Interestingly, I used median averging too. Gave me a boost of 0.03LB in one of the models that I used in my final ensemble! 0.34LB to 0.31LB."
    },
    {
      "id": 795725,
      "postDate": "2020-04-03T01:41:12.213Z",
      "content": "<p>Thanks for sharing your approach. \nI wanted to ask about this:</p>\n\n<blockquote>\n  <p>Unmasking DeepFakes with simple Features\n  Based on analysis of high frequency components in fourier spectra. I've tested the method on a part of dfdc dataset, found out that it works, yet not as good to beat single CNN model.</p>\n</blockquote>\n\n<p>I tried that as well, but it didn't work for me (I shared the plot here <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134216#768859\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134216#768859</a>)\nWhat was the crop size or any pre-processing that made it work for you ? </p>",
      "rawMarkdown": "Thanks for sharing your approach. \nI wanted to ask about this:\n&gt; Unmasking DeepFakes with simple Features\n&gt; Based on analysis of high frequency components in fourier spectra. I've tested the method on a part of dfdc dataset, found out that it works, yet not as good to beat single CNN model.\n\nI tried that as well, but it didn't work for me (I shared the plot here https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134216#768859)\nWhat was the crop size or any pre-processing that made it work for you ? ",
      "replies": [
        {
          "id": 796038,
          "postDate": "2020-04-03T08:25:58.753Z",
          "content": "<p>Nothing for preprocessing, just used the codes from authors. By 'it worked' I mean that it worked well for some deepfakes, but for majority it didn't. So I guess it can be used for some specific df method where the pixels change is significant on the face borders or eyes.</p>",
          "rawMarkdown": "Nothing for preprocessing, just used the codes from authors. By 'it worked' I mean that it worked well for some deepfakes, but for majority it didn't. So I guess it can be used for some specific df method where the pixels change is significant on the face borders or eyes.",
          "votes": 1
        }
      ]
    },
    {
      "id": 802574,
      "postDate": "2020-04-09T16:18:33.863Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 802565,
      "postDate": "2020-04-09T16:13:04.950Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 795401,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-04-02T17:24:58.413000",
      "content": "<p>Interesting to see what did not work for you. I started the competition with capsule networks and I got the same conclusions as you, my results on around half data (folder 0 to 25):\n- VGG19/Capsulesx10: LB=0.62 (CV4 with Augmentation, 32 first frames, 256x256)\n- ResNeXt/Capsulesx10: LB=0.63 (CV4, 23 frames, 256x256)</p>\n\n<p>Many papers claim they've the best approach and give comparisons to other approaches but none have the truth. It always depends on data, quality, hyper-parameters ...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 795543,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-04-02T20:05:46.160000",
          "content": "<p>A lot of papers also use accuracy as the metric, but logloss is much less forgiving, so the Kaggle competition is much harder than the problem these papers are solving.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 795569,
          "author_name": "Vladislav Ostankovich",
          "author_url": "",
          "post_date": "2020-04-02T20:44:19.660000",
          "content": "<p>Agree</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 795622,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-04-02T21:57:37.930000",
          "content": "<p>True, but even high accuracy is difficult to reach.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 796079,
          "author_name": "yuanzhe zhou",
          "author_url": "",
          "post_date": "2020-04-03T09:18:02.463000",
          "content": "<p>you may have 90% accuracy with infinite logloss :/</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1325436,
      "author_name": "Pratim Ugale",
      "author_url": "",
      "post_date": "2021-05-27T18:41:07.753000",
      "content": "<p>Hi, thank you for sharing your detailed study!<br>\nCan you please share the notebook or give an idea of what preprocessing you did for this paper \"Unmasking DeepFakes with simple Features\". Somehow I can't reproduce the results that they say should appear and end up getting overlapping  plots.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 795857,
      "author_name": "Debanga Raj Neog",
      "author_url": "",
      "post_date": "2020-04-03T05:06:06.517000",
      "content": "<p>Interestingly, I used median averging too. Gave me a boost of 0.03LB in one of the models that I used in my final ensemble! 0.34LB to 0.31LB.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 795725,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2020-04-03T01:41:12.213000",
      "content": "<p>Thanks for sharing your approach. \nI wanted to ask about this:</p>\n\n<blockquote>\n  <p>Unmasking DeepFakes with simple Features\n  Based on analysis of high frequency components in fourier spectra. I've tested the method on a part of dfdc dataset, found out that it works, yet not as good to beat single CNN model.</p>\n</blockquote>\n\n<p>I tried that as well, but it didn't work for me (I shared the plot here <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134216#768859\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134216#768859</a>)\nWhat was the crop size or any pre-processing that made it work for you ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 796038,
          "author_name": "Vladislav Ostankovich",
          "author_url": "",
          "post_date": "2020-04-03T08:25:58.753000",
          "content": "<p>Nothing for preprocessing, just used the codes from authors. By 'it worked' I mean that it worked well for some deepfakes, but for majority it didn't. So I guess it can be used for some specific df method where the pixels change is significant on the face borders or eyes.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 802574,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-09T16:18:33.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 802565,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-09T16:13:04.950000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "795384": "Hi to all the competitors and organizers of DFDC, this was my first Kaggle competition and here I want to share the methods I've experimented with during the competition and share the details of how I've finally crossed the 0.3 border. \n\nFirstly I would like to write a list of methods that didn't work for me (or they worked but not as good as my final solution).\n\n# Classical methods\n- [Exposing Deep Fakes Using Inconsistent Head Poses](https://arxiv.org/abs/1811.00661)\n\nCode: https://bitbucket.org/ericyang3721/headpose_forensic/src/master/\nIn few words this method uses difference between face and head poses. Firstly I've implemented it myself, then found the code from authors on bitbucket. Approach didn't work for the majority of cases. I guess it only works good for FaceSwap faces.\n\nDidn't test it against Public LB\n\n- [Unmasking DeepFakes with simple Features](https://arxiv.org/abs/1911.00686)\n\nCode: https://github.com/cc-hpc-itwm/DeepFakeDetection\nBased on analysis of high frequency components in fourier spectra. I've tested the method on a part of dfdc dataset, found out that it works, yet not as good to beat single CNN model.\n\nDidn't test it against Public LB\n\n# Deep Learning-based methods\n- [Exposing DeepFake Videos By Detecting Face Warping Artifacts](https://arxiv.org/abs/1811.00656)\n\nCode: https://github.com/danmohaha/CVPRW2019_Face_Artifacts\nThis was my first submission to the competition. I've used pretrained model provided by authors.\n\n**Public Score: 1.16260**\n\n- [Detecting Face2Face Facial Reenactment in Videos](https://arxiv.org/abs/2001.07444)\n\nI've liked this approach and it even improved my score at some point. I've used ResNet-18 models as stated in the paper and compared it to single ResNet-18 model that was trained on the same data. The improvement in public score was about 0.045. This is either because the approach really works and paying attention to different face regions really matters or because this is just a kind of ensembling (5 similar models, but trained on different image regions). Or maybe both, I don't know 😃. But this approach didn't work with EfficientNet models, which I've used later. \n\n**Public Score: 0.55683 (for single ResNet-18)**\n**Public Score: 0.51062 (for ResNets-18)**\n**Public Score: 0.38760 (for EFNets-4)**\n\n- [MesoNet: a Compact Facial Video Forgery Detection Network](https://arxiv.org/abs/1809.00888)\n\nCode: https://github.com/DariusAf/MesoNet\nI've tested only pretrained models. Didn't train it on dfdc dataset.\n\n**Public Score: 1.54918**\n\n- [CNN-generated images are surprisingly easy to spot... for now](https://arxiv.org/abs/1912.11035)\n\nCode: https://github.com/peterwang512/CNNDetection\nI've already gave a thought about this approach here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134295#766094\n\nDidn't test it against Public LB\n\n- [Deepfake Video Detection Using Recurrent Neural Networks](https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf)\n\nThe method which took the majority of my time, for some reason I've thought it has to outperform single CNN model. I've spend several weeks experimenting with Conv-LSTM models. \n\nI've firstly tried to use CNN model as feature extractor (removing last prediction layer) and train LSTM model on those features.\n\nI've tried CNN-LSTM model (similar to https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference), experimented with dropouts, num of parameters, training generators, undersampling, class imbalance, custom losses, augmentations, etc.\n\nI've experimented with number of frames (16 and 32) and with changing the gap between them (sequential frames from different parts of video, every third frame and 2-3 frames per second).\n\nI've sad that it didn't perform even close to single CNN model, I don't know the reason why. Maybe I needed to experiment more or did some mistakes in architecture or training progress.\n\n**Public Score: 0.44083 (The best one across many experimental tries. Based on EFNet-4)**\n\n- [Use of a Capsule Network to Detect Fake Images and Videos](https://arxiv.org/abs/1910.12467)\n\nCode: https://github.com/nii-yamagishilab/Capsule-Forensics-v2\n\nTrained it for few epochs on subset of DFDC dataset (1:1 real/fake). I suggest it could perform much better with more experiments, but I didn't have time for it. This is the method which I could regret I didn't pay much attention to.\n\n**Public Score: 0.60641**\n\n- [MARS: Motion-Augmented RGB Stream for Action Recognition](https://hal.inria.fr/hal-02140558/document)\n\nCode: https://github.com/craston/MARS\n\nThis is the only method I didn't have time to fully experiment with. The problem is in preparation of optical flows for training. Even with CUDA compiled OpenCV, it takes around 100 hours to extract x and y direction optical flows from 32 frames (the method requires sequences) on my PC. I so much wanted to train/test this approach, but due to the time constraints I've decided to stop it. And this is the second and the last method I could regret I didn't test.\n\n# My Solution\n\n**Public score: 0.29477**\n\nThe notebook is available here: https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15\n\nNow I will describe step by step how I've been improving single CNN model and ended up with ensemble. I'm kind of disappointed that after experimenting with other approaches not any of those came closer to single CNN model. But okay, let's continue.\n\n### Data preparation\nVideo reader: https://www.kaggle.com/humananalog/deepfakes-inference-demo (Big thanks to @humananalog)\nFace extractor: https://github.com/1adrianb/face-alignment\n\nI've used this specific face extractor, because I wanted to train my model on face only, not the whole head. Since the lib gives landmarks, I could crop face region easily. After some time I've realised that landmarks extraction takes much time (even though it fits to extract 16-20 frames and process them within 8 secs) and decided to keep only detection part, removing all the stuff related to landmarks. The detector used in the lib is S3FD. I've additionally used two more face detection models (MTCNN and BlazeFace) for better confidence. Using additional detectors helped me to remove some FPs and save some FNs. This is especially true for very dark images. Finally I've came up with extracting 32 frames for each person from each video. The reason for that is because I had many experiments with Conv-LSTM models. After all I had around 4M .png images of 224x224 size, but I didn't use so big number of faces for training my final single CNN model. I've randomly picked half of those (so I had 16 faces per person).\n\n### Training\nTraining EFNet-4 with setting Keras `class_weights` during training worked well for me. But I've found out better and what's more important much faster way to train my model. The idea is to balance fake and real data on each epoch. Since the training data has around 5:1 imbalance, each epoch we take all real examples but only 0.2 part of fake examples. In such approach 5 epochs is enough for model to see all examples. But i've experimentally found out that 6 epochs gave the best results. During the 6th epoch I've randomly picked fake examples. \n\nI was afraid that my model would overfit to real faces, but seems like it didn't. Augmentations helped to solve this issue. Couple of words about augmentations. I've used different albumentations functions, but during training I've checked if the image is too dark and applied more brightness with larger aug probability.\n\nIf someone needs the generator, here is the code: (I'm too lazy now to clean up my codes and release on github). Note that here `fake_images` list is 5 times larger than `real_images`. Amount of images for each epoch was around 750K (350K real and 350K fake).\n\n    def generator(real_images, fake_images, preprocess_input_fn, batch_size=1, input_shape=(224,224,3), do_aug=False):\n    \n        i = 0\n        k = 0\n    \n        images = real_images + fake_images[k*len(real_images):(k+1)*len(real_images)]\n        random.shuffle(images)\n\n        while True:\n        \n            x_batch = np.zeros((batch_size, input_shape[0], input_shape[1], 3))\n            y_batch = np.zeros((batch_size), dtype=np.single)\n        \n            for b in range(batch_size):\n            \n                if i == len(images):\n                    i = 0\n                    k += 1\n                    if k &gt; (len(fake_images)//len(real_images) - 1):\n                        images = real_images + [str(i) for i in np.random.choice(fake_images, size = len(real_images), replace = False)]\n                    else:\n                        images = real_images + fake_images[k*len(real_images):(k+1)*len(real_images)]\n                    random.shuffle(images)\n                \n                x = Image.open(images[i])\n                x = np.array(x) \n            \n                if do_aug:\n                    x = augment(x)\n\n                y = 1. if '/fake/' in images[i] else 0.\n            \n                x_batch[b] = x\n                y_batch[b] = y             \n                \n                i += 1\n            \n            x_batch = preprocess_input_fn(x_batch)\n    \n            yield (x_batch, y_batch)\n\nAfter I've trained several EFNet-4 models, I've found out that all models perform best after 6 epochs. Then I've started to train models for ensemble without validation. I've used RAdam optimizer with initial learning rate of 0.001. This learning rate was changed for 4th and 5th epochs to 0.0005 and for the last 6th epoch I've used 0.0001. Loss is binary crossentropy.\n\n### Inference\nIn the last two or three days I decided to switch from my EFNet-4 ensemble of 13 models to EFNet-6 ensemble of 5 models and trained them. Boost was around 0.005. Inference could be completely viewed in my released notebook here: https://www.kaggle.com/vostankovich/efnet6-ensemble-5-clip-1e-15.\n\n### Tricks that worked\n- Using `read_frames_at_indices()` along with `capture.get(cv2.CAP_PROP_FRAME_COUNT)` instead of `read_frames()`. I don't know the reason why, but it works a little bit faster.\n- Resizing video for the sake of saving time for detector.\n- Removing small detected objects. These are mainly false positives and don't contain faces. I've used the threshold of `min(video.shape)//20`.\n- Zooming out of face. Didn't find a reason why, but it works better.\n- Median averaging across frames. I was surpised, but it gave a boost of more than 0.01 compared to mean averaging.\n\n### Tricks that didn't work\n- **Clipping**. It worked for single models, but for ensembles of 5+ models it always made worse. Even though for safety I've submitted two versions, one with 0.01-0.99 clip and the second one without it (well there was a clip but very small one of 1e-15).\n- **TTA**. I've tried with complex TTAs as well as with simple Horizontal Flip. But it turned out that it is better to process more frames without TTA, then apply TTA to less frames.\n- **Judge by one person**. I've tested the approach of giving FAKE prediction if one of the persons in video is FAKE. I've sorted the multiple faces by x-coordinates, and then gave prediction to each person, finally giving FAKE prediction if one is FAKE. It performs slightly worse than judge all persons equally.\n\n**Public score: 0.34 (for single EFNet-4 model)**\n**Public score: 0.30 (for ensemble of 13 EFNet-4 models)**\n**Public score: 0.295 (for ensemble of 5 EFNet-6 models)**",
    "795401": "Interesting to see what did not work for you. I started the competition with capsule networks and I got the same conclusions as you, my results on around half data (folder 0 to 25):\n- VGG19/Capsulesx10: LB=0.62 (CV4 with Augmentation, 32 first frames, 256x256)\n- ResNeXt/Capsulesx10: LB=0.63 (CV4, 23 frames, 256x256)\n\nMany papers claim they've the best approach and give comparisons to other approaches but none have the truth. It always depends on data, quality, hyper-parameters ...",
    "1325436": "Hi, thank you for sharing your detailed study!\nCan you please share the notebook or give an idea of what preprocessing you did for this paper \"Unmasking DeepFakes with simple Features\". Somehow I can't reproduce the results that they say should appear and end up getting overlapping  plots.\n",
    "795857": "Interestingly, I used median averging too. Gave me a boost of 0.03LB in one of the models that I used in my final ensemble! 0.34LB to 0.31LB.",
    "795725": "Thanks for sharing your approach. \nI wanted to ask about this:\n&gt; Unmasking DeepFakes with simple Features\n&gt; Based on analysis of high frequency components in fourier spectra. I've tested the method on a part of dfdc dataset, found out that it works, yet not as good to beat single CNN model.\n\nI tried that as well, but it didn't work for me (I shared the plot here https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134216#768859)\nWhat was the crop size or any pre-processing that made it work for you ? ",
    "802574": "",
    "802565": ""
  }
}