{
  "id": 158506,
  "title": "2nd Place Removed Solution -QH Team ",
  "url": "/competitions/deepfake-detection-challenge/discussion/158506",
  "author_name": "",
  "post_date": "2020-06-14T13:51:31.367512Z",
  "votes": 70,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Interested in this task and of course also greatly attracted by the prize, we (five newbie kagglers) formed QH-Team. We are friends and all have computer vision backgrounds. This competition is very different from what we have faced in many other computer vision tasks. Many tricks were tried but proven to be not working on the public leaderboard (not sure if they work on the private leaderboard). Indeed, most of our time was spent in handling overfitting of the offline validation set. We struggled a lot to maintain our competency (Rank 4th~47th) in the public leaderboard throughout the entire competition period. </p>\n\n<p>It was a great surprise to us that we were ranked 2nd place (Private Leaderboard Scores:  0.42490&amp; 0.44431, Public Leaderboard Score: 0.29421) when the private leaderboard was released. We felt that the days and nights spent were well rewarded. Unfortunately, both results were removed due to the usage of faceforensic++ data even though this dataset can be accessed by all participants (we later confirmed this with dataset owners). We didn’t realize that this would be an issue during the competition. The rules are complicated to new Kagglers and we believe it would be much clearer to state that no external data is allowed for this competition at the very beginning. We are deeply disappointed by the final decision that both our scores are removed from the leaderboards. This simply erased all our efforts devoted to this competition. </p>\n\n<p>We would still like to thank Facebook and Kaggle for organizing this event. This competition draws huge attention from both industry and academia, it will definitely boost the development of deepfake detection techniques to fight against the fake videos. Here, we would like to present our solution to the people who are interested in this task. Hope it will inspire you if you are working in the related field. </p>\n\n<p>Meisong Zheng, Chuanrui Hu, Walton Tao Wang, Wenju Huang, Shengtao Xiao\nQH Team</p>\n\n<hr>\n\n<h1>1. Experiments</h1>\n\n<p>SGD optimizer with an initial learning rate of 0.01 and  x0.1 at epoch 2,6,10,14,17 is used to train our model. We used FocalLoss here. It took 10~20 hours to train a single model with 2 P40 GPUs. PyTorch was used during the training process.  (During fine-tuning process learning rate of 0.001 and  x0.1 at epoch 3,6,10,13,18 is used to tune our model)</p>\n\n<h2>1.1. Baseline Single Model</h2>\n\n<p>a) IR-SE-50[1-3] (a 50 layer IR-SE model with the similar structure as IR-SE-152 ) was trained with Part-22 as validation set and the rest parts as training set. Its LB score was 0.51639 with 30 frames used per video. </p>\n\n<p>b) We changed the backbone to a deeper CNN structure IR-SE-152. This boosted our performance to 0.436 from 0.51639.</p>\n\n<p>c) Removing the noisy cases with mask filtering,  IR-SE-152’s loss was further reduced to 0.40.</p>\n\n<p>d) Random cropping was revised a bit to introduce more perturbations during the training process.  BoundingBox X 1.3 -&gt; resize to 291x291-&gt; random cropping to 224x224. Center cropping was then used during inference. Our single model’s performance was improved to 0.376.</p>\n\n<p>e) We found that RA-92[4] behaves better than IR-SE-152 on various tasks and it has a smaller model size(190MB vs. 223MB). We changed our single model to RA-92 and got a new high score of 0.362.  </p>\n\n<p>f) We tried various powerful backbones but got little progress. </p>\n\n<p>g) We realized that using Part-22 as validation may hardly reflect the true performance of our model when tested online. This is mainly because many models from Part-22 also appear in other folders.  There are many suggestions about splitting training/validation sets from comments of experienced Kagglers and we then decided to use 20% of the folders as the validation set to prevent overfitting. Part0-9 were selected as the offline validation set. Our single RA-92 model loss was further reduced to 0.342. (We also tried many other combinations, Part0-9 gave us the best Public LB performance)</p>\n\n<p>h) To solve the multi-face issues, face tracking(see details in 4.3) was introduced during the inference process. More faces per video(120 frames uniformly extracted from 30th~270th frames) were also used for prediction. Our single model score was then improved to 0.331.</p>\n\n<p>i) We further fine-tuned the 0.331 single model with additional training data from Part5~9 (Part5~9 were offline validation set earlier) and FF++[7].  Random erasing[8] was used when we fine-tuned the model. In the end, we further reduced the LB loss to  0.323 (The best single model and used for final submission). \n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2F2ec5ee874012a6e00423e43913e6e6f0%2FRE.jpg?generation=1592141564255571&amp;alt=media\" alt=\"\"></p>\n\n<p>j) We tried sphere and water augmentation from DALI with a probe of 0.2 in the training process. This led to about 0.01 loss reduction on certain models. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fc1250ded525f812f1190462ecf141334%2Fwater_effect.png?generation=1592141734353937&amp;alt=media\" alt=\"\"></p>\n\n<p>k) We implemented a multi-cropping strategy during inference. It showed some performance enhancement. This trick was not used in the final submissions due to the inference time restriction. Using additional frames from video is more cost-effective and can also improve the performance.</p>\n\n<h2>1.2.  Model Ensembling</h2>\n\n<h3>1.2.1 Diversity of network structures and training data</h3>\n\n<p>How to effectively aggregate features of various models is important in our final submission. We had attempted to combine a few models 1) different models trained with identical data, 2) same models trained with different data, and 3) different models trained with slightly different data. The third one outperformed the rest and was used in the final submission.</p>\n\n<p>a): We trained RA-92、IR-SE-152 and SE-RESNEXT101-32x8d [5] with Part0~49 (except Part22 which was used for validation). Merging these three models  led to  0.025 performance enhancement from the best single model.\n| Model train with Part0-49(excluding Part22) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|RA-92 <br> IR-SE-152  <br>SE-RENEXT101-32x8d| 0.36528<br>0.37623<br>0.38449|<br>0.34058|</p>\n\n<p>b): We trained three RA-92 models with different training data. Ensembling of these models led to negligible performance enhancement. \n| Validation data(the rest as train data) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|Part0~9 <br> Part20~29  <br>Part30~39| 0.33166<br>0.3504<br>0.34027|<br>0.33145|</p>\n\n<p>c): We used RA-92, IR-SE-152 and SE_RESNEXT101_32x8d (trained with different fake data sampled from Part10-49).  In this case, model and data diversities were all preserved. It reached our best LB score. These models were then ensembled and submitted. \n| Model train with Part0-49(excluding Part22) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|RA-92 <br> IR-SE-152  <br>SE_RESNEXT101_32x8dV2| 0.323<br>0.321<br>0.34 (estimated)|<br>0.294|</p>\n\n<ul>\n<li>Remark: The SE_RESNEXT101_32x8dV2 model was trained the last day, and we didn’t have time to submit it alone, but it behaves better offline than a 0.351 model.  </li>\n</ul>\n\n<h3>1.2.2 Model Ensembling Method</h3>\n\n<p>Our final submission(0.294@LB) aggregates results of three models (0.32~0.34@LB in Table 4-2 ) in a frame-by-frame manner. The score of face from a typical frame is determined as below:</p>\n\n<pre><code>        frame_face_score = average( 2 largest scores); if 2 scores &gt;= 0.5\n        frame_face_score = average( 2 smallest scores ); if 2 scores &lt; 0.5\n</code></pre>\n\n<p>The above strategy was adopted as it showed strong robustness compared with other model ensembling techniques.  The conclusion was drawn when we experimented on 4 models RA-92, IR-SE-152, EfficientNet-B7 and SE-ResNext101_32x8d with LB scores around 0.37~0.39.  Offline and online results of ensembled models are shown in Table 4-2. (These are intermediate results only as the models haven’t reached their best performance at the moment when we tried model ensembling)</p>\n\n<p>|  |CV-RAW|CV-C40|CV-RES4|LB|\n| --- | --- |--- |--- |--- |\n|  V1-Eff |0.139|  0.333|  0.576|  0.343|\n|V1-SE| 0.156 | 0.318 | 0.516 |  0.34058|\n|V2| 0.136 | 0.335 | 0.621 | 0.391 |\n|V3| 0.148 | 0.347 | 0.549 | Not submitted |</p>\n\n<ul>\n<li>Results of model ensembling on raw, c40 compression and ¼ resolution validation set of part0~9. </li>\n<li>V1-Eff: Final ensembling strategy with (RA-92, IR-SE-152, EfficientNet-B7)；</li>\n<li>V1-SE: Final ensembling strategy with (SE-RESNEXT101_32x8d, RA-92, IR-SE-152)；</li>\n<li>V2: max( scores); if 2 scores &gt;= 0.5;min( scores ); if 2 scores &lt; 0.5；</li>\n<li>V3: max(scores) if all scores &gt;0.5; min(scores) if all scores &lt;0.5; else average(scores)；</li>\n</ul>\n\n<p>We also tried other ensembling methods, such as simple average. But they were not as good as V1-SE. </p>\n\n<h2>1.3  Inference</h2>\n\n<p>Our inference process is shown as Figure below. It mainly consists of 5 modules.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fde6d63cdd7625ceb078548d81c5c4f3a%2Fflow_chart.png?generation=1592141837895073&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li><p>VideoCapture: Our framework uniformly extracts 120 frames from 30-th to 270-th frames of every video.</p></li>\n<li><p>FaceDetector: Face detection is performed via CenterNet with batchsize of 60 frames. Our framework selects at most 2 faces with top detection confidence if more than 2 faces are detected within a single frame. Faces with low confidence are abandoned. </p></li>\n<li><p>FaceTracker: Face tracking is performed via direct IOU matching between bounding boxes from current and previous frames. Tracking sequences with less than 20 faces are ignored (most likely false detection). </p></li>\n<li><p>FaceClassifier: The system then predicts classification scores of all selected faces with a batch size of 120. Results of three classifiers are ensembled here for each face.   </p></li>\n<li><p>PostProcess: The video score is calculated as a weighted average of the frame scores with weights being the detection confidences. If there are multiple tracking sequences in a video, the maximum score is used as the final prediction of the entire video.  </p></li>\n</ul>\n\n<h1>2.What did not help</h1>\n\n<ul>\n<li><p>SlowFast: We tried the slowfast[9] network with slow-branch (mainly a 3D-convolutional network).  Even though it can reach comparable offline performance as the image-based 2D convolution method (e.g. RA92), the best LB score is only 0.477. It seems SlowFast  is easily overfitted. We didn’t spend too much time in this direction. </p></li>\n<li><p>Crop152: We had modified a CROP152 model as below. This model requires lots of computing  resources but leads to marginal performance improvement. We didn’t use these models later. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fab002e71ced88c23c98394d5b807c784%2Fcrop152.png?generation=1592141868496743&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Audio Classifier: As audio may also be faked, we trained a CNN-based audio classifier with fake audio clips from the last five folders and all real audios. Melgram is used as input to the network. Validation accuracy on the offline testset is about 0.98. However, the score on LB is only 0.6929 (slightly better than 0.5 prediction). Considering that  there may be very limited fake-audios in the Public Set, we abandoned the audio classifier.  </p></li>\n<li><p>Multi-task Learning with Mask Info: Lots of studies have shown that multi-task learning can boost performance of the core task. We designed a multi-task network which can distinguish real/fake samples and categorize the size of swapped areas (e.g. 0: no swapping; 1: 5% changed, 2:5%~20% changed ...). We believed that the auxiliary face area information can guide the CNN model to learn features which are closely related to the manipulated face areas.  Though sounds very attractive, the results showed negligible difference from the original classifier model. </p></li>\n<li><p>Face alignment: We also tried to pre-process the face images with face alignment. However, several submissions showed the scores deteriorate slightly if alignment is used. Facial landmark detection with Dlib also requires extra computational resources and landmark points are inaccurate when the faces are with large poses and under challenging illumination conditions. Considering these facts, we didn’t use face alignment later. Fig. 5-b below shows failure cases of face alignment when the face is under large poses. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2F3153a69fa728b2f75a40d68db630684f%2Falignment_failure.png?generation=1592141895213430&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Video-level Data Augmentations: We also tried the following video-level data augmentation methods[6]: re-encode the video as 15FPS,  reduce the resolution of the video to 1/4 of its original size  and increase the compression ratio to generate low quality video. Results are not improved on the public LB when models are trained with these data. These generated samples were therefore not used. </p></li>\n<li><p>Hard Sample Mining: Hard sample mining is widely adopted in many detection and classification tasks. We randomly selected those failure cases from the left-over faces (previously only 400K faces are used from the 1.1million cleaned faces) and added these faces into the fine-tuning process.  However, the score worsens after fine-tuning. </p></li>\n<li><p>LSTM:  We used the convolutional features from the trained image-based deepfake classifier as input to LSTM and tuned the LSTM parameters with features of face sequences. The LSTM hardly outperformed its corresponding image-based classifier on the offline validation set. We didn’t submit it. </p></li>\n<li><p>Distance based method: Triplet and softmax loss were combined to perform metric learning. This, however, didn’t work for us as well.</p></li>\n<li><p>Postprocess with Score Clipping and Mapping：Score clipping slightly improved our single model at the early stage of the competition when our LB score was about 0.4. However, it worsened our performance for the ensembled models.  We also tried to map the score such that the threshold cutting point was at 0.5 on the offline validation set. However, the score mapping trick led to very inconsistent results on the Public LB. Therefore, it was not used in the final submission. </p></li>\n<li><p>Data Size and Sampling Methods: Down-sampling of faked faces is better than up-sampling of the real faces. We tried to train our models with more training samples. This worsened our models’ scores. </p></li>\n<li><p>Training with Clusters: Like most of the participants, we also realized the overlapping ID issue. Thus, we attempted to split the dataset by ID clusters and train our model with it. However, the submitted score showed no improvement.</p></li>\n<li><p>Training tricks: Frozen/warm up/labelSmoothing deteriorates our training loss for our models.\n300x300:  Larger input size leads to more overfitting of our model. They behaved better on CV but worse on LB.  We thus fixed our input size to 224.   </p></li>\n<li><p>Reducing Cropping Margin: Hoping the network can learn more features from the facial region, we reduced the cropping margin. Validation accuracy was improved by this method. However, the loss increased to  0.411 from 0.34 on the Public LB.</p></li>\n<li><p>Error Level Analysis: We followed ELA[https://fotoforensics.com/tutorial-ela.php] to conduct Error Level Analysis. However, no obvious characteristic of the swapped face was identified. We didn’t try this method further.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fc37c8285ed977b705aa6a64c4ca7ba96%2Ffailure_cases.png?generation=1592141930858740&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n\n<h1>3. Reference</h1>\n\n<p>[1] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.\n[2] Hu J, Shen L, Sun G. Squeeze-and-excitation networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7132-7141.\n[3] Han D, Kim J, Kim J. Deep pyramidal residual networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 5927-5935.\n[4] Wang F, Jiang M, Qian C, et al. Residual attention network for image classification[C]. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 3156-3164.\n[5] Xie S, Girshick R, Dollár P, et al. Aggregated residual transformations for deep neural networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 1492-1500.\n[6] Dolhansky B, Howes R, Pflaum B, et al. The Deepfake Detection Challenge (DFDC) Preview Dataset[J]. arXiv preprint arXiv:1910.08854, 2019.\n[7] Rossler A, Cozzolino D, Verdoliva L, et al. Faceforensics++: Learning to detect manipulated facial images[C]//Proceedings of the IEEE International Conference on Computer Vision. 2019: 1-11.\n[8] Zhong Z, Zheng L, Kang G, et al. Random erasing data augmentation[J]. arXiv preprint arXiv:1708.04896, 2017.\n[9] Feichtenhofer C, Fan H, Malik J, et al. Slowfast networks for video recognition[C]. In Proceedings of the IEEE International Conference on Computer Vision. 2019: 6202-6211.</p>\n\n<hr>\n\n<h3>If you feel the above solution is helpful and hope to discuss more offline with us, you may contact Meisong via zhengmeisong@gmail.com</h3>",
  "messages": [
    {
      "id": "885812",
      "postDate": "06/14/2020 13:51:31",
      "content": "<p>Interested in this task and of course also greatly attracted by the prize, we (five newbie kagglers) formed QH-Team. We are friends and all have computer vision backgrounds. This competition is very different from what we have faced in many other computer vision tasks. Many tricks were tried but proven to be not working on the public leaderboard (not sure if they work on the private leaderboard). Indeed, most of our time was spent in handling overfitting of the offline validation set. We struggled a lot to maintain our competency (Rank 4th~47th) in the public leaderboard throughout the entire competition period. </p>\n\n<p>It was a great surprise to us that we were ranked 2nd place (Private Leaderboard Scores:  0.42490&amp; 0.44431, Public Leaderboard Score: 0.29421) when the private leaderboard was released. We felt that the days and nights spent were well rewarded. Unfortunately, both results were removed due to the usage of faceforensic++ data even though this dataset can be accessed by all participants (we later confirmed this with dataset owners). We didn’t realize that this would be an issue during the competition. The rules are complicated to new Kagglers and we believe it would be much clearer to state that no external data is allowed for this competition at the very beginning. We are deeply disappointed by the final decision that both our scores are removed from the leaderboards. This simply erased all our efforts devoted to this competition. </p>\n\n<p>We would still like to thank Facebook and Kaggle for organizing this event. This competition draws huge attention from both industry and academia, it will definitely boost the development of deepfake detection techniques to fight against the fake videos. Here, we would like to present our solution to the people who are interested in this task. Hope it will inspire you if you are working in the related field. </p>\n\n<p>Meisong Zheng, Chuanrui Hu, Walton Tao Wang, Wenju Huang, Shengtao Xiao\nQH Team</p>\n\n<hr>\n\n<h1>1. Experiments</h1>\n\n<p>SGD optimizer with an initial learning rate of 0.01 and  x0.1 at epoch 2,6,10,14,17 is used to train our model. We used FocalLoss here. It took 10~20 hours to train a single model with 2 P40 GPUs. PyTorch was used during the training process.  (During fine-tuning process learning rate of 0.001 and  x0.1 at epoch 3,6,10,13,18 is used to tune our model)</p>\n\n<h2>1.1. Baseline Single Model</h2>\n\n<p>a) IR-SE-50[1-3] (a 50 layer IR-SE model with the similar structure as IR-SE-152 ) was trained with Part-22 as validation set and the rest parts as training set. Its LB score was 0.51639 with 30 frames used per video. </p>\n\n<p>b) We changed the backbone to a deeper CNN structure IR-SE-152. This boosted our performance to 0.436 from 0.51639.</p>\n\n<p>c) Removing the noisy cases with mask filtering,  IR-SE-152’s loss was further reduced to 0.40.</p>\n\n<p>d) Random cropping was revised a bit to introduce more perturbations during the training process.  BoundingBox X 1.3 -&gt; resize to 291x291-&gt; random cropping to 224x224. Center cropping was then used during inference. Our single model’s performance was improved to 0.376.</p>\n\n<p>e) We found that RA-92[4] behaves better than IR-SE-152 on various tasks and it has a smaller model size(190MB vs. 223MB). We changed our single model to RA-92 and got a new high score of 0.362.  </p>\n\n<p>f) We tried various powerful backbones but got little progress. </p>\n\n<p>g) We realized that using Part-22 as validation may hardly reflect the true performance of our model when tested online. This is mainly because many models from Part-22 also appear in other folders.  There are many suggestions about splitting training/validation sets from comments of experienced Kagglers and we then decided to use 20% of the folders as the validation set to prevent overfitting. Part0-9 were selected as the offline validation set. Our single RA-92 model loss was further reduced to 0.342. (We also tried many other combinations, Part0-9 gave us the best Public LB performance)</p>\n\n<p>h) To solve the multi-face issues, face tracking(see details in 4.3) was introduced during the inference process. More faces per video(120 frames uniformly extracted from 30th~270th frames) were also used for prediction. Our single model score was then improved to 0.331.</p>\n\n<p>i) We further fine-tuned the 0.331 single model with additional training data from Part5~9 (Part5~9 were offline validation set earlier) and FF++[7].  Random erasing[8] was used when we fine-tuned the model. In the end, we further reduced the LB loss to  0.323 (The best single model and used for final submission). \n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2F2ec5ee874012a6e00423e43913e6e6f0%2FRE.jpg?generation=1592141564255571&amp;alt=media\" alt=\"\"></p>\n\n<p>j) We tried sphere and water augmentation from DALI with a probe of 0.2 in the training process. This led to about 0.01 loss reduction on certain models. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fc1250ded525f812f1190462ecf141334%2Fwater_effect.png?generation=1592141734353937&amp;alt=media\" alt=\"\"></p>\n\n<p>k) We implemented a multi-cropping strategy during inference. It showed some performance enhancement. This trick was not used in the final submissions due to the inference time restriction. Using additional frames from video is more cost-effective and can also improve the performance.</p>\n\n<h2>1.2.  Model Ensembling</h2>\n\n<h3>1.2.1 Diversity of network structures and training data</h3>\n\n<p>How to effectively aggregate features of various models is important in our final submission. We had attempted to combine a few models 1) different models trained with identical data, 2) same models trained with different data, and 3) different models trained with slightly different data. The third one outperformed the rest and was used in the final submission.</p>\n\n<p>a): We trained RA-92、IR-SE-152 and SE-RESNEXT101-32x8d [5] with Part0~49 (except Part22 which was used for validation). Merging these three models  led to  0.025 performance enhancement from the best single model.\n| Model train with Part0-49(excluding Part22) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|RA-92 <br> IR-SE-152  <br>SE-RENEXT101-32x8d| 0.36528<br>0.37623<br>0.38449|<br>0.34058|</p>\n\n<p>b): We trained three RA-92 models with different training data. Ensembling of these models led to negligible performance enhancement. \n| Validation data(the rest as train data) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|Part0~9 <br> Part20~29  <br>Part30~39| 0.33166<br>0.3504<br>0.34027|<br>0.33145|</p>\n\n<p>c): We used RA-92, IR-SE-152 and SE_RESNEXT101_32x8d (trained with different fake data sampled from Part10-49).  In this case, model and data diversities were all preserved. It reached our best LB score. These models were then ensembled and submitted. \n| Model train with Part0-49(excluding Part22) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|RA-92 <br> IR-SE-152  <br>SE_RESNEXT101_32x8dV2| 0.323<br>0.321<br>0.34 (estimated)|<br>0.294|</p>\n\n<ul>\n<li>Remark: The SE_RESNEXT101_32x8dV2 model was trained the last day, and we didn’t have time to submit it alone, but it behaves better offline than a 0.351 model.  </li>\n</ul>\n\n<h3>1.2.2 Model Ensembling Method</h3>\n\n<p>Our final submission(0.294@LB) aggregates results of three models (0.32~0.34@LB in Table 4-2 ) in a frame-by-frame manner. The score of face from a typical frame is determined as below:</p>\n\n<pre><code>        frame_face_score = average( 2 largest scores); if 2 scores &gt;= 0.5\n        frame_face_score = average( 2 smallest scores ); if 2 scores &lt; 0.5\n</code></pre>\n\n<p>The above strategy was adopted as it showed strong robustness compared with other model ensembling techniques.  The conclusion was drawn when we experimented on 4 models RA-92, IR-SE-152, EfficientNet-B7 and SE-ResNext101_32x8d with LB scores around 0.37~0.39.  Offline and online results of ensembled models are shown in Table 4-2. (These are intermediate results only as the models haven’t reached their best performance at the moment when we tried model ensembling)</p>\n\n<p>|  |CV-RAW|CV-C40|CV-RES4|LB|\n| --- | --- |--- |--- |--- |\n|  V1-Eff |0.139|  0.333|  0.576|  0.343|\n|V1-SE| 0.156 | 0.318 | 0.516 |  0.34058|\n|V2| 0.136 | 0.335 | 0.621 | 0.391 |\n|V3| 0.148 | 0.347 | 0.549 | Not submitted |</p>\n\n<ul>\n<li>Results of model ensembling on raw, c40 compression and ¼ resolution validation set of part0~9. </li>\n<li>V1-Eff: Final ensembling strategy with (RA-92, IR-SE-152, EfficientNet-B7)；</li>\n<li>V1-SE: Final ensembling strategy with (SE-RESNEXT101_32x8d, RA-92, IR-SE-152)；</li>\n<li>V2: max( scores); if 2 scores &gt;= 0.5;min( scores ); if 2 scores &lt; 0.5；</li>\n<li>V3: max(scores) if all scores &gt;0.5; min(scores) if all scores &lt;0.5; else average(scores)；</li>\n</ul>\n\n<p>We also tried other ensembling methods, such as simple average. But they were not as good as V1-SE. </p>\n\n<h2>1.3  Inference</h2>\n\n<p>Our inference process is shown as Figure below. It mainly consists of 5 modules.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fde6d63cdd7625ceb078548d81c5c4f3a%2Fflow_chart.png?generation=1592141837895073&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li><p>VideoCapture: Our framework uniformly extracts 120 frames from 30-th to 270-th frames of every video.</p></li>\n<li><p>FaceDetector: Face detection is performed via CenterNet with batchsize of 60 frames. Our framework selects at most 2 faces with top detection confidence if more than 2 faces are detected within a single frame. Faces with low confidence are abandoned. </p></li>\n<li><p>FaceTracker: Face tracking is performed via direct IOU matching between bounding boxes from current and previous frames. Tracking sequences with less than 20 faces are ignored (most likely false detection). </p></li>\n<li><p>FaceClassifier: The system then predicts classification scores of all selected faces with a batch size of 120. Results of three classifiers are ensembled here for each face.   </p></li>\n<li><p>PostProcess: The video score is calculated as a weighted average of the frame scores with weights being the detection confidences. If there are multiple tracking sequences in a video, the maximum score is used as the final prediction of the entire video.  </p></li>\n</ul>\n\n<h1>2.What did not help</h1>\n\n<ul>\n<li><p>SlowFast: We tried the slowfast[9] network with slow-branch (mainly a 3D-convolutional network).  Even though it can reach comparable offline performance as the image-based 2D convolution method (e.g. RA92), the best LB score is only 0.477. It seems SlowFast  is easily overfitted. We didn’t spend too much time in this direction. </p></li>\n<li><p>Crop152: We had modified a CROP152 model as below. This model requires lots of computing  resources but leads to marginal performance improvement. We didn’t use these models later. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fab002e71ced88c23c98394d5b807c784%2Fcrop152.png?generation=1592141868496743&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Audio Classifier: As audio may also be faked, we trained a CNN-based audio classifier with fake audio clips from the last five folders and all real audios. Melgram is used as input to the network. Validation accuracy on the offline testset is about 0.98. However, the score on LB is only 0.6929 (slightly better than 0.5 prediction). Considering that  there may be very limited fake-audios in the Public Set, we abandoned the audio classifier.  </p></li>\n<li><p>Multi-task Learning with Mask Info: Lots of studies have shown that multi-task learning can boost performance of the core task. We designed a multi-task network which can distinguish real/fake samples and categorize the size of swapped areas (e.g. 0: no swapping; 1: 5% changed, 2:5%~20% changed ...). We believed that the auxiliary face area information can guide the CNN model to learn features which are closely related to the manipulated face areas.  Though sounds very attractive, the results showed negligible difference from the original classifier model. </p></li>\n<li><p>Face alignment: We also tried to pre-process the face images with face alignment. However, several submissions showed the scores deteriorate slightly if alignment is used. Facial landmark detection with Dlib also requires extra computational resources and landmark points are inaccurate when the faces are with large poses and under challenging illumination conditions. Considering these facts, we didn’t use face alignment later. Fig. 5-b below shows failure cases of face alignment when the face is under large poses. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2F3153a69fa728b2f75a40d68db630684f%2Falignment_failure.png?generation=1592141895213430&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Video-level Data Augmentations: We also tried the following video-level data augmentation methods[6]: re-encode the video as 15FPS,  reduce the resolution of the video to 1/4 of its original size  and increase the compression ratio to generate low quality video. Results are not improved on the public LB when models are trained with these data. These generated samples were therefore not used. </p></li>\n<li><p>Hard Sample Mining: Hard sample mining is widely adopted in many detection and classification tasks. We randomly selected those failure cases from the left-over faces (previously only 400K faces are used from the 1.1million cleaned faces) and added these faces into the fine-tuning process.  However, the score worsens after fine-tuning. </p></li>\n<li><p>LSTM:  We used the convolutional features from the trained image-based deepfake classifier as input to LSTM and tuned the LSTM parameters with features of face sequences. The LSTM hardly outperformed its corresponding image-based classifier on the offline validation set. We didn’t submit it. </p></li>\n<li><p>Distance based method: Triplet and softmax loss were combined to perform metric learning. This, however, didn’t work for us as well.</p></li>\n<li><p>Postprocess with Score Clipping and Mapping：Score clipping slightly improved our single model at the early stage of the competition when our LB score was about 0.4. However, it worsened our performance for the ensembled models.  We also tried to map the score such that the threshold cutting point was at 0.5 on the offline validation set. However, the score mapping trick led to very inconsistent results on the Public LB. Therefore, it was not used in the final submission. </p></li>\n<li><p>Data Size and Sampling Methods: Down-sampling of faked faces is better than up-sampling of the real faces. We tried to train our models with more training samples. This worsened our models’ scores. </p></li>\n<li><p>Training with Clusters: Like most of the participants, we also realized the overlapping ID issue. Thus, we attempted to split the dataset by ID clusters and train our model with it. However, the submitted score showed no improvement.</p></li>\n<li><p>Training tricks: Frozen/warm up/labelSmoothing deteriorates our training loss for our models.\n300x300:  Larger input size leads to more overfitting of our model. They behaved better on CV but worse on LB.  We thus fixed our input size to 224.   </p></li>\n<li><p>Reducing Cropping Margin: Hoping the network can learn more features from the facial region, we reduced the cropping margin. Validation accuracy was improved by this method. However, the loss increased to  0.411 from 0.34 on the Public LB.</p></li>\n<li><p>Error Level Analysis: We followed ELA[https://fotoforensics.com/tutorial-ela.php] to conduct Error Level Analysis. However, no obvious characteristic of the swapped face was identified. We didn’t try this method further.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fc37c8285ed977b705aa6a64c4ca7ba96%2Ffailure_cases.png?generation=1592141930858740&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n\n<h1>3. Reference</h1>\n\n<p>[1] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.\n[2] Hu J, Shen L, Sun G. Squeeze-and-excitation networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7132-7141.\n[3] Han D, Kim J, Kim J. Deep pyramidal residual networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 5927-5935.\n[4] Wang F, Jiang M, Qian C, et al. Residual attention network for image classification[C]. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 3156-3164.\n[5] Xie S, Girshick R, Dollár P, et al. Aggregated residual transformations for deep neural networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 1492-1500.\n[6] Dolhansky B, Howes R, Pflaum B, et al. The Deepfake Detection Challenge (DFDC) Preview Dataset[J]. arXiv preprint arXiv:1910.08854, 2019.\n[7] Rossler A, Cozzolino D, Verdoliva L, et al. Faceforensics++: Learning to detect manipulated facial images[C]//Proceedings of the IEEE International Conference on Computer Vision. 2019: 1-11.\n[8] Zhong Z, Zheng L, Kang G, et al. Random erasing data augmentation[J]. arXiv preprint arXiv:1708.04896, 2017.\n[9] Feichtenhofer C, Fan H, Malik J, et al. Slowfast networks for video recognition[C]. In Proceedings of the IEEE International Conference on Computer Vision. 2019: 6202-6211.</p>\n\n<hr>\n\n<h3>If you feel the above solution is helpful and hope to discuss more offline with us, you may contact Meisong via zhengmeisong@gmail.com</h3>",
      "rawMarkdown": "Interested in this task and of course also greatly attracted by the prize, we (five newbie kagglers) formed QH-Team. We are friends and all have computer vision backgrounds. This competition is very different from what we have faced in many other computer vision tasks. Many tricks were tried but proven to be not working on the public leaderboard (not sure if they work on the private leaderboard). Indeed, most of our time was spent in handling overfitting of the offline validation set. We struggled a lot to maintain our competency (Rank 4th~47th) in the public leaderboard throughout the entire competition period. \n\nIt was a great surprise to us that we were ranked 2nd place (Private Leaderboard Scores:  0.42490&amp; 0.44431, Public Leaderboard Score: 0.29421) when the private leaderboard was released. We felt that the days and nights spent were well rewarded. Unfortunately, both results were removed due to the usage of faceforensic++ data even though this dataset can be accessed by all participants (we later confirmed this with dataset owners). We didn’t realize that this would be an issue during the competition. The rules are complicated to new Kagglers and we believe it would be much clearer to state that no external data is allowed for this competition at the very beginning. We are deeply disappointed by the final decision that both our scores are removed from the leaderboards. This simply erased all our efforts devoted to this competition. \n\nWe would still like to thank Facebook and Kaggle for organizing this event. This competition draws huge attention from both industry and academia, it will definitely boost the development of deepfake detection techniques to fight against the fake videos. Here, we would like to present our solution to the people who are interested in this task. Hope it will inspire you if you are working in the related field. \n\nMeisong Zheng, Chuanrui Hu, Walton Tao Wang, Wenju Huang, Shengtao Xiao\nQH Team\n\n--------------------------------------------------------------------------------------------------\n# 1. Experiments\nSGD optimizer with an initial learning rate of 0.01 and  x0.1 at epoch 2,6,10,14,17 is used to train our model. We used FocalLoss here. It took 10~20 hours to train a single model with 2 P40 GPUs. PyTorch was used during the training process.  (During fine-tuning process learning rate of 0.001 and  x0.1 at epoch 3,6,10,13,18 is used to tune our model)\n\n## 1.1. Baseline Single Model\n\na) IR-SE-50[1-3] (a 50 layer IR-SE model with the similar structure as IR-SE-152 ) was trained with Part-22 as validation set and the rest parts as training set. Its LB score was 0.51639 with 30 frames used per video. \n\n\nb) We changed the backbone to a deeper CNN structure IR-SE-152. This boosted our performance to 0.436 from 0.51639.\n\n\nc) Removing the noisy cases with mask filtering,  IR-SE-152’s loss was further reduced to 0.40.\n\n\nd) Random cropping was revised a bit to introduce more perturbations during the training process.  BoundingBox X 1.3 -&gt; resize to 291x291-&gt; random cropping to 224x224. Center cropping was then used during inference. Our single model’s performance was improved to 0.376.\n\n\ne) We found that RA-92[4] behaves better than IR-SE-152 on various tasks and it has a smaller model size(190MB vs. 223MB). We changed our single model to RA-92 and got a new high score of 0.362.  \n\n\nf) We tried various powerful backbones but got little progress. \n\n\ng) We realized that using Part-22 as validation may hardly reflect the true performance of our model when tested online. This is mainly because many models from Part-22 also appear in other folders.  There are many suggestions about splitting training/validation sets from comments of experienced Kagglers and we then decided to use 20% of the folders as the validation set to prevent overfitting. Part0-9 were selected as the offline validation set. Our single RA-92 model loss was further reduced to 0.342. (We also tried many other combinations, Part0-9 gave us the best Public LB performance)\n\n\nh) To solve the multi-face issues, face tracking(see details in 4.3) was introduced during the inference process. More faces per video(120 frames uniformly extracted from 30th~270th frames) were also used for prediction. Our single model score was then improved to 0.331.\n\n\ni) We further fine-tuned the 0.331 single model with additional training data from Part5~9 (Part5~9 were offline validation set earlier) and FF++[7].  Random erasing[8] was used when we fine-tuned the model. In the end, we further reduced the LB loss to  0.323 (The best single model and used for final submission). \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2F2ec5ee874012a6e00423e43913e6e6f0%2FRE.jpg?generation=1592141564255571&amp;alt=media)\n\n\nj) We tried sphere and water augmentation from DALI with a probe of 0.2 in the training process. This led to about 0.01 loss reduction on certain models.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fc1250ded525f812f1190462ecf141334%2Fwater_effect.png?generation=1592141734353937&amp;alt=media)\n\n\nk) We implemented a multi-cropping strategy during inference. It showed some performance enhancement. This trick was not used in the final submissions due to the inference time restriction. Using additional frames from video is more cost-effective and can also improve the performance.\n\n\n##1.2.  Model Ensembling\n\n###1.2.1 Diversity of network structures and training data\n\nHow to effectively aggregate features of various models is important in our final submission. We had attempted to combine a few models 1) different models trained with identical data, 2) same models trained with different data, and 3) different models trained with slightly different data. The third one outperformed the rest and was used in the final submission.\n\n\na): We trained RA-92、IR-SE-152 and SE-RESNEXT101-32x8d [5] with Part0~49 (except Part22 which was used for validation). Merging these three models  led to  0.025 performance enhancement from the best single model.\n| Model train with Part0-49(excluding Part22) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|RA-92 <br> IR-SE-152  <br>SE-RENEXT101-32x8d| 0.36528<br>0.37623<br>0.38449|<br>0.34058|\n\nb): We trained three RA-92 models with different training data. Ensembling of these models led to negligible performance enhancement. \n| Validation data(the rest as train data) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|Part0~9 <br> Part20~29  <br>Part30~39| 0.33166<br>0.3504<br>0.34027|<br>0.33145|\n\nc): We used RA-92, IR-SE-152 and SE_RESNEXT101_32x8d (trained with different fake data sampled from Part10-49).  In this case, model and data diversities were all preserved. It reached our best LB score. These models were then ensembled and submitted. \n| Model train with Part0-49(excluding Part22) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|RA-92 <br> IR-SE-152  <br>SE_RESNEXT101_32x8dV2| 0.323<br>0.321<br>0.34 (estimated)|<br>0.294|\n\n- Remark: The SE_RESNEXT101_32x8dV2 model was trained the last day, and we didn’t have time to submit it alone, but it behaves better offline than a 0.351 model.  \n\n###1.2.2 Model Ensembling Method\n\nOur final submission(0.294@LB) aggregates results of three models (0.32~0.34@LB in Table 4-2 ) in a frame-by-frame manner. The score of face from a typical frame is determined as below:\n\n            frame_face_score = average( 2 largest scores); if 2 scores &gt;= 0.5\n            frame_face_score = average( 2 smallest scores ); if 2 scores &lt; 0.5\n\nThe above strategy was adopted as it showed strong robustness compared with other model ensembling techniques.  The conclusion was drawn when we experimented on 4 models RA-92, IR-SE-152, EfficientNet-B7 and SE-ResNext101_32x8d with LB scores around 0.37~0.39.  Offline and online results of ensembled models are shown in Table 4-2. (These are intermediate results only as the models haven’t reached their best performance at the moment when we tried model ensembling)\n\n|  |CV-RAW|CV-C40|CV-RES4|LB|\n| --- | --- |--- |--- |--- |\n|  V1-Eff |0.139|  0.333|  0.576|  0.343|\n|V1-SE| 0.156 | 0.318 | 0.516 |  0.34058|\n|V2| 0.136 | 0.335 | 0.621 | 0.391 |\n|V3| 0.148 | 0.347 | 0.549 | Not submitted |\n\n- Results of model ensembling on raw, c40 compression and ¼ resolution validation set of part0~9. \n- V1-Eff: Final ensembling strategy with (RA-92, IR-SE-152, EfficientNet-B7)；\n- V1-SE: Final ensembling strategy with (SE-RESNEXT101_32x8d, RA-92, IR-SE-152)；\n- V2: max( scores); if 2 scores &gt;= 0.5;min( scores ); if 2 scores &lt; 0.5；\n- V3: max(scores) if all scores &gt;0.5; min(scores) if all scores &lt;0.5; else average(scores)；\n\nWe also tried other ensembling methods, such as simple average. But they were not as good as V1-SE. \n\n##1.3  Inference\n\nOur inference process is shown as Figure below. It mainly consists of 5 modules.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fde6d63cdd7625ceb078548d81c5c4f3a%2Fflow_chart.png?generation=1592141837895073&amp;alt=media)\n\n- VideoCapture: Our framework uniformly extracts 120 frames from 30-th to 270-th frames of every video.\n\n\n- FaceDetector: Face detection is performed via CenterNet with batchsize of 60 frames. Our framework selects at most 2 faces with top detection confidence if more than 2 faces are detected within a single frame. Faces with low confidence are abandoned. \n\n\n- FaceTracker: Face tracking is performed via direct IOU matching between bounding boxes from current and previous frames. Tracking sequences with less than 20 faces are ignored (most likely false detection). \n\n\n- FaceClassifier: The system then predicts classification scores of all selected faces with a batch size of 120. Results of three classifiers are ensembled here for each face.   \n\n\n- PostProcess: The video score is calculated as a weighted average of the frame scores with weights being the detection confidences. If there are multiple tracking sequences in a video, the maximum score is used as the final prediction of the entire video.  \n\n\n#2.What did not help\n\n- SlowFast: We tried the slowfast[9] network with slow-branch (mainly a 3D-convolutional network).  Even though it can reach comparable offline performance as the image-based 2D convolution method (e.g. RA92), the best LB score is only 0.477. It seems SlowFast  is easily overfitted. We didn’t spend too much time in this direction. \n\n- Crop152: We had modified a CROP152 model as below. This model requires lots of computing  resources but leads to marginal performance improvement. We didn’t use these models later. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fab002e71ced88c23c98394d5b807c784%2Fcrop152.png?generation=1592141868496743&amp;alt=media)\n\n\n- Audio Classifier: As audio may also be faked, we trained a CNN-based audio classifier with fake audio clips from the last five folders and all real audios. Melgram is used as input to the network. Validation accuracy on the offline testset is about 0.98. However, the score on LB is only 0.6929 (slightly better than 0.5 prediction). Considering that  there may be very limited fake-audios in the Public Set, we abandoned the audio classifier.  \n\n\n- Multi-task Learning with Mask Info: Lots of studies have shown that multi-task learning can boost performance of the core task. We designed a multi-task network which can distinguish real/fake samples and categorize the size of swapped areas (e.g. 0: no swapping; 1: 5% changed, 2:5%~20% changed ...). We believed that the auxiliary face area information can guide the CNN model to learn features which are closely related to the manipulated face areas.  Though sounds very attractive, the results showed negligible difference from the original classifier model. \n\n\n- Face alignment: We also tried to pre-process the face images with face alignment. However, several submissions showed the scores deteriorate slightly if alignment is used. Facial landmark detection with Dlib also requires extra computational resources and landmark points are inaccurate when the faces are with large poses and under challenging illumination conditions. Considering these facts, we didn’t use face alignment later. Fig. 5-b below shows failure cases of face alignment when the face is under large poses. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2F3153a69fa728b2f75a40d68db630684f%2Falignment_failure.png?generation=1592141895213430&amp;alt=media)\n\n\n- Video-level Data Augmentations: We also tried the following video-level data augmentation methods[6]: re-encode the video as 15FPS,  reduce the resolution of the video to 1/4 of its original size  and increase the compression ratio to generate low quality video. Results are not improved on the public LB when models are trained with these data. These generated samples were therefore not used. \n\n\n- Hard Sample Mining: Hard sample mining is widely adopted in many detection and classification tasks. We randomly selected those failure cases from the left-over faces (previously only 400K faces are used from the 1.1million cleaned faces) and added these faces into the fine-tuning process.  However, the score worsens after fine-tuning. \n\n\n- LSTM:  We used the convolutional features from the trained image-based deepfake classifier as input to LSTM and tuned the LSTM parameters with features of face sequences. The LSTM hardly outperformed its corresponding image-based classifier on the offline validation set. We didn’t submit it. \n\n\n- Distance based method: Triplet and softmax loss were combined to perform metric learning. This, however, didn’t work for us as well.\n\n\n- Postprocess with Score Clipping and Mapping：Score clipping slightly improved our single model at the early stage of the competition when our LB score was about 0.4. However, it worsened our performance for the ensembled models.  We also tried to map the score such that the threshold cutting point was at 0.5 on the offline validation set. However, the score mapping trick led to very inconsistent results on the Public LB. Therefore, it was not used in the final submission. \n\n\n- Data Size and Sampling Methods: Down-sampling of faked faces is better than up-sampling of the real faces. We tried to train our models with more training samples. This worsened our models’ scores. \n\n\n- Training with Clusters: Like most of the participants, we also realized the overlapping ID issue. Thus, we attempted to split the dataset by ID clusters and train our model with it. However, the submitted score showed no improvement.\n\n\n- Training tricks: Frozen/warm up/labelSmoothing deteriorates our training loss for our models.\n300x300:  Larger input size leads to more overfitting of our model. They behaved better on CV but worse on LB.  We thus fixed our input size to 224.   \n\n- Reducing Cropping Margin: Hoping the network can learn more features from the facial region, we reduced the cropping margin. Validation accuracy was improved by this method. However, the loss increased to  0.411 from 0.34 on the Public LB.\n\n\n- Error Level Analysis: We followed ELA[https://fotoforensics.com/tutorial-ela.php] to conduct Error Level Analysis. However, no obvious characteristic of the swapped face was identified. We didn’t try this method further.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fc37c8285ed977b705aa6a64c4ca7ba96%2Ffailure_cases.png?generation=1592141930858740&amp;alt=media)\n\n\n\n#3. Reference\n[1] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.\n[2] Hu J, Shen L, Sun G. Squeeze-and-excitation networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7132-7141.\n[3] Han D, Kim J, Kim J. Deep pyramidal residual networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 5927-5935.\n[4] Wang F, Jiang M, Qian C, et al. Residual attention network for image classification[C]. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 3156-3164.\n[5] Xie S, Girshick R, Dollár P, et al. Aggregated residual transformations for deep neural networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 1492-1500.\n[6] Dolhansky B, Howes R, Pflaum B, et al. The Deepfake Detection Challenge (DFDC) Preview Dataset[J]. arXiv preprint arXiv:1910.08854, 2019.\n[7] Rossler A, Cozzolino D, Verdoliva L, et al. Faceforensics++: Learning to detect manipulated facial images[C]//Proceedings of the IEEE International Conference on Computer Vision. 2019: 1-11.\n[8] Zhong Z, Zheng L, Kang G, et al. Random erasing data augmentation[J]. arXiv preprint arXiv:1708.04896, 2017.\n[9] Feichtenhofer C, Fan H, Malik J, et al. Slowfast networks for video recognition[C]. In Proceedings of the IEEE International Conference on Computer Vision. 2019: 6202-6211.\n\n------------------------------------------------------------------------------------------------\n###If you feel the above solution is helpful and hope to discuss more offline with us, you may contact Meisong via zhengmeisong@gmail.com",
      "votes": null
    },
    {
      "id": "885890",
      "postDate": "06/14/2020 14:42:16",
      "content": "<blockquote>\n  <p>both results were removed due to the usage of faceforensic++ data even though this dataset can be accessed by all participants (we later confirmed this with dataset owners).</p>\n</blockquote>\n\n<p>I think it was made pretty clear throughout the competition that this dataset was not allowed because it required participants to get permission to use it, which means in theory that some people could have been denied permission (even if that never happened).</p>",
      "rawMarkdown": "&gt; both results were removed due to the usage of faceforensic++ data even though this dataset can be accessed by all participants (we later confirmed this with dataset owners).\n\nI think it was made pretty clear throughout the competition that this dataset was not allowed because it required participants to get permission to use it, which means in theory that some people could have been denied permission (even if that never happened).",
      "votes": null
    },
    {
      "id": "885955",
      "postDate": "06/14/2020 15:42:39",
      "content": "<p>We thought it can be used as many participants claimed this in the official data disclosure thread and there was no restrictions for all participants to use the data.  Only when we were told disqualified, we managed to dig out that  FF++ was banned by Kaggle in a discussion thread.  Anyway, hope this will be a lesson for all new Kagglers. If you are not 100% sure if the data is allowed, just don't use it. Or simply don't use any external data.  </p>",
      "rawMarkdown": "We thought it can be used as many participants claimed this in the official data disclosure thread and there was no restrictions for all participants to use the data.  Only when we were told disqualified, we managed to dig out that  FF++ was banned by Kaggle in a discussion thread.  Anyway, hope this will be a lesson for all new Kagglers. If you are not 100% sure if the data is allowed, just don't use it. Or simply don't use any external data.",
      "votes": null
    },
    {
      "id": "885989",
      "postDate": "06/14/2020 16:11:40",
      "content": "<p>Thanks for the detailed summary, and sorry to see the DQ of your time.  Your solution looks SOLID, and I can see the effort and time you have spent here to provide the quality write-up too. </p>\n\n<p>I can completely understand that the competition rules can be very confusing, and more often than not the one bit of useful information is buried deep inside long threads.   </p>\n\n<p>I am not saying this is the case of this specific team, but for new kagglers whose native language is not English,  trying to comprehend what is said by kaggle admin regarding the competition rules can be very difficult.  Especially when those rules also evidently can be interpreted differently between different competitions.</p>",
      "rawMarkdown": "Thanks for the detailed summary, and sorry to see the DQ of your time.  Your solution looks SOLID, and I can see the effort and time you have spent here to provide the quality write-up too. \n\nI can completely understand that the competition rules can be very confusing, and more often than not the one bit of useful information is buried deep inside long threads.   \n\nI am not saying this is the case of this specific team, but for new kagglers whose native language is not English,  trying to comprehend what is said by kaggle admin regarding the competition rules can be very difficult.  Especially when those rules also evidently can be interpreted differently between different competitions.",
      "votes": null
    },
    {
      "id": "886359",
      "postDate": "06/15/2020 01:32:14",
      "content": "<p>FaceForensics consists of Youtube videos with given links. Your case may be the same as the 1st removed team? </p>",
      "rawMarkdown": "FaceForensics consists of Youtube videos with given links. Your case may be the same as the 1st removed team?",
      "votes": null
    },
    {
      "id": "886405",
      "postDate": "06/15/2020 02:57:58",
      "content": "<p>Thanks for your understanding, Yifan. The rules are complicated. It's extremely unfriendly for new kaggler. Hopefully, Kaggle could simplify its DATA USE POLICY. just like  <code>no external data is allowed for this competition at the very beginning</code></p>",
      "rawMarkdown": "Thanks for your understanding, Yifan. The rules are complicated. It's extremely unfriendly for new kaggler. Hopefully, Kaggle could simplify its DATA USE POLICY. just like  ` no external data is allowed for this competition at the very beginning`",
      "votes": null
    },
    {
      "id": "886414",
      "postDate": "06/15/2020 03:08:30",
      "content": "<p>we are not very sure about this. </p>",
      "rawMarkdown": "we are not very sure about this.",
      "votes": null
    },
    {
      "id": "886513",
      "postDate": "06/15/2020 05:08:52",
      "content": "<p>Congrats! Very impressive result especially for your first kaggle competition!</p>",
      "rawMarkdown": "Congrats! Very impressive result especially for your first kaggle competition!",
      "votes": null
    },
    {
      "id": "887412",
      "postDate": "06/15/2020 16:59:03",
      "content": "<p>Impressive! and inspiring for a novice such me!</p>",
      "rawMarkdown": "Impressive! and inspiring for a novice such me!",
      "votes": null
    },
    {
      "id": "887829",
      "postDate": "06/15/2020 23:07:36",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "887853",
      "postDate": "06/15/2020 23:48:21",
      "content": "<p>Congrats! </p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "888175",
      "postDate": "06/16/2020 07:02:49",
      "content": "<p>What is ra-92? Google brings some strange stuff... </p>",
      "rawMarkdown": "What is ra-92? Google brings some strange stuff...",
      "votes": null
    },
    {
      "id": "888243",
      "postDate": "06/16/2020 08:13:40",
      "content": "<p>pls read the this paper [4] Wang F, Jiang M, Qian C, et al. Residual attention network for image classification[C]. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 3156-3164.</p>",
      "rawMarkdown": "pls read the this paper [4] Wang F, Jiang M, Qian C, et al. Residual attention network for image classification[C]. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 3156-3164.",
      "votes": null
    },
    {
      "id": "888292",
      "postDate": "06/16/2020 09:01:53",
      "content": "<p>Great God, such a thorough explanation. It's quite sad that you applied so much effort and didn't get to reap your rewards. Still, congratulations! It's surprising to see novice Kagglers come up with such a detailed solution, that too with reference to research papers. You guys should be proud of yourselves. </p>",
      "rawMarkdown": "Great God, such a thorough explanation. It's quite sad that you applied so much effort and didn't get to reap your rewards. Still, congratulations! It's surprising to see novice Kagglers come up with such a detailed solution, that too with reference to research papers. You guys should be proud of yourselves.",
      "votes": null
    },
    {
      "id": "888957",
      "postDate": "06/16/2020 17:14:29",
      "content": "<p>Congrats!!</p>",
      "rawMarkdown": "Congrats!!",
      "votes": null
    },
    {
      "id": "889644",
      "postDate": "06/17/2020 03:46:02",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "890185",
      "postDate": "06/17/2020 11:18:47",
      "content": "<p>thanks for sharing. really enjoyed reading your solution with the detailed log for each step/experiment.</p>",
      "rawMarkdown": "thanks for sharing. really enjoyed reading your solution with the detailed log for each step/experiment.",
      "votes": null
    },
    {
      "id": "957165",
      "postDate": "08/04/2020 05:38:01",
      "content": "<p>very good write-up, sorry for your removal.</p>",
      "rawMarkdown": "very good write-up, sorry for your removal.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 885890,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "06/14/2020 14:42:16",
      "content": "<blockquote>\n  <p>both results were removed due to the usage of faceforensic++ data even though this dataset can be accessed by all participants (we later confirmed this with dataset owners).</p>\n</blockquote>\n\n<p>I think it was made pretty clear throughout the competition that this dataset was not allowed because it required participants to get permission to use it, which means in theory that some people could have been denied permission (even if that never happened).</p>",
      "votes": null,
      "replies": [
        {
          "id": 885955,
          "author_name": "xiaoshengtao",
          "author_url": "",
          "post_date": "06/14/2020 15:42:39",
          "content": "<p>We thought it can be used as many participants claimed this in the official data disclosure thread and there was no restrictions for all participants to use the data.  Only when we were told disqualified, we managed to dig out that  FF++ was banned by Kaggle in a discussion thread.  Anyway, hope this will be a lesson for all new Kagglers. If you are not 100% sure if the data is allowed, just don't use it. Or simply don't use any external data.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 885989,
      "author_name": "yifanxie",
      "author_url": "",
      "post_date": "06/14/2020 16:11:40",
      "content": "<p>Thanks for the detailed summary, and sorry to see the DQ of your time.  Your solution looks SOLID, and I can see the effort and time you have spent here to provide the quality write-up too. </p>\n\n<p>I can completely understand that the competition rules can be very confusing, and more often than not the one bit of useful information is buried deep inside long threads.   </p>\n\n<p>I am not saying this is the case of this specific team, but for new kagglers whose native language is not English,  trying to comprehend what is said by kaggle admin regarding the competition rules can be very difficult.  Especially when those rules also evidently can be interpreted differently between different competitions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 886405,
          "author_name": "wangtao360",
          "author_url": "",
          "post_date": "06/15/2020 02:57:58",
          "content": "<p>Thanks for your understanding, Yifan. The rules are complicated. It's extremely unfriendly for new kaggler. Hopefully, Kaggle could simplify its DATA USE POLICY. just like  <code>no external data is allowed for this competition at the very beginning</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 886359,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "06/15/2020 01:32:14",
      "content": "<p>FaceForensics consists of Youtube videos with given links. Your case may be the same as the 1st removed team? </p>",
      "votes": null,
      "replies": [
        {
          "id": 886414,
          "author_name": "wangtao360",
          "author_url": "",
          "post_date": "06/15/2020 03:08:30",
          "content": "<p>we are not very sure about this. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 886513,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "06/15/2020 05:08:52",
      "content": "<p>Congrats! Very impressive result especially for your first kaggle competition!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 887412,
      "author_name": "olivierfoudras",
      "author_url": "",
      "post_date": "06/15/2020 16:59:03",
      "content": "<p>Impressive! and inspiring for a novice such me!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 887829,
      "author_name": "rashidulhasanhridoy",
      "author_url": "",
      "post_date": "06/15/2020 23:07:36",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 887853,
      "author_name": "rashidulhasanhridoy",
      "author_url": "",
      "post_date": "06/15/2020 23:48:21",
      "content": "<p>Congrats! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 888175,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "06/16/2020 07:02:49",
      "content": "<p>What is ra-92? Google brings some strange stuff... </p>",
      "votes": null,
      "replies": [
        {
          "id": 888243,
          "author_name": "wangtao360",
          "author_url": "",
          "post_date": "06/16/2020 08:13:40",
          "content": "<p>pls read the this paper [4] Wang F, Jiang M, Qian C, et al. Residual attention network for image classification[C]. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 3156-3164.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 889644,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "06/17/2020 03:46:02",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 888292,
      "author_name": "danoozy44",
      "author_url": "",
      "post_date": "06/16/2020 09:01:53",
      "content": "<p>Great God, such a thorough explanation. It's quite sad that you applied so much effort and didn't get to reap your rewards. Still, congratulations! It's surprising to see novice Kagglers come up with such a detailed solution, that too with reference to research papers. You guys should be proud of yourselves. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 888957,
      "author_name": "ramonmf",
      "author_url": "",
      "post_date": "06/16/2020 17:14:29",
      "content": "<p>Congrats!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 890185,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "06/17/2020 11:18:47",
      "content": "<p>thanks for sharing. really enjoyed reading your solution with the detailed log for each step/experiment.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 957165,
      "author_name": "fuzhuolin",
      "author_url": "",
      "post_date": "08/04/2020 05:38:01",
      "content": "<p>very good write-up, sorry for your removal.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "885812": "Interested in this task and of course also greatly attracted by the prize, we (five newbie kagglers) formed QH-Team. We are friends and all have computer vision backgrounds. This competition is very different from what we have faced in many other computer vision tasks. Many tricks were tried but proven to be not working on the public leaderboard (not sure if they work on the private leaderboard). Indeed, most of our time was spent in handling overfitting of the offline validation set. We struggled a lot to maintain our competency (Rank 4th~47th) in the public leaderboard throughout the entire competition period. \n\nIt was a great surprise to us that we were ranked 2nd place (Private Leaderboard Scores:  0.42490&amp; 0.44431, Public Leaderboard Score: 0.29421) when the private leaderboard was released. We felt that the days and nights spent were well rewarded. Unfortunately, both results were removed due to the usage of faceforensic++ data even though this dataset can be accessed by all participants (we later confirmed this with dataset owners). We didn’t realize that this would be an issue during the competition. The rules are complicated to new Kagglers and we believe it would be much clearer to state that no external data is allowed for this competition at the very beginning. We are deeply disappointed by the final decision that both our scores are removed from the leaderboards. This simply erased all our efforts devoted to this competition. \n\nWe would still like to thank Facebook and Kaggle for organizing this event. This competition draws huge attention from both industry and academia, it will definitely boost the development of deepfake detection techniques to fight against the fake videos. Here, we would like to present our solution to the people who are interested in this task. Hope it will inspire you if you are working in the related field. \n\nMeisong Zheng, Chuanrui Hu, Walton Tao Wang, Wenju Huang, Shengtao Xiao\nQH Team\n\n--------------------------------------------------------------------------------------------------\n# 1. Experiments\nSGD optimizer with an initial learning rate of 0.01 and  x0.1 at epoch 2,6,10,14,17 is used to train our model. We used FocalLoss here. It took 10~20 hours to train a single model with 2 P40 GPUs. PyTorch was used during the training process.  (During fine-tuning process learning rate of 0.001 and  x0.1 at epoch 3,6,10,13,18 is used to tune our model)\n\n## 1.1. Baseline Single Model\n\na) IR-SE-50[1-3] (a 50 layer IR-SE model with the similar structure as IR-SE-152 ) was trained with Part-22 as validation set and the rest parts as training set. Its LB score was 0.51639 with 30 frames used per video. \n\n\nb) We changed the backbone to a deeper CNN structure IR-SE-152. This boosted our performance to 0.436 from 0.51639.\n\n\nc) Removing the noisy cases with mask filtering,  IR-SE-152’s loss was further reduced to 0.40.\n\n\nd) Random cropping was revised a bit to introduce more perturbations during the training process.  BoundingBox X 1.3 -&gt; resize to 291x291-&gt; random cropping to 224x224. Center cropping was then used during inference. Our single model’s performance was improved to 0.376.\n\n\ne) We found that RA-92[4] behaves better than IR-SE-152 on various tasks and it has a smaller model size(190MB vs. 223MB). We changed our single model to RA-92 and got a new high score of 0.362.  \n\n\nf) We tried various powerful backbones but got little progress. \n\n\ng) We realized that using Part-22 as validation may hardly reflect the true performance of our model when tested online. This is mainly because many models from Part-22 also appear in other folders.  There are many suggestions about splitting training/validation sets from comments of experienced Kagglers and we then decided to use 20% of the folders as the validation set to prevent overfitting. Part0-9 were selected as the offline validation set. Our single RA-92 model loss was further reduced to 0.342. (We also tried many other combinations, Part0-9 gave us the best Public LB performance)\n\n\nh) To solve the multi-face issues, face tracking(see details in 4.3) was introduced during the inference process. More faces per video(120 frames uniformly extracted from 30th~270th frames) were also used for prediction. Our single model score was then improved to 0.331.\n\n\ni) We further fine-tuned the 0.331 single model with additional training data from Part5~9 (Part5~9 were offline validation set earlier) and FF++[7].  Random erasing[8] was used when we fine-tuned the model. In the end, we further reduced the LB loss to  0.323 (The best single model and used for final submission). \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2F2ec5ee874012a6e00423e43913e6e6f0%2FRE.jpg?generation=1592141564255571&amp;alt=media)\n\n\nj) We tried sphere and water augmentation from DALI with a probe of 0.2 in the training process. This led to about 0.01 loss reduction on certain models.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fc1250ded525f812f1190462ecf141334%2Fwater_effect.png?generation=1592141734353937&amp;alt=media)\n\n\nk) We implemented a multi-cropping strategy during inference. It showed some performance enhancement. This trick was not used in the final submissions due to the inference time restriction. Using additional frames from video is more cost-effective and can also improve the performance.\n\n\n##1.2.  Model Ensembling\n\n###1.2.1 Diversity of network structures and training data\n\nHow to effectively aggregate features of various models is important in our final submission. We had attempted to combine a few models 1) different models trained with identical data, 2) same models trained with different data, and 3) different models trained with slightly different data. The third one outperformed the rest and was used in the final submission.\n\n\na): We trained RA-92、IR-SE-152 and SE-RESNEXT101-32x8d [5] with Part0~49 (except Part22 which was used for validation). Merging these three models  led to  0.025 performance enhancement from the best single model.\n| Model train with Part0-49(excluding Part22) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|RA-92 <br> IR-SE-152  <br>SE-RENEXT101-32x8d| 0.36528<br>0.37623<br>0.38449|<br>0.34058|\n\nb): We trained three RA-92 models with different training data. Ensembling of these models led to negligible performance enhancement. \n| Validation data(the rest as train data) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|Part0~9 <br> Part20~29  <br>Part30~39| 0.33166<br>0.3504<br>0.34027|<br>0.33145|\n\nc): We used RA-92, IR-SE-152 and SE_RESNEXT101_32x8d (trained with different fake data sampled from Part10-49).  In this case, model and data diversities were all preserved. It reached our best LB score. These models were then ensembled and submitted. \n| Model train with Part0-49(excluding Part22) | Public LB score | Ensembling Result |\n| --- | --- | --- |\n|RA-92 <br> IR-SE-152  <br>SE_RESNEXT101_32x8dV2| 0.323<br>0.321<br>0.34 (estimated)|<br>0.294|\n\n- Remark: The SE_RESNEXT101_32x8dV2 model was trained the last day, and we didn’t have time to submit it alone, but it behaves better offline than a 0.351 model.  \n\n###1.2.2 Model Ensembling Method\n\nOur final submission(0.294@LB) aggregates results of three models (0.32~0.34@LB in Table 4-2 ) in a frame-by-frame manner. The score of face from a typical frame is determined as below:\n\n            frame_face_score = average( 2 largest scores); if 2 scores &gt;= 0.5\n            frame_face_score = average( 2 smallest scores ); if 2 scores &lt; 0.5\n\nThe above strategy was adopted as it showed strong robustness compared with other model ensembling techniques.  The conclusion was drawn when we experimented on 4 models RA-92, IR-SE-152, EfficientNet-B7 and SE-ResNext101_32x8d with LB scores around 0.37~0.39.  Offline and online results of ensembled models are shown in Table 4-2. (These are intermediate results only as the models haven’t reached their best performance at the moment when we tried model ensembling)\n\n|  |CV-RAW|CV-C40|CV-RES4|LB|\n| --- | --- |--- |--- |--- |\n|  V1-Eff |0.139|  0.333|  0.576|  0.343|\n|V1-SE| 0.156 | 0.318 | 0.516 |  0.34058|\n|V2| 0.136 | 0.335 | 0.621 | 0.391 |\n|V3| 0.148 | 0.347 | 0.549 | Not submitted |\n\n- Results of model ensembling on raw, c40 compression and ¼ resolution validation set of part0~9. \n- V1-Eff: Final ensembling strategy with (RA-92, IR-SE-152, EfficientNet-B7)；\n- V1-SE: Final ensembling strategy with (SE-RESNEXT101_32x8d, RA-92, IR-SE-152)；\n- V2: max( scores); if 2 scores &gt;= 0.5;min( scores ); if 2 scores &lt; 0.5；\n- V3: max(scores) if all scores &gt;0.5; min(scores) if all scores &lt;0.5; else average(scores)；\n\nWe also tried other ensembling methods, such as simple average. But they were not as good as V1-SE. \n\n##1.3  Inference\n\nOur inference process is shown as Figure below. It mainly consists of 5 modules.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fde6d63cdd7625ceb078548d81c5c4f3a%2Fflow_chart.png?generation=1592141837895073&amp;alt=media)\n\n- VideoCapture: Our framework uniformly extracts 120 frames from 30-th to 270-th frames of every video.\n\n\n- FaceDetector: Face detection is performed via CenterNet with batchsize of 60 frames. Our framework selects at most 2 faces with top detection confidence if more than 2 faces are detected within a single frame. Faces with low confidence are abandoned. \n\n\n- FaceTracker: Face tracking is performed via direct IOU matching between bounding boxes from current and previous frames. Tracking sequences with less than 20 faces are ignored (most likely false detection). \n\n\n- FaceClassifier: The system then predicts classification scores of all selected faces with a batch size of 120. Results of three classifiers are ensembled here for each face.   \n\n\n- PostProcess: The video score is calculated as a weighted average of the frame scores with weights being the detection confidences. If there are multiple tracking sequences in a video, the maximum score is used as the final prediction of the entire video.  \n\n\n#2.What did not help\n\n- SlowFast: We tried the slowfast[9] network with slow-branch (mainly a 3D-convolutional network).  Even though it can reach comparable offline performance as the image-based 2D convolution method (e.g. RA92), the best LB score is only 0.477. It seems SlowFast  is easily overfitted. We didn’t spend too much time in this direction. \n\n- Crop152: We had modified a CROP152 model as below. This model requires lots of computing  resources but leads to marginal performance improvement. We didn’t use these models later. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fab002e71ced88c23c98394d5b807c784%2Fcrop152.png?generation=1592141868496743&amp;alt=media)\n\n\n- Audio Classifier: As audio may also be faked, we trained a CNN-based audio classifier with fake audio clips from the last five folders and all real audios. Melgram is used as input to the network. Validation accuracy on the offline testset is about 0.98. However, the score on LB is only 0.6929 (slightly better than 0.5 prediction). Considering that  there may be very limited fake-audios in the Public Set, we abandoned the audio classifier.  \n\n\n- Multi-task Learning with Mask Info: Lots of studies have shown that multi-task learning can boost performance of the core task. We designed a multi-task network which can distinguish real/fake samples and categorize the size of swapped areas (e.g. 0: no swapping; 1: 5% changed, 2:5%~20% changed ...). We believed that the auxiliary face area information can guide the CNN model to learn features which are closely related to the manipulated face areas.  Though sounds very attractive, the results showed negligible difference from the original classifier model. \n\n\n- Face alignment: We also tried to pre-process the face images with face alignment. However, several submissions showed the scores deteriorate slightly if alignment is used. Facial landmark detection with Dlib also requires extra computational resources and landmark points are inaccurate when the faces are with large poses and under challenging illumination conditions. Considering these facts, we didn’t use face alignment later. Fig. 5-b below shows failure cases of face alignment when the face is under large poses. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2F3153a69fa728b2f75a40d68db630684f%2Falignment_failure.png?generation=1592141895213430&amp;alt=media)\n\n\n- Video-level Data Augmentations: We also tried the following video-level data augmentation methods[6]: re-encode the video as 15FPS,  reduce the resolution of the video to 1/4 of its original size  and increase the compression ratio to generate low quality video. Results are not improved on the public LB when models are trained with these data. These generated samples were therefore not used. \n\n\n- Hard Sample Mining: Hard sample mining is widely adopted in many detection and classification tasks. We randomly selected those failure cases from the left-over faces (previously only 400K faces are used from the 1.1million cleaned faces) and added these faces into the fine-tuning process.  However, the score worsens after fine-tuning. \n\n\n- LSTM:  We used the convolutional features from the trained image-based deepfake classifier as input to LSTM and tuned the LSTM parameters with features of face sequences. The LSTM hardly outperformed its corresponding image-based classifier on the offline validation set. We didn’t submit it. \n\n\n- Distance based method: Triplet and softmax loss were combined to perform metric learning. This, however, didn’t work for us as well.\n\n\n- Postprocess with Score Clipping and Mapping：Score clipping slightly improved our single model at the early stage of the competition when our LB score was about 0.4. However, it worsened our performance for the ensembled models.  We also tried to map the score such that the threshold cutting point was at 0.5 on the offline validation set. However, the score mapping trick led to very inconsistent results on the Public LB. Therefore, it was not used in the final submission. \n\n\n- Data Size and Sampling Methods: Down-sampling of faked faces is better than up-sampling of the real faces. We tried to train our models with more training samples. This worsened our models’ scores. \n\n\n- Training with Clusters: Like most of the participants, we also realized the overlapping ID issue. Thus, we attempted to split the dataset by ID clusters and train our model with it. However, the submitted score showed no improvement.\n\n\n- Training tricks: Frozen/warm up/labelSmoothing deteriorates our training loss for our models.\n300x300:  Larger input size leads to more overfitting of our model. They behaved better on CV but worse on LB.  We thus fixed our input size to 224.   \n\n- Reducing Cropping Margin: Hoping the network can learn more features from the facial region, we reduced the cropping margin. Validation accuracy was improved by this method. However, the loss increased to  0.411 from 0.34 on the Public LB.\n\n\n- Error Level Analysis: We followed ELA[https://fotoforensics.com/tutorial-ela.php] to conduct Error Level Analysis. However, no obvious characteristic of the swapped face was identified. We didn’t try this method further.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4193778%2Fc37c8285ed977b705aa6a64c4ca7ba96%2Ffailure_cases.png?generation=1592141930858740&amp;alt=media)\n\n\n\n#3. Reference\n[1] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.\n[2] Hu J, Shen L, Sun G. Squeeze-and-excitation networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7132-7141.\n[3] Han D, Kim J, Kim J. Deep pyramidal residual networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 5927-5935.\n[4] Wang F, Jiang M, Qian C, et al. Residual attention network for image classification[C]. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 3156-3164.\n[5] Xie S, Girshick R, Dollár P, et al. Aggregated residual transformations for deep neural networks[C]. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 1492-1500.\n[6] Dolhansky B, Howes R, Pflaum B, et al. The Deepfake Detection Challenge (DFDC) Preview Dataset[J]. arXiv preprint arXiv:1910.08854, 2019.\n[7] Rossler A, Cozzolino D, Verdoliva L, et al. Faceforensics++: Learning to detect manipulated facial images[C]//Proceedings of the IEEE International Conference on Computer Vision. 2019: 1-11.\n[8] Zhong Z, Zheng L, Kang G, et al. Random erasing data augmentation[J]. arXiv preprint arXiv:1708.04896, 2017.\n[9] Feichtenhofer C, Fan H, Malik J, et al. Slowfast networks for video recognition[C]. In Proceedings of the IEEE International Conference on Computer Vision. 2019: 6202-6211.\n\n------------------------------------------------------------------------------------------------\n###If you feel the above solution is helpful and hope to discuss more offline with us, you may contact Meisong via zhengmeisong@gmail.com",
    "885890": "&gt; both results were removed due to the usage of faceforensic++ data even though this dataset can be accessed by all participants (we later confirmed this with dataset owners).\n\nI think it was made pretty clear throughout the competition that this dataset was not allowed because it required participants to get permission to use it, which means in theory that some people could have been denied permission (even if that never happened).",
    "885955": "We thought it can be used as many participants claimed this in the official data disclosure thread and there was no restrictions for all participants to use the data.  Only when we were told disqualified, we managed to dig out that  FF++ was banned by Kaggle in a discussion thread.  Anyway, hope this will be a lesson for all new Kagglers. If you are not 100% sure if the data is allowed, just don't use it. Or simply don't use any external data.",
    "885989": "Thanks for the detailed summary, and sorry to see the DQ of your time.  Your solution looks SOLID, and I can see the effort and time you have spent here to provide the quality write-up too. \n\nI can completely understand that the competition rules can be very confusing, and more often than not the one bit of useful information is buried deep inside long threads.   \n\nI am not saying this is the case of this specific team, but for new kagglers whose native language is not English,  trying to comprehend what is said by kaggle admin regarding the competition rules can be very difficult.  Especially when those rules also evidently can be interpreted differently between different competitions.",
    "886359": "FaceForensics consists of Youtube videos with given links. Your case may be the same as the 1st removed team?",
    "886405": "Thanks for your understanding, Yifan. The rules are complicated. It's extremely unfriendly for new kaggler. Hopefully, Kaggle could simplify its DATA USE POLICY. just like  ` no external data is allowed for this competition at the very beginning`",
    "886414": "we are not very sure about this.",
    "886513": "Congrats! Very impressive result especially for your first kaggle competition!",
    "887412": "Impressive! and inspiring for a novice such me!",
    "887829": "Congrats!",
    "887853": "Congrats!",
    "888175": "What is ra-92? Google brings some strange stuff...",
    "888243": "pls read the this paper [4] Wang F, Jiang M, Qian C, et al. Residual attention network for image classification[C]. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 3156-3164.",
    "888292": "Great God, such a thorough explanation. It's quite sad that you applied so much effort and didn't get to reap your rewards. Still, congratulations! It's surprising to see novice Kagglers come up with such a detailed solution, that too with reference to research papers. You guys should be proud of yourselves.",
    "888957": "Congrats!!",
    "889644": "Thank you!",
    "890185": "thanks for sharing. really enjoyed reading your solution with the detailed log for each step/experiment.",
    "957165": "very good write-up, sorry for your removal."
  },
  "source": "meta"
}