{
  "id": 140829,
  "title": "EDIT) Private 1089th Place Approach (BIG DROP from public 7th)",
  "url": "/competitions/deepfake-detection-challenge/discussion/140829",
  "author_name": "Jinhan",
  "post_date": "2020-04-03T10:43:22.167000",
  "votes": 71,
  "comment_count": 30,
  "views": 0,
  "content": "<p>EDIT) We shared our approach with concern about shake-up before private leaderboard was revealed. Unfortunately, we experienced way bigger shake-up than we expected in private LB. Since we have no objective evidence about our bad private score, we cannot specify which part was wrong in our submissions. So please take care while reading our approach.</p>\n\n<p>+) Congratulations for winners and those who did overcome shake-up.   </p>\n\n<p>================ Our Sharing Before Private LB ===============</p>\n\n<h1>0. Intro</h1>\n\n<p>It was great experience for us to participate in this competition. So we want to say thank you to organizers of this challenge and others helped us a lot.   </p>\n\n<p><a href=\"/pudae81\">@pudae81</a> Your great introductory seminar about kaggle and code competition helped us a lot in joining this competition. We'll keep our fingers crossed for your journey to new career. 🤞   </p>\n\n<p><a href=\"/humananalog\">@humananalog</a> We couldn't make it this far without great jobs shared by you. We really appreciates your sharing of awesome training and inference kernels. And that's why we share our humble approach to community even if we're not sure whether it's overfitted to public LB or not.</p>\n\n<h1>1. Overview of Our Approach</h1>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png\" alt=\"Method Overview\"> <br>\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png</a></p>\n\n<p>Figure above shows our approach. We used MTCNN as face detector and attention-based models as prediction model.</p>\n\n<h1>2. Face Detector : MTCNN</h1>\n\n<p>We used MTCNN to detect face in video frame. We resized frame to 1280x720 before we pass it to MTCNN. And parameters passed to MTCNN is\n- thresholds of [0.6, 0.7. 0.95]\n- min_face_size of (720/30)</p>\n\n<h1>3. Model Architecture</h1>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png\" alt=\"Model Architecture\"> <br>\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png</a></p>\n\n<p>Model architecture we used is shown above. We first extract feature vectors from each face crop, then apply self-attention based model to sequence of feature vectors to aggregate information.   </p>\n\n<p>ImageNet pretrained efficient-b4 was used as our face crop feature extractor.\nWe froze first 4 blocks due to our GPU memory capability.  </p>\n\n<p>We have two similar models to get prediction of video from sequence of sampled face features.</p>\n\n<h3>3.2.1. Model 1</h3>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png\" alt=\"Model 1 figure\"> <br>\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png</a></p>\n\n<p>Our first model uses self-attention between all sampled faces. \"self-attention module\" in figure contains KQV projection, self-attention, and residual connection. Then fully-connected layer follows to get prediction for each face crop. Finally, we take max prediction value for video prediction values.   </p>\n\n<p>The best score of this model in LB was 0.24827 without any TTA.</p>\n\n<h3>3.2.2. Model 2</h3>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png\" alt=\"Model 2 figure\"> <br>\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png</a></p>\n\n<p>As figure shows, our second model takes self-attention between faces in the same frame first, then takes self-attention between all sampled faces. After that, the procedure is same with Model 1.   </p>\n\n<p>The best score of this model in LB was 0.25178 without any TTA.</p>\n\n<h1>4. Training Method</h1>\n\n<h2>4.1. Face Crops Preparation</h2>\n\n<ul>\n<li>sample 100 frames uniformly for each videos.</li>\n<li>use MTCNN to detect face in frame.</li>\n<li>2x extra margin added to face crop box(so the size is 9 times bigger in area sense) for further augmentation at train time   </li>\n<li>face aligned with five face landmarks.</li>\n</ul>\n\n<h2>4.2. Train/Valid Split</h2>\n\n<ul>\n<li><p>One of our teammates clustered all real videos so that no actor appears in the videos from different clusters at the same time. After that, he splitted some clusters as validation set so that validation set has videos of different resolutions.   </p></li>\n<li><p>With this cluster based real video split, their corresponding fake videos are splitted to the group of their original videos.</p></li>\n<li><p>Every train epoch, we used all real videos and sampled fake videos randomly to balance the number of real and fake samples.</p></li>\n</ul>\n\n<h2>4.3. Augmentation at train time</h2>\n\n<ul>\n<li>Sample 1 to 20 frames from previsously sampled 100 frames and load corresponding face crops.</li>\n<li>Videowise (i.e. applied to all face crops samely)\n<ul><li>DFDC preview dataset paper based quality augmentation\n<ul><li><a href=\"https://arxiv.org/abs/1910.08854\"></a><a href=\"https://arxiv.org/abs/1910.08854\"></a><a href=\"https://arxiv.org/abs/1910.08854\">https://arxiv.org/abs/1910.08854</a></li></ul></li>\n<li>Various image transforms such as rotate, equalize, solarize, posterize etc...</li>\n<li>Random horizontal flip</li></ul></li>\n<li>Imagewise\n<ul><li>Random Crop\n<ul><li>Randomness occures in margin and face-center translation</li></ul></li>\n<li>Resize Interpolation\n<ul><li>we use three different resize interpolation : bilinear, bicubic, nearest</li></ul></li></ul></li>\n</ul>\n\n<h2>4.4. Optimizer</h2>\n\n<p>We used same optimizer settings for both model.\n- SGD with momentum 0.9\n- Learning Rate Schedule\n    - 5 epochs of linear warm-up from 0 to 0.001\n    - Cosine annealing from 0.001 to 0 through 35 epochs (total 40 epochs)</p>\n\n<h2>4.5. Further Regularization</h2>\n\n<ul>\n<li>weight decay: 0.0001   </li>\n<li>dropout layer right before final fully connected layer: p=0.4</li>\n<li>label smoothing epsilon: 0.01</li>\n</ul>\n\n<h2>4.6. Validation losses of model 1 while training</h2>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png\" alt=\"Model 1 loss\">\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png</a></p>\n\n<h2>4.7. Validation losses of model 2 while training</h2>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png\" alt=\"Model 2 loss\">\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png</a></p>\n\n<h1>5. Inference Strategy</h1>\n\n<p>With our inference strategy below, we got 0.23086 in LB</p>\n\n<h2>5.1. Sample Faces</h2>\n\n<ul>\n<li>Sample 20 frames per video uniformly.   </li>\n<li>Detect and crop at most 3 faces from each frame with threshold stated above.</li>\n</ul>\n\n<h2>5.2. Ensemble</h2>\n\n<p>We averaged predictions from two models</p>\n\n<h2>5.3. Test Time Augmentation for each model</h2>\n\n<p>For each face detected, make 8 different crops of different margin and offset\n- 2 of 8 : two different margins of 1.1 and 1.15 without offsets.\n- 2 of 8 : two different y-axis offset of 0.15, -0.15 with margin of 1.15\n- 4 of 8 : horizontal flips of above 4 crops</p>\n\n<h1>6. Outro</h1>\n\n<p>Thank you for reading. All questions are welcome.</p>",
  "messages": [
    {
      "id": 796142,
      "postDate": "2020-04-03T10:43:22.167Z",
      "content": "<p>EDIT) We shared our approach with concern about shake-up before private leaderboard was revealed. Unfortunately, we experienced way bigger shake-up than we expected in private LB. Since we have no objective evidence about our bad private score, we cannot specify which part was wrong in our submissions. So please take care while reading our approach.</p>\n\n<p>+) Congratulations for winners and those who did overcome shake-up.   </p>\n\n<p>================ Our Sharing Before Private LB ===============</p>\n\n<h1>0. Intro</h1>\n\n<p>It was great experience for us to participate in this competition. So we want to say thank you to organizers of this challenge and others helped us a lot.   </p>\n\n<p><a href=\"/pudae81\">@pudae81</a> Your great introductory seminar about kaggle and code competition helped us a lot in joining this competition. We'll keep our fingers crossed for your journey to new career. 🤞   </p>\n\n<p><a href=\"/humananalog\">@humananalog</a> We couldn't make it this far without great jobs shared by you. We really appreciates your sharing of awesome training and inference kernels. And that's why we share our humble approach to community even if we're not sure whether it's overfitted to public LB or not.</p>\n\n<h1>1. Overview of Our Approach</h1>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png\" alt=\"Method Overview\"> <br>\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png</a></p>\n\n<p>Figure above shows our approach. We used MTCNN as face detector and attention-based models as prediction model.</p>\n\n<h1>2. Face Detector : MTCNN</h1>\n\n<p>We used MTCNN to detect face in video frame. We resized frame to 1280x720 before we pass it to MTCNN. And parameters passed to MTCNN is\n- thresholds of [0.6, 0.7. 0.95]\n- min_face_size of (720/30)</p>\n\n<h1>3. Model Architecture</h1>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png\" alt=\"Model Architecture\"> <br>\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png</a></p>\n\n<p>Model architecture we used is shown above. We first extract feature vectors from each face crop, then apply self-attention based model to sequence of feature vectors to aggregate information.   </p>\n\n<p>ImageNet pretrained efficient-b4 was used as our face crop feature extractor.\nWe froze first 4 blocks due to our GPU memory capability.  </p>\n\n<p>We have two similar models to get prediction of video from sequence of sampled face features.</p>\n\n<h3>3.2.1. Model 1</h3>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png\" alt=\"Model 1 figure\"> <br>\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png</a></p>\n\n<p>Our first model uses self-attention between all sampled faces. \"self-attention module\" in figure contains KQV projection, self-attention, and residual connection. Then fully-connected layer follows to get prediction for each face crop. Finally, we take max prediction value for video prediction values.   </p>\n\n<p>The best score of this model in LB was 0.24827 without any TTA.</p>\n\n<h3>3.2.2. Model 2</h3>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png\" alt=\"Model 2 figure\"> <br>\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png</a></p>\n\n<p>As figure shows, our second model takes self-attention between faces in the same frame first, then takes self-attention between all sampled faces. After that, the procedure is same with Model 1.   </p>\n\n<p>The best score of this model in LB was 0.25178 without any TTA.</p>\n\n<h1>4. Training Method</h1>\n\n<h2>4.1. Face Crops Preparation</h2>\n\n<ul>\n<li>sample 100 frames uniformly for each videos.</li>\n<li>use MTCNN to detect face in frame.</li>\n<li>2x extra margin added to face crop box(so the size is 9 times bigger in area sense) for further augmentation at train time   </li>\n<li>face aligned with five face landmarks.</li>\n</ul>\n\n<h2>4.2. Train/Valid Split</h2>\n\n<ul>\n<li><p>One of our teammates clustered all real videos so that no actor appears in the videos from different clusters at the same time. After that, he splitted some clusters as validation set so that validation set has videos of different resolutions.   </p></li>\n<li><p>With this cluster based real video split, their corresponding fake videos are splitted to the group of their original videos.</p></li>\n<li><p>Every train epoch, we used all real videos and sampled fake videos randomly to balance the number of real and fake samples.</p></li>\n</ul>\n\n<h2>4.3. Augmentation at train time</h2>\n\n<ul>\n<li>Sample 1 to 20 frames from previsously sampled 100 frames and load corresponding face crops.</li>\n<li>Videowise (i.e. applied to all face crops samely)\n<ul><li>DFDC preview dataset paper based quality augmentation\n<ul><li><a href=\"https://arxiv.org/abs/1910.08854\"></a><a href=\"https://arxiv.org/abs/1910.08854\"></a><a href=\"https://arxiv.org/abs/1910.08854\">https://arxiv.org/abs/1910.08854</a></li></ul></li>\n<li>Various image transforms such as rotate, equalize, solarize, posterize etc...</li>\n<li>Random horizontal flip</li></ul></li>\n<li>Imagewise\n<ul><li>Random Crop\n<ul><li>Randomness occures in margin and face-center translation</li></ul></li>\n<li>Resize Interpolation\n<ul><li>we use three different resize interpolation : bilinear, bicubic, nearest</li></ul></li></ul></li>\n</ul>\n\n<h2>4.4. Optimizer</h2>\n\n<p>We used same optimizer settings for both model.\n- SGD with momentum 0.9\n- Learning Rate Schedule\n    - 5 epochs of linear warm-up from 0 to 0.001\n    - Cosine annealing from 0.001 to 0 through 35 epochs (total 40 epochs)</p>\n\n<h2>4.5. Further Regularization</h2>\n\n<ul>\n<li>weight decay: 0.0001   </li>\n<li>dropout layer right before final fully connected layer: p=0.4</li>\n<li>label smoothing epsilon: 0.01</li>\n</ul>\n\n<h2>4.6. Validation losses of model 1 while training</h2>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png\" alt=\"Model 1 loss\">\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png</a></p>\n\n<h2>4.7. Validation losses of model 2 while training</h2>\n\n<p><img src=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png\" alt=\"Model 2 loss\">\noriginal image: <a href=\"https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png\">https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png</a></p>\n\n<h1>5. Inference Strategy</h1>\n\n<p>With our inference strategy below, we got 0.23086 in LB</p>\n\n<h2>5.1. Sample Faces</h2>\n\n<ul>\n<li>Sample 20 frames per video uniformly.   </li>\n<li>Detect and crop at most 3 faces from each frame with threshold stated above.</li>\n</ul>\n\n<h2>5.2. Ensemble</h2>\n\n<p>We averaged predictions from two models</p>\n\n<h2>5.3. Test Time Augmentation for each model</h2>\n\n<p>For each face detected, make 8 different crops of different margin and offset\n- 2 of 8 : two different margins of 1.1 and 1.15 without offsets.\n- 2 of 8 : two different y-axis offset of 0.15, -0.15 with margin of 1.15\n- 4 of 8 : horizontal flips of above 4 crops</p>\n\n<h1>6. Outro</h1>\n\n<p>Thank you for reading. All questions are welcome.</p>",
      "rawMarkdown": "EDIT) We shared our approach with concern about shake-up before private leaderboard was revealed. Unfortunately, we experienced way bigger shake-up than we expected in private LB. Since we have no objective evidence about our bad private score, we cannot specify which part was wrong in our submissions. So please take care while reading our approach.\n\n+) Congratulations for winners and those who did overcome shake-up.   \n   \n\n\n\n================ Our Sharing Before Private LB ===============\n\n# 0. Intro\nIt was great experience for us to participate in this competition. So we want to say thank you to organizers of this challenge and others helped us a lot.   \n\n@pudae81 Your great introductory seminar about kaggle and code competition helped us a lot in joining this competition. We'll keep our fingers crossed for your journey to new career. 🤞   \n\n@humananalog We couldn't make it this far without great jobs shared by you. We really appreciates your sharing of awesome training and inference kernels. And that's why we share our humble approach to community even if we're not sure whether it's overfitted to public LB or not.\n\n# 1. Overview of Our Approach\n![Method Overview](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png)   \noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png\n\nFigure above shows our approach. We used MTCNN as face detector and attention-based models as prediction model.\n\n# 2. Face Detector : MTCNN \nWe used MTCNN to detect face in video frame. We resized frame to 1280x720 before we pass it to MTCNN. And parameters passed to MTCNN is\n- thresholds of [0.6, 0.7. 0.95]\n- min\\_face\\_size of (720/30)\n\n# 3. Model Architecture\n![Model Architecture](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png)   \noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png\n\nModel architecture we used is shown above. We first extract feature vectors from each face crop, then apply self-attention based model to sequence of feature vectors to aggregate information.   \n\nImageNet pretrained efficient-b4 was used as our face crop feature extractor.\nWe froze first 4 blocks due to our GPU memory capability.  \n\nWe have two similar models to get prediction of video from sequence of sampled face features.\n\n### 3.2.1. Model 1\n![Model 1 figure](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png)   \noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png\n\nOur first model uses self-attention between all sampled faces. \"self-attention module\" in figure contains KQV projection, self-attention, and residual connection. Then fully-connected layer follows to get prediction for each face crop. Finally, we take max prediction value for video prediction values.   \n\nThe best score of this model in LB was 0.24827 without any TTA.\n\n\n### 3.2.2. Model 2\n![Model 2 figure](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png)   \noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png\n\nAs figure shows, our second model takes self-attention between faces in the same frame first, then takes self-attention between all sampled faces. After that, the procedure is same with Model 1.   \n\nThe best score of this model in LB was 0.25178 without any TTA.\n\n\n# 4. Training Method\n\n## 4.1. Face Crops Preparation\n- sample 100 frames uniformly for each videos.\n- use MTCNN to detect face in frame.\n- 2x extra margin added to face crop box(so the size is 9 times bigger in area sense) for further augmentation at train time   \n- face aligned with five face landmarks.\n\n## 4.2. Train/Valid Split\n- One of our teammates clustered all real videos so that no actor appears in the videos from different clusters at the same time. After that, he splitted some clusters as validation set so that validation set has videos of different resolutions.   \n\n- With this cluster based real video split, their corresponding fake videos are splitted to the group of their original videos.\n\n- Every train epoch, we used all real videos and sampled fake videos randomly to balance the number of real and fake samples.\n\n## 4.3. Augmentation at train time\n- Sample 1 to 20 frames from previsously sampled 100 frames and load corresponding face crops.\n- Videowise (i.e. applied to all face crops samely)\n    - DFDC preview dataset paper based quality augmentation\n        - https://arxiv.org/abs/1910.08854\n    - Various image transforms such as rotate, equalize, solarize, posterize etc...\n    - Random horizontal flip\n- Imagewise\n    - Random Crop\n        - Randomness occures in margin and face-center translation\n    - Resize Interpolation\n        - we use three different resize interpolation : bilinear, bicubic, nearest\n\n\n## 4.4. Optimizer\nWe used same optimizer settings for both model.\n- SGD with momentum 0.9\n- Learning Rate Schedule\n    - 5 epochs of linear warm-up from 0 to 0.001\n    - Cosine annealing from 0.001 to 0 through 35 epochs (total 40 epochs)\n\n\n## 4.5. Further Regularization\n- weight decay: 0.0001   \n- dropout layer right before final fully connected layer: p=0.4\n- label smoothing epsilon: 0.01\n\n## 4.6. Validation losses of model 1 while training\n![Model 1 loss](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png)\noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png\n\n## 4.7. Validation losses of model 2 while training\n![Model 2 loss](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png)\noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png\n\n# 5. Inference Strategy\nWith our inference strategy below, we got 0.23086 in LB\n\n## 5.1. Sample Faces\n- Sample 20 frames per video uniformly.   \n- Detect and crop at most 3 faces from each frame with threshold stated above.\n\n\n## 5.2. Ensemble\nWe averaged predictions from two models\n\n## 5.3. Test Time Augmentation for each model\nFor each face detected, make 8 different crops of different margin and offset\n- 2 of 8 : two different margins of 1.1 and 1.15 without offsets.\n- 2 of 8 : two different y-axis offset of 0.15, -0.15 with margin of 1.15\n- 4 of 8 : horizontal flips of above 4 crops\n\n\n# 6. Outro\nThank you for reading. All questions are welcome.",
      "votes": 71
    },
    {
      "id": 796209,
      "postDate": "2020-04-03T11:56:45.120Z",
      "content": "<p>Great write-up. Happy to have been of assistance. 😄 </p>",
      "rawMarkdown": "Great write-up. Happy to have been of assistance. 😄 ",
      "votes": 3
    },
    {
      "id": 796335,
      "postDate": "2020-04-03T13:57:40.623Z",
      "content": "<p>Thank you for sharing!\nI notice that you use a large face margin (2x). Will this margin parameter have a great impact on your model performance? We have tried a similar attention method with a small face margin 0.05, but the performance is poor.</p>",
      "rawMarkdown": "Thank you for sharing!\nI notice that you use a large face margin (2x). Will this margin parameter have a great impact on your model performance? We have tried a similar attention method with a small face margin 0.05, but the performance is poor.",
      "votes": 1,
      "replies": [
        {
          "id": 796447,
          "postDate": "2020-04-03T15:33:22.817Z",
          "content": "<p>We cropped face with 2x extra margin. But at the train time, we cropped it again so that extra margin is randomly chosen in [0.8, 1.2] and random translation offsets for two axes in [-0.2, 0.2]. They are distributed as truncated normal.\nAfter some EDA, we thought that detected crop box is too tight to contain whole modified area in video frames. So after we changed our face detector from BlazeFace to MTCNN, we always used margins bigger than 1/2.\nWe also wanted to test different margins but we didn't have enough time. So analysis on margin itself was not enough. (we changed multiple things at the same time when we experiment our settings)</p>",
          "rawMarkdown": "We cropped face with 2x extra margin. But at the train time, we cropped it again so that extra margin is randomly chosen in [0.8, 1.2] and random translation offsets for two axes in [-0.2, 0.2]. They are distributed as truncated normal.\nAfter some EDA, we thought that detected crop box is too tight to contain whole modified area in video frames. So after we changed our face detector from BlazeFace to MTCNN, we always used margins bigger than 1/2.\nWe also wanted to test different margins but we didn't have enough time. So analysis on margin itself was not enough. (we changed multiple things at the same time when we experiment our settings)",
          "votes": 1
        }
      ]
    },
    {
      "id": 796187,
      "postDate": "2020-04-03T11:34:37.030Z",
      "content": "<p>Nice work and congratulation! Could you share the video ids which clustered by actors? And did you compare the validation splitted by folder? There is large gap in my CV/LB. I guess maybe the better validation can figure out it.</p>",
      "rawMarkdown": "Nice work and congratulation! Could you share the video ids which clustered by actors? And did you compare the validation splitted by folder? There is large gap in my CV/LB. I guess maybe the better validation can figure out it.",
      "votes": 1,
      "replies": [
        {
          "id": 796801,
          "postDate": "2020-04-03T23:38:19.363Z",
          "content": "<p>It depends on <a href=\"/woodykwon\">@woodykwon</a> 's choice.\nBut  he think his work is not perfect. I don't agree. The outcome is excellent and helpful</p>",
          "rawMarkdown": "It depends on @woodykwon 's choice.\nBut  he think his work is not perfect. I don't agree. The outcome is excellent and helpful"
        }
      ]
    },
    {
      "id": 796161,
      "postDate": "2020-04-03T11:07:40.273Z",
      "content": "<p>Congratulation on your position. Thank you for sharing, a very nice write up indeed. Wonder how do you get the idea of using self attention, it worked really well, wonderful!\nQuestion: what is really <strong>minfacesize of (720/30)</strong> ?</p>",
      "rawMarkdown": "Congratulation on your position. Thank you for sharing, a very nice write up indeed. Wonder how do you get the idea of using self attention, it worked really well, wonderful!\nQuestion: what is really **minfacesize of (720/30)** ?",
      "votes": 1,
      "replies": [
        {
          "id": 796184,
          "postDate": "2020-04-03T11:31:47.757Z",
          "content": "<p>We want our model to be able to aggregate information from similar faces so that it doesn't get confused by unmodified faces in the videos (or image patch sampled by face detector which is not a face). That's why we chose to use self-attention mechanism and take max prediction between them.</p>\n\n<p>question about min_face_size will be answered later</p>",
          "rawMarkdown": "We want our model to be able to aggregate information from similar faces so that it doesn't get confused by unmodified faces in the videos (or image patch sampled by face detector which is not a face). That's why we chose to use self-attention mechanism and take max prediction between them.\n\nquestion about min\\_face\\_size will be answered later"
        }
      ]
    },
    {
      "id": 796152,
      "postDate": "2020-04-03T10:58:32.573Z",
      "content": "<p>Very impressive! Especially impressive LB score for single models, &lt;0.25 for Model 1.</p>\n\n<p>Do you think the self-attention model will generalise well or badly to deepfakes made with different methods from the 'wild' that will be in the final private LB?</p>",
      "rawMarkdown": "Very impressive! Especially impressive LB score for single models, &lt;0.25 for Model 1.\n\nDo you think the self-attention model will generalise well or badly to deepfakes made with different methods from the 'wild' that will be in the final private LB?",
      "votes": 1,
      "replies": [
        {
          "id": 796179,
          "postDate": "2020-04-03T11:27:31.397Z",
          "content": "<p>To be honest, I'm not sure whether our models generalize well or not. (totally hope so)\nWe've never tested our models to \"organic videos with and without deepfakes\"</p>",
          "rawMarkdown": "To be honest, I'm not sure whether our models generalize well or not. (totally hope so)\nWe've never tested our models to \"organic videos with and without deepfakes\""
        },
        {
          "id": 796976,
          "postDate": "2020-04-04T06:35:52.530Z",
          "content": "<p><a href=\"/jinhan23\">@jinhan23</a>  thanks for sharing work.Good work\nCould u point me to the attention model you are talking about . . I dint use any before so want to see.</p>",
          "rawMarkdown": "@jinhan23  thanks for sharing work.Good work\nCould u point me to the attention model you are talking about . . I dint use any before so want to see."
        },
        {
          "id": 797674,
          "postDate": "2020-04-04T19:24:29.120Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> I guess this paper would be helpful.\n<a href=\"https://arxiv.org/abs/1706.03762\">https://arxiv.org/abs/1706.03762</a>\nour attention looks similar to scaled dot-product attention in this paper.</p>",
          "rawMarkdown": "@jaideepvalani I guess this paper would be helpful.\nhttps://arxiv.org/abs/1706.03762\nour attention looks similar to scaled dot-product attention in this paper."
        }
      ]
    },
    {
      "id": 818457,
      "postDate": "2020-04-23T22:42:32.557Z",
      "content": "<p>Biggest RIP?</p>",
      "rawMarkdown": "Biggest RIP?",
      "votes": 2
    },
    {
      "id": 796566,
      "postDate": "2020-04-03T17:38:03.497Z",
      "content": "<p>Are you guys planning to share the code? Amazing architecture!</p>",
      "rawMarkdown": "Are you guys planning to share the code? Amazing architecture!",
      "votes": 2,
      "replies": [
        {
          "id": 796797,
          "postDate": "2020-04-03T23:27:31.277Z",
          "content": "<ol>\n<li>We should check kaggle's code sharing rule.  (during competition period) </li>\n<li>Code sharing depends on <a href=\"/caffeinism\">@caffeinism</a>  and <a href=\"/snowyunee\">@snowyunee</a>' s amount of free time  </li>\n</ol>",
          "rawMarkdown": " 1. We should check kaggle's code sharing rule.  (during competition period) \n 2. Code sharing depends on @caffeinism  and @snowyunee' s amount of free time  "
        },
        {
          "id": 796842,
          "postDate": "2020-04-04T01:27:31.973Z",
          "content": "<p>I think we should check these also.\n- Generalization of our work (private LB score)\n- Reproducibility (We couldn’t do the same experiment multiple times due to lacking time)</p>",
          "rawMarkdown": "I think we should check these also.\n- Generalization of our work (private LB score)\n- Reproducibility (We couldn’t do the same experiment multiple times due to lacking time)"
        },
        {
          "id": 796967,
          "postDate": "2020-04-04T06:05:40.460Z",
          "content": "<p>Thanks for your replies. I am of particular interest on your self-attention architecture. It seems from the diagram that you do not have a decoder, and merely using the output at each time step as the query, am I right? Do you mind to elaborate more on how you implemented your Model 1 and 2 ? </p>",
          "rawMarkdown": "Thanks for your replies. I am of particular interest on your self-attention architecture. It seems from the diagram that you do not have a decoder, and merely using the output at each time step as the query, am I right? Do you mind to elaborate more on how you implemented your Model 1 and 2 ? "
        },
        {
          "id": 797666,
          "postDate": "2020-04-04T19:18:37.650Z",
          "content": "<p>As I wrote in the article, \"self-attention module\" in our figure includes\n- Query, Key, Value projection(from input)\n- Self-attention output with projected Q, K, V (dot-product attention)\n- Residual Connection (which means input + self-attention output is final output of this module)</p>\n\n<p>We didn't use decoder with attention mechanism in model 1 and 2.</p>",
          "rawMarkdown": "As I wrote in the article, \"self-attention module\" in our figure includes\n- Query, Key, Value projection(from input)\n- Self-attention output with projected Q, K, V (dot-product attention)\n- Residual Connection (which means input + self-attention output is final output of this module)\n\nWe didn't use decoder with attention mechanism in model 1 and 2."
        },
        {
          "id": 800531,
          "postDate": "2020-04-07T13:49:58.910Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2799186%2Faf2ba2298848dbe037996ed4965d2cdd%2FScreenshot%202020-04-07%20at%209.50.52%20PM.png?generation=1586267466346639&amp;alt=media\" alt=\"\"></p>\n\n<p>Hey, may I check if this is the same as your architecture for the self-attention part? </p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2799186%2Faf2ba2298848dbe037996ed4965d2cdd%2FScreenshot%202020-04-07%20at%209.50.52%20PM.png?generation=1586267466346639&amp;alt=media)\n\nHey, may I check if this is the same as your architecture for the self-attention part? "
        }
      ]
    },
    {
      "id": 824037,
      "postDate": "2020-04-28T05:47:35.383Z",
      "content": "<p>Attention model is defeated by deepfake! Actually, almost all top models are defeated with a 0.2x drop or more.</p>",
      "rawMarkdown": "Attention model is defeated by deepfake! Actually, almost all top models are defeated with a 0.2x drop or more.",
      "votes": -1
    },
    {
      "id": 823723,
      "postDate": "2020-04-27T20:53:40.413Z",
      "content": "<p>At least you won a gold for this thread!</p>",
      "rawMarkdown": "At least you won a gold for this thread!"
    },
    {
      "id": 814629,
      "postDate": "2020-04-20T20:25:59.663Z",
      "content": "<p>Ótima redação. Parabéns</p>",
      "rawMarkdown": "\nÓtima redação. Parabéns"
    },
    {
      "id": 812773,
      "postDate": "2020-04-19T04:34:40.670Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!"
    },
    {
      "id": 811478,
      "postDate": "2020-04-18T00:45:36.893Z",
      "content": "<p>good</p>",
      "rawMarkdown": "good"
    },
    {
      "id": 799219,
      "postDate": "2020-04-06T09:17:28.920Z",
      "content": "<p>Thank you this detailed feedback and congratulations on your attention architecture!\nCan you share which tool/software did you use to generate those great and self explained figures?</p>",
      "rawMarkdown": "Thank you this detailed feedback and congratulations on your attention architecture!\nCan you share which tool/software did you use to generate those great and self explained figures?",
      "replies": [
        {
          "id": 800224,
          "postDate": "2020-04-07T07:30:50.197Z",
          "content": "<p>One of my teammates generated those figures with MS PowerPoint.</p>",
          "rawMarkdown": "One of my teammates generated those figures with MS PowerPoint."
        }
      ]
    },
    {
      "id": 796512,
      "postDate": "2020-04-03T16:44:20.070Z",
      "content": "<p>Amazing work!! Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Amazing work!! Congrats and thanks for sharing."
    },
    {
      "id": 819536,
      "postDate": "2020-04-24T16:51:36.130Z",
      "content": "<p>Nice job!! Gonna study your work.</p>",
      "rawMarkdown": "Nice job!! Gonna study your work.",
      "isDeleted": true,
      "replies": [
        {
          "id": 831712,
          "postDate": "2020-05-03T15:01:09.030Z",
          "content": "<p>Please don't :)</p>",
          "rawMarkdown": "Please don't :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 803209,
      "postDate": "2020-04-10T08:42:55.940Z",
      "content": "<p>Thanks for sharing. It was useful.</p>",
      "rawMarkdown": "Thanks for sharing. It was useful."
    },
    {
      "id": 798708,
      "postDate": "2020-04-05T18:58:54.877Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !"
    }
  ],
  "comments": [
    {
      "id": 796209,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2020-04-03T11:56:45.120000",
      "content": "<p>Great write-up. Happy to have been of assistance. 😄 </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 796335,
      "author_name": "Chason",
      "author_url": "",
      "post_date": "2020-04-03T13:57:40.623000",
      "content": "<p>Thank you for sharing!\nI notice that you use a large face margin (2x). Will this margin parameter have a great impact on your model performance? We have tried a similar attention method with a small face margin 0.05, but the performance is poor.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 796447,
          "author_name": "Jinhan",
          "author_url": "",
          "post_date": "2020-04-03T15:33:22.817000",
          "content": "<p>We cropped face with 2x extra margin. But at the train time, we cropped it again so that extra margin is randomly chosen in [0.8, 1.2] and random translation offsets for two axes in [-0.2, 0.2]. They are distributed as truncated normal.\nAfter some EDA, we thought that detected crop box is too tight to contain whole modified area in video frames. So after we changed our face detector from BlazeFace to MTCNN, we always used margins bigger than 1/2.\nWe also wanted to test different margins but we didn't have enough time. So analysis on margin itself was not enough. (we changed multiple things at the same time when we experiment our settings)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 796187,
      "author_name": "Xavier_lin",
      "author_url": "",
      "post_date": "2020-04-03T11:34:37.030000",
      "content": "<p>Nice work and congratulation! Could you share the video ids which clustered by actors? And did you compare the validation splitted by folder? There is large gap in my CV/LB. I guess maybe the better validation can figure out it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 796801,
          "author_name": "taeksoon.kwon",
          "author_url": "",
          "post_date": "2020-04-03T23:38:19.363000",
          "content": "<p>It depends on <a href=\"/woodykwon\">@woodykwon</a> 's choice.\nBut  he think his work is not perfect. I don't agree. The outcome is excellent and helpful</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 796161,
      "author_name": "Zungmann",
      "author_url": "",
      "post_date": "2020-04-03T11:07:40.273000",
      "content": "<p>Congratulation on your position. Thank you for sharing, a very nice write up indeed. Wonder how do you get the idea of using self attention, it worked really well, wonderful!\nQuestion: what is really <strong>minfacesize of (720/30)</strong> ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 796184,
          "author_name": "Jinhan",
          "author_url": "",
          "post_date": "2020-04-03T11:31:47.757000",
          "content": "<p>We want our model to be able to aggregate information from similar faces so that it doesn't get confused by unmodified faces in the videos (or image patch sampled by face detector which is not a face). That's why we chose to use self-attention mechanism and take max prediction between them.</p>\n\n<p>question about min_face_size will be answered later</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 796152,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-04-03T10:58:32.573000",
      "content": "<p>Very impressive! Especially impressive LB score for single models, &lt;0.25 for Model 1.</p>\n\n<p>Do you think the self-attention model will generalise well or badly to deepfakes made with different methods from the 'wild' that will be in the final private LB?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 796179,
          "author_name": "Jinhan",
          "author_url": "",
          "post_date": "2020-04-03T11:27:31.397000",
          "content": "<p>To be honest, I'm not sure whether our models generalize well or not. (totally hope so)\nWe've never tested our models to \"organic videos with and without deepfakes\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 796976,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-04-04T06:35:52.530000",
          "content": "<p><a href=\"/jinhan23\">@jinhan23</a>  thanks for sharing work.Good work\nCould u point me to the attention model you are talking about . . I dint use any before so want to see.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 797674,
          "author_name": "Jinhan",
          "author_url": "",
          "post_date": "2020-04-04T19:24:29.120000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> I guess this paper would be helpful.\n<a href=\"https://arxiv.org/abs/1706.03762\">https://arxiv.org/abs/1706.03762</a>\nour attention looks similar to scaled dot-product attention in this paper.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 818457,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-04-23T22:42:32.557000",
      "content": "<p>Biggest RIP?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 796566,
      "author_name": "Kar Kin",
      "author_url": "",
      "post_date": "2020-04-03T17:38:03.497000",
      "content": "<p>Are you guys planning to share the code? Amazing architecture!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 796797,
          "author_name": "taeksoon.kwon",
          "author_url": "",
          "post_date": "2020-04-03T23:27:31.277000",
          "content": "<ol>\n<li>We should check kaggle's code sharing rule.  (during competition period) </li>\n<li>Code sharing depends on <a href=\"/caffeinism\">@caffeinism</a>  and <a href=\"/snowyunee\">@snowyunee</a>' s amount of free time  </li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 796842,
          "author_name": "Jinhan",
          "author_url": "",
          "post_date": "2020-04-04T01:27:31.973000",
          "content": "<p>I think we should check these also.\n- Generalization of our work (private LB score)\n- Reproducibility (We couldn’t do the same experiment multiple times due to lacking time)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 796967,
          "author_name": "Kar Kin",
          "author_url": "",
          "post_date": "2020-04-04T06:05:40.460000",
          "content": "<p>Thanks for your replies. I am of particular interest on your self-attention architecture. It seems from the diagram that you do not have a decoder, and merely using the output at each time step as the query, am I right? Do you mind to elaborate more on how you implemented your Model 1 and 2 ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 797666,
          "author_name": "Jinhan",
          "author_url": "",
          "post_date": "2020-04-04T19:18:37.650000",
          "content": "<p>As I wrote in the article, \"self-attention module\" in our figure includes\n- Query, Key, Value projection(from input)\n- Self-attention output with projected Q, K, V (dot-product attention)\n- Residual Connection (which means input + self-attention output is final output of this module)</p>\n\n<p>We didn't use decoder with attention mechanism in model 1 and 2.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 800531,
          "author_name": "Kar Kin",
          "author_url": "",
          "post_date": "2020-04-07T13:49:58.910000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2799186%2Faf2ba2298848dbe037996ed4965d2cdd%2FScreenshot%202020-04-07%20at%209.50.52%20PM.png?generation=1586267466346639&amp;alt=media\" alt=\"\"></p>\n\n<p>Hey, may I check if this is the same as your architecture for the self-attention part? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 824037,
      "author_name": "ant1",
      "author_url": "",
      "post_date": "2020-04-28T05:47:35.383000",
      "content": "<p>Attention model is defeated by deepfake! Actually, almost all top models are defeated with a 0.2x drop or more.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 823723,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-27T20:53:40.413000",
      "content": "<p>At least you won a gold for this thread!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 814629,
      "author_name": "Elias Feitoza Feitoza",
      "author_url": "",
      "post_date": "2020-04-20T20:25:59.663000",
      "content": "<p>Ótima redação. Parabéns</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 812773,
      "author_name": "Mohib Ayub",
      "author_url": "",
      "post_date": "2020-04-19T04:34:40.670000",
      "content": "<p>Great work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 811478,
      "author_name": "Takos1112",
      "author_url": "",
      "post_date": "2020-04-18T00:45:36.893000",
      "content": "<p>good</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 799219,
      "author_name": "Maxime Churin",
      "author_url": "",
      "post_date": "2020-04-06T09:17:28.920000",
      "content": "<p>Thank you this detailed feedback and congratulations on your attention architecture!\nCan you share which tool/software did you use to generate those great and self explained figures?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 800224,
          "author_name": "Jinhan",
          "author_url": "",
          "post_date": "2020-04-07T07:30:50.197000",
          "content": "<p>One of my teammates generated those figures with MS PowerPoint.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 796512,
      "author_name": "Mont3z Claro5",
      "author_url": "",
      "post_date": "2020-04-03T16:44:20.070000",
      "content": "<p>Amazing work!! Congrats and thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819536,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T16:51:36.130000",
      "content": "<p>Nice job!! Gonna study your work.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 831712,
          "author_name": "TestTrained",
          "author_url": "",
          "post_date": "2020-05-03T15:01:09.030000",
          "content": "<p>Please don't :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 803209,
      "author_name": "Moniga",
      "author_url": "",
      "post_date": "2020-04-10T08:42:55.940000",
      "content": "<p>Thanks for sharing. It was useful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 798708,
      "author_name": "Dr. Hemanth Kumar",
      "author_url": "",
      "post_date": "2020-04-05T18:58:54.877000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "796142": "EDIT) We shared our approach with concern about shake-up before private leaderboard was revealed. Unfortunately, we experienced way bigger shake-up than we expected in private LB. Since we have no objective evidence about our bad private score, we cannot specify which part was wrong in our submissions. So please take care while reading our approach.\n\n+) Congratulations for winners and those who did overcome shake-up.   \n   \n\n\n\n================ Our Sharing Before Private LB ===============\n\n# 0. Intro\nIt was great experience for us to participate in this competition. So we want to say thank you to organizers of this challenge and others helped us a lot.   \n\n@pudae81 Your great introductory seminar about kaggle and code competition helped us a lot in joining this competition. We'll keep our fingers crossed for your journey to new career. 🤞   \n\n@humananalog We couldn't make it this far without great jobs shared by you. We really appreciates your sharing of awesome training and inference kernels. And that's why we share our humble approach to community even if we're not sure whether it's overfitted to public LB or not.\n\n# 1. Overview of Our Approach\n![Method Overview](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png)   \noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure1.png\n\nFigure above shows our approach. We used MTCNN as face detector and attention-based models as prediction model.\n\n# 2. Face Detector : MTCNN \nWe used MTCNN to detect face in video frame. We resized frame to 1280x720 before we pass it to MTCNN. And parameters passed to MTCNN is\n- thresholds of [0.6, 0.7. 0.95]\n- min\\_face\\_size of (720/30)\n\n# 3. Model Architecture\n![Model Architecture](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png)   \noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure2.png\n\nModel architecture we used is shown above. We first extract feature vectors from each face crop, then apply self-attention based model to sequence of feature vectors to aggregate information.   \n\nImageNet pretrained efficient-b4 was used as our face crop feature extractor.\nWe froze first 4 blocks due to our GPU memory capability.  \n\nWe have two similar models to get prediction of video from sequence of sampled face features.\n\n### 3.2.1. Model 1\n![Model 1 figure](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png)   \noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure3.png\n\nOur first model uses self-attention between all sampled faces. \"self-attention module\" in figure contains KQV projection, self-attention, and residual connection. Then fully-connected layer follows to get prediction for each face crop. Finally, we take max prediction value for video prediction values.   \n\nThe best score of this model in LB was 0.24827 without any TTA.\n\n\n### 3.2.2. Model 2\n![Model 2 figure](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png)   \noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/figure4.png\n\nAs figure shows, our second model takes self-attention between faces in the same frame first, then takes self-attention between all sampled faces. After that, the procedure is same with Model 1.   \n\nThe best score of this model in LB was 0.25178 without any TTA.\n\n\n# 4. Training Method\n\n## 4.1. Face Crops Preparation\n- sample 100 frames uniformly for each videos.\n- use MTCNN to detect face in frame.\n- 2x extra margin added to face crop box(so the size is 9 times bigger in area sense) for further augmentation at train time   \n- face aligned with five face landmarks.\n\n## 4.2. Train/Valid Split\n- One of our teammates clustered all real videos so that no actor appears in the videos from different clusters at the same time. After that, he splitted some clusters as validation set so that validation set has videos of different resolutions.   \n\n- With this cluster based real video split, their corresponding fake videos are splitted to the group of their original videos.\n\n- Every train epoch, we used all real videos and sampled fake videos randomly to balance the number of real and fake samples.\n\n## 4.3. Augmentation at train time\n- Sample 1 to 20 frames from previsously sampled 100 frames and load corresponding face crops.\n- Videowise (i.e. applied to all face crops samely)\n    - DFDC preview dataset paper based quality augmentation\n        - https://arxiv.org/abs/1910.08854\n    - Various image transforms such as rotate, equalize, solarize, posterize etc...\n    - Random horizontal flip\n- Imagewise\n    - Random Crop\n        - Randomness occures in margin and face-center translation\n    - Resize Interpolation\n        - we use three different resize interpolation : bilinear, bicubic, nearest\n\n\n## 4.4. Optimizer\nWe used same optimizer settings for both model.\n- SGD with momentum 0.9\n- Learning Rate Schedule\n    - 5 epochs of linear warm-up from 0 to 0.001\n    - Cosine annealing from 0.001 to 0 through 35 epochs (total 40 epochs)\n\n\n## 4.5. Further Regularization\n- weight decay: 0.0001   \n- dropout layer right before final fully connected layer: p=0.4\n- label smoothing epsilon: 0.01\n\n## 4.6. Validation losses of model 1 while training\n![Model 1 loss](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png)\noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model1.png\n\n## 4.7. Validation losses of model 2 while training\n![Model 2 loss](https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png)\noriginal image: https://raw.githubusercontent.com/jinhanpark/Deepfake-Detection-Challenge/master/assets/model2.png\n\n# 5. Inference Strategy\nWith our inference strategy below, we got 0.23086 in LB\n\n## 5.1. Sample Faces\n- Sample 20 frames per video uniformly.   \n- Detect and crop at most 3 faces from each frame with threshold stated above.\n\n\n## 5.2. Ensemble\nWe averaged predictions from two models\n\n## 5.3. Test Time Augmentation for each model\nFor each face detected, make 8 different crops of different margin and offset\n- 2 of 8 : two different margins of 1.1 and 1.15 without offsets.\n- 2 of 8 : two different y-axis offset of 0.15, -0.15 with margin of 1.15\n- 4 of 8 : horizontal flips of above 4 crops\n\n\n# 6. Outro\nThank you for reading. All questions are welcome.",
    "796209": "Great write-up. Happy to have been of assistance. 😄 ",
    "796335": "Thank you for sharing!\nI notice that you use a large face margin (2x). Will this margin parameter have a great impact on your model performance? We have tried a similar attention method with a small face margin 0.05, but the performance is poor.",
    "796187": "Nice work and congratulation! Could you share the video ids which clustered by actors? And did you compare the validation splitted by folder? There is large gap in my CV/LB. I guess maybe the better validation can figure out it.",
    "796161": "Congratulation on your position. Thank you for sharing, a very nice write up indeed. Wonder how do you get the idea of using self attention, it worked really well, wonderful!\nQuestion: what is really **minfacesize of (720/30)** ?",
    "796152": "Very impressive! Especially impressive LB score for single models, &lt;0.25 for Model 1.\n\nDo you think the self-attention model will generalise well or badly to deepfakes made with different methods from the 'wild' that will be in the final private LB?",
    "818457": "Biggest RIP?",
    "796566": "Are you guys planning to share the code? Amazing architecture!",
    "824037": "Attention model is defeated by deepfake! Actually, almost all top models are defeated with a 0.2x drop or more.",
    "823723": "At least you won a gold for this thread!",
    "814629": "\nÓtima redação. Parabéns",
    "812773": "Great work!",
    "811478": "good",
    "799219": "Thank you this detailed feedback and congratulations on your attention architecture!\nCan you share which tool/software did you use to generate those great and self explained figures?",
    "796512": "Amazing work!! Congrats and thanks for sharing.",
    "819536": "Nice job!! Gonna study your work.",
    "803209": "Thanks for sharing. It was useful.",
    "798708": "Thanks for sharing !"
  }
}