{
  "id": 146864,
  "title": "Our solution",
  "url": "/competitions/deepfake-detection-challenge/discussion/146864",
  "author_name": "deepware.ai",
  "post_date": "2020-04-28T18:14:35.333000",
  "votes": 19,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, we congratulate the winners of the Deepfake Detection competition and also would like to thank every competitor who worked hard to solve the deepfakes problem. We learned a lot throughout this amazing journey with you all, and we hope to see you again in the upcoming Deepfake challenges.</p>\n\n<p>Before intensely experimenting with deep neural network models, we decided to inspect the deepfake data with established computer vision techniques, to see if there's a simple cue which can reveal a deepfake, having these features in our mind for later feature-engineering purposes. Initially, we could observe indicators of deepfakes in frame inspections; such small artifacts were too hard to generalize for a novel solution. Although these experiments haven't resulted with great success, they provided a better understanding of deepfakes for us. Later on, we started experimenting with deep learnig based models, which then showed significant performance increase in detecting fake faces.</p>\n\n<h2>Face Extraction</h2>\n\n<p>For face extraction and processing purposes, we used FaceNet-Pytorch,</p>\n\n<ul>\n<li>Extracted landmarks and bounding boxes via MTCNN, then faces were cropped and saved in PNG format for high-quality image representation.</li>\n<li><p>Also, we used InceptionResNET features to differentiate individuals in the same frame from each other.</p>\n\n<h2>Data Preparation</h2></li>\n<li><p><strong>Margin</strong>: After square cropping faces with respect to bounding boxes, we’ve scaled each face with “1.2”, which is determined by experimentation.</p></li>\n<li><strong>Resize</strong>: Scaled faces resized to 224x224, with BICUBIC sampling.</li>\n<li><strong>Augmentations</strong>: We’ve experimented with “kornia” and “albumentations”.</li>\n<li>Used techniques such as Adaptive Histogram Equalization, Object Detection algorithms to increase detection robustness.</li>\n<li>Fine-tuned MTCNN for robust and rapid face detection.</li>\n<li>Used specific methods such as Compression, Edge-detection, Landmark Extraction, and Low-Quality to generate several datasets.</li>\n<li>Additionally extracted features such as Quality Metrics, Optical Flow, FFT, DoG, Glcm, Lbp</li>\n</ul>\n\n<h2>Augmentations and generalization</h2>\n\n<p>During our experiments, we have used different levels of augmentation techniques to increase robustness towards uncertainty in the wild. </p>\n\n<pre><code>def chain_aug(img):\n    return A.Compose([A.HorizontalFlip(p=1),\n                A.JpegCompression(quality_lower=70, quality_upper=100, p=1.),\n                A.Blur(blur_limit=7, p=1.),\n                A.IAASharpen(alpha=(0.2, 0.5), lightness=(0.5, 1.0), p=1.)\n                ], p=1)(image=img)['image']\n</code></pre>\n\n<hr>\n\n<pre><code>def light_aug(img):\n    return A.Compose([A.HorizontalFlip(p=1),\n                A.RandomBrightnessContrast(p=1),\n                A.RandomGamma(p=1),\n                A.CLAHE(p=1), ], p=1)(image=img)['image']\n</code></pre>\n\n<hr>\n\n<pre><code>def draw_ellipse(img):\n    img = Image.fromarray(img)\n    draw = ImageDraw.Draw(img)\n    w, h = img.size\n    cx, cy = w//2, h//2\n    x0 = np.random.randint(0, cx)\n    y0 = np.random.randint(0, cy)\n    x1 = np.random.randint(cx, w)\n    y1 = np.random.randint(cy, h)\n    draw.ellipse((x0, y0, x1, y1), fill='black')\n    return np.asarray(img)\n</code></pre>\n\n<h2>Models</h2>\n\n<p>Early studies preferred backbones such as <strong>ResNet</strong>, <strong>Xception</strong>, <strong>VGG</strong> whereas we have updated our backbone choice as <strong>EfficientNet</strong>. The main idea with the backbone network was to train a simple binary classifier and to discriminate between real and fake faces. Experiments for binary-classification models started from EfficientNet-b0 and scaled up to EfficientNet-b5. Here’s a couple of things we’ve tried.\n- Low res effnet-b3 + Effnet-b5 (224x224) 0.30656\n- Effnet-b4 0.38793 (224x224)\n- Effnet-b2 0.43068 (224x224)\n- Triplet loss models (224x224)\n- An ensemble of XGB, LGB, CATBoost, KNN models</p>\n\n<p>For final submissions:\n- An ensemble of EfficientNets: 0.31499 [<strong>FAILED</strong>] (Although scripts worked Kaggle submissions, as well as tests, sadly, our best performing model failed on private board scoring.)\n- Ensembles of EfficientNets + Other models mentioned above: 0.33552</p>\n\n<h2>Training schedule</h2>\n\n<p>In our development environment, we used Nvidia-docker, with a 4x 2080TI Lambda workstation.</p>\n\n<ul>\n<li>A balanced sampling of reals and fakes from the dataset.</li>\n<li>Batch size of 16 per GPU (4x), 625 iterations per epoch. </li>\n<li>12,500 iterations in total.</li>\n<li>Adam with weight_decay =10e-6</li>\n<li>OneCycleLR, with lr set to min 3e-6 and max 3e-4</li>\n<li>FP16 training</li>\n</ul>\n\n<h2>Validation</h2>\n\n<p>At first, like many competitors, we measured our performance with 400 public videos, after some time, we switched to folders 0-1-2. Later on, our team developed a validation split strategy based on similarity scores.  </p>\n\n<p>During the competition, we were tracking;\n- <strong>Log loss</strong> score for each class\n- Accuracy, confusion matrix scores\n- Error, gradient map analysis with custom debug scripts</p>\n\n<h2>Things that didn’t work</h2>\n\n<ul>\n<li>Experimented with masks with different models such as U-Net.</li>\n<li>SVM, Decision tree algorithms used EfficientNet bottleneck features as well as flatten face images but we didn’t see any improvement in results.</li>\n<li>LightGBM with Quality Metrics dataset.</li>\n<li>SVM with DFT, mentioned in Unmasking DeepFake with simple Features</li>\n<li>LightGBM with DoG + GLCM + LBP</li>\n<li>Temporal Model with LSTM, for capturing temporal inconsistency</li>\n<li>Optical flow features, however, the DFDC dataset was too noisy for robust optical flow estimations.</li>\n<li>We experimented with an  idea where we subtract frame n from n+1 </li>\n</ul>\n\n<p>Through the DFDC journey, we have had many challenges of generalizing our models to deepfakes in the wild. Training with only the DFDC dataset, our models resulted in high performance with \"face swap\" auto-encoder based generation techniques. Still, they failed to generalize other variations of deepfakes and generation techniques. We are very excited for future contributions, and we are happy to announce a deployed version of our solution at <a href=\"https://deepware.ai\">https://deepware.ai</a>. Once again, we would like to thank the amazing Kaggle community and especially the Facebook team and everyone who worked hard on the Deepfakes problem to provide novel solutions.</p>",
  "messages": [
    {
      "id": 825006,
      "postDate": "2020-04-28T18:14:35.333Z",
      "content": "<p>First of all, we congratulate the winners of the Deepfake Detection competition and also would like to thank every competitor who worked hard to solve the deepfakes problem. We learned a lot throughout this amazing journey with you all, and we hope to see you again in the upcoming Deepfake challenges.</p>\n\n<p>Before intensely experimenting with deep neural network models, we decided to inspect the deepfake data with established computer vision techniques, to see if there's a simple cue which can reveal a deepfake, having these features in our mind for later feature-engineering purposes. Initially, we could observe indicators of deepfakes in frame inspections; such small artifacts were too hard to generalize for a novel solution. Although these experiments haven't resulted with great success, they provided a better understanding of deepfakes for us. Later on, we started experimenting with deep learnig based models, which then showed significant performance increase in detecting fake faces.</p>\n\n<h2>Face Extraction</h2>\n\n<p>For face extraction and processing purposes, we used FaceNet-Pytorch,</p>\n\n<ul>\n<li>Extracted landmarks and bounding boxes via MTCNN, then faces were cropped and saved in PNG format for high-quality image representation.</li>\n<li><p>Also, we used InceptionResNET features to differentiate individuals in the same frame from each other.</p>\n\n<h2>Data Preparation</h2></li>\n<li><p><strong>Margin</strong>: After square cropping faces with respect to bounding boxes, we’ve scaled each face with “1.2”, which is determined by experimentation.</p></li>\n<li><strong>Resize</strong>: Scaled faces resized to 224x224, with BICUBIC sampling.</li>\n<li><strong>Augmentations</strong>: We’ve experimented with “kornia” and “albumentations”.</li>\n<li>Used techniques such as Adaptive Histogram Equalization, Object Detection algorithms to increase detection robustness.</li>\n<li>Fine-tuned MTCNN for robust and rapid face detection.</li>\n<li>Used specific methods such as Compression, Edge-detection, Landmark Extraction, and Low-Quality to generate several datasets.</li>\n<li>Additionally extracted features such as Quality Metrics, Optical Flow, FFT, DoG, Glcm, Lbp</li>\n</ul>\n\n<h2>Augmentations and generalization</h2>\n\n<p>During our experiments, we have used different levels of augmentation techniques to increase robustness towards uncertainty in the wild. </p>\n\n<pre><code>def chain_aug(img):\n    return A.Compose([A.HorizontalFlip(p=1),\n                A.JpegCompression(quality_lower=70, quality_upper=100, p=1.),\n                A.Blur(blur_limit=7, p=1.),\n                A.IAASharpen(alpha=(0.2, 0.5), lightness=(0.5, 1.0), p=1.)\n                ], p=1)(image=img)['image']\n</code></pre>\n\n<hr>\n\n<pre><code>def light_aug(img):\n    return A.Compose([A.HorizontalFlip(p=1),\n                A.RandomBrightnessContrast(p=1),\n                A.RandomGamma(p=1),\n                A.CLAHE(p=1), ], p=1)(image=img)['image']\n</code></pre>\n\n<hr>\n\n<pre><code>def draw_ellipse(img):\n    img = Image.fromarray(img)\n    draw = ImageDraw.Draw(img)\n    w, h = img.size\n    cx, cy = w//2, h//2\n    x0 = np.random.randint(0, cx)\n    y0 = np.random.randint(0, cy)\n    x1 = np.random.randint(cx, w)\n    y1 = np.random.randint(cy, h)\n    draw.ellipse((x0, y0, x1, y1), fill='black')\n    return np.asarray(img)\n</code></pre>\n\n<h2>Models</h2>\n\n<p>Early studies preferred backbones such as <strong>ResNet</strong>, <strong>Xception</strong>, <strong>VGG</strong> whereas we have updated our backbone choice as <strong>EfficientNet</strong>. The main idea with the backbone network was to train a simple binary classifier and to discriminate between real and fake faces. Experiments for binary-classification models started from EfficientNet-b0 and scaled up to EfficientNet-b5. Here’s a couple of things we’ve tried.\n- Low res effnet-b3 + Effnet-b5 (224x224) 0.30656\n- Effnet-b4 0.38793 (224x224)\n- Effnet-b2 0.43068 (224x224)\n- Triplet loss models (224x224)\n- An ensemble of XGB, LGB, CATBoost, KNN models</p>\n\n<p>For final submissions:\n- An ensemble of EfficientNets: 0.31499 [<strong>FAILED</strong>] (Although scripts worked Kaggle submissions, as well as tests, sadly, our best performing model failed on private board scoring.)\n- Ensembles of EfficientNets + Other models mentioned above: 0.33552</p>\n\n<h2>Training schedule</h2>\n\n<p>In our development environment, we used Nvidia-docker, with a 4x 2080TI Lambda workstation.</p>\n\n<ul>\n<li>A balanced sampling of reals and fakes from the dataset.</li>\n<li>Batch size of 16 per GPU (4x), 625 iterations per epoch. </li>\n<li>12,500 iterations in total.</li>\n<li>Adam with weight_decay =10e-6</li>\n<li>OneCycleLR, with lr set to min 3e-6 and max 3e-4</li>\n<li>FP16 training</li>\n</ul>\n\n<h2>Validation</h2>\n\n<p>At first, like many competitors, we measured our performance with 400 public videos, after some time, we switched to folders 0-1-2. Later on, our team developed a validation split strategy based on similarity scores.  </p>\n\n<p>During the competition, we were tracking;\n- <strong>Log loss</strong> score for each class\n- Accuracy, confusion matrix scores\n- Error, gradient map analysis with custom debug scripts</p>\n\n<h2>Things that didn’t work</h2>\n\n<ul>\n<li>Experimented with masks with different models such as U-Net.</li>\n<li>SVM, Decision tree algorithms used EfficientNet bottleneck features as well as flatten face images but we didn’t see any improvement in results.</li>\n<li>LightGBM with Quality Metrics dataset.</li>\n<li>SVM with DFT, mentioned in Unmasking DeepFake with simple Features</li>\n<li>LightGBM with DoG + GLCM + LBP</li>\n<li>Temporal Model with LSTM, for capturing temporal inconsistency</li>\n<li>Optical flow features, however, the DFDC dataset was too noisy for robust optical flow estimations.</li>\n<li>We experimented with an  idea where we subtract frame n from n+1 </li>\n</ul>\n\n<p>Through the DFDC journey, we have had many challenges of generalizing our models to deepfakes in the wild. Training with only the DFDC dataset, our models resulted in high performance with \"face swap\" auto-encoder based generation techniques. Still, they failed to generalize other variations of deepfakes and generation techniques. We are very excited for future contributions, and we are happy to announce a deployed version of our solution at <a href=\"https://deepware.ai\">https://deepware.ai</a>. Once again, we would like to thank the amazing Kaggle community and especially the Facebook team and everyone who worked hard on the Deepfakes problem to provide novel solutions.</p>",
      "rawMarkdown": "First of all, we congratulate the winners of the Deepfake Detection competition and also would like to thank every competitor who worked hard to solve the deepfakes problem. We learned a lot throughout this amazing journey with you all, and we hope to see you again in the upcoming Deepfake challenges.\n\nBefore intensely experimenting with deep neural network models, we decided to inspect the deepfake data with established computer vision techniques, to see if there's a simple cue which can reveal a deepfake, having these features in our mind for later feature-engineering purposes. Initially, we could observe indicators of deepfakes in frame inspections; such small artifacts were too hard to generalize for a novel solution. Although these experiments haven't resulted with great success, they provided a better understanding of deepfakes for us. Later on, we started experimenting with deep learnig based models, which then showed significant performance increase in detecting fake faces.\n## Face Extraction\nFor face extraction and processing purposes, we used FaceNet-Pytorch,\n\n- Extracted landmarks and bounding boxes via MTCNN, then faces were cropped and saved in PNG format for high-quality image representation.\n- Also, we used InceptionResNET features to differentiate individuals in the same frame from each other.\n## Data Preparation\n- **Margin**: After square cropping faces with respect to bounding boxes, we’ve scaled each face with “1.2”, which is determined by experimentation.\n- **Resize**: Scaled faces resized to 224x224, with BICUBIC sampling.\n- **Augmentations**: We’ve experimented with “kornia” and “albumentations”.\n- Used techniques such as Adaptive Histogram Equalization, Object Detection algorithms to increase detection robustness.\n- Fine-tuned MTCNN for robust and rapid face detection.\n- Used specific methods such as Compression, Edge-detection, Landmark Extraction, and Low-Quality to generate several datasets.\n- Additionally extracted features such as Quality Metrics, Optical Flow, FFT, DoG, Glcm, Lbp\n\n## Augmentations and generalization\nDuring our experiments, we have used different levels of augmentation techniques to increase robustness towards uncertainty in the wild. \n\n    def chain_aug(img):\n        return A.Compose([A.HorizontalFlip(p=1),\n                    A.JpegCompression(quality_lower=70, quality_upper=100, p=1.),\n                    A.Blur(blur_limit=7, p=1.),\n                    A.IAASharpen(alpha=(0.2, 0.5), lightness=(0.5, 1.0), p=1.)\n                    ], p=1)(image=img)['image']\n<hr>\n\n    def light_aug(img):\n        return A.Compose([A.HorizontalFlip(p=1),\n                    A.RandomBrightnessContrast(p=1),\n                    A.RandomGamma(p=1),\n                    A.CLAHE(p=1), ], p=1)(image=img)['image']\n<hr>\n\n    def draw_ellipse(img):\n        img = Image.fromarray(img)\n        draw = ImageDraw.Draw(img)\n        w, h = img.size\n        cx, cy = w//2, h//2\n        x0 = np.random.randint(0, cx)\n        y0 = np.random.randint(0, cy)\n        x1 = np.random.randint(cx, w)\n        y1 = np.random.randint(cy, h)\n        draw.ellipse((x0, y0, x1, y1), fill='black')\n        return np.asarray(img)\n\n## Models\nEarly studies preferred backbones such as **ResNet**, **Xception**, **VGG** whereas we have updated our backbone choice as **EfficientNet**. The main idea with the backbone network was to train a simple binary classifier and to discriminate between real and fake faces. Experiments for binary-classification models started from EfficientNet-b0 and scaled up to EfficientNet-b5. Here’s a couple of things we’ve tried.\n- Low res effnet-b3 + Effnet-b5 (224x224) 0.30656\n- Effnet-b4 0.38793 (224x224)\n- Effnet-b2 0.43068 (224x224)\n- Triplet loss models (224x224)\n- An ensemble of XGB, LGB, CATBoost, KNN models\n\nFor final submissions:\n- An ensemble of EfficientNets: 0.31499 [**FAILED**] (Although scripts worked Kaggle submissions, as well as tests, sadly, our best performing model failed on private board scoring.)\n- Ensembles of EfficientNets + Other models mentioned above: 0.33552\n## Training schedule\nIn our development environment, we used Nvidia-docker, with a 4x 2080TI Lambda workstation.\n\n- A balanced sampling of reals and fakes from the dataset.\n- Batch size of 16 per GPU (4x), 625 iterations per epoch. \n- 12,500 iterations in total.\n- Adam with weight_decay =10e-6\n- OneCycleLR, with lr set to min 3e-6 and max 3e-4\n- FP16 training\n\n## Validation\nAt first, like many competitors, we measured our performance with 400 public videos, after some time, we switched to folders 0-1-2. Later on, our team developed a validation split strategy based on similarity scores.  \n\nDuring the competition, we were tracking;\n- **Log loss** score for each class\n- Accuracy, confusion matrix scores\n- Error, gradient map analysis with custom debug scripts\n\n## Things that didn’t work\n- Experimented with masks with different models such as U-Net.\n- SVM, Decision tree algorithms used EfficientNet bottleneck features as well as flatten face images but we didn’t see any improvement in results.\n- LightGBM with Quality Metrics dataset.\n- SVM with DFT, mentioned in Unmasking DeepFake with simple Features\n- LightGBM with DoG + GLCM + LBP\n- Temporal Model with LSTM, for capturing temporal inconsistency\n- Optical flow features, however, the DFDC dataset was too noisy for robust optical flow estimations.\n- We experimented with an  idea where we subtract frame n from n+1 \n\n\nThrough the DFDC journey, we have had many challenges of generalizing our models to deepfakes in the wild. Training with only the DFDC dataset, our models resulted in high performance with \"face swap\" auto-encoder based generation techniques. Still, they failed to generalize other variations of deepfakes and generation techniques. We are very excited for future contributions, and we are happy to announce a deployed version of our solution at https://deepware.ai. Once again, we would like to thank the amazing Kaggle community and especially the Facebook team and everyone who worked hard on the Deepfakes problem to provide novel solutions.\n",
      "votes": 18
    },
    {
      "id": 826771,
      "postDate": "2020-04-29T21:11:59.350Z",
      "content": "<p><a href=\"/deepware\">@deepware</a> Thanks for sharing! Many things to learn from it.</p>",
      "rawMarkdown": "@deepware Thanks for sharing! Many things to learn from it.",
      "votes": 1
    },
    {
      "id": 827222,
      "postDate": "2020-04-30T06:33:44.210Z",
      "content": "<p>Thanks for sharing! Gonna test your agumentations. </p>",
      "rawMarkdown": "Thanks for sharing! Gonna test your agumentations. ",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 826771,
      "author_name": "Debanga Raj Neog",
      "author_url": "",
      "post_date": "2020-04-29T21:11:59.350000",
      "content": "<p><a href=\"/deepware\">@deepware</a> Thanks for sharing! Many things to learn from it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 827222,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-30T06:33:44.210000",
      "content": "<p>Thanks for sharing! Gonna test your agumentations. </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "825006": "First of all, we congratulate the winners of the Deepfake Detection competition and also would like to thank every competitor who worked hard to solve the deepfakes problem. We learned a lot throughout this amazing journey with you all, and we hope to see you again in the upcoming Deepfake challenges.\n\nBefore intensely experimenting with deep neural network models, we decided to inspect the deepfake data with established computer vision techniques, to see if there's a simple cue which can reveal a deepfake, having these features in our mind for later feature-engineering purposes. Initially, we could observe indicators of deepfakes in frame inspections; such small artifacts were too hard to generalize for a novel solution. Although these experiments haven't resulted with great success, they provided a better understanding of deepfakes for us. Later on, we started experimenting with deep learnig based models, which then showed significant performance increase in detecting fake faces.\n## Face Extraction\nFor face extraction and processing purposes, we used FaceNet-Pytorch,\n\n- Extracted landmarks and bounding boxes via MTCNN, then faces were cropped and saved in PNG format for high-quality image representation.\n- Also, we used InceptionResNET features to differentiate individuals in the same frame from each other.\n## Data Preparation\n- **Margin**: After square cropping faces with respect to bounding boxes, we’ve scaled each face with “1.2”, which is determined by experimentation.\n- **Resize**: Scaled faces resized to 224x224, with BICUBIC sampling.\n- **Augmentations**: We’ve experimented with “kornia” and “albumentations”.\n- Used techniques such as Adaptive Histogram Equalization, Object Detection algorithms to increase detection robustness.\n- Fine-tuned MTCNN for robust and rapid face detection.\n- Used specific methods such as Compression, Edge-detection, Landmark Extraction, and Low-Quality to generate several datasets.\n- Additionally extracted features such as Quality Metrics, Optical Flow, FFT, DoG, Glcm, Lbp\n\n## Augmentations and generalization\nDuring our experiments, we have used different levels of augmentation techniques to increase robustness towards uncertainty in the wild. \n\n    def chain_aug(img):\n        return A.Compose([A.HorizontalFlip(p=1),\n                    A.JpegCompression(quality_lower=70, quality_upper=100, p=1.),\n                    A.Blur(blur_limit=7, p=1.),\n                    A.IAASharpen(alpha=(0.2, 0.5), lightness=(0.5, 1.0), p=1.)\n                    ], p=1)(image=img)['image']\n<hr>\n\n    def light_aug(img):\n        return A.Compose([A.HorizontalFlip(p=1),\n                    A.RandomBrightnessContrast(p=1),\n                    A.RandomGamma(p=1),\n                    A.CLAHE(p=1), ], p=1)(image=img)['image']\n<hr>\n\n    def draw_ellipse(img):\n        img = Image.fromarray(img)\n        draw = ImageDraw.Draw(img)\n        w, h = img.size\n        cx, cy = w//2, h//2\n        x0 = np.random.randint(0, cx)\n        y0 = np.random.randint(0, cy)\n        x1 = np.random.randint(cx, w)\n        y1 = np.random.randint(cy, h)\n        draw.ellipse((x0, y0, x1, y1), fill='black')\n        return np.asarray(img)\n\n## Models\nEarly studies preferred backbones such as **ResNet**, **Xception**, **VGG** whereas we have updated our backbone choice as **EfficientNet**. The main idea with the backbone network was to train a simple binary classifier and to discriminate between real and fake faces. Experiments for binary-classification models started from EfficientNet-b0 and scaled up to EfficientNet-b5. Here’s a couple of things we’ve tried.\n- Low res effnet-b3 + Effnet-b5 (224x224) 0.30656\n- Effnet-b4 0.38793 (224x224)\n- Effnet-b2 0.43068 (224x224)\n- Triplet loss models (224x224)\n- An ensemble of XGB, LGB, CATBoost, KNN models\n\nFor final submissions:\n- An ensemble of EfficientNets: 0.31499 [**FAILED**] (Although scripts worked Kaggle submissions, as well as tests, sadly, our best performing model failed on private board scoring.)\n- Ensembles of EfficientNets + Other models mentioned above: 0.33552\n## Training schedule\nIn our development environment, we used Nvidia-docker, with a 4x 2080TI Lambda workstation.\n\n- A balanced sampling of reals and fakes from the dataset.\n- Batch size of 16 per GPU (4x), 625 iterations per epoch. \n- 12,500 iterations in total.\n- Adam with weight_decay =10e-6\n- OneCycleLR, with lr set to min 3e-6 and max 3e-4\n- FP16 training\n\n## Validation\nAt first, like many competitors, we measured our performance with 400 public videos, after some time, we switched to folders 0-1-2. Later on, our team developed a validation split strategy based on similarity scores.  \n\nDuring the competition, we were tracking;\n- **Log loss** score for each class\n- Accuracy, confusion matrix scores\n- Error, gradient map analysis with custom debug scripts\n\n## Things that didn’t work\n- Experimented with masks with different models such as U-Net.\n- SVM, Decision tree algorithms used EfficientNet bottleneck features as well as flatten face images but we didn’t see any improvement in results.\n- LightGBM with Quality Metrics dataset.\n- SVM with DFT, mentioned in Unmasking DeepFake with simple Features\n- LightGBM with DoG + GLCM + LBP\n- Temporal Model with LSTM, for capturing temporal inconsistency\n- Optical flow features, however, the DFDC dataset was too noisy for robust optical flow estimations.\n- We experimented with an  idea where we subtract frame n from n+1 \n\n\nThrough the DFDC journey, we have had many challenges of generalizing our models to deepfakes in the wild. Training with only the DFDC dataset, our models resulted in high performance with \"face swap\" auto-encoder based generation techniques. Still, they failed to generalize other variations of deepfakes and generation techniques. We are very excited for future contributions, and we are happy to announce a deployed version of our solution at https://deepware.ai. Once again, we would like to thank the amazing Kaggle community and especially the Facebook team and everyone who worked hard on the Deepfakes problem to provide novel solutions.\n",
    "826771": "@deepware Thanks for sharing! Many things to learn from it.",
    "827222": "Thanks for sharing! Gonna test your agumentations. "
  }
}