{
  "id": 145840,
  "title": "79st Place solution: A Hotchpotch of EfficientNets",
  "url": "/competitions/deepfake-detection-challenge/writeups/headhunters-79st-place-solution-a-hotchpotch-of-ef",
  "author_name": "",
  "post_date": "2020-05-02T11:01:45.697Z",
  "votes": 20,
  "comment_count": 14,
  "views": 0,
  "content": "<h3><strong>Introduction:</strong></h3>\n\n<p>Hi Kagglers! </p>\n\n<p>Congratulations to everyone who survived the leaderboard shakeup. Some of you probably expected that it would be hard to generalize to new videos and especially DeepFake videos “in the wild”. In essence, we tried to include as much model diversity in our ensemble as possible in order to survive a leaderboard shakeup.</p>\n\n<p>For us ( <a href=\"/koenbotermans\">@koenbotermans</a> , <a href=\"/kevindelnoye\">@kevindelnoye</a>  and I) this competition was 3 months of blood, sweat and tears with ups and downs and a radical switch from Keras to Pytorch halfway in the competition. I hope the insights in this post will be helpful to other competitors.</p>\n\n<p>As many other competitors we also use a frame-by-frame classification approach for detected faces. We agonized over adding the audio data in our model or doing sequence modeling (e.g. LSTM cells), but ended up not experimenting with it. This gave us a lot of time to focus on creating datasets and training classification models.</p>\n\n<p>The final kernel that we used for our submission can be found here:\n<a href=\"https://www.kaggle.com/carlolepelaars/efficientnet2xb5-b6200-b6finetuned-b4-2xres-01clip\">https://www.kaggle.com/carlolepelaars/efficientnet2xb5-b6200-b6finetuned-b4-2xres-01clip</a></p>\n\n<h3><strong>Data Preparation:</strong></h3>\n\n<p>Over the course of the competition we created several image datasets of face detections. The first baseline (ImagesFaces1) only contained the first face detection in a video using MTCNN as our face detector. Fine-tuning EfficientNetB5 on this dataset already gave us a pretty reasonable baseline (0.23083 local validation, 0.39680 on public leaderboard). Half-way we switched to RetinaFace as face detector as it had better performance and was faster. The first dataset after this switch (ImagesFaces3) contains 200K faces as training data. The final dataset (ImagesFaces6) featured 2M face detections from the training videos. Our final solution has both models that are trained on ImagesFaces3 and ImagesFaces6. It is arguable that 2M face detections is overkill, but luckily we had sufficient computational resources to experiment with it. In all datasets we took a 20% sample as validation data avoiding data leakage due to the limited number of actors. The datasets we created were balanced so we had 50% real and 50% fake faces.</p>\n\n<h3><strong>Face Detection:</strong></h3>\n\n<p>We started out with MTCNN as our face detector, but we thought it was too slow and that we could have better performance. Eventually we settled on RetinaFace as it had better performance and faster inference time.</p>\n\n<p>The MTCNN library we used is built on Tensorflow so naturally our classification models where based on Tensorflow/Keras. We did experiments using XCeptionNet and EfficientNetB0-B8.\nWhen we switched to RetinaFace we could only find a Pytorch implementation and as both Tensorflow and Pytorch allocate all GPU memory we felt forced to choose one framework for face detection and classification models. With 1.5 months left we therefore chose to go with RetinaFace and start experimenting with Pytorch models.</p>\n\n<h3><strong>Models and Ensembling strategy:</strong></h3>\n\n<p>Our final 81st place solution is a hotchpotch of EfficientNet Models. This haphazard combination was intentional to include diversity and to have a good chance at generalizing well to new data. The final models that were used:</p>\n\n<ol>\n<li>EfficientNetB6, 200x200 resolution, ImagesFaces6</li>\n<li>EfficientNetB5, 224x224 resolution, ImagesFaces3</li>\n<li>EfficientNetB5, 224x224 resolution, ImagesFaces6</li>\n<li>EfficientNetB6, 224x224 resolution, ImagesFaces6, finetuned for one epoch on our validation data.</li>\n<li>EfficientNetB4, 224x224 resolution, ImagesFaces6, data augmentation and label smoothing.</li>\n</ol>\n\n<p>All models are trained with a starting learning rate of 0.001 and using a “Reduce on Plateau” learning rate schedule with a patience of 2 and with a multiplier of 0.5. All models have a dropout of 0.4 before the final layer and group normalization so we don’t lose performance with small batch sizes (e.g. 20).</p>\n\n<p>We settled on a simple mean from all models for each video. We made 25 predictions per video on evenly spread frames. No weighting of models was done. The predictions were clipped between 0.01 and 0.99.</p>\n\n<p>In order to speed up the inference process we decreased the frame resolution by a factor of 2. Most original video frames were already high resolution so the face detector was still effective on a frame with a reduced resolution.</p>\n\n<h3><strong>Generalization:</strong></h3>\n\n<p>Throughout the competition we generally had a large correlation between our local validation scores and public leaderboard scores. A lower log loss on the local validation generally meant a better leaderboard score, but our final validation scores were extremely low (approx. 0.068 log loss). This means that we most likely had some data leakage in our training data, but we don’t know exactly where it was coming from.</p>\n\n<h3><strong>Data Augmentation:</strong></h3>\n\n<p>One thing I regret is that we underestimated the power of data augmentation for this competition. With only a month left we started using rotation (15 degrees) and flipping to augment the data, but looking at other public solutions we probably could have gotten a much higher score with more extensive data augmentation. </p>\n\n<h3><strong>What we would have liked to try but didn’t:</strong></h3>\n\n<ul>\n<li>LSTM cells</li>\n<li>Mixed precision (with NVIDIA DALI or apex).</li>\n<li>Include audio data and extract features using LSTM cells.</li>\n<li>Including ResNeXT101_wsl models. We trained a few models, but the weight files were too large to add it to the ensemble.</li>\n<li>Stochastic Weight Averaging (SWA).</li>\n<li>Fine-tuning “Noisy Student” weights (The Noisy Student paper was fairly new and not implemented in the libraries we used yet.)</li>\n<li>Training with Mixup.</li>\n<li>Taking the difference between frames and include it as a channel.</li>\n</ul>\n\n<h3><strong>What didn’t work:</strong></h3>\n\n<ul>\n<li>Taking the median of predictions for a video.</li>\n<li>Naïve postprocessing (Changing a prediction of 0.8 to 0.95, etc.)</li>\n<li>Clipping more than 0.01 (Sometimes there was a leaderboard improvement, but we decided it was too risky).</li>\n<li>Test Time Augmentation (TTA) with rotations. </li>\n</ul>\n\n<p>I hope you got some insights from this solution overview! Feel free to ask questions or give feedback on this approach!</p>",
  "messages": [
    {
      "id": "819595",
      "postDate": "04/24/2020 17:44:51",
      "content": "<h3><strong>Introduction:</strong></h3>\n\n<p>Hi Kagglers! </p>\n\n<p>Congratulations to everyone who survived the leaderboard shakeup. Some of you probably expected that it would be hard to generalize to new videos and especially DeepFake videos “in the wild”. In essence, we tried to include as much model diversity in our ensemble as possible in order to survive a leaderboard shakeup.</p>\n\n<p>For us ( <a href=\"/koenbotermans\">@koenbotermans</a> , <a href=\"/kevindelnoye\">@kevindelnoye</a>  and I) this competition was 3 months of blood, sweat and tears with ups and downs and a radical switch from Keras to Pytorch halfway in the competition. I hope the insights in this post will be helpful to other competitors.</p>\n\n<p>As many other competitors we also use a frame-by-frame classification approach for detected faces. We agonized over adding the audio data in our model or doing sequence modeling (e.g. LSTM cells), but ended up not experimenting with it. This gave us a lot of time to focus on creating datasets and training classification models.</p>\n\n<p>The final kernel that we used for our submission can be found here:\n<a href=\"https://www.kaggle.com/carlolepelaars/efficientnet2xb5-b6200-b6finetuned-b4-2xres-01clip\">https://www.kaggle.com/carlolepelaars/efficientnet2xb5-b6200-b6finetuned-b4-2xres-01clip</a></p>\n\n<h3><strong>Data Preparation:</strong></h3>\n\n<p>Over the course of the competition we created several image datasets of face detections. The first baseline (ImagesFaces1) only contained the first face detection in a video using MTCNN as our face detector. Fine-tuning EfficientNetB5 on this dataset already gave us a pretty reasonable baseline (0.23083 local validation, 0.39680 on public leaderboard). Half-way we switched to RetinaFace as face detector as it had better performance and was faster. The first dataset after this switch (ImagesFaces3) contains 200K faces as training data. The final dataset (ImagesFaces6) featured 2M face detections from the training videos. Our final solution has both models that are trained on ImagesFaces3 and ImagesFaces6. It is arguable that 2M face detections is overkill, but luckily we had sufficient computational resources to experiment with it. In all datasets we took a 20% sample as validation data avoiding data leakage due to the limited number of actors. The datasets we created were balanced so we had 50% real and 50% fake faces.</p>\n\n<h3><strong>Face Detection:</strong></h3>\n\n<p>We started out with MTCNN as our face detector, but we thought it was too slow and that we could have better performance. Eventually we settled on RetinaFace as it had better performance and faster inference time.</p>\n\n<p>The MTCNN library we used is built on Tensorflow so naturally our classification models where based on Tensorflow/Keras. We did experiments using XCeptionNet and EfficientNetB0-B8.\nWhen we switched to RetinaFace we could only find a Pytorch implementation and as both Tensorflow and Pytorch allocate all GPU memory we felt forced to choose one framework for face detection and classification models. With 1.5 months left we therefore chose to go with RetinaFace and start experimenting with Pytorch models.</p>\n\n<h3><strong>Models and Ensembling strategy:</strong></h3>\n\n<p>Our final 81st place solution is a hotchpotch of EfficientNet Models. This haphazard combination was intentional to include diversity and to have a good chance at generalizing well to new data. The final models that were used:</p>\n\n<ol>\n<li>EfficientNetB6, 200x200 resolution, ImagesFaces6</li>\n<li>EfficientNetB5, 224x224 resolution, ImagesFaces3</li>\n<li>EfficientNetB5, 224x224 resolution, ImagesFaces6</li>\n<li>EfficientNetB6, 224x224 resolution, ImagesFaces6, finetuned for one epoch on our validation data.</li>\n<li>EfficientNetB4, 224x224 resolution, ImagesFaces6, data augmentation and label smoothing.</li>\n</ol>\n\n<p>All models are trained with a starting learning rate of 0.001 and using a “Reduce on Plateau” learning rate schedule with a patience of 2 and with a multiplier of 0.5. All models have a dropout of 0.4 before the final layer and group normalization so we don’t lose performance with small batch sizes (e.g. 20).</p>\n\n<p>We settled on a simple mean from all models for each video. We made 25 predictions per video on evenly spread frames. No weighting of models was done. The predictions were clipped between 0.01 and 0.99.</p>\n\n<p>In order to speed up the inference process we decreased the frame resolution by a factor of 2. Most original video frames were already high resolution so the face detector was still effective on a frame with a reduced resolution.</p>\n\n<h3><strong>Generalization:</strong></h3>\n\n<p>Throughout the competition we generally had a large correlation between our local validation scores and public leaderboard scores. A lower log loss on the local validation generally meant a better leaderboard score, but our final validation scores were extremely low (approx. 0.068 log loss). This means that we most likely had some data leakage in our training data, but we don’t know exactly where it was coming from.</p>\n\n<h3><strong>Data Augmentation:</strong></h3>\n\n<p>One thing I regret is that we underestimated the power of data augmentation for this competition. With only a month left we started using rotation (15 degrees) and flipping to augment the data, but looking at other public solutions we probably could have gotten a much higher score with more extensive data augmentation. </p>\n\n<h3><strong>What we would have liked to try but didn’t:</strong></h3>\n\n<ul>\n<li>LSTM cells</li>\n<li>Mixed precision (with NVIDIA DALI or apex).</li>\n<li>Include audio data and extract features using LSTM cells.</li>\n<li>Including ResNeXT101_wsl models. We trained a few models, but the weight files were too large to add it to the ensemble.</li>\n<li>Stochastic Weight Averaging (SWA).</li>\n<li>Fine-tuning “Noisy Student” weights (The Noisy Student paper was fairly new and not implemented in the libraries we used yet.)</li>\n<li>Training with Mixup.</li>\n<li>Taking the difference between frames and include it as a channel.</li>\n</ul>\n\n<h3><strong>What didn’t work:</strong></h3>\n\n<ul>\n<li>Taking the median of predictions for a video.</li>\n<li>Naïve postprocessing (Changing a prediction of 0.8 to 0.95, etc.)</li>\n<li>Clipping more than 0.01 (Sometimes there was a leaderboard improvement, but we decided it was too risky).</li>\n<li>Test Time Augmentation (TTA) with rotations. </li>\n</ul>\n\n<p>I hope you got some insights from this solution overview! Feel free to ask questions or give feedback on this approach!</p>",
      "rawMarkdown": "### **Introduction:**\n\nHi Kagglers! \n\nCongratulations to everyone who survived the leaderboard shakeup. Some of you probably expected that it would be hard to generalize to new videos and especially DeepFake videos “in the wild”. In essence, we tried to include as much model diversity in our ensemble as possible in order to survive a leaderboard shakeup.\n\nFor us ( @koenbotermans , @kevindelnoye  and I) this competition was 3 months of blood, sweat and tears with ups and downs and a radical switch from Keras to Pytorch halfway in the competition. I hope the insights in this post will be helpful to other competitors.\n\nAs many other competitors we also use a frame-by-frame classification approach for detected faces. We agonized over adding the audio data in our model or doing sequence modeling (e.g. LSTM cells), but ended up not experimenting with it. This gave us a lot of time to focus on creating datasets and training classification models.\n\nThe final kernel that we used for our submission can be found here:\nhttps://www.kaggle.com/carlolepelaars/efficientnet2xb5-b6200-b6finetuned-b4-2xres-01clip\n\n### **Data Preparation:**\n\nOver the course of the competition we created several image datasets of face detections. The first baseline (ImagesFaces1) only contained the first face detection in a video using MTCNN as our face detector. Fine-tuning EfficientNetB5 on this dataset already gave us a pretty reasonable baseline (0.23083 local validation, 0.39680 on public leaderboard). Half-way we switched to RetinaFace as face detector as it had better performance and was faster. The first dataset after this switch (ImagesFaces3) contains 200K faces as training data. The final dataset (ImagesFaces6) featured 2M face detections from the training videos. Our final solution has both models that are trained on ImagesFaces3 and ImagesFaces6. It is arguable that 2M face detections is overkill, but luckily we had sufficient computational resources to experiment with it. In all datasets we took a 20% sample as validation data avoiding data leakage due to the limited number of actors. The datasets we created were balanced so we had 50% real and 50% fake faces.\n\n\n### **Face Detection:**\n\nWe started out with MTCNN as our face detector, but we thought it was too slow and that we could have better performance. Eventually we settled on RetinaFace as it had better performance and faster inference time.\n\nThe MTCNN library we used is built on Tensorflow so naturally our classification models where based on Tensorflow/Keras. We did experiments using XCeptionNet and EfficientNetB0-B8.\nWhen we switched to RetinaFace we could only find a Pytorch implementation and as both Tensorflow and Pytorch allocate all GPU memory we felt forced to choose one framework for face detection and classification models. With 1.5 months left we therefore chose to go with RetinaFace and start experimenting with Pytorch models.\n\n### **Models and Ensembling strategy:**\n\nOur final 81st place solution is a hotchpotch of EfficientNet Models. This haphazard combination was intentional to include diversity and to have a good chance at generalizing well to new data. The final models that were used:\n\n1. EfficientNetB6, 200x200 resolution, ImagesFaces6\n2. EfficientNetB5, 224x224 resolution, ImagesFaces3\n3. EfficientNetB5, 224x224 resolution, ImagesFaces6\n4. EfficientNetB6, 224x224 resolution, ImagesFaces6, finetuned for one epoch on our validation data.\n5. EfficientNetB4, 224x224 resolution, ImagesFaces6, data augmentation and label smoothing.\n\nAll models are trained with a starting learning rate of 0.001 and using a “Reduce on Plateau” learning rate schedule with a patience of 2 and with a multiplier of 0.5. All models have a dropout of 0.4 before the final layer and group normalization so we don’t lose performance with small batch sizes (e.g. 20).\n\nWe settled on a simple mean from all models for each video. We made 25 predictions per video on evenly spread frames. No weighting of models was done. The predictions were clipped between 0.01 and 0.99.\n\nIn order to speed up the inference process we decreased the frame resolution by a factor of 2. Most original video frames were already high resolution so the face detector was still effective on a frame with a reduced resolution.\n\n### **Generalization:**\n\nThroughout the competition we generally had a large correlation between our local validation scores and public leaderboard scores. A lower log loss on the local validation generally meant a better leaderboard score, but our final validation scores were extremely low (approx. 0.068 log loss). This means that we most likely had some data leakage in our training data, but we don’t know exactly where it was coming from.\n\n### **Data Augmentation:**\n\nOne thing I regret is that we underestimated the power of data augmentation for this competition. With only a month left we started using rotation (15 degrees) and flipping to augment the data, but looking at other public solutions we probably could have gotten a much higher score with more extensive data augmentation. \n\n### **What we would have liked to try but didn’t:**\n\n- LSTM cells\n- Mixed precision (with NVIDIA DALI or apex).\n- Include audio data and extract features using LSTM cells.\n- Including ResNeXT101_wsl models. We trained a few models, but the weight files were too large to add it to the ensemble.\n- Stochastic Weight Averaging (SWA).\n- Fine-tuning “Noisy Student” weights (The Noisy Student paper was fairly new and not implemented in the libraries we used yet.)\n- Training with Mixup.\n- Taking the difference between frames and include it as a channel.\n\n### **What didn’t work:**\n\n- Taking the median of predictions for a video.\n- Naïve postprocessing (Changing a prediction of 0.8 to 0.95, etc.)\n- Clipping more than 0.01 (Sometimes there was a leaderboard improvement, but we decided it was too risky).\n- Test Time Augmentation (TTA) with rotations. \n\nI hope you got some insights from this solution overview! Feel free to ask questions or give feedback on this approach!",
      "votes": null
    },
    {
      "id": "819629",
      "postDate": "04/24/2020 18:15:34",
      "content": "<p>thanks for sharing :)</p>",
      "rawMarkdown": "thanks for sharing :)",
      "votes": null
    },
    {
      "id": "819655",
      "postDate": "04/24/2020 18:50:09",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍",
      "votes": null
    },
    {
      "id": "820599",
      "postDate": "04/25/2020 15:02:49",
      "content": "<p>Thank you very much for insights.\nAnd warm congratulations for the nice achievement! Great job! </p>",
      "rawMarkdown": "Thank you very much for insights.\nAnd warm congratulations for the nice achievement! Great job!",
      "votes": null
    },
    {
      "id": "820772",
      "postDate": "04/25/2020 17:39:12",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "820979",
      "postDate": "04/25/2020 20:40:19",
      "content": "<p>Dear Carlo,\nFirst congratulations\nSecond thanks for sharing</p>",
      "rawMarkdown": "Dear Carlo,\nFirst congratulations\nSecond thanks for sharing",
      "votes": null
    },
    {
      "id": "821135",
      "postDate": "04/26/2020 00:00:25",
      "content": "<p>Thank you for the kind words!</p>",
      "rawMarkdown": "Thank you for the kind words!",
      "votes": null
    },
    {
      "id": "821703",
      "postDate": "04/26/2020 11:01:55",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "822309",
      "postDate": "04/26/2020 20:20:55",
      "content": "<p>Congratulations on your medal +1</p>",
      "rawMarkdown": "Congratulations on your medal +1",
      "votes": null
    },
    {
      "id": "822403",
      "postDate": "04/26/2020 21:55:37",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "823057",
      "postDate": "04/27/2020 11:19:53",
      "content": "<p>Congrats <a href=\"/carlolepelaars\">@carlolepelaars</a> .. </p>",
      "rawMarkdown": "Congrats @carlolepelaars ..",
      "votes": null
    },
    {
      "id": "823426",
      "postDate": "04/27/2020 16:19:02",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": null
    },
    {
      "id": "824620",
      "postDate": "04/28/2020 13:49:58",
      "content": "<p>How much margin did you extend for a detected face?</p>",
      "rawMarkdown": "How much margin did you extend for a detected face?",
      "votes": null
    },
    {
      "id": "830137",
      "postDate": "05/02/2020 11:01:04",
      "content": "<p>Hi, we used the raw bounding box predictions without margin. This has probably been a weakness in our approach.</p>",
      "rawMarkdown": "Hi, we used the raw bounding box predictions without margin. This has probably been a weakness in our approach.",
      "votes": null
    },
    {
      "id": "3483115",
      "postDate": "06/28/2026 13:15:10",
      "content": "<p>Hi Carlo! First of all, thank you for sharing such a detailed write-up. I really enjoyed reading about your approach to the competition.</p>\n<p>I have an AI project idea related to water resource management and intelligent monitoring that I've been working on. I think it has some exciting potential, and I'd love to get your thoughts on it.</p>\n<p>If you're open to it, would you be interested in discussing the idea or possibly exploring whether there's an opportunity to collaborate? I completely understand if you're busy, but I'd really appreciate your insights.</p>\n<p>Thanks again, and congratulations on the great work! <a href=\"https://www.kaggle.com/carlolepelaars\" target=\"_blank\">@carlolepelaars</a> </p>",
      "rawMarkdown": "Hi Carlo! First of all, thank you for sharing such a detailed write-up. I really enjoyed reading about your approach to the competition.\n\nI have an AI project idea related to water resource management and intelligent monitoring that I've been working on. I think it has some exciting potential, and I'd love to get your thoughts on it.\n\nIf you're open to it, would you be interested in discussing the idea or possibly exploring whether there's an opportunity to collaborate? I completely understand if you're busy, but I'd really appreciate your insights.\n\nThanks again, and congratulations on the great work! @carlolepelaars",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3483115,
      "author_name": "",
      "author_url": "",
      "post_date": "06/28/2026 13:15:10",
      "content": "<p>Hi Carlo! First of all, thank you for sharing such a detailed write-up. I really enjoyed reading about your approach to the competition.</p>\n<p>I have an AI project idea related to water resource management and intelligent monitoring that I've been working on. I think it has some exciting potential, and I'd love to get your thoughts on it.</p>\n<p>If you're open to it, would you be interested in discussing the idea or possibly exploring whether there's an opportunity to collaborate? I completely understand if you're busy, but I'd really appreciate your insights.</p>\n<p>Thanks again, and congratulations on the great work! <a href=\"https://www.kaggle.com/carlolepelaars\" target=\"_blank\">@carlolepelaars</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819629,
      "author_name": "albeffe",
      "author_url": "",
      "post_date": "04/24/2020 18:15:34",
      "content": "<p>thanks for sharing :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 819655,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "04/24/2020 18:50:09",
          "content": "<p>👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 820599,
      "author_name": "muhakabartay",
      "author_url": "",
      "post_date": "04/25/2020 15:02:49",
      "content": "<p>Thank you very much for insights.\nAnd warm congratulations for the nice achievement! Great job! </p>",
      "votes": null,
      "replies": [
        {
          "id": 820772,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "04/25/2020 17:39:12",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 820979,
      "author_name": "jafarib",
      "author_url": "",
      "post_date": "04/25/2020 20:40:19",
      "content": "<p>Dear Carlo,\nFirst congratulations\nSecond thanks for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 821135,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "04/26/2020 00:00:25",
          "content": "<p>Thank you for the kind words!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 821703,
      "author_name": "shivan118",
      "author_url": "",
      "post_date": "04/26/2020 11:01:55",
      "content": "<p>Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 822309,
      "author_name": "",
      "author_url": "",
      "post_date": "04/26/2020 20:20:55",
      "content": "<p>Congratulations on your medal +1</p>",
      "votes": null,
      "replies": [
        {
          "id": 822403,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "04/26/2020 21:55:37",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 823057,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "04/27/2020 11:19:53",
      "content": "<p>Congrats <a href=\"/carlolepelaars\">@carlolepelaars</a> .. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 823426,
      "author_name": "soham1024",
      "author_url": "",
      "post_date": "04/27/2020 16:19:02",
      "content": "<p>thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 824620,
      "author_name": "cookiecs",
      "author_url": "",
      "post_date": "04/28/2020 13:49:58",
      "content": "<p>How much margin did you extend for a detected face?</p>",
      "votes": null,
      "replies": [
        {
          "id": 830137,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "05/02/2020 11:01:04",
          "content": "<p>Hi, we used the raw bounding box predictions without margin. This has probably been a weakness in our approach.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "819595": "### **Introduction:**\n\nHi Kagglers! \n\nCongratulations to everyone who survived the leaderboard shakeup. Some of you probably expected that it would be hard to generalize to new videos and especially DeepFake videos “in the wild”. In essence, we tried to include as much model diversity in our ensemble as possible in order to survive a leaderboard shakeup.\n\nFor us ( @koenbotermans , @kevindelnoye  and I) this competition was 3 months of blood, sweat and tears with ups and downs and a radical switch from Keras to Pytorch halfway in the competition. I hope the insights in this post will be helpful to other competitors.\n\nAs many other competitors we also use a frame-by-frame classification approach for detected faces. We agonized over adding the audio data in our model or doing sequence modeling (e.g. LSTM cells), but ended up not experimenting with it. This gave us a lot of time to focus on creating datasets and training classification models.\n\nThe final kernel that we used for our submission can be found here:\nhttps://www.kaggle.com/carlolepelaars/efficientnet2xb5-b6200-b6finetuned-b4-2xres-01clip\n\n### **Data Preparation:**\n\nOver the course of the competition we created several image datasets of face detections. The first baseline (ImagesFaces1) only contained the first face detection in a video using MTCNN as our face detector. Fine-tuning EfficientNetB5 on this dataset already gave us a pretty reasonable baseline (0.23083 local validation, 0.39680 on public leaderboard). Half-way we switched to RetinaFace as face detector as it had better performance and was faster. The first dataset after this switch (ImagesFaces3) contains 200K faces as training data. The final dataset (ImagesFaces6) featured 2M face detections from the training videos. Our final solution has both models that are trained on ImagesFaces3 and ImagesFaces6. It is arguable that 2M face detections is overkill, but luckily we had sufficient computational resources to experiment with it. In all datasets we took a 20% sample as validation data avoiding data leakage due to the limited number of actors. The datasets we created were balanced so we had 50% real and 50% fake faces.\n\n\n### **Face Detection:**\n\nWe started out with MTCNN as our face detector, but we thought it was too slow and that we could have better performance. Eventually we settled on RetinaFace as it had better performance and faster inference time.\n\nThe MTCNN library we used is built on Tensorflow so naturally our classification models where based on Tensorflow/Keras. We did experiments using XCeptionNet and EfficientNetB0-B8.\nWhen we switched to RetinaFace we could only find a Pytorch implementation and as both Tensorflow and Pytorch allocate all GPU memory we felt forced to choose one framework for face detection and classification models. With 1.5 months left we therefore chose to go with RetinaFace and start experimenting with Pytorch models.\n\n### **Models and Ensembling strategy:**\n\nOur final 81st place solution is a hotchpotch of EfficientNet Models. This haphazard combination was intentional to include diversity and to have a good chance at generalizing well to new data. The final models that were used:\n\n1. EfficientNetB6, 200x200 resolution, ImagesFaces6\n2. EfficientNetB5, 224x224 resolution, ImagesFaces3\n3. EfficientNetB5, 224x224 resolution, ImagesFaces6\n4. EfficientNetB6, 224x224 resolution, ImagesFaces6, finetuned for one epoch on our validation data.\n5. EfficientNetB4, 224x224 resolution, ImagesFaces6, data augmentation and label smoothing.\n\nAll models are trained with a starting learning rate of 0.001 and using a “Reduce on Plateau” learning rate schedule with a patience of 2 and with a multiplier of 0.5. All models have a dropout of 0.4 before the final layer and group normalization so we don’t lose performance with small batch sizes (e.g. 20).\n\nWe settled on a simple mean from all models for each video. We made 25 predictions per video on evenly spread frames. No weighting of models was done. The predictions were clipped between 0.01 and 0.99.\n\nIn order to speed up the inference process we decreased the frame resolution by a factor of 2. Most original video frames were already high resolution so the face detector was still effective on a frame with a reduced resolution.\n\n### **Generalization:**\n\nThroughout the competition we generally had a large correlation between our local validation scores and public leaderboard scores. A lower log loss on the local validation generally meant a better leaderboard score, but our final validation scores were extremely low (approx. 0.068 log loss). This means that we most likely had some data leakage in our training data, but we don’t know exactly where it was coming from.\n\n### **Data Augmentation:**\n\nOne thing I regret is that we underestimated the power of data augmentation for this competition. With only a month left we started using rotation (15 degrees) and flipping to augment the data, but looking at other public solutions we probably could have gotten a much higher score with more extensive data augmentation. \n\n### **What we would have liked to try but didn’t:**\n\n- LSTM cells\n- Mixed precision (with NVIDIA DALI or apex).\n- Include audio data and extract features using LSTM cells.\n- Including ResNeXT101_wsl models. We trained a few models, but the weight files were too large to add it to the ensemble.\n- Stochastic Weight Averaging (SWA).\n- Fine-tuning “Noisy Student” weights (The Noisy Student paper was fairly new and not implemented in the libraries we used yet.)\n- Training with Mixup.\n- Taking the difference between frames and include it as a channel.\n\n### **What didn’t work:**\n\n- Taking the median of predictions for a video.\n- Naïve postprocessing (Changing a prediction of 0.8 to 0.95, etc.)\n- Clipping more than 0.01 (Sometimes there was a leaderboard improvement, but we decided it was too risky).\n- Test Time Augmentation (TTA) with rotations. \n\nI hope you got some insights from this solution overview! Feel free to ask questions or give feedback on this approach!",
    "819629": "thanks for sharing :)",
    "819655": "👍",
    "820599": "Thank you very much for insights.\nAnd warm congratulations for the nice achievement! Great job!",
    "820772": "Thank you!",
    "820979": "Dear Carlo,\nFirst congratulations\nSecond thanks for sharing",
    "821135": "Thank you for the kind words!",
    "821703": "Thanks for sharing :)",
    "822309": "Congratulations on your medal +1",
    "822403": "Thank you!",
    "823057": "Congrats @carlolepelaars ..",
    "823426": "thanks for sharing",
    "824620": "How much margin did you extend for a detected face?",
    "830137": "Hi, we used the raw bounding box predictions without margin. This has probably been a weakness in our approach.",
    "3483115": "Hi Carlo! First of all, thank you for sharing such a detailed write-up. I really enjoyed reading about your approach to the competition.\n\nI have an AI project idea related to water resource management and intelligent monitoring that I've been working on. I think it has some exciting potential, and I'd love to get your thoughts on it.\n\nIf you're open to it, would you be interested in discussing the idea or possibly exploring whether there's an opportunity to collaborate? I completely understand if you're busy, but I'd really appreciate your insights.\n\nThanks again, and congratulations on the great work! @carlolepelaars"
  },
  "source": "meta"
}