{
  "id": 145841,
  "title": "43rd place private LB solution",
  "url": "/competitions/deepfake-detection-challenge/writeups/ispl-43rd-place-private-lb-solution",
  "author_name": "",
  "post_date": "2020-04-24T17:45:41.169752Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Here is a brief description of the solution that led us to the 43rd place on the private LB!\nThe core idea is described in our paper \"Video Face Manipulation Detection Through Ensemble of CNNs\" available on <a href=\"https://arxiv.org/pdf/2004.07676.pdf\">arXiv</a>. To reproduce the paper results, please refer to our <a href=\"https://github.com/polimi-ispl/icpr2020dfdc\">GitHub repository</a>. Kaggle notebook for inference is also now available <a href=\"https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model\">here</a>.</p>\n\n<h1>Model</h1>\n\n<p>We started from EfficientNetB4 and we tweaked it a bit, by adding an attention mechanism in the middle of its convolutional blocks chain. We call it EfficientNetB4Att.</p>\n\n<h1>Training</h1>\n\n<p>We resorted to two different training paradigm, end-to-end and Siamese training with triplet loss, both considering frames as samples.</p>\n\n<p>We then trained 5 instances for each model in a 5-fold strategy based on folders. For each fold, we selected 40 consecutive folders for training and the remaning 10 for validation. There was no overlap between the folds.\nBy doing this we ended up with 10 models, 5 trained in a end-to-end fashion and 5 trained in a siamese fashion.</p>\n\n<p>During training, we considered only frames with one face, keeping the best face found by Blazeface.\nAs augmentation we used addictive noise, change of saturation, brightness, downscale and jpeg\ncompression. We trained using Adam for 20k iterations max (more details on this in the <a href=\"https://arxiv.org/pdf/2004.07676.pdf\">paper</a>).</p>\n\n<h1>Inference</h1>\n\n<p>At inference time, we considered 72 frames per video and looked at all the faces found by Blazeface, keeping only those with a score above a certain threshold. In case a frame had more than one face above the threshold, but with discordant scores, we took the maximum score above them. The rationale is that if we have multiple faces and just one face is fake, we want to classify the frame as fake. We then averaged the scores of all the networks, the scores of all the frame of the video and computed the sigmoid.\nYou can look into the inference code in details <a href=\"https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model\">here</a>.</p>\n\n<p>What a journey! Big big thanks to all my teammates from <a href=\"http://ispl.deib.polimi.it/\">Image and Sound Processing Lab (ISPL)</a> of Politecnico di Milano.\n@unklb197 @edoardodanielecannas @saramandelli @bestagini</p>",
  "messages": [
    {
      "id": "819597",
      "postDate": "04/24/2020 17:45:41",
      "content": "<p>Here is a brief description of the solution that led us to the 43rd place on the private LB!\nThe core idea is described in our paper \"Video Face Manipulation Detection Through Ensemble of CNNs\" available on <a href=\"https://arxiv.org/pdf/2004.07676.pdf\">arXiv</a>. To reproduce the paper results, please refer to our <a href=\"https://github.com/polimi-ispl/icpr2020dfdc\">GitHub repository</a>. Kaggle notebook for inference is also now available <a href=\"https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model\">here</a>.</p>\n\n<h1>Model</h1>\n\n<p>We started from EfficientNetB4 and we tweaked it a bit, by adding an attention mechanism in the middle of its convolutional blocks chain. We call it EfficientNetB4Att.</p>\n\n<h1>Training</h1>\n\n<p>We resorted to two different training paradigm, end-to-end and Siamese training with triplet loss, both considering frames as samples.</p>\n\n<p>We then trained 5 instances for each model in a 5-fold strategy based on folders. For each fold, we selected 40 consecutive folders for training and the remaning 10 for validation. There was no overlap between the folds.\nBy doing this we ended up with 10 models, 5 trained in a end-to-end fashion and 5 trained in a siamese fashion.</p>\n\n<p>During training, we considered only frames with one face, keeping the best face found by Blazeface.\nAs augmentation we used addictive noise, change of saturation, brightness, downscale and jpeg\ncompression. We trained using Adam for 20k iterations max (more details on this in the <a href=\"https://arxiv.org/pdf/2004.07676.pdf\">paper</a>).</p>\n\n<h1>Inference</h1>\n\n<p>At inference time, we considered 72 frames per video and looked at all the faces found by Blazeface, keeping only those with a score above a certain threshold. In case a frame had more than one face above the threshold, but with discordant scores, we took the maximum score above them. The rationale is that if we have multiple faces and just one face is fake, we want to classify the frame as fake. We then averaged the scores of all the networks, the scores of all the frame of the video and computed the sigmoid.\nYou can look into the inference code in details <a href=\"https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model\">here</a>.</p>\n\n<p>What a journey! Big big thanks to all my teammates from <a href=\"http://ispl.deib.polimi.it/\">Image and Sound Processing Lab (ISPL)</a> of Politecnico di Milano.\n@unklb197 @edoardodanielecannas @saramandelli @bestagini</p>",
      "rawMarkdown": "Here is a brief description of the solution that led us to the 43rd place on the private LB!\nThe core idea is described in our paper \"Video Face Manipulation Detection Through Ensemble of CNNs\" available on [arXiv](https://arxiv.org/pdf/2004.07676.pdf). To reproduce the paper results, please refer to our [GitHub repository](https://github.com/polimi-ispl/icpr2020dfdc). Kaggle notebook for inference is also now available [here](https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model).\n\n#Model\nWe started from EfficientNetB4 and we tweaked it a bit, by adding an attention mechanism in the middle of its convolutional blocks chain. We call it EfficientNetB4Att.\n\n#Training\nWe resorted to two different training paradigm, end-to-end and Siamese training with triplet loss, both considering frames as samples.\n\nWe then trained 5 instances for each model in a 5-fold strategy based on folders. For each fold, we selected 40 consecutive folders for training and the remaning 10 for validation. There was no overlap between the folds.\nBy doing this we ended up with 10 models, 5 trained in a end-to-end fashion and 5 trained in a siamese fashion.\n\nDuring training, we considered only frames with one face, keeping the best face found by Blazeface.\nAs augmentation we used addictive noise, change of saturation, brightness, downscale and jpeg\ncompression. We trained using Adam for 20k iterations max (more details on this in the [paper](https://arxiv.org/pdf/2004.07676.pdf)).\n\n\n#Inference\nAt inference time, we considered 72 frames per video and looked at all the faces found by Blazeface, keeping only those with a score above a certain threshold. In case a frame had more than one face above the threshold, but with discordant scores, we took the maximum score above them. The rationale is that if we have multiple faces and just one face is fake, we want to classify the frame as fake. We then averaged the scores of all the networks, the scores of all the frame of the video and computed the sigmoid.\nYou can look into the inference code in details [here](https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model).\n\n\nWhat a journey! Big big thanks to all my teammates from [Image and Sound Processing Lab (ISPL)](http://ispl.deib.polimi.it/) of Politecnico di Milano.\n@unklb197 @edoardodanielecannas @saramandelli @bestagini",
      "votes": null
    },
    {
      "id": "819689",
      "postDate": "04/24/2020 19:14:40",
      "content": "<p>Very nice that you wrote a full paper about your solution! Could you still use ImageNet weights for your architecture or did you have to train your model from scratch?</p>",
      "rawMarkdown": "Very nice that you wrote a full paper about your solution! Could you still use ImageNet weights for your architecture or did you have to train your model from scratch?",
      "votes": null
    },
    {
      "id": "819790",
      "postDate": "04/24/2020 21:38:31",
      "content": "<p>Thank you! We initialize the architecture with ImageNet weights and then we add the attention layer.</p>",
      "rawMarkdown": "Thank you! We initialize the architecture with ImageNet weights and then we add the attention layer.",
      "votes": null
    },
    {
      "id": "819823",
      "postDate": "04/24/2020 22:52:06",
      "content": "<p>Nice!</p>",
      "rawMarkdown": "Nice!",
      "votes": null
    },
    {
      "id": "819835",
      "postDate": "04/24/2020 23:11:26",
      "content": "<p>Are you sure that efficientnet  without SE blocks can't select area like eyes, nose, mouth (your fig.4 in article)? </p>",
      "rawMarkdown": "Are you sure that efficientnet  without SE blocks can't select area like eyes, nose, mouth (your fig.4 in article)?",
      "votes": null
    },
    {
      "id": "820864",
      "postDate": "04/25/2020 19:10:59",
      "content": "<p>That’s a good point, we don’t have any experimental evidence of what vanilla EfficientNet is considering. We’ll look into it, thank you.</p>",
      "rawMarkdown": "That’s a good point, we don’t have any experimental evidence of what vanilla EfficientNet is considering. We’ll look into it, thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 819689,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "04/24/2020 19:14:40",
      "content": "<p>Very nice that you wrote a full paper about your solution! Could you still use ImageNet weights for your architecture or did you have to train your model from scratch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 819790,
          "author_name": "nicobonne",
          "author_url": "",
          "post_date": "04/24/2020 21:38:31",
          "content": "<p>Thank you! We initialize the architecture with ImageNet weights and then we add the attention layer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 819823,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "04/24/2020 22:52:06",
          "content": "<p>Nice!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 819835,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "04/24/2020 23:11:26",
      "content": "<p>Are you sure that efficientnet  without SE blocks can't select area like eyes, nose, mouth (your fig.4 in article)? </p>",
      "votes": null,
      "replies": [
        {
          "id": 820864,
          "author_name": "nicobonne",
          "author_url": "",
          "post_date": "04/25/2020 19:10:59",
          "content": "<p>That’s a good point, we don’t have any experimental evidence of what vanilla EfficientNet is considering. We’ll look into it, thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "819597": "Here is a brief description of the solution that led us to the 43rd place on the private LB!\nThe core idea is described in our paper \"Video Face Manipulation Detection Through Ensemble of CNNs\" available on [arXiv](https://arxiv.org/pdf/2004.07676.pdf). To reproduce the paper results, please refer to our [GitHub repository](https://github.com/polimi-ispl/icpr2020dfdc). Kaggle notebook for inference is also now available [here](https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model).\n\n#Model\nWe started from EfficientNetB4 and we tweaked it a bit, by adding an attention mechanism in the middle of its convolutional blocks chain. We call it EfficientNetB4Att.\n\n#Training\nWe resorted to two different training paradigm, end-to-end and Siamese training with triplet loss, both considering frames as samples.\n\nWe then trained 5 instances for each model in a 5-fold strategy based on folders. For each fold, we selected 40 consecutive folders for training and the remaning 10 for validation. There was no overlap between the folds.\nBy doing this we ended up with 10 models, 5 trained in a end-to-end fashion and 5 trained in a siamese fashion.\n\nDuring training, we considered only frames with one face, keeping the best face found by Blazeface.\nAs augmentation we used addictive noise, change of saturation, brightness, downscale and jpeg\ncompression. We trained using Adam for 20k iterations max (more details on this in the [paper](https://arxiv.org/pdf/2004.07676.pdf)).\n\n\n#Inference\nAt inference time, we considered 72 frames per video and looked at all the faces found by Blazeface, keeping only those with a score above a certain threshold. In case a frame had more than one face above the threshold, but with discordant scores, we took the maximum score above them. The rationale is that if we have multiple faces and just one face is fake, we want to classify the frame as fake. We then averaged the scores of all the networks, the scores of all the frame of the video and computed the sigmoid.\nYou can look into the inference code in details [here](https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model).\n\n\nWhat a journey! Big big thanks to all my teammates from [Image and Sound Processing Lab (ISPL)](http://ispl.deib.polimi.it/) of Politecnico di Milano.\n@unklb197 @edoardodanielecannas @saramandelli @bestagini",
    "819689": "Very nice that you wrote a full paper about your solution! Could you still use ImageNet weights for your architecture or did you have to train your model from scratch?",
    "819790": "Thank you! We initialize the architecture with ImageNet weights and then we add the attention layer.",
    "819823": "Nice!",
    "819835": "Are you sure that efficientnet  without SE blocks can't select area like eyes, nose, mouth (your fig.4 in article)?",
    "820864": "That’s a good point, we don’t have any experimental evidence of what vanilla EfficientNet is considering. We’ll look into it, thank you."
  },
  "source": "meta"
}