{
  "id": 145965,
  "title": "27th place solution",
  "url": "/competitions/deepfake-detection-challenge/discussion/145965",
  "author_name": "Jan Bre",
  "post_date": "2020-04-25T09:58:51.968000",
  "votes": 18,
  "comment_count": 8,
  "views": 0,
  "content": "<h2>Approach</h2>\n\n<p>We used a simple frame-by-frame classification approach.</p>\n\n<h2>Face-Detector</h2>\n\n<p>We used simple MTCNN detector.\nInput size for the face detector was scaled down by 50% to significantly reduce inference time and memory usage without much loss in detection accuracy. We then used the bounding box to extract a square shaped face with image center coordinates and height being the same, but adjusted width.</p>\n\n<h2>Final Ensemble</h2>\n\n<p>Our final ensemble consists of B5/B6/B7 single model ensemble.\n <a href=\"/ratthachat\">@ratthachat</a> will give further details on model training in the comment section.</p>\n\n<h2>Models</h2>\n\n<p>We mostly used EfficientNets as we did not have much success with others Models.\n(1 Xception was used in the worse performing final essemble)\nWe played around with input sizes and sticked with 224*224 as most cropped faces in the training dataset were around that size.</p>\n\n<p>We started with B0 - B2 with some success. \nBut our final solutions of B5-B7 outperformed them.</p>\n\n<h2>Memory Usage and Inference Time</h2>\n\n<p>Due to memory constraints we used a batch generator that extracts and predicts only on two frames from a video at once. This also enabled us to use parallel inference to significantly reduce run-time and predict on every 27th frame in the end.</p>\n\n<h2>Things that may have worked</h2>\n\n<p>We considered organizers willl introduce organic videos and different deepfake methods in the test set.\nThats why we constantly evaluated our models on organic videos using Predictions and GradCam and used face crops without big margin, we thought this may generalise better. We used a margin of 14 pixels from the MTCNN face detector. </p>\n\n<p>The final prediction was a simple average of all predictions per video. We had better results splitting predictions per actor, but due to time reasons and as we considered organizers introducing &gt; 3 actors videos and multiple deep faked faces per video and did not want to run in these edge cases in the test set, we decided against the idea.</p>\n\n<p>Using more frames per video was superior to using TTA in our experiments.\nWe clipped prediction using 0.01 and 0.99.</p>\n\n<h2>Acknowledgments</h2>\n\n<p>Thank you to <a href=\"/ratthachat\">@ratthachat</a> <a href=\"/hmendonca\">@hmendonca</a> <a href=\"/hmmoghaddam\">@hmmoghaddam</a> for the incredible experience and Congratulations to <a href=\"/hmmoghaddam\">@hmmoghaddam</a> on becoming Kaggle expert. </p>\n\n<p>Thank you to the organizers for hosting this competition.</p>",
  "messages": [
    {
      "id": 820288,
      "postDate": "2020-04-25T09:58:51.967Z",
      "content": "<h2>Approach</h2>\n\n<p>We used a simple frame-by-frame classification approach.</p>\n\n<h2>Face-Detector</h2>\n\n<p>We used simple MTCNN detector.\nInput size for the face detector was scaled down by 50% to significantly reduce inference time and memory usage without much loss in detection accuracy. We then used the bounding box to extract a square shaped face with image center coordinates and height being the same, but adjusted width.</p>\n\n<h2>Final Ensemble</h2>\n\n<p>Our final ensemble consists of B5/B6/B7 single model ensemble.\n <a href=\"/ratthachat\">@ratthachat</a> will give further details on model training in the comment section.</p>\n\n<h2>Models</h2>\n\n<p>We mostly used EfficientNets as we did not have much success with others Models.\n(1 Xception was used in the worse performing final essemble)\nWe played around with input sizes and sticked with 224*224 as most cropped faces in the training dataset were around that size.</p>\n\n<p>We started with B0 - B2 with some success. \nBut our final solutions of B5-B7 outperformed them.</p>\n\n<h2>Memory Usage and Inference Time</h2>\n\n<p>Due to memory constraints we used a batch generator that extracts and predicts only on two frames from a video at once. This also enabled us to use parallel inference to significantly reduce run-time and predict on every 27th frame in the end.</p>\n\n<h2>Things that may have worked</h2>\n\n<p>We considered organizers willl introduce organic videos and different deepfake methods in the test set.\nThats why we constantly evaluated our models on organic videos using Predictions and GradCam and used face crops without big margin, we thought this may generalise better. We used a margin of 14 pixels from the MTCNN face detector. </p>\n\n<p>The final prediction was a simple average of all predictions per video. We had better results splitting predictions per actor, but due to time reasons and as we considered organizers introducing &gt; 3 actors videos and multiple deep faked faces per video and did not want to run in these edge cases in the test set, we decided against the idea.</p>\n\n<p>Using more frames per video was superior to using TTA in our experiments.\nWe clipped prediction using 0.01 and 0.99.</p>\n\n<h2>Acknowledgments</h2>\n\n<p>Thank you to <a href=\"/ratthachat\">@ratthachat</a> <a href=\"/hmendonca\">@hmendonca</a> <a href=\"/hmmoghaddam\">@hmmoghaddam</a> for the incredible experience and Congratulations to <a href=\"/hmmoghaddam\">@hmmoghaddam</a> on becoming Kaggle expert. </p>\n\n<p>Thank you to the organizers for hosting this competition.</p>",
      "rawMarkdown": "## Approach\n\nWe used a simple frame-by-frame classification approach.\n\n## Face-Detector\n\nWe used simple MTCNN detector.\nInput size for the face detector was scaled down by 50% to significantly reduce inference time and memory usage without much loss in detection accuracy. We then used the bounding box to extract a square shaped face with image center coordinates and height being the same, but adjusted width.\n\n## Final Ensemble\nOur final ensemble consists of B5/B6/B7 single model ensemble.\n @ratthachat will give further details on model training in the comment section.\n\n## Models\nWe mostly used EfficientNets as we did not have much success with others Models.\n(1 Xception was used in the worse performing final essemble)\nWe played around with input sizes and sticked with 224*224 as most cropped faces in the training dataset were around that size.\n\nWe started with B0 - B2 with some success. \nBut our final solutions of B5-B7 outperformed them.\n\n## Memory Usage and Inference Time\n\nDue to memory constraints we used a batch generator that extracts and predicts only on two frames from a video at once. This also enabled us to use parallel inference to significantly reduce run-time and predict on every 27th frame in the end.\n\n## Things that may have worked\n\nWe considered organizers willl introduce organic videos and different deepfake methods in the test set.\nThats why we constantly evaluated our models on organic videos using Predictions and GradCam and used face crops without big margin, we thought this may generalise better. We used a margin of 14 pixels from the MTCNN face detector. \n\nThe final prediction was a simple average of all predictions per video. We had better results splitting predictions per actor, but due to time reasons and as we considered organizers introducing &gt; 3 actors videos and multiple deep faked faces per video and did not want to run in these edge cases in the test set, we decided against the idea.\n\nUsing more frames per video was superior to using TTA in our experiments.\nWe clipped prediction using 0.01 and 0.99.\n\n## Acknowledgments\nThank you to @ratthachat @hmendonca @hmmoghaddam for the incredible experience and Congratulations to @hmmoghaddam on becoming Kaggle expert. \n\nThank you to the organizers for hosting this competition.",
      "votes": 18
    },
    {
      "id": 821129,
      "postDate": "2020-04-25T23:52:55.267Z",
      "content": "<p>Thanks for the organizers to host this competition, and thanks every participant for sharing valuable knowledge!! <br>\nI would like to than Jan <a href=\"/jpbremer\">@jpbremer</a> for inviting me to the team, and thanks Hamid and Henrique <br>\n<a href=\"/hmmoghaddam\">@hmmoghaddam</a> <a href=\"/hmendonca\">@hmendonca</a> for their hard works. I was lucky but these three guys are real experts. </p>\n\n<p>I would like to add some details of our team to Jan’s writeup. I think we divide the tasks nicely : </p>\n\n<h2>Big picture of our team</h2>\n\n<p><strong>Jan</strong> : Designed and wrote our team main training / inferencing pipelines, which were the center of our team. In the pipeline, Jan carefully designed to make sure that the twin pair of  fake-real were always in the same batch, so that it can challenge the model to differentiate between the two (similar ideas to Siamese net), but also make sure that \"non-sensible\" twin (where fake and real are almost the same due to ineffcient DeepFake) should not be input to the model.</p>\n\n<p><strong>Hamid</strong> : Professionally transferred all Jan’s code into TPU with improvement, so reduced training time 10-20 folds and so made us possible to utilize biggest models such as B5 - B7 within several hours of training. Details below.</p>\n\n<p><strong>Henrique</strong> : conducted extensive data exploration (e.g. see his <a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda\">public kernel</a> and <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140467\">discussions</a>), and made the best solid validation splits to our team just in time in the final phase . Please see details in the attached links.</p>\n\n<p><strong>Me</strong> : just gathered everything the three guys doing together </p>\n\n<p>Each of us also tried many different ideas e.g. incorporating time-domain, incorporating segmentation loss with face-only picture / full frame pictures, many valid external data, many kinds of face-detectors etc. but none of them improved from Jan’s base pipeline. I wrote my own 3-4 pipelines which ended up not using at all ...</p>\n\n<h2>Models</h2>\n\n<p>Around the end, Hamid came up with the a bit fancy EffNet which seems to work better than plain one (unfortunately, we don't have enough time to investigate the real differences).  I am sorry that I could not find original source/kernel of this inspiration for proper credits. Hamid also made a fix on face cropping to make it more accurate on non-square faces.</p>\n\n<p>We employed RectifiedAdam+Lookahead and used BCE with small label smoothing parameters 0.03 , found empirically . </p>\n\n<p>```</p>\n\n<p>def efficientAttention():\n    in_lay = Input(shape=(FLAGS['H'],FLAGS['W'],3))</p>\n\n<pre><code>base_model = efn.EfficientNetB5(\n    input_shape=( FLAGS['H'],FLAGS['W'], 3),\n    weights='noisy-student',\n    include_top=False\n )\n\npt_depth = base_model.get_output_shape_at(0)[-1]\npt_features = base_model(in_lay)\nbn_features = BatchNormalization()(pt_features)\n# here we do an attention mechanism to turn pixels in the GAP on an off\nattn_layer = Conv2D(64, kernel_size = (1,1), padding = 'same', activation = 'relu')(Dropout(0.5)(bn_features))\nattn_layer = Conv2D(16, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\nattn_layer = Conv2D(8, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\nattn_layer = Conv2D(1, kernel_size = (1,1), padding = 'valid', activation = 'sigmoid')(attn_layer)\n# fan it out to all of the channels\nup_c2_w = np.ones((1, 1, 1, pt_depth))\nup_c2 = Conv2D(pt_depth, kernel_size = (1,1), padding = 'same', \n               activation = 'linear', use_bias = False, weights = [up_c2_w])\nup_c2.trainable = False\nattn_layer = up_c2(attn_layer)\n\nmask_features = multiply([attn_layer, bn_features])\ngap_features = GlobalAveragePooling2D()(mask_features)\ngap_mask = GlobalAveragePooling2D()(attn_layer)\n# to account for missing values from the attention model\ngap = Lambda(lambda x: x[0]/x[1], name = 'RescaleGAP')([gap_features, gap_mask])\ngap_dr = Dropout(0.25)(gap)\ndr_steps = Dropout(0.25)(Dense(128, activation = 'relu')(gap_dr))\nout_layer = Dense(1, activation = 'sigmoid')(dr_steps)\nmodel = tf.keras.Model(inputs = [in_lay], outputs = [out_layer])\nreturn model\n</code></pre>\n\n<p>```</p>\n\n<p>B5 was our best single model, and I tried my best to incorporate B6 and B7 in the final ensemble within the time limit. Really thanks to <a href=\"/timesler\">@timesler</a> for his contribution on MTCNN and fast-MTCNN <a href=\"https://www.kaggle.com/timesler/fast-mtcnn-detector-55-fps-at-full-resolution\">public kernels</a>. Thanks <a href=\"/zaharch\">@zaharch</a>  for his brilliant <a href=\"https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata\">data analysis</a>, and finally would like to thank <a href=\"/humananalog\">@humananalog</a> for inspiring me on <a href=\"https://www.kaggle.com/humananalog/inference-demo\">inference codes</a> .</p>",
      "rawMarkdown": "Thanks for the organizers to host this competition, and thanks every participant for sharing valuable knowledge!!  \nI would like to than Jan @jpbremer for inviting me to the team, and thanks Hamid and Henrique  \n@hmmoghaddam @hmendonca for their hard works. I was lucky but these three guys are real experts. \n\nI would like to add some details of our team to Jan’s writeup. I think we divide the tasks nicely : \n\n## Big picture of our team\n\n**Jan** : Designed and wrote our team main training / inferencing pipelines, which were the center of our team. In the pipeline, Jan carefully designed to make sure that the twin pair of  fake-real were always in the same batch, so that it can challenge the model to differentiate between the two (similar ideas to Siamese net), but also make sure that \"non-sensible\" twin (where fake and real are almost the same due to ineffcient DeepFake) should not be input to the model.\n\n**Hamid** : Professionally transferred all Jan’s code into TPU with improvement, so reduced training time 10-20 folds and so made us possible to utilize biggest models such as B5 - B7 within several hours of training. Details below.\n\n**Henrique** : conducted extensive data exploration (e.g. see his [public kernel](https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda) and [discussions](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140467)), and made the best solid validation splits to our team just in time in the final phase . Please see details in the attached links.\n\n**Me** : just gathered everything the three guys doing together \n\nEach of us also tried many different ideas e.g. incorporating time-domain, incorporating segmentation loss with face-only picture / full frame pictures, many valid external data, many kinds of face-detectors etc. but none of them improved from Jan’s base pipeline. I wrote my own 3-4 pipelines which ended up not using at all ...\n\n## Models\nAround the end, Hamid came up with the a bit fancy EffNet which seems to work better than plain one (unfortunately, we don't have enough time to investigate the real differences).  I am sorry that I could not find original source/kernel of this inspiration for proper credits. Hamid also made a fix on face cropping to make it more accurate on non-square faces.\n\nWe employed RectifiedAdam+Lookahead and used BCE with small label smoothing parameters 0.03 , found empirically . \n\n```\n\ndef efficientAttention():\n    in_lay = Input(shape=(FLAGS['H'],FLAGS['W'],3))\n    \n    base_model = efn.EfficientNetB5(\n        input_shape=( FLAGS['H'],FLAGS['W'], 3),\n        weights='noisy-student',\n        include_top=False\n     )\n\n    pt_depth = base_model.get_output_shape_at(0)[-1]\n    pt_features = base_model(in_lay)\n    bn_features = BatchNormalization()(pt_features)\n    # here we do an attention mechanism to turn pixels in the GAP on an off\n    attn_layer = Conv2D(64, kernel_size = (1,1), padding = 'same', activation = 'relu')(Dropout(0.5)(bn_features))\n    attn_layer = Conv2D(16, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\n    attn_layer = Conv2D(8, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\n    attn_layer = Conv2D(1, kernel_size = (1,1), padding = 'valid', activation = 'sigmoid')(attn_layer)\n    # fan it out to all of the channels\n    up_c2_w = np.ones((1, 1, 1, pt_depth))\n    up_c2 = Conv2D(pt_depth, kernel_size = (1,1), padding = 'same', \n                   activation = 'linear', use_bias = False, weights = [up_c2_w])\n    up_c2.trainable = False\n    attn_layer = up_c2(attn_layer)\n\n    mask_features = multiply([attn_layer, bn_features])\n    gap_features = GlobalAveragePooling2D()(mask_features)\n    gap_mask = GlobalAveragePooling2D()(attn_layer)\n    # to account for missing values from the attention model\n    gap = Lambda(lambda x: x[0]/x[1], name = 'RescaleGAP')([gap_features, gap_mask])\n    gap_dr = Dropout(0.25)(gap)\n    dr_steps = Dropout(0.25)(Dense(128, activation = 'relu')(gap_dr))\n    out_layer = Dense(1, activation = 'sigmoid')(dr_steps)\n    model = tf.keras.Model(inputs = [in_lay], outputs = [out_layer])\n    return model\n\n```\n\n B5 was our best single model, and I tried my best to incorporate B6 and B7 in the final ensemble within the time limit. Really thanks to @timesler for his contribution on MTCNN and fast-MTCNN [public kernels](https://www.kaggle.com/timesler/fast-mtcnn-detector-55-fps-at-full-resolution). Thanks @zaharch  for his brilliant [data analysis](https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata), and finally would like to thank @humananalog for inspiring me on [inference codes](https://www.kaggle.com/humananalog/inference-demo) .\n",
      "votes": 6
    },
    {
      "id": 873179,
      "postDate": "2020-06-03T21:57:34.087Z",
      "content": "<p>Hey! Congratulations for the 27th position! </p>\n\n<p>Could you please explain what do you mean by \"organic videos\"?  </p>\n\n<p>Tagging my old friend <a href=\"/ratthachat\">@ratthachat</a> for response :)</p>",
      "rawMarkdown": "Hey! Congratulations for the 27th position! \n\nCould you please explain what do you mean by \"organic videos\"?  \n\nTagging my old friend @ratthachat for response :)",
      "replies": [
        {
          "id": 877770,
          "postDate": "2020-06-08T00:36:21.620Z",
          "content": "<p>Hope you are doing well <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> !\nBy organic videos , I think Jan just meant \"real video without any modifications\" ... <a href=\"/jpbremer\">@jpbremer</a> </p>",
          "rawMarkdown": "Hope you are doing well @rishabhiitbhu !\nBy organic videos , I think Jan just meant \"real video without any modifications\" ... @jpbremer ",
          "votes": 1
        }
      ]
    },
    {
      "id": 828427,
      "postDate": "2020-05-01T02:29:57.450Z",
      "content": "<p>Thanks for the detailed explaination!! You are very generous!!</p>",
      "rawMarkdown": "Thanks for the detailed explaination!! You are very generous!!",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 825581,
      "postDate": "2020-04-29T05:44:18.240Z",
      "content": "<p>Thanks for sharing. Could you clarifying some questions.\n1) What is frame-by-frame classification? I know many people talked about this, but is this about predict every frame? I saw many was predicting specific frames and get mean or median\n2) Could you share your albumentation code since I found this quiet influencial. \nThanks</p>",
      "rawMarkdown": "Thanks for sharing. Could you clarifying some questions.\n1) What is frame-by-frame classification? I know many people talked about this, but is this about predict every frame? I saw many was predicting specific frames and get mean or median\n2) Could you share your albumentation code since I found this quiet influencial. \nThanks",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 827690,
          "postDate": "2020-04-30T13:15:01.517Z",
          "content": "<p>Hi <a href=\"/muerbingsha\">@muerbingsha</a> , \n1) yes frame-by-frame classification is just simply classifiy like independent images and just average the predictions -- although sound not so good, it's the most effective in our experiments</p>\n\n<p>2) since our final is TPU code, we didn't use albumentation, but rather tf primitive codes like this (credit :  Hamid)\n```\ndef data_augment(self,image, label, seed=2020):</p>\n\n<pre><code>    image = tf.cast(image, tf.float32)\n    #Gaussian blur\n    if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['blur_prob']:\n        image = self.augmentor.gaussian_blur(image, self.AugParams['blur_ksize'], self.AugParams['blur_sigma'])\n\n    #Random block out\n    if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['blockout_prob']:\n        image = self.augmentor.random_blockout(image, self.AugParams['blockout_sl'], self.AugParams['blockout_sh'], self.AugParams['blockout_rl'])\n\n    #Random scale\n    if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['scale_prob']:\n        if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; 0.5:\n            image = self.augmentor.zoom_in(image, self.AugParams['scale_factor'])\n        else:\n            image = self.augmentor.zoom_out(image, self.AugParams['scale_factor'])\n    #Random rotate\n    if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['rot_prob']:\n        angle = tf.random.uniform(shape=[], minval=-self.AugParams['rot_range'], maxval=self.AugParams['rot_range'], dtype=tf.int32)\n        image = self.augmentor.image_rotate(image,angle)\n     return tf.cast(image, tf.uint8), label \n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "Hi @muerbingsha , \n1) yes frame-by-frame classification is just simply classifiy like independent images and just average the predictions -- although sound not so good, it's the most effective in our experiments\n\n2) since our final is TPU code, we didn't use albumentation, but rather tf primitive codes like this (credit :  Hamid)\n```\ndef data_augment(self,image, label, seed=2020):\n\n        image = tf.cast(image, tf.float32)\n        #Gaussian blur\n        if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['blur_prob']:\n            image = self.augmentor.gaussian_blur(image, self.AugParams['blur_ksize'], self.AugParams['blur_sigma'])\n\n        #Random block out\n        if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['blockout_prob']:\n            image = self.augmentor.random_blockout(image, self.AugParams['blockout_sl'], self.AugParams['blockout_sh'], self.AugParams['blockout_rl'])\n\n        #Random scale\n        if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['scale_prob']:\n            if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; 0.5:\n                image = self.augmentor.zoom_in(image, self.AugParams['scale_factor'])\n            else:\n                image = self.augmentor.zoom_out(image, self.AugParams['scale_factor'])\n        #Random rotate\n        if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['rot_prob']:\n            angle = tf.random.uniform(shape=[], minval=-self.AugParams['rot_range'], maxval=self.AugParams['rot_range'], dtype=tf.int32)\n            image = self.augmentor.image_rotate(image,angle)\n         return tf.cast(image, tf.uint8), label \n```",
          "votes": 2
        }
      ]
    },
    {
      "id": 837906,
      "postDate": "2020-05-08T06:16:34.847Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing\n"
    },
    {
      "id": 832998,
      "postDate": "2020-05-04T14:43:39.070Z",
      "content": "<p>Congrats and Thanks for sharing.</p>",
      "rawMarkdown": "Congrats and Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 821129,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2020-04-25T23:52:55.267000",
      "content": "<p>Thanks for the organizers to host this competition, and thanks every participant for sharing valuable knowledge!! <br>\nI would like to than Jan <a href=\"/jpbremer\">@jpbremer</a> for inviting me to the team, and thanks Hamid and Henrique <br>\n<a href=\"/hmmoghaddam\">@hmmoghaddam</a> <a href=\"/hmendonca\">@hmendonca</a> for their hard works. I was lucky but these three guys are real experts. </p>\n\n<p>I would like to add some details of our team to Jan’s writeup. I think we divide the tasks nicely : </p>\n\n<h2>Big picture of our team</h2>\n\n<p><strong>Jan</strong> : Designed and wrote our team main training / inferencing pipelines, which were the center of our team. In the pipeline, Jan carefully designed to make sure that the twin pair of  fake-real were always in the same batch, so that it can challenge the model to differentiate between the two (similar ideas to Siamese net), but also make sure that \"non-sensible\" twin (where fake and real are almost the same due to ineffcient DeepFake) should not be input to the model.</p>\n\n<p><strong>Hamid</strong> : Professionally transferred all Jan’s code into TPU with improvement, so reduced training time 10-20 folds and so made us possible to utilize biggest models such as B5 - B7 within several hours of training. Details below.</p>\n\n<p><strong>Henrique</strong> : conducted extensive data exploration (e.g. see his <a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda\">public kernel</a> and <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140467\">discussions</a>), and made the best solid validation splits to our team just in time in the final phase . Please see details in the attached links.</p>\n\n<p><strong>Me</strong> : just gathered everything the three guys doing together </p>\n\n<p>Each of us also tried many different ideas e.g. incorporating time-domain, incorporating segmentation loss with face-only picture / full frame pictures, many valid external data, many kinds of face-detectors etc. but none of them improved from Jan’s base pipeline. I wrote my own 3-4 pipelines which ended up not using at all ...</p>\n\n<h2>Models</h2>\n\n<p>Around the end, Hamid came up with the a bit fancy EffNet which seems to work better than plain one (unfortunately, we don't have enough time to investigate the real differences).  I am sorry that I could not find original source/kernel of this inspiration for proper credits. Hamid also made a fix on face cropping to make it more accurate on non-square faces.</p>\n\n<p>We employed RectifiedAdam+Lookahead and used BCE with small label smoothing parameters 0.03 , found empirically . </p>\n\n<p>```</p>\n\n<p>def efficientAttention():\n    in_lay = Input(shape=(FLAGS['H'],FLAGS['W'],3))</p>\n\n<pre><code>base_model = efn.EfficientNetB5(\n    input_shape=( FLAGS['H'],FLAGS['W'], 3),\n    weights='noisy-student',\n    include_top=False\n )\n\npt_depth = base_model.get_output_shape_at(0)[-1]\npt_features = base_model(in_lay)\nbn_features = BatchNormalization()(pt_features)\n# here we do an attention mechanism to turn pixels in the GAP on an off\nattn_layer = Conv2D(64, kernel_size = (1,1), padding = 'same', activation = 'relu')(Dropout(0.5)(bn_features))\nattn_layer = Conv2D(16, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\nattn_layer = Conv2D(8, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\nattn_layer = Conv2D(1, kernel_size = (1,1), padding = 'valid', activation = 'sigmoid')(attn_layer)\n# fan it out to all of the channels\nup_c2_w = np.ones((1, 1, 1, pt_depth))\nup_c2 = Conv2D(pt_depth, kernel_size = (1,1), padding = 'same', \n               activation = 'linear', use_bias = False, weights = [up_c2_w])\nup_c2.trainable = False\nattn_layer = up_c2(attn_layer)\n\nmask_features = multiply([attn_layer, bn_features])\ngap_features = GlobalAveragePooling2D()(mask_features)\ngap_mask = GlobalAveragePooling2D()(attn_layer)\n# to account for missing values from the attention model\ngap = Lambda(lambda x: x[0]/x[1], name = 'RescaleGAP')([gap_features, gap_mask])\ngap_dr = Dropout(0.25)(gap)\ndr_steps = Dropout(0.25)(Dense(128, activation = 'relu')(gap_dr))\nout_layer = Dense(1, activation = 'sigmoid')(dr_steps)\nmodel = tf.keras.Model(inputs = [in_lay], outputs = [out_layer])\nreturn model\n</code></pre>\n\n<p>```</p>\n\n<p>B5 was our best single model, and I tried my best to incorporate B6 and B7 in the final ensemble within the time limit. Really thanks to <a href=\"/timesler\">@timesler</a> for his contribution on MTCNN and fast-MTCNN <a href=\"https://www.kaggle.com/timesler/fast-mtcnn-detector-55-fps-at-full-resolution\">public kernels</a>. Thanks <a href=\"/zaharch\">@zaharch</a>  for his brilliant <a href=\"https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata\">data analysis</a>, and finally would like to thank <a href=\"/humananalog\">@humananalog</a> for inspiring me on <a href=\"https://www.kaggle.com/humananalog/inference-demo\">inference codes</a> .</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 873179,
      "author_name": "Rishabh Agrahari",
      "author_url": "",
      "post_date": "2020-06-03T21:57:34.087000",
      "content": "<p>Hey! Congratulations for the 27th position! </p>\n\n<p>Could you please explain what do you mean by \"organic videos\"?  </p>\n\n<p>Tagging my old friend <a href=\"/ratthachat\">@ratthachat</a> for response :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 877770,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2020-06-08T00:36:21.620000",
          "content": "<p>Hope you are doing well <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> !\nBy organic videos , I think Jan just meant \"real video without any modifications\" ... <a href=\"/jpbremer\">@jpbremer</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 828427,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-01T02:29:57.450000",
      "content": "<p>Thanks for the detailed explaination!! You are very generous!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 825581,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-29T05:44:18.240000",
      "content": "<p>Thanks for sharing. Could you clarifying some questions.\n1) What is frame-by-frame classification? I know many people talked about this, but is this about predict every frame? I saw many was predicting specific frames and get mean or median\n2) Could you share your albumentation code since I found this quiet influencial. \nThanks</p>",
      "votes": 2,
      "replies": [
        {
          "id": 827690,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2020-04-30T13:15:01.517000",
          "content": "<p>Hi <a href=\"/muerbingsha\">@muerbingsha</a> , \n1) yes frame-by-frame classification is just simply classifiy like independent images and just average the predictions -- although sound not so good, it's the most effective in our experiments</p>\n\n<p>2) since our final is TPU code, we didn't use albumentation, but rather tf primitive codes like this (credit :  Hamid)\n```\ndef data_augment(self,image, label, seed=2020):</p>\n\n<pre><code>    image = tf.cast(image, tf.float32)\n    #Gaussian blur\n    if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['blur_prob']:\n        image = self.augmentor.gaussian_blur(image, self.AugParams['blur_ksize'], self.AugParams['blur_sigma'])\n\n    #Random block out\n    if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['blockout_prob']:\n        image = self.augmentor.random_blockout(image, self.AugParams['blockout_sl'], self.AugParams['blockout_sh'], self.AugParams['blockout_rl'])\n\n    #Random scale\n    if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['scale_prob']:\n        if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; 0.5:\n            image = self.augmentor.zoom_in(image, self.AugParams['scale_factor'])\n        else:\n            image = self.augmentor.zoom_out(image, self.AugParams['scale_factor'])\n    #Random rotate\n    if tf.random.uniform(shape=[], minval=0.0, maxval=1.0) &gt; self.AugParams['rot_prob']:\n        angle = tf.random.uniform(shape=[], minval=-self.AugParams['rot_range'], maxval=self.AugParams['rot_range'], dtype=tf.int32)\n        image = self.augmentor.image_rotate(image,angle)\n     return tf.cast(image, tf.uint8), label \n</code></pre>\n\n<p>```</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 837906,
      "author_name": "sravan kothacheruvu",
      "author_url": "",
      "post_date": "2020-05-08T06:16:34.847000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 832998,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-04T14:43:39.070000",
      "content": "<p>Congrats and Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "820288": "## Approach\n\nWe used a simple frame-by-frame classification approach.\n\n## Face-Detector\n\nWe used simple MTCNN detector.\nInput size for the face detector was scaled down by 50% to significantly reduce inference time and memory usage without much loss in detection accuracy. We then used the bounding box to extract a square shaped face with image center coordinates and height being the same, but adjusted width.\n\n## Final Ensemble\nOur final ensemble consists of B5/B6/B7 single model ensemble.\n @ratthachat will give further details on model training in the comment section.\n\n## Models\nWe mostly used EfficientNets as we did not have much success with others Models.\n(1 Xception was used in the worse performing final essemble)\nWe played around with input sizes and sticked with 224*224 as most cropped faces in the training dataset were around that size.\n\nWe started with B0 - B2 with some success. \nBut our final solutions of B5-B7 outperformed them.\n\n## Memory Usage and Inference Time\n\nDue to memory constraints we used a batch generator that extracts and predicts only on two frames from a video at once. This also enabled us to use parallel inference to significantly reduce run-time and predict on every 27th frame in the end.\n\n## Things that may have worked\n\nWe considered organizers willl introduce organic videos and different deepfake methods in the test set.\nThats why we constantly evaluated our models on organic videos using Predictions and GradCam and used face crops without big margin, we thought this may generalise better. We used a margin of 14 pixels from the MTCNN face detector. \n\nThe final prediction was a simple average of all predictions per video. We had better results splitting predictions per actor, but due to time reasons and as we considered organizers introducing &gt; 3 actors videos and multiple deep faked faces per video and did not want to run in these edge cases in the test set, we decided against the idea.\n\nUsing more frames per video was superior to using TTA in our experiments.\nWe clipped prediction using 0.01 and 0.99.\n\n## Acknowledgments\nThank you to @ratthachat @hmendonca @hmmoghaddam for the incredible experience and Congratulations to @hmmoghaddam on becoming Kaggle expert. \n\nThank you to the organizers for hosting this competition.",
    "821129": "Thanks for the organizers to host this competition, and thanks every participant for sharing valuable knowledge!!  \nI would like to than Jan @jpbremer for inviting me to the team, and thanks Hamid and Henrique  \n@hmmoghaddam @hmendonca for their hard works. I was lucky but these three guys are real experts. \n\nI would like to add some details of our team to Jan’s writeup. I think we divide the tasks nicely : \n\n## Big picture of our team\n\n**Jan** : Designed and wrote our team main training / inferencing pipelines, which were the center of our team. In the pipeline, Jan carefully designed to make sure that the twin pair of  fake-real were always in the same batch, so that it can challenge the model to differentiate between the two (similar ideas to Siamese net), but also make sure that \"non-sensible\" twin (where fake and real are almost the same due to ineffcient DeepFake) should not be input to the model.\n\n**Hamid** : Professionally transferred all Jan’s code into TPU with improvement, so reduced training time 10-20 folds and so made us possible to utilize biggest models such as B5 - B7 within several hours of training. Details below.\n\n**Henrique** : conducted extensive data exploration (e.g. see his [public kernel](https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda) and [discussions](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140467)), and made the best solid validation splits to our team just in time in the final phase . Please see details in the attached links.\n\n**Me** : just gathered everything the three guys doing together \n\nEach of us also tried many different ideas e.g. incorporating time-domain, incorporating segmentation loss with face-only picture / full frame pictures, many valid external data, many kinds of face-detectors etc. but none of them improved from Jan’s base pipeline. I wrote my own 3-4 pipelines which ended up not using at all ...\n\n## Models\nAround the end, Hamid came up with the a bit fancy EffNet which seems to work better than plain one (unfortunately, we don't have enough time to investigate the real differences).  I am sorry that I could not find original source/kernel of this inspiration for proper credits. Hamid also made a fix on face cropping to make it more accurate on non-square faces.\n\nWe employed RectifiedAdam+Lookahead and used BCE with small label smoothing parameters 0.03 , found empirically . \n\n```\n\ndef efficientAttention():\n    in_lay = Input(shape=(FLAGS['H'],FLAGS['W'],3))\n    \n    base_model = efn.EfficientNetB5(\n        input_shape=( FLAGS['H'],FLAGS['W'], 3),\n        weights='noisy-student',\n        include_top=False\n     )\n\n    pt_depth = base_model.get_output_shape_at(0)[-1]\n    pt_features = base_model(in_lay)\n    bn_features = BatchNormalization()(pt_features)\n    # here we do an attention mechanism to turn pixels in the GAP on an off\n    attn_layer = Conv2D(64, kernel_size = (1,1), padding = 'same', activation = 'relu')(Dropout(0.5)(bn_features))\n    attn_layer = Conv2D(16, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\n    attn_layer = Conv2D(8, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\n    attn_layer = Conv2D(1, kernel_size = (1,1), padding = 'valid', activation = 'sigmoid')(attn_layer)\n    # fan it out to all of the channels\n    up_c2_w = np.ones((1, 1, 1, pt_depth))\n    up_c2 = Conv2D(pt_depth, kernel_size = (1,1), padding = 'same', \n                   activation = 'linear', use_bias = False, weights = [up_c2_w])\n    up_c2.trainable = False\n    attn_layer = up_c2(attn_layer)\n\n    mask_features = multiply([attn_layer, bn_features])\n    gap_features = GlobalAveragePooling2D()(mask_features)\n    gap_mask = GlobalAveragePooling2D()(attn_layer)\n    # to account for missing values from the attention model\n    gap = Lambda(lambda x: x[0]/x[1], name = 'RescaleGAP')([gap_features, gap_mask])\n    gap_dr = Dropout(0.25)(gap)\n    dr_steps = Dropout(0.25)(Dense(128, activation = 'relu')(gap_dr))\n    out_layer = Dense(1, activation = 'sigmoid')(dr_steps)\n    model = tf.keras.Model(inputs = [in_lay], outputs = [out_layer])\n    return model\n\n```\n\n B5 was our best single model, and I tried my best to incorporate B6 and B7 in the final ensemble within the time limit. Really thanks to @timesler for his contribution on MTCNN and fast-MTCNN [public kernels](https://www.kaggle.com/timesler/fast-mtcnn-detector-55-fps-at-full-resolution). Thanks @zaharch  for his brilliant [data analysis](https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata), and finally would like to thank @humananalog for inspiring me on [inference codes](https://www.kaggle.com/humananalog/inference-demo) .\n",
    "873179": "Hey! Congratulations for the 27th position! \n\nCould you please explain what do you mean by \"organic videos\"?  \n\nTagging my old friend @ratthachat for response :)",
    "828427": "Thanks for the detailed explaination!! You are very generous!!",
    "825581": "Thanks for sharing. Could you clarifying some questions.\n1) What is frame-by-frame classification? I know many people talked about this, but is this about predict every frame? I saw many was predicting specific frames and get mean or median\n2) Could you share your albumentation code since I found this quiet influencial. \nThanks",
    "837906": "Thanks for sharing\n",
    "832998": "Congrats and Thanks for sharing."
  }
}