{
  "id": 140364,
  "title": "5th (from 20th) place approach",
  "url": "/competitions/deepfake-detection-challenge/discussion/140364",
  "author_name": "James Howard",
  "post_date": "2020-04-01T13:51:58.267000",
  "votes": 100,
  "comment_count": 27,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3108186%2Fc48aafd2dcdc7ff90d51c894946eb8d3%2Fapproach.PNG?generation=1585749236643739&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://james.dev/img/approach_large.PNG\">Full size/res picture available here.</a></p>\n\n<h1>Intro</h1>\n\n<p>Our solution is largely based around 3D CNNs which we felt would generalise well to 'different' methods of deepfakes. Instead of focusing on specific pixel patterns in frames, we hope they might identify temporal problems like the often tried-and-failed 2D CNN-&gt;RNN models, but with a much lower tendency to overfit. The 3D models have similar numbers of parameters to 2D models, and yet also the kernels have a 'depth', so they are relatively modest in their ability to fit in the 2D plane.</p>\n\n<p>Only a month ago we had a solution using a few 3D CNNs, a 2D CNN-&gt;LSTM and 2D CNN and we were 3rd on the leader board. However, we found it hard to keep up over the last couple of weeks and have dropped to 22nd.</p>\n\n<p>Fingers crossed for the private leader board result, and good luck everyone!</p>\n\n<h1>Face extraction</h1>\n\n<p>The basis of the pipeline is the face extractor. This was written within the first week of the competition and since then it's undergone only minor tweaks. In summary, every nth frame (we settled on 10) is passed through MTCNN. The bounding box co-ordinates are then used to set a mask in a 3D array. Faces which are contiguous (overlapping) in 3D are assumed to be a single person's face moving through face and time. We then extract a bounding box which includes this entire face over time and create a video from this region of interest, including every frame (not just every 10th). </p>\n\n<p>One nice thing about this method, even ignoring the video aspect, is it greatly reduced false positives, because the 'face' had to be present for a long period of time of the video to count as a face.</p>\n\n<p>Separately, we use the bounding boxes of MTCNN for a traditional 2D CNN inference.</p>\n\n<h1>3D CNNs</h1>\n\n<p>We trained 7 different 3D CNNs across 4 different architectures (I3D, 3D ResNet34, MC3 &amp; R2+1D) and 2 different resolutions (224 x 224 &amp; 112 x 112). The validation losses of the 3D CNNs ranged between 0.1374 and 0.1905. I initially struggled to fit models with spatial-level augmentation, finding they were so disruptive that the model struggled to tell real apart from fake. I had therefore only been using only pixel-level augmentation such as brightness (and Hflip) and some random cropping. However, Ian developed what we hope is a very successful 3D cutmix approach which led to models with good CV scores to be trained and were added to the ensemble.</p>\n\n<p>To put the benefits of ensembling with these models into context, when we last tried the I3D model (lowest validation loss) alone we ended up with a public LB score of 0.341. With ensembling we’ve improved to 0.253. </p>\n\n<p>Both these networks and the 2D CNNs, these were trained with AdamW (or sometimes Ranger) with a 1-cycle learning rate scheduler.</p>\n\n<h1>2D CNNs</h1>\n\n<p>We ended up including a single 2D CNN in the ensemble in the end. It was trained with quite aggressive augmentation. Ian settled on SE-ResNeXT50 for the final model, though we tried multiple others including Xceptions and EfficientNets.</p>\n\n<h1>Things that didn't work, or didn't work enough to be included.</h1>\n\n<ul>\n<li><p>Using a series of PNGs for videos instead of MP4s. I think this is interesting. We trained these models by generating videos using an identical face extraction pipeline and saving the faces as MP4s (~216,000 fake faces, ~42,000 real, 17Gb of video). Towards the end we wondered if saving the faces as MP4s led to a loss of fidelity and so we saved every frame of these videos as a PNG and tried retraining the networks that way. This took almost a week, resulted in over 500 Gb of data, and resulted in the 3D CNNs overfitting even with reference to the validation data very early. It was a very time-consuming experiment that we still find surprising, given the PNG data is more like what the feature extractor is seeing at test time.</p></li>\n<li><p>Using fewer frames. This might seem obvious but it's a bit more nuanced. We only used faces for videos if they were contiguous over at least 30 frames, up to a maximum of 100. That 100 was set by time constraints. However, using fewer frames would also mean a face is able to move around less. We found in some videos someone would walk across the entire screen in a second. This meant that the face only took up a tiny proportion of the entire region of interest. We therefore tested reducing the maximum frame number to, say, 64 ensure we still got a couple of seconds of video (we trained on chunks of 64 frames) but faces moved around less. However, this gave a worse LB score. We also tried more intelligent methods, such as using at least 30 frames, but then ‘stopping’ as soon as a face had moved &gt; 1 face diameter away from its start point. Again, this was worse (this was particularly surprising). It probably shows that deepfakes are more and less detectable through different stages of the video, and so more frames is always better.</p></li>\n<li><p>CNN -&gt; RNNs did work a <em>little</em>, but not enough to use. Before Ian joined and we became a team, my pipeline did include one of these, but with his better models we dropped it. This seems in contrast to a lot of people who seemed to have 0 success. I suspect this is because we trained these models with the MP4 files rather than image files (see the previous point).</p></li>\n<li><p>A segmentation network to identify altered pixels. Ian developed masked by 'diffing' the real and fake videos and trained a network to identify altered pixels. It had some success, but overfitting meant it had an LB score of ~0.5 and wasn't enough to provide benefit in ensembling.</p></li>\n<li><p>Skipping alternate frames as a form of test-time-augmentation to amplify deltas between frames certainly didn't help, and possibly hindered.</p></li>\n<li><p>Training a fused 3D/2D model which allowed the ensemble weighting to be 'learned' on validation videos did worse than a simple 50:50 pooling.</p></li>\n<li><p>Using more complicated methods of pooling predictions across faces than taking a mean. I previously thought the 'fakest' value should be used, for example, if we were using very high thresholds for detection.</p></li>\n<li><p>Some 3D models. For example, architectures such as 'slowfast' were very easy to train but overfitted profoundly, akin to a 2D CNN. Other networks such as MiCT we struggled to fit at all!</p></li>\n</ul>",
  "messages": [
    {
      "id": 794027,
      "postDate": "2020-04-01T13:51:58.267Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3108186%2Fc48aafd2dcdc7ff90d51c894946eb8d3%2Fapproach.PNG?generation=1585749236643739&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://james.dev/img/approach_large.PNG\">Full size/res picture available here.</a></p>\n\n<h1>Intro</h1>\n\n<p>Our solution is largely based around 3D CNNs which we felt would generalise well to 'different' methods of deepfakes. Instead of focusing on specific pixel patterns in frames, we hope they might identify temporal problems like the often tried-and-failed 2D CNN-&gt;RNN models, but with a much lower tendency to overfit. The 3D models have similar numbers of parameters to 2D models, and yet also the kernels have a 'depth', so they are relatively modest in their ability to fit in the 2D plane.</p>\n\n<p>Only a month ago we had a solution using a few 3D CNNs, a 2D CNN-&gt;LSTM and 2D CNN and we were 3rd on the leader board. However, we found it hard to keep up over the last couple of weeks and have dropped to 22nd.</p>\n\n<p>Fingers crossed for the private leader board result, and good luck everyone!</p>\n\n<h1>Face extraction</h1>\n\n<p>The basis of the pipeline is the face extractor. This was written within the first week of the competition and since then it's undergone only minor tweaks. In summary, every nth frame (we settled on 10) is passed through MTCNN. The bounding box co-ordinates are then used to set a mask in a 3D array. Faces which are contiguous (overlapping) in 3D are assumed to be a single person's face moving through face and time. We then extract a bounding box which includes this entire face over time and create a video from this region of interest, including every frame (not just every 10th). </p>\n\n<p>One nice thing about this method, even ignoring the video aspect, is it greatly reduced false positives, because the 'face' had to be present for a long period of time of the video to count as a face.</p>\n\n<p>Separately, we use the bounding boxes of MTCNN for a traditional 2D CNN inference.</p>\n\n<h1>3D CNNs</h1>\n\n<p>We trained 7 different 3D CNNs across 4 different architectures (I3D, 3D ResNet34, MC3 &amp; R2+1D) and 2 different resolutions (224 x 224 &amp; 112 x 112). The validation losses of the 3D CNNs ranged between 0.1374 and 0.1905. I initially struggled to fit models with spatial-level augmentation, finding they were so disruptive that the model struggled to tell real apart from fake. I had therefore only been using only pixel-level augmentation such as brightness (and Hflip) and some random cropping. However, Ian developed what we hope is a very successful 3D cutmix approach which led to models with good CV scores to be trained and were added to the ensemble.</p>\n\n<p>To put the benefits of ensembling with these models into context, when we last tried the I3D model (lowest validation loss) alone we ended up with a public LB score of 0.341. With ensembling we’ve improved to 0.253. </p>\n\n<p>Both these networks and the 2D CNNs, these were trained with AdamW (or sometimes Ranger) with a 1-cycle learning rate scheduler.</p>\n\n<h1>2D CNNs</h1>\n\n<p>We ended up including a single 2D CNN in the ensemble in the end. It was trained with quite aggressive augmentation. Ian settled on SE-ResNeXT50 for the final model, though we tried multiple others including Xceptions and EfficientNets.</p>\n\n<h1>Things that didn't work, or didn't work enough to be included.</h1>\n\n<ul>\n<li><p>Using a series of PNGs for videos instead of MP4s. I think this is interesting. We trained these models by generating videos using an identical face extraction pipeline and saving the faces as MP4s (~216,000 fake faces, ~42,000 real, 17Gb of video). Towards the end we wondered if saving the faces as MP4s led to a loss of fidelity and so we saved every frame of these videos as a PNG and tried retraining the networks that way. This took almost a week, resulted in over 500 Gb of data, and resulted in the 3D CNNs overfitting even with reference to the validation data very early. It was a very time-consuming experiment that we still find surprising, given the PNG data is more like what the feature extractor is seeing at test time.</p></li>\n<li><p>Using fewer frames. This might seem obvious but it's a bit more nuanced. We only used faces for videos if they were contiguous over at least 30 frames, up to a maximum of 100. That 100 was set by time constraints. However, using fewer frames would also mean a face is able to move around less. We found in some videos someone would walk across the entire screen in a second. This meant that the face only took up a tiny proportion of the entire region of interest. We therefore tested reducing the maximum frame number to, say, 64 ensure we still got a couple of seconds of video (we trained on chunks of 64 frames) but faces moved around less. However, this gave a worse LB score. We also tried more intelligent methods, such as using at least 30 frames, but then ‘stopping’ as soon as a face had moved &gt; 1 face diameter away from its start point. Again, this was worse (this was particularly surprising). It probably shows that deepfakes are more and less detectable through different stages of the video, and so more frames is always better.</p></li>\n<li><p>CNN -&gt; RNNs did work a <em>little</em>, but not enough to use. Before Ian joined and we became a team, my pipeline did include one of these, but with his better models we dropped it. This seems in contrast to a lot of people who seemed to have 0 success. I suspect this is because we trained these models with the MP4 files rather than image files (see the previous point).</p></li>\n<li><p>A segmentation network to identify altered pixels. Ian developed masked by 'diffing' the real and fake videos and trained a network to identify altered pixels. It had some success, but overfitting meant it had an LB score of ~0.5 and wasn't enough to provide benefit in ensembling.</p></li>\n<li><p>Skipping alternate frames as a form of test-time-augmentation to amplify deltas between frames certainly didn't help, and possibly hindered.</p></li>\n<li><p>Training a fused 3D/2D model which allowed the ensemble weighting to be 'learned' on validation videos did worse than a simple 50:50 pooling.</p></li>\n<li><p>Using more complicated methods of pooling predictions across faces than taking a mean. I previously thought the 'fakest' value should be used, for example, if we were using very high thresholds for detection.</p></li>\n<li><p>Some 3D models. For example, architectures such as 'slowfast' were very easy to train but overfitted profoundly, akin to a 2D CNN. Other networks such as MiCT we struggled to fit at all!</p></li>\n</ul>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3108186%2Fc48aafd2dcdc7ff90d51c894946eb8d3%2Fapproach.PNG?generation=1585749236643739&amp;alt=media)\n\n[Full size/res picture available here.](https://james.dev/img/approach_large.PNG)\n\n# Intro\n\nOur solution is largely based around 3D CNNs which we felt would generalise well to 'different' methods of deepfakes. Instead of focusing on specific pixel patterns in frames, we hope they might identify temporal problems like the often tried-and-failed 2D CNN-&gt;RNN models, but with a much lower tendency to overfit. The 3D models have similar numbers of parameters to 2D models, and yet also the kernels have a 'depth', so they are relatively modest in their ability to fit in the 2D plane.\n\n\nOnly a month ago we had a solution using a few 3D CNNs, a 2D CNN-&gt;LSTM and 2D CNN and we were 3rd on the leader board. However, we found it hard to keep up over the last couple of weeks and have dropped to 22nd.\n\n\nFingers crossed for the private leader board result, and good luck everyone!\n\n# Face extraction\n\nThe basis of the pipeline is the face extractor. This was written within the first week of the competition and since then it's undergone only minor tweaks. In summary, every nth frame (we settled on 10) is passed through MTCNN. The bounding box co-ordinates are then used to set a mask in a 3D array. Faces which are contiguous (overlapping) in 3D are assumed to be a single person's face moving through face and time. We then extract a bounding box which includes this entire face over time and create a video from this region of interest, including every frame (not just every 10th). \n\n\nOne nice thing about this method, even ignoring the video aspect, is it greatly reduced false positives, because the 'face' had to be present for a long period of time of the video to count as a face.\n\n\nSeparately, we use the bounding boxes of MTCNN for a traditional 2D CNN inference.\n\n# 3D CNNs\n\nWe trained 7 different 3D CNNs across 4 different architectures (I3D, 3D ResNet34, MC3 &amp; R2+1D) and 2 different resolutions (224 x 224 &amp; 112 x 112). The validation losses of the 3D CNNs ranged between 0.1374 and 0.1905. I initially struggled to fit models with spatial-level augmentation, finding they were so disruptive that the model struggled to tell real apart from fake. I had therefore only been using only pixel-level augmentation such as brightness (and Hflip) and some random cropping. However, Ian developed what we hope is a very successful 3D cutmix approach which led to models with good CV scores to be trained and were added to the ensemble.\n\n\nTo put the benefits of ensembling with these models into context, when we last tried the I3D model (lowest validation loss) alone we ended up with a public LB score of 0.341. With ensembling we’ve improved to 0.253. \n\n\nBoth these networks and the 2D CNNs, these were trained with AdamW (or sometimes Ranger) with a 1-cycle learning rate scheduler.\n\n\n# 2D CNNs\n\nWe ended up including a single 2D CNN in the ensemble in the end. It was trained with quite aggressive augmentation. Ian settled on SE-ResNeXT50 for the final model, though we tried multiple others including Xceptions and EfficientNets.\n\n\n# Things that didn't work, or didn't work enough to be included.\n\n- Using a series of PNGs for videos instead of MP4s. I think this is interesting. We trained these models by generating videos using an identical face extraction pipeline and saving the faces as MP4s (~216,000 fake faces, ~42,000 real, 17Gb of video). Towards the end we wondered if saving the faces as MP4s led to a loss of fidelity and so we saved every frame of these videos as a PNG and tried retraining the networks that way. This took almost a week, resulted in over 500 Gb of data, and resulted in the 3D CNNs overfitting even with reference to the validation data very early. It was a very time-consuming experiment that we still find surprising, given the PNG data is more like what the feature extractor is seeing at test time.\n\n\n- Using fewer frames. This might seem obvious but it's a bit more nuanced. We only used faces for videos if they were contiguous over at least 30 frames, up to a maximum of 100. That 100 was set by time constraints. However, using fewer frames would also mean a face is able to move around less. We found in some videos someone would walk across the entire screen in a second. This meant that the face only took up a tiny proportion of the entire region of interest. We therefore tested reducing the maximum frame number to, say, 64 ensure we still got a couple of seconds of video (we trained on chunks of 64 frames) but faces moved around less. However, this gave a worse LB score. We also tried more intelligent methods, such as using at least 30 frames, but then ‘stopping’ as soon as a face had moved &gt; 1 face diameter away from its start point. Again, this was worse (this was particularly surprising). It probably shows that deepfakes are more and less detectable through different stages of the video, and so more frames is always better.\n\n\n- CNN -&gt; RNNs did work a _little_, but not enough to use. Before Ian joined and we became a team, my pipeline did include one of these, but with his better models we dropped it. This seems in contrast to a lot of people who seemed to have 0 success. I suspect this is because we trained these models with the MP4 files rather than image files (see the previous point).\n\n\n- A segmentation network to identify altered pixels. Ian developed masked by 'diffing' the real and fake videos and trained a network to identify altered pixels. It had some success, but overfitting meant it had an LB score of ~0.5 and wasn't enough to provide benefit in ensembling.\n\n\n- Skipping alternate frames as a form of test-time-augmentation to amplify deltas between frames certainly didn't help, and possibly hindered.\n\n\n- Training a fused 3D/2D model which allowed the ensemble weighting to be 'learned' on validation videos did worse than a simple 50:50 pooling.\n\n\n- Using more complicated methods of pooling predictions across faces than taking a mean. I previously thought the 'fakest' value should be used, for example, if we were using very high thresholds for detection.\n\n\n- Some 3D models. For example, architectures such as 'slowfast' were very easy to train but overfitted profoundly, akin to a 2D CNN. Other networks such as MiCT we struggled to fit at all!",
      "votes": 100
    },
    {
      "id": 794963,
      "postDate": "2020-04-02T08:49:02.063Z",
      "content": "<p>Good job Well done James and Ian.</p>",
      "rawMarkdown": "Good job Well done James and Ian.",
      "votes": 9
    },
    {
      "id": 1311648,
      "postDate": "2021-05-17T14:27:33.333Z",
      "content": "<p>That was very interesting, I will follow your next contributions. Thanks for this awesome work. :)</p>",
      "rawMarkdown": "That was very interesting, I will follow your next contributions. Thanks for this awesome work. :)",
      "votes": 1
    },
    {
      "id": 818588,
      "postDate": "2020-04-24T01:35:44.227Z",
      "content": "<p>Happy we survived the shakeup! Kudos to <a href=\"/jamesphoward\">@jamesphoward</a> for leading this gold medal finish.</p>\n\n<p>Edit: Our kernel is now public: <a href=\"https://www.kaggle.com/vaillant/dfdc-3d-2d-inc-cutmix-with-3d-model-fix\">https://www.kaggle.com/vaillant/dfdc-3d-2d-inc-cutmix-with-3d-model-fix</a></p>",
      "rawMarkdown": "Happy we survived the shakeup! Kudos to @jamesphoward for leading this gold medal finish.\n\nEdit: Our kernel is now public: https://www.kaggle.com/vaillant/dfdc-3d-2d-inc-cutmix-with-3d-model-fix",
      "votes": 4
    },
    {
      "id": 823216,
      "postDate": "2020-04-27T13:47:33.387Z",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> congrats. thanks for your sharing. I also tried 3D CNN method like slowfast, the offline results performance was better than other CNN methods, but LB was worse. would you mind how you split train/val data offline? </p>",
      "rawMarkdown": "@jamesphoward congrats. thanks for your sharing. I also tried 3D CNN method like slowfast, the offline results performance was better than other CNN methods, but LB was worse. would you mind how you split train/val data offline? ",
      "votes": 1,
      "replies": [
        {
          "id": 823232,
          "postDate": "2020-04-27T13:59:47.387Z",
          "content": "<p>Oh interesting! You are the only other person who I've seen use 3D CNNs. Did you use 3D in your final submission?</p>\n\n<p>We just used chunks 45-49 as validation. This is because a couple of these chunks seemed to contain errors in later frames, and because we only used the first n frames for validation, it sense to use these for validation so every video could still be useful.</p>",
          "rawMarkdown": "Oh interesting! You are the only other person who I've seen use 3D CNNs. Did you use 3D in your final submission?\n\nWe just used chunks 45-49 as validation. This is because a couple of these chunks seemed to contain errors in later frames, and because we only used the first n frames for validation, it sense to use these for validation so every video could still be useful."
        },
        {
          "id": 823284,
          "postDate": "2020-04-27T14:36:12.057Z",
          "content": "<p>Thanks for your details, James.  In final submission, we haven't used 3D CNN, and I tried slowfast many times, results were not good. so transfered to 2D CNN.  later we will make our solution public. 🤝 🤝  </p>",
          "rawMarkdown": "Thanks for your details, James.  In final submission, we haven't used 3D CNN, and I tried slowfast many times, results were not good. so transfered to 2D CNN.  later we will make our solution public. 🤝 🤝  ",
          "votes": 1
        },
        {
          "id": 823298,
          "postDate": "2020-04-27T14:45:38.973Z",
          "content": "<p>It's interesting you only tried slowfast, as we found it was the only network (other than MiCT) which didn't work for this task!</p>",
          "rawMarkdown": "It's interesting you only tried slowfast, as we found it was the only network (other than MiCT) which didn't work for this task!"
        },
        {
          "id": 823313,
          "postDate": "2020-04-27T14:56:54.023Z",
          "content": "<p>😷😭 😭  hahaha very regret not to try other methods. coz I used slowfast before in other sequence tasks.... performed very well than other methods. </p>",
          "rawMarkdown": "😷😭 😭  hahaha very regret not to try other methods. coz I used slowfast before in other sequence tasks.... performed very well than other methods. "
        }
      ]
    },
    {
      "id": 798268,
      "postDate": "2020-04-05T11:06:44.833Z",
      "content": "<p>Thanks for sharing, and fingers crossed :)</p>",
      "rawMarkdown": "Thanks for sharing, and fingers crossed :)",
      "votes": 1
    },
    {
      "id": 795644,
      "postDate": "2020-04-02T22:41:57.003Z",
      "content": "<p>Congrats on your amazing work! \nAnd thanks for sharing it with this great explanation. :D\nSome of your comments on the implementation of LSTM inspired me on the implementation of this Transformer model (forward method).\n<a href=\"https://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986\">https://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986</a></p>",
      "rawMarkdown": "Congrats on your amazing work! \nAnd thanks for sharing it with this great explanation. :D\nSome of your comments on the implementation of LSTM inspired me on the implementation of this Transformer model (forward method).\nhttps://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986",
      "votes": 1
    },
    {
      "id": 794568,
      "postDate": "2020-04-01T22:40:19.110Z",
      "content": "<p>Very interesting about the mp4's vs png's. I did several trials of packing everything into an hdf5 array and considered what you had done with capturing each faces zone and treating it as it's own video stream, but never got around to it because I was stuck doing the week-long runs like you described. Lots of things to learn from this competition. </p>",
      "rawMarkdown": "Very interesting about the mp4's vs png's. I did several trials of packing everything into an hdf5 array and considered what you had done with capturing each faces zone and treating it as it's own video stream, but never got around to it because I was stuck doing the week-long runs like you described. Lots of things to learn from this competition. ",
      "votes": 1
    },
    {
      "id": 794481,
      "postDate": "2020-04-01T21:02:12.037Z",
      "content": "<p>20th now. LB already getting cleaned up it seems. </p>",
      "rawMarkdown": "20th now. LB already getting cleaned up it seems. ",
      "votes": 1
    },
    {
      "id": 794376,
      "postDate": "2020-04-01T19:06:13.297Z",
      "content": "<p>Grat job. 1 vote from me 💪 </p>",
      "rawMarkdown": "Grat job. 1 vote from me 💪 ",
      "votes": 1
    },
    {
      "id": 818619,
      "postDate": "2020-04-24T02:27:00.323Z",
      "content": "<p><a href=\"/vaillant\">@vaillant</a> <a href=\"/jamesphoward\">@jamesphoward</a> congratulation for your finish - really enjoyed reading your solution</p>",
      "rawMarkdown": "@vaillant @jamesphoward congratulation for your finish - really enjoyed reading your solution",
      "votes": 2
    },
    {
      "id": 818453,
      "postDate": "2020-04-23T22:38:37.793Z",
      "content": "<p>Congrats on the jump! You were killing this competition from the very start! Well deserved!</p>",
      "rawMarkdown": "Congrats on the jump! You were killing this competition from the very start! Well deserved!",
      "votes": 2,
      "replies": [
        {
          "id": 818463,
          "postDate": "2020-04-23T22:49:29.910Z",
          "content": "<p>Thank you! Congrats to you too!</p>",
          "rawMarkdown": "Thank you! Congrats to you too!",
          "votes": 2
        }
      ]
    },
    {
      "id": 883563,
      "postDate": "2020-06-12T18:25:41.210Z",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> I guess you should rename the discussion. ;) \nCongratulations again!</p>",
      "rawMarkdown": "@jamesphoward I guess you should rename the discussion. ;) \nCongratulations again!",
      "replies": [
        {
          "id": 883579,
          "postDate": "2020-06-12T18:39:08.467Z",
          "content": "<p>Ha, thanks, and done :)</p>",
          "rawMarkdown": "Ha, thanks, and done :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 824775,
      "postDate": "2020-04-28T15:44:27.743Z",
      "content": "<p>Great work man</p>",
      "rawMarkdown": "Great work man"
    },
    {
      "id": 822313,
      "postDate": "2020-04-26T20:23:34Z",
      "content": "<p>Congrats on your medal and thanks for sharing your insights +1</p>",
      "rawMarkdown": "Congrats on your medal and thanks for sharing your insights +1"
    },
    {
      "id": 819615,
      "postDate": "2020-04-24T17:57:13.630Z",
      "content": "<p>Well done <a href=\"/jamesphoward\">@jamesphoward</a>! Couldn't find the \"second upvote\" button though. :P \nAre you planning to release a notebook?</p>",
      "rawMarkdown": "Well done @jamesphoward! Couldn't find the \"second upvote\" button though. :P \nAre you planning to release a notebook?"
    },
    {
      "id": 794498,
      "postDate": "2020-04-01T21:22:03.060Z",
      "content": "<p>Very nice. Well done James and Ian.</p>",
      "rawMarkdown": "Very nice. Well done James and Ian."
    },
    {
      "id": 794446,
      "postDate": "2020-04-01T20:30:26.267Z",
      "content": "<p>The png stuff is very interesting... I also tried lstm and gru with pre extracted png with zero success.... Very strange that mp4 worked!\nThanks for sharing this! </p>",
      "rawMarkdown": "The png stuff is very interesting... I also tried lstm and gru with pre extracted png with zero success.... Very strange that mp4 worked!\nThanks for sharing this! ",
      "replies": [
        {
          "id": 795635,
          "postDate": "2020-04-02T22:24:27.790Z",
          "content": "<p>In retrospect, and upon revisiting my code, it seems that i converted twice from bgr to rgb. I loaded the image using opencv, converted to rgb and savedas png. In the generator I loaded the png and converted again to rgb. Perhaps this was the problem. </p>",
          "rawMarkdown": "In retrospect, and upon revisiting my code, it seems that i converted twice from bgr to rgb. I loaded the image using opencv, converted to rgb and savedas png. In the generator I loaded the png and converted again to rgb. Perhaps this was the problem. "
        }
      ]
    },
    {
      "id": 794282,
      "postDate": "2020-04-01T17:55:48.283Z",
      "content": "<p>Thanks for sharing your approach. Great diagram by the way! </p>",
      "rawMarkdown": "Thanks for sharing your approach. Great diagram by the way! "
    },
    {
      "id": 819675,
      "postDate": "2020-04-24T19:00:47.680Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    },
    {
      "id": 794471,
      "postDate": "2020-04-01T20:53:15.230Z",
      "content": "<p>Great post! Thanks for sharing! </p>",
      "rawMarkdown": "Great post! Thanks for sharing! "
    }
  ],
  "comments": [
    {
      "id": 794963,
      "author_name": "shivan kumar",
      "author_url": "",
      "post_date": "2020-04-02T08:49:02.063000",
      "content": "<p>Good job Well done James and Ian.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1311648,
      "author_name": "MatthiRou",
      "author_url": "",
      "post_date": "2021-05-17T14:27:33.333000",
      "content": "<p>That was very interesting, I will follow your next contributions. Thanks for this awesome work. :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 818588,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-04-24T01:35:44.227000",
      "content": "<p>Happy we survived the shakeup! Kudos to <a href=\"/jamesphoward\">@jamesphoward</a> for leading this gold medal finish.</p>\n\n<p>Edit: Our kernel is now public: <a href=\"https://www.kaggle.com/vaillant/dfdc-3d-2d-inc-cutmix-with-3d-model-fix\">https://www.kaggle.com/vaillant/dfdc-3d-2d-inc-cutmix-with-3d-model-fix</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 823216,
      "author_name": "walton",
      "author_url": "",
      "post_date": "2020-04-27T13:47:33.387000",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> congrats. thanks for your sharing. I also tried 3D CNN method like slowfast, the offline results performance was better than other CNN methods, but LB was worse. would you mind how you split train/val data offline? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 823232,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-04-27T13:59:47.387000",
          "content": "<p>Oh interesting! You are the only other person who I've seen use 3D CNNs. Did you use 3D in your final submission?</p>\n\n<p>We just used chunks 45-49 as validation. This is because a couple of these chunks seemed to contain errors in later frames, and because we only used the first n frames for validation, it sense to use these for validation so every video could still be useful.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 823284,
          "author_name": "walton",
          "author_url": "",
          "post_date": "2020-04-27T14:36:12.057000",
          "content": "<p>Thanks for your details, James.  In final submission, we haven't used 3D CNN, and I tried slowfast many times, results were not good. so transfered to 2D CNN.  later we will make our solution public. 🤝 🤝  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 823298,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-04-27T14:45:38.973000",
          "content": "<p>It's interesting you only tried slowfast, as we found it was the only network (other than MiCT) which didn't work for this task!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 823313,
          "author_name": "walton",
          "author_url": "",
          "post_date": "2020-04-27T14:56:54.023000",
          "content": "<p>😷😭 😭  hahaha very regret not to try other methods. coz I used slowfast before in other sequence tasks.... performed very well than other methods. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 798268,
      "author_name": "thekokonuts",
      "author_url": "",
      "post_date": "2020-04-05T11:06:44.833000",
      "content": "<p>Thanks for sharing, and fingers crossed :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 795644,
      "author_name": "Mont3z Claro5",
      "author_url": "",
      "post_date": "2020-04-02T22:41:57.003000",
      "content": "<p>Congrats on your amazing work! \nAnd thanks for sharing it with this great explanation. :D\nSome of your comments on the implementation of LSTM inspired me on the implementation of this Transformer model (forward method).\n<a href=\"https://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986\">https://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 794568,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-04-01T22:40:19.110000",
      "content": "<p>Very interesting about the mp4's vs png's. I did several trials of packing everything into an hdf5 array and considered what you had done with capturing each faces zone and treating it as it's own video stream, but never got around to it because I was stuck doing the week-long runs like you described. Lots of things to learn from this competition. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 794481,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-04-01T21:02:12.037000",
      "content": "<p>20th now. LB already getting cleaned up it seems. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 794376,
      "author_name": "podsyp",
      "author_url": "",
      "post_date": "2020-04-01T19:06:13.297000",
      "content": "<p>Grat job. 1 vote from me 💪 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 818619,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-04-24T02:27:00.323000",
      "content": "<p><a href=\"/vaillant\">@vaillant</a> <a href=\"/jamesphoward\">@jamesphoward</a> congratulation for your finish - really enjoyed reading your solution</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 818453,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-04-23T22:38:37.793000",
      "content": "<p>Congrats on the jump! You were killing this competition from the very start! Well deserved!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 818463,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-04-23T22:49:29.910000",
          "content": "<p>Thank you! Congrats to you too!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 883563,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-06-12T18:25:41.210000",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> I guess you should rename the discussion. ;) \nCongratulations again!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 883579,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-06-12T18:39:08.467000",
          "content": "<p>Ha, thanks, and done :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 824775,
      "author_name": "Vinayak Tyagi",
      "author_url": "",
      "post_date": "2020-04-28T15:44:27.743000",
      "content": "<p>Great work man</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 822313,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T20:23:34",
      "content": "<p>Congrats on your medal and thanks for sharing your insights +1</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819615,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-04-24T17:57:13.630000",
      "content": "<p>Well done <a href=\"/jamesphoward\">@jamesphoward</a>! Couldn't find the \"second upvote\" button though. :P \nAre you planning to release a notebook?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 794498,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "2020-04-01T21:22:03.060000",
      "content": "<p>Very nice. Well done James and Ian.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 794446,
      "author_name": "Moshel",
      "author_url": "",
      "post_date": "2020-04-01T20:30:26.267000",
      "content": "<p>The png stuff is very interesting... I also tried lstm and gru with pre extracted png with zero success.... Very strange that mp4 worked!\nThanks for sharing this! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 795635,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-04-02T22:24:27.790000",
          "content": "<p>In retrospect, and upon revisiting my code, it seems that i converted twice from bgr to rgb. I loaded the image using opencv, converted to rgb and savedas png. In the generator I loaded the png and converted again to rgb. Perhaps this was the problem. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 794282,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-04-01T17:55:48.283000",
      "content": "<p>Thanks for sharing your approach. Great diagram by the way! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819675,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T19:00:47.680000",
      "content": "",
      "votes": -3,
      "replies": []
    },
    {
      "id": 794471,
      "author_name": "Akash",
      "author_url": "",
      "post_date": "2020-04-01T20:53:15.230000",
      "content": "<p>Great post! Thanks for sharing! </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "794027": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3108186%2Fc48aafd2dcdc7ff90d51c894946eb8d3%2Fapproach.PNG?generation=1585749236643739&amp;alt=media)\n\n[Full size/res picture available here.](https://james.dev/img/approach_large.PNG)\n\n# Intro\n\nOur solution is largely based around 3D CNNs which we felt would generalise well to 'different' methods of deepfakes. Instead of focusing on specific pixel patterns in frames, we hope they might identify temporal problems like the often tried-and-failed 2D CNN-&gt;RNN models, but with a much lower tendency to overfit. The 3D models have similar numbers of parameters to 2D models, and yet also the kernels have a 'depth', so they are relatively modest in their ability to fit in the 2D plane.\n\n\nOnly a month ago we had a solution using a few 3D CNNs, a 2D CNN-&gt;LSTM and 2D CNN and we were 3rd on the leader board. However, we found it hard to keep up over the last couple of weeks and have dropped to 22nd.\n\n\nFingers crossed for the private leader board result, and good luck everyone!\n\n# Face extraction\n\nThe basis of the pipeline is the face extractor. This was written within the first week of the competition and since then it's undergone only minor tweaks. In summary, every nth frame (we settled on 10) is passed through MTCNN. The bounding box co-ordinates are then used to set a mask in a 3D array. Faces which are contiguous (overlapping) in 3D are assumed to be a single person's face moving through face and time. We then extract a bounding box which includes this entire face over time and create a video from this region of interest, including every frame (not just every 10th). \n\n\nOne nice thing about this method, even ignoring the video aspect, is it greatly reduced false positives, because the 'face' had to be present for a long period of time of the video to count as a face.\n\n\nSeparately, we use the bounding boxes of MTCNN for a traditional 2D CNN inference.\n\n# 3D CNNs\n\nWe trained 7 different 3D CNNs across 4 different architectures (I3D, 3D ResNet34, MC3 &amp; R2+1D) and 2 different resolutions (224 x 224 &amp; 112 x 112). The validation losses of the 3D CNNs ranged between 0.1374 and 0.1905. I initially struggled to fit models with spatial-level augmentation, finding they were so disruptive that the model struggled to tell real apart from fake. I had therefore only been using only pixel-level augmentation such as brightness (and Hflip) and some random cropping. However, Ian developed what we hope is a very successful 3D cutmix approach which led to models with good CV scores to be trained and were added to the ensemble.\n\n\nTo put the benefits of ensembling with these models into context, when we last tried the I3D model (lowest validation loss) alone we ended up with a public LB score of 0.341. With ensembling we’ve improved to 0.253. \n\n\nBoth these networks and the 2D CNNs, these were trained with AdamW (or sometimes Ranger) with a 1-cycle learning rate scheduler.\n\n\n# 2D CNNs\n\nWe ended up including a single 2D CNN in the ensemble in the end. It was trained with quite aggressive augmentation. Ian settled on SE-ResNeXT50 for the final model, though we tried multiple others including Xceptions and EfficientNets.\n\n\n# Things that didn't work, or didn't work enough to be included.\n\n- Using a series of PNGs for videos instead of MP4s. I think this is interesting. We trained these models by generating videos using an identical face extraction pipeline and saving the faces as MP4s (~216,000 fake faces, ~42,000 real, 17Gb of video). Towards the end we wondered if saving the faces as MP4s led to a loss of fidelity and so we saved every frame of these videos as a PNG and tried retraining the networks that way. This took almost a week, resulted in over 500 Gb of data, and resulted in the 3D CNNs overfitting even with reference to the validation data very early. It was a very time-consuming experiment that we still find surprising, given the PNG data is more like what the feature extractor is seeing at test time.\n\n\n- Using fewer frames. This might seem obvious but it's a bit more nuanced. We only used faces for videos if they were contiguous over at least 30 frames, up to a maximum of 100. That 100 was set by time constraints. However, using fewer frames would also mean a face is able to move around less. We found in some videos someone would walk across the entire screen in a second. This meant that the face only took up a tiny proportion of the entire region of interest. We therefore tested reducing the maximum frame number to, say, 64 ensure we still got a couple of seconds of video (we trained on chunks of 64 frames) but faces moved around less. However, this gave a worse LB score. We also tried more intelligent methods, such as using at least 30 frames, but then ‘stopping’ as soon as a face had moved &gt; 1 face diameter away from its start point. Again, this was worse (this was particularly surprising). It probably shows that deepfakes are more and less detectable through different stages of the video, and so more frames is always better.\n\n\n- CNN -&gt; RNNs did work a _little_, but not enough to use. Before Ian joined and we became a team, my pipeline did include one of these, but with his better models we dropped it. This seems in contrast to a lot of people who seemed to have 0 success. I suspect this is because we trained these models with the MP4 files rather than image files (see the previous point).\n\n\n- A segmentation network to identify altered pixels. Ian developed masked by 'diffing' the real and fake videos and trained a network to identify altered pixels. It had some success, but overfitting meant it had an LB score of ~0.5 and wasn't enough to provide benefit in ensembling.\n\n\n- Skipping alternate frames as a form of test-time-augmentation to amplify deltas between frames certainly didn't help, and possibly hindered.\n\n\n- Training a fused 3D/2D model which allowed the ensemble weighting to be 'learned' on validation videos did worse than a simple 50:50 pooling.\n\n\n- Using more complicated methods of pooling predictions across faces than taking a mean. I previously thought the 'fakest' value should be used, for example, if we were using very high thresholds for detection.\n\n\n- Some 3D models. For example, architectures such as 'slowfast' were very easy to train but overfitted profoundly, akin to a 2D CNN. Other networks such as MiCT we struggled to fit at all!",
    "794963": "Good job Well done James and Ian.",
    "1311648": "That was very interesting, I will follow your next contributions. Thanks for this awesome work. :)",
    "818588": "Happy we survived the shakeup! Kudos to @jamesphoward for leading this gold medal finish.\n\nEdit: Our kernel is now public: https://www.kaggle.com/vaillant/dfdc-3d-2d-inc-cutmix-with-3d-model-fix",
    "823216": "@jamesphoward congrats. thanks for your sharing. I also tried 3D CNN method like slowfast, the offline results performance was better than other CNN methods, but LB was worse. would you mind how you split train/val data offline? ",
    "798268": "Thanks for sharing, and fingers crossed :)",
    "795644": "Congrats on your amazing work! \nAnd thanks for sharing it with this great explanation. :D\nSome of your comments on the implementation of LSTM inspired me on the implementation of this Transformer model (forward method).\nhttps://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986",
    "794568": "Very interesting about the mp4's vs png's. I did several trials of packing everything into an hdf5 array and considered what you had done with capturing each faces zone and treating it as it's own video stream, but never got around to it because I was stuck doing the week-long runs like you described. Lots of things to learn from this competition. ",
    "794481": "20th now. LB already getting cleaned up it seems. ",
    "794376": "Grat job. 1 vote from me 💪 ",
    "818619": "@vaillant @jamesphoward congratulation for your finish - really enjoyed reading your solution",
    "818453": "Congrats on the jump! You were killing this competition from the very start! Well deserved!",
    "883563": "@jamesphoward I guess you should rename the discussion. ;) \nCongratulations again!",
    "824775": "Great work man",
    "822313": "Congrats on your medal and thanks for sharing your insights +1",
    "819615": "Well done @jamesphoward! Couldn't find the \"second upvote\" button though. :P \nAre you planning to release a notebook?",
    "794498": "Very nice. Well done James and Ian.",
    "794446": "The png stuff is very interesting... I also tried lstm and gru with pre extracted png with zero success.... Very strange that mp4 worked!\nThanks for sharing this! ",
    "794282": "Thanks for sharing your approach. Great diagram by the way! ",
    "819675": "",
    "794471": "Great post! Thanks for sharing! "
  }
}