{
  "id": 145732,
  "title": "39th Private LB: small and simple",
  "url": "/competitions/deepfake-detection-challenge/writeups/fergusoci-39th-private-lb-small-and-simple",
  "author_name": "",
  "post_date": "2020-04-24T10:01:54.347884100Z",
  "votes": 28,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Congratulations to everyone, and I hope that those who had scores that errored out get that sorted… Very frustrating.</p>\n\n<p><strong>My approach in a nutshell - keep it simple and write off the hard deep fakes:</strong> </p>\n\n<p>I figured that a complex model would not get these right anyway, and would end up messing up the easy ones too because it would be prone to overfitting. I ended up with 39th on the private LB (0.477); around 100th public LB (0.328).</p>\n\n<p><strong>Details:</strong></p>\n\n<p>My main worry was that models were not going to generalise well, primarily because of the training data we had available to us. My reading of the rules was that any external data was at best in a grey area, and more probably out of bounds, so I only used the data given to us.</p>\n\n<p>I built as simple and as small a model as possible, knowing that this wouldn’t be able to detect hard deep fakes, but figuring that it might detect the easier ones more robustly.</p>\n\n<p><strong>Dataset:</strong></p>\n\n<p>I used @humananalog  Blazeface solution to save off faces in 32 frames from each video to the hard drive (evenly spaced throughout the video). I screened the videos, only allowing training cases where all of the face parts (eyes, nose, mouth, etc) were visible in all frames. This reduced the training set by c.25% and I think this was important in getting the models to converge properly. </p>\n\n<p><strong>Sampling:</strong></p>\n\n<p>I sampled 50-50 real/fake videos in each epoch (all real videos + 1 sampled corresponding fake, each epoch choosing a random rake). Once a video was sampled, I randomly sampled 1 frame from the 10 frames with the highest confidence in that video. Overall, I’m not sure how much the sampling scheme affected the score. It didn’t have a huge impact on my validation set or the public leaderboard, but maybe it counted for something in the private leaderboard.</p>\n\n<p><strong>Augmentation:</strong></p>\n\n<p>Random horizontal flips, jpeg compression, brightness contrast. Randomly selecting frames as outlined above also did some of this job.</p>\n\n<p><strong>Models:</strong></p>\n\n<p>I used EfficientNet-b0 as it was the smallest model that gave reasonable results.</p>\n\n<p><strong>Resolution:</strong></p>\n\n<p>I used three different input sizes (separate models for each): 180, 196, 224. I used smaller rather than larger resolutions in a bid to make things more robust (reasons outlined above).</p>\n\n<p><strong>Ensemble:</strong></p>\n\n<p>Each model was fit over 5-fold samples (folder split). The end result averaged these.</p>",
  "messages": [
    {
      "id": "819042",
      "postDate": "04/24/2020 10:01:54",
      "content": "<p>Congratulations to everyone, and I hope that those who had scores that errored out get that sorted… Very frustrating.</p>\n\n<p><strong>My approach in a nutshell - keep it simple and write off the hard deep fakes:</strong> </p>\n\n<p>I figured that a complex model would not get these right anyway, and would end up messing up the easy ones too because it would be prone to overfitting. I ended up with 39th on the private LB (0.477); around 100th public LB (0.328).</p>\n\n<p><strong>Details:</strong></p>\n\n<p>My main worry was that models were not going to generalise well, primarily because of the training data we had available to us. My reading of the rules was that any external data was at best in a grey area, and more probably out of bounds, so I only used the data given to us.</p>\n\n<p>I built as simple and as small a model as possible, knowing that this wouldn’t be able to detect hard deep fakes, but figuring that it might detect the easier ones more robustly.</p>\n\n<p><strong>Dataset:</strong></p>\n\n<p>I used @humananalog  Blazeface solution to save off faces in 32 frames from each video to the hard drive (evenly spaced throughout the video). I screened the videos, only allowing training cases where all of the face parts (eyes, nose, mouth, etc) were visible in all frames. This reduced the training set by c.25% and I think this was important in getting the models to converge properly. </p>\n\n<p><strong>Sampling:</strong></p>\n\n<p>I sampled 50-50 real/fake videos in each epoch (all real videos + 1 sampled corresponding fake, each epoch choosing a random rake). Once a video was sampled, I randomly sampled 1 frame from the 10 frames with the highest confidence in that video. Overall, I’m not sure how much the sampling scheme affected the score. It didn’t have a huge impact on my validation set or the public leaderboard, but maybe it counted for something in the private leaderboard.</p>\n\n<p><strong>Augmentation:</strong></p>\n\n<p>Random horizontal flips, jpeg compression, brightness contrast. Randomly selecting frames as outlined above also did some of this job.</p>\n\n<p><strong>Models:</strong></p>\n\n<p>I used EfficientNet-b0 as it was the smallest model that gave reasonable results.</p>\n\n<p><strong>Resolution:</strong></p>\n\n<p>I used three different input sizes (separate models for each): 180, 196, 224. I used smaller rather than larger resolutions in a bid to make things more robust (reasons outlined above).</p>\n\n<p><strong>Ensemble:</strong></p>\n\n<p>Each model was fit over 5-fold samples (folder split). The end result averaged these.</p>",
      "rawMarkdown": "Congratulations to everyone, and I hope that those who had scores that errored out get that sorted… Very frustrating.\n\n**My approach in a nutshell - keep it simple and write off the hard deep fakes:** \n\nI figured that a complex model would not get these right anyway, and would end up messing up the easy ones too because it would be prone to overfitting. I ended up with 39th on the private LB (0.477); around 100th public LB (0.328).\n\n**Details:**\n\nMy main worry was that models were not going to generalise well, primarily because of the training data we had available to us. My reading of the rules was that any external data was at best in a grey area, and more probably out of bounds, so I only used the data given to us.\n\nI built as simple and as small a model as possible, knowing that this wouldn’t be able to detect hard deep fakes, but figuring that it might detect the easier ones more robustly.\n\n**Dataset:**\n\nI used @humananalog  Blazeface solution to save off faces in 32 frames from each video to the hard drive (evenly spaced throughout the video). I screened the videos, only allowing training cases where all of the face parts (eyes, nose, mouth, etc) were visible in all frames. This reduced the training set by c.25% and I think this was important in getting the models to converge properly. \n\n**Sampling:**\n\nI sampled 50-50 real/fake videos in each epoch (all real videos + 1 sampled corresponding fake, each epoch choosing a random rake). Once a video was sampled, I randomly sampled 1 frame from the 10 frames with the highest confidence in that video. Overall, I’m not sure how much the sampling scheme affected the score. It didn’t have a huge impact on my validation set or the public leaderboard, but maybe it counted for something in the private leaderboard.\n\n**Augmentation:**\n\nRandom horizontal flips, jpeg compression, brightness contrast. Randomly selecting frames as outlined above also did some of this job.\n\n**Models:**\n\nI used EfficientNet-b0 as it was the smallest model that gave reasonable results.\n\n**Resolution:**\n\nI used three different input sizes (separate models for each): 180, 196, 224. I used smaller rather than larger resolutions in a bid to make things more robust (reasons outlined above).\n\n**Ensemble:**\n\nEach model was fit over 5-fold samples (folder split). The end result averaged these.",
      "votes": null
    },
    {
      "id": "819061",
      "postDate": "04/24/2020 10:22:02",
      "content": "<p>Very nice. Our approach was very similar— except we used effb3, did not use various image sizes (which I wish I had) but only 224, and 3 folds ensemble (with median). For our final sub we just run it many times to add more ensembles to rank 118. Added a few xception, resnext models in the ensemble (50% effb3, 50% others) for safe side and resulted in Rank 96.</p>",
      "rawMarkdown": "Very nice. Our approach was very similar— except we used effb3, did not use various image sizes (which I wish I had) but only 224, and 3 folds ensemble (with median). For our final sub we just run it many times to add more ensembles to rank 118. Added a few xception, resnext models in the ensemble (50% effb3, 50% others) for safe side and resulted in Rank 96.",
      "votes": null
    },
    {
      "id": "819064",
      "postDate": "04/24/2020 10:27:49",
      "content": "<p>Well done! The simple things that I think made a difference for me were ensembling on different folds (rather than different seeds), and image sizes.</p>",
      "rawMarkdown": "Well done! The simple things that I think made a difference for me were ensembling on different folds (rather than different seeds), and image sizes.",
      "votes": null
    },
    {
      "id": "819093",
      "postDate": "04/24/2020 10:54:29",
      "content": "<p>Very elegant, probably one of the most useful posts on the forum for people getting started. Really underlines the strength of regularisation by using minimalist models with minimal assumption in problems where overfitting is going to be a large issue.</p>",
      "rawMarkdown": "Very elegant, probably one of the most useful posts on the forum for people getting started. Really underlines the strength of regularisation by using minimalist models with minimal assumption in problems where overfitting is going to be a large issue.",
      "votes": null
    },
    {
      "id": "819189",
      "postDate": "04/24/2020 12:15:08",
      "content": "<p>Thank you for sharing your approach. Cheers!</p>",
      "rawMarkdown": "Thank you for sharing your approach. Cheers!",
      "votes": null
    },
    {
      "id": "819710",
      "postDate": "04/24/2020 19:36:58",
      "content": "<p>Amazing! Would you like to collaborate on your training process for EfficientNetB0? Did you use any learning rate scheduling, did your local validation score also come from 5-fold cross validation and did your validation scores correlate well with leaderboard scores? As for the image screening, did you do this manually or was there some automation involved?</p>",
      "rawMarkdown": "Amazing! Would you like to collaborate on your training process for EfficientNetB0? Did you use any learning rate scheduling, did your local validation score also come from 5-fold cross validation and did your validation scores correlate well with leaderboard scores? As for the image screening, did you do this manually or was there some automation involved?",
      "votes": null
    },
    {
      "id": "819731",
      "postDate": "04/24/2020 19:59:16",
      "content": "<p>I used Adam as the optimizer, reduced LR on plateau (reduction factor = 0.1). In order to speed things up I just used 1 of the 5 folds for local validation (folders 41 - 49). This correlated reasonably well with the leaderboard. However I was wary of both the LB and local validation because I thought they were giving too positive a picture... Hence why I tried to underfit them.</p>\n\n<p>For image screening, this was automated. I extracted face crops and facepart locations using Blazeface. In some instances it couldn't find a face, or it found the face but couldn't identify both eyes, etc. I screened the videos based on this: for a video to be included in training, all of the frames had to have at least one face, as well as all face parts.</p>",
      "rawMarkdown": "I used Adam as the optimizer, reduced LR on plateau (reduction factor = 0.1). In order to speed things up I just used 1 of the 5 folds for local validation (folders 41 - 49). This correlated reasonably well with the leaderboard. However I was wary of both the LB and local validation because I thought they were giving too positive a picture... Hence why I tried to underfit them.\n\nFor image screening, this was automated. I extracted face crops and facepart locations using Blazeface. In some instances it couldn't find a face, or it found the face but couldn't identify both eyes, etc. I screened the videos based on this: for a video to be included in training, all of the frames had to have at least one face, as well as all face parts.",
      "votes": null
    },
    {
      "id": "819741",
      "postDate": "04/24/2020 20:08:39",
      "content": "<p>Very nice, thank you for the clarification!</p>",
      "rawMarkdown": "Very nice, thank you for the clarification!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 819061,
      "author_name": "debanga",
      "author_url": "",
      "post_date": "04/24/2020 10:22:02",
      "content": "<p>Very nice. Our approach was very similar— except we used effb3, did not use various image sizes (which I wish I had) but only 224, and 3 folds ensemble (with median). For our final sub we just run it many times to add more ensembles to rank 118. Added a few xception, resnext models in the ensemble (50% effb3, 50% others) for safe side and resulted in Rank 96.</p>",
      "votes": null,
      "replies": [
        {
          "id": 819064,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "04/24/2020 10:27:49",
          "content": "<p>Well done! The simple things that I think made a difference for me were ensembling on different folds (rather than different seeds), and image sizes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 819093,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "04/24/2020 10:54:29",
      "content": "<p>Very elegant, probably one of the most useful posts on the forum for people getting started. Really underlines the strength of regularisation by using minimalist models with minimal assumption in problems where overfitting is going to be a large issue.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819189,
      "author_name": "xincuimath",
      "author_url": "",
      "post_date": "04/24/2020 12:15:08",
      "content": "<p>Thank you for sharing your approach. Cheers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819710,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "04/24/2020 19:36:58",
      "content": "<p>Amazing! Would you like to collaborate on your training process for EfficientNetB0? Did you use any learning rate scheduling, did your local validation score also come from 5-fold cross validation and did your validation scores correlate well with leaderboard scores? As for the image screening, did you do this manually or was there some automation involved?</p>",
      "votes": null,
      "replies": [
        {
          "id": 819731,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "04/24/2020 19:59:16",
          "content": "<p>I used Adam as the optimizer, reduced LR on plateau (reduction factor = 0.1). In order to speed things up I just used 1 of the 5 folds for local validation (folders 41 - 49). This correlated reasonably well with the leaderboard. However I was wary of both the LB and local validation because I thought they were giving too positive a picture... Hence why I tried to underfit them.</p>\n\n<p>For image screening, this was automated. I extracted face crops and facepart locations using Blazeface. In some instances it couldn't find a face, or it found the face but couldn't identify both eyes, etc. I screened the videos based on this: for a video to be included in training, all of the frames had to have at least one face, as well as all face parts.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 819741,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "04/24/2020 20:08:39",
          "content": "<p>Very nice, thank you for the clarification!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "819042": "Congratulations to everyone, and I hope that those who had scores that errored out get that sorted… Very frustrating.\n\n**My approach in a nutshell - keep it simple and write off the hard deep fakes:** \n\nI figured that a complex model would not get these right anyway, and would end up messing up the easy ones too because it would be prone to overfitting. I ended up with 39th on the private LB (0.477); around 100th public LB (0.328).\n\n**Details:**\n\nMy main worry was that models were not going to generalise well, primarily because of the training data we had available to us. My reading of the rules was that any external data was at best in a grey area, and more probably out of bounds, so I only used the data given to us.\n\nI built as simple and as small a model as possible, knowing that this wouldn’t be able to detect hard deep fakes, but figuring that it might detect the easier ones more robustly.\n\n**Dataset:**\n\nI used @humananalog  Blazeface solution to save off faces in 32 frames from each video to the hard drive (evenly spaced throughout the video). I screened the videos, only allowing training cases where all of the face parts (eyes, nose, mouth, etc) were visible in all frames. This reduced the training set by c.25% and I think this was important in getting the models to converge properly. \n\n**Sampling:**\n\nI sampled 50-50 real/fake videos in each epoch (all real videos + 1 sampled corresponding fake, each epoch choosing a random rake). Once a video was sampled, I randomly sampled 1 frame from the 10 frames with the highest confidence in that video. Overall, I’m not sure how much the sampling scheme affected the score. It didn’t have a huge impact on my validation set or the public leaderboard, but maybe it counted for something in the private leaderboard.\n\n**Augmentation:**\n\nRandom horizontal flips, jpeg compression, brightness contrast. Randomly selecting frames as outlined above also did some of this job.\n\n**Models:**\n\nI used EfficientNet-b0 as it was the smallest model that gave reasonable results.\n\n**Resolution:**\n\nI used three different input sizes (separate models for each): 180, 196, 224. I used smaller rather than larger resolutions in a bid to make things more robust (reasons outlined above).\n\n**Ensemble:**\n\nEach model was fit over 5-fold samples (folder split). The end result averaged these.",
    "819061": "Very nice. Our approach was very similar— except we used effb3, did not use various image sizes (which I wish I had) but only 224, and 3 folds ensemble (with median). For our final sub we just run it many times to add more ensembles to rank 118. Added a few xception, resnext models in the ensemble (50% effb3, 50% others) for safe side and resulted in Rank 96.",
    "819064": "Well done! The simple things that I think made a difference for me were ensembling on different folds (rather than different seeds), and image sizes.",
    "819093": "Very elegant, probably one of the most useful posts on the forum for people getting started. Really underlines the strength of regularisation by using minimalist models with minimal assumption in problems where overfitting is going to be a large issue.",
    "819189": "Thank you for sharing your approach. Cheers!",
    "819710": "Amazing! Would you like to collaborate on your training process for EfficientNetB0? Did you use any learning rate scheduling, did your local validation score also come from 5-fold cross validation and did your validation scores correlate well with leaderboard scores? As for the image screening, did you do this manually or was there some automation involved?",
    "819731": "I used Adam as the optimizer, reduced LR on plateau (reduction factor = 0.1). In order to speed things up I just used 1 of the 5 folds for local validation (folders 41 - 49). This correlated reasonably well with the leaderboard. However I was wary of both the LB and local validation because I thought they were giving too positive a picture... Hence why I tried to underfit them.\n\nFor image screening, this was automated. I extracted face crops and facepart locations using Blazeface. In some instances it couldn't find a face, or it found the face but couldn't identify both eyes, etc. I screened the videos based on this: for a video to be included in training, all of the frames had to have at least one face, as well as all face parts.",
    "819741": "Very nice, thank you for the clarification!"
  },
  "source": "meta"
}