{
  "id": 138249,
  "title": "Some last minute insights",
  "url": "/competitions/deepfake-detection-challenge/discussion/138249",
  "author_name": "",
  "post_date": "2020-03-24T09:36:35.133360Z",
  "votes": 13,
  "comment_count": 5,
  "views": 0,
  "content": "<h3>Pre-processing</h3>\n\n<p>This is soooo crucial. Really, I lost entire months because of this going from crappy models to crappy results. I assumed quantity over quality would be best, and that lots of faces would even out the few bad ones, but it was wrong. The noise introduced by bad faces is significant: it worsens everything in the pipeline: the training, the predictions... even with the best models, you'll get crap.</p>\n\n<p>I'm convinced pre-processing is actually more important than the model itself. I highly advise the mobile face detector of <a href=\"/unkownhihi\">@unkownhihi</a> and <a href=\"/harshitsheoran\">@harshitsheoran</a></p>\n\n<p><a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison\">https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison</a></p>\n\n<p>It's really good, and reasonably fast.</p>\n\n<h3>Models</h3>\n\n<p>I advice first starting with pre-trained known models. Like efficient net or <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a> . Dunno why, but when training them, it seems better to initialize them with the pre-trained weights rather than random weights. Most of the time, when trying the latter, the training was just going nowhere. Dunno why. Perhaps the differences are too subtle and you need a meaningful starting point.</p>\n\n<h3>Optimizer</h3>\n\n<p>Optimizers like Adam, Nadam and such tend, from my experience, to converge fairly faster. However, they tend to have more fluctuations in validation scores, apparently overfitting on some details.</p>\n\n<p>Simpler algorithms like SGD tend to converge slower, however, from my experience they tend to generalize better and have more stable validation scores.</p>\n\n<h3>Validation set</h3>\n\n<p>For validation, it is IMHO best and simplest to take a whole directory (or a few).</p>\n\n<p>Since the directories tend to have just a bunch of actors, this reduces the amount you have seen them elsewhere, making it a reasonable validation set.</p>\n\n<p>I personally picked the training directory 28 as validation set. The advantage is that you can verify your pipeline since the training data in the kaggle train/test sample of 400 videos is also from this directory.</p>\n\n<h3>About the difference between validation score and leaderboard score</h3>\n\n<p>One of the reason is probably a set of new actors appearing, making predictions less reliable. However, I think one other important reason is that the test set is \"augmented\".</p>\n\n<p>In case you missed it, the test set is different than the training set, as mentioned here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013</a></p>\n\n<p>Basically, the test set is augmented as follows:</p>\n\n<p>a) 2/9 of videos: reduce the FPS of the video to 15;\nb) 2/9 of videos: reduce the resolution of the video to 1/4 of its original size;\nc) 2/9 of videos: reduce the overall encoding quality.\nd) 3/9 of videos: unmodified</p>\n\n<p>I'm pretty sure that points (b) and (c) can \"destroy\" many predictions and worsen the score. I think it is likely that the excessive zooming and compression artifacts can easily be confused as fakes.</p>\n\n<p>To boost the score, I suppose countermeasures must be taken. Especially point (c) is a bit tricky since we do not know how exactly the \"overall encoding quality\" was reduced.</p>",
  "messages": [
    {
      "id": "784528",
      "postDate": "03/24/2020 09:36:35",
      "content": "<h3>Pre-processing</h3>\n\n<p>This is soooo crucial. Really, I lost entire months because of this going from crappy models to crappy results. I assumed quantity over quality would be best, and that lots of faces would even out the few bad ones, but it was wrong. The noise introduced by bad faces is significant: it worsens everything in the pipeline: the training, the predictions... even with the best models, you'll get crap.</p>\n\n<p>I'm convinced pre-processing is actually more important than the model itself. I highly advise the mobile face detector of <a href=\"/unkownhihi\">@unkownhihi</a> and <a href=\"/harshitsheoran\">@harshitsheoran</a></p>\n\n<p><a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison\">https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison</a></p>\n\n<p>It's really good, and reasonably fast.</p>\n\n<h3>Models</h3>\n\n<p>I advice first starting with pre-trained known models. Like efficient net or <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a> . Dunno why, but when training them, it seems better to initialize them with the pre-trained weights rather than random weights. Most of the time, when trying the latter, the training was just going nowhere. Dunno why. Perhaps the differences are too subtle and you need a meaningful starting point.</p>\n\n<h3>Optimizer</h3>\n\n<p>Optimizers like Adam, Nadam and such tend, from my experience, to converge fairly faster. However, they tend to have more fluctuations in validation scores, apparently overfitting on some details.</p>\n\n<p>Simpler algorithms like SGD tend to converge slower, however, from my experience they tend to generalize better and have more stable validation scores.</p>\n\n<h3>Validation set</h3>\n\n<p>For validation, it is IMHO best and simplest to take a whole directory (or a few).</p>\n\n<p>Since the directories tend to have just a bunch of actors, this reduces the amount you have seen them elsewhere, making it a reasonable validation set.</p>\n\n<p>I personally picked the training directory 28 as validation set. The advantage is that you can verify your pipeline since the training data in the kaggle train/test sample of 400 videos is also from this directory.</p>\n\n<h3>About the difference between validation score and leaderboard score</h3>\n\n<p>One of the reason is probably a set of new actors appearing, making predictions less reliable. However, I think one other important reason is that the test set is \"augmented\".</p>\n\n<p>In case you missed it, the test set is different than the training set, as mentioned here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013</a></p>\n\n<p>Basically, the test set is augmented as follows:</p>\n\n<p>a) 2/9 of videos: reduce the FPS of the video to 15;\nb) 2/9 of videos: reduce the resolution of the video to 1/4 of its original size;\nc) 2/9 of videos: reduce the overall encoding quality.\nd) 3/9 of videos: unmodified</p>\n\n<p>I'm pretty sure that points (b) and (c) can \"destroy\" many predictions and worsen the score. I think it is likely that the excessive zooming and compression artifacts can easily be confused as fakes.</p>\n\n<p>To boost the score, I suppose countermeasures must be taken. Especially point (c) is a bit tricky since we do not know how exactly the \"overall encoding quality\" was reduced.</p>",
      "rawMarkdown": "### Pre-processing\n \nThis is soooo crucial. Really, I lost entire months because of this going from crappy models to crappy results. I assumed quantity over quality would be best, and that lots of faces would even out the few bad ones, but it was wrong. The noise introduced by bad faces is significant: it worsens everything in the pipeline: the training, the predictions... even with the best models, you'll get crap.\n\nI'm convinced pre-processing is actually more important than the model itself. I highly advise the mobile face detector of @unkownhihi and @harshitsheoran\n\nhttps://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison\n\nIt's really good, and reasonably fast.\n\n### Models\n\nI advice first starting with pre-trained known models. Like efficient net or https://keras.io/applications/ . Dunno why, but when training them, it seems better to initialize them with the pre-trained weights rather than random weights. Most of the time, when trying the latter, the training was just going nowhere. Dunno why. Perhaps the differences are too subtle and you need a meaningful starting point.\n\n### Optimizer\n\nOptimizers like Adam, Nadam and such tend, from my experience, to converge fairly faster. However, they tend to have more fluctuations in validation scores, apparently overfitting on some details.\n\nSimpler algorithms like SGD tend to converge slower, however, from my experience they tend to generalize better and have more stable validation scores.\n\n### Validation set\n\nFor validation, it is IMHO best and simplest to take a whole directory (or a few).\n\nSince the directories tend to have just a bunch of actors, this reduces the amount you have seen them elsewhere, making it a reasonable validation set.\n\nI personally picked the training directory 28 as validation set. The advantage is that you can verify your pipeline since the training data in the kaggle train/test sample of 400 videos is also from this directory.\n\n### About the difference between validation score and leaderboard score\n\nOne of the reason is probably a set of new actors appearing, making predictions less reliable. However, I think one other important reason is that the test set is \"augmented\".\n\nIn case you missed it, the test set is different than the training set, as mentioned here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013\n\nBasically, the test set is augmented as follows:\n\na) 2/9 of videos: reduce the FPS of the video to 15;\nb) 2/9 of videos: reduce the resolution of the video to 1/4 of its original size;\nc) 2/9 of videos: reduce the overall encoding quality.\nd) 3/9 of videos: unmodified\n\nI'm pretty sure that points (b) and (c) can \"destroy\" many predictions and worsen the score. I think it is likely that the excessive zooming and compression artifacts can easily be confused as fakes.\n\nTo boost the score, I suppose countermeasures must be taken. Especially point (c) is a bit tricky since we do not know how exactly the \"overall encoding quality\" was reduced.",
      "votes": null
    },
    {
      "id": "784544",
      "postDate": "03/24/2020 09:58:50",
      "content": "<p>I'm sure that the compression affect the prediction, even 5%. I'm wondering what augmentation was applied by the top tiers.</p>",
      "rawMarkdown": "I'm sure that the compression affect the prediction, even 5%. I'm wondering what augmentation was applied by the top tiers.",
      "votes": null
    },
    {
      "id": "788874",
      "postDate": "03/28/2020 05:42:35",
      "content": "<p><a href=\"/dagnelies\">@dagnelies</a> <br>\n1)what do you mean by 2/9 here\n2) how fps affects teh video.. video would be too slow moving ?</p>",
      "rawMarkdown": "dagnelies  \n1)what do you mean by 2/9 here\n2) how fps affects teh video.. video would be too slow moving ?",
      "votes": null
    },
    {
      "id": "793036",
      "postDate": "03/31/2020 18:45:38",
      "content": "<p>1) just \"2/9\" ...it's described in detail in the link's article\n2) the \"speed\" would be the same, it's just recorded at 15 FPS instead of 30 FPS</p>",
      "rawMarkdown": "1) just \"2/9\" ...it's described in detail in the link's article\n2) the \"speed\" would be the same, it's just recorded at 15 FPS instead of 30 FPS",
      "votes": null
    },
    {
      "id": "793052",
      "postDate": "03/31/2020 19:01:07",
      "content": "<p>I didn't get to participate much in this competition but I played around with the data a lot. It occurred to me to use only 1-face videos as a training set and 2-face videos as a holdout test set. Why? Because we're given the original videos. Face detectors will generally work very well on the real faces. If you get all the real videos with a high probability of 1-face, you can assume their associated deepfakes will only have one face. So you can make a training set of real/fake videos where you only take the highest-probability box from the face detector. This would reduce noise drastically.</p>\n\n<p>This looked a lot of fun, I wish I'd had more time...</p>",
      "rawMarkdown": "I didn't get to participate much in this competition but I played around with the data a lot. It occurred to me to use only 1-face videos as a training set and 2-face videos as a holdout test set. Why? Because we're given the original videos. Face detectors will generally work very well on the real faces. If you get all the real videos with a high probability of 1-face, you can assume their associated deepfakes will only have one face. So you can make a training set of real/fake videos where you only take the highest-probability box from the face detector. This would reduce noise drastically.\n\nThis looked a lot of fun, I wish I'd had more time...",
      "votes": null
    },
    {
      "id": "793284",
      "postDate": "03/31/2020 23:27:49",
      "content": "<p>Definitely custom and clever data augmentation will be a huge contributor to top solutions.</p>",
      "rawMarkdown": "Definitely custom and clever data augmentation will be a huge contributor to top solutions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 784544,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "03/24/2020 09:58:50",
      "content": "<p>I'm sure that the compression affect the prediction, even 5%. I'm wondering what augmentation was applied by the top tiers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 788874,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "03/28/2020 05:42:35",
      "content": "<p><a href=\"/dagnelies\">@dagnelies</a> <br>\n1)what do you mean by 2/9 here\n2) how fps affects teh video.. video would be too slow moving ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 793036,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "03/31/2020 18:45:38",
          "content": "<p>1) just \"2/9\" ...it's described in detail in the link's article\n2) the \"speed\" would be the same, it's just recorded at 15 FPS instead of 30 FPS</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 793052,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "03/31/2020 19:01:07",
      "content": "<p>I didn't get to participate much in this competition but I played around with the data a lot. It occurred to me to use only 1-face videos as a training set and 2-face videos as a holdout test set. Why? Because we're given the original videos. Face detectors will generally work very well on the real faces. If you get all the real videos with a high probability of 1-face, you can assume their associated deepfakes will only have one face. So you can make a training set of real/fake videos where you only take the highest-probability box from the face detector. This would reduce noise drastically.</p>\n\n<p>This looked a lot of fun, I wish I'd had more time...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 793284,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "03/31/2020 23:27:49",
      "content": "<p>Definitely custom and clever data augmentation will be a huge contributor to top solutions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "784528": "### Pre-processing\n \nThis is soooo crucial. Really, I lost entire months because of this going from crappy models to crappy results. I assumed quantity over quality would be best, and that lots of faces would even out the few bad ones, but it was wrong. The noise introduced by bad faces is significant: it worsens everything in the pipeline: the training, the predictions... even with the best models, you'll get crap.\n\nI'm convinced pre-processing is actually more important than the model itself. I highly advise the mobile face detector of @unkownhihi and @harshitsheoran\n\nhttps://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison\n\nIt's really good, and reasonably fast.\n\n### Models\n\nI advice first starting with pre-trained known models. Like efficient net or https://keras.io/applications/ . Dunno why, but when training them, it seems better to initialize them with the pre-trained weights rather than random weights. Most of the time, when trying the latter, the training was just going nowhere. Dunno why. Perhaps the differences are too subtle and you need a meaningful starting point.\n\n### Optimizer\n\nOptimizers like Adam, Nadam and such tend, from my experience, to converge fairly faster. However, they tend to have more fluctuations in validation scores, apparently overfitting on some details.\n\nSimpler algorithms like SGD tend to converge slower, however, from my experience they tend to generalize better and have more stable validation scores.\n\n### Validation set\n\nFor validation, it is IMHO best and simplest to take a whole directory (or a few).\n\nSince the directories tend to have just a bunch of actors, this reduces the amount you have seen them elsewhere, making it a reasonable validation set.\n\nI personally picked the training directory 28 as validation set. The advantage is that you can verify your pipeline since the training data in the kaggle train/test sample of 400 videos is also from this directory.\n\n### About the difference between validation score and leaderboard score\n\nOne of the reason is probably a set of new actors appearing, making predictions less reliable. However, I think one other important reason is that the test set is \"augmented\".\n\nIn case you missed it, the test set is different than the training set, as mentioned here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013\n\nBasically, the test set is augmented as follows:\n\na) 2/9 of videos: reduce the FPS of the video to 15;\nb) 2/9 of videos: reduce the resolution of the video to 1/4 of its original size;\nc) 2/9 of videos: reduce the overall encoding quality.\nd) 3/9 of videos: unmodified\n\nI'm pretty sure that points (b) and (c) can \"destroy\" many predictions and worsen the score. I think it is likely that the excessive zooming and compression artifacts can easily be confused as fakes.\n\nTo boost the score, I suppose countermeasures must be taken. Especially point (c) is a bit tricky since we do not know how exactly the \"overall encoding quality\" was reduced.",
    "784544": "I'm sure that the compression affect the prediction, even 5%. I'm wondering what augmentation was applied by the top tiers.",
    "788874": "dagnelies  \n1)what do you mean by 2/9 here\n2) how fps affects teh video.. video would be too slow moving ?",
    "793036": "1) just \"2/9\" ...it's described in detail in the link's article\n2) the \"speed\" would be the same, it's just recorded at 15 FPS instead of 30 FPS",
    "793052": "I didn't get to participate much in this competition but I played around with the data a lot. It occurred to me to use only 1-face videos as a training set and 2-face videos as a holdout test set. Why? Because we're given the original videos. Face detectors will generally work very well on the real faces. If you get all the real videos with a high probability of 1-face, you can assume their associated deepfakes will only have one face. So you can make a training set of real/fake videos where you only take the highest-probability box from the face detector. This would reduce noise drastically.\n\nThis looked a lot of fun, I wish I'd had more time...",
    "793284": "Definitely custom and clever data augmentation will be a huge contributor to top solutions."
  },
  "source": "meta"
}