{
  "id": 134733,
  "title": "The fight for generalization",
  "url": "/competitions/deepfake-detection-challenge/discussion/134733",
  "author_name": "ryches",
  "post_date": "2020-03-10T03:38:51.308000",
  "votes": 30,
  "comment_count": 11,
  "views": 0,
  "content": "<p>One of the things I've seen people discussing is the gap between train, validation and leaderboard scores. I wanted to pose a hypothesis as to what exactly is going on with this competition. My belief is that many models are failing by devolving into a facial recognition system. </p>\n\n<p>As discussed here <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832</a> and in the original DFDC paper a finite pool of actors were used for the videos we have been provided for this competition. The problem that arises is that we need to generalize to unseen faces outside of the actors and the videos they created. We need a method that locates the properties of fakes rather than simply recognizing faces and putting them in a fake and real bin. </p>\n\n<p>This likely happens because the faces in the real videos don't necessarily overlap with those that were swapped in for the fakes. So that's to say the model might be learning faces present in the real videos vs faces present in the fake videos. That means regardless of whether you split based on original source video ID or folder or even face embedding of the reals you will likely fall trap to the mismatch between real and fake. </p>\n\n<p>This likely explains the mismatch between validation and test and it likely also explains the rapid overfitting many people are encountering. It is also likely the case that any augmentation steps and repeated exposures to the same faces will possibly only strengthen our model's knowledge of those specific faces rather than help generalization.</p>\n\n<p>So the question is how to combat this and create a model that generalizes beyond just learning a few faces. This will be crucial for the private test set when our models are exposed to non-actor videos.  What I have seen from the discussion of some of the top-ranking people is that they are likely doing something smart with either selecting a good subset of the videos, making sure that there isnt repeated exposure of faces, or selecting only a subset of frames which are good unique viewpoints and preventing overexposure to certain videos. Alternatively, there may be some way to anonymize the faces and prevent models from learning facial features and learn fake characteristics instead. </p>",
  "messages": [
    {
      "id": 767741,
      "postDate": "2020-03-10T03:38:51.307Z",
      "content": "<p>One of the things I've seen people discussing is the gap between train, validation and leaderboard scores. I wanted to pose a hypothesis as to what exactly is going on with this competition. My belief is that many models are failing by devolving into a facial recognition system. </p>\n\n<p>As discussed here <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832</a> and in the original DFDC paper a finite pool of actors were used for the videos we have been provided for this competition. The problem that arises is that we need to generalize to unseen faces outside of the actors and the videos they created. We need a method that locates the properties of fakes rather than simply recognizing faces and putting them in a fake and real bin. </p>\n\n<p>This likely happens because the faces in the real videos don't necessarily overlap with those that were swapped in for the fakes. So that's to say the model might be learning faces present in the real videos vs faces present in the fake videos. That means regardless of whether you split based on original source video ID or folder or even face embedding of the reals you will likely fall trap to the mismatch between real and fake. </p>\n\n<p>This likely explains the mismatch between validation and test and it likely also explains the rapid overfitting many people are encountering. It is also likely the case that any augmentation steps and repeated exposures to the same faces will possibly only strengthen our model's knowledge of those specific faces rather than help generalization.</p>\n\n<p>So the question is how to combat this and create a model that generalizes beyond just learning a few faces. This will be crucial for the private test set when our models are exposed to non-actor videos.  What I have seen from the discussion of some of the top-ranking people is that they are likely doing something smart with either selecting a good subset of the videos, making sure that there isnt repeated exposure of faces, or selecting only a subset of frames which are good unique viewpoints and preventing overexposure to certain videos. Alternatively, there may be some way to anonymize the faces and prevent models from learning facial features and learn fake characteristics instead. </p>",
      "rawMarkdown": "One of the things I've seen people discussing is the gap between train, validation and leaderboard scores. I wanted to pose a hypothesis as to what exactly is going on with this competition. My belief is that many models are failing by devolving into a facial recognition system. \n\nAs discussed here https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832 and in the original DFDC paper a finite pool of actors were used for the videos we have been provided for this competition. The problem that arises is that we need to generalize to unseen faces outside of the actors and the videos they created. We need a method that locates the properties of fakes rather than simply recognizing faces and putting them in a fake and real bin. \n\nThis likely happens because the faces in the real videos don't necessarily overlap with those that were swapped in for the fakes. So that's to say the model might be learning faces present in the real videos vs faces present in the fake videos. That means regardless of whether you split based on original source video ID or folder or even face embedding of the reals you will likely fall trap to the mismatch between real and fake. \n\nThis likely explains the mismatch between validation and test and it likely also explains the rapid overfitting many people are encountering. It is also likely the case that any augmentation steps and repeated exposures to the same faces will possibly only strengthen our model's knowledge of those specific faces rather than help generalization.\n\nSo the question is how to combat this and create a model that generalizes beyond just learning a few faces. This will be crucial for the private test set when our models are exposed to non-actor videos.  What I have seen from the discussion of some of the top-ranking people is that they are likely doing something smart with either selecting a good subset of the videos, making sure that there isnt repeated exposure of faces, or selecting only a subset of frames which are good unique viewpoints and preventing overexposure to certain videos. Alternatively, there may be some way to anonymize the faces and prevent models from learning facial features and learn fake characteristics instead. ",
      "votes": 30
    },
    {
      "id": 767934,
      "postDate": "2020-03-10T09:15:03.177Z",
      "content": "<p>It turns out my image classifier model is good at detecting really bad fakes, such as faces with huge blurs on them. It's not good at detecting more subtle fakes. The training data contains a lot of bad fakes that won't fool anyone, but I'm assuming the test data is of higher quality.</p>",
      "rawMarkdown": "It turns out my image classifier model is good at detecting really bad fakes, such as faces with huge blurs on them. It's not good at detecting more subtle fakes. The training data contains a lot of bad fakes that won't fool anyone, but I'm assuming the test data is of higher quality.",
      "votes": 10,
      "replies": [
        {
          "id": 768366,
          "postDate": "2020-03-10T17:38:38.633Z",
          "content": "<p>Interesting. I did some error analysis and found the largest contributors to my loss on the validation set seemed to primarily be coming from videos where the fake was flickering on and off and when there were two faces on screen</p>",
          "rawMarkdown": "Interesting. I did some error analysis and found the largest contributors to my loss on the validation set seemed to primarily be coming from videos where the fake was flickering on and off and when there were two faces on screen",
          "votes": 7
        },
        {
          "id": 768520,
          "postDate": "2020-03-10T22:53:25.817Z",
          "content": "<p>Yes. If there are multiple faces taking the average of all faces doesn't make sense as your likely to end up with 0.5. But there is another approach. Give it some thought.</p>",
          "rawMarkdown": "Yes. If there are multiple faces taking the average of all faces doesn't make sense as your likely to end up with 0.5. But there is another approach. Give it some thought.",
          "votes": 1
        }
      ]
    },
    {
      "id": 767761,
      "postDate": "2020-03-10T04:14:49.323Z",
      "content": "<p>I am building an LGBM model for this competition because of overfitting. It may be unusual but I hope it'll work.</p>",
      "rawMarkdown": "I am building an LGBM model for this competition because of overfitting. It may be unusual but I hope it'll work.",
      "votes": 6
    },
    {
      "id": 769647,
      "postDate": "2020-03-12T05:36:07.727Z",
      "content": "<p>Do we really need the real and corresponding fake videos for training? \nCan the real and fake videos be independent of each other?</p>",
      "rawMarkdown": "Do we really need the real and corresponding fake videos for training? \nCan the real and fake videos be independent of each other?",
      "replies": [
        {
          "id": 769738,
          "postDate": "2020-03-12T07:55:04.977Z",
          "content": "<p>depends... if you are using faces for training, the only way to know which face is fake is by using the real/fake pair</p>",
          "rawMarkdown": "depends... if you are using faces for training, the only way to know which face is fake is by using the real/fake pair"
        },
        {
          "id": 770279,
          "postDate": "2020-03-12T18:54:56.757Z",
          "content": "<p>Hi - when you say that you are training with a real/fake pair. Are you modifing the dataloader to select a real/fake pair or is it done when the dataset is initialized?  I am speaking in terms of pytorch implementation.</p>",
          "rawMarkdown": "Hi - when you say that you are training with a real/fake pair. Are you modifing the dataloader to select a real/fake pair or is it done when the dataset is initialized?  I am speaking in terms of pytorch implementation."
        },
        {
          "id": 770461,
          "postDate": "2020-03-13T01:25:36.507Z",
          "content": "<p>Some of the videos have more then one face. If you just crop faces and mark them real/fake according to the video label, you will get bad training data. The only way to know which face is actually fake is by comparing the frames from the real and fake video. its not exactly trivial as there is a lot of compression noise but doable</p>",
          "rawMarkdown": "Some of the videos have more then one face. If you just crop faces and mark them real/fake according to the video label, you will get bad training data. The only way to know which face is actually fake is by comparing the frames from the real and fake video. its not exactly trivial as there is a lot of compression noise but doable"
        },
        {
          "id": 770501,
          "postDate": "2020-03-13T02:54:49.157Z",
          "content": "<p>its not exactly trivial as there is a lot of compression noise but doable -&gt; Are you referring to psnr, ssim and nrmse on the extracted faces along this direction?</p>",
          "rawMarkdown": "its not exactly trivial as there is a lot of compression noise but doable -&gt; Are you referring to psnr, ssim and nrmse on the extracted faces along this direction?"
        },
        {
          "id": 770625,
          "postDate": "2020-03-13T07:01:00.953Z",
          "content": "<p>I have no idea what are all these 4 letters :)</p>\n\n<p>You just subtract or absdiff the real from the fake. You get almost zero everywhere other than the fake face</p>",
          "rawMarkdown": "I have no idea what are all these 4 letters :)\n\nYou just subtract or absdiff the real from the fake. You get almost zero everywhere other than the fake face",
          "votes": 2
        }
      ]
    },
    {
      "id": 768061,
      "postDate": "2020-03-10T12:11:33.613Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 767934,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2020-03-10T09:15:03.177000",
      "content": "<p>It turns out my image classifier model is good at detecting really bad fakes, such as faces with huge blurs on them. It's not good at detecting more subtle fakes. The training data contains a lot of bad fakes that won't fool anyone, but I'm assuming the test data is of higher quality.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 768366,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-03-10T17:38:38.633000",
          "content": "<p>Interesting. I did some error analysis and found the largest contributors to my loss on the validation set seemed to primarily be coming from videos where the fake was flickering on and off and when there were two faces on screen</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 768520,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "2020-03-10T22:53:25.817000",
          "content": "<p>Yes. If there are multiple faces taking the average of all faces doesn't make sense as your likely to end up with 0.5. But there is another approach. Give it some thought.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 767761,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-03-10T04:14:49.323000",
      "content": "<p>I am building an LGBM model for this competition because of overfitting. It may be unusual but I hope it'll work.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 769647,
      "author_name": "A*DF",
      "author_url": "",
      "post_date": "2020-03-12T05:36:07.727000",
      "content": "<p>Do we really need the real and corresponding fake videos for training? \nCan the real and fake videos be independent of each other?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 769738,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-12T07:55:04.977000",
          "content": "<p>depends... if you are using faces for training, the only way to know which face is fake is by using the real/fake pair</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770279,
          "author_name": "SkyLord",
          "author_url": "",
          "post_date": "2020-03-12T18:54:56.757000",
          "content": "<p>Hi - when you say that you are training with a real/fake pair. Are you modifing the dataloader to select a real/fake pair or is it done when the dataset is initialized?  I am speaking in terms of pytorch implementation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770461,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-13T01:25:36.507000",
          "content": "<p>Some of the videos have more then one face. If you just crop faces and mark them real/fake according to the video label, you will get bad training data. The only way to know which face is actually fake is by comparing the frames from the real and fake video. its not exactly trivial as there is a lot of compression noise but doable</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770501,
          "author_name": "A*DF",
          "author_url": "",
          "post_date": "2020-03-13T02:54:49.157000",
          "content": "<p>its not exactly trivial as there is a lot of compression noise but doable -&gt; Are you referring to psnr, ssim and nrmse on the extracted faces along this direction?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770625,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-13T07:01:00.953000",
          "content": "<p>I have no idea what are all these 4 letters :)</p>\n\n<p>You just subtract or absdiff the real from the fake. You get almost zero everywhere other than the fake face</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 768061,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-10T12:11:33.613000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "767741": "One of the things I've seen people discussing is the gap between train, validation and leaderboard scores. I wanted to pose a hypothesis as to what exactly is going on with this competition. My belief is that many models are failing by devolving into a facial recognition system. \n\nAs discussed here https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832 and in the original DFDC paper a finite pool of actors were used for the videos we have been provided for this competition. The problem that arises is that we need to generalize to unseen faces outside of the actors and the videos they created. We need a method that locates the properties of fakes rather than simply recognizing faces and putting them in a fake and real bin. \n\nThis likely happens because the faces in the real videos don't necessarily overlap with those that were swapped in for the fakes. So that's to say the model might be learning faces present in the real videos vs faces present in the fake videos. That means regardless of whether you split based on original source video ID or folder or even face embedding of the reals you will likely fall trap to the mismatch between real and fake. \n\nThis likely explains the mismatch between validation and test and it likely also explains the rapid overfitting many people are encountering. It is also likely the case that any augmentation steps and repeated exposures to the same faces will possibly only strengthen our model's knowledge of those specific faces rather than help generalization.\n\nSo the question is how to combat this and create a model that generalizes beyond just learning a few faces. This will be crucial for the private test set when our models are exposed to non-actor videos.  What I have seen from the discussion of some of the top-ranking people is that they are likely doing something smart with either selecting a good subset of the videos, making sure that there isnt repeated exposure of faces, or selecting only a subset of frames which are good unique viewpoints and preventing overexposure to certain videos. Alternatively, there may be some way to anonymize the faces and prevent models from learning facial features and learn fake characteristics instead. ",
    "767934": "It turns out my image classifier model is good at detecting really bad fakes, such as faces with huge blurs on them. It's not good at detecting more subtle fakes. The training data contains a lot of bad fakes that won't fool anyone, but I'm assuming the test data is of higher quality.",
    "767761": "I am building an LGBM model for this competition because of overfitting. It may be unusual but I hope it'll work.",
    "769647": "Do we really need the real and corresponding fake videos for training? \nCan the real and fake videos be independent of each other?",
    "768061": ""
  }
}