{
  "id": 133684,
  "title": "Face-detection based model limitations?",
  "url": "/competitions/deepfake-detection-challenge/discussion/133684",
  "author_name": "Yassine Alouini",
  "post_date": "2020-03-03T20:54:37.442000",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>For now, my main strategy has been the following:</p>\n\n<ul>\n<li>Extract and crop faces from videos</li>\n<li>Save these cropped faces </li>\n<li>Train a model to predict the FAKE or TRUE label using these cropped faces</li>\n</ul>\n\n<p>However, I have noticed that some hard to classify (by human-standards at least) FAKE videos contain \"artifacts\" outside the face. For example, the below screenshot has been extracted from the atrteuctxr.mp4 FAKE video (from part_0). I have circled in red what I think makes the video FAKE (let me know in the comments if I am missing something obvious).  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fa47e97cb006f02414033352318023597%2Fatrteuctxr_annotated.png?generation=1583268826562663&amp;alt=media\" alt=\"\"></p>\n\n<p>How do you deal with such samples? Do you add these regions? Maybe as metadata? Is there a way to learn what the deepfake generator is trying to do there? </p>\n\n<p>Thanks in advance for your tips and help!</p>",
  "messages": [
    {
      "id": 762874,
      "postDate": "2020-03-03T22:13:10.947Z",
      "content": "<p>take the frame from real video, take the frame from the fake video, lets say you are taking first frame from both of them, and using numpy see what thing make the difference, it is often seen that we can not just see with our eyes but the pixels were changed. We do have many videos where a fake face looks absolutely real, but when we compare using np array and look at the pixel values, things change, and most of the people are commenting everywhere in the discussion threads that sometimes the video is labeled \"fake\" but have all \"real\" faces, Not true my friends, its the pixels which change and not our vision.</p>",
      "rawMarkdown": "take the frame from real video, take the frame from the fake video, lets say you are taking first frame from both of them, and using numpy see what thing make the difference, it is often seen that we can not just see with our eyes but the pixels were changed. We do have many videos where a fake face looks absolutely real, but when we compare using np array and look at the pixel values, things change, and most of the people are commenting everywhere in the discussion threads that sometimes the video is labeled \"fake\" but have all \"real\" faces, Not true my friends, its the pixels which change and not our vision.",
      "votes": 5,
      "replies": [
        {
          "id": 762877,
          "postDate": "2020-03-03T22:20:12.917Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 762878,
          "postDate": "2020-03-03T22:23:22.470Z",
          "content": "<p>Not necessary, they can use different fps and quality, but it does not mean, so, then in this case, take a real frame, resize to say 512x512, get the bounding box for that, take the fake frame resize it to 512x512, crop the bounding box at exact same place as real video, then compare the faces.</p>",
          "rawMarkdown": "Not necessary, they can use different fps and quality, but it does not mean, so, then in this case, take a real frame, resize to say 512x512, get the bounding box for that, take the fake frame resize it to 512x512, crop the bounding box at exact same place as real video, then compare the faces."
        },
        {
          "id": 762881,
          "postDate": "2020-03-03T22:32:48.443Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 762882,
          "postDate": "2020-03-03T22:37:57.770Z",
          "content": "<p>I used that technique like in the very start of the competition and it did not work out for me. I am currently using a similar approach to code but a lot different in output.</p>",
          "rawMarkdown": "I used that technique like in the very start of the competition and it did not work out for me. I am currently using a similar approach to code but a lot different in output.",
          "votes": 1
        },
        {
          "id": 762889,
          "postDate": "2020-03-03T22:53:21.753Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 763071,
          "postDate": "2020-03-04T05:30:16.870Z",
          "content": "<p>I don't believe translating bounding boxes from original to fake frame is a good idea as it assumes the face detector will detect the same bounding box on fake frames.</p>",
          "rawMarkdown": "I don't believe translating bounding boxes from original to fake frame is a good idea as it assumes the face detector will detect the same bounding box on fake frames."
        },
        {
          "id": 763122,
          "postDate": "2020-03-04T06:50:03.390Z",
          "content": "<p><a href=\"/maralski\">@maralski</a> that is true, but that was just for testing purposes.</p>",
          "rawMarkdown": "@maralski that is true, but that was just for testing purposes."
        }
      ]
    },
    {
      "id": 762831,
      "postDate": "2020-03-03T20:54:37.443Z",
      "content": "<p>For now, my main strategy has been the following:</p>\n\n<ul>\n<li>Extract and crop faces from videos</li>\n<li>Save these cropped faces </li>\n<li>Train a model to predict the FAKE or TRUE label using these cropped faces</li>\n</ul>\n\n<p>However, I have noticed that some hard to classify (by human-standards at least) FAKE videos contain \"artifacts\" outside the face. For example, the below screenshot has been extracted from the atrteuctxr.mp4 FAKE video (from part_0). I have circled in red what I think makes the video FAKE (let me know in the comments if I am missing something obvious).  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fa47e97cb006f02414033352318023597%2Fatrteuctxr_annotated.png?generation=1583268826562663&amp;alt=media\" alt=\"\"></p>\n\n<p>How do you deal with such samples? Do you add these regions? Maybe as metadata? Is there a way to learn what the deepfake generator is trying to do there? </p>\n\n<p>Thanks in advance for your tips and help!</p>",
      "rawMarkdown": "For now, my main strategy has been the following:\n\n- Extract and crop faces from videos\n- Save these cropped faces \n- Train a model to predict the FAKE or TRUE label using these cropped faces\n\nHowever, I have noticed that some hard to classify (by human-standards at least) FAKE videos contain \"artifacts\" outside the face. For example, the below screenshot has been extracted from the atrteuctxr.mp4 FAKE video (from part_0). I have circled in red what I think makes the video FAKE (let me know in the comments if I am missing something obvious).  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fa47e97cb006f02414033352318023597%2Fatrteuctxr_annotated.png?generation=1583268826562663&amp;alt=media)\n\n\n\nHow do you deal with such samples? Do you add these regions? Maybe as metadata? Is there a way to learn what the deepfake generator is trying to do there? \n\nThanks in advance for your tips and help!\n",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 762874,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2020-03-03T22:13:10.947000",
      "content": "<p>take the frame from real video, take the frame from the fake video, lets say you are taking first frame from both of them, and using numpy see what thing make the difference, it is often seen that we can not just see with our eyes but the pixels were changed. We do have many videos where a fake face looks absolutely real, but when we compare using np array and look at the pixel values, things change, and most of the people are commenting everywhere in the discussion threads that sometimes the video is labeled \"fake\" but have all \"real\" faces, Not true my friends, its the pixels which change and not our vision.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 762877,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-03T22:20:12.917000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762878,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-03T22:23:22.470000",
          "content": "<p>Not necessary, they can use different fps and quality, but it does not mean, so, then in this case, take a real frame, resize to say 512x512, get the bounding box for that, take the fake frame resize it to 512x512, crop the bounding box at exact same place as real video, then compare the faces.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762881,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-03T22:32:48.443000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 762882,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-03T22:37:57.770000",
          "content": "<p>I used that technique like in the very start of the competition and it did not work out for me. I am currently using a similar approach to code but a lot different in output.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762889,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-03T22:53:21.753000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763071,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "2020-03-04T05:30:16.870000",
          "content": "<p>I don't believe translating bounding boxes from original to fake frame is a good idea as it assumes the face detector will detect the same bounding box on fake frames.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763122,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-04T06:50:03.390000",
          "content": "<p><a href=\"/maralski\">@maralski</a> that is true, but that was just for testing purposes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "762874": "take the frame from real video, take the frame from the fake video, lets say you are taking first frame from both of them, and using numpy see what thing make the difference, it is often seen that we can not just see with our eyes but the pixels were changed. We do have many videos where a fake face looks absolutely real, but when we compare using np array and look at the pixel values, things change, and most of the people are commenting everywhere in the discussion threads that sometimes the video is labeled \"fake\" but have all \"real\" faces, Not true my friends, its the pixels which change and not our vision.",
    "762831": "For now, my main strategy has been the following:\n\n- Extract and crop faces from videos\n- Save these cropped faces \n- Train a model to predict the FAKE or TRUE label using these cropped faces\n\nHowever, I have noticed that some hard to classify (by human-standards at least) FAKE videos contain \"artifacts\" outside the face. For example, the below screenshot has been extracted from the atrteuctxr.mp4 FAKE video (from part_0). I have circled in red what I think makes the video FAKE (let me know in the comments if I am missing something obvious).  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fa47e97cb006f02414033352318023597%2Fatrteuctxr_annotated.png?generation=1583268826562663&amp;alt=media)\n\n\n\nHow do you deal with such samples? Do you add these regions? Maybe as metadata? Is there a way to learn what the deepfake generator is trying to do there? \n\nThanks in advance for your tips and help!\n"
  }
}