{
  "id": 123000,
  "title": "Possible Mislabelled data?",
  "url": "/competitions/deepfake-detection-challenge/discussion/123000",
  "author_name": "Nikhil Peri",
  "post_date": "2019-12-24T03:47:09.896000",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I found the following example of what I believe to be a mislabeled sample.</p>\n\n<p>in <code>dfdc_train_part_3/gfcspejcib.mp4</code> i found no significant image or audio difference between the faces of the two actors.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2128634%2F243ee7a3611a1f52c43a128026c61c19%2FScreen%20Shot%202019-12-23%20at%2010.42.54%20PM.png?generation=1577158996107693&amp;alt=media\" alt=\"\">\n^ this is a sample comparing both actors from the real and the fake</p>\n\n<p>the mean absolute pixel difference between the faces in real and fake is ~110 which is fairly small when compared to other faked videos (typically in the ~150 range)...note due to compression some change should always be expected</p>\n\n<p>Attached is <code>gfcspejcib.mp4</code> and a few of the \"fakes\" let me know if you see a difference because i cannot see it, or measure it with my metrics ... I will be running a pipeline to find more of these \"real fakes\" I think it will be important to filter these out of the training set</p>",
  "messages": [
    {
      "id": 704703,
      "postDate": "2019-12-27T21:41:10.740Z",
      "content": "<p>Hello everyone,</p>\n\n<p>In a dataset of such size, it is expected that, sometimes, training data may contain noise. In this dataset, noise may come from a number of sources and among them there may be failures in the generation of face swaps. The amount of failed swaps is expected to be very low and the participants are expected to keep this in mind. In some other cases, face swap may be subtle and, in some cases, almost imperceptible. </p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hello everyone,\n\nIn a dataset of such size, it is expected that, sometimes, training data may contain noise. In this dataset, noise may come from a number of sources and among them there may be failures in the generation of face swaps. The amount of failed swaps is expected to be very low and the participants are expected to keep this in mind. In some other cases, face swap may be subtle and, in some cases, almost imperceptible. \n\nThanks!",
      "votes": 8,
      "replies": [
        {
          "id": 704875,
          "postDate": "2019-12-28T05:46:25.547Z",
          "content": "<p>Totally agree ;-) That's why I personally started by cleaning up data.</p>",
          "rawMarkdown": "Totally agree ;-) That's why I personally started by cleaning up data."
        },
        {
          "id": 771540,
          "postDate": "2020-03-14T09:49:49.873Z",
          "content": "<p>How do you clean up the data?</p>",
          "rawMarkdown": "How do you clean up the data?",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 701920,
      "postDate": "2019-12-24T03:47:09.897Z",
      "content": "<p>I found the following example of what I believe to be a mislabeled sample.</p>\n\n<p>in <code>dfdc_train_part_3/gfcspejcib.mp4</code> i found no significant image or audio difference between the faces of the two actors.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2128634%2F243ee7a3611a1f52c43a128026c61c19%2FScreen%20Shot%202019-12-23%20at%2010.42.54%20PM.png?generation=1577158996107693&amp;alt=media\" alt=\"\">\n^ this is a sample comparing both actors from the real and the fake</p>\n\n<p>the mean absolute pixel difference between the faces in real and fake is ~110 which is fairly small when compared to other faked videos (typically in the ~150 range)...note due to compression some change should always be expected</p>\n\n<p>Attached is <code>gfcspejcib.mp4</code> and a few of the \"fakes\" let me know if you see a difference because i cannot see it, or measure it with my metrics ... I will be running a pipeline to find more of these \"real fakes\" I think it will be important to filter these out of the training set</p>",
      "rawMarkdown": "I found the following example of what I believe to be a mislabeled sample.\n\nin `dfdc_train_part_3/gfcspejcib.mp4` i found no significant image or audio difference between the faces of the two actors.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2128634%2F243ee7a3611a1f52c43a128026c61c19%2FScreen%20Shot%202019-12-23%20at%2010.42.54%20PM.png?generation=1577158996107693&amp;alt=media)\n^ this is a sample comparing both actors from the real and the fake\n\nthe mean absolute pixel difference between the faces in real and fake is ~110 which is fairly small when compared to other faked videos (typically in the ~150 range)...note due to compression some change should always be expected\n\nAttached is `gfcspejcib.mp4` and a few of the \"fakes\" let me know if you see a difference because i cannot see it, or measure it with my metrics ... I will be running a pipeline to find more of these \"real fakes\" I think it will be important to filter these out of the training set",
      "votes": 5
    },
    {
      "id": 703493,
      "postDate": "2019-12-26T08:50:00.280Z",
      "content": "<p>I could probably add some myself.</p>\n\n<p>I'm doing frame comparison using ssim, psnr and nrmse and look at odd results.</p>\n\n<p>Folder <code>dfdc_train_part_40</code>: \nreal: upowrtzdno.mp4\nfake: kgsszrmscq.mp4</p>\n\n<p>real: hioffslzjp.mp4\nfake: zmkpfhjxru.mp4</p>\n\n<p>real: srmsbgbeqi.mp4\nfake: rnoaglzwbo.mp4</p>\n\n<p>real: vnqpzwmohz.mp4\nfake: nyayqkqtre.mp4</p>\n\n<p>real: rzqzpdcust.mp4\nfake: sabdeisrai.mp4</p>\n\n<p>etc ...\nAll of the clips of this actor do not seem to be fake. \nThe only difference I've notices is that there's a slight out-of-sync between frames of original and the one from fake fake.  But that doesn't make a video fake ... </p>\n\n<p>I could go on, but I would be interesting what other findings do you guys have.</p>\n\n<p>Later edit:\nBe careful about videos like the following from folder <code>dfdc_train_part_1</code>:\nreal: rmhsahyvta.mp4\nfakes: \nrmanmzhjkm.mp4\nhjezlaklbj.mp4\nyydzfachgi.mp4\ncloymwmsgb.mp4\nuvlynrnmgb.mp4\nnlbbatlhwh.mp4\nksdqodntwc.mp4\naelsfznuqw.mp4\nwjewnsibqo.mp4\npxvxurfyuy.mp4\ndztvxrzuta.mp4\nnsvdrlugkt.mp4\nlwaavueeng.mp4\nmuzcmfrzqa.mp4\ntrhqrdommy.mp4\nseonqpzbco.mp4\neuqtwhhwwr.mp4\nalfidaevix.mp4\nyhqwrblsvr.mp4\nogxilhjehh.mp4\nyitsrqeiih.mp4\nwrijnggwum.mp4</p>\n\n<p>Where just a few frames have changes ... I wouldn't consider them to bee \"deep fake\" videos ... but yet because they have a few blurry face frames they seem to be in that category. </p>",
      "rawMarkdown": "I could probably add some myself.\n\nI'm doing frame comparison using ssim, psnr and nrmse and look at odd results.\n\nFolder `dfdc_train_part_40`: \nreal: upowrtzdno.mp4\nfake: kgsszrmscq.mp4\n\nreal: hioffslzjp.mp4\nfake: zmkpfhjxru.mp4\n\nreal: srmsbgbeqi.mp4\nfake: rnoaglzwbo.mp4\n\nreal: vnqpzwmohz.mp4\nfake: nyayqkqtre.mp4\n\nreal: rzqzpdcust.mp4\nfake: sabdeisrai.mp4\n\netc ...\nAll of the clips of this actor do not seem to be fake. \nThe only difference I've notices is that there's a slight out-of-sync between frames of original and the one from fake fake.  But that doesn't make a video fake ... \n\nI could go on, but I would be interesting what other findings do you guys have.\n\n\nLater edit:\nBe careful about videos like the following from folder `dfdc_train_part_1`:\nreal: rmhsahyvta.mp4\nfakes: \nrmanmzhjkm.mp4\nhjezlaklbj.mp4\nyydzfachgi.mp4\ncloymwmsgb.mp4\nuvlynrnmgb.mp4\nnlbbatlhwh.mp4\nksdqodntwc.mp4\naelsfznuqw.mp4\nwjewnsibqo.mp4\npxvxurfyuy.mp4\ndztvxrzuta.mp4\nnsvdrlugkt.mp4\nlwaavueeng.mp4\nmuzcmfrzqa.mp4\ntrhqrdommy.mp4\nseonqpzbco.mp4\neuqtwhhwwr.mp4\nalfidaevix.mp4\nyhqwrblsvr.mp4\nogxilhjehh.mp4\nyitsrqeiih.mp4\nwrijnggwum.mp4\n\nWhere just a few frames have changes ... I wouldn't consider them to bee \"deep fake\" videos ... but yet because they have a few blurry face frames they seem to be in that category. ",
      "votes": 3,
      "replies": [
        {
          "id": 704710,
          "postDate": "2019-12-27T22:06:17.407Z",
          "content": "<p>Remember that fake doesn't just mean video, it can be audio as well</p>",
          "rawMarkdown": "Remember that fake doesn't just mean video, it can be audio as well",
          "votes": 1
        },
        {
          "id": 704884,
          "postDate": "2019-12-28T06:00:21.287Z",
          "content": "<p>Yeap. Somehow I feel is going to be a bit more difficult in figuring out that the audio is \"fake\" when the image is not ... Hopefully I'll manage to find more than enough videos with only audio \"fake\" to be able to make something out of them.</p>",
          "rawMarkdown": "Yeap. Somehow I feel is going to be a bit more difficult in figuring out that the audio is \"fake\" when the image is not ... Hopefully I'll manage to find more than enough videos with only audio \"fake\" to be able to make something out of them."
        }
      ]
    },
    {
      "id": 765534,
      "postDate": "2020-03-06T18:51:27.473Z",
      "content": "<p>As far as I have seen, deepfake techniques usually manipulate either the face or audio in the input video frames. That's why training on cropped faces makes sense. While I have seen videos in the training dataset where some random patch in the video appears to be blur (labelled as FAKE). Should we expect same cases in private test set? I mean not only the face or audio, but a random patch in the frame is manipulated (blurred). I really appreciate any comment on this. While training I can consider these videos as outliers, but if the same trend follows in test set then I may have to retrain my model with different techniques. Can anybody please tell me if the test set DeepFakes are really DeepFakes?</p>",
      "rawMarkdown": "As far as I have seen, deepfake techniques usually manipulate either the face or audio in the input video frames. That's why training on cropped faces makes sense. While I have seen videos in the training dataset where some random patch in the video appears to be blur (labelled as FAKE). Should we expect same cases in private test set? I mean not only the face or audio, but a random patch in the frame is manipulated (blurred). I really appreciate any comment on this. While training I can consider these videos as outliers, but if the same trend follows in test set then I may have to retrain my model with different techniques. Can anybody please tell me if the test set DeepFakes are really DeepFakes?"
    }
  ],
  "comments": [
    {
      "id": 704703,
      "author_name": "Mozaic",
      "author_url": "",
      "post_date": "2019-12-27T21:41:10.740000",
      "content": "<p>Hello everyone,</p>\n\n<p>In a dataset of such size, it is expected that, sometimes, training data may contain noise. In this dataset, noise may come from a number of sources and among them there may be failures in the generation of face swaps. The amount of failed swaps is expected to be very low and the participants are expected to keep this in mind. In some other cases, face swap may be subtle and, in some cases, almost imperceptible. </p>\n\n<p>Thanks!</p>",
      "votes": 8,
      "replies": [
        {
          "id": 704875,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2019-12-28T05:46:25.547000",
          "content": "<p>Totally agree ;-) That's why I personally started by cleaning up data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 771540,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-14T09:49:49.873000",
          "content": "<p>How do you clean up the data?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 703493,
      "author_name": "Mihai Cvasnievschi",
      "author_url": "",
      "post_date": "2019-12-26T08:50:00.280000",
      "content": "<p>I could probably add some myself.</p>\n\n<p>I'm doing frame comparison using ssim, psnr and nrmse and look at odd results.</p>\n\n<p>Folder <code>dfdc_train_part_40</code>: \nreal: upowrtzdno.mp4\nfake: kgsszrmscq.mp4</p>\n\n<p>real: hioffslzjp.mp4\nfake: zmkpfhjxru.mp4</p>\n\n<p>real: srmsbgbeqi.mp4\nfake: rnoaglzwbo.mp4</p>\n\n<p>real: vnqpzwmohz.mp4\nfake: nyayqkqtre.mp4</p>\n\n<p>real: rzqzpdcust.mp4\nfake: sabdeisrai.mp4</p>\n\n<p>etc ...\nAll of the clips of this actor do not seem to be fake. \nThe only difference I've notices is that there's a slight out-of-sync between frames of original and the one from fake fake.  But that doesn't make a video fake ... </p>\n\n<p>I could go on, but I would be interesting what other findings do you guys have.</p>\n\n<p>Later edit:\nBe careful about videos like the following from folder <code>dfdc_train_part_1</code>:\nreal: rmhsahyvta.mp4\nfakes: \nrmanmzhjkm.mp4\nhjezlaklbj.mp4\nyydzfachgi.mp4\ncloymwmsgb.mp4\nuvlynrnmgb.mp4\nnlbbatlhwh.mp4\nksdqodntwc.mp4\naelsfznuqw.mp4\nwjewnsibqo.mp4\npxvxurfyuy.mp4\ndztvxrzuta.mp4\nnsvdrlugkt.mp4\nlwaavueeng.mp4\nmuzcmfrzqa.mp4\ntrhqrdommy.mp4\nseonqpzbco.mp4\neuqtwhhwwr.mp4\nalfidaevix.mp4\nyhqwrblsvr.mp4\nogxilhjehh.mp4\nyitsrqeiih.mp4\nwrijnggwum.mp4</p>\n\n<p>Where just a few frames have changes ... I wouldn't consider them to bee \"deep fake\" videos ... but yet because they have a few blurry face frames they seem to be in that category. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 704710,
          "author_name": "David Austin",
          "author_url": "",
          "post_date": "2019-12-27T22:06:17.407000",
          "content": "<p>Remember that fake doesn't just mean video, it can be audio as well</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 704884,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2019-12-28T06:00:21.287000",
          "content": "<p>Yeap. Somehow I feel is going to be a bit more difficult in figuring out that the audio is \"fake\" when the image is not ... Hopefully I'll manage to find more than enough videos with only audio \"fake\" to be able to make something out of them.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 765534,
      "author_name": "Aditya Killekar",
      "author_url": "",
      "post_date": "2020-03-06T18:51:27.473000",
      "content": "<p>As far as I have seen, deepfake techniques usually manipulate either the face or audio in the input video frames. That's why training on cropped faces makes sense. While I have seen videos in the training dataset where some random patch in the video appears to be blur (labelled as FAKE). Should we expect same cases in private test set? I mean not only the face or audio, but a random patch in the frame is manipulated (blurred). I really appreciate any comment on this. While training I can consider these videos as outliers, but if the same trend follows in test set then I may have to retrain my model with different techniques. Can anybody please tell me if the test set DeepFakes are really DeepFakes?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "704703": "Hello everyone,\n\nIn a dataset of such size, it is expected that, sometimes, training data may contain noise. In this dataset, noise may come from a number of sources and among them there may be failures in the generation of face swaps. The amount of failed swaps is expected to be very low and the participants are expected to keep this in mind. In some other cases, face swap may be subtle and, in some cases, almost imperceptible. \n\nThanks!",
    "701920": "I found the following example of what I believe to be a mislabeled sample.\n\nin `dfdc_train_part_3/gfcspejcib.mp4` i found no significant image or audio difference between the faces of the two actors.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2128634%2F243ee7a3611a1f52c43a128026c61c19%2FScreen%20Shot%202019-12-23%20at%2010.42.54%20PM.png?generation=1577158996107693&amp;alt=media)\n^ this is a sample comparing both actors from the real and the fake\n\nthe mean absolute pixel difference between the faces in real and fake is ~110 which is fairly small when compared to other faked videos (typically in the ~150 range)...note due to compression some change should always be expected\n\nAttached is `gfcspejcib.mp4` and a few of the \"fakes\" let me know if you see a difference because i cannot see it, or measure it with my metrics ... I will be running a pipeline to find more of these \"real fakes\" I think it will be important to filter these out of the training set",
    "703493": "I could probably add some myself.\n\nI'm doing frame comparison using ssim, psnr and nrmse and look at odd results.\n\nFolder `dfdc_train_part_40`: \nreal: upowrtzdno.mp4\nfake: kgsszrmscq.mp4\n\nreal: hioffslzjp.mp4\nfake: zmkpfhjxru.mp4\n\nreal: srmsbgbeqi.mp4\nfake: rnoaglzwbo.mp4\n\nreal: vnqpzwmohz.mp4\nfake: nyayqkqtre.mp4\n\nreal: rzqzpdcust.mp4\nfake: sabdeisrai.mp4\n\netc ...\nAll of the clips of this actor do not seem to be fake. \nThe only difference I've notices is that there's a slight out-of-sync between frames of original and the one from fake fake.  But that doesn't make a video fake ... \n\nI could go on, but I would be interesting what other findings do you guys have.\n\n\nLater edit:\nBe careful about videos like the following from folder `dfdc_train_part_1`:\nreal: rmhsahyvta.mp4\nfakes: \nrmanmzhjkm.mp4\nhjezlaklbj.mp4\nyydzfachgi.mp4\ncloymwmsgb.mp4\nuvlynrnmgb.mp4\nnlbbatlhwh.mp4\nksdqodntwc.mp4\naelsfznuqw.mp4\nwjewnsibqo.mp4\npxvxurfyuy.mp4\ndztvxrzuta.mp4\nnsvdrlugkt.mp4\nlwaavueeng.mp4\nmuzcmfrzqa.mp4\ntrhqrdommy.mp4\nseonqpzbco.mp4\neuqtwhhwwr.mp4\nalfidaevix.mp4\nyhqwrblsvr.mp4\nogxilhjehh.mp4\nyitsrqeiih.mp4\nwrijnggwum.mp4\n\nWhere just a few frames have changes ... I wouldn't consider them to bee \"deep fake\" videos ... but yet because they have a few blurry face frames they seem to be in that category. ",
    "765534": "As far as I have seen, deepfake techniques usually manipulate either the face or audio in the input video frames. That's why training on cropped faces makes sense. While I have seen videos in the training dataset where some random patch in the video appears to be blur (labelled as FAKE). Should we expect same cases in private test set? I mean not only the face or audio, but a random patch in the frame is manipulated (blurred). I really appreciate any comment on this. While training I can consider these videos as outliers, but if the same trend follows in test set then I may have to retrain my model with different techniques. Can anybody please tell me if the test set DeepFakes are really DeepFakes?"
  }
}