{
  "id": 124417,
  "title": "Hash conclusions",
  "url": "/competitions/deepfake-detection-challenge/discussion/124417",
  "author_name": "",
  "post_date": "2020-01-04T01:33:53.389582600Z",
  "votes": 69,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Want to share with you some results from looking at the metadata from the hashes point of view. I calculated 3 hashes:</p>\n\n<ol>\n<li>md5 of full files by <code>hashlib</code> package</li>\n<li>hash on audio files numpy data by <code>xxhash</code> package</li>\n<li>hash on pixel data by <code>xxhash</code> package, 5 fixed points and average values per image for 3 frames (first, last and median), on 18 numbers per video overall.</li>\n</ol>\n\n<p>I updated <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">the dataset</a> that I maintain with these columns, and there is the <a href=\"https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata\">accompanying kernel</a> for the dataset.</p>\n\n<p>The numbers are\n1. <strong>10152 fakes are with modified audio</strong>, about 10%. 5249 of them are caused by modified audio sample rate, specifically under-sampled to 1/16000. This last group is not interesting, probably an artifact of one of the fake generation methods. All of them are in folders 45-49. Other 4903 \"real\" audio fakes are distributed equally across the folders. \n2. <strong>30 files are without audio</strong>, 13 reals and 17 fakes (having those 13 as their originals). I haven't noticed anything else special about them, just a missing audio.\n3. All fakes have pixel hashes different from their originals. In other words, <strong>all fakes are also video fakes</strong> (not judging the quality here). \n4. <strong>There are duplicate files</strong>! That is, equal md5. There are two groups of them. The first is couples of real videos which are duplicates. 79 couples completely identical and 23 more with duplicate audio and almost identical video (invisible noise). All of them in folders 22 and 23. The second group are fakes, which are fakes for the same originals but happen to be duplicates of one another. There are 347 of them, spread across the folders. Some of them in groups of dozens of identical files. Better to clean away before any training.</p>\n\n<p>Hopefully it saves you some time.</p>",
  "messages": [
    {
      "id": "709836",
      "postDate": "01/04/2020 01:33:53",
      "content": "<p>Want to share with you some results from looking at the metadata from the hashes point of view. I calculated 3 hashes:</p>\n\n<ol>\n<li>md5 of full files by <code>hashlib</code> package</li>\n<li>hash on audio files numpy data by <code>xxhash</code> package</li>\n<li>hash on pixel data by <code>xxhash</code> package, 5 fixed points and average values per image for 3 frames (first, last and median), on 18 numbers per video overall.</li>\n</ol>\n\n<p>I updated <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">the dataset</a> that I maintain with these columns, and there is the <a href=\"https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata\">accompanying kernel</a> for the dataset.</p>\n\n<p>The numbers are\n1. <strong>10152 fakes are with modified audio</strong>, about 10%. 5249 of them are caused by modified audio sample rate, specifically under-sampled to 1/16000. This last group is not interesting, probably an artifact of one of the fake generation methods. All of them are in folders 45-49. Other 4903 \"real\" audio fakes are distributed equally across the folders. \n2. <strong>30 files are without audio</strong>, 13 reals and 17 fakes (having those 13 as their originals). I haven't noticed anything else special about them, just a missing audio.\n3. All fakes have pixel hashes different from their originals. In other words, <strong>all fakes are also video fakes</strong> (not judging the quality here). \n4. <strong>There are duplicate files</strong>! That is, equal md5. There are two groups of them. The first is couples of real videos which are duplicates. 79 couples completely identical and 23 more with duplicate audio and almost identical video (invisible noise). All of them in folders 22 and 23. The second group are fakes, which are fakes for the same originals but happen to be duplicates of one another. There are 347 of them, spread across the folders. Some of them in groups of dozens of identical files. Better to clean away before any training.</p>\n\n<p>Hopefully it saves you some time.</p>",
      "rawMarkdown": "Want to share with you some results from looking at the metadata from the hashes point of view. I calculated 3 hashes:\n\n1. md5 of full files by `hashlib` package\n2. hash on audio files numpy data by `xxhash` package\n3. hash on pixel data by `xxhash` package, 5 fixed points and average values per image for 3 frames (first, last and median), on 18 numbers per video overall.\n\nI updated [the dataset](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc) that I maintain with these columns, and there is the [accompanying kernel](https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata) for the dataset.\n\nThe numbers are\n1. **10152 fakes are with modified audio**, about 10%. 5249 of them are caused by modified audio sample rate, specifically under-sampled to 1/16000. This last group is not interesting, probably an artifact of one of the fake generation methods. All of them are in folders 45-49. Other 4903 \"real\" audio fakes are distributed equally across the folders. \n2. **30 files are without audio**, 13 reals and 17 fakes (having those 13 as their originals). I haven't noticed anything else special about them, just a missing audio.\n3. All fakes have pixel hashes different from their originals. In other words, **all fakes are also video fakes** (not judging the quality here). \n4. **There are duplicate files**! That is, equal md5. There are two groups of them. The first is couples of real videos which are duplicates. 79 couples completely identical and 23 more with duplicate audio and almost identical video (invisible noise). All of them in folders 22 and 23. The second group are fakes, which are fakes for the same originals but happen to be duplicates of one another. There are 347 of them, spread across the folders. Some of them in groups of dozens of identical files. Better to clean away before any training.\n\nHopefully it saves you some time.",
      "votes": null
    },
    {
      "id": "709842",
      "postDate": "01/04/2020 01:56:36",
      "content": "<p>In a post one of the hosts commented that down sampled audio are NOT fake.  So your 5249 are down sampled.  I think this means that not all 5249 should be considered as fake.  Do you agree?</p>",
      "rawMarkdown": "In a post one of the hosts commented that down sampled audio are NOT fake.  So your 5249 are down sampled.  I think this means that not all 5249 should be considered as fake.  Do you agree?",
      "votes": null
    },
    {
      "id": "709850",
      "postDate": "01/04/2020 02:09:41",
      "content": "<p>Yeah, I saw that post. I think he wrote it in a confusing way, my understanding is that he implied that the down-sampling itself was not intentional, but rather, as I mentioned, it is an artifact, a side-effect. All those videos also have pixel changes, so they are real fakes, real video fakes, with that unintentional audio change.</p>\n\n<p>So, to your question, no, there is no reason not to consider them as fakes.</p>",
      "rawMarkdown": "Yeah, I saw that post. I think he wrote it in a confusing way, my understanding is that he implied that the down-sampling itself was not intentional, but rather, as I mentioned, it is an artifact, a side-effect. All those videos also have pixel changes, so they are real fakes, real video fakes, with that unintentional audio change.\n\nSo, to your question, no, there is no reason not to consider them as fakes.",
      "votes": null
    },
    {
      "id": "710088",
      "postDate": "01/04/2020 10:02:42",
      "content": "<p>You are an angel! </p>",
      "rawMarkdown": "You are an angel!",
      "votes": null
    },
    {
      "id": "710225",
      "postDate": "01/04/2020 13:29:59",
      "content": "<p>You maybe <strong>nosound</strong> but we do feel your vibration. Thank you for your insights</p>",
      "rawMarkdown": "You maybe **nosound** but we do feel your vibration. Thank you for your insights",
      "votes": null
    },
    {
      "id": "710377",
      "postDate": "01/04/2020 16:54:02",
      "content": "<p>Can explain further on what hashes you mean? </p>",
      "rawMarkdown": "Can explain further on what hashes you mean?",
      "votes": null
    },
    {
      "id": "710401",
      "postDate": "01/04/2020 17:33:01",
      "content": "<p>Sure, md5:</p>\n\n<p><code>hashlib.md5(open(str(filepath),'rb').read()).hexdigest()</code></p>\n\n<p>audio and pixels:</p>\n\n<p><code>\nh = xxhash.xxh64(); h.update(data); myhash = h.intdigest()\n</code></p>",
      "rawMarkdown": "Sure, md5:\n\n`hashlib.md5(open(str(filepath),'rb').read()).hexdigest()`\n\naudio and pixels:\n\n`    \nh = xxhash.xxh64(); h.update(data); myhash = h.intdigest()\n`",
      "votes": null
    },
    {
      "id": "710551",
      "postDate": "01/04/2020 21:54:18",
      "content": "<p>would be nice to split the training set labeling into video FAKE/REAL labels and a different audio FAKE/REAL label, which column in your dataset tells if the audio is \"intentional\" audio fake  ? </p>",
      "rawMarkdown": "would be nice to split the training set labeling into video FAKE/REAL labels and a different audio FAKE/REAL label, which column in your dataset tells if the audio is \"intentional\" audio fake  ?",
      "votes": null
    },
    {
      "id": "710575",
      "postDate": "01/04/2020 23:05:38",
      "content": "<p>You are right, there is no dedicate column for it yet. It is <code>audio.@codec_time_base</code> not equal <code>1/16000</code> and <code>wav.hash != wav.hash.orig</code></p>",
      "rawMarkdown": "You are right, there is no dedicate column for it yet. It is `audio.@codec_time_base` not equal `1/16000` and `wav.hash != wav.hash.orig`",
      "votes": null
    },
    {
      "id": "710615",
      "postDate": "01/05/2020 01:23:03",
      "content": "<p>I created a minimalist metadata based on this: <a href=\"https://www.kaggle.com/basharallabadi/dfdc-video-audio-labels\">https://www.kaggle.com/basharallabadi/dfdc-video-audio-labels</a></p>\n\n<p>thanks for sharing your analysis</p>",
      "rawMarkdown": "I created a minimalist metadata based on this: https://www.kaggle.com/basharallabadi/dfdc-video-audio-labels\n\nthanks for sharing your analysis",
      "votes": null
    },
    {
      "id": "710938",
      "postDate": "01/05/2020 12:57:19",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> \nI don't understand why:  It is <code>label equal FAKE</code> and <code>audio.@codec_time_base not equal 1/16000</code> and <code>wav.hash != wav.hash.orig</code> that means fake audio ? It just means modified audio, it's a necessary condition but not sufficient. One FAKE video could have faces + audio modified. Only if video (pxl.hash) is not modified and audio modified and label = FAKE then we know it's a fake audio. Do you agree?\nTools used in this competition to modify videos/audio could have not introduced fake but just re-encoding for any reason.</p>\n\n<p>and it looks <code>pxl.hash</code> always different than <code>pxl.orig.hash</code> so we cannot conclude. We just know what you said in item#1: 10152 fakes are with modified audio.</p>",
      "rawMarkdown": "zaharch \nI don't understand why:  It is `label equal FAKE` and `audio.@codec_time_base not equal 1/16000` and `wav.hash != wav.hash.orig` that means fake audio ? It just means modified audio, it's a necessary condition but not sufficient. One FAKE video could have faces + audio modified. Only if video (pxl.hash) is not modified and audio modified and label = FAKE then we know it's a fake audio. Do you agree?\nTools used in this competition to modify videos/audio could have not introduced fake but just re-encoding for any reason.\n\nand it looks `pxl.hash` always different than `pxl.orig.hash` so we cannot conclude. We just know what you said in item#1: 10152 fakes are with modified audio.",
      "votes": null
    },
    {
      "id": "710969",
      "postDate": "01/05/2020 13:45:13",
      "content": "<p><a href=\"/mpware\">@mpware</a> well they said the 1/16000 down sampling are not fake audio so we can exclude those out of the 10152 .. which leaves you with ~4k audio files.. what you said could apply to those, but so far no way to tell .. I'm doing sound analysis on some examples to see the audio modifications in those 4k </p>",
      "rawMarkdown": "mpware well they said the 1/16000 down sampling are not fake audio so we can exclude those out of the 10152 .. which leaves you with ~4k audio files.. what you said could apply to those, but so far no way to tell .. I'm doing sound analysis on some examples to see the audio modifications in those 4k",
      "votes": null
    },
    {
      "id": "711180",
      "postDate": "01/05/2020 19:28:29",
      "content": "<p><a href=\"/mpware\">@mpware</a> , I kept it short for brevity, the issue that you raised has originally been <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785\">discussed here</a>. <a href=\"/ke8ctn\">@ke8ctn</a> showed that it is just a compression artifacts. And then what <a href=\"/basharallabadi\">@basharallabadi</a>  mentioned, - the organizer said they are not real fakes.</p>",
      "rawMarkdown": "mpware , I kept it short for brevity, the issue that you raised has originally been [discussed here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785). @ke8ctn showed that it is just a compression artifacts. And then what @basharallabadi  mentioned, - the organizer said they are not real fakes.",
      "votes": null
    },
    {
      "id": "712244",
      "postDate": "01/07/2020 01:35:54",
      "content": "<p>Wow what a info!\nThank you for sharing.</p>",
      "rawMarkdown": "Wow what a info!\nThank you for sharing.",
      "votes": null
    },
    {
      "id": "712299",
      "postDate": "01/07/2020 03:38:46",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> hi, could you please indicate which post said down sampled audio are NOT fake?</p>",
      "rawMarkdown": "zaharch hi, could you please indicate which post said down sampled audio are NOT fake?",
      "votes": null
    },
    {
      "id": "712311",
      "postDate": "01/07/2020 04:28:12",
      "content": "<p>it is <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121694\">here</a></p>",
      "rawMarkdown": "it is [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121694)",
      "votes": null
    },
    {
      "id": "713528",
      "postDate": "01/08/2020 11:38:39",
      "content": "<p>This was really helpful thanks!</p>",
      "rawMarkdown": "This was really helpful thanks!",
      "votes": null
    },
    {
      "id": "722639",
      "postDate": "01/18/2020 21:09:57",
      "content": "<p>Being lazy, did you happen to actually detect what parts of each video is fake? not all frames on the video are faked and while training this could be a problem. I was going to write this but if you already have it ready...</p>",
      "rawMarkdown": "Being lazy, did you happen to actually detect what parts of each video is fake? not all frames on the video are faked and while training this could be a problem. I was going to write this but if you already have it ready...",
      "votes": null
    },
    {
      "id": "722689",
      "postDate": "01/18/2020 23:29:53",
      "content": "<p>Actually no, I have not looked into separating frames into reals and fakes, I have done only whole-video analysis. And I don't think it is a good direction to look into, not in my todo list. But being lazy is great, I am totally with you on that.</p>",
      "rawMarkdown": "Actually no, I have not looked into separating frames into reals and fakes, I have done only whole-video analysis. And I don't think it is a good direction to look into, not in my todo list. But being lazy is great, I am totally with you on that.",
      "votes": null
    },
    {
      "id": "722735",
      "postDate": "01/19/2020 02:41:17",
      "content": "<p>well, I overcome my laziness and looked into part_0. had some interesting videos from there. have a look at dfdc_train_part_0/zxyvcnkeiz.mp4\nhas fake segments only at the beginning and the end\nBut, if you only do whole video analysis... its really irrelevant</p>",
      "rawMarkdown": "well, I overcome my laziness and looked into part_0. had some interesting videos from there. have a look at dfdc_train_part_0/zxyvcnkeiz.mp4\nhas fake segments only at the beginning and the end\nBut, if you only do whole video analysis... its really irrelevant",
      "votes": null
    },
    {
      "id": "722744",
      "postDate": "01/19/2020 02:56:41",
      "content": "<p>not even sure if my detection is correct, the fake and real are pretty close.... It seems almost all videos are doctored start to finish, so probably not important anyways</p>",
      "rawMarkdown": "not even sure if my detection is correct, the fake and real are pretty close.... It seems almost all videos are doctored start to finish, so probably not important anyways",
      "votes": null
    },
    {
      "id": "747397",
      "postDate": "02/16/2020 11:32:35",
      "content": "<p>Thanks for collecting these metadata and sharing the insights. \nBased on this observation, </p>\n\n<blockquote>\n  <p>All fakes have pixel hashes different from their originals. In other words, all fakes are also video fakes (not judging the quality here).</p>\n</blockquote>\n\n<p>is it possible to successfully build a model that:\n1. Identifies the original \"mother\" video\n2. Computes the pixel hash of the video to be identified\n3. If it is different than the original, flag as FAKE. </p>\n\n<p>Maybe this approach won't work if the private dataset only contains unobserved videos? \nWhat are your thoughts? </p>",
      "rawMarkdown": "Thanks for collecting these metadata and sharing the insights. \nBased on this observation, \n\n&gt; All fakes have pixel hashes different from their originals. In other words, all fakes are also video fakes (not judging the quality here).\n\nis it possible to successfully build a model that:\n1. Identifies the original \"mother\" video\n2. Computes the pixel hash of the video to be identified\n3. If it is different than the original, flag as FAKE. \n\nMaybe this approach won't work if the private dataset only contains unobserved videos? \nWhat are your thoughts?",
      "votes": null
    },
    {
      "id": "747693",
      "postDate": "02/16/2020 18:18:15",
      "content": "<p>Will work if the private test will contain such pairs, but I think it will not.</p>",
      "rawMarkdown": "Will work if the private test will contain such pairs, but I think it will not.",
      "votes": null
    },
    {
      "id": "751692",
      "postDate": "02/20/2020 12:12:27",
      "content": "<p>Thanks..</p>",
      "rawMarkdown": "Thanks..",
      "votes": null
    },
    {
      "id": "753274",
      "postDate": "02/22/2020 00:58:17",
      "content": "<p>hey, do you mind explaining how you found the difference between real and fake video frames</p>",
      "rawMarkdown": "hey, do you mind explaining how you found the difference between real and fake video frames",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 709842,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/04/2020 01:56:36",
      "content": "<p>In a post one of the hosts commented that down sampled audio are NOT fake.  So your 5249 are down sampled.  I think this means that not all 5249 should be considered as fake.  Do you agree?</p>",
      "votes": null,
      "replies": [
        {
          "id": 709850,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/04/2020 02:09:41",
          "content": "<p>Yeah, I saw that post. I think he wrote it in a confusing way, my understanding is that he implied that the down-sampling itself was not intentional, but rather, as I mentioned, it is an artifact, a side-effect. All those videos also have pixel changes, so they are real fakes, real video fakes, with that unintentional audio change.</p>\n\n<p>So, to your question, no, there is no reason not to consider them as fakes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712299,
          "author_name": "bomb2peng",
          "author_url": "",
          "post_date": "01/07/2020 03:38:46",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> hi, could you please indicate which post said down sampled audio are NOT fake?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712311,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/07/2020 04:28:12",
          "content": "<p>it is <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121694\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 710088,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "01/04/2020 10:02:42",
      "content": "<p>You are an angel! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 710225,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "01/04/2020 13:29:59",
      "content": "<p>You maybe <strong>nosound</strong> but we do feel your vibration. Thank you for your insights</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 710377,
      "author_name": "hermesbonilla",
      "author_url": "",
      "post_date": "01/04/2020 16:54:02",
      "content": "<p>Can explain further on what hashes you mean? </p>",
      "votes": null,
      "replies": [
        {
          "id": 710401,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/04/2020 17:33:01",
          "content": "<p>Sure, md5:</p>\n\n<p><code>hashlib.md5(open(str(filepath),'rb').read()).hexdigest()</code></p>\n\n<p>audio and pixels:</p>\n\n<p><code>\nh = xxhash.xxh64(); h.update(data); myhash = h.intdigest()\n</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 710551,
      "author_name": "basharallabadi",
      "author_url": "",
      "post_date": "01/04/2020 21:54:18",
      "content": "<p>would be nice to split the training set labeling into video FAKE/REAL labels and a different audio FAKE/REAL label, which column in your dataset tells if the audio is \"intentional\" audio fake  ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 710575,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/04/2020 23:05:38",
          "content": "<p>You are right, there is no dedicate column for it yet. It is <code>audio.@codec_time_base</code> not equal <code>1/16000</code> and <code>wav.hash != wav.hash.orig</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710615,
          "author_name": "basharallabadi",
          "author_url": "",
          "post_date": "01/05/2020 01:23:03",
          "content": "<p>I created a minimalist metadata based on this: <a href=\"https://www.kaggle.com/basharallabadi/dfdc-video-audio-labels\">https://www.kaggle.com/basharallabadi/dfdc-video-audio-labels</a></p>\n\n<p>thanks for sharing your analysis</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710938,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "01/05/2020 12:57:19",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> \nI don't understand why:  It is <code>label equal FAKE</code> and <code>audio.@codec_time_base not equal 1/16000</code> and <code>wav.hash != wav.hash.orig</code> that means fake audio ? It just means modified audio, it's a necessary condition but not sufficient. One FAKE video could have faces + audio modified. Only if video (pxl.hash) is not modified and audio modified and label = FAKE then we know it's a fake audio. Do you agree?\nTools used in this competition to modify videos/audio could have not introduced fake but just re-encoding for any reason.</p>\n\n<p>and it looks <code>pxl.hash</code> always different than <code>pxl.orig.hash</code> so we cannot conclude. We just know what you said in item#1: 10152 fakes are with modified audio.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710969,
          "author_name": "basharallabadi",
          "author_url": "",
          "post_date": "01/05/2020 13:45:13",
          "content": "<p><a href=\"/mpware\">@mpware</a> well they said the 1/16000 down sampling are not fake audio so we can exclude those out of the 10152 .. which leaves you with ~4k audio files.. what you said could apply to those, but so far no way to tell .. I'm doing sound analysis on some examples to see the audio modifications in those 4k </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 711180,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/05/2020 19:28:29",
          "content": "<p><a href=\"/mpware\">@mpware</a> , I kept it short for brevity, the issue that you raised has originally been <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785\">discussed here</a>. <a href=\"/ke8ctn\">@ke8ctn</a> showed that it is just a compression artifacts. And then what <a href=\"/basharallabadi\">@basharallabadi</a>  mentioned, - the organizer said they are not real fakes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 712244,
      "author_name": "gwsong",
      "author_url": "",
      "post_date": "01/07/2020 01:35:54",
      "content": "<p>Wow what a info!\nThank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 713528,
      "author_name": "emrebayram",
      "author_url": "",
      "post_date": "01/08/2020 11:38:39",
      "content": "<p>This was really helpful thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 722639,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "01/18/2020 21:09:57",
      "content": "<p>Being lazy, did you happen to actually detect what parts of each video is fake? not all frames on the video are faked and while training this could be a problem. I was going to write this but if you already have it ready...</p>",
      "votes": null,
      "replies": [
        {
          "id": 722689,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/18/2020 23:29:53",
          "content": "<p>Actually no, I have not looked into separating frames into reals and fakes, I have done only whole-video analysis. And I don't think it is a good direction to look into, not in my todo list. But being lazy is great, I am totally with you on that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722735,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/19/2020 02:41:17",
          "content": "<p>well, I overcome my laziness and looked into part_0. had some interesting videos from there. have a look at dfdc_train_part_0/zxyvcnkeiz.mp4\nhas fake segments only at the beginning and the end\nBut, if you only do whole video analysis... its really irrelevant</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722744,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/19/2020 02:56:41",
          "content": "<p>not even sure if my detection is correct, the fake and real are pretty close.... It seems almost all videos are doctored start to finish, so probably not important anyways</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 753274,
          "author_name": "harisanthanam",
          "author_url": "",
          "post_date": "02/22/2020 00:58:17",
          "content": "<p>hey, do you mind explaining how you found the difference between real and fake video frames</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 747397,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "02/16/2020 11:32:35",
      "content": "<p>Thanks for collecting these metadata and sharing the insights. \nBased on this observation, </p>\n\n<blockquote>\n  <p>All fakes have pixel hashes different from their originals. In other words, all fakes are also video fakes (not judging the quality here).</p>\n</blockquote>\n\n<p>is it possible to successfully build a model that:\n1. Identifies the original \"mother\" video\n2. Computes the pixel hash of the video to be identified\n3. If it is different than the original, flag as FAKE. </p>\n\n<p>Maybe this approach won't work if the private dataset only contains unobserved videos? \nWhat are your thoughts? </p>",
      "votes": null,
      "replies": [
        {
          "id": 747693,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "02/16/2020 18:18:15",
          "content": "<p>Will work if the private test will contain such pairs, but I think it will not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 751692,
      "author_name": "hansels",
      "author_url": "",
      "post_date": "02/20/2020 12:12:27",
      "content": "<p>Thanks..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "709836": "Want to share with you some results from looking at the metadata from the hashes point of view. I calculated 3 hashes:\n\n1. md5 of full files by `hashlib` package\n2. hash on audio files numpy data by `xxhash` package\n3. hash on pixel data by `xxhash` package, 5 fixed points and average values per image for 3 frames (first, last and median), on 18 numbers per video overall.\n\nI updated [the dataset](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc) that I maintain with these columns, and there is the [accompanying kernel](https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata) for the dataset.\n\nThe numbers are\n1. **10152 fakes are with modified audio**, about 10%. 5249 of them are caused by modified audio sample rate, specifically under-sampled to 1/16000. This last group is not interesting, probably an artifact of one of the fake generation methods. All of them are in folders 45-49. Other 4903 \"real\" audio fakes are distributed equally across the folders. \n2. **30 files are without audio**, 13 reals and 17 fakes (having those 13 as their originals). I haven't noticed anything else special about them, just a missing audio.\n3. All fakes have pixel hashes different from their originals. In other words, **all fakes are also video fakes** (not judging the quality here). \n4. **There are duplicate files**! That is, equal md5. There are two groups of them. The first is couples of real videos which are duplicates. 79 couples completely identical and 23 more with duplicate audio and almost identical video (invisible noise). All of them in folders 22 and 23. The second group are fakes, which are fakes for the same originals but happen to be duplicates of one another. There are 347 of them, spread across the folders. Some of them in groups of dozens of identical files. Better to clean away before any training.\n\nHopefully it saves you some time.",
    "709842": "In a post one of the hosts commented that down sampled audio are NOT fake.  So your 5249 are down sampled.  I think this means that not all 5249 should be considered as fake.  Do you agree?",
    "709850": "Yeah, I saw that post. I think he wrote it in a confusing way, my understanding is that he implied that the down-sampling itself was not intentional, but rather, as I mentioned, it is an artifact, a side-effect. All those videos also have pixel changes, so they are real fakes, real video fakes, with that unintentional audio change.\n\nSo, to your question, no, there is no reason not to consider them as fakes.",
    "710088": "You are an angel!",
    "710225": "You maybe **nosound** but we do feel your vibration. Thank you for your insights",
    "710377": "Can explain further on what hashes you mean?",
    "710401": "Sure, md5:\n\n`hashlib.md5(open(str(filepath),'rb').read()).hexdigest()`\n\naudio and pixels:\n\n`    \nh = xxhash.xxh64(); h.update(data); myhash = h.intdigest()\n`",
    "710551": "would be nice to split the training set labeling into video FAKE/REAL labels and a different audio FAKE/REAL label, which column in your dataset tells if the audio is \"intentional\" audio fake  ?",
    "710575": "You are right, there is no dedicate column for it yet. It is `audio.@codec_time_base` not equal `1/16000` and `wav.hash != wav.hash.orig`",
    "710615": "I created a minimalist metadata based on this: https://www.kaggle.com/basharallabadi/dfdc-video-audio-labels\n\nthanks for sharing your analysis",
    "710938": "zaharch \nI don't understand why:  It is `label equal FAKE` and `audio.@codec_time_base not equal 1/16000` and `wav.hash != wav.hash.orig` that means fake audio ? It just means modified audio, it's a necessary condition but not sufficient. One FAKE video could have faces + audio modified. Only if video (pxl.hash) is not modified and audio modified and label = FAKE then we know it's a fake audio. Do you agree?\nTools used in this competition to modify videos/audio could have not introduced fake but just re-encoding for any reason.\n\nand it looks `pxl.hash` always different than `pxl.orig.hash` so we cannot conclude. We just know what you said in item#1: 10152 fakes are with modified audio.",
    "710969": "mpware well they said the 1/16000 down sampling are not fake audio so we can exclude those out of the 10152 .. which leaves you with ~4k audio files.. what you said could apply to those, but so far no way to tell .. I'm doing sound analysis on some examples to see the audio modifications in those 4k",
    "711180": "mpware , I kept it short for brevity, the issue that you raised has originally been [discussed here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785). @ke8ctn showed that it is just a compression artifacts. And then what @basharallabadi  mentioned, - the organizer said they are not real fakes.",
    "712244": "Wow what a info!\nThank you for sharing.",
    "712299": "zaharch hi, could you please indicate which post said down sampled audio are NOT fake?",
    "712311": "it is [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121694)",
    "713528": "This was really helpful thanks!",
    "722639": "Being lazy, did you happen to actually detect what parts of each video is fake? not all frames on the video are faked and while training this could be a problem. I was going to write this but if you already have it ready...",
    "722689": "Actually no, I have not looked into separating frames into reals and fakes, I have done only whole-video analysis. And I don't think it is a good direction to look into, not in my todo list. But being lazy is great, I am totally with you on that.",
    "722735": "well, I overcome my laziness and looked into part_0. had some interesting videos from there. have a look at dfdc_train_part_0/zxyvcnkeiz.mp4\nhas fake segments only at the beginning and the end\nBut, if you only do whole video analysis... its really irrelevant",
    "722744": "not even sure if my detection is correct, the fake and real are pretty close.... It seems almost all videos are doctored start to finish, so probably not important anyways",
    "747397": "Thanks for collecting these metadata and sharing the insights. \nBased on this observation, \n\n&gt; All fakes have pixel hashes different from their originals. In other words, all fakes are also video fakes (not judging the quality here).\n\nis it possible to successfully build a model that:\n1. Identifies the original \"mother\" video\n2. Computes the pixel hash of the video to be identified\n3. If it is different than the original, flag as FAKE. \n\nMaybe this approach won't work if the private dataset only contains unobserved videos? \nWhat are your thoughts?",
    "747693": "Will work if the private test will contain such pairs, but I think it will not.",
    "751692": "Thanks..",
    "753274": "hey, do you mind explaining how you found the difference between real and fake video frames"
  },
  "source": "meta"
}