{
  "id": 121861,
  "title": "~8% audio altered",
  "url": "/competitions/deepfake-detection-challenge/discussion/121861",
  "author_name": "",
  "post_date": "2019-12-16T07:24:35.994488Z",
  "votes": 52,
  "comment_count": 40,
  "views": 0,
  "content": "<p>Looked through the audio data and found about 8% have been altered in the training data in some way. This accurately maps to exactly 10152 that were marked as fake and zero that were marked as real. </p>",
  "messages": [
    {
      "id": "696149",
      "postDate": "12/16/2019 07:24:35",
      "content": "<p>Looked through the audio data and found about 8% have been altered in the training data in some way. This accurately maps to exactly 10152 that were marked as fake and zero that were marked as real. </p>",
      "rawMarkdown": "Looked through the audio data and found about 8% have been altered in the training data in some way. This accurately maps to exactly 10152 that were marked as fake and zero that were marked as real.",
      "votes": null
    },
    {
      "id": "696222",
      "postDate": "12/16/2019 10:07:31",
      "content": "<p>what do you consider as altered audio ? Would it be considered fake if it just goes through transcoding ? </p>",
      "rawMarkdown": "what do you consider as altered audio ? Would it be considered fake if it just goes through transcoding ?",
      "votes": null
    },
    {
      "id": "696236",
      "postDate": "12/16/2019 10:25:50",
      "content": "<p>By altered I mean that the audio is not the exact same as the source video. How exactly it's different, I do not know, but it does perfectly map to fake labels so I am assuming they are considering any audio change at all to be a fake. </p>",
      "rawMarkdown": "By altered I mean that the audio is not the exact same as the source video. How exactly it's different, I do not know, but it does perfectly map to fake labels so I am assuming they are considering any audio change at all to be a fake.",
      "votes": null
    },
    {
      "id": "696444",
      "postDate": "12/16/2019 16:04:05",
      "content": "<p>Could you give examples of altered audios in the training/sample_training files?</p>",
      "rawMarkdown": "Could you give examples of altered audios in the training/sample_training files?",
      "votes": null
    },
    {
      "id": "696525",
      "postDate": "12/16/2019 18:17:11",
      "content": "<p>Didn't look in sample_training. xvzsbxrlxq.mp4 has altered audio in the training set. </p>",
      "rawMarkdown": "Didn't look in sample_training. xvzsbxrlxq.mp4 has altered audio in the training set.",
      "votes": null
    },
    {
      "id": "696588",
      "postDate": "12/16/2019 19:56:41",
      "content": "<p>Oops. Sorry, I think I accidentally deleted my earlier comment. Edited it back to the best of my memory. Thanks! I found very blatant audio tamperings (eg: aomdezxgzc.mp4) in manual inspection. </p>\n\n<p>It will be interesting to categorize what sortof audio tampering has been done in the data. I will reply here when I get more insight into this.</p>",
      "rawMarkdown": "Oops. Sorry, I think I accidentally deleted my earlier comment. Edited it back to the best of my memory. Thanks! I found very blatant audio tamperings (eg: aomdezxgzc.mp4) in manual inspection. \n\nIt will be interesting to categorize what sortof audio tampering has been done in the data. I will reply here when I get more insight into this.",
      "votes": null
    },
    {
      "id": "697419",
      "postDate": "12/17/2019 22:05:13",
      "content": "<p>Tell me what I’m missing because I’m new here... but wouldn’t a deepfake alter both the video and the audio as part of the fakery?  i.e. it’s true that if you always had the original video available identifying trickery would be trivial.  are you implying that altered video streams do NOT similarly perfectly call out the fakes?</p>",
      "rawMarkdown": "Tell me what I’m missing because I’m new here... but wouldn’t a deepfake alter both the video and the audio as part of the fakery?  i.e. it’s true that if you always had the original video available identifying trickery would be trivial.  are you implying that altered video streams do NOT similarly perfectly call out the fakes?",
      "votes": null
    },
    {
      "id": "697710",
      "postDate": "12/18/2019 09:55:28",
      "content": "<p>I did't find any .mp4 named aomdezxgzc.mp4? Can you share its location? For example, the belonging set?</p>",
      "rawMarkdown": "I did't find any .mp4 named aomdezxgzc.mp4? Can you share its location? For example, the belonging set?",
      "votes": null
    },
    {
      "id": "697719",
      "postDate": "12/18/2019 10:05:54",
      "content": "<p>On my side, comparing byte-wise the first second with the source audio, I get about 4.5% of modified audio...\nI thought any re-encoding would have (statistically) changed 99.99% of the values (since my 1s tryout)... Did you compare the entire 10s audio files?</p>",
      "rawMarkdown": "On my side, comparing byte-wise the first second with the source audio, I get about 4.5% of modified audio...\nI thought any re-encoding would have (statistically) changed 99.99% of the values (since my 1s tryout)... Did you compare the entire 10s audio files?",
      "votes": null
    },
    {
      "id": "697927",
      "postDate": "12/18/2019 14:58:37",
      "content": "<p>It's in dfdc train part 49/ in the train set. Were able to locate it?</p>",
      "rawMarkdown": "It's in dfdc train part 49/ in the train set. Were able to locate it?",
      "votes": null
    },
    {
      "id": "697945",
      "postDate": "12/18/2019 15:17:27",
      "content": "<p>Found it. Thank you very much.</p>",
      "rawMarkdown": "Found it. Thank you very much.",
      "votes": null
    },
    {
      "id": "698282",
      "postDate": "12/19/2019 02:33:33",
      "content": "<p>May I ask how do you judge that the audio of two mp4 files is different?\nI wanna hear the altered audio.\nThx in advance.</p>",
      "rawMarkdown": "May I ask how do you judge that the audio of two mp4 files is different?\nI wanna hear the altered audio.\nThx in advance.",
      "votes": null
    },
    {
      "id": "698291",
      "postDate": "12/19/2019 02:56:48",
      "content": "<p>I loaded the original audio and then compared to the videos that marked that video as the original. Then just applied a np array comparison to see if they were equal. </p>",
      "rawMarkdown": "I loaded the original audio and then compared to the videos that marked that video as the original. Then just applied a np array comparison to see if they were equal.",
      "votes": null
    },
    {
      "id": "698421",
      "postDate": "12/19/2019 07:53:07",
      "content": "<p>Just found out that the average is 4.5% fake audio from directory 0 to 44. Dir 45 to 49 have 55% fake audio. Might be interesting to use 45-49 for a balanced set.</p>",
      "rawMarkdown": "Just found out that the average is 4.5% fake audio from directory 0 to 44. Dir 45 to 49 have 55% fake audio. Might be interesting to use 45-49 for a balanced set.",
      "votes": null
    },
    {
      "id": "698436",
      "postDate": "12/19/2019 08:17:33",
      "content": "<p>so you mean all the videos labeled as FAKE in the training set have audio altered?</p>",
      "rawMarkdown": "so you mean all the videos labeled as FAKE in the training set have audio altered?",
      "votes": null
    },
    {
      "id": "698574",
      "postDate": "12/19/2019 12:13:23",
      "content": "<p>Is there any instance when the audio corresponding to the video is altered but the video(face) is not? \nIdeally this is a possible case, but does the data have such instances and even if it does, model can only detect any altered-generated audio and not the one when the audio is pasted unaltered from some other speaker or source.</p>",
      "rawMarkdown": "Is there any instance when the audio corresponding to the video is altered but the video(face) is not? \nIdeally this is a possible case, but does the data have such instances and even if it does, model can only detect any altered-generated audio and not the one when the audio is pasted unaltered from some other speaker or source.",
      "votes": null
    },
    {
      "id": "698704",
      "postDate": "12/19/2019 15:47:20",
      "content": "<p>Interesting point. While my code was running I had it print every 1000 files the percent and I saw huge fluctuations during run. That might be the explanation </p>",
      "rawMarkdown": "Interesting point. While my code was running I had it print every 1000 files the percent and I saw huge fluctuations during run. That might be the explanation",
      "votes": null
    },
    {
      "id": "699078",
      "postDate": "12/20/2019 03:11:11",
      "content": "<p>Thx. It helps a lot.\nI compare the audio of some videos labeled  as 'fake'. Indeed, I find some pair which have different audio np array. BUT when I listen to them, the audio of the pair hears the same.\nWould you please show me a pair whose audio hears different?\nThx in advance.</p>",
      "rawMarkdown": "Thx. It helps a lot.\nI compare the audio of some videos labeled  as 'fake'. Indeed, I find some pair which have different audio np array. BUT when I listen to them, the audio of the pair hears the same.\nWould you please show me a pair whose audio hears different?\nThx in advance.",
      "votes": null
    },
    {
      "id": "699090",
      "postDate": "12/20/2019 03:28:58",
      "content": "<p>Please read all the responses in this topic carefully, you will find the answer.</p>",
      "rawMarkdown": "Please read all the responses in this topic carefully, you will find the answer.",
      "votes": null
    },
    {
      "id": "699103",
      "postDate": "12/20/2019 03:45:55",
      "content": "<p>Don't think thats what he said - he said 8% of files had audio altered and they were all part of the fake set.  There are more than 8% fake in the full dataset.</p>",
      "rawMarkdown": "Don't think thats what he said - he said 8% of files had audio altered and they were all part of the fake set.  There are more than 8% fake in the full dataset.",
      "votes": null
    },
    {
      "id": "699119",
      "postDate": "12/20/2019 04:04:34",
      "content": "<p>yep, I find it. thx</p>",
      "rawMarkdown": "yep, I find it. thx",
      "votes": null
    },
    {
      "id": "699193",
      "postDate": "12/20/2019 06:19:17",
      "content": "<p>I started a <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121268\">similar discussion</a> and <a href=\"/phoenix9032\">@phoenix9032</a> had an interesting question.\n&gt; How do you understand that without prior knowledge ? How do you know how someone speaks ?</p>",
      "rawMarkdown": "I started a [similar discussion](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121268) and @phoenix9032 had an interesting question.\n&gt; How do you understand that without prior knowledge ? How do you know how someone speaks ?",
      "votes": null
    },
    {
      "id": "700961",
      "postDate": "12/22/2019 21:53:10",
      "content": "<p>I also took a look at the audio differences in the fake clips. See here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785</a></p>",
      "rawMarkdown": "I also took a look at the audio differences in the fake clips. See here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785",
      "votes": null
    },
    {
      "id": "701112",
      "postDate": "12/23/2019 05:22:26",
      "content": "<p>If you visit the papers that you have mentioned in the discussion and dig deeper into this problem statement (see <a href=\"https://www.asvspoof.org/\">https://www.asvspoof.org/</a> ), you will find really interesting analysis of Spectographs by different people which show the differences between the original audios and faked audios (for example see <a href=\"https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35\">https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35</a>? ).</p>\n\n<p>The important point to remember I think in this case is that it is NOT an audio swap or a lip sync where someone else spoke the exact same thing. Instead it's an audio altercation or some other form of fake audio creation. I understand this kind of altercation is not as easily distinguishable just by listening but still if you do a more thorough analysis (like plot the spectographs) you will find the audios which are altered are clearly different from the originals. (see <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785</a> )</p>",
      "rawMarkdown": "If you visit the papers that you have mentioned in the discussion and dig deeper into this problem statement (see https://www.asvspoof.org/ ), you will find really interesting analysis of Spectographs by different people which show the differences between the original audios and faked audios (for example see https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35? ).\n\nThe important point to remember I think in this case is that it is NOT an audio swap or a lip sync where someone else spoke the exact same thing. Instead it's an audio altercation or some other form of fake audio creation. I understand this kind of altercation is not as easily distinguishable just by listening but still if you do a more thorough analysis (like plot the spectographs) you will find the audios which are altered are clearly different from the originals. (see https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785 )",
      "votes": null
    },
    {
      "id": "708955",
      "postDate": "01/02/2020 21:45:36",
      "content": "<p>Hi <a href=\"/ryches\">@ryches</a> , I can't understand something very basic about what you are saying here. Original is defined only for the fake videos, so haven't you checked audio alteration only for the fakes subset? If so, I don't understand the second sentence, that group is fakes by definition. Am I missing something?</p>",
      "rawMarkdown": "Hi @ryches , I can't understand something very basic about what you are saying here. Original is defined only for the fake videos, so haven't you checked audio alteration only for the fakes subset? If so, I don't understand the second sentence, that group is fakes by definition. Am I missing something?",
      "votes": null
    },
    {
      "id": "708959",
      "postDate": "01/02/2020 21:51:11",
      "content": "<p>I iterated through the full csv I have created with the name of every video and the video it references (none if it is an original video). While iterating through I would load the audio for that row and the video it was altered from. Compare the audio of those two. If they are equal or there is no source video (like it was an original, real video) then I marked it as 0 (unaltered audio), but if there was a source video and after comparing the two they were not the same then I marked it as 1 (altered audio). </p>\n\n<p>Not sure if that clears up any confusion</p>",
      "rawMarkdown": "I iterated through the full csv I have created with the name of every video and the video it references (none if it is an original video). While iterating through I would load the audio for that row and the video it was altered from. Compare the audio of those two. If they are equal or there is no source video (like it was an original, real video) then I marked it as 0 (unaltered audio), but if there was a source video and after comparing the two they were not the same then I marked it as 1 (altered audio). \n\nNot sure if that clears up any confusion",
      "votes": null
    },
    {
      "id": "708966",
      "postDate": "01/02/2020 22:03:13",
      "content": "<p>It doesn't clear my confusion yet. You say \"if there was a source video and after comparing the two they were not the same then I marked it as 1\", so every video marked as 1 had a source video. But only fakes videos have a source video. And this is my confusion, what makes the second sentence in your original post to be non-trivial?</p>",
      "rawMarkdown": "It doesn't clear my confusion yet. You say \"if there was a source video and after comparing the two they were not the same then I marked it as 1\", so every video marked as 1 had a source video. But only fakes videos have a source video. And this is my confusion, what makes the second sentence in your original post to be non-trivial?",
      "votes": null
    },
    {
      "id": "708979",
      "postDate": "01/02/2020 22:16:02",
      "content": "<p>I guess that part is trivial because it is only really looking at the subset of fakes, but it is just confirmation the method did correctly label some videos audio as identical and others non-identical. </p>",
      "rawMarkdown": "I guess that part is trivial because it is only really looking at the subset of fakes, but it is just confirmation the method did correctly label some videos audio as identical and others non-identical.",
      "votes": null
    },
    {
      "id": "708983",
      "postDate": "01/02/2020 22:19:25",
      "content": "<p>I see, thank you for the clarification.</p>",
      "rawMarkdown": "I see, thank you for the clarification.",
      "votes": null
    },
    {
      "id": "709183",
      "postDate": "01/03/2020 06:12:06",
      "content": "<p><a href=\"/ryches\">@ryches</a> I used the same approach to find fake videos with modified audio, but I've found exactly 10,157 items, not 10,152.</p>",
      "rawMarkdown": "ryches I used the same approach to find fake videos with modified audio, but I've found exactly 10,157 items, not 10,152.",
      "votes": null
    },
    {
      "id": "710668",
      "postDate": "01/05/2020 03:58:19",
      "content": "<p><a href=\"/raibektussupbekov\">@raibektussupbekov</a> , thanks for sharing the csv.</p>\n\n<p>As per the csv, kpmieleslf.mp4 is an audio fake of dhjnjkzuhq.mp4. But if you listen to the audio of these two clips, the audio in kpmieleslf.mp4 seems to be same as that in its original. </p>\n\n<p>Am i missing something here ?</p>",
      "rawMarkdown": "raibektussupbekov , thanks for sharing the csv.\n\nAs per the csv, kpmieleslf.mp4 is an audio fake of dhjnjkzuhq.mp4. But if you listen to the audio of these two clips, the audio in kpmieleslf.mp4 seems to be same as that in its original. \n\nAm i missing something here ?",
      "votes": null
    },
    {
      "id": "710707",
      "postDate": "01/05/2020 05:17:12",
      "content": "<p>Some people here have noticed that it is the case with some modified audios. A fake and its original might sound similar despite the fact that they actually differ as data.</p>",
      "rawMarkdown": "Some people here have noticed that it is the case with some modified audios. A fake and its original might sound similar despite the fact that they actually differ as data.",
      "votes": null
    },
    {
      "id": "715062",
      "postDate": "01/10/2020 05:00:10",
      "content": "<p>In the 10,152 videos, is there any situation where the audio is FAKE but the face is REAL?</p>",
      "rawMarkdown": "In the 10,152 videos, is there any situation where the audio is FAKE but the face is REAL?",
      "votes": null
    },
    {
      "id": "715075",
      "postDate": "01/10/2020 05:20:49",
      "content": "<p>In all 10152 videos there is at least a minor change in pixel data, compared with the originals. In fact, this is true for all fakes. Some fakes barely change anything, some change not the face, there are duplicates between fakes, and duplicates between reals. But in general, the answer to your question is no.</p>",
      "rawMarkdown": "In all 10152 videos there is at least a minor change in pixel data, compared with the originals. In fact, this is true for all fakes. Some fakes barely change anything, some change not the face, there are duplicates between fakes, and duplicates between reals. But in general, the answer to your question is no.",
      "votes": null
    },
    {
      "id": "715091",
      "postDate": "01/10/2020 05:50:45",
      "content": "<p>Thank you. So, when training an image-based binary classification model, there is no need to delete these audio-modified data, right?</p>",
      "rawMarkdown": "Thank you. So, when training an image-based binary classification model, there is no need to delete these audio-modified data, right?",
      "votes": null
    },
    {
      "id": "715136",
      "postDate": "01/10/2020 06:47:34",
      "content": "<p>Right, but to be cautious I would say I don't see a reason to exclude them, so far. It was not publicly proven so far that their pixel level changes any different from all other fakes. </p>",
      "rawMarkdown": "Right, but to be cautious I would say I don't see a reason to exclude them, so far. It was not publicly proven so far that their pixel level changes any different from all other fakes.",
      "votes": null
    },
    {
      "id": "715149",
      "postDate": "01/10/2020 07:23:41",
      "content": "<p>Thx</p>",
      "rawMarkdown": "Thx",
      "votes": null
    },
    {
      "id": "731781",
      "postDate": "01/29/2020 03:32:52",
      "content": "<p>Great, Thanks, How do you find the exact num of fake audio？</p>",
      "rawMarkdown": "Great, Thanks, How do you find the exact num of fake audio？",
      "votes": null
    },
    {
      "id": "756932",
      "postDate": "02/26/2020 08:29:58",
      "content": "<p>Did someone succeed to label fake audios only?</p>",
      "rawMarkdown": "Did someone succeed to label fake audios only?",
      "votes": null
    },
    {
      "id": "779067",
      "postDate": "03/19/2020 02:19:44",
      "content": "<p><a href=\"/raibektussupbekov\">@raibektussupbekov</a>. thank you for the csv file. Is it possible for you to publish a script that can produce this csv file?</p>",
      "rawMarkdown": "raibektussupbekov. thank you for the csv file. Is it possible for you to publish a script that can produce this csv file?",
      "votes": null
    },
    {
      "id": "779532",
      "postDate": "03/19/2020 12:55:41",
      "content": "<p><a href=\"/projdev\">@projdev</a> I used this function.</p>\n\n<p>```\nimport glob\nimport json\nimport subprocess\nimport librosa\nimport numpy as np</p>\n\n<p>dfdc_train_wav_path = './disk1/dfdc_train_wav/'</p>\n\n<p>def audio_altered(fake_path, fake_video, orig_path, orig_video):</p>\n\n<pre><code>\"\"\"Finds out if audio of fake_video was altered.\n\n# Arguments\n    fake_path: fake mp4 video path name\n    fake_video: fake mp4 video name\n    orig_path: original mp4 video path name\n    orig_video: original mp4 video name\n\n# Returns\n    True - if audio of fake mp4 video was altered\n    False - otherwise\n\"\"\"\n\nfake_wav = fake_video[:-4] + '.wav'\nfake_wav_path = dfdc_train_wav_path + fake_wav\ntry:\n    # in case if .wav has already been extracted\n    fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\nexcept FileNotFoundError:\n    # extract fake_path audio\n    # .wav audio format is used because librosa.load() doesn't work with .aac\n    command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (fake_path, fake_wav_path)\n    subprocess.run(command, shell=True) \n\n    try:\n        fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n    except FileNotFoundError:\n        # if fake video has no audio than its audio is not altered\n        return False\n\norig_wav = orig_video[:-4] + '.wav'\norig_wav_path = dfdc_train_wav_path + orig_wav\ntry:\n    # in case if .wav has already been extracted\n    orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\nexcept FileNotFoundError:\n    # extract orig_path audio\n    # .wav audio format is used because librosa.load() doesn't work with .aac\n    command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (orig_path, orig_wav_path)\n    subprocess.run(command, shell=True)\n\n    try:\n        orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n    except FileNotFoundError:\n        # if original video has no audio but fake video does than audio is altered\n        return True\n\nreturn fake_rate != orig_rate or not np.array_equal(fake_data, orig_data)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "projdev I used this function.\n\n\n```\nimport glob\nimport json\nimport subprocess\nimport librosa\nimport numpy as np\n\ndfdc_train_wav_path = './disk1/dfdc_train_wav/'\n\ndef audio_altered(fake_path, fake_video, orig_path, orig_video):\n    \n    \"\"\"Finds out if audio of fake_video was altered.\n    \n    # Arguments\n        fake_path: fake mp4 video path name\n        fake_video: fake mp4 video name\n        orig_path: original mp4 video path name\n        orig_video: original mp4 video name\n        \n    # Returns\n        True - if audio of fake mp4 video was altered\n        False - otherwise\n    \"\"\"\n\n    fake_wav = fake_video[:-4] + '.wav'\n    fake_wav_path = dfdc_train_wav_path + fake_wav\n    try:\n        # in case if .wav has already been extracted\n        fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n    except FileNotFoundError:\n        # extract fake_path audio\n        # .wav audio format is used because librosa.load() doesn't work with .aac\n        command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (fake_path, fake_wav_path)\n        subprocess.run(command, shell=True) \n\n        try:\n            fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n        except FileNotFoundError:\n            # if fake video has no audio than its audio is not altered\n            return False\n    \n    orig_wav = orig_video[:-4] + '.wav'\n    orig_wav_path = dfdc_train_wav_path + orig_wav\n    try:\n        # in case if .wav has already been extracted\n        orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n    except FileNotFoundError:\n        # extract orig_path audio\n        # .wav audio format is used because librosa.load() doesn't work with .aac\n        command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (orig_path, orig_wav_path)\n        subprocess.run(command, shell=True)\n\n        try:\n            orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n        except FileNotFoundError:\n            # if original video has no audio but fake video does than audio is altered\n            return True\n    \n    return fake_rate != orig_rate or not np.array_equal(fake_data, orig_data)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 696222,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "12/16/2019 10:07:31",
      "content": "<p>what do you consider as altered audio ? Would it be considered fake if it just goes through transcoding ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 696236,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "12/16/2019 10:25:50",
          "content": "<p>By altered I mean that the audio is not the exact same as the source video. How exactly it's different, I do not know, but it does perfectly map to fake labels so I am assuming they are considering any audio change at all to be a fake. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 697719,
          "author_name": "simoninparis",
          "author_url": "",
          "post_date": "12/18/2019 10:05:54",
          "content": "<p>On my side, comparing byte-wise the first second with the source audio, I get about 4.5% of modified audio...\nI thought any re-encoding would have (statistically) changed 99.99% of the values (since my 1s tryout)... Did you compare the entire 10s audio files?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 698421,
          "author_name": "simoninparis",
          "author_url": "",
          "post_date": "12/19/2019 07:53:07",
          "content": "<p>Just found out that the average is 4.5% fake audio from directory 0 to 44. Dir 45 to 49 have 55% fake audio. Might be interesting to use 45-49 for a balanced set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 698704,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "12/19/2019 15:47:20",
          "content": "<p>Interesting point. While my code was running I had it print every 1000 files the percent and I saw huge fluctuations during run. That might be the explanation </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 696444,
      "author_name": "bayesianconspirator",
      "author_url": "",
      "post_date": "12/16/2019 16:04:05",
      "content": "<p>Could you give examples of altered audios in the training/sample_training files?</p>",
      "votes": null,
      "replies": [
        {
          "id": 696525,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "12/16/2019 18:17:11",
          "content": "<p>Didn't look in sample_training. xvzsbxrlxq.mp4 has altered audio in the training set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 696588,
          "author_name": "bayesianconspirator",
          "author_url": "",
          "post_date": "12/16/2019 19:56:41",
          "content": "<p>Oops. Sorry, I think I accidentally deleted my earlier comment. Edited it back to the best of my memory. Thanks! I found very blatant audio tamperings (eg: aomdezxgzc.mp4) in manual inspection. </p>\n\n<p>It will be interesting to categorize what sortof audio tampering has been done in the data. I will reply here when I get more insight into this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 697710,
          "author_name": "fionalxd",
          "author_url": "",
          "post_date": "12/18/2019 09:55:28",
          "content": "<p>I did't find any .mp4 named aomdezxgzc.mp4? Can you share its location? For example, the belonging set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 697927,
          "author_name": "bayesianconspirator",
          "author_url": "",
          "post_date": "12/18/2019 14:58:37",
          "content": "<p>It's in dfdc train part 49/ in the train set. Were able to locate it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 697945,
          "author_name": "fionalxd",
          "author_url": "",
          "post_date": "12/18/2019 15:17:27",
          "content": "<p>Found it. Thank you very much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 697419,
      "author_name": "jeffplourde",
      "author_url": "",
      "post_date": "12/17/2019 22:05:13",
      "content": "<p>Tell me what I’m missing because I’m new here... but wouldn’t a deepfake alter both the video and the audio as part of the fakery?  i.e. it’s true that if you always had the original video available identifying trickery would be trivial.  are you implying that altered video streams do NOT similarly perfectly call out the fakes?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 698282,
      "author_name": "daddyjin",
      "author_url": "",
      "post_date": "12/19/2019 02:33:33",
      "content": "<p>May I ask how do you judge that the audio of two mp4 files is different?\nI wanna hear the altered audio.\nThx in advance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 698291,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "12/19/2019 02:56:48",
          "content": "<p>I loaded the original audio and then compared to the videos that marked that video as the original. Then just applied a np array comparison to see if they were equal. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699078,
          "author_name": "daddyjin",
          "author_url": "",
          "post_date": "12/20/2019 03:11:11",
          "content": "<p>Thx. It helps a lot.\nI compare the audio of some videos labeled  as 'fake'. Indeed, I find some pair which have different audio np array. BUT when I listen to them, the audio of the pair hears the same.\nWould you please show me a pair whose audio hears different?\nThx in advance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699090,
          "author_name": "fionalxd",
          "author_url": "",
          "post_date": "12/20/2019 03:28:58",
          "content": "<p>Please read all the responses in this topic carefully, you will find the answer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699119,
          "author_name": "daddyjin",
          "author_url": "",
          "post_date": "12/20/2019 04:04:34",
          "content": "<p>yep, I find it. thx</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 698436,
      "author_name": "moewie94",
      "author_url": "",
      "post_date": "12/19/2019 08:17:33",
      "content": "<p>so you mean all the videos labeled as FAKE in the training set have audio altered?</p>",
      "votes": null,
      "replies": [
        {
          "id": 699103,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/20/2019 03:45:55",
          "content": "<p>Don't think thats what he said - he said 8% of files had audio altered and they were all part of the fake set.  There are more than 8% fake in the full dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 698574,
      "author_name": "ameypatil",
      "author_url": "",
      "post_date": "12/19/2019 12:13:23",
      "content": "<p>Is there any instance when the audio corresponding to the video is altered but the video(face) is not? \nIdeally this is a possible case, but does the data have such instances and even if it does, model can only detect any altered-generated audio and not the one when the audio is pasted unaltered from some other speaker or source.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 699193,
      "author_name": "thanatoz",
      "author_url": "",
      "post_date": "12/20/2019 06:19:17",
      "content": "<p>I started a <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121268\">similar discussion</a> and <a href=\"/phoenix9032\">@phoenix9032</a> had an interesting question.\n&gt; How do you understand that without prior knowledge ? How do you know how someone speaks ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 701112,
          "author_name": "prakharg24",
          "author_url": "",
          "post_date": "12/23/2019 05:22:26",
          "content": "<p>If you visit the papers that you have mentioned in the discussion and dig deeper into this problem statement (see <a href=\"https://www.asvspoof.org/\">https://www.asvspoof.org/</a> ), you will find really interesting analysis of Spectographs by different people which show the differences between the original audios and faked audios (for example see <a href=\"https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35\">https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35</a>? ).</p>\n\n<p>The important point to remember I think in this case is that it is NOT an audio swap or a lip sync where someone else spoke the exact same thing. Instead it's an audio altercation or some other form of fake audio creation. I understand this kind of altercation is not as easily distinguishable just by listening but still if you do a more thorough analysis (like plot the spectographs) you will find the audios which are altered are clearly different from the originals. (see <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785</a> )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 700961,
      "author_name": "ke8ctn",
      "author_url": "",
      "post_date": "12/22/2019 21:53:10",
      "content": "<p>I also took a look at the audio differences in the fake clips. See here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 708955,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "01/02/2020 21:45:36",
      "content": "<p>Hi <a href=\"/ryches\">@ryches</a> , I can't understand something very basic about what you are saying here. Original is defined only for the fake videos, so haven't you checked audio alteration only for the fakes subset? If so, I don't understand the second sentence, that group is fakes by definition. Am I missing something?</p>",
      "votes": null,
      "replies": [
        {
          "id": 708959,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "01/02/2020 21:51:11",
          "content": "<p>I iterated through the full csv I have created with the name of every video and the video it references (none if it is an original video). While iterating through I would load the audio for that row and the video it was altered from. Compare the audio of those two. If they are equal or there is no source video (like it was an original, real video) then I marked it as 0 (unaltered audio), but if there was a source video and after comparing the two they were not the same then I marked it as 1 (altered audio). </p>\n\n<p>Not sure if that clears up any confusion</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708966,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/02/2020 22:03:13",
          "content": "<p>It doesn't clear my confusion yet. You say \"if there was a source video and after comparing the two they were not the same then I marked it as 1\", so every video marked as 1 had a source video. But only fakes videos have a source video. And this is my confusion, what makes the second sentence in your original post to be non-trivial?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708979,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "01/02/2020 22:16:02",
          "content": "<p>I guess that part is trivial because it is only really looking at the subset of fakes, but it is just confirmation the method did correctly label some videos audio as identical and others non-identical. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708983,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/02/2020 22:19:25",
          "content": "<p>I see, thank you for the clarification.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 709183,
      "author_name": "raibektussupbekov",
      "author_url": "",
      "post_date": "01/03/2020 06:12:06",
      "content": "<p><a href=\"/ryches\">@ryches</a> I used the same approach to find fake videos with modified audio, but I've found exactly 10,157 items, not 10,152.</p>",
      "votes": null,
      "replies": [
        {
          "id": 710668,
          "author_name": "ravivadapalli",
          "author_url": "",
          "post_date": "01/05/2020 03:58:19",
          "content": "<p><a href=\"/raibektussupbekov\">@raibektussupbekov</a> , thanks for sharing the csv.</p>\n\n<p>As per the csv, kpmieleslf.mp4 is an audio fake of dhjnjkzuhq.mp4. But if you listen to the audio of these two clips, the audio in kpmieleslf.mp4 seems to be same as that in its original. </p>\n\n<p>Am i missing something here ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710707,
          "author_name": "raibektussupbekov",
          "author_url": "",
          "post_date": "01/05/2020 05:17:12",
          "content": "<p>Some people here have noticed that it is the case with some modified audios. A fake and its original might sound similar despite the fact that they actually differ as data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 779067,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "03/19/2020 02:19:44",
          "content": "<p><a href=\"/raibektussupbekov\">@raibektussupbekov</a>. thank you for the csv file. Is it possible for you to publish a script that can produce this csv file?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 779532,
          "author_name": "raibektussupbekov",
          "author_url": "",
          "post_date": "03/19/2020 12:55:41",
          "content": "<p><a href=\"/projdev\">@projdev</a> I used this function.</p>\n\n<p>```\nimport glob\nimport json\nimport subprocess\nimport librosa\nimport numpy as np</p>\n\n<p>dfdc_train_wav_path = './disk1/dfdc_train_wav/'</p>\n\n<p>def audio_altered(fake_path, fake_video, orig_path, orig_video):</p>\n\n<pre><code>\"\"\"Finds out if audio of fake_video was altered.\n\n# Arguments\n    fake_path: fake mp4 video path name\n    fake_video: fake mp4 video name\n    orig_path: original mp4 video path name\n    orig_video: original mp4 video name\n\n# Returns\n    True - if audio of fake mp4 video was altered\n    False - otherwise\n\"\"\"\n\nfake_wav = fake_video[:-4] + '.wav'\nfake_wav_path = dfdc_train_wav_path + fake_wav\ntry:\n    # in case if .wav has already been extracted\n    fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\nexcept FileNotFoundError:\n    # extract fake_path audio\n    # .wav audio format is used because librosa.load() doesn't work with .aac\n    command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (fake_path, fake_wav_path)\n    subprocess.run(command, shell=True) \n\n    try:\n        fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n    except FileNotFoundError:\n        # if fake video has no audio than its audio is not altered\n        return False\n\norig_wav = orig_video[:-4] + '.wav'\norig_wav_path = dfdc_train_wav_path + orig_wav\ntry:\n    # in case if .wav has already been extracted\n    orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\nexcept FileNotFoundError:\n    # extract orig_path audio\n    # .wav audio format is used because librosa.load() doesn't work with .aac\n    command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (orig_path, orig_wav_path)\n    subprocess.run(command, shell=True)\n\n    try:\n        orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n    except FileNotFoundError:\n        # if original video has no audio but fake video does than audio is altered\n        return True\n\nreturn fake_rate != orig_rate or not np.array_equal(fake_data, orig_data)\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 715062,
      "author_name": "chenshen03",
      "author_url": "",
      "post_date": "01/10/2020 05:00:10",
      "content": "<p>In the 10,152 videos, is there any situation where the audio is FAKE but the face is REAL?</p>",
      "votes": null,
      "replies": [
        {
          "id": 715075,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/10/2020 05:20:49",
          "content": "<p>In all 10152 videos there is at least a minor change in pixel data, compared with the originals. In fact, this is true for all fakes. Some fakes barely change anything, some change not the face, there are duplicates between fakes, and duplicates between reals. But in general, the answer to your question is no.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715091,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "01/10/2020 05:50:45",
          "content": "<p>Thank you. So, when training an image-based binary classification model, there is no need to delete these audio-modified data, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715136,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/10/2020 06:47:34",
          "content": "<p>Right, but to be cautious I would say I don't see a reason to exclude them, so far. It was not publicly proven so far that their pixel level changes any different from all other fakes. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715149,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "01/10/2020 07:23:41",
          "content": "<p>Thx</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 731781,
      "author_name": "beeaware",
      "author_url": "",
      "post_date": "01/29/2020 03:32:52",
      "content": "<p>Great, Thanks, How do you find the exact num of fake audio？</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 756932,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "02/26/2020 08:29:58",
      "content": "<p>Did someone succeed to label fake audios only?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "696149": "Looked through the audio data and found about 8% have been altered in the training data in some way. This accurately maps to exactly 10152 that were marked as fake and zero that were marked as real.",
    "696222": "what do you consider as altered audio ? Would it be considered fake if it just goes through transcoding ?",
    "696236": "By altered I mean that the audio is not the exact same as the source video. How exactly it's different, I do not know, but it does perfectly map to fake labels so I am assuming they are considering any audio change at all to be a fake.",
    "696444": "Could you give examples of altered audios in the training/sample_training files?",
    "696525": "Didn't look in sample_training. xvzsbxrlxq.mp4 has altered audio in the training set.",
    "696588": "Oops. Sorry, I think I accidentally deleted my earlier comment. Edited it back to the best of my memory. Thanks! I found very blatant audio tamperings (eg: aomdezxgzc.mp4) in manual inspection. \n\nIt will be interesting to categorize what sortof audio tampering has been done in the data. I will reply here when I get more insight into this.",
    "697419": "Tell me what I’m missing because I’m new here... but wouldn’t a deepfake alter both the video and the audio as part of the fakery?  i.e. it’s true that if you always had the original video available identifying trickery would be trivial.  are you implying that altered video streams do NOT similarly perfectly call out the fakes?",
    "697710": "I did't find any .mp4 named aomdezxgzc.mp4? Can you share its location? For example, the belonging set?",
    "697719": "On my side, comparing byte-wise the first second with the source audio, I get about 4.5% of modified audio...\nI thought any re-encoding would have (statistically) changed 99.99% of the values (since my 1s tryout)... Did you compare the entire 10s audio files?",
    "697927": "It's in dfdc train part 49/ in the train set. Were able to locate it?",
    "697945": "Found it. Thank you very much.",
    "698282": "May I ask how do you judge that the audio of two mp4 files is different?\nI wanna hear the altered audio.\nThx in advance.",
    "698291": "I loaded the original audio and then compared to the videos that marked that video as the original. Then just applied a np array comparison to see if they were equal.",
    "698421": "Just found out that the average is 4.5% fake audio from directory 0 to 44. Dir 45 to 49 have 55% fake audio. Might be interesting to use 45-49 for a balanced set.",
    "698436": "so you mean all the videos labeled as FAKE in the training set have audio altered?",
    "698574": "Is there any instance when the audio corresponding to the video is altered but the video(face) is not? \nIdeally this is a possible case, but does the data have such instances and even if it does, model can only detect any altered-generated audio and not the one when the audio is pasted unaltered from some other speaker or source.",
    "698704": "Interesting point. While my code was running I had it print every 1000 files the percent and I saw huge fluctuations during run. That might be the explanation",
    "699078": "Thx. It helps a lot.\nI compare the audio of some videos labeled  as 'fake'. Indeed, I find some pair which have different audio np array. BUT when I listen to them, the audio of the pair hears the same.\nWould you please show me a pair whose audio hears different?\nThx in advance.",
    "699090": "Please read all the responses in this topic carefully, you will find the answer.",
    "699103": "Don't think thats what he said - he said 8% of files had audio altered and they were all part of the fake set.  There are more than 8% fake in the full dataset.",
    "699119": "yep, I find it. thx",
    "699193": "I started a [similar discussion](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121268) and @phoenix9032 had an interesting question.\n&gt; How do you understand that without prior knowledge ? How do you know how someone speaks ?",
    "700961": "I also took a look at the audio differences in the fake clips. See here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785",
    "701112": "If you visit the papers that you have mentioned in the discussion and dig deeper into this problem statement (see https://www.asvspoof.org/ ), you will find really interesting analysis of Spectographs by different people which show the differences between the original audios and faked audios (for example see https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35? ).\n\nThe important point to remember I think in this case is that it is NOT an audio swap or a lip sync where someone else spoke the exact same thing. Instead it's an audio altercation or some other form of fake audio creation. I understand this kind of altercation is not as easily distinguishable just by listening but still if you do a more thorough analysis (like plot the spectographs) you will find the audios which are altered are clearly different from the originals. (see https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122785 )",
    "708955": "Hi @ryches , I can't understand something very basic about what you are saying here. Original is defined only for the fake videos, so haven't you checked audio alteration only for the fakes subset? If so, I don't understand the second sentence, that group is fakes by definition. Am I missing something?",
    "708959": "I iterated through the full csv I have created with the name of every video and the video it references (none if it is an original video). While iterating through I would load the audio for that row and the video it was altered from. Compare the audio of those two. If they are equal or there is no source video (like it was an original, real video) then I marked it as 0 (unaltered audio), but if there was a source video and after comparing the two they were not the same then I marked it as 1 (altered audio). \n\nNot sure if that clears up any confusion",
    "708966": "It doesn't clear my confusion yet. You say \"if there was a source video and after comparing the two they were not the same then I marked it as 1\", so every video marked as 1 had a source video. But only fakes videos have a source video. And this is my confusion, what makes the second sentence in your original post to be non-trivial?",
    "708979": "I guess that part is trivial because it is only really looking at the subset of fakes, but it is just confirmation the method did correctly label some videos audio as identical and others non-identical.",
    "708983": "I see, thank you for the clarification.",
    "709183": "ryches I used the same approach to find fake videos with modified audio, but I've found exactly 10,157 items, not 10,152.",
    "710668": "raibektussupbekov , thanks for sharing the csv.\n\nAs per the csv, kpmieleslf.mp4 is an audio fake of dhjnjkzuhq.mp4. But if you listen to the audio of these two clips, the audio in kpmieleslf.mp4 seems to be same as that in its original. \n\nAm i missing something here ?",
    "710707": "Some people here have noticed that it is the case with some modified audios. A fake and its original might sound similar despite the fact that they actually differ as data.",
    "715062": "In the 10,152 videos, is there any situation where the audio is FAKE but the face is REAL?",
    "715075": "In all 10152 videos there is at least a minor change in pixel data, compared with the originals. In fact, this is true for all fakes. Some fakes barely change anything, some change not the face, there are duplicates between fakes, and duplicates between reals. But in general, the answer to your question is no.",
    "715091": "Thank you. So, when training an image-based binary classification model, there is no need to delete these audio-modified data, right?",
    "715136": "Right, but to be cautious I would say I don't see a reason to exclude them, so far. It was not publicly proven so far that their pixel level changes any different from all other fakes.",
    "715149": "Thx",
    "731781": "Great, Thanks, How do you find the exact num of fake audio？",
    "756932": "Did someone succeed to label fake audios only?",
    "779067": "raibektussupbekov. thank you for the csv file. Is it possible for you to publish a script that can produce this csv file?",
    "779532": "projdev I used this function.\n\n\n```\nimport glob\nimport json\nimport subprocess\nimport librosa\nimport numpy as np\n\ndfdc_train_wav_path = './disk1/dfdc_train_wav/'\n\ndef audio_altered(fake_path, fake_video, orig_path, orig_video):\n    \n    \"\"\"Finds out if audio of fake_video was altered.\n    \n    # Arguments\n        fake_path: fake mp4 video path name\n        fake_video: fake mp4 video name\n        orig_path: original mp4 video path name\n        orig_video: original mp4 video name\n        \n    # Returns\n        True - if audio of fake mp4 video was altered\n        False - otherwise\n    \"\"\"\n\n    fake_wav = fake_video[:-4] + '.wav'\n    fake_wav_path = dfdc_train_wav_path + fake_wav\n    try:\n        # in case if .wav has already been extracted\n        fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n    except FileNotFoundError:\n        # extract fake_path audio\n        # .wav audio format is used because librosa.load() doesn't work with .aac\n        command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (fake_path, fake_wav_path)\n        subprocess.run(command, shell=True) \n\n        try:\n            fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n        except FileNotFoundError:\n            # if fake video has no audio than its audio is not altered\n            return False\n    \n    orig_wav = orig_video[:-4] + '.wav'\n    orig_wav_path = dfdc_train_wav_path + orig_wav\n    try:\n        # in case if .wav has already been extracted\n        orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n    except FileNotFoundError:\n        # extract orig_path audio\n        # .wav audio format is used because librosa.load() doesn't work with .aac\n        command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (orig_path, orig_wav_path)\n        subprocess.run(command, shell=True)\n\n        try:\n            orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n        except FileNotFoundError:\n            # if original video has no audio but fake video does than audio is altered\n            return True\n    \n    return fake_rate != orig_rate or not np.array_equal(fake_data, orig_data)\n```"
  },
  "source": "meta"
}