{
  "id": 122785,
  "title": "Finding the fake audio",
  "url": "/competitions/deepfake-detection-challenge/discussion/122785",
  "author_name": "",
  "post_date": "2019-12-22T21:51:17.146763200Z",
  "votes": 58,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I decided to do a bit of analysis on the fake clips in the training set so I could add a label for fake audio. </p>\n\n<p>For each video marked as fake in the original metadata, I compare the audio between the fake video and the original. I quickly noticed lots of sample rate mismatches and missing audio tracks between the originals and fakes, so I take this into consideration. If any of the audio samples differ, or if the sample rates do not match, or if the fake video contains audio and the original does not, I mark the video as containing modified audio. If the samples are all identical, or if the fake video contains no audio, I mark the video as NOT containing modified audio. Using these criteria, I find exactly 10,152 videos with modified audio, which agrees with ryches' findings (see <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861</a>). Of these, 4,209 have unchanged sample rates but some difference in the samples, while 5,249 have different rates.</p>\n\n<p>After finding the clips with differing audio, I did a quick check, and the differences in the audio seem to be subtle in many cases. For the videos where the sample rate does not change and both the original and fake video contain audio, I independently normalize each video's audio track and then do three sets of comparisons to find the pairs with 1) highest peak absolute error, 2) highest mean absolute error, and 3) lowest correlation. When I listen to the clips with the greatest deviation from their sources, I perceive almost no difference, though plotting the error does show that a difference exists. This to me suggests that these clips have differing audio due solely to transcoding, dynamics compression, noise removal, or other \"benign\" sorts of processing as opposed to faked audio. If you want to compare them yourself, the pair with the max peak absolute error is (qdwcydrmpl.mp4, tgkjhdyxww.mp4), the pair with the max mean absolute error is (svdjmhpwnb.mp4, xrluwubrhs.mp4), and the pair with the minimum corellation  is (dvzwkvoefx.mp4, odlidgmzqf.mp4). None of these seems to have significant perceivable differences, and if you listen to the error signal, it sounds like noise or compression artifacts.</p>\n\n<p>By contrast, when I listen to the clips with sample rates that differ from the original clips, the difference is much more striking. There are noticeable distortions, changes in pitch, and other artifacts that are very different from typical compression artifacts. Check out (tazlckybzu.mp4, cukedccddh.mp4) and (sazcnznkjo.mp4, bcoydxioxf.mp4) for examples. Interestingly, <strong>all of these rate-changed clips appear to be located in the higher-numbered subdirectories</strong> of the training archive, starting with #45. That means that there's a significant imbalance in the training set with respect to the organization of the files, which is important to note when training with only a subset of the data. If you're training on subsets of the training set due to hardware limitations, you probably want to have many of the subdirectories represented in each batch/epoch, rather than training on the data from one or two subdirectories at a time.</p>",
  "messages": [
    {
      "id": "700957",
      "postDate": "12/22/2019 21:51:17",
      "content": "<p>I decided to do a bit of analysis on the fake clips in the training set so I could add a label for fake audio. </p>\n\n<p>For each video marked as fake in the original metadata, I compare the audio between the fake video and the original. I quickly noticed lots of sample rate mismatches and missing audio tracks between the originals and fakes, so I take this into consideration. If any of the audio samples differ, or if the sample rates do not match, or if the fake video contains audio and the original does not, I mark the video as containing modified audio. If the samples are all identical, or if the fake video contains no audio, I mark the video as NOT containing modified audio. Using these criteria, I find exactly 10,152 videos with modified audio, which agrees with ryches' findings (see <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861</a>). Of these, 4,209 have unchanged sample rates but some difference in the samples, while 5,249 have different rates.</p>\n\n<p>After finding the clips with differing audio, I did a quick check, and the differences in the audio seem to be subtle in many cases. For the videos where the sample rate does not change and both the original and fake video contain audio, I independently normalize each video's audio track and then do three sets of comparisons to find the pairs with 1) highest peak absolute error, 2) highest mean absolute error, and 3) lowest correlation. When I listen to the clips with the greatest deviation from their sources, I perceive almost no difference, though plotting the error does show that a difference exists. This to me suggests that these clips have differing audio due solely to transcoding, dynamics compression, noise removal, or other \"benign\" sorts of processing as opposed to faked audio. If you want to compare them yourself, the pair with the max peak absolute error is (qdwcydrmpl.mp4, tgkjhdyxww.mp4), the pair with the max mean absolute error is (svdjmhpwnb.mp4, xrluwubrhs.mp4), and the pair with the minimum corellation  is (dvzwkvoefx.mp4, odlidgmzqf.mp4). None of these seems to have significant perceivable differences, and if you listen to the error signal, it sounds like noise or compression artifacts.</p>\n\n<p>By contrast, when I listen to the clips with sample rates that differ from the original clips, the difference is much more striking. There are noticeable distortions, changes in pitch, and other artifacts that are very different from typical compression artifacts. Check out (tazlckybzu.mp4, cukedccddh.mp4) and (sazcnznkjo.mp4, bcoydxioxf.mp4) for examples. Interestingly, <strong>all of these rate-changed clips appear to be located in the higher-numbered subdirectories</strong> of the training archive, starting with #45. That means that there's a significant imbalance in the training set with respect to the organization of the files, which is important to note when training with only a subset of the data. If you're training on subsets of the training set due to hardware limitations, you probably want to have many of the subdirectories represented in each batch/epoch, rather than training on the data from one or two subdirectories at a time.</p>",
      "rawMarkdown": "I decided to do a bit of analysis on the fake clips in the training set so I could add a label for fake audio. \n\nFor each video marked as fake in the original metadata, I compare the audio between the fake video and the original. I quickly noticed lots of sample rate mismatches and missing audio tracks between the originals and fakes, so I take this into consideration. If any of the audio samples differ, or if the sample rates do not match, or if the fake video contains audio and the original does not, I mark the video as containing modified audio. If the samples are all identical, or if the fake video contains no audio, I mark the video as NOT containing modified audio. Using these criteria, I find exactly 10,152 videos with modified audio, which agrees with ryches' findings (see https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861). Of these, 4,209 have unchanged sample rates but some difference in the samples, while 5,249 have different rates.\n\nAfter finding the clips with differing audio, I did a quick check, and the differences in the audio seem to be subtle in many cases. For the videos where the sample rate does not change and both the original and fake video contain audio, I independently normalize each video's audio track and then do three sets of comparisons to find the pairs with 1) highest peak absolute error, 2) highest mean absolute error, and 3) lowest correlation. When I listen to the clips with the greatest deviation from their sources, I perceive almost no difference, though plotting the error does show that a difference exists. This to me suggests that these clips have differing audio due solely to transcoding, dynamics compression, noise removal, or other \"benign\" sorts of processing as opposed to faked audio. If you want to compare them yourself, the pair with the max peak absolute error is (qdwcydrmpl.mp4, tgkjhdyxww.mp4), the pair with the max mean absolute error is (svdjmhpwnb.mp4, xrluwubrhs.mp4), and the pair with the minimum corellation  is (dvzwkvoefx.mp4, odlidgmzqf.mp4). None of these seems to have significant perceivable differences, and if you listen to the error signal, it sounds like noise or compression artifacts.\n\nBy contrast, when I listen to the clips with sample rates that differ from the original clips, the difference is much more striking. There are noticeable distortions, changes in pitch, and other artifacts that are very different from typical compression artifacts. Check out (tazlckybzu.mp4, cukedccddh.mp4) and (sazcnznkjo.mp4, bcoydxioxf.mp4) for examples. Interestingly, **all of these rate-changed clips appear to be located in the higher-numbered subdirectories** of the training archive, starting with #45. That means that there's a significant imbalance in the training set with respect to the organization of the files, which is important to note when training with only a subset of the data. If you're training on subsets of the training set due to hardware limitations, you probably want to have many of the subdirectories represented in each batch/epoch, rather than training on the data from one or two subdirectories at a time.",
      "votes": null
    },
    {
      "id": "701044",
      "postDate": "12/23/2019 02:48:33",
      "content": "<p>Coooooool man, nice work.\nDoes the unmodified audio have a common sample rate?\nIf so, it may help the detection.</p>",
      "rawMarkdown": "Coooooool man, nice work.\nDoes the unmodified audio have a common sample rate?\nIf so, it may help the detection.",
      "votes": null
    },
    {
      "id": "701074",
      "postDate": "12/23/2019 03:40:23",
      "content": "<p>Yes, the exact imbalance is shown in this thread. <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861#698421\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861#698421</a></p>",
      "rawMarkdown": "Yes, the exact imbalance is shown in this thread. https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861#698421",
      "votes": null
    },
    {
      "id": "701080",
      "postDate": "12/23/2019 03:52:12",
      "content": "<p>In the training data, yes. All of the fake audio seems to be 16 kHz, while the original is either 44.1 or 48 kHz.</p>\n\n<p>It's certainly something to be aware of during training, since none of the modified audio will have frequency content above 8 kHz, but you probably shouldn't rely too much on this \"feature\" for the test set.</p>",
      "rawMarkdown": "In the training data, yes. All of the fake audio seems to be 16 kHz, while the original is either 44.1 or 48 kHz.\n\nIt's certainly something to be aware of during training, since none of the modified audio will have frequency content above 8 kHz, but you probably shouldn't rely too much on this \"feature\" for the test set.",
      "votes": null
    },
    {
      "id": "701083",
      "postDate": "12/23/2019 03:56:41",
      "content": "<p>I claim that the imbalance is more extreme than that post suggests, with <em>all</em> of the faked audio in the last few subdirectories.</p>",
      "rawMarkdown": "I claim that the imbalance is more extreme than that post suggests, with *all* of the faked audio in the last few subdirectories.",
      "votes": null
    },
    {
      "id": "701096",
      "postDate": "12/23/2019 04:45:04",
      "content": "<p>Thx, friend. \nI found some videos with different audios, but I cannot hear the difference. Now thanks to your analysis, I know they share the same sample rate.\nI shall look into those audios with different sample rate.\nAnd again, thank you for sharing!</p>",
      "rawMarkdown": "Thx, friend. \nI found some videos with different audios, but I cannot hear the difference. Now thanks to your analysis, I know they share the same sample rate.\nI shall look into those audios with different sample rate.\nAnd again, thank you for sharing!",
      "votes": null
    },
    {
      "id": "701143",
      "postDate": "12/23/2019 06:38:05",
      "content": "<p>Thank you for the share. Just curious what do you think would be the best direction when it comes to audio. Is it better to combine the audio features together with the video features then feed it to the network, or to have two separate detector for vid and audio then merge at the end ? </p>",
      "rawMarkdown": "Thank you for the share. Just curious what do you think would be the best direction when it comes to audio. Is it better to combine the audio features together with the video features then feed it to the network, or to have two separate detector for vid and audio then merge at the end ?",
      "votes": null
    },
    {
      "id": "701153",
      "postDate": "12/23/2019 06:52:35",
      "content": "<p>Even when you check the audio samples of Tacotron which are under-trained, you will notice that the undertrained model, rather than creating audible artefacts in sounds, creates more muffled audio. Is this a result of variable audio sample rate and will be treated as noise too? </p>",
      "rawMarkdown": "Even when you check the audio samples of Tacotron which are under-trained, you will notice that the undertrained model, rather than creating audible artefacts in sounds, creates more muffled audio. Is this a result of variable audio sample rate and will be treated as noise too?",
      "votes": null
    },
    {
      "id": "705877",
      "postDate": "12/29/2019 15:54:51",
      "content": "<p><a href=\"/ke8ctn\">@ke8ctn</a> what library and output format did you use for conversion? ffmpeg and mp3?</p>",
      "rawMarkdown": "ke8ctn what library and output format did you use for conversion? ffmpeg and mp3?",
      "votes": null
    },
    {
      "id": "705892",
      "postDate": "12/29/2019 16:31:27",
      "content": "<p>I used ffmpeg to extract the audio as WAV files.</p>",
      "rawMarkdown": "I used ffmpeg to extract the audio as WAV files.",
      "votes": null
    },
    {
      "id": "709182",
      "postDate": "01/03/2020 06:11:12",
      "content": "<p><a href=\"/ke8ctn\">@ke8ctn</a> I used the same approach to find fake videos with modified audio, but I've found exactly 10,15*<em>7</em>* items, not 10,15*<em>2</em>*.</p>",
      "rawMarkdown": "ke8ctn I used the same approach to find fake videos with modified audio, but I've found exactly 10,15**7** items, not 10,15**2**.",
      "votes": null
    },
    {
      "id": "710034",
      "postDate": "01/04/2020 08:42:09",
      "content": "<p>Thanks. I just wanted to ask if anyone can upload or give a link to all metadata files. Due to hardware limitations it is much hard for me. I'll really appreciate the help.</p>",
      "rawMarkdown": "Thanks. I just wanted to ask if anyone can upload or give a link to all metadata files. Due to hardware limitations it is much hard for me. I'll really appreciate the help.",
      "votes": null
    },
    {
      "id": "710514",
      "postDate": "01/04/2020 20:32:10",
      "content": "<p>Interesting.  Are you sure you're handling missing audio the same way? If so, I'll try to have a look soon to see if I can account for the differences in our results</p>",
      "rawMarkdown": "Interesting.  Are you sure you're handling missing audio the same way? If so, I'll try to have a look soon to see if I can account for the differences in our results",
      "votes": null
    },
    {
      "id": "710521",
      "postDate": "01/04/2020 20:54:31",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc</a></p>",
      "rawMarkdown": "https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc",
      "votes": null
    },
    {
      "id": "710711",
      "postDate": "01/05/2020 05:40:33",
      "content": "<p>I dropped 8 missing mp4 files only after extracting and comparing audios. But it seems to me that doesn't influence the result. Anyway this is my code, maybe it's wrong in some cases.</p>\n\n<p>```\nimport glob\nimport json\nimport subprocess\nimport librosa\nimport numpy as np</p>\n\n<p>dfdc_train_wav_path = \"./disk1/dfdc_train_wav/\"</p>\n\n<p>def audio_altered(fake_path, fake_video, orig_path, orig_video):\n    \"\"\"Finds out if audio of fake_video was altered.</p>\n\n<pre><code># Arguments\n    fake_path: fake mp4 video path name\n    fake_video: fake mp4 video name\n    orig_path: original mp4 video path name\n    orig_video: original mp4 video name\n\n# Returns\n    True - if audio of fake mp4 video was altered\n    False - otherwise\n\"\"\"\nfake_wav = fake_video[:-4] + '.wav'\nfake_wav_path = dfdc_train_wav_path + fake_wav\ntry:\n    # in case if .wav has already been extracted\n    fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\nexcept FileNotFoundError:\n    # extract fake_path audio\n    # .wav audio format is used because librosa.load() doesn't work with .aac\n    command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (fake_path, fake_wav_path)\n    subprocess.run(command, shell=True) \n\n    try:\n        fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n    except FileNotFoundError:\n        # if fake video has no audio than its audio is not altered\n        return False\n\norig_wav = orig_video[:-4] + '.wav'\norig_wav_path = dfdc_train_wav_path + orig_wav\ntry:\n    # in case if .wav has already been extracted\n    orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\nexcept FileNotFoundError:\n    # extract orig_path audio\n    # .wav audio format is used because librosa.load() doesn't work with .aac\n    command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (orig_path, orig_wav_path)\n    subprocess.run(command, shell=True)\n\n    try:\n        orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n    except FileNotFoundError:\n        # if original video has no audio but fake video does than audio is altered\n        return True\n\nreturn fake_rate != orig_rate or not np.array_equal(fake_data, orig_data)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I dropped 8 missing mp4 files only after extracting and comparing audios. But it seems to me that doesn't influence the result. Anyway this is my code, maybe it's wrong in some cases.\n\n```\nimport glob\nimport json\nimport subprocess\nimport librosa\nimport numpy as np\n\ndfdc_train_wav_path = \"./disk1/dfdc_train_wav/\"\n\ndef audio_altered(fake_path, fake_video, orig_path, orig_video):\n    \"\"\"Finds out if audio of fake_video was altered.\n    \n    # Arguments\n        fake_path: fake mp4 video path name\n        fake_video: fake mp4 video name\n        orig_path: original mp4 video path name\n        orig_video: original mp4 video name\n        \n    # Returns\n        True - if audio of fake mp4 video was altered\n        False - otherwise\n    \"\"\"\n    fake_wav = fake_video[:-4] + '.wav'\n    fake_wav_path = dfdc_train_wav_path + fake_wav\n    try:\n        # in case if .wav has already been extracted\n        fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n    except FileNotFoundError:\n        # extract fake_path audio\n        # .wav audio format is used because librosa.load() doesn't work with .aac\n        command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (fake_path, fake_wav_path)\n        subprocess.run(command, shell=True) \n\n        try:\n            fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n        except FileNotFoundError:\n            # if fake video has no audio than its audio is not altered\n            return False\n    \n    orig_wav = orig_video[:-4] + '.wav'\n    orig_wav_path = dfdc_train_wav_path + orig_wav\n    try:\n        # in case if .wav has already been extracted\n        orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n    except FileNotFoundError:\n        # extract orig_path audio\n        # .wav audio format is used because librosa.load() doesn't work with .aac\n        command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (orig_path, orig_wav_path)\n        subprocess.run(command, shell=True)\n\n        try:\n            orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n        except FileNotFoundError:\n            # if original video has no audio but fake video does than audio is altered\n            return True\n    \n    return fake_rate != orig_rate or not np.array_equal(fake_data, orig_data)\n```",
      "votes": null
    },
    {
      "id": "731774",
      "postDate": "01/29/2020 03:13:54",
      "content": "<p>Great idea, audio is important.</p>",
      "rawMarkdown": "Great idea, audio is important.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 701044,
      "author_name": "daddyjin",
      "author_url": "",
      "post_date": "12/23/2019 02:48:33",
      "content": "<p>Coooooool man, nice work.\nDoes the unmodified audio have a common sample rate?\nIf so, it may help the detection.</p>",
      "votes": null,
      "replies": [
        {
          "id": 701080,
          "author_name": "ke8ctn",
          "author_url": "",
          "post_date": "12/23/2019 03:52:12",
          "content": "<p>In the training data, yes. All of the fake audio seems to be 16 kHz, while the original is either 44.1 or 48 kHz.</p>\n\n<p>It's certainly something to be aware of during training, since none of the modified audio will have frequency content above 8 kHz, but you probably shouldn't rely too much on this \"feature\" for the test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 701096,
          "author_name": "daddyjin",
          "author_url": "",
          "post_date": "12/23/2019 04:45:04",
          "content": "<p>Thx, friend. \nI found some videos with different audios, but I cannot hear the difference. Now thanks to your analysis, I know they share the same sample rate.\nI shall look into those audios with different sample rate.\nAnd again, thank you for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 701074,
      "author_name": "skrish13",
      "author_url": "",
      "post_date": "12/23/2019 03:40:23",
      "content": "<p>Yes, the exact imbalance is shown in this thread. <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861#698421\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861#698421</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 701083,
          "author_name": "ke8ctn",
          "author_url": "",
          "post_date": "12/23/2019 03:56:41",
          "content": "<p>I claim that the imbalance is more extreme than that post suggests, with <em>all</em> of the faked audio in the last few subdirectories.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 701143,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "12/23/2019 06:38:05",
      "content": "<p>Thank you for the share. Just curious what do you think would be the best direction when it comes to audio. Is it better to combine the audio features together with the video features then feed it to the network, or to have two separate detector for vid and audio then merge at the end ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 701153,
      "author_name": "thanatoz",
      "author_url": "",
      "post_date": "12/23/2019 06:52:35",
      "content": "<p>Even when you check the audio samples of Tacotron which are under-trained, you will notice that the undertrained model, rather than creating audible artefacts in sounds, creates more muffled audio. Is this a result of variable audio sample rate and will be treated as noise too? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 705877,
      "author_name": "drsilicon",
      "author_url": "",
      "post_date": "12/29/2019 15:54:51",
      "content": "<p><a href=\"/ke8ctn\">@ke8ctn</a> what library and output format did you use for conversion? ffmpeg and mp3?</p>",
      "votes": null,
      "replies": [
        {
          "id": 705892,
          "author_name": "ke8ctn",
          "author_url": "",
          "post_date": "12/29/2019 16:31:27",
          "content": "<p>I used ffmpeg to extract the audio as WAV files.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 709182,
      "author_name": "raibektussupbekov",
      "author_url": "",
      "post_date": "01/03/2020 06:11:12",
      "content": "<p><a href=\"/ke8ctn\">@ke8ctn</a> I used the same approach to find fake videos with modified audio, but I've found exactly 10,15*<em>7</em>* items, not 10,15*<em>2</em>*.</p>",
      "votes": null,
      "replies": [
        {
          "id": 710514,
          "author_name": "ke8ctn",
          "author_url": "",
          "post_date": "01/04/2020 20:32:10",
          "content": "<p>Interesting.  Are you sure you're handling missing audio the same way? If so, I'll try to have a look soon to see if I can account for the differences in our results</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710711,
          "author_name": "raibektussupbekov",
          "author_url": "",
          "post_date": "01/05/2020 05:40:33",
          "content": "<p>I dropped 8 missing mp4 files only after extracting and comparing audios. But it seems to me that doesn't influence the result. Anyway this is my code, maybe it's wrong in some cases.</p>\n\n<p>```\nimport glob\nimport json\nimport subprocess\nimport librosa\nimport numpy as np</p>\n\n<p>dfdc_train_wav_path = \"./disk1/dfdc_train_wav/\"</p>\n\n<p>def audio_altered(fake_path, fake_video, orig_path, orig_video):\n    \"\"\"Finds out if audio of fake_video was altered.</p>\n\n<pre><code># Arguments\n    fake_path: fake mp4 video path name\n    fake_video: fake mp4 video name\n    orig_path: original mp4 video path name\n    orig_video: original mp4 video name\n\n# Returns\n    True - if audio of fake mp4 video was altered\n    False - otherwise\n\"\"\"\nfake_wav = fake_video[:-4] + '.wav'\nfake_wav_path = dfdc_train_wav_path + fake_wav\ntry:\n    # in case if .wav has already been extracted\n    fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\nexcept FileNotFoundError:\n    # extract fake_path audio\n    # .wav audio format is used because librosa.load() doesn't work with .aac\n    command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (fake_path, fake_wav_path)\n    subprocess.run(command, shell=True) \n\n    try:\n        fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n    except FileNotFoundError:\n        # if fake video has no audio than its audio is not altered\n        return False\n\norig_wav = orig_video[:-4] + '.wav'\norig_wav_path = dfdc_train_wav_path + orig_wav\ntry:\n    # in case if .wav has already been extracted\n    orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\nexcept FileNotFoundError:\n    # extract orig_path audio\n    # .wav audio format is used because librosa.load() doesn't work with .aac\n    command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (orig_path, orig_wav_path)\n    subprocess.run(command, shell=True)\n\n    try:\n        orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n    except FileNotFoundError:\n        # if original video has no audio but fake video does than audio is altered\n        return True\n\nreturn fake_rate != orig_rate or not np.array_equal(fake_data, orig_data)\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 710034,
      "author_name": "ankitsainiankit",
      "author_url": "",
      "post_date": "01/04/2020 08:42:09",
      "content": "<p>Thanks. I just wanted to ask if anyone can upload or give a link to all metadata files. Due to hardware limitations it is much hard for me. I'll really appreciate the help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 710521,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/04/2020 20:54:31",
          "content": "<p><a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 731774,
      "author_name": "beeaware",
      "author_url": "",
      "post_date": "01/29/2020 03:13:54",
      "content": "<p>Great idea, audio is important.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "700957": "I decided to do a bit of analysis on the fake clips in the training set so I could add a label for fake audio. \n\nFor each video marked as fake in the original metadata, I compare the audio between the fake video and the original. I quickly noticed lots of sample rate mismatches and missing audio tracks between the originals and fakes, so I take this into consideration. If any of the audio samples differ, or if the sample rates do not match, or if the fake video contains audio and the original does not, I mark the video as containing modified audio. If the samples are all identical, or if the fake video contains no audio, I mark the video as NOT containing modified audio. Using these criteria, I find exactly 10,152 videos with modified audio, which agrees with ryches' findings (see https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861). Of these, 4,209 have unchanged sample rates but some difference in the samples, while 5,249 have different rates.\n\nAfter finding the clips with differing audio, I did a quick check, and the differences in the audio seem to be subtle in many cases. For the videos where the sample rate does not change and both the original and fake video contain audio, I independently normalize each video's audio track and then do three sets of comparisons to find the pairs with 1) highest peak absolute error, 2) highest mean absolute error, and 3) lowest correlation. When I listen to the clips with the greatest deviation from their sources, I perceive almost no difference, though plotting the error does show that a difference exists. This to me suggests that these clips have differing audio due solely to transcoding, dynamics compression, noise removal, or other \"benign\" sorts of processing as opposed to faked audio. If you want to compare them yourself, the pair with the max peak absolute error is (qdwcydrmpl.mp4, tgkjhdyxww.mp4), the pair with the max mean absolute error is (svdjmhpwnb.mp4, xrluwubrhs.mp4), and the pair with the minimum corellation  is (dvzwkvoefx.mp4, odlidgmzqf.mp4). None of these seems to have significant perceivable differences, and if you listen to the error signal, it sounds like noise or compression artifacts.\n\nBy contrast, when I listen to the clips with sample rates that differ from the original clips, the difference is much more striking. There are noticeable distortions, changes in pitch, and other artifacts that are very different from typical compression artifacts. Check out (tazlckybzu.mp4, cukedccddh.mp4) and (sazcnznkjo.mp4, bcoydxioxf.mp4) for examples. Interestingly, **all of these rate-changed clips appear to be located in the higher-numbered subdirectories** of the training archive, starting with #45. That means that there's a significant imbalance in the training set with respect to the organization of the files, which is important to note when training with only a subset of the data. If you're training on subsets of the training set due to hardware limitations, you probably want to have many of the subdirectories represented in each batch/epoch, rather than training on the data from one or two subdirectories at a time.",
    "701044": "Coooooool man, nice work.\nDoes the unmodified audio have a common sample rate?\nIf so, it may help the detection.",
    "701074": "Yes, the exact imbalance is shown in this thread. https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121861#698421",
    "701080": "In the training data, yes. All of the fake audio seems to be 16 kHz, while the original is either 44.1 or 48 kHz.\n\nIt's certainly something to be aware of during training, since none of the modified audio will have frequency content above 8 kHz, but you probably shouldn't rely too much on this \"feature\" for the test set.",
    "701083": "I claim that the imbalance is more extreme than that post suggests, with *all* of the faked audio in the last few subdirectories.",
    "701096": "Thx, friend. \nI found some videos with different audios, but I cannot hear the difference. Now thanks to your analysis, I know they share the same sample rate.\nI shall look into those audios with different sample rate.\nAnd again, thank you for sharing!",
    "701143": "Thank you for the share. Just curious what do you think would be the best direction when it comes to audio. Is it better to combine the audio features together with the video features then feed it to the network, or to have two separate detector for vid and audio then merge at the end ?",
    "701153": "Even when you check the audio samples of Tacotron which are under-trained, you will notice that the undertrained model, rather than creating audible artefacts in sounds, creates more muffled audio. Is this a result of variable audio sample rate and will be treated as noise too?",
    "705877": "ke8ctn what library and output format did you use for conversion? ffmpeg and mp3?",
    "705892": "I used ffmpeg to extract the audio as WAV files.",
    "709182": "ke8ctn I used the same approach to find fake videos with modified audio, but I've found exactly 10,15**7** items, not 10,15**2**.",
    "710034": "Thanks. I just wanted to ask if anyone can upload or give a link to all metadata files. Due to hardware limitations it is much hard for me. I'll really appreciate the help.",
    "710514": "Interesting.  Are you sure you're handling missing audio the same way? If so, I'll try to have a look soon to see if I can account for the differences in our results",
    "710521": "https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc",
    "710711": "I dropped 8 missing mp4 files only after extracting and comparing audios. But it seems to me that doesn't influence the result. Anyway this is my code, maybe it's wrong in some cases.\n\n```\nimport glob\nimport json\nimport subprocess\nimport librosa\nimport numpy as np\n\ndfdc_train_wav_path = \"./disk1/dfdc_train_wav/\"\n\ndef audio_altered(fake_path, fake_video, orig_path, orig_video):\n    \"\"\"Finds out if audio of fake_video was altered.\n    \n    # Arguments\n        fake_path: fake mp4 video path name\n        fake_video: fake mp4 video name\n        orig_path: original mp4 video path name\n        orig_video: original mp4 video name\n        \n    # Returns\n        True - if audio of fake mp4 video was altered\n        False - otherwise\n    \"\"\"\n    fake_wav = fake_video[:-4] + '.wav'\n    fake_wav_path = dfdc_train_wav_path + fake_wav\n    try:\n        # in case if .wav has already been extracted\n        fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n    except FileNotFoundError:\n        # extract fake_path audio\n        # .wav audio format is used because librosa.load() doesn't work with .aac\n        command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (fake_path, fake_wav_path)\n        subprocess.run(command, shell=True) \n\n        try:\n            fake_data, fake_rate = librosa.load(fake_wav_path, sr=None)\n        except FileNotFoundError:\n            # if fake video has no audio than its audio is not altered\n            return False\n    \n    orig_wav = orig_video[:-4] + '.wav'\n    orig_wav_path = dfdc_train_wav_path + orig_wav\n    try:\n        # in case if .wav has already been extracted\n        orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n    except FileNotFoundError:\n        # extract orig_path audio\n        # .wav audio format is used because librosa.load() doesn't work with .aac\n        command = \"./ffmpeg-git-amd64-static/ffmpeg -i %s -vn -f wav %s\" % (orig_path, orig_wav_path)\n        subprocess.run(command, shell=True)\n\n        try:\n            orig_data, orig_rate = librosa.load(orig_wav_path, sr=None)\n        except FileNotFoundError:\n            # if original video has no audio but fake video does than audio is altered\n            return True\n    \n    return fake_rate != orig_rate or not np.array_equal(fake_data, orig_data)\n```",
    "731774": "Great idea, audio is important."
  },
  "source": "meta"
}