{
  "id": 130502,
  "title": "Fake Audio Detection Broad Guidance",
  "url": "/competitions/deepfake-detection-challenge/discussion/130502",
  "author_name": "student",
  "post_date": "2020-02-14T14:39:56.761000",
  "votes": 11,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Folks: </p>\n\n<p>There is a lot of literature on Fake video; looking for some broad guidance on fake audio detection. I would really appreciate any pointers on literature or even cryptic clues:) Thank you so much.</p>",
  "messages": [
    {
      "id": 746045,
      "postDate": "2020-02-14T14:39:56.760Z",
      "content": "<p>Folks: </p>\n\n<p>There is a lot of literature on Fake video; looking for some broad guidance on fake audio detection. I would really appreciate any pointers on literature or even cryptic clues:) Thank you so much.</p>",
      "rawMarkdown": "Folks: \n\nThere is a lot of literature on Fake video; looking for some broad guidance on fake audio detection. I would really appreciate any pointers on literature or even cryptic clues:) Thank you so much.",
      "votes": 11
    },
    {
      "id": 746098,
      "postDate": "2020-02-14T16:07:29.857Z",
      "content": "<p>This article might help. <a href=\"https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35\">https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35</a></p>\n\n<p>The general idea is to convert audio to spectrograms and train an image classifier on the audio spectrograms. </p>",
      "rawMarkdown": "This article might help. https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35\n\nThe general idea is to convert audio to spectrograms and train an image classifier on the audio spectrograms. ",
      "votes": 2,
      "replies": [
        {
          "id": 746115,
          "postDate": "2020-02-14T16:37:32.523Z",
          "content": "<p>So I tried this, even tried using their code to extract the spectrograms, of various resolutions.</p>\n\n<p>I tried <em>so</em> many methods then to classify the spectrograms, including:\n- Classification using 2D CNNs \n- Classification using 1D CNNs (like they used)\n- Segmentation of the abnormal part of the spectrograms using a 2D HRNet\n- Segmentation of the abnormal part of the spectrograms using a UNet with 1D convolutions</p>\n\n<p>I tried doing the above with monochrome pictures, i.e. 1 input channel, and also the coloured spectrograms their code provides.</p>\n\n<p>All of them were incredibly underwhelming with ~65% accuracy at best. Given only a small proportion of videos in the training set even had altered audio, it just didn't seem worth pursuing further. However, keen to heard if anyone else has found otherwise?</p>",
          "rawMarkdown": "So I tried this, even tried using their code to extract the spectrograms, of various resolutions.\n\nI tried *so* many methods then to classify the spectrograms, including:\n- Classification using 2D CNNs \n- Classification using 1D CNNs (like they used)\n- Segmentation of the abnormal part of the spectrograms using a 2D HRNet\n- Segmentation of the abnormal part of the spectrograms using a UNet with 1D convolutions\n\nI tried doing the above with monochrome pictures, i.e. 1 input channel, and also the coloured spectrograms their code provides.\n\nAll of them were incredibly underwhelming with ~65% accuracy at best. Given only a small proportion of videos in the training set even had altered audio, it just didn't seem worth pursuing further. However, keen to heard if anyone else has found otherwise?\n\n",
          "votes": 18
        },
        {
          "id": 746131,
          "postDate": "2020-02-14T16:57:54.563Z",
          "content": "<p>Thank you so much. Will pursue it and report back if it helps. Thanks again</p>",
          "rawMarkdown": "Thank you so much. Will pursue it and report back if it helps. Thanks again",
          "votes": 2
        },
        {
          "id": 746171,
          "postDate": "2020-02-14T17:46:46.197Z",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> thanks for sharing your experience with this. I had tried combining audio spectrum images with frames from the video and running them through a 3dCNN but the model loss plateaued at ~0.69 so essentially it was just predicting random. Interesting to see you can get to 0.3 without accounting for audio. Gives me a better idea where to focus my efforts and what I can expect without using the audio. </p>",
          "rawMarkdown": "@jamesphoward thanks for sharing your experience with this. I had tried combining audio spectrum images with frames from the video and running them through a 3dCNN but the model loss plateaued at ~0.69 so essentially it was just predicting random. Interesting to see you can get to 0.3 without accounting for audio. Gives me a better idea where to focus my efforts and what I can expect without using the audio. ",
          "votes": 1
        },
        {
          "id": 746473,
          "postDate": "2020-02-15T04:00:48.517Z",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> what was your training data. All the training parts??</p>",
          "rawMarkdown": "@jamesphoward what was your training data. All the training parts??"
        },
        {
          "id": 746479,
          "postDate": "2020-02-15T04:21:21.057Z",
          "content": "<p>I went through every fake video and diff'ed its spectrogram versus that of the original video; any differnce and I called it fake audio, otherwise it was real. I ended up with 4k fake spectrograms and 80+k real spectrograms. I then undersampled the real ones so they were 1:1 before training as above.</p>\n\n<p>EDIT: Also, when using the segmentation method, I tried randomly cropping the spectrogram and the mask to synthesise a bit more data. Still didn't work.</p>",
          "rawMarkdown": "I went through every fake video and diff'ed its spectrogram versus that of the original video; any differnce and I called it fake audio, otherwise it was real. I ended up with 4k fake spectrograms and 80+k real spectrograms. I then undersampled the real ones so they were 1:1 before training as above.\n\nEDIT: Also, when using the segmentation method, I tried randomly cropping the spectrogram and the mask to synthesise a bit more data. Still didn't work.",
          "votes": 2
        },
        {
          "id": 746495,
          "postDate": "2020-02-15T05:12:33.237Z",
          "content": "<p>I think we should use some external dataset to fit the model and then train it on the competition dataset. as I don't see any other promising model structures for fake audio detection.</p>",
          "rawMarkdown": "I think we should use some external dataset to fit the model and then train it on the competition dataset. as I don't see any other promising model structures for fake audio detection."
        },
        {
          "id": 748301,
          "postDate": "2020-02-17T11:09:13.107Z",
          "content": "<p>Thanks for sharing your experience!!</p>",
          "rawMarkdown": "Thanks for sharing your experience!!",
          "isDeleted": true
        },
        {
          "id": 759815,
          "postDate": "2020-02-29T13:46:58.990Z",
          "content": "<p>I tried similar approaches with similar results. One key element in this : better results when using extreme power curves : like pow(x, 0.1) when drawing graphical FFT results, or even clipping results. The answer often seems to lie in the background noise, not the \"strong\" signals... ? anyone ?</p>",
          "rawMarkdown": "I tried similar approaches with similar results. One key element in this : better results when using extreme power curves : like pow(x, 0.1) when drawing graphical FFT results, or even clipping results. The answer often seems to lie in the background noise, not the \"strong\" signals... ? anyone ?"
        }
      ]
    },
    {
      "id": 759596,
      "postDate": "2020-02-29T08:53:08.137Z",
      "content": "<p>I simply use the following extract to wav files from video (couldn't find a clea way to do it  in memory). Keep in mind that ffmpeg needs to be a dataset because internet is not allowed.\n<code>\nfor o in tqdm(train_video_files):\n    command = f\"ffmpeg -i {o} -vn audio/{o.stem}.wav\"\n    subprocess.call(command, shell=True)\n</code></p>",
      "rawMarkdown": "I simply use the following extract to wav files from video (couldn't find a clea way to do it  in memory). Keep in mind that ffmpeg needs to be a dataset because internet is not allowed.\n```\nfor o in tqdm(train_video_files):\n    command = f\"ffmpeg -i {o} -vn audio/{o.stem}.wav\"\n    subprocess.call(command, shell=True)\n```"
    },
    {
      "id": 748037,
      "postDate": "2020-02-17T06:11:25.473Z",
      "content": "<p>hey,guys. I wanna ask about how you extract .wav file from .mp4 file in kaggle notebook and run successfully after submitting? Thanks!</p>",
      "rawMarkdown": "hey,guys. I wanna ask about how you extract .wav file from .mp4 file in kaggle notebook and run successfully after submitting? Thanks!"
    }
  ],
  "comments": [
    {
      "id": 746098,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2020-02-14T16:07:29.857000",
      "content": "<p>This article might help. <a href=\"https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35\">https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35</a></p>\n\n<p>The general idea is to convert audio to spectrograms and train an image classifier on the audio spectrograms. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 746115,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-02-14T16:37:32.523000",
          "content": "<p>So I tried this, even tried using their code to extract the spectrograms, of various resolutions.</p>\n\n<p>I tried <em>so</em> many methods then to classify the spectrograms, including:\n- Classification using 2D CNNs \n- Classification using 1D CNNs (like they used)\n- Segmentation of the abnormal part of the spectrograms using a 2D HRNet\n- Segmentation of the abnormal part of the spectrograms using a UNet with 1D convolutions</p>\n\n<p>I tried doing the above with monochrome pictures, i.e. 1 input channel, and also the coloured spectrograms their code provides.</p>\n\n<p>All of them were incredibly underwhelming with ~65% accuracy at best. Given only a small proportion of videos in the training set even had altered audio, it just didn't seem worth pursuing further. However, keen to heard if anyone else has found otherwise?</p>",
          "votes": 18,
          "replies": []
        },
        {
          "id": 746131,
          "author_name": "student",
          "author_url": "",
          "post_date": "2020-02-14T16:57:54.563000",
          "content": "<p>Thank you so much. Will pursue it and report back if it helps. Thanks again</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 746171,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2020-02-14T17:46:46.197000",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> thanks for sharing your experience with this. I had tried combining audio spectrum images with frames from the video and running them through a 3dCNN but the model loss plateaued at ~0.69 so essentially it was just predicting random. Interesting to see you can get to 0.3 without accounting for audio. Gives me a better idea where to focus my efforts and what I can expect without using the audio. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 746473,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-02-15T04:00:48.517000",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> what was your training data. All the training parts??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 746479,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-02-15T04:21:21.057000",
          "content": "<p>I went through every fake video and diff'ed its spectrogram versus that of the original video; any differnce and I called it fake audio, otherwise it was real. I ended up with 4k fake spectrograms and 80+k real spectrograms. I then undersampled the real ones so they were 1:1 before training as above.</p>\n\n<p>EDIT: Also, when using the segmentation method, I tried randomly cropping the spectrogram and the mask to synthesise a bit more data. Still didn't work.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 746495,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-02-15T05:12:33.237000",
          "content": "<p>I think we should use some external dataset to fit the model and then train it on the competition dataset. as I don't see any other promising model structures for fake audio detection.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 748301,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-17T11:09:13.107000",
          "content": "<p>Thanks for sharing your experience!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759815,
          "author_name": "Simon Caby",
          "author_url": "",
          "post_date": "2020-02-29T13:46:58.990000",
          "content": "<p>I tried similar approaches with similar results. One key element in this : better results when using extreme power curves : like pow(x, 0.1) when drawing graphical FFT results, or even clipping results. The answer often seems to lie in the background noise, not the \"strong\" signals... ? anyone ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 759596,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-02-29T08:53:08.137000",
      "content": "<p>I simply use the following extract to wav files from video (couldn't find a clea way to do it  in memory). Keep in mind that ffmpeg needs to be a dataset because internet is not allowed.\n<code>\nfor o in tqdm(train_video_files):\n    command = f\"ffmpeg -i {o} -vn audio/{o.stem}.wav\"\n    subprocess.call(command, shell=True)\n</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 748037,
      "author_name": "Emmettj",
      "author_url": "",
      "post_date": "2020-02-17T06:11:25.473000",
      "content": "<p>hey,guys. I wanna ask about how you extract .wav file from .mp4 file in kaggle notebook and run successfully after submitting? Thanks!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "746045": "Folks: \n\nThere is a lot of literature on Fake video; looking for some broad guidance on fake audio detection. I would really appreciate any pointers on literature or even cryptic clues:) Thank you so much.",
    "746098": "This article might help. https://medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35\n\nThe general idea is to convert audio to spectrograms and train an image classifier on the audio spectrograms. ",
    "759596": "I simply use the following extract to wav files from video (couldn't find a clea way to do it  in memory). Keep in mind that ffmpeg needs to be a dataset because internet is not allowed.\n```\nfor o in tqdm(train_video_files):\n    command = f\"ffmpeg -i {o} -vn audio/{o.stem}.wav\"\n    subprocess.call(command, shell=True)\n```",
    "748037": "hey,guys. I wanna ask about how you extract .wav file from .mp4 file in kaggle notebook and run successfully after submitting? Thanks!"
  }
}