{
  "id": 145977,
  "title": "Audio features",
  "url": "/competitions/deepfake-detection-challenge/discussion/145977",
  "author_name": "",
  "post_date": "2020-04-25T10:50:45.960626100Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>What about audio features? Anybody used them? Did it help to improve quality?</p>",
  "messages": [
    {
      "id": "820334",
      "postDate": "04/25/2020 10:50:45",
      "content": "<p>What about audio features? Anybody used them? Did it help to improve quality?</p>",
      "rawMarkdown": "What about audio features? Anybody used them? Did it help to improve quality?",
      "votes": null
    },
    {
      "id": "820341",
      "postDate": "04/25/2020 10:57:08",
      "content": "<p>There's only one person I've seen who used them, <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145648#818785\">and he fell over 900 places from 1st place</a>. Poor guy.</p>",
      "rawMarkdown": "There's only one person I've seen who used them, [and he fell over 900 places from 1st place](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145648#818785). Poor guy.",
      "votes": null
    },
    {
      "id": "820368",
      "postDate": "04/25/2020 11:44:07",
      "content": "<p>Hmm. Looks like it was a bad idea</p>",
      "rawMarkdown": "Hmm. Looks like it was a bad idea",
      "votes": null
    },
    {
      "id": "822298",
      "postDate": "04/26/2020 20:09:44",
      "content": "<p>I suspect that audio severely hurt my score in the private set but I can’t be certain of it until they release the dataset.</p>\n\n<p>Audio analysis did provide marginal benefits in the provided training set and on public LB but there are severe limitations in using the given audio (such as no audio labels, lack of diversity in audio manipulations, and very severe inbalance – only a small percentage of videos have audio manipulated). I spent about 2-3 weeks finding a solution for audio analysis. My CV1 for my single audio models ranged from .64 to .66 on a balanced DFDC set (balanced by given labels and not audio pseudo-labels). I only submitted 1 of my early single audio models on the public LB and it got ~.67. The mentioned scores assume I predict a .5 for any video that is not reasonably confident that the audio is fake. For my final ensembles on the public LB the audio helped by about .008-.01 when my log loss was in the .19 to .23 range.</p>\n\n<p>There are a lot of things that could be improved in my audio model but I kept things very simple to give more compute time to the video analysis. Better audio sampling techniques, more powerful models, more careful augmentation, and ensembling should further boost audio scores but probably not by a huge margin.</p>\n\n<p>I’m curious how well other people’s audio models performed. </p>",
      "rawMarkdown": "I suspect that audio severely hurt my score in the private set but I can’t be certain of it until they release the dataset.\n\nAudio analysis did provide marginal benefits in the provided training set and on public LB but there are severe limitations in using the given audio (such as no audio labels, lack of diversity in audio manipulations, and very severe inbalance – only a small percentage of videos have audio manipulated). I spent about 2-3 weeks finding a solution for audio analysis. My CV1 for my single audio models ranged from .64 to .66 on a balanced DFDC set (balanced by given labels and not audio pseudo-labels). I only submitted 1 of my early single audio models on the public LB and it got ~.67. The mentioned scores assume I predict a .5 for any video that is not reasonably confident that the audio is fake. For my final ensembles on the public LB the audio helped by about .008-.01 when my log loss was in the .19 to .23 range.\n\nThere are a lot of things that could be improved in my audio model but I kept things very simple to give more compute time to the video analysis. Better audio sampling techniques, more powerful models, more careful augmentation, and ensembling should further boost audio scores but probably not by a huge margin.\n\nI’m curious how well other people’s audio models performed.",
      "votes": null
    },
    {
      "id": "822414",
      "postDate": "04/26/2020 22:13:42",
      "content": "<p>I suppose the public set includes exactly 27 fake audio samples, as mentioned in this <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128650\">leak</a>. Anyway, we did not use any leaks or metadata, and our audio models gave about 0.68 public score.</p>",
      "rawMarkdown": "I suppose the public set includes exactly 27 fake audio samples, as mentioned in this [leak](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128650). Anyway, we did not use any leaks or metadata, and our audio models gave about 0.68 public score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 820341,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "04/25/2020 10:57:08",
      "content": "<p>There's only one person I've seen who used them, <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145648#818785\">and he fell over 900 places from 1st place</a>. Poor guy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 820368,
          "author_name": "loopdigga",
          "author_url": "",
          "post_date": "04/25/2020 11:44:07",
          "content": "<p>Hmm. Looks like it was a bad idea</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 822298,
          "author_name": "davidmilam",
          "author_url": "",
          "post_date": "04/26/2020 20:09:44",
          "content": "<p>I suspect that audio severely hurt my score in the private set but I can’t be certain of it until they release the dataset.</p>\n\n<p>Audio analysis did provide marginal benefits in the provided training set and on public LB but there are severe limitations in using the given audio (such as no audio labels, lack of diversity in audio manipulations, and very severe inbalance – only a small percentage of videos have audio manipulated). I spent about 2-3 weeks finding a solution for audio analysis. My CV1 for my single audio models ranged from .64 to .66 on a balanced DFDC set (balanced by given labels and not audio pseudo-labels). I only submitted 1 of my early single audio models on the public LB and it got ~.67. The mentioned scores assume I predict a .5 for any video that is not reasonably confident that the audio is fake. For my final ensembles on the public LB the audio helped by about .008-.01 when my log loss was in the .19 to .23 range.</p>\n\n<p>There are a lot of things that could be improved in my audio model but I kept things very simple to give more compute time to the video analysis. Better audio sampling techniques, more powerful models, more careful augmentation, and ensembling should further boost audio scores but probably not by a huge margin.</p>\n\n<p>I’m curious how well other people’s audio models performed. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 822414,
          "author_name": "sorokin",
          "author_url": "",
          "post_date": "04/26/2020 22:13:42",
          "content": "<p>I suppose the public set includes exactly 27 fake audio samples, as mentioned in this <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128650\">leak</a>. Anyway, we did not use any leaks or metadata, and our audio models gave about 0.68 public score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "820334": "What about audio features? Anybody used them? Did it help to improve quality?",
    "820341": "There's only one person I've seen who used them, [and he fell over 900 places from 1st place](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145648#818785). Poor guy.",
    "820368": "Hmm. Looks like it was a bad idea",
    "822298": "I suspect that audio severely hurt my score in the private set but I can’t be certain of it until they release the dataset.\n\nAudio analysis did provide marginal benefits in the provided training set and on public LB but there are severe limitations in using the given audio (such as no audio labels, lack of diversity in audio manipulations, and very severe inbalance – only a small percentage of videos have audio manipulated). I spent about 2-3 weeks finding a solution for audio analysis. My CV1 for my single audio models ranged from .64 to .66 on a balanced DFDC set (balanced by given labels and not audio pseudo-labels). I only submitted 1 of my early single audio models on the public LB and it got ~.67. The mentioned scores assume I predict a .5 for any video that is not reasonably confident that the audio is fake. For my final ensembles on the public LB the audio helped by about .008-.01 when my log loss was in the .19 to .23 range.\n\nThere are a lot of things that could be improved in my audio model but I kept things very simple to give more compute time to the video analysis. Better audio sampling techniques, more powerful models, more careful augmentation, and ensembling should further boost audio scores but probably not by a huge margin.\n\nI’m curious how well other people’s audio models performed.",
    "822414": "I suppose the public set includes exactly 27 fake audio samples, as mentioned in this [leak](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128650). Anyway, we did not use any leaks or metadata, and our audio models gave about 0.68 public score."
  },
  "source": "meta"
}