{
  "id": 164230,
  "title": "Why is MFCC still so iconic for audio embedding ?",
  "url": "/competitions/birdsong-recognition/discussion/164230",
  "author_name": "",
  "post_date": "2020-07-05T09:28:12.956528900Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>As fas as I can say, all the main public kernels are using MFCC like features. I mean, they're using classical audio feature extraction methods. This is a bit surprising when we know that best embedding features for various tasks (NLP, Computer Vision) are from neural networks (Bert, ResNet, ...). <br>\n<strong>I'm deeply sure that neural embedding could make the difference during this competition ...</strong></p>",
  "messages": [
    {
      "id": "916015",
      "postDate": "07/05/2020 09:28:12",
      "content": "<p>As fas as I can say, all the main public kernels are using MFCC like features. I mean, they're using classical audio feature extraction methods. This is a bit surprising when we know that best embedding features for various tasks (NLP, Computer Vision) are from neural networks (Bert, ResNet, ...). <br>\n<strong>I'm deeply sure that neural embedding could make the difference during this competition ...</strong></p>",
      "rawMarkdown": "As fas as I can say, all the main public kernels are using MFCC like features. I mean, they're using classical audio feature extraction methods. This is a bit surprising when we know that best embedding features for various tasks (NLP, Computer Vision) are from neural networks (Bert, ResNet, ...).   \n**I'm deeply sure that neural embedding could make the difference during this competition ...**",
      "votes": null
    },
    {
      "id": "916019",
      "postDate": "07/05/2020 09:31:15",
      "content": "<p>I found this paper which seems to be interresting [wav2vec: Unsupervised Pre-training for Speech Recognition] (<a href=\"https://arxiv.org/abs/1904.05862\">https://arxiv.org/abs/1904.05862</a>) . Quoting their abstract:  </p>\n\n<blockquote>\n  <p>We explore unsupervised pre-training for speech recognition by learning representations of raw audio</p>\n</blockquote>",
      "rawMarkdown": "I found this paper which seems to be interresting [wav2vec: Unsupervised Pre-training for Speech Recognition] (https://arxiv.org/abs/1904.05862) . Quoting their abstract:  \n&gt;We explore unsupervised pre-training for speech recognition by learning representations of raw audio",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 916019,
      "author_name": "kneroma",
      "author_url": "",
      "post_date": "07/05/2020 09:31:15",
      "content": "<p>I found this paper which seems to be interresting [wav2vec: Unsupervised Pre-training for Speech Recognition] (<a href=\"https://arxiv.org/abs/1904.05862\">https://arxiv.org/abs/1904.05862</a>) . Quoting their abstract:  </p>\n\n<blockquote>\n  <p>We explore unsupervised pre-training for speech recognition by learning representations of raw audio</p>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "916015": "As fas as I can say, all the main public kernels are using MFCC like features. I mean, they're using classical audio feature extraction methods. This is a bit surprising when we know that best embedding features for various tasks (NLP, Computer Vision) are from neural networks (Bert, ResNet, ...).   \n**I'm deeply sure that neural embedding could make the difference during this competition ...**",
    "916019": "I found this paper which seems to be interresting [wav2vec: Unsupervised Pre-training for Speech Recognition] (https://arxiv.org/abs/1904.05862) . Quoting their abstract:  \n&gt;We explore unsupervised pre-training for speech recognition by learning representations of raw audio"
  },
  "source": "meta"
}