{
  "id": 159431,
  "title": "Bert Equivalent for Audio Classification ?",
  "url": "/competitions/birdsong-recognition/discussion/159431",
  "author_name": "",
  "post_date": "2020-06-17T12:41:38.481646300Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Completely new to Audio related tasks, but is there a go-to architecture that is used for audio classification, like the BERT family for text based tasks.</p>",
  "messages": [
    {
      "id": "890305",
      "postDate": "06/17/2020 12:41:38",
      "content": "<p>Completely new to Audio related tasks, but is there a go-to architecture that is used for audio classification, like the BERT family for text based tasks.</p>",
      "rawMarkdown": "Completely new to Audio related tasks, but is there a go-to architecture that is used for audio classification, like the BERT family for text based tasks.",
      "votes": null
    },
    {
      "id": "891479",
      "postDate": "06/18/2020 07:46:46",
      "content": "<p>Not that I am aware of. Right now, audio processing is mostly stuck with spectrograms and image classifiers (CNNs). Despite their success, there has been some criticism about this approach: <a href=\"https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\">https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd</a></p>\n\n<p>I saw researchers attempt audio event recognition using RNN, but almost all of those approaches performed worse than CNNs. We know that birds encode species identity in their vocalizations and spectrograms have been used by ornithologists for decades. Yet, I would love to see some approaches that think beyond the convolution :)</p>",
      "rawMarkdown": "Not that I am aware of. Right now, audio processing is mostly stuck with spectrograms and image classifiers (CNNs). Despite their success, there has been some criticism about this approach: https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\n\nI saw researchers attempt audio event recognition using RNN, but almost all of those approaches performed worse than CNNs. We know that birds encode species identity in their vocalizations and spectrograms have been used by ornithologists for decades. Yet, I would love to see some approaches that think beyond the convolution :)",
      "votes": null
    },
    {
      "id": "891500",
      "postDate": "06/18/2020 08:05:00",
      "content": "<p>This stuff shows some very promising results for speech recognition: <a href=\"https://arxiv.org/abs/1904.05862\">https://arxiv.org/abs/1904.05862</a></p>",
      "rawMarkdown": "This stuff shows some very promising results for speech recognition: https://arxiv.org/abs/1904.05862",
      "votes": null
    },
    {
      "id": "892103",
      "postDate": "06/18/2020 16:52:11",
      "content": "<p>A couple things to add here...</p>\n\n<ul>\n<li><p>TASNet denoisers have been having good results working with learned and gammatone filterbanks instead of the usual spectrograms. The ConvTasNet paper has some nice observations that the learned filterbanks for speech denoising have a similar frequency distribution to the Mel scale (used when making melspectrograms): this re-confirms that the Mel distribution nicely matches human speech, as it was designed for. Now, it turns out that bird vocalizations tend to overlap human frequency ranges, but it could be interesting to train some TASNet-like filterbanks for birdsong and see what the learned frequency distribution is like. For a multi-species classifier, I wouldn't expect a huge difference, but the distribution could shift if focused on (say) just owls or just songbirds.</p></li>\n<li><p>In some cases, people use 1d-convolutions over (mel)spectrograms using the filters as the 'depth' dimension, instead of treating the spectrogram as a single-channel image. This means there's no weight sharing across the filterbank, which does tend to look a bit different in the higher vs lower frequencies. (This usually doesn't lead to a huge change in model performance, but may be a nice thing to try.)</p></li>\n</ul>",
      "rawMarkdown": "A couple things to add here...\n\n* TASNet denoisers have been having good results working with learned and gammatone filterbanks instead of the usual spectrograms. The ConvTasNet paper has some nice observations that the learned filterbanks for speech denoising have a similar frequency distribution to the Mel scale (used when making melspectrograms): this re-confirms that the Mel distribution nicely matches human speech, as it was designed for. Now, it turns out that bird vocalizations tend to overlap human frequency ranges, but it could be interesting to train some TASNet-like filterbanks for birdsong and see what the learned frequency distribution is like. For a multi-species classifier, I wouldn't expect a huge difference, but the distribution could shift if focused on (say) just owls or just songbirds.\n\n* In some cases, people use 1d-convolutions over (mel)spectrograms using the filters as the 'depth' dimension, instead of treating the spectrogram as a single-channel image. This means there's no weight sharing across the filterbank, which does tend to look a bit different in the higher vs lower frequencies. (This usually doesn't lead to a huge change in model performance, but may be a nice thing to try.)",
      "votes": null
    },
    {
      "id": "967185",
      "postDate": "08/12/2020 04:21:07",
      "content": "<p>I would recommend WaveNet. It might need a fair bit of tuning, but it's efficient and much of the infrastructure is already there from previous Kaggle kernels</p>",
      "rawMarkdown": "I would recommend WaveNet. It might need a fair bit of tuning, but it's efficient and much of the infrastructure is already there from previous Kaggle kernels",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 967185,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "08/12/2020 04:21:07",
      "content": "<p>I would recommend WaveNet. It might need a fair bit of tuning, but it's efficient and much of the infrastructure is already there from previous Kaggle kernels</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 891479,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "06/18/2020 07:46:46",
      "content": "<p>Not that I am aware of. Right now, audio processing is mostly stuck with spectrograms and image classifiers (CNNs). Despite their success, there has been some criticism about this approach: <a href=\"https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\">https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd</a></p>\n\n<p>I saw researchers attempt audio event recognition using RNN, but almost all of those approaches performed worse than CNNs. We know that birds encode species identity in their vocalizations and spectrograms have been used by ornithologists for decades. Yet, I would love to see some approaches that think beyond the convolution :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 892103,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "06/18/2020 16:52:11",
          "content": "<p>A couple things to add here...</p>\n\n<ul>\n<li><p>TASNet denoisers have been having good results working with learned and gammatone filterbanks instead of the usual spectrograms. The ConvTasNet paper has some nice observations that the learned filterbanks for speech denoising have a similar frequency distribution to the Mel scale (used when making melspectrograms): this re-confirms that the Mel distribution nicely matches human speech, as it was designed for. Now, it turns out that bird vocalizations tend to overlap human frequency ranges, but it could be interesting to train some TASNet-like filterbanks for birdsong and see what the learned frequency distribution is like. For a multi-species classifier, I wouldn't expect a huge difference, but the distribution could shift if focused on (say) just owls or just songbirds.</p></li>\n<li><p>In some cases, people use 1d-convolutions over (mel)spectrograms using the filters as the 'depth' dimension, instead of treating the spectrogram as a single-channel image. This means there's no weight sharing across the filterbank, which does tend to look a bit different in the higher vs lower frequencies. (This usually doesn't lead to a huge change in model performance, but may be a nice thing to try.)</p></li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 891500,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "06/18/2020 08:05:00",
      "content": "<p>This stuff shows some very promising results for speech recognition: <a href=\"https://arxiv.org/abs/1904.05862\">https://arxiv.org/abs/1904.05862</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "890305": "Completely new to Audio related tasks, but is there a go-to architecture that is used for audio classification, like the BERT family for text based tasks.",
    "891479": "Not that I am aware of. Right now, audio processing is mostly stuck with spectrograms and image classifiers (CNNs). Despite their success, there has been some criticism about this approach: https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\n\nI saw researchers attempt audio event recognition using RNN, but almost all of those approaches performed worse than CNNs. We know that birds encode species identity in their vocalizations and spectrograms have been used by ornithologists for decades. Yet, I would love to see some approaches that think beyond the convolution :)",
    "891500": "This stuff shows some very promising results for speech recognition: https://arxiv.org/abs/1904.05862",
    "892103": "A couple things to add here...\n\n* TASNet denoisers have been having good results working with learned and gammatone filterbanks instead of the usual spectrograms. The ConvTasNet paper has some nice observations that the learned filterbanks for speech denoising have a similar frequency distribution to the Mel scale (used when making melspectrograms): this re-confirms that the Mel distribution nicely matches human speech, as it was designed for. Now, it turns out that bird vocalizations tend to overlap human frequency ranges, but it could be interesting to train some TASNet-like filterbanks for birdsong and see what the learned frequency distribution is like. For a multi-species classifier, I wouldn't expect a huge difference, but the distribution could shift if focused on (say) just owls or just songbirds.\n\n* In some cases, people use 1d-convolutions over (mel)spectrograms using the filters as the 'depth' dimension, instead of treating the spectrogram as a single-channel image. This means there's no weight sharing across the filterbank, which does tend to look a bit different in the higher vs lower frequencies. (This usually doesn't lead to a huge change in model performance, but may be a nice thing to try.)",
    "967185": "I would recommend WaveNet. It might need a fair bit of tuning, but it's efficient and much of the infrastructure is already there from previous Kaggle kernels"
  },
  "source": "meta"
}