{
  "id": 310749,
  "title": "RNNs for audio - line of research status?",
  "url": "/competitions/birdclef-2022/discussion/310749",
  "author_name": "",
  "post_date": "2022-03-02T20:23:23.785311900Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello, community!</p>\n<p>I'm trying to get the big picture of Audio Processing. Some quick introductions mentioned RNN over time-varying features. Even though I could find some resources along those lines, competition solutions seem to rely much more on \"the (MEL) Spectrogram pivot\" to the Computer Vision world.</p>\n<p>The question I have is for someone less rookie than me and it's the following one:</p>\n<p><strong>What is the status of the RNN line of research with regards to audio processing?</strong> </p>\n<p>Did it just fall behind and the SOTA is CNN-only?  Are there any scenarios where it is relevant?</p>\n<p>Any information is much appreciated. I have been digging around but it wasn't as fast as with the Spectrogram + CNN to find useful resources.</p>\n<p>Thanks!</p>\n<p>dataista0</p>",
  "messages": [
    {
      "id": "1710225",
      "postDate": "03/02/2022 20:23:23",
      "content": "<p>Hello, community!</p>\n<p>I'm trying to get the big picture of Audio Processing. Some quick introductions mentioned RNN over time-varying features. Even though I could find some resources along those lines, competition solutions seem to rely much more on \"the (MEL) Spectrogram pivot\" to the Computer Vision world.</p>\n<p>The question I have is for someone less rookie than me and it's the following one:</p>\n<p><strong>What is the status of the RNN line of research with regards to audio processing?</strong> </p>\n<p>Did it just fall behind and the SOTA is CNN-only?  Are there any scenarios where it is relevant?</p>\n<p>Any information is much appreciated. I have been digging around but it wasn't as fast as with the Spectrogram + CNN to find useful resources.</p>\n<p>Thanks!</p>\n<p>dataista0</p>",
      "rawMarkdown": "Hello, community!\n\nI'm trying to get the big picture of Audio Processing. Some quick introductions mentioned RNN over time-varying features. Even though I could find some resources along those lines, competition solutions seem to rely much more on \"the (MEL) Spectrogram pivot\" to the Computer Vision world.\n\nThe question I have is for someone less rookie than me and it's the following one:\n\n **What is the status of the RNN line of research with regards to audio processing?** \n\nDid it just fall behind and the SOTA is CNN-only?  Are there any scenarios where it is relevant?\n\nAny information is much appreciated. I have been digging around but it wasn't as fast as with the Spectrogram + CNN to find useful resources.\n\nThanks!\n\ndataista0",
      "votes": null
    },
    {
      "id": "1710235",
      "postDate": "03/02/2022 20:34:28",
      "content": "<p>I don't know about current SOTA in audio (probably transformers ;)) but you can also combine both CNNs and recurrent cells. I was suprised to read so much about \"visual solutions\".</p>",
      "rawMarkdown": "I don't know about current SOTA in audio (probably transformers ;)) but you can also combine both CNNs and recurrent cells. I was suprised to read so much about \"visual solutions\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1710235,
      "author_name": "m02ph3u5",
      "author_url": "",
      "post_date": "03/02/2022 20:34:28",
      "content": "<p>I don't know about current SOTA in audio (probably transformers ;)) but you can also combine both CNNs and recurrent cells. I was suprised to read so much about \"visual solutions\".</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1710225": "Hello, community!\n\nI'm trying to get the big picture of Audio Processing. Some quick introductions mentioned RNN over time-varying features. Even though I could find some resources along those lines, competition solutions seem to rely much more on \"the (MEL) Spectrogram pivot\" to the Computer Vision world.\n\nThe question I have is for someone less rookie than me and it's the following one:\n\n **What is the status of the RNN line of research with regards to audio processing?** \n\nDid it just fall behind and the SOTA is CNN-only?  Are there any scenarios where it is relevant?\n\nAny information is much appreciated. I have been digging around but it wasn't as fast as with the Spectrogram + CNN to find useful resources.\n\nThanks!\n\ndataista0",
    "1710235": "I don't know about current SOTA in audio (probably transformers ;)) but you can also combine both CNNs and recurrent cells. I was suprised to read so much about \"visual solutions\"."
  },
  "source": "meta"
}