{
  "id": 199619,
  "title": "🐤🦆 🦅🐸 Extensive Resource Compilation for Audio Data ",
  "url": "/competitions/rfcx-species-audio-detection/discussion/199619",
  "author_name": "Tensor Girl",
  "post_date": "2020-11-26T13:30:36.357000",
  "votes": 84,
  "comment_count": 10,
  "views": 0,
  "content": "<p><img src=\"https://drive.google.com/uc?id=1FSBysbTXREKgKyLkVeqIuBvwB8nlOhrK\" alt=\"\"></p>\n<p>Special Thanks to <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> for sharing a great youtube tutorials by Valerio Velardo to learn about <strong>Audio Signal Processing for Machine Learning</strong> . Great resource for beginners to get started.</p>\n<p><a href=\"https://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0\" target=\"_blank\">https://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0</a></p>\n<p>Thanks to <a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a> for sharing great one stop resources for beginners </p>\n<p><strong>First-stop for audio analysis :</strong> <br>\n<a href=\"https://www.audiocontentanalysis.org/teaching/\" target=\"_blank\">https://www.audiocontentanalysis.org/teaching/</a><br>\nThe videos are very well explained and it has code implementations as well. </p>\n<p><strong>Mel Frequency Cepstral Coefficients Detailed Explanation:</strong> <br>\n<a href=\"https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\" target=\"_blank\">https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd</a></p>\n<p>I came across this pdf file from TU Berlin which has extensive resource compilation for audio data </p>\n<p>Link : <a href=\"https://www.ak.tu-berlin.de/fileadmin/a0135/downloads/resources_aed4dl.pdf\" target=\"_blank\">https://www.ak.tu-berlin.de/fileadmin/a0135/downloads/resources_aed4dl.pdf</a></p>\n<p>Some of the highlights from the pdf are below </p>\n<p><strong>Datasets:</strong></p>\n<p>• <a href=\"http://www.cs.tut.fi/~heittolt/datasets\" target=\"_blank\">http://www.cs.tut.fi/~heittolt/datasets</a> [Collection]<br>\n• <a href=\"https://github.com/ybayle/awesome-deep-learning-music/blob/master/datasets.md\" target=\"_blank\">https://github.com/ybayle/awesome-deep-learning-music/blob/master/datasets.md</a><br>\n• <a href=\"https://annotator.freesound.org/fsd/downloads/\" target=\"_blank\">https://annotator.freesound.org/fsd/downloads/</a> [FSD]<br>\n• <a href=\"https://zenodo.org/record/3384388#.XaXsWeczbOQ\" target=\"_blank\">https://zenodo.org/record/3384388#.XaXsWeczbOQ</a> [MIMII Dataset]</p>\n<p><strong>Most often used in research:</strong></p>\n<p>• <a href=\"http://dcase.community/challenge2018/task-acoustic-scene-classification#audio-dataset\" target=\"_blank\">http://dcase.community/challenge2018/task-acoustic-scene-classification#audio-dataset</a><br>\n• <a href=\"https://urbansounddataset.weebly.com/urbansound8k.html\" target=\"_blank\">https://urbansounddataset.weebly.com/urbansound8k.html</a><br>\n• <a href=\"https://github.com/karoldvl/ESC-50\" target=\"_blank\">https://github.com/karoldvl/ESC-50</a><br>\n• <a href=\"https://github.com/karoldvl/ESC-10\" target=\"_blank\">https://github.com/karoldvl/ESC-10</a></p>\n<p><strong>Data Augmentation in time-domain:</strong></p>\n<p>• Salamon, J. and Bello, J.P., 2017. Deep convolutional neural networks and data augmentation for<br>\nenvironmental sound classification. IEEE Signal Processing Letters, 24(3), pp.279-283.</p>\n<p><strong>Data Augmentation in frequency-domain:</strong></p>\n<p>• Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D. and Le, Q.V., 2019. Specaugment:<br>\nA simple data augmentation method for automatic speech recognition.</p>\n<p>• Zhang, Z., Xu, S., Cao, S. and Zhang, S., 2018, November. Deep convolutional neural network with<br>\nmixup for environmental sound classification. In Chinese Conference on Pattern Recognition and<br>\nComputer Vision (PRCV) (pp. 356-367). Springer, Cham.</p>\n<p><strong>Data Augmentation with Adverserial Networks:</strong></p>\n<p>• Donahue, C., McAuley, J. and Puckette, M., 2018. Adversarial audio synthesis.<br>\n[<a href=\"https://github.com/chrisdonahue/wavegan\" target=\"_blank\">https://github.com/chrisdonahue/wavegan</a>]<br>\n[<a href=\"https://chrisdonahue.com/wavegan_examples/\" target=\"_blank\">https://chrisdonahue.com/wavegan_examples/</a>]<br>\n[<a href=\"https://www.youtube.com/watch?v=BA-Z0KJIyJs\" target=\"_blank\">https://www.youtube.com/watch?v=BA-Z0KJIyJs</a>]</p>\n<p>• Engel, J., Agrawal, K.K., Chen, S., Gulrajani, I., Donahue, C. and Roberts, A., 2019. Gansynth:<br>\nAdversarial neural audio synthesis.</p>\n<p>• Unknown Authors, DDSP: Differentiable Digital Signal Processing<br>\n[<a href=\"https://openreview.net/forum?id=B1x1ma4tDr\" target=\"_blank\">https://openreview.net/forum?id=B1x1ma4tDr</a>]</p>\n<p>• Yamamoto, R., Song, E. and Kim, J.M., 2019. Parallel WaveGAN: A fast waveform generation<br>\nmodel based on generative adversarial networks with multi-resolution spectrogram. arXiv preprint<br>\narXiv:1910.11480. [<a href=\"https://r9y9.github.io/demos/projects/icassp2020/\" target=\"_blank\">https://r9y9.github.io/demos/projects/icassp2020/</a>]</p>\n<p><strong>End-to-end learning on audio:</strong></p>\n<p>• Dieleman, S. and Schrauwen, B., 2014, May. End-to-end learning for music audio. In 2014 IEEE<br>\nInternational Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6964-6968). IEEE.</p>\n<p>• Pons, J., Nieto, O., Prockup, M., Schmidt, E., Ehmann, A. and Serra, X., 2017. End-to-end learning for music audio tagging at scale.</p>\n<p>• Abdoli, S., Cardinal, P. and Koerich, A.L., 2019. End-to-End Environmental Sound Classification using a 1D Convolutional Neural Network. Expert Systems with Applications.</p>\n<p><strong>Lectures/Presentations:</strong></p>\n<p>• <a href=\"https://www.youtube.com/watch?v=7B1WBa3sC3I\" target=\"_blank\">https://www.youtube.com/watch?v=7B1WBa3sC3I</a><br>\n[Deep Neural Networks for Sound Event Detection]</p>\n<p>• <a href=\"https://www.youtube.com/watch?v=zvccOFz2KxI\" target=\"_blank\">https://www.youtube.com/watch?v=zvccOFz2KxI</a><br>\n[Robust Sound Event Detection in Acoustic Sensor Networks]</p>\n<p>• <a href=\"https://www.youtube.com/watch?v=9X66iwEQSyI\" target=\"_blank\">https://www.youtube.com/watch?v=9X66iwEQSyI</a><br>\n[Audio Event Detection w/Deep Learning at at Stanley Black &amp; Decker]</p>\n<p>• <a href=\"https://www.youtube.com/watch?v=IzRCC2-IBwU\" target=\"_blank\">https://www.youtube.com/watch?v=IzRCC2-IBwU</a><br>\n[Machine Listening of Everyday Soundscapes]</p>\n<p>Hope you like the compilation from TU Berlin . If you come across resources like this , please share it in comments and I can include it in the post .</p>",
  "messages": [
    {
      "id": 1092017,
      "postDate": "2020-11-26T13:30:36.357Z",
      "content": "<p><img src=\"https://drive.google.com/uc?id=1FSBysbTXREKgKyLkVeqIuBvwB8nlOhrK\" alt=\"\"></p>\n<p>Special Thanks to <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> for sharing a great youtube tutorials by Valerio Velardo to learn about <strong>Audio Signal Processing for Machine Learning</strong> . Great resource for beginners to get started.</p>\n<p><a href=\"https://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0\" target=\"_blank\">https://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0</a></p>\n<p>Thanks to <a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a> for sharing great one stop resources for beginners </p>\n<p><strong>First-stop for audio analysis :</strong> <br>\n<a href=\"https://www.audiocontentanalysis.org/teaching/\" target=\"_blank\">https://www.audiocontentanalysis.org/teaching/</a><br>\nThe videos are very well explained and it has code implementations as well. </p>\n<p><strong>Mel Frequency Cepstral Coefficients Detailed Explanation:</strong> <br>\n<a href=\"https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\" target=\"_blank\">https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd</a></p>\n<p>I came across this pdf file from TU Berlin which has extensive resource compilation for audio data </p>\n<p>Link : <a href=\"https://www.ak.tu-berlin.de/fileadmin/a0135/downloads/resources_aed4dl.pdf\" target=\"_blank\">https://www.ak.tu-berlin.de/fileadmin/a0135/downloads/resources_aed4dl.pdf</a></p>\n<p>Some of the highlights from the pdf are below </p>\n<p><strong>Datasets:</strong></p>\n<p>• <a href=\"http://www.cs.tut.fi/~heittolt/datasets\" target=\"_blank\">http://www.cs.tut.fi/~heittolt/datasets</a> [Collection]<br>\n• <a href=\"https://github.com/ybayle/awesome-deep-learning-music/blob/master/datasets.md\" target=\"_blank\">https://github.com/ybayle/awesome-deep-learning-music/blob/master/datasets.md</a><br>\n• <a href=\"https://annotator.freesound.org/fsd/downloads/\" target=\"_blank\">https://annotator.freesound.org/fsd/downloads/</a> [FSD]<br>\n• <a href=\"https://zenodo.org/record/3384388#.XaXsWeczbOQ\" target=\"_blank\">https://zenodo.org/record/3384388#.XaXsWeczbOQ</a> [MIMII Dataset]</p>\n<p><strong>Most often used in research:</strong></p>\n<p>• <a href=\"http://dcase.community/challenge2018/task-acoustic-scene-classification#audio-dataset\" target=\"_blank\">http://dcase.community/challenge2018/task-acoustic-scene-classification#audio-dataset</a><br>\n• <a href=\"https://urbansounddataset.weebly.com/urbansound8k.html\" target=\"_blank\">https://urbansounddataset.weebly.com/urbansound8k.html</a><br>\n• <a href=\"https://github.com/karoldvl/ESC-50\" target=\"_blank\">https://github.com/karoldvl/ESC-50</a><br>\n• <a href=\"https://github.com/karoldvl/ESC-10\" target=\"_blank\">https://github.com/karoldvl/ESC-10</a></p>\n<p><strong>Data Augmentation in time-domain:</strong></p>\n<p>• Salamon, J. and Bello, J.P., 2017. Deep convolutional neural networks and data augmentation for<br>\nenvironmental sound classification. IEEE Signal Processing Letters, 24(3), pp.279-283.</p>\n<p><strong>Data Augmentation in frequency-domain:</strong></p>\n<p>• Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D. and Le, Q.V., 2019. Specaugment:<br>\nA simple data augmentation method for automatic speech recognition.</p>\n<p>• Zhang, Z., Xu, S., Cao, S. and Zhang, S., 2018, November. Deep convolutional neural network with<br>\nmixup for environmental sound classification. In Chinese Conference on Pattern Recognition and<br>\nComputer Vision (PRCV) (pp. 356-367). Springer, Cham.</p>\n<p><strong>Data Augmentation with Adverserial Networks:</strong></p>\n<p>• Donahue, C., McAuley, J. and Puckette, M., 2018. Adversarial audio synthesis.<br>\n[<a href=\"https://github.com/chrisdonahue/wavegan\" target=\"_blank\">https://github.com/chrisdonahue/wavegan</a>]<br>\n[<a href=\"https://chrisdonahue.com/wavegan_examples/\" target=\"_blank\">https://chrisdonahue.com/wavegan_examples/</a>]<br>\n[<a href=\"https://www.youtube.com/watch?v=BA-Z0KJIyJs\" target=\"_blank\">https://www.youtube.com/watch?v=BA-Z0KJIyJs</a>]</p>\n<p>• Engel, J., Agrawal, K.K., Chen, S., Gulrajani, I., Donahue, C. and Roberts, A., 2019. Gansynth:<br>\nAdversarial neural audio synthesis.</p>\n<p>• Unknown Authors, DDSP: Differentiable Digital Signal Processing<br>\n[<a href=\"https://openreview.net/forum?id=B1x1ma4tDr\" target=\"_blank\">https://openreview.net/forum?id=B1x1ma4tDr</a>]</p>\n<p>• Yamamoto, R., Song, E. and Kim, J.M., 2019. Parallel WaveGAN: A fast waveform generation<br>\nmodel based on generative adversarial networks with multi-resolution spectrogram. arXiv preprint<br>\narXiv:1910.11480. [<a href=\"https://r9y9.github.io/demos/projects/icassp2020/\" target=\"_blank\">https://r9y9.github.io/demos/projects/icassp2020/</a>]</p>\n<p><strong>End-to-end learning on audio:</strong></p>\n<p>• Dieleman, S. and Schrauwen, B., 2014, May. End-to-end learning for music audio. In 2014 IEEE<br>\nInternational Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6964-6968). IEEE.</p>\n<p>• Pons, J., Nieto, O., Prockup, M., Schmidt, E., Ehmann, A. and Serra, X., 2017. End-to-end learning for music audio tagging at scale.</p>\n<p>• Abdoli, S., Cardinal, P. and Koerich, A.L., 2019. End-to-End Environmental Sound Classification using a 1D Convolutional Neural Network. Expert Systems with Applications.</p>\n<p><strong>Lectures/Presentations:</strong></p>\n<p>• <a href=\"https://www.youtube.com/watch?v=7B1WBa3sC3I\" target=\"_blank\">https://www.youtube.com/watch?v=7B1WBa3sC3I</a><br>\n[Deep Neural Networks for Sound Event Detection]</p>\n<p>• <a href=\"https://www.youtube.com/watch?v=zvccOFz2KxI\" target=\"_blank\">https://www.youtube.com/watch?v=zvccOFz2KxI</a><br>\n[Robust Sound Event Detection in Acoustic Sensor Networks]</p>\n<p>• <a href=\"https://www.youtube.com/watch?v=9X66iwEQSyI\" target=\"_blank\">https://www.youtube.com/watch?v=9X66iwEQSyI</a><br>\n[Audio Event Detection w/Deep Learning at at Stanley Black &amp; Decker]</p>\n<p>• <a href=\"https://www.youtube.com/watch?v=IzRCC2-IBwU\" target=\"_blank\">https://www.youtube.com/watch?v=IzRCC2-IBwU</a><br>\n[Machine Listening of Everyday Soundscapes]</p>\n<p>Hope you like the compilation from TU Berlin . If you come across resources like this , please share it in comments and I can include it in the post .</p>",
      "rawMarkdown": "![](https://drive.google.com/uc?id=1FSBysbTXREKgKyLkVeqIuBvwB8nlOhrK)\n\nSpecial Thanks to @slawekbiel for sharing a great youtube tutorials by Valerio Velardo to learn about **Audio Signal Processing for Machine Learning** . Great resource for beginners to get started.\n\nhttps://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0\n\nThanks to @mrutyunjaybiswal for sharing great one stop resources for beginners \n\n**First-stop for audio analysis :** \nhttps://www.audiocontentanalysis.org/teaching/\nThe videos are very well explained and it has code implementations as well. \n\n**Mel Frequency Cepstral Coefficients Detailed Explanation:** \nhttps://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\n\nI came across this pdf file from TU Berlin which has extensive resource compilation for audio data \n\nLink : https://www.ak.tu-berlin.de/fileadmin/a0135/downloads/resources_aed4dl.pdf\n\nSome of the highlights from the pdf are below \n\n\n**Datasets:**\n\n• http://www.cs.tut.fi/~heittolt/datasets [Collection]\n• https://github.com/ybayle/awesome-deep-learning-music/blob/master/datasets.md\n• https://annotator.freesound.org/fsd/downloads/ [FSD]\n• https://zenodo.org/record/3384388#.XaXsWeczbOQ [MIMII Dataset]\n\n**Most often used in research:**\n\n• http://dcase.community/challenge2018/task-acoustic-scene-classification#audio-dataset\n• https://urbansounddataset.weebly.com/urbansound8k.html\n• https://github.com/karoldvl/ESC-50\n• https://github.com/karoldvl/ESC-10\n\n**Data Augmentation in time-domain:**\n\n• Salamon, J. and Bello, J.P., 2017. Deep convolutional neural networks and data augmentation for\nenvironmental sound classification. IEEE Signal Processing Letters, 24(3), pp.279-283.\n\n\n**Data Augmentation in frequency-domain:**\n\n• Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D. and Le, Q.V., 2019. Specaugment:\nA simple data augmentation method for automatic speech recognition.\n\n• Zhang, Z., Xu, S., Cao, S. and Zhang, S., 2018, November. Deep convolutional neural network with\nmixup for environmental sound classification. In Chinese Conference on Pattern Recognition and\nComputer Vision (PRCV) (pp. 356-367). Springer, Cham.\n\n**Data Augmentation with Adverserial Networks:**\n\n• Donahue, C., McAuley, J. and Puckette, M., 2018. Adversarial audio synthesis.\n[https://github.com/chrisdonahue/wavegan]\n[https://chrisdonahue.com/wavegan_examples/]\n[https://www.youtube.com/watch?v=BA-Z0KJIyJs]\n\n• Engel, J., Agrawal, K.K., Chen, S., Gulrajani, I., Donahue, C. and Roberts, A., 2019. Gansynth:\nAdversarial neural audio synthesis.\n\n• Unknown Authors, DDSP: Differentiable Digital Signal Processing\n[https://openreview.net/forum?id=B1x1ma4tDr]\n\n• Yamamoto, R., Song, E. and Kim, J.M., 2019. Parallel WaveGAN: A fast waveform generation\nmodel based on generative adversarial networks with multi-resolution spectrogram. arXiv preprint\narXiv:1910.11480. [https://r9y9.github.io/demos/projects/icassp2020/]\n\n\n**End-to-end learning on audio:**\n\n• Dieleman, S. and Schrauwen, B., 2014, May. End-to-end learning for music audio. In 2014 IEEE\nInternational Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6964-6968). IEEE.\n\n• Pons, J., Nieto, O., Prockup, M., Schmidt, E., Ehmann, A. and Serra, X., 2017. End-to-end learning for music audio tagging at scale.\n\n• Abdoli, S., Cardinal, P. and Koerich, A.L., 2019. End-to-End Environmental Sound Classification using a 1D Convolutional Neural Network. Expert Systems with Applications.\n\n\n\n**Lectures/Presentations:**\n\n• https://www.youtube.com/watch?v=7B1WBa3sC3I\n[Deep Neural Networks for Sound Event Detection]\n\n• https://www.youtube.com/watch?v=zvccOFz2KxI\n[Robust Sound Event Detection in Acoustic Sensor Networks]\n\n• https://www.youtube.com/watch?v=9X66iwEQSyI\n[Audio Event Detection w/Deep Learning at at Stanley Black & Decker]\n\n• https://www.youtube.com/watch?v=IzRCC2-IBwU\n[Machine Listening of Everyday Soundscapes]\n\nHope you like the compilation from TU Berlin . If you come across resources like this , please share it in comments and I can include it in the post .",
      "votes": 84
    },
    {
      "id": 1092143,
      "postDate": "2020-11-26T15:11:45.733Z",
      "content": "<p>I stumbled upon this series <a href=\"https://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0\" target=\"_blank\">Audio Signal Processing for Machine Learning</a> by Valerio Velardo. I found it a terrific resource for a novice to get up to speed with audio.</p>",
      "rawMarkdown": "I stumbled upon this series [Audio Signal Processing for Machine Learning](https://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0) by Valerio Velardo. I found it a terrific resource for a novice to get up to speed with audio.",
      "votes": 5,
      "replies": [
        {
          "id": 1092160,
          "postDate": "2020-11-26T15:20:05.053Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1100751,
      "postDate": "2020-12-03T09:56:34.447Z",
      "content": "<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> As I have started this competition as my research project, I will be adding contents, articles, video explanations to this thread, hope It would help the community.</p>\n<ol>\n<li>Here is an awesome, well-organized, <strong>first-stop for audio analysis</strong> : <a href=\"https://www.audiocontentanalysis.org/teaching/\" target=\"_blank\">https://www.audiocontentanalysis.org/teaching/</a><br>\nThe videos are very well explained and it has code implementations as well. Do check out.</li>\n<li><strong>Mel Frequency Cepstral Coefficients</strong> Detailed Explanation: <a href=\"https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\" target=\"_blank\">https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd</a></li>\n</ol>",
      "rawMarkdown": "@usharengaraju As I have started this competition as my research project, I will be adding contents, articles, video explanations to this thread, hope It would help the community.\n\n1. Here is an awesome, well-organized, **first-stop for audio analysis** : https://www.audiocontentanalysis.org/teaching/\nThe videos are very well explained and it has code implementations as well. Do check out.\n2. **Mel Frequency Cepstral Coefficients** Detailed Explanation: https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd",
      "votes": 2,
      "replies": [
        {
          "id": 1105093,
          "postDate": "2020-12-07T14:12:36.870Z",
          "content": "<p><a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a>  Thank you so much ..I have updated the post with your links </p>",
          "rawMarkdown": "@mrutyunjaybiswal  Thank you so much ..I have updated the post with your links ",
          "votes": 1
        },
        {
          "id": 1105295,
          "postDate": "2020-12-07T18:27:17.597Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> ma'am, I will add more resources gradually.</p>",
          "rawMarkdown": "Thanks, @usharengaraju ma'am, I will add more resources gradually.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1183128,
      "postDate": "2021-02-02T19:06:49.540Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> , another important reference might be the Cornell birdcall competition which was held recently. </p>\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition</a></p>",
      "rawMarkdown": "Thanks for sharing @usharengaraju , another important reference might be the Cornell birdcall competition which was held recently. \n\nhttps://www.kaggle.com/c/birdsong-recognition"
    },
    {
      "id": 1137886,
      "postDate": "2021-01-04T08:54:45.823Z",
      "content": "<p>Thanks for this! ^_^</p>",
      "rawMarkdown": "Thanks for this! ^_^",
      "votes": 1
    },
    {
      "id": 1105024,
      "postDate": "2020-12-07T13:18:23.987Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1094780,
      "postDate": "2020-11-28T23:06:22.627Z",
      "content": "<p>Great info, thanks for sharing !!</p>",
      "rawMarkdown": "Great info, thanks for sharing !!",
      "votes": 1
    },
    {
      "id": 1092108,
      "postDate": "2020-11-26T14:51:03.377Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1092143,
      "author_name": "Slawek Biel",
      "author_url": "",
      "post_date": "2020-11-26T15:11:45.733000",
      "content": "<p>I stumbled upon this series <a href=\"https://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0\" target=\"_blank\">Audio Signal Processing for Machine Learning</a> by Valerio Velardo. I found it a terrific resource for a novice to get up to speed with audio.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1092160,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-26T15:20:05.053000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1100751,
      "author_name": "Ultron",
      "author_url": "",
      "post_date": "2020-12-03T09:56:34.447000",
      "content": "<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> As I have started this competition as my research project, I will be adding contents, articles, video explanations to this thread, hope It would help the community.</p>\n<ol>\n<li>Here is an awesome, well-organized, <strong>first-stop for audio analysis</strong> : <a href=\"https://www.audiocontentanalysis.org/teaching/\" target=\"_blank\">https://www.audiocontentanalysis.org/teaching/</a><br>\nThe videos are very well explained and it has code implementations as well. Do check out.</li>\n<li><strong>Mel Frequency Cepstral Coefficients</strong> Detailed Explanation: <a href=\"https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\" target=\"_blank\">https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd</a></li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 1105093,
          "author_name": "Tensor Girl",
          "author_url": "",
          "post_date": "2020-12-07T14:12:36.870000",
          "content": "<p><a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a>  Thank you so much ..I have updated the post with your links </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1105295,
          "author_name": "Ultron",
          "author_url": "",
          "post_date": "2020-12-07T18:27:17.597000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> ma'am, I will add more resources gradually.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1183128,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2021-02-02T19:06:49.540000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> , another important reference might be the Cornell birdcall competition which was held recently. </p>\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1137886,
      "author_name": "Aritra Roy Gosthipaty",
      "author_url": "",
      "post_date": "2021-01-04T08:54:45.823000",
      "content": "<p>Thanks for this! ^_^</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1105024,
      "author_name": "Siwei Luo",
      "author_url": "",
      "post_date": "2020-12-07T13:18:23.987000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1094780,
      "author_name": "Mohamed Hany",
      "author_url": "",
      "post_date": "2020-11-28T23:06:22.627000",
      "content": "<p>Great info, thanks for sharing !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1092108,
      "author_name": "Utkarsh Gupta",
      "author_url": "",
      "post_date": "2020-11-26T14:51:03.377000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1092017": "![](https://drive.google.com/uc?id=1FSBysbTXREKgKyLkVeqIuBvwB8nlOhrK)\n\nSpecial Thanks to @slawekbiel for sharing a great youtube tutorials by Valerio Velardo to learn about **Audio Signal Processing for Machine Learning** . Great resource for beginners to get started.\n\nhttps://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0\n\nThanks to @mrutyunjaybiswal for sharing great one stop resources for beginners \n\n**First-stop for audio analysis :** \nhttps://www.audiocontentanalysis.org/teaching/\nThe videos are very well explained and it has code implementations as well. \n\n**Mel Frequency Cepstral Coefficients Detailed Explanation:** \nhttps://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\n\nI came across this pdf file from TU Berlin which has extensive resource compilation for audio data \n\nLink : https://www.ak.tu-berlin.de/fileadmin/a0135/downloads/resources_aed4dl.pdf\n\nSome of the highlights from the pdf are below \n\n\n**Datasets:**\n\n• http://www.cs.tut.fi/~heittolt/datasets [Collection]\n• https://github.com/ybayle/awesome-deep-learning-music/blob/master/datasets.md\n• https://annotator.freesound.org/fsd/downloads/ [FSD]\n• https://zenodo.org/record/3384388#.XaXsWeczbOQ [MIMII Dataset]\n\n**Most often used in research:**\n\n• http://dcase.community/challenge2018/task-acoustic-scene-classification#audio-dataset\n• https://urbansounddataset.weebly.com/urbansound8k.html\n• https://github.com/karoldvl/ESC-50\n• https://github.com/karoldvl/ESC-10\n\n**Data Augmentation in time-domain:**\n\n• Salamon, J. and Bello, J.P., 2017. Deep convolutional neural networks and data augmentation for\nenvironmental sound classification. IEEE Signal Processing Letters, 24(3), pp.279-283.\n\n\n**Data Augmentation in frequency-domain:**\n\n• Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D. and Le, Q.V., 2019. Specaugment:\nA simple data augmentation method for automatic speech recognition.\n\n• Zhang, Z., Xu, S., Cao, S. and Zhang, S., 2018, November. Deep convolutional neural network with\nmixup for environmental sound classification. In Chinese Conference on Pattern Recognition and\nComputer Vision (PRCV) (pp. 356-367). Springer, Cham.\n\n**Data Augmentation with Adverserial Networks:**\n\n• Donahue, C., McAuley, J. and Puckette, M., 2018. Adversarial audio synthesis.\n[https://github.com/chrisdonahue/wavegan]\n[https://chrisdonahue.com/wavegan_examples/]\n[https://www.youtube.com/watch?v=BA-Z0KJIyJs]\n\n• Engel, J., Agrawal, K.K., Chen, S., Gulrajani, I., Donahue, C. and Roberts, A., 2019. Gansynth:\nAdversarial neural audio synthesis.\n\n• Unknown Authors, DDSP: Differentiable Digital Signal Processing\n[https://openreview.net/forum?id=B1x1ma4tDr]\n\n• Yamamoto, R., Song, E. and Kim, J.M., 2019. Parallel WaveGAN: A fast waveform generation\nmodel based on generative adversarial networks with multi-resolution spectrogram. arXiv preprint\narXiv:1910.11480. [https://r9y9.github.io/demos/projects/icassp2020/]\n\n\n**End-to-end learning on audio:**\n\n• Dieleman, S. and Schrauwen, B., 2014, May. End-to-end learning for music audio. In 2014 IEEE\nInternational Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6964-6968). IEEE.\n\n• Pons, J., Nieto, O., Prockup, M., Schmidt, E., Ehmann, A. and Serra, X., 2017. End-to-end learning for music audio tagging at scale.\n\n• Abdoli, S., Cardinal, P. and Koerich, A.L., 2019. End-to-End Environmental Sound Classification using a 1D Convolutional Neural Network. Expert Systems with Applications.\n\n\n\n**Lectures/Presentations:**\n\n• https://www.youtube.com/watch?v=7B1WBa3sC3I\n[Deep Neural Networks for Sound Event Detection]\n\n• https://www.youtube.com/watch?v=zvccOFz2KxI\n[Robust Sound Event Detection in Acoustic Sensor Networks]\n\n• https://www.youtube.com/watch?v=9X66iwEQSyI\n[Audio Event Detection w/Deep Learning at at Stanley Black & Decker]\n\n• https://www.youtube.com/watch?v=IzRCC2-IBwU\n[Machine Listening of Everyday Soundscapes]\n\nHope you like the compilation from TU Berlin . If you come across resources like this , please share it in comments and I can include it in the post .",
    "1092143": "I stumbled upon this series [Audio Signal Processing for Machine Learning](https://www.youtube.com/playlist?list=PL-wATfeyAMNqIee7cH3q1bh4QJFAaeNv0) by Valerio Velardo. I found it a terrific resource for a novice to get up to speed with audio.",
    "1100751": "@usharengaraju As I have started this competition as my research project, I will be adding contents, articles, video explanations to this thread, hope It would help the community.\n\n1. Here is an awesome, well-organized, **first-stop for audio analysis** : https://www.audiocontentanalysis.org/teaching/\nThe videos are very well explained and it has code implementations as well. Do check out.\n2. **Mel Frequency Cepstral Coefficients** Detailed Explanation: https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd",
    "1183128": "Thanks for sharing @usharengaraju , another important reference might be the Cornell birdcall competition which was held recently. \n\nhttps://www.kaggle.com/c/birdsong-recognition",
    "1137886": "Thanks for this! ^_^",
    "1105024": "Thanks for sharing",
    "1094780": "Great info, thanks for sharing !!",
    "1092108": "Thank you for sharing"
  }
}