{
  "id": 308393,
  "title": "Papers on Soundscapes and Animal Sound detection",
  "url": "/competitions/birdclef-2022/discussion/308393",
  "author_name": "",
  "post_date": "2022-02-18T13:04:31.662472700Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey everyone!</p>\n<p>I wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have no domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!</p>\n<p><strong>Research Papers on Soundscapes:</strong></p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/2103.03483\" target=\"_blank\">Environmental Sound Classification on the Edge: Deep Acoustic Networks for Extremely Resource-Constrained Devices</a> - SoTA work on ESC50 dataset.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2011.00801\" target=\"_blank\">Sound Event Detection and Separation: a Benchmark on Desed Synthetic Soundscapes</a> - We show that the localization in time of sound events is still a problem for SED systems. We also show that reverberation and non-target sound events are severely degrading the performance of the SED systems. In the latter case, sound separation seems like a promising solution.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2007.03931\" target=\"_blank\">Training Sound Event Detection On A Heterogeneous Dataset</a> - We propose to perform a detailed analysis of DCASE 2020 task 4 sound event detection baseline with regards to several aspects such as the type of data used for training, the parameters of the mean-teacher or the transformations applied while generating the synthetic soundscapes. Some of the parameters that are usually used as default are shown to be sub-optimal.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1907.12812\" target=\"_blank\">An artifcial life approach to studying niche differentiation in soundscape ecology</a> -  The experiments in this paper test that hypothesis in a simulated soundscape in order to verify the feasibility of intraspecies communication as a driver of acoustic niche differentiation.</p></li>\n</ul>\n<p><strong>Research Papers on discerning animal sounds:</strong></p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/2011.00175\" target=\"_blank\">Convolutional Neural Networks Based System for Urban Sound Tagging with Spatiotemporal Context</a> - In this paper, we proposed convolutional neural networks (CNNs) based system for UST with spatiotemporal context. In our system, multiple features and spatiotemporal context are combined, and fed into a residual CNN to predict whether noise of pollution is present in a 10-second recording. </p></li>\n<li><p><a href=\"https://arxiv.org/abs/1902.09069\" target=\"_blank\">Automatic Detection and Compression for Passive Acoustic Monitoring of the African Forest Elephant</a> - In collaboration with conservation efforts, we construct a large labeled dataset of passive acoustic recordings of the African Forest Elephant via crowdsourcing, compromising thousands of hours of recordings in the wild. Using state-of-the-art techniques in artificial intelligence we improve upon previously proposed methods for passive acoustic monitoring for classification and segmentation.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1810.09273\" target=\"_blank\">Automatic acoustic identification of individual animals: Improving generalisation across species and recording conditions</a> - Here we present a general automatic identification method, that can work across multiple animal species with various levels of complexity in their communication systems. We further introduce new analysis techniques based on dataset manipulations that can evaluate the robustness and generality of a classifier.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1810.09078\" target=\"_blank\">Our Practice Of Using Machine Learning To Recognize Species By Voice</a> - The best way to monitor those species are through audio recognition. Classifying sound can be a difficult task even for humans. Powerful audio signals and their processing techniques make it possible to detect audio of various species. </p></li>\n<li><p><a href=\"https://arxiv.org/abs/1804.05502\" target=\"_blank\">Automatic Rain and Cicada Chorus Filtering of Bird Acoustic Data</a> - In this paper, we address the challenge of filtering noise from rain and cicada choruses from recordings containing bird sound. We improve upon previously established classification approaches using acoustic indices and Mel Frequency Cepstral Coefficients (MFCCs) as acoustic features to detect these noise sources, approaching the problem with the motivation of removing these sounds. We investigate the use of acoustic indices, and machine learning classifiers to find the most effective filters.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1909.07526\" target=\"_blank\">Data-Efficient Classification of Birdcall Through Convolutional Neural Networks Transfer Learning</a> - Deep learning Convolutional Neural Network (CNN) models are powerful classification models but require a large amount of training data. In niche domains such as bird acoustics, it is expensive and difficult to obtain a large number of training samples. One method of classifying data with a limited number of training samples is to employ transfer learning. In this research, we evaluated the effectiveness of birdcall classification using transfer learning from a larger base dataset (2814 samples in 46 classes) to a smaller target dataset (351 samples in 10 classes) using the ResNet-50 CNN. We obtained 79% average validation accuracy on the target dataset in 5-fold cross-validation. The methodology of transfer learning from an ImageNet-trained CNN to a project-specific and a much smaller set of classes and images was extended to the domain of spectrogram images, where the base dataset effectively played the role of the ImageNet.</p></li>\n</ul>",
  "messages": [
    {
      "id": "1695914",
      "postDate": "02/18/2022 13:04:31",
      "content": "<p>Hey everyone!</p>\n<p>I wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have no domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!</p>\n<p><strong>Research Papers on Soundscapes:</strong></p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/2103.03483\" target=\"_blank\">Environmental Sound Classification on the Edge: Deep Acoustic Networks for Extremely Resource-Constrained Devices</a> - SoTA work on ESC50 dataset.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2011.00801\" target=\"_blank\">Sound Event Detection and Separation: a Benchmark on Desed Synthetic Soundscapes</a> - We show that the localization in time of sound events is still a problem for SED systems. We also show that reverberation and non-target sound events are severely degrading the performance of the SED systems. In the latter case, sound separation seems like a promising solution.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2007.03931\" target=\"_blank\">Training Sound Event Detection On A Heterogeneous Dataset</a> - We propose to perform a detailed analysis of DCASE 2020 task 4 sound event detection baseline with regards to several aspects such as the type of data used for training, the parameters of the mean-teacher or the transformations applied while generating the synthetic soundscapes. Some of the parameters that are usually used as default are shown to be sub-optimal.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1907.12812\" target=\"_blank\">An artifcial life approach to studying niche differentiation in soundscape ecology</a> -  The experiments in this paper test that hypothesis in a simulated soundscape in order to verify the feasibility of intraspecies communication as a driver of acoustic niche differentiation.</p></li>\n</ul>\n<p><strong>Research Papers on discerning animal sounds:</strong></p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/2011.00175\" target=\"_blank\">Convolutional Neural Networks Based System for Urban Sound Tagging with Spatiotemporal Context</a> - In this paper, we proposed convolutional neural networks (CNNs) based system for UST with spatiotemporal context. In our system, multiple features and spatiotemporal context are combined, and fed into a residual CNN to predict whether noise of pollution is present in a 10-second recording. </p></li>\n<li><p><a href=\"https://arxiv.org/abs/1902.09069\" target=\"_blank\">Automatic Detection and Compression for Passive Acoustic Monitoring of the African Forest Elephant</a> - In collaboration with conservation efforts, we construct a large labeled dataset of passive acoustic recordings of the African Forest Elephant via crowdsourcing, compromising thousands of hours of recordings in the wild. Using state-of-the-art techniques in artificial intelligence we improve upon previously proposed methods for passive acoustic monitoring for classification and segmentation.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1810.09273\" target=\"_blank\">Automatic acoustic identification of individual animals: Improving generalisation across species and recording conditions</a> - Here we present a general automatic identification method, that can work across multiple animal species with various levels of complexity in their communication systems. We further introduce new analysis techniques based on dataset manipulations that can evaluate the robustness and generality of a classifier.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1810.09078\" target=\"_blank\">Our Practice Of Using Machine Learning To Recognize Species By Voice</a> - The best way to monitor those species are through audio recognition. Classifying sound can be a difficult task even for humans. Powerful audio signals and their processing techniques make it possible to detect audio of various species. </p></li>\n<li><p><a href=\"https://arxiv.org/abs/1804.05502\" target=\"_blank\">Automatic Rain and Cicada Chorus Filtering of Bird Acoustic Data</a> - In this paper, we address the challenge of filtering noise from rain and cicada choruses from recordings containing bird sound. We improve upon previously established classification approaches using acoustic indices and Mel Frequency Cepstral Coefficients (MFCCs) as acoustic features to detect these noise sources, approaching the problem with the motivation of removing these sounds. We investigate the use of acoustic indices, and machine learning classifiers to find the most effective filters.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1909.07526\" target=\"_blank\">Data-Efficient Classification of Birdcall Through Convolutional Neural Networks Transfer Learning</a> - Deep learning Convolutional Neural Network (CNN) models are powerful classification models but require a large amount of training data. In niche domains such as bird acoustics, it is expensive and difficult to obtain a large number of training samples. One method of classifying data with a limited number of training samples is to employ transfer learning. In this research, we evaluated the effectiveness of birdcall classification using transfer learning from a larger base dataset (2814 samples in 46 classes) to a smaller target dataset (351 samples in 10 classes) using the ResNet-50 CNN. We obtained 79% average validation accuracy on the target dataset in 5-fold cross-validation. The methodology of transfer learning from an ImageNet-trained CNN to a project-specific and a much smaller set of classes and images was extended to the domain of spectrogram images, where the base dataset effectively played the role of the ImageNet.</p></li>\n</ul>",
      "rawMarkdown": "Hey everyone!\n\nI wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have no domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!\n\n**Research Papers on Soundscapes:**\n\n* [Environmental Sound Classification on the Edge: Deep Acoustic Networks for Extremely Resource-Constrained Devices](https://arxiv.org/abs/2103.03483) - SoTA work on ESC50 dataset.\n\n* [Sound Event Detection and Separation: a Benchmark on Desed Synthetic Soundscapes](https://arxiv.org/abs/2011.00801) - We show that the localization in time of sound events is still a problem for SED systems. We also show that reverberation and non-target sound events are severely degrading the performance of the SED systems. In the latter case, sound separation seems like a promising solution.\n\n* [Training Sound Event Detection On A Heterogeneous Dataset](https://arxiv.org/abs/2007.03931) - We propose to perform a detailed analysis of DCASE 2020 task 4 sound event detection baseline with regards to several aspects such as the type of data used for training, the parameters of the mean-teacher or the transformations applied while generating the synthetic soundscapes. Some of the parameters that are usually used as default are shown to be sub-optimal.\n\n* [An artifcial life approach to studying niche differentiation in soundscape ecology](https://arxiv.org/abs/1907.12812) -  The experiments in this paper test that hypothesis in a simulated soundscape in order to verify the feasibility of intraspecies communication as a driver of acoustic niche differentiation.\n\n**Research Papers on discerning animal sounds:**\n\n* [Convolutional Neural Networks Based System for Urban Sound Tagging with Spatiotemporal Context](https://arxiv.org/abs/2011.00175) - In this paper, we proposed convolutional neural networks (CNNs) based system for UST with spatiotemporal context. In our system, multiple features and spatiotemporal context are combined, and fed into a residual CNN to predict whether noise of pollution is present in a 10-second recording. \n\n* [Automatic Detection and Compression for Passive Acoustic Monitoring of the African Forest Elephant](https://arxiv.org/abs/1902.09069) - In collaboration with conservation efforts, we construct a large labeled dataset of passive acoustic recordings of the African Forest Elephant via crowdsourcing, compromising thousands of hours of recordings in the wild. Using state-of-the-art techniques in artificial intelligence we improve upon previously proposed methods for passive acoustic monitoring for classification and segmentation.\n\n* [Automatic acoustic identification of individual animals: Improving generalisation across species and recording conditions](https://arxiv.org/abs/1810.09273) - Here we present a general automatic identification method, that can work across multiple animal species with various levels of complexity in their communication systems. We further introduce new analysis techniques based on dataset manipulations that can evaluate the robustness and generality of a classifier.\n\n* [Our Practice Of Using Machine Learning To Recognize Species By Voice](https://arxiv.org/abs/1810.09078) - The best way to monitor those species are through audio recognition. Classifying sound can be a difficult task even for humans. Powerful audio signals and their processing techniques make it possible to detect audio of various species. \n\n* [Automatic Rain and Cicada Chorus Filtering of Bird Acoustic Data](https://arxiv.org/abs/1804.05502) - In this paper, we address the challenge of filtering noise from rain and cicada choruses from recordings containing bird sound. We improve upon previously established classification approaches using acoustic indices and Mel Frequency Cepstral Coefficients (MFCCs) as acoustic features to detect these noise sources, approaching the problem with the motivation of removing these sounds. We investigate the use of acoustic indices, and machine learning classifiers to find the most effective filters.\n\n* [Data-Efficient Classification of Birdcall Through Convolutional Neural Networks Transfer Learning](https://arxiv.org/abs/1909.07526) - Deep learning Convolutional Neural Network (CNN) models are powerful classification models but require a large amount of training data. In niche domains such as bird acoustics, it is expensive and difficult to obtain a large number of training samples. One method of classifying data with a limited number of training samples is to employ transfer learning. In this research, we evaluated the effectiveness of birdcall classification using transfer learning from a larger base dataset (2814 samples in 46 classes) to a smaller target dataset (351 samples in 10 classes) using the ResNet-50 CNN. We obtained 79% average validation accuracy on the target dataset in 5-fold cross-validation. The methodology of transfer learning from an ImageNet-trained CNN to a project-specific and a much smaller set of classes and images was extended to the domain of spectrogram images, where the base dataset effectively played the role of the ImageNet.",
      "votes": null
    },
    {
      "id": "1763482",
      "postDate": "04/21/2022 14:59:39",
      "content": "<p>nice work! i am also new in this area.  these papers may help me a lot .</p>",
      "rawMarkdown": "nice work! i am also new in this area.  these papers may help me a lot .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1763482,
      "author_name": "xuxiaodong",
      "author_url": "",
      "post_date": "04/21/2022 14:59:39",
      "content": "<p>nice work! i am also new in this area.  these papers may help me a lot .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1695914": "Hey everyone!\n\nI wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have no domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!\n\n**Research Papers on Soundscapes:**\n\n* [Environmental Sound Classification on the Edge: Deep Acoustic Networks for Extremely Resource-Constrained Devices](https://arxiv.org/abs/2103.03483) - SoTA work on ESC50 dataset.\n\n* [Sound Event Detection and Separation: a Benchmark on Desed Synthetic Soundscapes](https://arxiv.org/abs/2011.00801) - We show that the localization in time of sound events is still a problem for SED systems. We also show that reverberation and non-target sound events are severely degrading the performance of the SED systems. In the latter case, sound separation seems like a promising solution.\n\n* [Training Sound Event Detection On A Heterogeneous Dataset](https://arxiv.org/abs/2007.03931) - We propose to perform a detailed analysis of DCASE 2020 task 4 sound event detection baseline with regards to several aspects such as the type of data used for training, the parameters of the mean-teacher or the transformations applied while generating the synthetic soundscapes. Some of the parameters that are usually used as default are shown to be sub-optimal.\n\n* [An artifcial life approach to studying niche differentiation in soundscape ecology](https://arxiv.org/abs/1907.12812) -  The experiments in this paper test that hypothesis in a simulated soundscape in order to verify the feasibility of intraspecies communication as a driver of acoustic niche differentiation.\n\n**Research Papers on discerning animal sounds:**\n\n* [Convolutional Neural Networks Based System for Urban Sound Tagging with Spatiotemporal Context](https://arxiv.org/abs/2011.00175) - In this paper, we proposed convolutional neural networks (CNNs) based system for UST with spatiotemporal context. In our system, multiple features and spatiotemporal context are combined, and fed into a residual CNN to predict whether noise of pollution is present in a 10-second recording. \n\n* [Automatic Detection and Compression for Passive Acoustic Monitoring of the African Forest Elephant](https://arxiv.org/abs/1902.09069) - In collaboration with conservation efforts, we construct a large labeled dataset of passive acoustic recordings of the African Forest Elephant via crowdsourcing, compromising thousands of hours of recordings in the wild. Using state-of-the-art techniques in artificial intelligence we improve upon previously proposed methods for passive acoustic monitoring for classification and segmentation.\n\n* [Automatic acoustic identification of individual animals: Improving generalisation across species and recording conditions](https://arxiv.org/abs/1810.09273) - Here we present a general automatic identification method, that can work across multiple animal species with various levels of complexity in their communication systems. We further introduce new analysis techniques based on dataset manipulations that can evaluate the robustness and generality of a classifier.\n\n* [Our Practice Of Using Machine Learning To Recognize Species By Voice](https://arxiv.org/abs/1810.09078) - The best way to monitor those species are through audio recognition. Classifying sound can be a difficult task even for humans. Powerful audio signals and their processing techniques make it possible to detect audio of various species. \n\n* [Automatic Rain and Cicada Chorus Filtering of Bird Acoustic Data](https://arxiv.org/abs/1804.05502) - In this paper, we address the challenge of filtering noise from rain and cicada choruses from recordings containing bird sound. We improve upon previously established classification approaches using acoustic indices and Mel Frequency Cepstral Coefficients (MFCCs) as acoustic features to detect these noise sources, approaching the problem with the motivation of removing these sounds. We investigate the use of acoustic indices, and machine learning classifiers to find the most effective filters.\n\n* [Data-Efficient Classification of Birdcall Through Convolutional Neural Networks Transfer Learning](https://arxiv.org/abs/1909.07526) - Deep learning Convolutional Neural Network (CNN) models are powerful classification models but require a large amount of training data. In niche domains such as bird acoustics, it is expensive and difficult to obtain a large number of training samples. One method of classifying data with a limited number of training samples is to employ transfer learning. In this research, we evaluated the effectiveness of birdcall classification using transfer learning from a larger base dataset (2814 samples in 46 classes) to a smaller target dataset (351 samples in 10 classes) using the ResNet-50 CNN. We obtained 79% average validation accuracy on the target dataset in 5-fold cross-validation. The methodology of transfer learning from an ImageNet-trained CNN to a project-specific and a much smaller set of classes and images was extended to the domain of spectrogram images, where the base dataset effectively played the role of the ImageNet.",
    "1763482": "nice work! i am also new in this area.  these papers may help me a lot ."
  },
  "source": "meta"
}