{
  "id": 166486,
  "title": "Detecting bird sound in unknown acoustic background using crowdsourced training data",
  "url": "/competitions/birdsong-recognition/discussion/166486",
  "author_name": "",
  "post_date": "2020-07-13T05:13:46.659650900Z",
  "votes": 25,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have been reviewing some of the existing literature related to bird sounds and writing summaries of the papers. Below is the summary of <a href=\"https://arxiv.org/abs/1505.06443\">Detecting bird sound in unknown acoustic background using crowdsourced training data</a></p>\n\n<p>Paper evaluates models performance at discriminating background audio vs bird sound periods. Most models have a strong assumption that bird audio will almost always be present and will output a bird class prediction even when no birds are present. When applying this to organic recordings this is problematic because it will trigger lots of false positives and confusion as bird songs are likely only present in a small portion of recordings. </p>\n\n<p>Utilizes data from xeno-canto and is applied to 15 bird species. Doesn’t require manual preprocessing and could be scaled to the 9k species available on xeno-canto in theory. </p>\n\n<p>Utilizes various other audio datasets as background noise in order to create a negative (no-bird sounds) class for the model to learn. \n- IEEE AASP Challenge: Detection and Classification of Acoustic Scenes and Events competition database</p>\n\n<ul>\n<li>Park, expected to be easy as large segments of near silence</li>\n<li>Open air market, expected to be much more difficult as it’s filled with human speech</li>\n</ul>\n\n<p>Selected data from Xeno-canto with an ‘A’ rating so there are less background bird sounds that might confuse the model. </p>\n\n<p>“We compute spectrograms for each recording (20ms frame length, 50% overlap, FFT length equal to the closest power of two that gives at least 93Hz frequency bin spacing, for the sampling rate of the specific recording). We only keep frequency bins in the range of 1-10kHz.”</p>\n\n<p>1-10khz in theory is a good range because it discards low frequency wind noise and focuses on most birds' vocal range.</p>\n\n<p>The bottom 90% of spectrogram frames based on power are discarded in hopes that that will eliminate periods of silence. 6 spectral statistic features are extracted to be used as features. This reduces the bird sound training data to 6000 feature vectors which represent 60-120 sec of bird only data. A gaussian mixture model is trained on these feature vectors. The same process is applied to the test data except the bottom 90% is not skipped so as to test against these quiet periods. </p>\n\n<p>Performance was shown to be good for some birds but poor for others. Above .9 AUC for in specific scenarios like the quiet park scenario, but very poor with certain species. Working hypothesis here is the poor performers may be from birds that operate in lower frequencies that have greater overlap with background noise profiles. </p>",
  "messages": [
    {
      "id": "926920",
      "postDate": "07/13/2020 05:13:46",
      "content": "<p>I have been reviewing some of the existing literature related to bird sounds and writing summaries of the papers. Below is the summary of <a href=\"https://arxiv.org/abs/1505.06443\">Detecting bird sound in unknown acoustic background using crowdsourced training data</a></p>\n\n<p>Paper evaluates models performance at discriminating background audio vs bird sound periods. Most models have a strong assumption that bird audio will almost always be present and will output a bird class prediction even when no birds are present. When applying this to organic recordings this is problematic because it will trigger lots of false positives and confusion as bird songs are likely only present in a small portion of recordings. </p>\n\n<p>Utilizes data from xeno-canto and is applied to 15 bird species. Doesn’t require manual preprocessing and could be scaled to the 9k species available on xeno-canto in theory. </p>\n\n<p>Utilizes various other audio datasets as background noise in order to create a negative (no-bird sounds) class for the model to learn. \n- IEEE AASP Challenge: Detection and Classification of Acoustic Scenes and Events competition database</p>\n\n<ul>\n<li>Park, expected to be easy as large segments of near silence</li>\n<li>Open air market, expected to be much more difficult as it’s filled with human speech</li>\n</ul>\n\n<p>Selected data from Xeno-canto with an ‘A’ rating so there are less background bird sounds that might confuse the model. </p>\n\n<p>“We compute spectrograms for each recording (20ms frame length, 50% overlap, FFT length equal to the closest power of two that gives at least 93Hz frequency bin spacing, for the sampling rate of the specific recording). We only keep frequency bins in the range of 1-10kHz.”</p>\n\n<p>1-10khz in theory is a good range because it discards low frequency wind noise and focuses on most birds' vocal range.</p>\n\n<p>The bottom 90% of spectrogram frames based on power are discarded in hopes that that will eliminate periods of silence. 6 spectral statistic features are extracted to be used as features. This reduces the bird sound training data to 6000 feature vectors which represent 60-120 sec of bird only data. A gaussian mixture model is trained on these feature vectors. The same process is applied to the test data except the bottom 90% is not skipped so as to test against these quiet periods. </p>\n\n<p>Performance was shown to be good for some birds but poor for others. Above .9 AUC for in specific scenarios like the quiet park scenario, but very poor with certain species. Working hypothesis here is the poor performers may be from birds that operate in lower frequencies that have greater overlap with background noise profiles. </p>",
      "rawMarkdown": "I have been reviewing some of the existing literature related to bird sounds and writing summaries of the papers. Below is the summary of [Detecting bird sound in unknown acoustic background using crowdsourced training data](https://arxiv.org/abs/1505.06443)\n\nPaper evaluates models performance at discriminating background audio vs bird sound periods. Most models have a strong assumption that bird audio will almost always be present and will output a bird class prediction even when no birds are present. When applying this to organic recordings this is problematic because it will trigger lots of false positives and confusion as bird songs are likely only present in a small portion of recordings. \n\nUtilizes data from xeno-canto and is applied to 15 bird species. Doesn’t require manual preprocessing and could be scaled to the 9k species available on xeno-canto in theory. \n\nUtilizes various other audio datasets as background noise in order to create a negative (no-bird sounds) class for the model to learn. \n- IEEE AASP Challenge: Detection and Classification of Acoustic Scenes and Events competition database\n\n  - Park, expected to be easy as large segments of near silence\n  - Open air market, expected to be much more difficult as it’s filled with human speech\n\nSelected data from Xeno-canto with an ‘A’ rating so there are less background bird sounds that might confuse the model. \n\n“We compute spectrograms for each recording (20ms frame length, 50% overlap, FFT length equal to the closest power of two that gives at least 93Hz frequency bin spacing, for the sampling rate of the specific recording). We only keep frequency bins in the range of 1-10kHz.”\n\n1-10khz in theory is a good range because it discards low frequency wind noise and focuses on most birds' vocal range.\n\nThe bottom 90% of spectrogram frames based on power are discarded in hopes that that will eliminate periods of silence. 6 spectral statistic features are extracted to be used as features. This reduces the bird sound training data to 6000 feature vectors which represent 60-120 sec of bird only data. A gaussian mixture model is trained on these feature vectors. The same process is applied to the test data except the bottom 90% is not skipped so as to test against these quiet periods. \n\nPerformance was shown to be good for some birds but poor for others. Above .9 AUC for in specific scenarios like the quiet park scenario, but very poor with certain species. Working hypothesis here is the poor performers may be from birds that operate in lower frequencies that have greater overlap with background noise profiles.",
      "votes": null
    },
    {
      "id": "950157",
      "postDate": "07/29/2020 08:04:47",
      "content": "<p>Thank you for the summary. I guess there will be \"easy\" bird and \"difficult\" bird to identify </p>",
      "rawMarkdown": "Thank you for the summary. I guess there will be \"easy\" bird and \"difficult\" bird to identify",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 950157,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "07/29/2020 08:04:47",
      "content": "<p>Thank you for the summary. I guess there will be \"easy\" bird and \"difficult\" bird to identify </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "926920": "I have been reviewing some of the existing literature related to bird sounds and writing summaries of the papers. Below is the summary of [Detecting bird sound in unknown acoustic background using crowdsourced training data](https://arxiv.org/abs/1505.06443)\n\nPaper evaluates models performance at discriminating background audio vs bird sound periods. Most models have a strong assumption that bird audio will almost always be present and will output a bird class prediction even when no birds are present. When applying this to organic recordings this is problematic because it will trigger lots of false positives and confusion as bird songs are likely only present in a small portion of recordings. \n\nUtilizes data from xeno-canto and is applied to 15 bird species. Doesn’t require manual preprocessing and could be scaled to the 9k species available on xeno-canto in theory. \n\nUtilizes various other audio datasets as background noise in order to create a negative (no-bird sounds) class for the model to learn. \n- IEEE AASP Challenge: Detection and Classification of Acoustic Scenes and Events competition database\n\n  - Park, expected to be easy as large segments of near silence\n  - Open air market, expected to be much more difficult as it’s filled with human speech\n\nSelected data from Xeno-canto with an ‘A’ rating so there are less background bird sounds that might confuse the model. \n\n“We compute spectrograms for each recording (20ms frame length, 50% overlap, FFT length equal to the closest power of two that gives at least 93Hz frequency bin spacing, for the sampling rate of the specific recording). We only keep frequency bins in the range of 1-10kHz.”\n\n1-10khz in theory is a good range because it discards low frequency wind noise and focuses on most birds' vocal range.\n\nThe bottom 90% of spectrogram frames based on power are discarded in hopes that that will eliminate periods of silence. 6 spectral statistic features are extracted to be used as features. This reduces the bird sound training data to 6000 feature vectors which represent 60-120 sec of bird only data. A gaussian mixture model is trained on these feature vectors. The same process is applied to the test data except the bottom 90% is not skipped so as to test against these quiet periods. \n\nPerformance was shown to be good for some birds but poor for others. Above .9 AUC for in specific scenarios like the quiet park scenario, but very poor with certain species. Working hypothesis here is the poor performers may be from birds that operate in lower frequencies that have greater overlap with background noise profiles.",
    "950157": "Thank you for the summary. I guess there will be \"easy\" bird and \"difficult\" bird to identify"
  },
  "source": "meta"
}