{
  "id": 169538,
  "title": "[placeholder] creating annotated soundscapes for  local validation set ",
  "url": "/competitions/birdsong-recognition/discussion/169538",
  "author_name": "",
  "post_date": "2020-07-24T06:17:16.766622400Z",
  "votes": 46,
  "comment_count": 48,
  "views": 0,
  "content": "<p>[deleted]</p>\n\n<p>please wait ... </p>\n\n<ul>\n<li><p>i am creating a procedure to read and playback annotation like \"example_test_audio_summary.csv\" using Raven+R+python ...</p></li>\n<li><p>i realize i make a mistake and is fixing it now. </p></li>\n<li><p>the post will be updated later</p></li>\n</ul>",
  "messages": [
    {
      "id": "943089",
      "postDate": "07/24/2020 06:17:16",
      "content": "<p>[deleted]</p>\n\n<p>please wait ... </p>\n\n<ul>\n<li><p>i am creating a procedure to read and playback annotation like \"example_test_audio_summary.csv\" using Raven+R+python ...</p></li>\n<li><p>i realize i make a mistake and is fixing it now. </p></li>\n<li><p>the post will be updated later</p></li>\n</ul>",
      "rawMarkdown": "[deleted]\n\nplease wait ... \n\n- i am creating a procedure to read and playback annotation like \"example_test_audio_summary.csv\" using Raven+R+python ...\n\n- i realize i make a mistake and is fixing it now. \n\n- the post will be updated later",
      "votes": null
    },
    {
      "id": "943130",
      "postDate": "07/24/2020 06:54:53",
      "content": "<p>Thanks for starting the thread!\nSo far I tried to use the <a href=\"https://www.kaggle.com/c/mlsp-2013-birds\">MLSP 2013 Bird Classification Challenge</a> data . I will create a kaggle dataset with useful labels.</p>\n\n<p><strong>Pros</strong></p>\n\n<ul>\n<li>645 x 10s soundscapes recorded in Oregon</li>\n<li>multiclass-multilabel dataset</li>\n</ul>\n\n<p><strong>Cons</strong></p>\n\n<ul>\n<li>Only 19 bird species</li>\n<li>16 kHz wav files</li>\n<li>7 year old recordings</li>\n</ul>",
      "rawMarkdown": "Thanks for starting the thread!\nSo far I tried to use the [MLSP 2013 Bird Classification Challenge](https://www.kaggle.com/c/mlsp-2013-birds) data . I will create a kaggle dataset with useful labels.\n\n**Pros**\n\n- 645 x 10s soundscapes recorded in Oregon\n- multiclass-multilabel dataset\n\n**Cons**\n\n- Only 19 bird species\n- 16 kHz wav files\n- 7 year old recordings",
      "votes": null
    },
    {
      "id": "943526",
      "postDate": "07/24/2020 12:15:22",
      "content": "<p>for a start, please refer to these slides<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd322e9eb5cf653be8cdfa1a0dc6ee21f%2FSlide1.png?generation=1595592879332532&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe5bd1ac4f56de723b523b51108eb5b1f%2FSlide2.png?generation=1595592896697819&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F873f3596985621ed74201b7a345c70df%2FSlide4.png?generation=1595592918527139&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "for a start, please refer to these slides![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd322e9eb5cf653be8cdfa1a0dc6ee21f%2FSlide1.png?generation=1595592879332532&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe5bd1ac4f56de723b523b51108eb5b1f%2FSlide2.png?generation=1595592896697819&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F873f3596985621ed74201b7a345c70df%2FSlide4.png?generation=1595592918527139&amp;alt=media)",
      "votes": null
    },
    {
      "id": "943619",
      "postDate": "07/24/2020 13:25:40",
      "content": "<p>is this a bug ??<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F59995c501cc1d135f011b89ff05bfbdf%2FSelection_026.png?generation=1595597138393031&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "is this a bug ??![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F59995c501cc1d135f011b89ff05bfbdf%2FSelection_026.png?generation=1595597138393031&amp;alt=media)",
      "votes": null
    },
    {
      "id": "944112",
      "postDate": "07/24/2020 20:53:33",
      "content": "<p>i attempt to visualize the annotations of BirdCLEF 2020 provided by the competition host at <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877\">https://www.kaggle.com/c/birdsong-recognition/discussion/158877</a></p>\n\n<p>they are very difficult! i think we may need to use train data with \"the lowest rating\" from <a href=\"https://ebird.org/media\">https://ebird.org/media</a>, etc. The quality of the train data we have is very high, compared to BirdCLEF 2020. I am not sure what we would be expecting for the test here ... anyway, do prepared for very bad test data.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F6a39ce2151aeb57d0b36c72be7f7156b%2FSelection_034.png?generation=1595623924948889&amp;alt=media\" alt=\"\"></p>\n\n<p>attached : raven-lite selection (annotation) file for  SSW51_20170819.wav</p>",
      "rawMarkdown": "i attempt to visualize the annotations of BirdCLEF 2020 provided by the competition host at https://www.kaggle.com/c/birdsong-recognition/discussion/158877\n\nthey are very difficult! i think we may need to use train data with \"the lowest rating\" from https://ebird.org/media, etc. The quality of the train data we have is very high, compared to BirdCLEF 2020. I am not sure what we would be expecting for the test here ... anyway, do prepared for very bad test data.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F6a39ce2151aeb57d0b36c72be7f7156b%2FSelection_034.png?generation=1595623924948889&amp;alt=media)\n\nattached : raven-lite selection (annotation) file for  SSW51_20170819.wav",
      "votes": null
    },
    {
      "id": "944246",
      "postDate": "07/25/2020 00:44:07",
      "content": "<p>check this work, it should be the baseline method for this challenge: \n<a href=\"http://ceur-ws.org/Vol-2380/paper_86.pdf\">http://ceur-ws.org/Vol-2380/paper_86.pdf</a></p>\n\n<p>Bird Species Identification in Soundscapes - Mario Lasseck</p>\n\n<p>\"Deep Convolutional Neural Networks are trained to classify 659 species. Different data augmentation techniques are applied to prevent overfitting and improve model accuracy and generalization. The proposed approach is evaluated in the BirdCLEF 2019 campaign and provides the best system to identify bird species in wildlife monitoring recordings. \"</p>\n\n<hr>\n\n<p>another good baseline\n<a href=\"https://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html\">https://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html</a></p>\n\n<p>Acoustic Detection of Humpback Whales Using a Convolutional Neural Network\n<a href=\"https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1e442a65981435576a6ac33e4c3178d6a62a06a1.pdf\">https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1e442a65981435576a6ac33e4c3178d6a62a06a1.pdf</a></p>\n\n<p><a href=\"https://medium.com/@kcimc/data-of-the-humpback-whale-9ef09c5920cd\">https://medium.com/@kcimc/data-of-the-humpback-whale-9ef09c5920cd</a></p>\n\n<p>\"Long-distance detection of bioacoustic events with per-channel energy normalization\"</p>",
      "rawMarkdown": "check this work, it should be the baseline method for this challenge: \nhttp://ceur-ws.org/Vol-2380/paper_86.pdf\n\nBird Species Identification in Soundscapes - Mario Lasseck\n\n\"Deep Convolutional Neural Networks are trained to classify 659 species. Different data augmentation techniques are applied to prevent overfitting and improve model accuracy and generalization. The proposed approach is evaluated in the BirdCLEF 2019 campaign and provides the best system to identify bird species in wildlife monitoring recordings. \"\n\n---\nanother good baseline\nhttps://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html\n\nAcoustic Detection of Humpback Whales Using a Convolutional Neural Network\nhttps://storage.googleapis.com/pub-tools-public-publication-data/pdf/1e442a65981435576a6ac33e4c3178d6a62a06a1.pdf\n\nhttps://medium.com/@kcimc/data-of-the-humpback-whale-9ef09c5920cd\n\n\"Long-distance detection of bioacoustic events with per-channel energy normalization\"",
      "votes": null
    },
    {
      "id": "944283",
      "postDate": "07/25/2020 01:57:20",
      "content": "<p>type of noise in bird recording:\n\"Birdsong Denoising Using Wavelets\" : <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4728069/\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4728069/</a></p>\n\n<p>it has a software:\n<a href=\"http://www.avianz.net/index.php/avianz-software/user-manual\">http://www.avianz.net/index.php/avianz-software/user-manual</a></p>",
      "rawMarkdown": "type of noise in bird recording:\n\"Birdsong Denoising Using Wavelets\" : https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4728069/\n\nit has a software:\nhttp://www.avianz.net/index.php/avianz-software/user-manual",
      "votes": null
    },
    {
      "id": "944706",
      "postDate": "07/25/2020 09:23:07",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=8xH2GjHKYj0\">https://www.youtube.com/watch?v=8xH2GjHKYj0</a></p>\n\n<p>\"Bird Song Hero: The song learning game for everyone\"\nCornell Lab of Ornithology</p>\n\n<p>learn how to use spectrogram to identify birds. seems that \"counts\" of the repeated tune is sometimes a feature</p>\n\n<p>related: \n<a href=\"https://www.youtube.com/watch?v=qp34UfXY_Wk\">https://www.youtube.com/watch?v=qp34UfXY_Wk</a>\n<a href=\"https://www.youtube.com/watch?v=04HpCmjlQs8\">https://www.youtube.com/watch?v=04HpCmjlQs8</a>\n<a href=\"https://www.youtube.com/watch?v=4_1zIwEENt8\">https://www.youtube.com/watch?v=4_1zIwEENt8</a></p>",
      "rawMarkdown": "https://www.youtube.com/watch?v=8xH2GjHKYj0\n \n\"Bird Song Hero: The song learning game for everyone\"\nCornell Lab of Ornithology\n\nlearn how to use spectrogram to identify birds. seems that \"counts\" of the repeated tune is sometimes a feature\n \nrelated: \nhttps://www.youtube.com/watch?v=qp34UfXY_Wk\nhttps://www.youtube.com/watch?v=04HpCmjlQs8\nhttps://www.youtube.com/watch?v=4_1zIwEENt8",
      "votes": null
    },
    {
      "id": "944732",
      "postDate": "07/25/2020 09:57:42",
      "content": "<p>Augmentation: this shows how ta real spectrogram changes.</p>\n\n<p><a href=\"http://earbirding.com/blog/archives/category/spectrograms\">http://earbirding.com/blog/archives/category/spectrograms</a></p>\n\n<p><img src=\"http://earbirding.com/blog/wp-content/uploads/2013/02/VESPslow.gif\" alt=\"\"></p>\n\n<p>dynamic-time-warping  seems to be applicable for augmentation (or even clustering)</p>",
      "rawMarkdown": "Augmentation: this shows how ta real spectrogram changes.\n\nhttp://earbirding.com/blog/archives/category/spectrograms\n\n![](http://earbirding.com/blog/wp-content/uploads/2013/02/VESPslow.gif)\n\ndynamic-time-warping  seems to be applicable for augmentation (or even clustering)",
      "votes": null
    },
    {
      "id": "944739",
      "postDate": "07/25/2020 10:07:20",
      "content": "<p><a href=\"https://experiments.withgoogle.com/ai/bird-sounds/view/\">https://experiments.withgoogle.com/ai/bird-sounds/view/</a>\n<a href=\"https://github.com/googlecreativelab/aiexperiments-bird-sounds\">https://github.com/googlecreativelab/aiexperiments-bird-sounds</a></p>\n\n<p>t-SNE visualization of \"Essential Set for North America\"</p>",
      "rawMarkdown": "https://experiments.withgoogle.com/ai/bird-sounds/view/\nhttps://github.com/googlecreativelab/aiexperiments-bird-sounds\n\nt-SNE visualization of \"Essential Set for North America\"",
      "votes": null
    },
    {
      "id": "946157",
      "postDate": "07/26/2020 11:58:53",
      "content": "<p>i manged to find the missing bird in the label. I use audcity noise reduction to process the audio, here is the target bird in annotation:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F55c3cc1076bd8e57b719f69b9fb5cfb5%2FSelection_027.png?generation=1595764731183429&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://www.macaulaylibrary.org/resources/audio-editing-tutorials/\">https://www.macaulaylibrary.org/resources/audio-editing-tutorials/</a>\nAudio editing tutorials for birdsong</p>",
      "rawMarkdown": "i manged to find the missing bird in the label. I use audcity noise reduction to process the audio, here is the target bird in annotation:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F55c3cc1076bd8e57b719f69b9fb5cfb5%2FSelection_027.png?generation=1595764731183429&amp;alt=media)\n\n\nhttps://www.macaulaylibrary.org/resources/audio-editing-tutorials/\nAudio editing tutorials for birdsong",
      "votes": null
    },
    {
      "id": "946854",
      "postDate": "07/26/2020 22:34:30",
      "content": "<p><a href=\"https://github.com/CrowdCurio/audio-annotator\">https://github.com/CrowdCurio/audio-annotator</a>\naudio-annotator is a web interface that allows users to annotate audio recordings.\n<img src=\"https://github.com/CrowdCurio/audio-annotator/raw/master/static/img/task-interface.png\" alt=\"\"></p>",
      "rawMarkdown": "https://github.com/CrowdCurio/audio-annotator\naudio-annotator is a web interface that allows users to annotate audio recordings.\n![](https://github.com/CrowdCurio/audio-annotator/raw/master/static/img/task-interface.png)",
      "votes": null
    },
    {
      "id": "946877",
      "postDate": "07/26/2020 23:14:43",
      "content": "<p><a href=\"https://github.com/justinsalamon/scaper\">https://github.com/justinsalamon/scaper</a></p>\n\n<p>Scaper: A library for soundscape synthesis and augmentation\n<a href=\"https://www.youtube.com/watch?v=zvccOFz2KxI\">https://www.youtube.com/watch?v=zvccOFz2KxI</a></p>\n\n<p>\" Scaper, an open-source library for soundscape synthesis and augmentation. Given a collection of isolated sound events, Scaper acts as a high-level sequencer that can generate multiple soundscapes from a single, probabilistically defined, “specification”. \"</p>",
      "rawMarkdown": "https://github.com/justinsalamon/scaper\n\nScaper: A library for soundscape synthesis and augmentation\nhttps://www.youtube.com/watch?v=zvccOFz2KxI\n\n\" Scaper, an open-source library for soundscape synthesis and augmentation. Given a collection of isolated sound events, Scaper acts as a high-level sequencer that can generate multiple soundscapes from a single, probabilistically defined, “specification”. \"",
      "votes": null
    },
    {
      "id": "948536",
      "postDate": "07/28/2020 03:41:41",
      "content": "<p>waiting patiently :) Thank you for the work as always</p>",
      "rawMarkdown": "waiting patiently :) Thank you for the work as always",
      "votes": null
    },
    {
      "id": "949811",
      "postDate": "07/29/2020 00:57:51",
      "content": "<p>i spends some days to hand annotate the time interval for the bird calls. Here is what I find:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F886f6b16839e952db30a1263ce52f437%2FSelection_036.png?generation=1595984240703288&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F64b6272f907bb77ce631e3cee79c3a62%2FSelection_037.png?generation=1595984239768601&amp;alt=media\" alt=\"\"></p>\n\n<p>in short, it is not only weak label, but also noisy label !</p>\n\n<p>also, if an interval is being process, it is likely to contain the bird (hence there is some time interval annotation)</p>",
      "rawMarkdown": "i spends some days to hand annotate the time interval for the bird calls. Here is what I find:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F886f6b16839e952db30a1263ce52f437%2FSelection_036.png?generation=1595984240703288&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F64b6272f907bb77ce631e3cee79c3a62%2FSelection_037.png?generation=1595984239768601&amp;alt=media)\n\nin short, it is not only weak label, but also noisy label !\n\nalso, if an interval is being process, it is likely to contain the bird (hence there is some time interval annotation)",
      "votes": null
    },
    {
      "id": "949992",
      "postDate": "07/29/2020 05:33:13",
      "content": "<ol>\n<li>Two-thirds of the clips don't have secondary labels, however, some of those actually contain calls of other species.</li>\n</ol>",
      "rawMarkdown": "5. Two-thirds of the clips don't have secondary labels, however, some of those actually contain calls of other species.",
      "votes": null
    },
    {
      "id": "950516",
      "postDate": "07/29/2020 12:51:43",
      "content": "<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> How to deal with labels that don't contain secondary labels? Do you try to output all 0's in a secondary head, or just don't compute the loss for those rows? ty</p>",
      "rawMarkdown": "hidehisaarai1213 How to deal with labels that don't contain secondary labels? Do you try to output all 0's in a secondary head, or just don't compute the loss for those rows? ty",
      "votes": null
    },
    {
      "id": "951093",
      "postDate": "07/29/2020 21:44:44",
      "content": "<p>So far I make my model output all 0's for <em>potential</em> secondary labels. I don't clearly separate primary / secondary labels, so I provide 264 dimension one-hot vector for each sample and put 1 to corresponding positions for primary / secondary labels. If the sample does not have secondary labels, then I make that vector whose elements are all 0 except for the position that corresponds to primary label. </p>",
      "rawMarkdown": "So far I make my model output all 0's for *potential* secondary labels. I don't clearly separate primary / secondary labels, so I provide 264 dimension one-hot vector for each sample and put 1 to corresponding positions for primary / secondary labels. If the sample does not have secondary labels, then I make that vector whose elements are all 0 except for the position that corresponds to primary label.",
      "votes": null
    },
    {
      "id": "952597",
      "postDate": "07/31/2020 04:30:51",
      "content": "<p>here is additional observation\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F03bde33e15603e1e5ba15159cceaad0f%2FSelection_048.png?generation=1596169848623517&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "here is additional observation\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F03bde33e15603e1e5ba15159cceaad0f%2FSelection_048.png?generation=1596169848623517&amp;alt=media)",
      "votes": null
    },
    {
      "id": "952600",
      "postDate": "07/31/2020 04:34:14",
      "content": "<p>if you want to collect your your data, consider microphone array. just like 3d vision, triangulation using microphone array can localized different bird call. e.g.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F610fe3a6fa1a7c5b81c5e226992d9780%2FSelection_047.png?generation=1596169986819962&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://github.com/HARKBird-project/HARKBird\">https://github.com/HARKBird-project/HARKBird</a>\n<a href=\"https://sites.google.com/view/alcore-suzuki/home/harkbird\">https://sites.google.com/view/alcore-suzuki/home/harkbird</a></p>",
      "rawMarkdown": "if you want to collect your your data, consider microphone array. just like 3d vision, triangulation using microphone array can localized different bird call. e.g.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F610fe3a6fa1a7c5b81c5e226992d9780%2FSelection_047.png?generation=1596169986819962&amp;alt=media)\n\nhttps://github.com/HARKBird-project/HARKBird\nhttps://sites.google.com/view/alcore-suzuki/home/harkbird",
      "votes": null
    },
    {
      "id": "953188",
      "postDate": "07/31/2020 15:35:46",
      "content": "<p><a href=\"/hengck23\">@hengck23</a>  would you mind sharing which file is this ? I wonder if mp3 would have switched to short window during these periods when higher time resolution is needed </p>",
      "rawMarkdown": "hengck23  would you mind sharing which file is this ? I wonder if mp3 would have switched to short window during these periods when higher time resolution is needed",
      "votes": null
    },
    {
      "id": "953272",
      "postDate": "07/31/2020 16:37:35",
      "content": "<p>see swamp sparrow (swaspa) : \n6kHz (window size 256): XC446809, XC116589, XC131033\n9kHz (window size 128): XC138153 \nthis site provide reference call files\n<code>\nhttps://www.audubon.org/field-guide/bird/swamp-sparrow\nSongs and Calls\nSweet, musical trill, all on one note.\n fast pulse-rate song\n slow pulse-rate song\n very slow pulse-rate song\n odd buzzy song\n chips #1\n chips #2\n</code></p>",
      "rawMarkdown": "see swamp sparrow (swaspa) : \n6kHz (window size 256): XC446809, XC116589, XC131033\n9kHz (window size 128): XC138153 \nthis site provide reference call files\n```\nhttps://www.audubon.org/field-guide/bird/swamp-sparrow\nSongs and Calls\nSweet, musical trill, all on one note.\n fast pulse-rate song\n slow pulse-rate song\n very slow pulse-rate song\n odd buzzy song\n chips #1\n chips #2\n```",
      "votes": null
    },
    {
      "id": "953298",
      "postDate": "07/31/2020 17:08:36",
      "content": "<p><a href=\"https://drive.google.com/drive/folders/1SExd8V-vFfTRrmTgjJ5PXN34TFmCx3Hd?usp=sharing\">https://drive.google.com/drive/folders/1SExd8V-vFfTRrmTgjJ5PXN34TFmCx3Hd?usp=sharing</a></p>\n\n<p>some  time-interval hand annotation for the ebird swaspa . Note that i am not sure if my annotation would be 100% correct:</p>\n\n<p>annotated region : high confident that it should contain the bird \nnon-annotated region :  i cannot identified the bird ... it could be a true negative or a false negative(i.e. miss)</p>",
      "rawMarkdown": "https://drive.google.com/drive/folders/1SExd8V-vFfTRrmTgjJ5PXN34TFmCx3Hd?usp=sharing\n \n\nsome  time-interval hand annotation for the ebird swaspa . Note that i am not sure if my annotation would be 100% correct:\n\nannotated region : high confident that it should contain the bird \nnon-annotated region :  i cannot identified the bird ... it could be a true negative or a false negative(i.e. miss)",
      "votes": null
    },
    {
      "id": "953663",
      "postDate": "08/01/2020 01:14:32",
      "content": "<p>segmented bird call event data</p>\n\n<p><a href=\"https://zenodo.org/record/1250690#.XyTA8zczbCJ\">https://zenodo.org/record/1250690#.XyTA8zczbCJ</a></p>\n\n<p>(some of the bird species overlap with ours, e.g.  Blue Jay ,Song Sparrow,Great Blue Heron)</p>\n\n<p><a href=\"https://figshare.com/articles/SwampSparrow_luscdb_zip/5625310\">https://figshare.com/articles/SwampSparrow_luscdb_zip/5625310</a>\n<a href=\"https://github.com/timsainb/avgn_paper/blob/vizmerge/notebooks/00.0-download-datasets/1.0-bird-db-download-dataset.ipynb\">https://github.com/timsainb/avgn_paper/blob/vizmerge/notebooks/00.0-download-datasets/1.0-bird-db-download-dataset.ipynb</a></p>",
      "rawMarkdown": "segmented bird call event data\n\nhttps://zenodo.org/record/1250690#.XyTA8zczbCJ\n\n(some of the bird species overlap with ours, e.g.  Blue Jay ,Song Sparrow,Great Blue Heron)\n\nhttps://figshare.com/articles/SwampSparrow_luscdb_zip/5625310\nhttps://github.com/timsainb/avgn_paper/blob/vizmerge/notebooks/00.0-download-datasets/1.0-bird-db-download-dataset.ipynb",
      "votes": null
    },
    {
      "id": "954428",
      "postDate": "08/01/2020 18:09:31",
      "content": "<p><a href=\"https://openreview.net/forum?id=_P9LyJ5pMDb\">https://openreview.net/forum?id=_P9LyJ5pMDb</a>\nUsing Self-Supervised Learning of Birdsong for Downstream Industrial Audio Classification\ndataset and code: <a href=\"https://github.com/SingingData/Birdsong\">https://github.com/SingingData/Birdsong</a></p>",
      "rawMarkdown": "https://openreview.net/forum?id=_P9LyJ5pMDb\nUsing Self-Supervised Learning of Birdsong for Downstream Industrial Audio Classification\ndataset and code: https://github.com/SingingData/Birdsong",
      "votes": null
    },
    {
      "id": "955655",
      "postDate": "08/02/2020 19:31:40",
      "content": "<p>i make the google drive link shareable. i will add more birds later. if you have problem accessing this, please comment here. thanks.</p>",
      "rawMarkdown": "i make the google drive link shareable. i will add more birds later. if you have problem accessing this, please comment here. thanks.",
      "votes": null
    },
    {
      "id": "955660",
      "postDate": "08/02/2020 19:40:30",
      "content": "<p>Thank you for sharing Heng. I have a question. For example, in swaspa/XC116589.Table.1.selections.txt we can see there is an empty interval in [1.814, 7.2129]. Do you think our models will perform better on the private LB if we take 5 seconds from empty segments like that, add it as \"nocall\" target so we have 264+1 total targets, and use it in our training? Thank you</p>",
      "rawMarkdown": "Thank you for sharing Heng. I have a question. For example, in swaspa/XC116589.Table.1.selections.txt we can see there is an empty interval in [1.814, 7.2129]. Do you think our models will perform better on the private LB if we take 5 seconds from empty segments like that, add it as \"nocall\" target so we have 264+1 total targets, and use it in our training? Thank you",
      "votes": null
    },
    {
      "id": "957183",
      "postDate": "08/04/2020 05:56:27",
      "content": "<p>i do not add them to training at first because i cannot be sure if the empty interval is background or not. (Actually, you can add \"background\" label to each wav, i.e. 2 label per class)</p>\n\n<p>rather, i would train with external noise first. then i will pseudo label in later iterations a shown below:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcdb33754345434fd168d50bde6f3f919%2FSelection_037.png?generation=1596520583850968&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "i do not add them to training at first because i cannot be sure if the empty interval is background or not. (Actually, you can add \"background\" label to each wav, i.e. 2 label per class)\n\nrather, i would train with external noise first. then i will pseudo label in later iterations a shown below:\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcdb33754345434fd168d50bde6f3f919%2FSelection_037.png?generation=1596520583850968&amp;alt=media)",
      "votes": null
    },
    {
      "id": "958374",
      "postDate": "08/05/2020 00:31:15",
      "content": "<p>One of my automatic event detection methods for training. see attached code for an illustration.\nyou can use ensemble of methods to create or score the annotations</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0ddfde356c58120ffa76c5769fa0b056%2FSelection_057.png?generation=1596587353555325&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F39700603673e25328a92bec23c16f579%2FSelection_058.png?generation=1596587351298156&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "One of my automatic event detection methods for training. see attached code for an illustration.\nyou can use ensemble of methods to create or score the annotations\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0ddfde356c58120ffa76c5769fa0b056%2FSelection_057.png?generation=1596587353555325&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F39700603673e25328a92bec23c16f579%2FSelection_058.png?generation=1596587351298156&amp;alt=media)",
      "votes": null
    },
    {
      "id": "960606",
      "postDate": "08/06/2020 14:33:48",
      "content": "<p>some insights on birds recording</p>\n\n<p><a href=\"https://www.avisoft.com/tutorials/measuring-sound-parameters-from-the-spectrogram-automatically/\">https://www.avisoft.com/tutorials/measuring-sound-parameters-from-the-spectrogram-automatically/</a>\n<a href=\"https://www.youtube.com/watch?v=9MYKnze7Zaw\">https://www.youtube.com/watch?v=9MYKnze7Zaw</a></p>",
      "rawMarkdown": "some insights on birds recording\n\nhttps://www.avisoft.com/tutorials/measuring-sound-parameters-from-the-spectrogram-automatically/\nhttps://www.youtube.com/watch?v=9MYKnze7Zaw",
      "votes": null
    },
    {
      "id": "960725",
      "postDate": "08/06/2020 16:19:40",
      "content": "<p>could be due to MP3 decoding delay</p>",
      "rawMarkdown": "could be due to MP3 decoding delay",
      "votes": null
    },
    {
      "id": "960900",
      "postDate": "08/06/2020 19:09:05",
      "content": "<p>i just realise that you can use BirdNET to label the clips</p>\n\n<p><a href=\"https://birdnet.cornell.edu/api/\">https://birdnet.cornell.edu/api/</a>\n<a href=\"https://github.com/kahst/BirdNET-Demo\">https://github.com/kahst/BirdNET-Demo</a>\n<a href=\"https://www.wildlabs.net/resources/community-announcements/wildlabs-virtual-meetup-recording-acoustic-monitoring\">https://www.wildlabs.net/resources/community-announcements/wildlabs-virtual-meetup-recording-acoustic-monitoring</a>\ndata and software from Australian Acoustic Observatory: <a href=\"https://data.acousticobservatory.org/listen\">https://data.acousticobservatory.org/listen</a>\n<a href=\"https://ap.qut.ecoacoustics.info/tutorials/01-usingap/practical?tabs=linux\">https://ap.qut.ecoacoustics.info/tutorials/01-usingap/practical?tabs=linux</a></p>",
      "rawMarkdown": "i just realise that you can use BirdNET to label the clips\n\nhttps://birdnet.cornell.edu/api/\nhttps://github.com/kahst/BirdNET-Demo\nhttps://www.wildlabs.net/resources/community-announcements/wildlabs-virtual-meetup-recording-acoustic-monitoring\ndata and software from Australian Acoustic Observatory: https://data.acousticobservatory.org/listen\nhttps://ap.qut.ecoacoustics.info/tutorials/01-usingap/practical?tabs=linux",
      "votes": null
    },
    {
      "id": "969648",
      "postDate": "08/13/2020 20:27:04",
      "content": "<p><a href=\"https://rpubs.com/marcelo-araya-salas/110155\" target=\"_blank\">https://rpubs.com/marcelo-araya-salas/110155</a></p>",
      "rawMarkdown": "https://rpubs.com/marcelo-araya-salas/110155",
      "votes": null
    },
    {
      "id": "971247",
      "postDate": "08/15/2020 10:22:26",
      "content": "<blockquote>\n  <p>one file cannot contains many bird type</p>\n</blockquote>\n<p>Are you referring to the training examples? If so, I understand. Otherwise, I don't understand - aren't the examples in the test set soundscapes containing possibly many different birds?</p>",
      "rawMarkdown": "> one file cannot contains many bird type\n\nAre you referring to the training examples? If so, I understand. Otherwise, I don't understand - aren't the examples in the test set soundscapes containing possibly many different birds?",
      "votes": null
    },
    {
      "id": "971252",
      "postDate": "08/15/2020 10:29:19",
      "content": "<p>you can expect maybe up to 4 to 8 birds. There will be some limit</p>",
      "rawMarkdown": "you can expect maybe up to 4 to 8 birds. There will be some limit",
      "votes": null
    },
    {
      "id": "971260",
      "postDate": "08/15/2020 10:33:35",
      "content": "<p>Ah, I see - you're saying there may be multiple birds, but not too many. Thanks for this and for all your posts here - very interesting resources, and a great way of sharing which (it seems) everyone is welcoming!</p>",
      "rawMarkdown": "Ah, I see - you're saying there may be multiple birds, but not too many. Thanks for this and for all your posts here - very interesting resources, and a great way of sharing which (it seems) everyone is welcoming!",
      "votes": null
    },
    {
      "id": "974496",
      "postDate": "08/18/2020 00:56:55",
      "content": "<p><a href=\"https://drive.google.com/drive/folders/14n_DlU0dGiXZE6Zl_AT78J2h7VZwZ405\" target=\"_blank\">https://drive.google.com/drive/folders/14n_DlU0dGiXZE6Zl_AT78J2h7VZwZ405</a></p>\n<p>pseudo labels results:</p>\n<ul>\n<li>efficientb2 : my trained model</li>\n<li>bird/_net : results predicted from <a href=\"https://github.com/kahst/BirdNET-Electron\" target=\"_blank\">https://github.com/kahst/BirdNET-Electron</a></li>\n<li>allaboutbirds : reference spectrogram from the website. </li>\n</ul>\n<p>this gives you an ideas how much improvement you can get. right now i am making an interface to hand correct these pseudo labels and then retrain the model with better strong labels</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6fa3be73ac4c74a79a09ae9d505e38ff%2FXC179124_20.png?generation=1597712354411244&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "https://drive.google.com/drive/folders/14n_DlU0dGiXZE6Zl_AT78J2h7VZwZ405\n\npseudo labels results:\n- efficientb2 : my trained model\n- bird/_net : results predicted from https://github.com/kahst/BirdNET-Electron\n- allaboutbirds : reference spectrogram from the website. \n\nthis gives you an ideas how much improvement you can get. right now i am making an interface to hand correct these pseudo labels and then retrain the model with better strong labels\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6fa3be73ac4c74a79a09ae9d505e38ff%2FXC179124_20.png?generation=1597712354411244&alt=media)",
      "votes": null
    },
    {
      "id": "981194",
      "postDate": "08/22/2020 08:54:56",
      "content": "<p>there is a time-annotated bird song dataset</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe36174a808fedc5f65bad172126f083c%2FSelection_034.png?generation=1598086493482977&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.sciencedirect.com/science/article/pii/S1574954115000151\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1574954115000151</a></p>\n<p><a href=\"http://taylor0.biology.ucla.edu/birdDBQuery/\" target=\"_blank\">http://taylor0.biology.ucla.edu/birdDBQuery/</a></p>\n<p>(emable Textgrid during search)<br>\nmaybe good for pretraining or event detection or controlled experiments?</p>",
      "rawMarkdown": "there is a time-annotated bird song dataset\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe36174a808fedc5f65bad172126f083c%2FSelection_034.png?generation=1598086493482977&alt=media)\nhttps://www.sciencedirect.com/science/article/pii/S1574954115000151\n\nhttp://taylor0.biology.ucla.edu/birdDBQuery/\n\n(emable Textgrid during search)\nmaybe good for pretraining or event detection or controlled experiments?",
      "votes": null
    },
    {
      "id": "983276",
      "postDate": "08/24/2020 06:38:43",
      "content": "<p>Thank you, this thread definitely helped me realise that I was wasting my time trying to solve the problem without better labeling of the data. I think I have probably realised this too late to put together a good model / submission but still at least I have learnt something and started to put together a model that can have some limited accuracy…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2Fb1d3ab630a397c403e0c5ecb06a8a22f%2Fexample_clipped.png?generation=1598251200823054&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thank you, this thread definitely helped me realise that I was wasting my time trying to solve the problem without better labeling of the data. I think I have probably realised this too late to put together a good model / submission but still at least I have learnt something and started to put together a model that can have some limited accuracy...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2Fb1d3ab630a397c403e0c5ecb06a8a22f%2Fexample_clipped.png?generation=1598251200823054&alt=media)",
      "votes": null
    },
    {
      "id": "983686",
      "postDate": "08/24/2020 14:04:54",
      "content": "<p>pseudo labels from softmax classifier train on samples without secondary labels.<br>\nclass activation map (cam) is used<br>\n(note that even when secondary = None, it is possible that there are still other background birds)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff58923760037923f0ab4054336984d33%2FXC109300_01.png?generation=1598277772831445&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "pseudo labels from softmax classifier train on samples without secondary labels.\nclass activation map (cam) is used\n(note that even when secondary = None, it is possible that there are still other background birds)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff58923760037923f0ab4054336984d33%2FXC109300_01.png?generation=1598277772831445&alt=media)",
      "votes": null
    },
    {
      "id": "983687",
      "postDate": "08/24/2020 14:06:51",
      "content": "<p>the difficulty is how to label  using ML methods automatically or semi automatically. <br>\ntotal label by hands is not practical or not possible.</p>",
      "rawMarkdown": "the difficulty is how to label  using ML methods automatically or semi automatically. \ntotal label by hands is not practical or not possible.",
      "votes": null
    },
    {
      "id": "985715",
      "postDate": "08/26/2020 01:05:03",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2Fab0559dff6661c1a6e0ac1d620b5eef0%2Fexample%20cropped.jpg?generation=1598413253515349&amp;alt=media\" alt=\"\"></p>\n<p>detection of similar calls across multiple audio tracks for a single bird type (note - my classifier is not trained on this bird type, though i did some manual work at earlier stages on a small number of other bird types)</p>\n<p>have to admit i had figured with secondary labels, if we can extract 'clean' (at least fairly clean) samples on the main bird type, can always create our own mix &amp; match examples for model training? and presumably a classifier that has been trained on 'clean' data should in turn be able to help with labeling original audio, if needed. i wasn't sure how critical it was to label all the original clips 100%. any thoughts welcome…</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2Fab0559dff6661c1a6e0ac1d620b5eef0%2Fexample%20cropped.jpg?generation=1598413253515349&alt=media)\n\ndetection of similar calls across multiple audio tracks for a single bird type (note - my classifier is not trained on this bird type, though i did some manual work at earlier stages on a small number of other bird types)\n\nhave to admit i had figured with secondary labels, if we can extract 'clean' (at least fairly clean) samples on the main bird type, can always create our own mix & match examples for model training? and presumably a classifier that has been trained on 'clean' data should in turn be able to help with labeling original audio, if needed. i wasn't sure how critical it was to label all the original clips 100%. any thoughts welcome...",
      "votes": null
    },
    {
      "id": "985886",
      "postDate": "08/26/2020 04:52:06",
      "content": "<p>Extract of some of the most 'similar' calls across multiple audio clips for aldfly</p>\n<p>Again assuming that the most typical call across multiple clips will help avoid 'secondary' birds to get clean data for classification. As the main bird will be in all the clips.</p>\n<p>However won't really solve the song/call mix from the same bird. Suspect I could try to work round to that but likely to be much more complicated and I'm already out of my depth in this comp!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2F142a1796f47bd49b3ce65aad68eb46bc%2Foutput%20example%204.jpg?generation=1598417385182818&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Extract of some of the most 'similar' calls across multiple audio clips for aldfly\n\nAgain assuming that the most typical call across multiple clips will help avoid 'secondary' birds to get clean data for classification. As the main bird will be in all the clips.\n\nHowever won't really solve the song/call mix from the same bird. Suspect I could try to work round to that but likely to be much more complicated and I'm already out of my depth in this comp!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2F142a1796f47bd49b3ce65aad68eb46bc%2Foutput%20example%204.jpg?generation=1598417385182818&alt=media)",
      "votes": null
    },
    {
      "id": "985938",
      "postDate": "08/26/2020 05:35:26",
      "content": "<p>\"have to admit i had figured with secondary labels, if we can extract 'clean' (at least fairly clean) samples on the main bird type, can always create our own mix &amp; match examples for model training? and presumably a classifier that has been trained on 'clean' data should in turn be able to help with labeling original audio, if needed. i wasn't sure how critical it was to label all the original clips 100%.\"</p>\n<p>100% clean data is not required. <br>\neven with 100% clean data there is a bigger issue of adding noise to the data to mimic the soundscape environment.</p>\n<p>i suggest you should quickly make a submission to get baseline results based on what you have now. you should also test your model on the two sample test clip and the given clefbird 2020 clips in the external data thread.</p>",
      "rawMarkdown": "\"have to admit i had figured with secondary labels, if we can extract 'clean' (at least fairly clean) samples on the main bird type, can always create our own mix & match examples for model training? and presumably a classifier that has been trained on 'clean' data should in turn be able to help with labeling original audio, if needed. i wasn't sure how critical it was to label all the original clips 100%.\"\n\n\n100% clean data is not required. \neven with 100% clean data there is a bigger issue of adding noise to the data to mimic the soundscape environment.\n\ni suggest you should quickly make a submission to get baseline results based on what you have now. you should also test your model on the two sample test clip and the given clefbird 2020 clips in the external data thread.",
      "votes": null
    },
    {
      "id": "987920",
      "postDate": "08/27/2020 16:11:09",
      "content": "<p>yet another time annotated dataset:<br>\n<a href=\"https://avocet.integrativebiology.natsci.msu.edu/species/1922\" target=\"_blank\">https://avocet.integrativebiology.natsci.msu.edu/species/1922</a></p>\n<p><img src=\"https://avocet.integrativebiology.natsci.msu.edu/recordings_data/1/169/16983/sonogram.jpg\" alt=\"\"></p>",
      "rawMarkdown": "yet another time annotated dataset:\nhttps://avocet.integrativebiology.natsci.msu.edu/species/1922\n\n![](https://avocet.integrativebiology.natsci.msu.edu/recordings_data/1/169/16983/sonogram.jpg)",
      "votes": null
    },
    {
      "id": "989740",
      "postDate": "08/29/2020 04:48:00",
      "content": "<p>Hey, I looked into Scaper. To me it doesn't seem useful since it requires the sounds it uses to match these descriptions:</p>\n<blockquote>\n  <ul>\n  <li><strong>Background files</strong>: are used to create the background of the soundscape, and<br>\n  should contain audio material that is perceived as a single holistic sound<br>\n  which is more distant, ambiguous, and texture-like (e.g. the \"hum\" or \"drone\"<br>\n  of an urban environment, or \"wind and rain\" sounds in a natural environment).<br>\n  Importantly, background files should not contain salient sound events.</li>\n  <li><strong>Foreground files</strong>: are used to create sound events. Each foreground audio<br>\n  file should contain a single sound event (short or long) such as a car honk,<br>\n  an animal vocalization, continuous speech, a siren or an idling engine.<br>\n  Foreground files should be as clean as possible with no background noise and<br>\n  no silence before/after the sound event.</li>\n  </ul>\n</blockquote>\n<p>And while we have access to some audio that matches the Background Files description, I'm pretty sure none of the audio we have access to matches the Foreground Files description. Am I missing something?</p>\n<p>Also, you outlined a really cool procedure for automatic event annotation and I want to try it myself. Did you use any particular library for determining Signal to Noise Ratio?</p>\n<p>Thanks for all your hard work in any case.</p>",
      "rawMarkdown": "Hey, I looked into Scaper. To me it doesn't seem useful since it requires the sounds it uses to match these descriptions:\n\n> * **Background files**: are used to create the background of the soundscape, and\n  should contain audio material that is perceived as a single holistic sound\n  which is more distant, ambiguous, and texture-like (e.g. the \"hum\" or \"drone\"\n  of an urban environment, or \"wind and rain\" sounds in a natural environment).\n  Importantly, background files should not contain salient sound events.\n* **Foreground files**: are used to create sound events. Each foreground audio\n  file should contain a single sound event (short or long) such as a car honk,\n  an animal vocalization, continuous speech, a siren or an idling engine.\n  Foreground files should be as clean as possible with no background noise and\n  no silence before/after the sound event.\n\nAnd while we have access to some audio that matches the Background Files description, I'm pretty sure none of the audio we have access to matches the Foreground Files description. Am I missing something?\n\nAlso, you outlined a really cool procedure for automatic event annotation and I want to try it myself. Did you use any particular library for determining Signal to Noise Ratio?\n\nThanks for all your hard work in any case.",
      "votes": null
    },
    {
      "id": "989844",
      "postDate": "08/29/2020 06:35:59",
      "content": "<p>at the moment I'm using </p>\n<pre><code>def signaltonoise(a, axis=0, ddof=0):\n    a = np.asanyarray(a)\n    return a.mean(axis)/a.std(axis=axis, ddof=ddof)\n\ndef segmentwise_snr(wave, segment_length):\n    snrs = []\n    start = 0\n    while start &lt; len(wave):\n        snrs.append(signaltonoise(wave[start:start+segment_length]))\n        start += segment_length\n    return snrs\n\ndef snr_windows(wave, segment_length, threshold, segment_windows):\n    segments = segmentwise_snr(wave, segment_length)\n\n    window_scores = {}\n    for w in segment_windows:\n        scores = []\n        start = 0\n        while start &lt; len(segments):\n            scores.append(np.mean(segments[start:start+w]) &gt; threshold)\n            start += 1\n        window_scores[w] = scores\n    return window_scores\n\nf = snr_windows(wave, int(sr / 8), 0.0008, [2, 5, 25])\n</code></pre>\n<p>But the results are utter garbage. Any advice for creating a rough automatic SNR-based bird-call event detector?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2Fe9753e01ca7fdc35cf3015b094460a5c%2FScreenshot%20from%202020-08-29%2017-07-56.png?generation=1598684960612534&amp;alt=media\" alt=\"\"> </p>",
      "rawMarkdown": "at the moment I'm using \n\n```\ndef signaltonoise(a, axis=0, ddof=0):\n    a = np.asanyarray(a)\n    return a.mean(axis)/a.std(axis=axis, ddof=ddof)\n\ndef segmentwise_snr(wave, segment_length):\n    snrs = []\n    start = 0\n    while start < len(wave):\n        snrs.append(signaltonoise(wave[start:start+segment_length]))\n        start += segment_length\n    return snrs\n\ndef snr_windows(wave, segment_length, threshold, segment_windows):\n    segments = segmentwise_snr(wave, segment_length)\n    \n    window_scores = {}\n    for w in segment_windows:\n        scores = []\n        start = 0\n        while start < len(segments):\n            scores.append(np.mean(segments[start:start+w]) > threshold)\n            start += 1\n        window_scores[w] = scores\n    return window_scores\n\nf = snr_windows(wave, int(sr / 8), 0.0008, [2, 5, 25])\n```\n\nBut the results are utter garbage. Any advice for creating a rough automatic SNR-based bird-call event detector?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2Fe9753e01ca7fdc35cf3015b094460a5c%2FScreenshot%20from%202020-08-29%2017-07-56.png?generation=1598684960612534&alt=media)",
      "votes": null
    },
    {
      "id": "992122",
      "postDate": "08/31/2020 00:53:37",
      "content": "<p>removing other birds help!</p>\n<p>i was investigating why i cannot detect brncre in the sample test audio. i did several noise augmentation and also download external data. i check visually that there is some training set + augmentation that resemble the test sample.</p>\n<p>it turns out that the problem is presence of other birds.<br>\n(this is results from trained softmax classifier)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc3dec9301bc2641e9319d77c4bda92bd%2FSelection_033.png?generation=1598835210055130&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "removing other birds help!\n\ni was investigating why i cannot detect brncre in the sample test audio. i did several noise augmentation and also download external data. i check visually that there is some training set + augmentation that resemble the test sample.\n\nit turns out that the problem is presence of other birds.\n(this is results from trained softmax classifier)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc3dec9301bc2641e9319d77c4bda92bd%2FSelection_033.png?generation=1598835210055130&alt=media)",
      "votes": null
    },
    {
      "id": "992406",
      "postDate": "08/31/2020 06:19:21",
      "content": "<p>Thanks for the info! This is really useful. Have you tried multi-label classification with sigmoid? The model can maybe learn from multiple birdcalls in the spectrogram</p>",
      "rawMarkdown": "Thanks for the info! This is really useful. Have you tried multi-label classification with sigmoid? The model can maybe learn from multiple birdcalls in the spectrogram",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969648,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/13/2020 20:27:04",
      "content": "<p><a href=\"https://rpubs.com/marcelo-araya-salas/110155\" target=\"_blank\">https://rpubs.com/marcelo-araya-salas/110155</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974496,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/18/2020 00:56:55",
      "content": "<p><a href=\"https://drive.google.com/drive/folders/14n_DlU0dGiXZE6Zl_AT78J2h7VZwZ405\" target=\"_blank\">https://drive.google.com/drive/folders/14n_DlU0dGiXZE6Zl_AT78J2h7VZwZ405</a></p>\n<p>pseudo labels results:</p>\n<ul>\n<li>efficientb2 : my trained model</li>\n<li>bird/_net : results predicted from <a href=\"https://github.com/kahst/BirdNET-Electron\" target=\"_blank\">https://github.com/kahst/BirdNET-Electron</a></li>\n<li>allaboutbirds : reference spectrogram from the website. </li>\n</ul>\n<p>this gives you an ideas how much improvement you can get. right now i am making an interface to hand correct these pseudo labels and then retrain the model with better strong labels</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6fa3be73ac4c74a79a09ae9d505e38ff%2FXC179124_20.png?generation=1597712354411244&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 981194,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/22/2020 08:54:56",
      "content": "<p>there is a time-annotated bird song dataset</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe36174a808fedc5f65bad172126f083c%2FSelection_034.png?generation=1598086493482977&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.sciencedirect.com/science/article/pii/S1574954115000151\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1574954115000151</a></p>\n<p><a href=\"http://taylor0.biology.ucla.edu/birdDBQuery/\" target=\"_blank\">http://taylor0.biology.ucla.edu/birdDBQuery/</a></p>\n<p>(emable Textgrid during search)<br>\nmaybe good for pretraining or event detection or controlled experiments?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 983276,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "08/24/2020 06:38:43",
      "content": "<p>Thank you, this thread definitely helped me realise that I was wasting my time trying to solve the problem without better labeling of the data. I think I have probably realised this too late to put together a good model / submission but still at least I have learnt something and started to put together a model that can have some limited accuracy…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2Fb1d3ab630a397c403e0c5ecb06a8a22f%2Fexample_clipped.png?generation=1598251200823054&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 983687,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/24/2020 14:06:51",
          "content": "<p>the difficulty is how to label  using ML methods automatically or semi automatically. <br>\ntotal label by hands is not practical or not possible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 985715,
          "author_name": "davidedwards1",
          "author_url": "",
          "post_date": "08/26/2020 01:05:03",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2Fab0559dff6661c1a6e0ac1d620b5eef0%2Fexample%20cropped.jpg?generation=1598413253515349&amp;alt=media\" alt=\"\"></p>\n<p>detection of similar calls across multiple audio tracks for a single bird type (note - my classifier is not trained on this bird type, though i did some manual work at earlier stages on a small number of other bird types)</p>\n<p>have to admit i had figured with secondary labels, if we can extract 'clean' (at least fairly clean) samples on the main bird type, can always create our own mix &amp; match examples for model training? and presumably a classifier that has been trained on 'clean' data should in turn be able to help with labeling original audio, if needed. i wasn't sure how critical it was to label all the original clips 100%. any thoughts welcome…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 985886,
          "author_name": "davidedwards1",
          "author_url": "",
          "post_date": "08/26/2020 04:52:06",
          "content": "<p>Extract of some of the most 'similar' calls across multiple audio clips for aldfly</p>\n<p>Again assuming that the most typical call across multiple clips will help avoid 'secondary' birds to get clean data for classification. As the main bird will be in all the clips.</p>\n<p>However won't really solve the song/call mix from the same bird. Suspect I could try to work round to that but likely to be much more complicated and I'm already out of my depth in this comp!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2F142a1796f47bd49b3ce65aad68eb46bc%2Foutput%20example%204.jpg?generation=1598417385182818&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 985938,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/26/2020 05:35:26",
          "content": "<p>\"have to admit i had figured with secondary labels, if we can extract 'clean' (at least fairly clean) samples on the main bird type, can always create our own mix &amp; match examples for model training? and presumably a classifier that has been trained on 'clean' data should in turn be able to help with labeling original audio, if needed. i wasn't sure how critical it was to label all the original clips 100%.\"</p>\n<p>100% clean data is not required. <br>\neven with 100% clean data there is a bigger issue of adding noise to the data to mimic the soundscape environment.</p>\n<p>i suggest you should quickly make a submission to get baseline results based on what you have now. you should also test your model on the two sample test clip and the given clefbird 2020 clips in the external data thread.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 983686,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/24/2020 14:04:54",
      "content": "<p>pseudo labels from softmax classifier train on samples without secondary labels.<br>\nclass activation map (cam) is used<br>\n(note that even when secondary = None, it is possible that there are still other background birds)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff58923760037923f0ab4054336984d33%2FXC109300_01.png?generation=1598277772831445&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 987920,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/27/2020 16:11:09",
      "content": "<p>yet another time annotated dataset:<br>\n<a href=\"https://avocet.integrativebiology.natsci.msu.edu/species/1922\" target=\"_blank\">https://avocet.integrativebiology.natsci.msu.edu/species/1922</a></p>\n<p><img src=\"https://avocet.integrativebiology.natsci.msu.edu/recordings_data/1/169/16983/sonogram.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 989740,
      "author_name": "lewington",
      "author_url": "",
      "post_date": "08/29/2020 04:48:00",
      "content": "<p>Hey, I looked into Scaper. To me it doesn't seem useful since it requires the sounds it uses to match these descriptions:</p>\n<blockquote>\n  <ul>\n  <li><strong>Background files</strong>: are used to create the background of the soundscape, and<br>\n  should contain audio material that is perceived as a single holistic sound<br>\n  which is more distant, ambiguous, and texture-like (e.g. the \"hum\" or \"drone\"<br>\n  of an urban environment, or \"wind and rain\" sounds in a natural environment).<br>\n  Importantly, background files should not contain salient sound events.</li>\n  <li><strong>Foreground files</strong>: are used to create sound events. Each foreground audio<br>\n  file should contain a single sound event (short or long) such as a car honk,<br>\n  an animal vocalization, continuous speech, a siren or an idling engine.<br>\n  Foreground files should be as clean as possible with no background noise and<br>\n  no silence before/after the sound event.</li>\n  </ul>\n</blockquote>\n<p>And while we have access to some audio that matches the Background Files description, I'm pretty sure none of the audio we have access to matches the Foreground Files description. Am I missing something?</p>\n<p>Also, you outlined a really cool procedure for automatic event annotation and I want to try it myself. Did you use any particular library for determining Signal to Noise Ratio?</p>\n<p>Thanks for all your hard work in any case.</p>",
      "votes": null,
      "replies": [
        {
          "id": 989844,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/29/2020 06:35:59",
          "content": "<p>at the moment I'm using </p>\n<pre><code>def signaltonoise(a, axis=0, ddof=0):\n    a = np.asanyarray(a)\n    return a.mean(axis)/a.std(axis=axis, ddof=ddof)\n\ndef segmentwise_snr(wave, segment_length):\n    snrs = []\n    start = 0\n    while start &lt; len(wave):\n        snrs.append(signaltonoise(wave[start:start+segment_length]))\n        start += segment_length\n    return snrs\n\ndef snr_windows(wave, segment_length, threshold, segment_windows):\n    segments = segmentwise_snr(wave, segment_length)\n\n    window_scores = {}\n    for w in segment_windows:\n        scores = []\n        start = 0\n        while start &lt; len(segments):\n            scores.append(np.mean(segments[start:start+w]) &gt; threshold)\n            start += 1\n        window_scores[w] = scores\n    return window_scores\n\nf = snr_windows(wave, int(sr / 8), 0.0008, [2, 5, 25])\n</code></pre>\n<p>But the results are utter garbage. Any advice for creating a rough automatic SNR-based bird-call event detector?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2Fe9753e01ca7fdc35cf3015b094460a5c%2FScreenshot%20from%202020-08-29%2017-07-56.png?generation=1598684960612534&amp;alt=media\" alt=\"\"> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 992122,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/31/2020 00:53:37",
      "content": "<p>removing other birds help!</p>\n<p>i was investigating why i cannot detect brncre in the sample test audio. i did several noise augmentation and also download external data. i check visually that there is some training set + augmentation that resemble the test sample.</p>\n<p>it turns out that the problem is presence of other birds.<br>\n(this is results from trained softmax classifier)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc3dec9301bc2641e9319d77c4bda92bd%2FSelection_033.png?generation=1598835210055130&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 992406,
          "author_name": "alanchn31",
          "author_url": "",
          "post_date": "08/31/2020 06:19:21",
          "content": "<p>Thanks for the info! This is really useful. Have you tried multi-label classification with sigmoid? The model can maybe learn from multiple birdcalls in the spectrogram</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 943130,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "07/24/2020 06:54:53",
      "content": "<p>Thanks for starting the thread!\nSo far I tried to use the <a href=\"https://www.kaggle.com/c/mlsp-2013-birds\">MLSP 2013 Bird Classification Challenge</a> data . I will create a kaggle dataset with useful labels.</p>\n\n<p><strong>Pros</strong></p>\n\n<ul>\n<li>645 x 10s soundscapes recorded in Oregon</li>\n<li>multiclass-multilabel dataset</li>\n</ul>\n\n<p><strong>Cons</strong></p>\n\n<ul>\n<li>Only 19 bird species</li>\n<li>16 kHz wav files</li>\n<li>7 year old recordings</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 943526,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2020 12:15:22",
      "content": "<p>for a start, please refer to these slides<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd322e9eb5cf653be8cdfa1a0dc6ee21f%2FSlide1.png?generation=1595592879332532&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe5bd1ac4f56de723b523b51108eb5b1f%2FSlide2.png?generation=1595592896697819&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F873f3596985621ed74201b7a345c70df%2FSlide4.png?generation=1595592918527139&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 971247,
          "author_name": "marcogorelli",
          "author_url": "",
          "post_date": "08/15/2020 10:22:26",
          "content": "<blockquote>\n  <p>one file cannot contains many bird type</p>\n</blockquote>\n<p>Are you referring to the training examples? If so, I understand. Otherwise, I don't understand - aren't the examples in the test set soundscapes containing possibly many different birds?</p>",
          "votes": null,
          "replies": [
            {
              "id": 971252,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "08/15/2020 10:29:19",
              "content": "<p>you can expect maybe up to 4 to 8 birds. There will be some limit</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 971260,
              "author_name": "marcogorelli",
              "author_url": "",
              "post_date": "08/15/2020 10:33:35",
              "content": "<p>Ah, I see - you're saying there may be multiple birds, but not too many. Thanks for this and for all your posts here - very interesting resources, and a great way of sharing which (it seems) everyone is welcoming!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 943619,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2020 13:25:40",
      "content": "<p>is this a bug ??<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F59995c501cc1d135f011b89ff05bfbdf%2FSelection_026.png?generation=1595597138393031&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 960725,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "08/06/2020 16:19:40",
          "content": "<p>could be due to MP3 decoding delay</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 944112,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2020 20:53:33",
      "content": "<p>i attempt to visualize the annotations of BirdCLEF 2020 provided by the competition host at <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877\">https://www.kaggle.com/c/birdsong-recognition/discussion/158877</a></p>\n\n<p>they are very difficult! i think we may need to use train data with \"the lowest rating\" from <a href=\"https://ebird.org/media\">https://ebird.org/media</a>, etc. The quality of the train data we have is very high, compared to BirdCLEF 2020. I am not sure what we would be expecting for the test here ... anyway, do prepared for very bad test data.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F6a39ce2151aeb57d0b36c72be7f7156b%2FSelection_034.png?generation=1595623924948889&amp;alt=media\" alt=\"\"></p>\n\n<p>attached : raven-lite selection (annotation) file for  SSW51_20170819.wav</p>",
      "votes": null,
      "replies": [
        {
          "id": 946157,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/26/2020 11:58:53",
          "content": "<p>i manged to find the missing bird in the label. I use audcity noise reduction to process the audio, here is the target bird in annotation:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F55c3cc1076bd8e57b719f69b9fb5cfb5%2FSelection_027.png?generation=1595764731183429&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://www.macaulaylibrary.org/resources/audio-editing-tutorials/\">https://www.macaulaylibrary.org/resources/audio-editing-tutorials/</a>\nAudio editing tutorials for birdsong</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 944246,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/25/2020 00:44:07",
      "content": "<p>check this work, it should be the baseline method for this challenge: \n<a href=\"http://ceur-ws.org/Vol-2380/paper_86.pdf\">http://ceur-ws.org/Vol-2380/paper_86.pdf</a></p>\n\n<p>Bird Species Identification in Soundscapes - Mario Lasseck</p>\n\n<p>\"Deep Convolutional Neural Networks are trained to classify 659 species. Different data augmentation techniques are applied to prevent overfitting and improve model accuracy and generalization. The proposed approach is evaluated in the BirdCLEF 2019 campaign and provides the best system to identify bird species in wildlife monitoring recordings. \"</p>\n\n<hr>\n\n<p>another good baseline\n<a href=\"https://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html\">https://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html</a></p>\n\n<p>Acoustic Detection of Humpback Whales Using a Convolutional Neural Network\n<a href=\"https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1e442a65981435576a6ac33e4c3178d6a62a06a1.pdf\">https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1e442a65981435576a6ac33e4c3178d6a62a06a1.pdf</a></p>\n\n<p><a href=\"https://medium.com/@kcimc/data-of-the-humpback-whale-9ef09c5920cd\">https://medium.com/@kcimc/data-of-the-humpback-whale-9ef09c5920cd</a></p>\n\n<p>\"Long-distance detection of bioacoustic events with per-channel energy normalization\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 944283,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/25/2020 01:57:20",
      "content": "<p>type of noise in bird recording:\n\"Birdsong Denoising Using Wavelets\" : <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4728069/\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4728069/</a></p>\n\n<p>it has a software:\n<a href=\"http://www.avianz.net/index.php/avianz-software/user-manual\">http://www.avianz.net/index.php/avianz-software/user-manual</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 944706,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/25/2020 09:23:07",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=8xH2GjHKYj0\">https://www.youtube.com/watch?v=8xH2GjHKYj0</a></p>\n\n<p>\"Bird Song Hero: The song learning game for everyone\"\nCornell Lab of Ornithology</p>\n\n<p>learn how to use spectrogram to identify birds. seems that \"counts\" of the repeated tune is sometimes a feature</p>\n\n<p>related: \n<a href=\"https://www.youtube.com/watch?v=qp34UfXY_Wk\">https://www.youtube.com/watch?v=qp34UfXY_Wk</a>\n<a href=\"https://www.youtube.com/watch?v=04HpCmjlQs8\">https://www.youtube.com/watch?v=04HpCmjlQs8</a>\n<a href=\"https://www.youtube.com/watch?v=4_1zIwEENt8\">https://www.youtube.com/watch?v=4_1zIwEENt8</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 944732,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/25/2020 09:57:42",
      "content": "<p>Augmentation: this shows how ta real spectrogram changes.</p>\n\n<p><a href=\"http://earbirding.com/blog/archives/category/spectrograms\">http://earbirding.com/blog/archives/category/spectrograms</a></p>\n\n<p><img src=\"http://earbirding.com/blog/wp-content/uploads/2013/02/VESPslow.gif\" alt=\"\"></p>\n\n<p>dynamic-time-warping  seems to be applicable for augmentation (or even clustering)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 944739,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/25/2020 10:07:20",
      "content": "<p><a href=\"https://experiments.withgoogle.com/ai/bird-sounds/view/\">https://experiments.withgoogle.com/ai/bird-sounds/view/</a>\n<a href=\"https://github.com/googlecreativelab/aiexperiments-bird-sounds\">https://github.com/googlecreativelab/aiexperiments-bird-sounds</a></p>\n\n<p>t-SNE visualization of \"Essential Set for North America\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 946854,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/26/2020 22:34:30",
      "content": "<p><a href=\"https://github.com/CrowdCurio/audio-annotator\">https://github.com/CrowdCurio/audio-annotator</a>\naudio-annotator is a web interface that allows users to annotate audio recordings.\n<img src=\"https://github.com/CrowdCurio/audio-annotator/raw/master/static/img/task-interface.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 946877,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/26/2020 23:14:43",
      "content": "<p><a href=\"https://github.com/justinsalamon/scaper\">https://github.com/justinsalamon/scaper</a></p>\n\n<p>Scaper: A library for soundscape synthesis and augmentation\n<a href=\"https://www.youtube.com/watch?v=zvccOFz2KxI\">https://www.youtube.com/watch?v=zvccOFz2KxI</a></p>\n\n<p>\" Scaper, an open-source library for soundscape synthesis and augmentation. Given a collection of isolated sound events, Scaper acts as a high-level sequencer that can generate multiple soundscapes from a single, probabilistically defined, “specification”. \"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 948536,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "07/28/2020 03:41:41",
      "content": "<p>waiting patiently :) Thank you for the work as always</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 949811,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/29/2020 00:57:51",
      "content": "<p>i spends some days to hand annotate the time interval for the bird calls. Here is what I find:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F886f6b16839e952db30a1263ce52f437%2FSelection_036.png?generation=1595984240703288&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F64b6272f907bb77ce631e3cee79c3a62%2FSelection_037.png?generation=1595984239768601&amp;alt=media\" alt=\"\"></p>\n\n<p>in short, it is not only weak label, but also noisy label !</p>\n\n<p>also, if an interval is being process, it is likely to contain the bird (hence there is some time interval annotation)</p>",
      "votes": null,
      "replies": [
        {
          "id": 949992,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/29/2020 05:33:13",
          "content": "<ol>\n<li>Two-thirds of the clips don't have secondary labels, however, some of those actually contain calls of other species.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 950516,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "07/29/2020 12:51:43",
          "content": "<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> How to deal with labels that don't contain secondary labels? Do you try to output all 0's in a secondary head, or just don't compute the loss for those rows? ty</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951093,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/29/2020 21:44:44",
          "content": "<p>So far I make my model output all 0's for <em>potential</em> secondary labels. I don't clearly separate primary / secondary labels, so I provide 264 dimension one-hot vector for each sample and put 1 to corresponding positions for primary / secondary labels. If the sample does not have secondary labels, then I make that vector whose elements are all 0 except for the position that corresponds to primary label. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 952597,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2020 04:30:51",
          "content": "<p>here is additional observation\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F03bde33e15603e1e5ba15159cceaad0f%2FSelection_048.png?generation=1596169848623517&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 953188,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "07/31/2020 15:35:46",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  would you mind sharing which file is this ? I wonder if mp3 would have switched to short window during these periods when higher time resolution is needed </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 953272,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2020 16:37:35",
          "content": "<p>see swamp sparrow (swaspa) : \n6kHz (window size 256): XC446809, XC116589, XC131033\n9kHz (window size 128): XC138153 \nthis site provide reference call files\n<code>\nhttps://www.audubon.org/field-guide/bird/swamp-sparrow\nSongs and Calls\nSweet, musical trill, all on one note.\n fast pulse-rate song\n slow pulse-rate song\n very slow pulse-rate song\n odd buzzy song\n chips #1\n chips #2\n</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 952600,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2020 04:34:14",
      "content": "<p>if you want to collect your your data, consider microphone array. just like 3d vision, triangulation using microphone array can localized different bird call. e.g.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F610fe3a6fa1a7c5b81c5e226992d9780%2FSelection_047.png?generation=1596169986819962&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://github.com/HARKBird-project/HARKBird\">https://github.com/HARKBird-project/HARKBird</a>\n<a href=\"https://sites.google.com/view/alcore-suzuki/home/harkbird\">https://sites.google.com/view/alcore-suzuki/home/harkbird</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953298,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2020 17:08:36",
      "content": "<p><a href=\"https://drive.google.com/drive/folders/1SExd8V-vFfTRrmTgjJ5PXN34TFmCx3Hd?usp=sharing\">https://drive.google.com/drive/folders/1SExd8V-vFfTRrmTgjJ5PXN34TFmCx3Hd?usp=sharing</a></p>\n\n<p>some  time-interval hand annotation for the ebird swaspa . Note that i am not sure if my annotation would be 100% correct:</p>\n\n<p>annotated region : high confident that it should contain the bird \nnon-annotated region :  i cannot identified the bird ... it could be a true negative or a false negative(i.e. miss)</p>",
      "votes": null,
      "replies": [
        {
          "id": 955655,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/02/2020 19:31:40",
          "content": "<p>i make the google drive link shareable. i will add more birds later. if you have problem accessing this, please comment here. thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955660,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "08/02/2020 19:40:30",
          "content": "<p>Thank you for sharing Heng. I have a question. For example, in swaspa/XC116589.Table.1.selections.txt we can see there is an empty interval in [1.814, 7.2129]. Do you think our models will perform better on the private LB if we take 5 seconds from empty segments like that, add it as \"nocall\" target so we have 264+1 total targets, and use it in our training? Thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 957183,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/04/2020 05:56:27",
          "content": "<p>i do not add them to training at first because i cannot be sure if the empty interval is background or not. (Actually, you can add \"background\" label to each wav, i.e. 2 label per class)</p>\n\n<p>rather, i would train with external noise first. then i will pseudo label in later iterations a shown below:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcdb33754345434fd168d50bde6f3f919%2FSelection_037.png?generation=1596520583850968&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 953663,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/01/2020 01:14:32",
      "content": "<p>segmented bird call event data</p>\n\n<p><a href=\"https://zenodo.org/record/1250690#.XyTA8zczbCJ\">https://zenodo.org/record/1250690#.XyTA8zczbCJ</a></p>\n\n<p>(some of the bird species overlap with ours, e.g.  Blue Jay ,Song Sparrow,Great Blue Heron)</p>\n\n<p><a href=\"https://figshare.com/articles/SwampSparrow_luscdb_zip/5625310\">https://figshare.com/articles/SwampSparrow_luscdb_zip/5625310</a>\n<a href=\"https://github.com/timsainb/avgn_paper/blob/vizmerge/notebooks/00.0-download-datasets/1.0-bird-db-download-dataset.ipynb\">https://github.com/timsainb/avgn_paper/blob/vizmerge/notebooks/00.0-download-datasets/1.0-bird-db-download-dataset.ipynb</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 954428,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/01/2020 18:09:31",
      "content": "<p><a href=\"https://openreview.net/forum?id=_P9LyJ5pMDb\">https://openreview.net/forum?id=_P9LyJ5pMDb</a>\nUsing Self-Supervised Learning of Birdsong for Downstream Industrial Audio Classification\ndataset and code: <a href=\"https://github.com/SingingData/Birdsong\">https://github.com/SingingData/Birdsong</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 958374,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/05/2020 00:31:15",
      "content": "<p>One of my automatic event detection methods for training. see attached code for an illustration.\nyou can use ensemble of methods to create or score the annotations</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0ddfde356c58120ffa76c5769fa0b056%2FSelection_057.png?generation=1596587353555325&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F39700603673e25328a92bec23c16f579%2FSelection_058.png?generation=1596587351298156&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 960606,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/06/2020 14:33:48",
      "content": "<p>some insights on birds recording</p>\n\n<p><a href=\"https://www.avisoft.com/tutorials/measuring-sound-parameters-from-the-spectrogram-automatically/\">https://www.avisoft.com/tutorials/measuring-sound-parameters-from-the-spectrogram-automatically/</a>\n<a href=\"https://www.youtube.com/watch?v=9MYKnze7Zaw\">https://www.youtube.com/watch?v=9MYKnze7Zaw</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 960900,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/06/2020 19:09:05",
      "content": "<p>i just realise that you can use BirdNET to label the clips</p>\n\n<p><a href=\"https://birdnet.cornell.edu/api/\">https://birdnet.cornell.edu/api/</a>\n<a href=\"https://github.com/kahst/BirdNET-Demo\">https://github.com/kahst/BirdNET-Demo</a>\n<a href=\"https://www.wildlabs.net/resources/community-announcements/wildlabs-virtual-meetup-recording-acoustic-monitoring\">https://www.wildlabs.net/resources/community-announcements/wildlabs-virtual-meetup-recording-acoustic-monitoring</a>\ndata and software from Australian Acoustic Observatory: <a href=\"https://data.acousticobservatory.org/listen\">https://data.acousticobservatory.org/listen</a>\n<a href=\"https://ap.qut.ecoacoustics.info/tutorials/01-usingap/practical?tabs=linux\">https://ap.qut.ecoacoustics.info/tutorials/01-usingap/practical?tabs=linux</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "943089": "[deleted]\n\nplease wait ... \n\n- i am creating a procedure to read and playback annotation like \"example_test_audio_summary.csv\" using Raven+R+python ...\n\n- i realize i make a mistake and is fixing it now. \n\n- the post will be updated later",
    "943130": "Thanks for starting the thread!\nSo far I tried to use the [MLSP 2013 Bird Classification Challenge](https://www.kaggle.com/c/mlsp-2013-birds) data . I will create a kaggle dataset with useful labels.\n\n**Pros**\n\n- 645 x 10s soundscapes recorded in Oregon\n- multiclass-multilabel dataset\n\n**Cons**\n\n- Only 19 bird species\n- 16 kHz wav files\n- 7 year old recordings",
    "943526": "for a start, please refer to these slides![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd322e9eb5cf653be8cdfa1a0dc6ee21f%2FSlide1.png?generation=1595592879332532&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe5bd1ac4f56de723b523b51108eb5b1f%2FSlide2.png?generation=1595592896697819&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F873f3596985621ed74201b7a345c70df%2FSlide4.png?generation=1595592918527139&amp;alt=media)",
    "943619": "is this a bug ??![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F59995c501cc1d135f011b89ff05bfbdf%2FSelection_026.png?generation=1595597138393031&amp;alt=media)",
    "944112": "i attempt to visualize the annotations of BirdCLEF 2020 provided by the competition host at https://www.kaggle.com/c/birdsong-recognition/discussion/158877\n\nthey are very difficult! i think we may need to use train data with \"the lowest rating\" from https://ebird.org/media, etc. The quality of the train data we have is very high, compared to BirdCLEF 2020. I am not sure what we would be expecting for the test here ... anyway, do prepared for very bad test data.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F6a39ce2151aeb57d0b36c72be7f7156b%2FSelection_034.png?generation=1595623924948889&amp;alt=media)\n\nattached : raven-lite selection (annotation) file for  SSW51_20170819.wav",
    "944246": "check this work, it should be the baseline method for this challenge: \nhttp://ceur-ws.org/Vol-2380/paper_86.pdf\n\nBird Species Identification in Soundscapes - Mario Lasseck\n\n\"Deep Convolutional Neural Networks are trained to classify 659 species. Different data augmentation techniques are applied to prevent overfitting and improve model accuracy and generalization. The proposed approach is evaluated in the BirdCLEF 2019 campaign and provides the best system to identify bird species in wildlife monitoring recordings. \"\n\n---\nanother good baseline\nhttps://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html\n\nAcoustic Detection of Humpback Whales Using a Convolutional Neural Network\nhttps://storage.googleapis.com/pub-tools-public-publication-data/pdf/1e442a65981435576a6ac33e4c3178d6a62a06a1.pdf\n\nhttps://medium.com/@kcimc/data-of-the-humpback-whale-9ef09c5920cd\n\n\"Long-distance detection of bioacoustic events with per-channel energy normalization\"",
    "944283": "type of noise in bird recording:\n\"Birdsong Denoising Using Wavelets\" : https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4728069/\n\nit has a software:\nhttp://www.avianz.net/index.php/avianz-software/user-manual",
    "944706": "https://www.youtube.com/watch?v=8xH2GjHKYj0\n \n\"Bird Song Hero: The song learning game for everyone\"\nCornell Lab of Ornithology\n\nlearn how to use spectrogram to identify birds. seems that \"counts\" of the repeated tune is sometimes a feature\n \nrelated: \nhttps://www.youtube.com/watch?v=qp34UfXY_Wk\nhttps://www.youtube.com/watch?v=04HpCmjlQs8\nhttps://www.youtube.com/watch?v=4_1zIwEENt8",
    "944732": "Augmentation: this shows how ta real spectrogram changes.\n\nhttp://earbirding.com/blog/archives/category/spectrograms\n\n![](http://earbirding.com/blog/wp-content/uploads/2013/02/VESPslow.gif)\n\ndynamic-time-warping  seems to be applicable for augmentation (or even clustering)",
    "944739": "https://experiments.withgoogle.com/ai/bird-sounds/view/\nhttps://github.com/googlecreativelab/aiexperiments-bird-sounds\n\nt-SNE visualization of \"Essential Set for North America\"",
    "946157": "i manged to find the missing bird in the label. I use audcity noise reduction to process the audio, here is the target bird in annotation:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F55c3cc1076bd8e57b719f69b9fb5cfb5%2FSelection_027.png?generation=1595764731183429&amp;alt=media)\n\n\nhttps://www.macaulaylibrary.org/resources/audio-editing-tutorials/\nAudio editing tutorials for birdsong",
    "946854": "https://github.com/CrowdCurio/audio-annotator\naudio-annotator is a web interface that allows users to annotate audio recordings.\n![](https://github.com/CrowdCurio/audio-annotator/raw/master/static/img/task-interface.png)",
    "946877": "https://github.com/justinsalamon/scaper\n\nScaper: A library for soundscape synthesis and augmentation\nhttps://www.youtube.com/watch?v=zvccOFz2KxI\n\n\" Scaper, an open-source library for soundscape synthesis and augmentation. Given a collection of isolated sound events, Scaper acts as a high-level sequencer that can generate multiple soundscapes from a single, probabilistically defined, “specification”. \"",
    "948536": "waiting patiently :) Thank you for the work as always",
    "949811": "i spends some days to hand annotate the time interval for the bird calls. Here is what I find:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F886f6b16839e952db30a1263ce52f437%2FSelection_036.png?generation=1595984240703288&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F64b6272f907bb77ce631e3cee79c3a62%2FSelection_037.png?generation=1595984239768601&amp;alt=media)\n\nin short, it is not only weak label, but also noisy label !\n\nalso, if an interval is being process, it is likely to contain the bird (hence there is some time interval annotation)",
    "949992": "5. Two-thirds of the clips don't have secondary labels, however, some of those actually contain calls of other species.",
    "950516": "hidehisaarai1213 How to deal with labels that don't contain secondary labels? Do you try to output all 0's in a secondary head, or just don't compute the loss for those rows? ty",
    "951093": "So far I make my model output all 0's for *potential* secondary labels. I don't clearly separate primary / secondary labels, so I provide 264 dimension one-hot vector for each sample and put 1 to corresponding positions for primary / secondary labels. If the sample does not have secondary labels, then I make that vector whose elements are all 0 except for the position that corresponds to primary label.",
    "952597": "here is additional observation\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F03bde33e15603e1e5ba15159cceaad0f%2FSelection_048.png?generation=1596169848623517&amp;alt=media)",
    "952600": "if you want to collect your your data, consider microphone array. just like 3d vision, triangulation using microphone array can localized different bird call. e.g.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F610fe3a6fa1a7c5b81c5e226992d9780%2FSelection_047.png?generation=1596169986819962&amp;alt=media)\n\nhttps://github.com/HARKBird-project/HARKBird\nhttps://sites.google.com/view/alcore-suzuki/home/harkbird",
    "953188": "hengck23  would you mind sharing which file is this ? I wonder if mp3 would have switched to short window during these periods when higher time resolution is needed",
    "953272": "see swamp sparrow (swaspa) : \n6kHz (window size 256): XC446809, XC116589, XC131033\n9kHz (window size 128): XC138153 \nthis site provide reference call files\n```\nhttps://www.audubon.org/field-guide/bird/swamp-sparrow\nSongs and Calls\nSweet, musical trill, all on one note.\n fast pulse-rate song\n slow pulse-rate song\n very slow pulse-rate song\n odd buzzy song\n chips #1\n chips #2\n```",
    "953298": "https://drive.google.com/drive/folders/1SExd8V-vFfTRrmTgjJ5PXN34TFmCx3Hd?usp=sharing\n \n\nsome  time-interval hand annotation for the ebird swaspa . Note that i am not sure if my annotation would be 100% correct:\n\nannotated region : high confident that it should contain the bird \nnon-annotated region :  i cannot identified the bird ... it could be a true negative or a false negative(i.e. miss)",
    "953663": "segmented bird call event data\n\nhttps://zenodo.org/record/1250690#.XyTA8zczbCJ\n\n(some of the bird species overlap with ours, e.g.  Blue Jay ,Song Sparrow,Great Blue Heron)\n\nhttps://figshare.com/articles/SwampSparrow_luscdb_zip/5625310\nhttps://github.com/timsainb/avgn_paper/blob/vizmerge/notebooks/00.0-download-datasets/1.0-bird-db-download-dataset.ipynb",
    "954428": "https://openreview.net/forum?id=_P9LyJ5pMDb\nUsing Self-Supervised Learning of Birdsong for Downstream Industrial Audio Classification\ndataset and code: https://github.com/SingingData/Birdsong",
    "955655": "i make the google drive link shareable. i will add more birds later. if you have problem accessing this, please comment here. thanks.",
    "955660": "Thank you for sharing Heng. I have a question. For example, in swaspa/XC116589.Table.1.selections.txt we can see there is an empty interval in [1.814, 7.2129]. Do you think our models will perform better on the private LB if we take 5 seconds from empty segments like that, add it as \"nocall\" target so we have 264+1 total targets, and use it in our training? Thank you",
    "957183": "i do not add them to training at first because i cannot be sure if the empty interval is background or not. (Actually, you can add \"background\" label to each wav, i.e. 2 label per class)\n\nrather, i would train with external noise first. then i will pseudo label in later iterations a shown below:\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcdb33754345434fd168d50bde6f3f919%2FSelection_037.png?generation=1596520583850968&amp;alt=media)",
    "958374": "One of my automatic event detection methods for training. see attached code for an illustration.\nyou can use ensemble of methods to create or score the annotations\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0ddfde356c58120ffa76c5769fa0b056%2FSelection_057.png?generation=1596587353555325&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F39700603673e25328a92bec23c16f579%2FSelection_058.png?generation=1596587351298156&amp;alt=media)",
    "960606": "some insights on birds recording\n\nhttps://www.avisoft.com/tutorials/measuring-sound-parameters-from-the-spectrogram-automatically/\nhttps://www.youtube.com/watch?v=9MYKnze7Zaw",
    "960725": "could be due to MP3 decoding delay",
    "960900": "i just realise that you can use BirdNET to label the clips\n\nhttps://birdnet.cornell.edu/api/\nhttps://github.com/kahst/BirdNET-Demo\nhttps://www.wildlabs.net/resources/community-announcements/wildlabs-virtual-meetup-recording-acoustic-monitoring\ndata and software from Australian Acoustic Observatory: https://data.acousticobservatory.org/listen\nhttps://ap.qut.ecoacoustics.info/tutorials/01-usingap/practical?tabs=linux",
    "969648": "https://rpubs.com/marcelo-araya-salas/110155",
    "971247": "> one file cannot contains many bird type\n\nAre you referring to the training examples? If so, I understand. Otherwise, I don't understand - aren't the examples in the test set soundscapes containing possibly many different birds?",
    "971252": "you can expect maybe up to 4 to 8 birds. There will be some limit",
    "971260": "Ah, I see - you're saying there may be multiple birds, but not too many. Thanks for this and for all your posts here - very interesting resources, and a great way of sharing which (it seems) everyone is welcoming!",
    "974496": "https://drive.google.com/drive/folders/14n_DlU0dGiXZE6Zl_AT78J2h7VZwZ405\n\npseudo labels results:\n- efficientb2 : my trained model\n- bird/_net : results predicted from https://github.com/kahst/BirdNET-Electron\n- allaboutbirds : reference spectrogram from the website. \n\nthis gives you an ideas how much improvement you can get. right now i am making an interface to hand correct these pseudo labels and then retrain the model with better strong labels\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6fa3be73ac4c74a79a09ae9d505e38ff%2FXC179124_20.png?generation=1597712354411244&alt=media)",
    "981194": "there is a time-annotated bird song dataset\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe36174a808fedc5f65bad172126f083c%2FSelection_034.png?generation=1598086493482977&alt=media)\nhttps://www.sciencedirect.com/science/article/pii/S1574954115000151\n\nhttp://taylor0.biology.ucla.edu/birdDBQuery/\n\n(emable Textgrid during search)\nmaybe good for pretraining or event detection or controlled experiments?",
    "983276": "Thank you, this thread definitely helped me realise that I was wasting my time trying to solve the problem without better labeling of the data. I think I have probably realised this too late to put together a good model / submission but still at least I have learnt something and started to put together a model that can have some limited accuracy...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2Fb1d3ab630a397c403e0c5ecb06a8a22f%2Fexample_clipped.png?generation=1598251200823054&alt=media)",
    "983686": "pseudo labels from softmax classifier train on samples without secondary labels.\nclass activation map (cam) is used\n(note that even when secondary = None, it is possible that there are still other background birds)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff58923760037923f0ab4054336984d33%2FXC109300_01.png?generation=1598277772831445&alt=media)",
    "983687": "the difficulty is how to label  using ML methods automatically or semi automatically. \ntotal label by hands is not practical or not possible.",
    "985715": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2Fab0559dff6661c1a6e0ac1d620b5eef0%2Fexample%20cropped.jpg?generation=1598413253515349&alt=media)\n\ndetection of similar calls across multiple audio tracks for a single bird type (note - my classifier is not trained on this bird type, though i did some manual work at earlier stages on a small number of other bird types)\n\nhave to admit i had figured with secondary labels, if we can extract 'clean' (at least fairly clean) samples on the main bird type, can always create our own mix & match examples for model training? and presumably a classifier that has been trained on 'clean' data should in turn be able to help with labeling original audio, if needed. i wasn't sure how critical it was to label all the original clips 100%. any thoughts welcome...",
    "985886": "Extract of some of the most 'similar' calls across multiple audio clips for aldfly\n\nAgain assuming that the most typical call across multiple clips will help avoid 'secondary' birds to get clean data for classification. As the main bird will be in all the clips.\n\nHowever won't really solve the song/call mix from the same bird. Suspect I could try to work round to that but likely to be much more complicated and I'm already out of my depth in this comp!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F575494%2F142a1796f47bd49b3ce65aad68eb46bc%2Foutput%20example%204.jpg?generation=1598417385182818&alt=media)",
    "985938": "\"have to admit i had figured with secondary labels, if we can extract 'clean' (at least fairly clean) samples on the main bird type, can always create our own mix & match examples for model training? and presumably a classifier that has been trained on 'clean' data should in turn be able to help with labeling original audio, if needed. i wasn't sure how critical it was to label all the original clips 100%.\"\n\n\n100% clean data is not required. \neven with 100% clean data there is a bigger issue of adding noise to the data to mimic the soundscape environment.\n\ni suggest you should quickly make a submission to get baseline results based on what you have now. you should also test your model on the two sample test clip and the given clefbird 2020 clips in the external data thread.",
    "987920": "yet another time annotated dataset:\nhttps://avocet.integrativebiology.natsci.msu.edu/species/1922\n\n![](https://avocet.integrativebiology.natsci.msu.edu/recordings_data/1/169/16983/sonogram.jpg)",
    "989740": "Hey, I looked into Scaper. To me it doesn't seem useful since it requires the sounds it uses to match these descriptions:\n\n> * **Background files**: are used to create the background of the soundscape, and\n  should contain audio material that is perceived as a single holistic sound\n  which is more distant, ambiguous, and texture-like (e.g. the \"hum\" or \"drone\"\n  of an urban environment, or \"wind and rain\" sounds in a natural environment).\n  Importantly, background files should not contain salient sound events.\n* **Foreground files**: are used to create sound events. Each foreground audio\n  file should contain a single sound event (short or long) such as a car honk,\n  an animal vocalization, continuous speech, a siren or an idling engine.\n  Foreground files should be as clean as possible with no background noise and\n  no silence before/after the sound event.\n\nAnd while we have access to some audio that matches the Background Files description, I'm pretty sure none of the audio we have access to matches the Foreground Files description. Am I missing something?\n\nAlso, you outlined a really cool procedure for automatic event annotation and I want to try it myself. Did you use any particular library for determining Signal to Noise Ratio?\n\nThanks for all your hard work in any case.",
    "989844": "at the moment I'm using \n\n```\ndef signaltonoise(a, axis=0, ddof=0):\n    a = np.asanyarray(a)\n    return a.mean(axis)/a.std(axis=axis, ddof=ddof)\n\ndef segmentwise_snr(wave, segment_length):\n    snrs = []\n    start = 0\n    while start < len(wave):\n        snrs.append(signaltonoise(wave[start:start+segment_length]))\n        start += segment_length\n    return snrs\n\ndef snr_windows(wave, segment_length, threshold, segment_windows):\n    segments = segmentwise_snr(wave, segment_length)\n    \n    window_scores = {}\n    for w in segment_windows:\n        scores = []\n        start = 0\n        while start < len(segments):\n            scores.append(np.mean(segments[start:start+w]) > threshold)\n            start += 1\n        window_scores[w] = scores\n    return window_scores\n\nf = snr_windows(wave, int(sr / 8), 0.0008, [2, 5, 25])\n```\n\nBut the results are utter garbage. Any advice for creating a rough automatic SNR-based bird-call event detector?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2Fe9753e01ca7fdc35cf3015b094460a5c%2FScreenshot%20from%202020-08-29%2017-07-56.png?generation=1598684960612534&alt=media)",
    "992122": "removing other birds help!\n\ni was investigating why i cannot detect brncre in the sample test audio. i did several noise augmentation and also download external data. i check visually that there is some training set + augmentation that resemble the test sample.\n\nit turns out that the problem is presence of other birds.\n(this is results from trained softmax classifier)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc3dec9301bc2641e9319d77c4bda92bd%2FSelection_033.png?generation=1598835210055130&alt=media)",
    "992406": "Thanks for the info! This is really useful. Have you tried multi-label classification with sigmoid? The model can maybe learn from multiple birdcalls in the spectrogram"
  },
  "source": "meta"
}