{
  "id": 166485,
  "title": "Recognizing Birds from Sound - The 2018 BirdCLEF Baseline System",
  "url": "/competitions/birdsong-recognition/discussion/166485",
  "author_name": "",
  "post_date": "2020-07-13T05:10:43.151858900Z",
  "votes": 21,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I have been reviewing some of the existing literature related to bird sounds and writing summaries of the papers. Below is the summary of <a href=\"https://arxiv.org/abs/1804.07177\">Recognizing Birds from Sound - The 2018 BirdCLEF Baseline System</a></p>\n\n<p>Paper proposes a method for preparing audio by creating features from 1s worth of audio converted to spectrograms and applying a CNN. Model uses data from 1500 birds from xeno-canto. Mel-spectrogram is used and a 300Hz and 15kHz high pass and low pass filter is used as this is the vocal range of birds loosely. </p>\n\n<p>A rule-based approach is applied to reject samples that don’t contain any bird sounds. Process is accepted to be imperfect but good enough to reduce the amount of unuseful data. After this cleaning process 521,873 spectrograms remain made up of these 1s audio clips. </p>\n\n<p>The architecture is a simple CNN with relatively few layers and filters. Preliminary results show that more layers would likely add value. (in my opinion borrowing pretrained imagenet weights and a simple resnet 18 would likely greatly improve performance).</p>\n\n<p>A small subset was set aside for hyperparameter tuning and a larger fully held out test set is used for final evaluation of the model. Evaluation used was MLRAP <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.label_ranking_average_precision_score.html\">https://scikit-learn.org/stable/modules/generated/sklearn.metrics.label_ranking_average_precision_score.html</a></p>\n\n<p>Potential further work is suggested in the domain of more recent architectures and introducing additional metadata as a feature to tune predictions or as a feature. This can help in localization of predictions. Many birds can be easily ruled out in many parts of the globe that they are not present in. </p>",
  "messages": [
    {
      "id": "926916",
      "postDate": "07/13/2020 05:10:43",
      "content": "<p>I have been reviewing some of the existing literature related to bird sounds and writing summaries of the papers. Below is the summary of <a href=\"https://arxiv.org/abs/1804.07177\">Recognizing Birds from Sound - The 2018 BirdCLEF Baseline System</a></p>\n\n<p>Paper proposes a method for preparing audio by creating features from 1s worth of audio converted to spectrograms and applying a CNN. Model uses data from 1500 birds from xeno-canto. Mel-spectrogram is used and a 300Hz and 15kHz high pass and low pass filter is used as this is the vocal range of birds loosely. </p>\n\n<p>A rule-based approach is applied to reject samples that don’t contain any bird sounds. Process is accepted to be imperfect but good enough to reduce the amount of unuseful data. After this cleaning process 521,873 spectrograms remain made up of these 1s audio clips. </p>\n\n<p>The architecture is a simple CNN with relatively few layers and filters. Preliminary results show that more layers would likely add value. (in my opinion borrowing pretrained imagenet weights and a simple resnet 18 would likely greatly improve performance).</p>\n\n<p>A small subset was set aside for hyperparameter tuning and a larger fully held out test set is used for final evaluation of the model. Evaluation used was MLRAP <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.label_ranking_average_precision_score.html\">https://scikit-learn.org/stable/modules/generated/sklearn.metrics.label_ranking_average_precision_score.html</a></p>\n\n<p>Potential further work is suggested in the domain of more recent architectures and introducing additional metadata as a feature to tune predictions or as a feature. This can help in localization of predictions. Many birds can be easily ruled out in many parts of the globe that they are not present in. </p>",
      "rawMarkdown": "I have been reviewing some of the existing literature related to bird sounds and writing summaries of the papers. Below is the summary of [Recognizing Birds from Sound - The 2018 BirdCLEF Baseline System](https://arxiv.org/abs/1804.07177)\n\nPaper proposes a method for preparing audio by creating features from 1s worth of audio converted to spectrograms and applying a CNN. Model uses data from 1500 birds from xeno-canto. Mel-spectrogram is used and a 300Hz and 15kHz high pass and low pass filter is used as this is the vocal range of birds loosely. \n\nA rule-based approach is applied to reject samples that don’t contain any bird sounds. Process is accepted to be imperfect but good enough to reduce the amount of unuseful data. After this cleaning process 521,873 spectrograms remain made up of these 1s audio clips. \n\nThe architecture is a simple CNN with relatively few layers and filters. Preliminary results show that more layers would likely add value. (in my opinion borrowing pretrained imagenet weights and a simple resnet 18 would likely greatly improve performance).\n\nA small subset was set aside for hyperparameter tuning and a larger fully held out test set is used for final evaluation of the model. Evaluation used was MLRAP https://scikit-learn.org/stable/modules/generated/sklearn.metrics.label_ranking_average_precision_score.html\n\nPotential further work is suggested in the domain of more recent architectures and introducing additional metadata as a feature to tune predictions or as a feature. This can help in localization of predictions. Many birds can be easily ruled out in many parts of the globe that they are not present in.",
      "votes": null
    },
    {
      "id": "926945",
      "postDate": "07/13/2020 05:45:44",
      "content": "<p>github repo for this work available here: <a href=\"https://github.com/kahst/BirdCLEF-Baseline\">https://github.com/kahst/BirdCLEF-Baseline</a></p>",
      "rawMarkdown": "github repo for this work available here: [https://github.com/kahst/BirdCLEF-Baseline](https://github.com/kahst/BirdCLEF-Baseline)",
      "votes": null
    },
    {
      "id": "929950",
      "postDate": "07/15/2020 05:20:36",
      "content": "<p>It was a very nice approach by the way.  Thanks.</p>\n\n<p>How about using existing libraries like Kaldi or Sphinx for using feature extraction here and then we can perform our own classification based on network. The existing libraries are used for ASR(Automatic Speech Recognition) system.  Can we use existing ASR libraries for as feature extraction part in the process and silence removal also important part isn't?</p>",
      "rawMarkdown": "It was a very nice approach by the way.  Thanks.\n\nHow about using existing libraries like Kaldi or Sphinx for using feature extraction here and then we can perform our own classification based on network. The existing libraries are used for ASR(Automatic Speech Recognition) system.  Can we use existing ASR libraries for as feature extraction part in the process and silence removal also important part isn't?",
      "votes": null
    },
    {
      "id": "931159",
      "postDate": "07/16/2020 03:06:39",
      "content": "<p>Thanks. In fact, the approach of using log-mel-spectograms as preprocessing, and CNNs or CRNNs for feature extraction is very common in audio applications. I have done some work with such models before, and they performed very well in different applications domains. If someone is interested going deep in this subject, I recommend this survey paper: <a href=\"https://ieeexplore.ieee.org/abstract/document/8336092\">https://ieeexplore.ieee.org/abstract/document/8336092</a>. </p>",
      "rawMarkdown": "Thanks. In fact, the approach of using log-mel-spectograms as preprocessing, and CNNs or CRNNs for feature extraction is very common in audio applications. I have done some work with such models before, and they performed very well in different applications domains. If someone is interested going deep in this subject, I recommend this survey paper: https://ieeexplore.ieee.org/abstract/document/8336092.",
      "votes": null
    },
    {
      "id": "931327",
      "postDate": "07/16/2020 05:55:33",
      "content": "<p>I wonder what score would it give on the LB :)</p>",
      "rawMarkdown": "I wonder what score would it give on the LB :)",
      "votes": null
    },
    {
      "id": "932174",
      "postDate": "07/16/2020 19:26:46",
      "content": "<p>I'm not sure exactly how applicable systems like those would be to bird speech. Maybe if you peeled off the layers that tried to do the decoding to the syllable level and just had it output some latent representation of the audio, but even then it had been trained on human speech and not birds. </p>",
      "rawMarkdown": "I'm not sure exactly how applicable systems like those would be to bird speech. Maybe if you peeled off the layers that tried to do the decoding to the syllable level and just had it output some latent representation of the audio, but even then it had been trained on human speech and not birds.",
      "votes": null
    },
    {
      "id": "932175",
      "postDate": "07/16/2020 19:27:07",
      "content": "<p>Very interesting. Thanks for the recommendation. I'll have to check this out. </p>",
      "rawMarkdown": "Very interesting. Thanks for the recommendation. I'll have to check this out.",
      "votes": null
    },
    {
      "id": "932176",
      "postDate": "07/16/2020 19:27:56",
      "content": "<p>I suspect significantly better than what the leaderboard and kaggle community has figured out to this point. </p>",
      "rawMarkdown": "I suspect significantly better than what the leaderboard and kaggle community has figured out to this point.",
      "votes": null
    },
    {
      "id": "932657",
      "postDate": "07/17/2020 07:31:09",
      "content": "<p><a href=\"/ryches\">@ryches</a>  I'm looking forward to you appearing at the top of LB :)</p>",
      "rawMarkdown": "ryches  I'm looking forward to you appearing at the top of LB :)",
      "votes": null
    },
    {
      "id": "932754",
      "postDate": "07/17/2020 08:57:19",
      "content": "<p>Thank you for sharing the paper!\nI read it and felt that it was essentially the same as Public Notebook's method for the following reasons.\nI would be happy if you could tell me if I had a misunderstanding.</p>\n\n<h3>Approach</h3>\n\n<p>It seems that the problems that Public Notebook may have are unsolved, since the basic approach is the same (mel-spectrograms + CNN).\n- Dealing with diverse crying species (in-class variation)\n- Dealing with similar crying species (between-class similarity)\n- Dealing with background noise other than stationary noise</p>\n\n<h3>Rule-based processing</h3>\n\n<p>It is possible that this kind of processing is implicitly performed in CNNs.\nThe preprocess might help CNNs, but conversely, missing information might cause adverse effects. </p>\n\n<h3>Birds' vocal range and the birds' hearing</h3>\n\n<p>In the paper, FMIN and FMAX are narrowed down according to birds' vocal range.\nI think this is somewhat effective.\nWhen I used FMAX=8kHz, the loss value got worse. It seems bad to limit too much.</p>\n\n<p>According to some websites, the hearing of birds has higher frequency resolution and time resolution than humans, and it is possible to hear weaker sounds. (Of course it depends on the species.)\nIt may be meaningful to select signal processing that matches not only the birds' vocal range but also the birds' hearing.</p>\n\n<p>However, since GTs are labeled by humans, I think it is appropriate to select signal processing that matches the humans' perception.\nIt's unclear if birds can distinguish the songs of other species, right?</p>",
      "rawMarkdown": "Thank you for sharing the paper!\nI read it and felt that it was essentially the same as Public Notebook's method for the following reasons.\nI would be happy if you could tell me if I had a misunderstanding.\n\n### Approach\nIt seems that the problems that Public Notebook may have are unsolved, since the basic approach is the same (mel-spectrograms + CNN).\n- Dealing with diverse crying species (in-class variation)\n- Dealing with similar crying species (between-class similarity)\n- Dealing with background noise other than stationary noise\n\n### Rule-based processing\nIt is possible that this kind of processing is implicitly performed in CNNs.\nThe preprocess might help CNNs, but conversely, missing information might cause adverse effects. \n\n### Birds' vocal range and the birds' hearing\nIn the paper, FMIN and FMAX are narrowed down according to birds' vocal range.\nI think this is somewhat effective.\nWhen I used FMAX=8kHz, the loss value got worse. It seems bad to limit too much.\n\nAccording to some websites, the hearing of birds has higher frequency resolution and time resolution than humans, and it is possible to hear weaker sounds. (Of course it depends on the species.)\nIt may be meaningful to select signal processing that matches not only the birds' vocal range but also the birds' hearing.\n\nHowever, since GTs are labeled by humans, I think it is appropriate to select signal processing that matches the humans' perception.\nIt's unclear if birds can distinguish the songs of other species, right?",
      "votes": null
    },
    {
      "id": "933289",
      "postDate": "07/17/2020 16:19:08",
      "content": "<p>After reading some of the hosts/paper authors comments my outlook has changed. They mention they haven't been able to solve the false positive problem so I am less bullish on a large leap in performance. </p>",
      "rawMarkdown": "After reading some of the hosts/paper authors comments my outlook has changed. They mention they haven't been able to solve the false positive problem so I am less bullish on a large leap in performance.",
      "votes": null
    },
    {
      "id": "933419",
      "postDate": "07/17/2020 17:32:54",
      "content": "<p>Dooling's '<a href=\"https://www.researchgate.net/publication/251412063_Hearing_in_Birds_and_Reptiles\">Hearing in Birds and Reptiles</a>' puts the top end of bird hearing somewhere around 12kHz, with sensitivity peaking in the range that a particular bird sings at. IIRC, most of the BirdClef literature ends up putting a top-end at about 10kHz, which is enough to get the golden crown kinglet and brown creeper. </p>\n\n<p>The higher time resolution is definitely A Thing. Chipping Sparrow and Junco are two species which humans sometimes have a hard time distinguishing (they both trill in a range of speeds, but the ranges overlap; a fast junco and slow chipping sparrow sound pretty similar). But they are typically pretty easy to tell apart in the spectrogram, as the Junco are usually alternating two notes faster than our ears can resolve. </p>\n\n<p>When you set up a DFT you've got a hop size and window length. Lowering the DFT hop size is also an excellent way to increase time resolution, and is generally just limited by processing capacity. The <a href=\"https://arxiv.org/abs/1912.10211\">PANNs paper</a> shows better results on Audioset classification with a 10ms window, for example, and I think didn't try shorter hops...</p>\n\n<p>Window length is a harder to get a single good answer for. A long window is good for resolving low-frequency sounds, but reduces locality: if you've got a 1-second window (this is ridiculously long, btw) and a short high-frequency burst at (say) 8kHz, you see a peak, but you're not sure where in the 1s window it happened. So you're typically looking for a 'goldilocks' length, which gets decent locality but still picks up the low frequencies well (for our owl friends). Another approach is to just have separate spectrograms for very low and/or very high frequencies: this is actually what you see in <a href=\"https://academy.allaboutbirds.org/peterson-field-guide-to-bird-sounds/\">Pieplow's books on bird song</a>, which have special 'owl' spectrograms. So this is definitely an interesting area to investigate!</p>\n\n<p>For your last question (do birds distinguish other species?), the answer is 'maybe sometimes' with a healthy dose of 'we don't understand bird thoughts.' Jays are known to mimic hawks (either to scare off other birds or to keep hawks away by proclaiming territory; no one's sure). Other mimics with large repertoires certainly reproduce other species sounds, but it's unclear whether they know what they're mimic'ing or just <a href=\"https://www.allaboutbirds.org/news/mockingbirds-can-learn-hundreds-of-songs-but-theres-a-limit/\">copying the environment</a>.\nThere are also <a href=\"https://www.audubon.org/news/a-wave-bird-alarm-calls-can-travel-100-miles-hour\">shared alarm calls</a>: short, high-pitched 'seet' calls recognized by many songbirds across species.</p>",
      "rawMarkdown": "Dooling's '[Hearing in Birds and Reptiles](https://www.researchgate.net/publication/251412063_Hearing_in_Birds_and_Reptiles)' puts the top end of bird hearing somewhere around 12kHz, with sensitivity peaking in the range that a particular bird sings at. IIRC, most of the BirdClef literature ends up putting a top-end at about 10kHz, which is enough to get the golden crown kinglet and brown creeper. \n\nThe higher time resolution is definitely A Thing. Chipping Sparrow and Junco are two species which humans sometimes have a hard time distinguishing (they both trill in a range of speeds, but the ranges overlap; a fast junco and slow chipping sparrow sound pretty similar). But they are typically pretty easy to tell apart in the spectrogram, as the Junco are usually alternating two notes faster than our ears can resolve. \n\nWhen you set up a DFT you've got a hop size and window length. Lowering the DFT hop size is also an excellent way to increase time resolution, and is generally just limited by processing capacity. The [PANNs paper](https://arxiv.org/abs/1912.10211) shows better results on Audioset classification with a 10ms window, for example, and I think didn't try shorter hops...\n\nWindow length is a harder to get a single good answer for. A long window is good for resolving low-frequency sounds, but reduces locality: if you've got a 1-second window (this is ridiculously long, btw) and a short high-frequency burst at (say) 8kHz, you see a peak, but you're not sure where in the 1s window it happened. So you're typically looking for a 'goldilocks' length, which gets decent locality but still picks up the low frequencies well (for our owl friends). Another approach is to just have separate spectrograms for very low and/or very high frequencies: this is actually what you see in [Pieplow's books on bird song](https://academy.allaboutbirds.org/peterson-field-guide-to-bird-sounds/), which have special 'owl' spectrograms. So this is definitely an interesting area to investigate!\n\nFor your last question (do birds distinguish other species?), the answer is 'maybe sometimes' with a healthy dose of 'we don't understand bird thoughts.' Jays are known to mimic hawks (either to scare off other birds or to keep hawks away by proclaiming territory; no one's sure). Other mimics with large repertoires certainly reproduce other species sounds, but it's unclear whether they know what they're mimic'ing or just [copying the environment](https://www.allaboutbirds.org/news/mockingbirds-can-learn-hundreds-of-songs-but-theres-a-limit/).\nThere are also [shared alarm calls](https://www.audubon.org/news/a-wave-bird-alarm-calls-can-travel-100-miles-hour): short, high-pitched 'seet' calls recognized by many songbirds across species.",
      "votes": null
    },
    {
      "id": "933677",
      "postDate": "07/17/2020 22:22:59",
      "content": "<p>Thank you for your very informative comments!\nI can confirm that there is a lot to do.</p>\n\n<p>To remove various trade-offs, I'm considering combining spectrograms obtained with different window functions ( or window lengths ).\nI will report somewhere when I get any results.</p>",
      "rawMarkdown": "Thank you for your very informative comments!\nI can confirm that there is a lot to do.\n\nTo remove various trade-offs, I'm considering combining spectrograms obtained with different window functions ( or window lengths ).\nI will report somewhere when I get any results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 926945,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "07/13/2020 05:45:44",
      "content": "<p>github repo for this work available here: <a href=\"https://github.com/kahst/BirdCLEF-Baseline\">https://github.com/kahst/BirdCLEF-Baseline</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 929950,
      "author_name": "narendrageek",
      "author_url": "",
      "post_date": "07/15/2020 05:20:36",
      "content": "<p>It was a very nice approach by the way.  Thanks.</p>\n\n<p>How about using existing libraries like Kaldi or Sphinx for using feature extraction here and then we can perform our own classification based on network. The existing libraries are used for ASR(Automatic Speech Recognition) system.  Can we use existing ASR libraries for as feature extraction part in the process and silence removal also important part isn't?</p>",
      "votes": null,
      "replies": [
        {
          "id": 932174,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "07/16/2020 19:26:46",
          "content": "<p>I'm not sure exactly how applicable systems like those would be to bird speech. Maybe if you peeled off the layers that tried to do the decoding to the syllable level and just had it output some latent representation of the audio, but even then it had been trained on human speech and not birds. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 931159,
      "author_name": "mauriciofigueiredo",
      "author_url": "",
      "post_date": "07/16/2020 03:06:39",
      "content": "<p>Thanks. In fact, the approach of using log-mel-spectograms as preprocessing, and CNNs or CRNNs for feature extraction is very common in audio applications. I have done some work with such models before, and they performed very well in different applications domains. If someone is interested going deep in this subject, I recommend this survey paper: <a href=\"https://ieeexplore.ieee.org/abstract/document/8336092\">https://ieeexplore.ieee.org/abstract/document/8336092</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 932175,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "07/16/2020 19:27:07",
          "content": "<p>Very interesting. Thanks for the recommendation. I'll have to check this out. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 931327,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "07/16/2020 05:55:33",
      "content": "<p>I wonder what score would it give on the LB :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 932176,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "07/16/2020 19:27:56",
          "content": "<p>I suspect significantly better than what the leaderboard and kaggle community has figured out to this point. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932657,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "07/17/2020 07:31:09",
          "content": "<p><a href=\"/ryches\">@ryches</a>  I'm looking forward to you appearing at the top of LB :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933289,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "07/17/2020 16:19:08",
          "content": "<p>After reading some of the hosts/paper authors comments my outlook has changed. They mention they haven't been able to solve the false positive problem so I am less bullish on a large leap in performance. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 932754,
      "author_name": "octpath0302",
      "author_url": "",
      "post_date": "07/17/2020 08:57:19",
      "content": "<p>Thank you for sharing the paper!\nI read it and felt that it was essentially the same as Public Notebook's method for the following reasons.\nI would be happy if you could tell me if I had a misunderstanding.</p>\n\n<h3>Approach</h3>\n\n<p>It seems that the problems that Public Notebook may have are unsolved, since the basic approach is the same (mel-spectrograms + CNN).\n- Dealing with diverse crying species (in-class variation)\n- Dealing with similar crying species (between-class similarity)\n- Dealing with background noise other than stationary noise</p>\n\n<h3>Rule-based processing</h3>\n\n<p>It is possible that this kind of processing is implicitly performed in CNNs.\nThe preprocess might help CNNs, but conversely, missing information might cause adverse effects. </p>\n\n<h3>Birds' vocal range and the birds' hearing</h3>\n\n<p>In the paper, FMIN and FMAX are narrowed down according to birds' vocal range.\nI think this is somewhat effective.\nWhen I used FMAX=8kHz, the loss value got worse. It seems bad to limit too much.</p>\n\n<p>According to some websites, the hearing of birds has higher frequency resolution and time resolution than humans, and it is possible to hear weaker sounds. (Of course it depends on the species.)\nIt may be meaningful to select signal processing that matches not only the birds' vocal range but also the birds' hearing.</p>\n\n<p>However, since GTs are labeled by humans, I think it is appropriate to select signal processing that matches the humans' perception.\nIt's unclear if birds can distinguish the songs of other species, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 933419,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "07/17/2020 17:32:54",
          "content": "<p>Dooling's '<a href=\"https://www.researchgate.net/publication/251412063_Hearing_in_Birds_and_Reptiles\">Hearing in Birds and Reptiles</a>' puts the top end of bird hearing somewhere around 12kHz, with sensitivity peaking in the range that a particular bird sings at. IIRC, most of the BirdClef literature ends up putting a top-end at about 10kHz, which is enough to get the golden crown kinglet and brown creeper. </p>\n\n<p>The higher time resolution is definitely A Thing. Chipping Sparrow and Junco are two species which humans sometimes have a hard time distinguishing (they both trill in a range of speeds, but the ranges overlap; a fast junco and slow chipping sparrow sound pretty similar). But they are typically pretty easy to tell apart in the spectrogram, as the Junco are usually alternating two notes faster than our ears can resolve. </p>\n\n<p>When you set up a DFT you've got a hop size and window length. Lowering the DFT hop size is also an excellent way to increase time resolution, and is generally just limited by processing capacity. The <a href=\"https://arxiv.org/abs/1912.10211\">PANNs paper</a> shows better results on Audioset classification with a 10ms window, for example, and I think didn't try shorter hops...</p>\n\n<p>Window length is a harder to get a single good answer for. A long window is good for resolving low-frequency sounds, but reduces locality: if you've got a 1-second window (this is ridiculously long, btw) and a short high-frequency burst at (say) 8kHz, you see a peak, but you're not sure where in the 1s window it happened. So you're typically looking for a 'goldilocks' length, which gets decent locality but still picks up the low frequencies well (for our owl friends). Another approach is to just have separate spectrograms for very low and/or very high frequencies: this is actually what you see in <a href=\"https://academy.allaboutbirds.org/peterson-field-guide-to-bird-sounds/\">Pieplow's books on bird song</a>, which have special 'owl' spectrograms. So this is definitely an interesting area to investigate!</p>\n\n<p>For your last question (do birds distinguish other species?), the answer is 'maybe sometimes' with a healthy dose of 'we don't understand bird thoughts.' Jays are known to mimic hawks (either to scare off other birds or to keep hawks away by proclaiming territory; no one's sure). Other mimics with large repertoires certainly reproduce other species sounds, but it's unclear whether they know what they're mimic'ing or just <a href=\"https://www.allaboutbirds.org/news/mockingbirds-can-learn-hundreds-of-songs-but-theres-a-limit/\">copying the environment</a>.\nThere are also <a href=\"https://www.audubon.org/news/a-wave-bird-alarm-calls-can-travel-100-miles-hour\">shared alarm calls</a>: short, high-pitched 'seet' calls recognized by many songbirds across species.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933677,
          "author_name": "octpath0302",
          "author_url": "",
          "post_date": "07/17/2020 22:22:59",
          "content": "<p>Thank you for your very informative comments!\nI can confirm that there is a lot to do.</p>\n\n<p>To remove various trade-offs, I'm considering combining spectrograms obtained with different window functions ( or window lengths ).\nI will report somewhere when I get any results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "926916": "I have been reviewing some of the existing literature related to bird sounds and writing summaries of the papers. Below is the summary of [Recognizing Birds from Sound - The 2018 BirdCLEF Baseline System](https://arxiv.org/abs/1804.07177)\n\nPaper proposes a method for preparing audio by creating features from 1s worth of audio converted to spectrograms and applying a CNN. Model uses data from 1500 birds from xeno-canto. Mel-spectrogram is used and a 300Hz and 15kHz high pass and low pass filter is used as this is the vocal range of birds loosely. \n\nA rule-based approach is applied to reject samples that don’t contain any bird sounds. Process is accepted to be imperfect but good enough to reduce the amount of unuseful data. After this cleaning process 521,873 spectrograms remain made up of these 1s audio clips. \n\nThe architecture is a simple CNN with relatively few layers and filters. Preliminary results show that more layers would likely add value. (in my opinion borrowing pretrained imagenet weights and a simple resnet 18 would likely greatly improve performance).\n\nA small subset was set aside for hyperparameter tuning and a larger fully held out test set is used for final evaluation of the model. Evaluation used was MLRAP https://scikit-learn.org/stable/modules/generated/sklearn.metrics.label_ranking_average_precision_score.html\n\nPotential further work is suggested in the domain of more recent architectures and introducing additional metadata as a feature to tune predictions or as a feature. This can help in localization of predictions. Many birds can be easily ruled out in many parts of the globe that they are not present in.",
    "926945": "github repo for this work available here: [https://github.com/kahst/BirdCLEF-Baseline](https://github.com/kahst/BirdCLEF-Baseline)",
    "929950": "It was a very nice approach by the way.  Thanks.\n\nHow about using existing libraries like Kaldi or Sphinx for using feature extraction here and then we can perform our own classification based on network. The existing libraries are used for ASR(Automatic Speech Recognition) system.  Can we use existing ASR libraries for as feature extraction part in the process and silence removal also important part isn't?",
    "931159": "Thanks. In fact, the approach of using log-mel-spectograms as preprocessing, and CNNs or CRNNs for feature extraction is very common in audio applications. I have done some work with such models before, and they performed very well in different applications domains. If someone is interested going deep in this subject, I recommend this survey paper: https://ieeexplore.ieee.org/abstract/document/8336092.",
    "931327": "I wonder what score would it give on the LB :)",
    "932174": "I'm not sure exactly how applicable systems like those would be to bird speech. Maybe if you peeled off the layers that tried to do the decoding to the syllable level and just had it output some latent representation of the audio, but even then it had been trained on human speech and not birds.",
    "932175": "Very interesting. Thanks for the recommendation. I'll have to check this out.",
    "932176": "I suspect significantly better than what the leaderboard and kaggle community has figured out to this point.",
    "932657": "ryches  I'm looking forward to you appearing at the top of LB :)",
    "932754": "Thank you for sharing the paper!\nI read it and felt that it was essentially the same as Public Notebook's method for the following reasons.\nI would be happy if you could tell me if I had a misunderstanding.\n\n### Approach\nIt seems that the problems that Public Notebook may have are unsolved, since the basic approach is the same (mel-spectrograms + CNN).\n- Dealing with diverse crying species (in-class variation)\n- Dealing with similar crying species (between-class similarity)\n- Dealing with background noise other than stationary noise\n\n### Rule-based processing\nIt is possible that this kind of processing is implicitly performed in CNNs.\nThe preprocess might help CNNs, but conversely, missing information might cause adverse effects. \n\n### Birds' vocal range and the birds' hearing\nIn the paper, FMIN and FMAX are narrowed down according to birds' vocal range.\nI think this is somewhat effective.\nWhen I used FMAX=8kHz, the loss value got worse. It seems bad to limit too much.\n\nAccording to some websites, the hearing of birds has higher frequency resolution and time resolution than humans, and it is possible to hear weaker sounds. (Of course it depends on the species.)\nIt may be meaningful to select signal processing that matches not only the birds' vocal range but also the birds' hearing.\n\nHowever, since GTs are labeled by humans, I think it is appropriate to select signal processing that matches the humans' perception.\nIt's unclear if birds can distinguish the songs of other species, right?",
    "933289": "After reading some of the hosts/paper authors comments my outlook has changed. They mention they haven't been able to solve the false positive problem so I am less bullish on a large leap in performance.",
    "933419": "Dooling's '[Hearing in Birds and Reptiles](https://www.researchgate.net/publication/251412063_Hearing_in_Birds_and_Reptiles)' puts the top end of bird hearing somewhere around 12kHz, with sensitivity peaking in the range that a particular bird sings at. IIRC, most of the BirdClef literature ends up putting a top-end at about 10kHz, which is enough to get the golden crown kinglet and brown creeper. \n\nThe higher time resolution is definitely A Thing. Chipping Sparrow and Junco are two species which humans sometimes have a hard time distinguishing (they both trill in a range of speeds, but the ranges overlap; a fast junco and slow chipping sparrow sound pretty similar). But they are typically pretty easy to tell apart in the spectrogram, as the Junco are usually alternating two notes faster than our ears can resolve. \n\nWhen you set up a DFT you've got a hop size and window length. Lowering the DFT hop size is also an excellent way to increase time resolution, and is generally just limited by processing capacity. The [PANNs paper](https://arxiv.org/abs/1912.10211) shows better results on Audioset classification with a 10ms window, for example, and I think didn't try shorter hops...\n\nWindow length is a harder to get a single good answer for. A long window is good for resolving low-frequency sounds, but reduces locality: if you've got a 1-second window (this is ridiculously long, btw) and a short high-frequency burst at (say) 8kHz, you see a peak, but you're not sure where in the 1s window it happened. So you're typically looking for a 'goldilocks' length, which gets decent locality but still picks up the low frequencies well (for our owl friends). Another approach is to just have separate spectrograms for very low and/or very high frequencies: this is actually what you see in [Pieplow's books on bird song](https://academy.allaboutbirds.org/peterson-field-guide-to-bird-sounds/), which have special 'owl' spectrograms. So this is definitely an interesting area to investigate!\n\nFor your last question (do birds distinguish other species?), the answer is 'maybe sometimes' with a healthy dose of 'we don't understand bird thoughts.' Jays are known to mimic hawks (either to scare off other birds or to keep hawks away by proclaiming territory; no one's sure). Other mimics with large repertoires certainly reproduce other species sounds, but it's unclear whether they know what they're mimic'ing or just [copying the environment](https://www.allaboutbirds.org/news/mockingbirds-can-learn-hundreds-of-songs-but-theres-a-limit/).\nThere are also [shared alarm calls](https://www.audubon.org/news/a-wave-bird-alarm-calls-can-travel-100-miles-hour): short, high-pitched 'seet' calls recognized by many songbirds across species.",
    "933677": "Thank you for your very informative comments!\nI can confirm that there is a lot to do.\n\nTo remove various trade-offs, I'm considering combining spectrograms obtained with different window functions ( or window lengths ).\nI will report somewhere when I get any results."
  },
  "source": "meta"
}