{
  "id": 236143,
  "title": "Spectrogram and image questions I have been thinking about",
  "url": "/competitions/birdclef-2021/discussion/236143",
  "author_name": "Steve Pousty",
  "post_date": "2021-05-03T04:12:57.867000",
  "votes": 12,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I tried looking at some calls in Sonic Visualizer<br>\n<a href=\"https://www.sonicvisualiser.org/\" target=\"_blank\">https://www.sonicvisualiser.org/</a><br>\nAnd I worked <br>\nbkcchi -&gt; xc121068.ogg <br>\nthrough this tutorial:<br>\n<a href=\"https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\" target=\"_blank\">https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0</a><br>\nand Sonic Visualizer</p>\n<p>I find the mel spectrogram actually ends up increasing the amount of noise highlighted in the picture. The plain power spectrogram in sonic visualizer seems to work better for features. I also don't think we need the mel scaling because we don't actually care about human hearing. What's more important is the actual spectrogram not what we can hear, especially since birds can hear outside of our range. </p>\n<p>Questions:<br>\n1)  I can't seem to get librosa to give me the clean image I am can get in Sonic Visualizer using the Simple Power Spectrum plugin. </p>\n<p>Plugin image<br>\n<img src=\"https://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing\" alt=\"\"><br>\n<a href=\"https://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing</a></p>\n<p>I did find out how to set the freq minimum for mel analysis but I can't find something that will do it for a simple power spectogram</p>\n<p><code>S = librosa.feature.melspectrogram(sound_array, sr=sample_rate, n_fft=n_fft, hop_length=hop_length, n_mels=n_mels, fmin=1536.0)</code></p>\n<p>Turns out some birds can hear infrasonic sounds. <br>\n<a href=\"https://escholarship.org/content/qt1kp2r437/qt1kp2r437_noSplash_8040d7507ac55f6c4c7582772f335f47.pdf?t=ptavtz\" target=\"_blank\">https://escholarship.org/content/qt1kp2r437/qt1kp2r437_noSplash_8040d7507ac55f6c4c7582772f335f47.pdf?t=ptavtz</a></p>\n<p>Does anyone know of a better way to get a reduction in the noise in the spectrogram (though the noise just may be due to the color ramp used)?</p>\n<p>2)  I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio. This seems like it would cause arbitrary issues with capturing the full call or song sequences. What is the reasoning for not doing analysis on the spectrogram of the full audio clip?</p>\n<p>I have at least one guess for why we need to chop into chunks:<br>\n      Images may be different sizes due to differing lengths or the patterns will get distorted if we try to make all the images the same size. Somehow the analysis tools are not happy when this happens. I guess we could solve this by making all the clips a certain size and those below the threshold throw away.</p>",
  "messages": [
    {
      "id": 1291447,
      "postDate": "2021-05-03T04:12:57.867Z",
      "content": "<p>I tried looking at some calls in Sonic Visualizer<br>\n<a href=\"https://www.sonicvisualiser.org/\" target=\"_blank\">https://www.sonicvisualiser.org/</a><br>\nAnd I worked <br>\nbkcchi -&gt; xc121068.ogg <br>\nthrough this tutorial:<br>\n<a href=\"https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\" target=\"_blank\">https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0</a><br>\nand Sonic Visualizer</p>\n<p>I find the mel spectrogram actually ends up increasing the amount of noise highlighted in the picture. The plain power spectrogram in sonic visualizer seems to work better for features. I also don't think we need the mel scaling because we don't actually care about human hearing. What's more important is the actual spectrogram not what we can hear, especially since birds can hear outside of our range. </p>\n<p>Questions:<br>\n1)  I can't seem to get librosa to give me the clean image I am can get in Sonic Visualizer using the Simple Power Spectrum plugin. </p>\n<p>Plugin image<br>\n<img src=\"https://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing\" alt=\"\"><br>\n<a href=\"https://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing</a></p>\n<p>I did find out how to set the freq minimum for mel analysis but I can't find something that will do it for a simple power spectogram</p>\n<p><code>S = librosa.feature.melspectrogram(sound_array, sr=sample_rate, n_fft=n_fft, hop_length=hop_length, n_mels=n_mels, fmin=1536.0)</code></p>\n<p>Turns out some birds can hear infrasonic sounds. <br>\n<a href=\"https://escholarship.org/content/qt1kp2r437/qt1kp2r437_noSplash_8040d7507ac55f6c4c7582772f335f47.pdf?t=ptavtz\" target=\"_blank\">https://escholarship.org/content/qt1kp2r437/qt1kp2r437_noSplash_8040d7507ac55f6c4c7582772f335f47.pdf?t=ptavtz</a></p>\n<p>Does anyone know of a better way to get a reduction in the noise in the spectrogram (though the noise just may be due to the color ramp used)?</p>\n<p>2)  I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio. This seems like it would cause arbitrary issues with capturing the full call or song sequences. What is the reasoning for not doing analysis on the spectrogram of the full audio clip?</p>\n<p>I have at least one guess for why we need to chop into chunks:<br>\n      Images may be different sizes due to differing lengths or the patterns will get distorted if we try to make all the images the same size. Somehow the analysis tools are not happy when this happens. I guess we could solve this by making all the clips a certain size and those below the threshold throw away.</p>",
      "rawMarkdown": "I tried looking at some calls in Sonic Visualizer\nhttps://www.sonicvisualiser.org/\nAnd I worked \nbkcchi -> xc121068.ogg \nthrough this tutorial:\nhttps://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\nand Sonic Visualizer\n\nI find the mel spectrogram actually ends up increasing the amount of noise highlighted in the picture. The plain power spectrogram in sonic visualizer seems to work better for features. I also don't think we need the mel scaling because we don't actually care about human hearing. What's more important is the actual spectrogram not what we can hear, especially since birds can hear outside of our range. \n\nQuestions:\n1)  I can't seem to get librosa to give me the clean image I am can get in Sonic Visualizer using the Simple Power Spectrum plugin. \n\nPlugin image\n![](https://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing)\nhttps://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing\n\nI did find out how to set the freq minimum for mel analysis but I can't find something that will do it for a simple power spectogram\n\n` S = librosa.feature.melspectrogram(sound_array, sr=sample_rate, n_fft=n_fft, hop_length=hop_length, n_mels=n_mels, fmin=1536.0)`\n\nTurns out some birds can hear infrasonic sounds. \nhttps://escholarship.org/content/qt1kp2r437/qt1kp2r437_noSplash_8040d7507ac55f6c4c7582772f335f47.pdf?t=ptavtz\n\nDoes anyone know of a better way to get a reduction in the noise in the spectrogram (though the noise just may be due to the color ramp used)?\n\n2)  I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio. This seems like it would cause arbitrary issues with capturing the full call or song sequences. What is the reasoning for not doing analysis on the spectrogram of the full audio clip?\n\nI have at least one guess for why we need to chop into chunks:\n      Images may be different sizes due to differing lengths or the patterns will get distorted if we try to make all the images the same size. Somehow the analysis tools are not happy when this happens. I guess we could solve this by making all the clips a certain size and those below the threshold throw away.\n   \n",
      "votes": 12
    },
    {
      "id": 1291831,
      "postDate": "2021-05-03T11:24:17.760Z",
      "content": "<blockquote>\n  <p>I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio.</p>\n</blockquote>\n<p>Probably because we are asked to predict on every 5 second period of train clips.  But, as was pointed out by host in a discussion in this forum (good idea to read everything hosts shared), models that predict using a larger window may be better.  One such model is the SED model used in some of the public notebooks.</p>\n<p>More generally, there is only one way to know which way is right: implement, test, and evaluate result.  </p>",
      "rawMarkdown": "> I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio.\n\nProbably because we are asked to predict on every 5 second period of train clips.  But, as was pointed out by host in a discussion in this forum (good idea to read everything hosts shared), models that predict using a larger window may be better.  One such model is the SED model used in some of the public notebooks.\n\nMore generally, there is only one way to know which way is right: implement, test, and evaluate result.  ",
      "votes": 3,
      "replies": [
        {
          "id": 1292261,
          "postDate": "2021-05-03T19:39:18.257Z",
          "content": "<p>Yeah I am in the middle of what you suggested but since this is a nice and helpful forum I was just trying to see the logic people were using. </p>\n<p>Thanks for the tips on the SED models - I am mostly the data munger, bird expert, and statistician NOT the Neural Net or audio/image recognition member on my team, so these discussions have been really good pointers for me.</p>",
          "rawMarkdown": "Yeah I am in the middle of what you suggested but since this is a nice and helpful forum I was just trying to see the logic people were using. \n\nThanks for the tips on the SED models - I am mostly the data munger, bird expert, and statistician NOT the Neural Net or audio/image recognition member on my team, so these discussions have been really good pointers for me."
        }
      ]
    },
    {
      "id": 1291703,
      "postDate": "2021-05-03T09:10:53.710Z",
      "content": "<blockquote>\n  <p>Does anyone know of a better way to get a reduction in the noise in the spectrogram (though the noise just may be due to the color ramp used)?</p>\n</blockquote>\n<p>I'll just outline a simple one: you can apply any modification to S you want, including power variation which will outline peaks better , sometimes these peaks may belong to someone screaming in a forest so I guess you have to filter them out. I've attached a simple example:</p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/61686934/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..v-__WzmQLNlIi67j4sdmNA.4Am33kZXvFZfrAPed0EzPFUf4C9J5MI03PNDnHpcpl-Qc3hMFPQXi4-dbhgOC5_zkZTlLax7WOqYeCdmLRf5WYlWABRWnXzOW-JCtOz9TEIiFv_GvwgJJEIa2jWthQrlWhJ32QJL185nqvnf165x7ihtM8sMCCempiNBf6tHtxDxBkEOBlB1aoxh-W72qyPO4Ea854XrJQkVne6BAbjsRwQa3HfZ9S9G94Fl_jzmjnGf-IART-RJfAxE7aTtdsMElq3BcqCj1Pe7F3GpxfpDLZXjaIGzRDspt9zRhRjbqe98_oDbQjlyf_FHQZ4wYaaXNRuohpeInj8lqUh7W6Xj8bvVmudgDoXgK7-Csupo6aSKMw3oBnA8hERBtKxuLXBC0-wi5mZpO1wQ26V3HQtbM1n66iypfO7cmDFoWHWtbm0iazMm9-nOVri7BJEzVHLgEzvdWLcsYA-HmwnTvHwjeS2S6PM07E-jWr9gxJURnGaQdBSnFmkcArUVQ9iUBw3rigOEKRpfgmzzeyf7j9aeM8J9lsrFlZlp_oVIGYp5PClvy56LWuZhHANm_7mBdzFV2aZ08aOagUrhlA2E_8eUXHBjOCQQYdQ29nOvGcgOAQfjegH4hx5PPqbLxsDB0fkg-XcV_g8C8AckvLXiC1BkLiaWxtmed1MpqDIx0Z4LewM.4LLR3OSpCWHpVgehRjuezQ/__results___files/__results___9_1.png\" alt=\"\"></p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/61686934/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..v-__WzmQLNlIi67j4sdmNA.4Am33kZXvFZfrAPed0EzPFUf4C9J5MI03PNDnHpcpl-Qc3hMFPQXi4-dbhgOC5_zkZTlLax7WOqYeCdmLRf5WYlWABRWnXzOW-JCtOz9TEIiFv_GvwgJJEIa2jWthQrlWhJ32QJL185nqvnf165x7ihtM8sMCCempiNBf6tHtxDxBkEOBlB1aoxh-W72qyPO4Ea854XrJQkVne6BAbjsRwQa3HfZ9S9G94Fl_jzmjnGf-IART-RJfAxE7aTtdsMElq3BcqCj1Pe7F3GpxfpDLZXjaIGzRDspt9zRhRjbqe98_oDbQjlyf_FHQZ4wYaaXNRuohpeInj8lqUh7W6Xj8bvVmudgDoXgK7-Csupo6aSKMw3oBnA8hERBtKxuLXBC0-wi5mZpO1wQ26V3HQtbM1n66iypfO7cmDFoWHWtbm0iazMm9-nOVri7BJEzVHLgEzvdWLcsYA-HmwnTvHwjeS2S6PM07E-jWr9gxJURnGaQdBSnFmkcArUVQ9iUBw3rigOEKRpfgmzzeyf7j9aeM8J9lsrFlZlp_oVIGYp5PClvy56LWuZhHANm_7mBdzFV2aZ08aOagUrhlA2E_8eUXHBjOCQQYdQ29nOvGcgOAQfjegH4hx5PPqbLxsDB0fkg-XcV_g8C8AckvLXiC1BkLiaWxtmed1MpqDIx0Z4LewM.4LLR3OSpCWHpVgehRjuezQ/__results___files/__results___10_0.png\" alt=\"\"></p>\n<blockquote>\n  <p>I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio.</p>\n</blockquote>\n<p>The hosts have decided that submission requires you to make a prediction of what is present ( if at all ) every 5 second interval. If you want to capture the entire call sequence you can load the audio with shifting for every 5 second segment. </p>",
      "rawMarkdown": "> Does anyone know of a better way to get a reduction in the noise in the spectrogram (though the noise just may be due to the color ramp used)?\n\nI'll just outline a simple one: you can apply any modification to S you want, including power variation which will outline peaks better , sometimes these peaks may belong to someone screaming in a forest so I guess you have to filter them out. I've attached a simple example:\n\n![](https://www.kaggleusercontent.com/kf/61686934/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..v-__WzmQLNlIi67j4sdmNA.4Am33kZXvFZfrAPed0EzPFUf4C9J5MI03PNDnHpcpl-Qc3hMFPQXi4-dbhgOC5_zkZTlLax7WOqYeCdmLRf5WYlWABRWnXzOW-JCtOz9TEIiFv_GvwgJJEIa2jWthQrlWhJ32QJL185nqvnf165x7ihtM8sMCCempiNBf6tHtxDxBkEOBlB1aoxh-W72qyPO4Ea854XrJQkVne6BAbjsRwQa3HfZ9S9G94Fl_jzmjnGf-IART-RJfAxE7aTtdsMElq3BcqCj1Pe7F3GpxfpDLZXjaIGzRDspt9zRhRjbqe98_oDbQjlyf_FHQZ4wYaaXNRuohpeInj8lqUh7W6Xj8bvVmudgDoXgK7-Csupo6aSKMw3oBnA8hERBtKxuLXBC0-wi5mZpO1wQ26V3HQtbM1n66iypfO7cmDFoWHWtbm0iazMm9-nOVri7BJEzVHLgEzvdWLcsYA-HmwnTvHwjeS2S6PM07E-jWr9gxJURnGaQdBSnFmkcArUVQ9iUBw3rigOEKRpfgmzzeyf7j9aeM8J9lsrFlZlp_oVIGYp5PClvy56LWuZhHANm_7mBdzFV2aZ08aOagUrhlA2E_8eUXHBjOCQQYdQ29nOvGcgOAQfjegH4hx5PPqbLxsDB0fkg-XcV_g8C8AckvLXiC1BkLiaWxtmed1MpqDIx0Z4LewM.4LLR3OSpCWHpVgehRjuezQ/__results___files/__results___9_1.png)\n\n![](https://www.kaggleusercontent.com/kf/61686934/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..v-__WzmQLNlIi67j4sdmNA.4Am33kZXvFZfrAPed0EzPFUf4C9J5MI03PNDnHpcpl-Qc3hMFPQXi4-dbhgOC5_zkZTlLax7WOqYeCdmLRf5WYlWABRWnXzOW-JCtOz9TEIiFv_GvwgJJEIa2jWthQrlWhJ32QJL185nqvnf165x7ihtM8sMCCempiNBf6tHtxDxBkEOBlB1aoxh-W72qyPO4Ea854XrJQkVne6BAbjsRwQa3HfZ9S9G94Fl_jzmjnGf-IART-RJfAxE7aTtdsMElq3BcqCj1Pe7F3GpxfpDLZXjaIGzRDspt9zRhRjbqe98_oDbQjlyf_FHQZ4wYaaXNRuohpeInj8lqUh7W6Xj8bvVmudgDoXgK7-Csupo6aSKMw3oBnA8hERBtKxuLXBC0-wi5mZpO1wQ26V3HQtbM1n66iypfO7cmDFoWHWtbm0iazMm9-nOVri7BJEzVHLgEzvdWLcsYA-HmwnTvHwjeS2S6PM07E-jWr9gxJURnGaQdBSnFmkcArUVQ9iUBw3rigOEKRpfgmzzeyf7j9aeM8J9lsrFlZlp_oVIGYp5PClvy56LWuZhHANm_7mBdzFV2aZ08aOagUrhlA2E_8eUXHBjOCQQYdQ29nOvGcgOAQfjegH4hx5PPqbLxsDB0fkg-XcV_g8C8AckvLXiC1BkLiaWxtmed1MpqDIx0Z4LewM.4LLR3OSpCWHpVgehRjuezQ/__results___files/__results___10_0.png)\n\n>  I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio.\n\nThe hosts have decided that submission requires you to make a prediction of what is present ( if at all ) every 5 second interval. If you want to capture the entire call sequence you can load the audio with shifting for every 5 second segment. \n",
      "votes": 3,
      "replies": [
        {
          "id": 1292284,
          "postDate": "2021-05-03T19:49:28.970Z",
          "content": "<p>Thanks so much. Right, in my remote sensing work you might use a Gaussian low pass filter to \"reduce\" noise or a high pass filter to find edges. I am just not as familiar with the techniques in audio processing and in Librosa in particular. </p>\n<p>I liked the power spectogram in sonic visualiser for highlighting calls and reducing noise but there is no way to automate the application and the source code is no longer available. Guess I will spend some time look at audio filtering and spectrogram rendering. </p>\n<p>Thanks again</p>",
          "rawMarkdown": "Thanks so much. Right, in my remote sensing work you might use a Gaussian low pass filter to \"reduce\" noise or a high pass filter to find edges. I am just not as familiar with the techniques in audio processing and in Librosa in particular. \n\nI liked the power spectogram in sonic visualiser for highlighting calls and reducing noise but there is no way to automate the application and the source code is no longer available. Guess I will spend some time look at audio filtering and spectrogram rendering. \n\nThanks again"
        }
      ]
    },
    {
      "id": 1291722,
      "postDate": "2021-05-03T09:33:26.553Z",
      "content": "<blockquote>\n  <p>I also don't think we need the mel scaling because we don't actually care about human hearing. What's more important is the actual spectrogram not what we can hear, especially since birds can hear outside of our range.</p>\n</blockquote>\n<p>We absolutely do care about human hearing, since we are comparing our model results to the annotations/labels given by human experts.</p>",
      "rawMarkdown": "> I also don't think we need the mel scaling because we don't actually care about human hearing. What's more important is the actual spectrogram not what we can hear, especially since birds can hear outside of our range.\n\nWe absolutely do care about human hearing, since we are comparing our model results to the annotations/labels given by human experts.",
      "votes": 1,
      "replies": [
        {
          "id": 1292219,
          "postDate": "2021-05-03T18:58:16.993Z",
          "content": "<p>However, I'm going to guess that the experts are probably looking at spectrograms (in addition to listening to the audio) while they annotate. And there's a decent chance they're looking at linear scales, as every publication of bird sounds I've looked at uses a linear scale, not a mel scale.</p>\n<p>Sources using a linear scale when displaying spectrograms include:  </p>\n<ul>\n<li><a href=\"http://www.thewarblerguide.com/\" target=\"_blank\">The Warbler Guide</a></li>\n<li><a href=\"https://academy.allaboutbirds.org/peterson-field-guide-to-bird-sounds/\" target=\"_blank\">Peterson Field Guide to the Birds Sounds of Eastern North America</a></li>\n<li><a href=\"https://www.macaulaylibrary.org/\" target=\"_blank\">Macaulay Library</a></li>\n<li><a href=\"https://www.xeno-canto.org/\" target=\"_blank\">xeno-canto</a></li>\n</ul>\n<p>Mel scales seem to be working for models; for example the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183269\" target=\"_blank\">2nd place solution from the Cornell Birdcall Identification</a> did note:</p>\n<blockquote>\n  <p>Spectrogram worked slightly worse than melspectrograms</p>\n</blockquote>\n<p>However, the question of why this would be, when linear scales are shown in so many publications, has been nagging at me. I have some ideas, but this comment is already long enough for now.</p>",
          "rawMarkdown": "However, I'm going to guess that the experts are probably looking at spectrograms (in addition to listening to the audio) while they annotate. And there's a decent chance they're looking at linear scales, as every publication of bird sounds I've looked at uses a linear scale, not a mel scale.\n\nSources using a linear scale when displaying spectrograms include:  \n* [The Warbler Guide](http://www.thewarblerguide.com/)\n* [Peterson Field Guide to the Birds Sounds of Eastern North America](https://academy.allaboutbirds.org/peterson-field-guide-to-bird-sounds/)\n* [Macaulay Library](https://www.macaulaylibrary.org/)\n* [xeno-canto](https://www.xeno-canto.org/)\n\nMel scales seem to be working for models; for example the [2nd place solution from the Cornell Birdcall Identification](https://www.kaggle.com/c/birdsong-recognition/discussion/183269) did note:\n> Spectrogram worked slightly worse than melspectrograms\n\nHowever, the question of why this would be, when linear scales are shown in so many publications, has been nagging at me. I have some ideas, but this comment is already long enough for now.\n",
          "votes": 5
        },
        {
          "id": 1292259,
          "postDate": "2021-05-03T19:36:55.593Z",
          "content": "<p>Not so, let me explain via an analogy. <br>\nIf you were trying to predict if someone had a heart attack, you would trust that the Dr. had rightly put the patients in the right category (heart attack or absence of heart attack). That doesn't mean you have to use the same methods the doctor used to diagnose the attack in your modeling. Usually you try to use all sort of information that doctor did not have access to. For example if you are experimenting on a new drug the doctor should not know if they patient was in the treatment or control group. </p>\n<p>For this competition, we are trusting that human interpreter correctly identifies the song in the audio and we don't even care about the method. The interpreter could even be deaf but actually see the bird they are recording and that would still be ok. We would still want to use all the information in the waveform/spectrogram, even if the recorder could not hear it.</p>\n<p>Hope that helps.</p>",
          "rawMarkdown": "Not so, let me explain via an analogy. \nIf you were trying to predict if someone had a heart attack, you would trust that the Dr. had rightly put the patients in the right category (heart attack or absence of heart attack). That doesn't mean you have to use the same methods the doctor used to diagnose the attack in your modeling. Usually you try to use all sort of information that doctor did not have access to. For example if you are experimenting on a new drug the doctor should not know if they patient was in the treatment or control group. \n\nFor this competition, we are trusting that human interpreter correctly identifies the song in the audio and we don't even care about the method. The interpreter could even be deaf but actually see the bird they are recording and that would still be ok. We would still want to use all the information in the waveform/spectrogram, even if the recorder could not hear it.\n\nHope that helps."
        },
        {
          "id": 1292278,
          "postDate": "2021-05-03T19:46:35.953Z",
          "content": "<blockquote>\n  <p>However, I'm going to guess that the experts are probably looking at spectrograms (in addition to listening to the audio) while they annotate. And there's a decent chance they're looking at linear scales, as every publication of bird sounds I've looked at uses a linear scale, not a mel scale.</p>\n</blockquote>\n<p>That's a good point that I had not thought of, thank you <a href=\"https://www.kaggle.com/jmreuter\" target=\"_blank\">@jmreuter</a>.</p>",
          "rawMarkdown": "> However, I'm going to guess that the experts are probably looking at spectrograms (in addition to listening to the audio) while they annotate. And there's a decent chance they're looking at linear scales, as every publication of bird sounds I've looked at uses a linear scale, not a mel scale.\n\nThat's a good point that I had not thought of, thank you @jmreuter."
        },
        {
          "id": 1292281,
          "postDate": "2021-05-03T19:48:27.797Z",
          "content": "<blockquote>\n  <p>For this competition, we are trusting that human interpreter correctly identifies the song in the audio.</p>\n</blockquote>\n<p>That definitely seems like one of the focal points in this competition; trying to predict what on earth these experts are paying attention to, there are so many misses in the training soundscapes, I've lost count.</p>",
          "rawMarkdown": "> For this competition, we are trusting that human interpreter correctly identifies the song in the audio.\n\nThat definitely seems like one of the focal points in this competition; trying to predict what on earth these experts are paying attention to, there are so many misses in the training soundscapes, I've lost count.",
          "votes": 2
        },
        {
          "id": 1292323,
          "postDate": "2021-05-03T20:45:27.890Z",
          "content": "<p>A couple thoughts on this front (entirely my own; just to add a couple different logs to the fire of discussion):</p>\n<ul>\n<li>Mel-scale (and related) capture the 'sorta-logarithmic-in-frequency' perceptual element. Birds have a similarly shaped audibility response curve as humans, though the sensitivity is greater.</li>\n<li>Log frequency scaling is 'natural' when you consider harmonics. The first harmonic for a frequency f is the octave, at frequency 2*f. If you want similar spacing between the fundamental and first harmonic regardless of f, you will reinvent log-scaling.</li>\n<li>Mel-scaling has an aggregation effect, as well, which should be helpful for reducing noisiness in the higher frequencies. It's maybe better to call it 'mel-pooling' in fact.</li>\n<li>The Petersen birdsong books have a separate spectrogram for owls, because it's hard to tell what's going on in the linear spectrogram. :)</li>\n</ul>\n<p>And as this all relates to neural networks:</p>\n<ul>\n<li>A log-frequency-scale should be helpful for reducing the variance of feature shapes in the low and high bands. If harmonics have roughly the same spacing at different parts of the spectrogram, input 2D-conv filters should be more reusable across the image.</li>\n<li>Any particular 'change-of-scaling' (ie, changing from mel-scale to bark-scale) can be emulated with a matrix multiplication. This should be very easy for a neural network to learn, if it needs to. (Given the mel-scaling matrix M, you can 'un-mel' with d·M^t, where d is a diagonal scaling by the inverse row-sums of M. M isn't invertible, but this works quite well in practice. Then rescaling by a different scaling matrix N gives you N·d·M^t, which reduces to a single matrix multiplication.)</li>\n<li>And this brings the nice philosophical question: What <strong>information</strong> is lost between the time-domain input and the scaled-spectrogram inputs that might be meaningful for classification? Does changing the frontend give the network <em>different</em> information, or does it just change the arrangement of what's already there?</li>\n</ul>\n<p>I believe there's some good prior work on spectrogram settings in past birdclef challenge writeups, as well. I remember the PANNs paper basically concludes that increasing time resolution (by reducing STFT hop size) is the best thing you can do, though it's obviously expensive.</p>",
          "rawMarkdown": "A couple thoughts on this front (entirely my own; just to add a couple different logs to the fire of discussion):\n\n- Mel-scale (and related) capture the 'sorta-logarithmic-in-frequency' perceptual element. Birds have a similarly shaped audibility response curve as humans, though the sensitivity is greater.\n- Log frequency scaling is 'natural' when you consider harmonics. The first harmonic for a frequency f is the octave, at frequency 2*f. If you want similar spacing between the fundamental and first harmonic regardless of f, you will reinvent log-scaling.\n- Mel-scaling has an aggregation effect, as well, which should be helpful for reducing noisiness in the higher frequencies. It's maybe better to call it 'mel-pooling' in fact.\n- The Petersen birdsong books have a separate spectrogram for owls, because it's hard to tell what's going on in the linear spectrogram. :)\n\nAnd as this all relates to neural networks:\n\n- A log-frequency-scale should be helpful for reducing the variance of feature shapes in the low and high bands. If harmonics have roughly the same spacing at different parts of the spectrogram, input 2D-conv filters should be more reusable across the image.\n- Any particular 'change-of-scaling' (ie, changing from mel-scale to bark-scale) can be emulated with a matrix multiplication. This should be very easy for a neural network to learn, if it needs to. (Given the mel-scaling matrix M, you can 'un-mel' with d·M^t, where d is a diagonal scaling by the inverse row-sums of M. M isn't invertible, but this works quite well in practice. Then rescaling by a different scaling matrix N gives you N·d·M^t, which reduces to a single matrix multiplication.)\n- And this brings the nice philosophical question: What **information** is lost between the time-domain input and the scaled-spectrogram inputs that might be meaningful for classification? Does changing the frontend give the network *different* information, or does it just change the arrangement of what's already there?\n\nI believe there's some good prior work on spectrogram settings in past birdclef challenge writeups, as well. I remember the PANNs paper basically concludes that increasing time resolution (by reducing STFT hop size) is the best thing you can do, though it's obviously expensive.",
          "votes": 7
        },
        {
          "id": 1292327,
          "postDate": "2021-05-03T20:48:39.507Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1292340,
          "postDate": "2021-05-03T20:56:46.060Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1292377,
          "postDate": "2021-05-03T21:43:38.287Z",
          "content": "<p>Really great info, thanks.<br>\nNow with all this new informatio, playing more with sonic visualiser, and working the call through this tutorial<br>\n<a href=\"https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\" target=\"_blank\">https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0</a></p>\n<p>I think the problem with running the chickadee call through that tutorial and saying the mel was not as good was a problem with needing to set a higher energy level for the values that appear black in the diagram. In sonic visualiser this would be \"lowering the gain\".  In other words - this would make the spectrogram image fed to the NN \"cleaner\".</p>\n<p>What has also been really interesting for me, coming from the GIS/spatial world is getting my head around looking at spectrograms and translating that to what I hear. For example <br>\nNorthern Flicker in NY<br>\n<a href=\"https://macaulaylibrary.org/asset/30597061\" target=\"_blank\">https://macaulaylibrary.org/asset/30597061</a><br>\nNorthern Flicker in Guatemala <br>\n<a href=\"https://macaulaylibrary.org/asset/523558\" target=\"_blank\">https://macaulaylibrary.org/asset/523558</a></p>\n<p>The higher pitch of the Guatemalan flicker is, I think, shown by the darker color (more energy) in the higher frequencies. I know this assumes that the energy shades are scaled the same between the two spectrograms. </p>\n<p>I know most of this seems obvious to people who have worked with sounds or spectrograms a lot, but it has taken the back and forth between hands on , reading, and discussion to start to see the obvious. </p>\n<p>Really fun stuff and I am getting a bit obsessive. </p>\n<p>I do think, if people choose to go the spectrogram with image recognition route, the production of the images may be just as or even more important than the modeling (assuming you get the general class of models correct)</p>",
          "rawMarkdown": "Really great info, thanks.\nNow with all this new informatio, playing more with sonic visualiser, and working the call through this tutorial\nhttps://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\n\nI think the problem with running the chickadee call through that tutorial and saying the mel was not as good was a problem with needing to set a higher energy level for the values that appear black in the diagram. In sonic visualiser this would be \"lowering the gain\".  In other words - this would make the spectrogram image fed to the NN \"cleaner\".\n\nWhat has also been really interesting for me, coming from the GIS/spatial world is getting my head around looking at spectrograms and translating that to what I hear. For example \nNorthern Flicker in NY\nhttps://macaulaylibrary.org/asset/30597061\nNorthern Flicker in Guatemala \nhttps://macaulaylibrary.org/asset/523558\n\nThe higher pitch of the Guatemalan flicker is, I think, shown by the darker color (more energy) in the higher frequencies. I know this assumes that the energy shades are scaled the same between the two spectrograms. \n\nI know most of this seems obvious to people who have worked with sounds or spectrograms a lot, but it has taken the back and forth between hands on , reading, and discussion to start to see the obvious. \n\nReally fun stuff and I am getting a bit obsessive. \n\nI do think, if people choose to go the spectrogram with image recognition route, the production of the images may be just as or even more important than the modeling (assuming you get the general class of models correct)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1490719,
      "postDate": "2021-08-25T20:00:52.233Z",
      "content": "<p><a href=\"https://github.com/Christoph-Lauer/Sonogram\" target=\"_blank\">Sonogram visible speech</a></p>",
      "rawMarkdown": "[Sonogram visible speech](https://github.com/Christoph-Lauer/Sonogram)"
    }
  ],
  "comments": [
    {
      "id": 1291831,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-05-03T11:24:17.760000",
      "content": "<blockquote>\n  <p>I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio.</p>\n</blockquote>\n<p>Probably because we are asked to predict on every 5 second period of train clips.  But, as was pointed out by host in a discussion in this forum (good idea to read everything hosts shared), models that predict using a larger window may be better.  One such model is the SED model used in some of the public notebooks.</p>\n<p>More generally, there is only one way to know which way is right: implement, test, and evaluate result.  </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1292261,
          "author_name": "Steve Pousty",
          "author_url": "",
          "post_date": "2021-05-03T19:39:18.257000",
          "content": "<p>Yeah I am in the middle of what you suggested but since this is a nice and helpful forum I was just trying to see the logic people were using. </p>\n<p>Thanks for the tips on the SED models - I am mostly the data munger, bird expert, and statistician NOT the Neural Net or audio/image recognition member on my team, so these discussions have been really good pointers for me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1291703,
      "author_name": "Andrey Shtrauss",
      "author_url": "",
      "post_date": "2021-05-03T09:10:53.710000",
      "content": "<blockquote>\n  <p>Does anyone know of a better way to get a reduction in the noise in the spectrogram (though the noise just may be due to the color ramp used)?</p>\n</blockquote>\n<p>I'll just outline a simple one: you can apply any modification to S you want, including power variation which will outline peaks better , sometimes these peaks may belong to someone screaming in a forest so I guess you have to filter them out. I've attached a simple example:</p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/61686934/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..v-__WzmQLNlIi67j4sdmNA.4Am33kZXvFZfrAPed0EzPFUf4C9J5MI03PNDnHpcpl-Qc3hMFPQXi4-dbhgOC5_zkZTlLax7WOqYeCdmLRf5WYlWABRWnXzOW-JCtOz9TEIiFv_GvwgJJEIa2jWthQrlWhJ32QJL185nqvnf165x7ihtM8sMCCempiNBf6tHtxDxBkEOBlB1aoxh-W72qyPO4Ea854XrJQkVne6BAbjsRwQa3HfZ9S9G94Fl_jzmjnGf-IART-RJfAxE7aTtdsMElq3BcqCj1Pe7F3GpxfpDLZXjaIGzRDspt9zRhRjbqe98_oDbQjlyf_FHQZ4wYaaXNRuohpeInj8lqUh7W6Xj8bvVmudgDoXgK7-Csupo6aSKMw3oBnA8hERBtKxuLXBC0-wi5mZpO1wQ26V3HQtbM1n66iypfO7cmDFoWHWtbm0iazMm9-nOVri7BJEzVHLgEzvdWLcsYA-HmwnTvHwjeS2S6PM07E-jWr9gxJURnGaQdBSnFmkcArUVQ9iUBw3rigOEKRpfgmzzeyf7j9aeM8J9lsrFlZlp_oVIGYp5PClvy56LWuZhHANm_7mBdzFV2aZ08aOagUrhlA2E_8eUXHBjOCQQYdQ29nOvGcgOAQfjegH4hx5PPqbLxsDB0fkg-XcV_g8C8AckvLXiC1BkLiaWxtmed1MpqDIx0Z4LewM.4LLR3OSpCWHpVgehRjuezQ/__results___files/__results___9_1.png\" alt=\"\"></p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/61686934/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..v-__WzmQLNlIi67j4sdmNA.4Am33kZXvFZfrAPed0EzPFUf4C9J5MI03PNDnHpcpl-Qc3hMFPQXi4-dbhgOC5_zkZTlLax7WOqYeCdmLRf5WYlWABRWnXzOW-JCtOz9TEIiFv_GvwgJJEIa2jWthQrlWhJ32QJL185nqvnf165x7ihtM8sMCCempiNBf6tHtxDxBkEOBlB1aoxh-W72qyPO4Ea854XrJQkVne6BAbjsRwQa3HfZ9S9G94Fl_jzmjnGf-IART-RJfAxE7aTtdsMElq3BcqCj1Pe7F3GpxfpDLZXjaIGzRDspt9zRhRjbqe98_oDbQjlyf_FHQZ4wYaaXNRuohpeInj8lqUh7W6Xj8bvVmudgDoXgK7-Csupo6aSKMw3oBnA8hERBtKxuLXBC0-wi5mZpO1wQ26V3HQtbM1n66iypfO7cmDFoWHWtbm0iazMm9-nOVri7BJEzVHLgEzvdWLcsYA-HmwnTvHwjeS2S6PM07E-jWr9gxJURnGaQdBSnFmkcArUVQ9iUBw3rigOEKRpfgmzzeyf7j9aeM8J9lsrFlZlp_oVIGYp5PClvy56LWuZhHANm_7mBdzFV2aZ08aOagUrhlA2E_8eUXHBjOCQQYdQ29nOvGcgOAQfjegH4hx5PPqbLxsDB0fkg-XcV_g8C8AckvLXiC1BkLiaWxtmed1MpqDIx0Z4LewM.4LLR3OSpCWHpVgehRjuezQ/__results___files/__results___10_0.png\" alt=\"\"></p>\n<blockquote>\n  <p>I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio.</p>\n</blockquote>\n<p>The hosts have decided that submission requires you to make a prediction of what is present ( if at all ) every 5 second interval. If you want to capture the entire call sequence you can load the audio with shifting for every 5 second segment. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1292284,
          "author_name": "Steve Pousty",
          "author_url": "",
          "post_date": "2021-05-03T19:49:28.970000",
          "content": "<p>Thanks so much. Right, in my remote sensing work you might use a Gaussian low pass filter to \"reduce\" noise or a high pass filter to find edges. I am just not as familiar with the techniques in audio processing and in Librosa in particular. </p>\n<p>I liked the power spectogram in sonic visualiser for highlighting calls and reducing noise but there is no way to automate the application and the source code is no longer available. Guess I will spend some time look at audio filtering and spectrogram rendering. </p>\n<p>Thanks again</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1291722,
      "author_name": "Agneev",
      "author_url": "",
      "post_date": "2021-05-03T09:33:26.553000",
      "content": "<blockquote>\n  <p>I also don't think we need the mel scaling because we don't actually care about human hearing. What's more important is the actual spectrogram not what we can hear, especially since birds can hear outside of our range.</p>\n</blockquote>\n<p>We absolutely do care about human hearing, since we are comparing our model results to the annotations/labels given by human experts.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1292219,
          "author_name": "undisclosed",
          "author_url": "",
          "post_date": "2021-05-03T18:58:16.993000",
          "content": "<p>However, I'm going to guess that the experts are probably looking at spectrograms (in addition to listening to the audio) while they annotate. And there's a decent chance they're looking at linear scales, as every publication of bird sounds I've looked at uses a linear scale, not a mel scale.</p>\n<p>Sources using a linear scale when displaying spectrograms include:  </p>\n<ul>\n<li><a href=\"http://www.thewarblerguide.com/\" target=\"_blank\">The Warbler Guide</a></li>\n<li><a href=\"https://academy.allaboutbirds.org/peterson-field-guide-to-bird-sounds/\" target=\"_blank\">Peterson Field Guide to the Birds Sounds of Eastern North America</a></li>\n<li><a href=\"https://www.macaulaylibrary.org/\" target=\"_blank\">Macaulay Library</a></li>\n<li><a href=\"https://www.xeno-canto.org/\" target=\"_blank\">xeno-canto</a></li>\n</ul>\n<p>Mel scales seem to be working for models; for example the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183269\" target=\"_blank\">2nd place solution from the Cornell Birdcall Identification</a> did note:</p>\n<blockquote>\n  <p>Spectrogram worked slightly worse than melspectrograms</p>\n</blockquote>\n<p>However, the question of why this would be, when linear scales are shown in so many publications, has been nagging at me. I have some ideas, but this comment is already long enough for now.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1292259,
          "author_name": "Steve Pousty",
          "author_url": "",
          "post_date": "2021-05-03T19:36:55.593000",
          "content": "<p>Not so, let me explain via an analogy. <br>\nIf you were trying to predict if someone had a heart attack, you would trust that the Dr. had rightly put the patients in the right category (heart attack or absence of heart attack). That doesn't mean you have to use the same methods the doctor used to diagnose the attack in your modeling. Usually you try to use all sort of information that doctor did not have access to. For example if you are experimenting on a new drug the doctor should not know if they patient was in the treatment or control group. </p>\n<p>For this competition, we are trusting that human interpreter correctly identifies the song in the audio and we don't even care about the method. The interpreter could even be deaf but actually see the bird they are recording and that would still be ok. We would still want to use all the information in the waveform/spectrogram, even if the recorder could not hear it.</p>\n<p>Hope that helps.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1292278,
          "author_name": "Agneev",
          "author_url": "",
          "post_date": "2021-05-03T19:46:35.953000",
          "content": "<blockquote>\n  <p>However, I'm going to guess that the experts are probably looking at spectrograms (in addition to listening to the audio) while they annotate. And there's a decent chance they're looking at linear scales, as every publication of bird sounds I've looked at uses a linear scale, not a mel scale.</p>\n</blockquote>\n<p>That's a good point that I had not thought of, thank you <a href=\"https://www.kaggle.com/jmreuter\" target=\"_blank\">@jmreuter</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1292281,
          "author_name": "Andrey Shtrauss",
          "author_url": "",
          "post_date": "2021-05-03T19:48:27.797000",
          "content": "<blockquote>\n  <p>For this competition, we are trusting that human interpreter correctly identifies the song in the audio.</p>\n</blockquote>\n<p>That definitely seems like one of the focal points in this competition; trying to predict what on earth these experts are paying attention to, there are so many misses in the training soundscapes, I've lost count.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1292323,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2021-05-03T20:45:27.890000",
          "content": "<p>A couple thoughts on this front (entirely my own; just to add a couple different logs to the fire of discussion):</p>\n<ul>\n<li>Mel-scale (and related) capture the 'sorta-logarithmic-in-frequency' perceptual element. Birds have a similarly shaped audibility response curve as humans, though the sensitivity is greater.</li>\n<li>Log frequency scaling is 'natural' when you consider harmonics. The first harmonic for a frequency f is the octave, at frequency 2*f. If you want similar spacing between the fundamental and first harmonic regardless of f, you will reinvent log-scaling.</li>\n<li>Mel-scaling has an aggregation effect, as well, which should be helpful for reducing noisiness in the higher frequencies. It's maybe better to call it 'mel-pooling' in fact.</li>\n<li>The Petersen birdsong books have a separate spectrogram for owls, because it's hard to tell what's going on in the linear spectrogram. :)</li>\n</ul>\n<p>And as this all relates to neural networks:</p>\n<ul>\n<li>A log-frequency-scale should be helpful for reducing the variance of feature shapes in the low and high bands. If harmonics have roughly the same spacing at different parts of the spectrogram, input 2D-conv filters should be more reusable across the image.</li>\n<li>Any particular 'change-of-scaling' (ie, changing from mel-scale to bark-scale) can be emulated with a matrix multiplication. This should be very easy for a neural network to learn, if it needs to. (Given the mel-scaling matrix M, you can 'un-mel' with d·M^t, where d is a diagonal scaling by the inverse row-sums of M. M isn't invertible, but this works quite well in practice. Then rescaling by a different scaling matrix N gives you N·d·M^t, which reduces to a single matrix multiplication.)</li>\n<li>And this brings the nice philosophical question: What <strong>information</strong> is lost between the time-domain input and the scaled-spectrogram inputs that might be meaningful for classification? Does changing the frontend give the network <em>different</em> information, or does it just change the arrangement of what's already there?</li>\n</ul>\n<p>I believe there's some good prior work on spectrogram settings in past birdclef challenge writeups, as well. I remember the PANNs paper basically concludes that increasing time resolution (by reducing STFT hop size) is the best thing you can do, though it's obviously expensive.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1292327,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T20:48:39.507000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1292340,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T20:56:46.060000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1292377,
          "author_name": "Steve Pousty",
          "author_url": "",
          "post_date": "2021-05-03T21:43:38.287000",
          "content": "<p>Really great info, thanks.<br>\nNow with all this new informatio, playing more with sonic visualiser, and working the call through this tutorial<br>\n<a href=\"https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\" target=\"_blank\">https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0</a></p>\n<p>I think the problem with running the chickadee call through that tutorial and saying the mel was not as good was a problem with needing to set a higher energy level for the values that appear black in the diagram. In sonic visualiser this would be \"lowering the gain\".  In other words - this would make the spectrogram image fed to the NN \"cleaner\".</p>\n<p>What has also been really interesting for me, coming from the GIS/spatial world is getting my head around looking at spectrograms and translating that to what I hear. For example <br>\nNorthern Flicker in NY<br>\n<a href=\"https://macaulaylibrary.org/asset/30597061\" target=\"_blank\">https://macaulaylibrary.org/asset/30597061</a><br>\nNorthern Flicker in Guatemala <br>\n<a href=\"https://macaulaylibrary.org/asset/523558\" target=\"_blank\">https://macaulaylibrary.org/asset/523558</a></p>\n<p>The higher pitch of the Guatemalan flicker is, I think, shown by the darker color (more energy) in the higher frequencies. I know this assumes that the energy shades are scaled the same between the two spectrograms. </p>\n<p>I know most of this seems obvious to people who have worked with sounds or spectrograms a lot, but it has taken the back and forth between hands on , reading, and discussion to start to see the obvious. </p>\n<p>Really fun stuff and I am getting a bit obsessive. </p>\n<p>I do think, if people choose to go the spectrogram with image recognition route, the production of the images may be just as or even more important than the modeling (assuming you get the general class of models correct)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1490719,
      "author_name": "ChristophLauer",
      "author_url": "",
      "post_date": "2021-08-25T20:00:52.233000",
      "content": "<p><a href=\"https://github.com/Christoph-Lauer/Sonogram\" target=\"_blank\">Sonogram visible speech</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1291447": "I tried looking at some calls in Sonic Visualizer\nhttps://www.sonicvisualiser.org/\nAnd I worked \nbkcchi -> xc121068.ogg \nthrough this tutorial:\nhttps://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\nand Sonic Visualizer\n\nI find the mel spectrogram actually ends up increasing the amount of noise highlighted in the picture. The plain power spectrogram in sonic visualizer seems to work better for features. I also don't think we need the mel scaling because we don't actually care about human hearing. What's more important is the actual spectrogram not what we can hear, especially since birds can hear outside of our range. \n\nQuestions:\n1)  I can't seem to get librosa to give me the clean image I am can get in Sonic Visualizer using the Simple Power Spectrum plugin. \n\nPlugin image\n![](https://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing)\nhttps://drive.google.com/file/d/1mTuk2yq_6NUqYOYVrx81jg7M89r5TnIb/view?usp=sharing\n\nI did find out how to set the freq minimum for mel analysis but I can't find something that will do it for a simple power spectogram\n\n` S = librosa.feature.melspectrogram(sound_array, sr=sample_rate, n_fft=n_fft, hop_length=hop_length, n_mels=n_mels, fmin=1536.0)`\n\nTurns out some birds can hear infrasonic sounds. \nhttps://escholarship.org/content/qt1kp2r437/qt1kp2r437_noSplash_8040d7507ac55f6c4c7582772f335f47.pdf?t=ptavtz\n\nDoes anyone know of a better way to get a reduction in the noise in the spectrogram (though the noise just may be due to the color ramp used)?\n\n2)  I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio. This seems like it would cause arbitrary issues with capturing the full call or song sequences. What is the reasoning for not doing analysis on the spectrogram of the full audio clip?\n\nI have at least one guess for why we need to chop into chunks:\n      Images may be different sizes due to differing lengths or the patterns will get distorted if we try to make all the images the same size. Somehow the analysis tools are not happy when this happens. I guess we could solve this by making all the clips a certain size and those below the threshold throw away.\n   \n",
    "1291831": "> I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio.\n\nProbably because we are asked to predict on every 5 second period of train clips.  But, as was pointed out by host in a discussion in this forum (good idea to read everything hosts shared), models that predict using a larger window may be better.  One such model is the SED model used in some of the public notebooks.\n\nMore generally, there is only one way to know which way is right: implement, test, and evaluate result.  ",
    "1291703": "> Does anyone know of a better way to get a reduction in the noise in the spectrogram (though the noise just may be due to the color ramp used)?\n\nI'll just outline a simple one: you can apply any modification to S you want, including power variation which will outline peaks better , sometimes these peaks may belong to someone screaming in a forest so I guess you have to filter them out. I've attached a simple example:\n\n![](https://www.kaggleusercontent.com/kf/61686934/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..v-__WzmQLNlIi67j4sdmNA.4Am33kZXvFZfrAPed0EzPFUf4C9J5MI03PNDnHpcpl-Qc3hMFPQXi4-dbhgOC5_zkZTlLax7WOqYeCdmLRf5WYlWABRWnXzOW-JCtOz9TEIiFv_GvwgJJEIa2jWthQrlWhJ32QJL185nqvnf165x7ihtM8sMCCempiNBf6tHtxDxBkEOBlB1aoxh-W72qyPO4Ea854XrJQkVne6BAbjsRwQa3HfZ9S9G94Fl_jzmjnGf-IART-RJfAxE7aTtdsMElq3BcqCj1Pe7F3GpxfpDLZXjaIGzRDspt9zRhRjbqe98_oDbQjlyf_FHQZ4wYaaXNRuohpeInj8lqUh7W6Xj8bvVmudgDoXgK7-Csupo6aSKMw3oBnA8hERBtKxuLXBC0-wi5mZpO1wQ26V3HQtbM1n66iypfO7cmDFoWHWtbm0iazMm9-nOVri7BJEzVHLgEzvdWLcsYA-HmwnTvHwjeS2S6PM07E-jWr9gxJURnGaQdBSnFmkcArUVQ9iUBw3rigOEKRpfgmzzeyf7j9aeM8J9lsrFlZlp_oVIGYp5PClvy56LWuZhHANm_7mBdzFV2aZ08aOagUrhlA2E_8eUXHBjOCQQYdQ29nOvGcgOAQfjegH4hx5PPqbLxsDB0fkg-XcV_g8C8AckvLXiC1BkLiaWxtmed1MpqDIx0Z4LewM.4LLR3OSpCWHpVgehRjuezQ/__results___files/__results___9_1.png)\n\n![](https://www.kaggleusercontent.com/kf/61686934/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..v-__WzmQLNlIi67j4sdmNA.4Am33kZXvFZfrAPed0EzPFUf4C9J5MI03PNDnHpcpl-Qc3hMFPQXi4-dbhgOC5_zkZTlLax7WOqYeCdmLRf5WYlWABRWnXzOW-JCtOz9TEIiFv_GvwgJJEIa2jWthQrlWhJ32QJL185nqvnf165x7ihtM8sMCCempiNBf6tHtxDxBkEOBlB1aoxh-W72qyPO4Ea854XrJQkVne6BAbjsRwQa3HfZ9S9G94Fl_jzmjnGf-IART-RJfAxE7aTtdsMElq3BcqCj1Pe7F3GpxfpDLZXjaIGzRDspt9zRhRjbqe98_oDbQjlyf_FHQZ4wYaaXNRuohpeInj8lqUh7W6Xj8bvVmudgDoXgK7-Csupo6aSKMw3oBnA8hERBtKxuLXBC0-wi5mZpO1wQ26V3HQtbM1n66iypfO7cmDFoWHWtbm0iazMm9-nOVri7BJEzVHLgEzvdWLcsYA-HmwnTvHwjeS2S6PM07E-jWr9gxJURnGaQdBSnFmkcArUVQ9iUBw3rigOEKRpfgmzzeyf7j9aeM8J9lsrFlZlp_oVIGYp5PClvy56LWuZhHANm_7mBdzFV2aZ08aOagUrhlA2E_8eUXHBjOCQQYdQ29nOvGcgOAQfjegH4hx5PPqbLxsDB0fkg-XcV_g8C8AckvLXiC1BkLiaWxtmed1MpqDIx0Z4LewM.4LLR3OSpCWHpVgehRjuezQ/__results___files/__results___10_0.png)\n\n>  I looked around and can't find a reason people are suggesting breaking the spectrogram into 5 second clips of audio.\n\nThe hosts have decided that submission requires you to make a prediction of what is present ( if at all ) every 5 second interval. If you want to capture the entire call sequence you can load the audio with shifting for every 5 second segment. \n",
    "1291722": "> I also don't think we need the mel scaling because we don't actually care about human hearing. What's more important is the actual spectrogram not what we can hear, especially since birds can hear outside of our range.\n\nWe absolutely do care about human hearing, since we are comparing our model results to the annotations/labels given by human experts.",
    "1490719": "[Sonogram visible speech](https://github.com/Christoph-Lauer/Sonogram)"
  }
}