{
  "id": 236502,
  "title": "CNN settings and structure to optimize pattern matching for bird sounds?",
  "url": "/competitions/birdclef-2021/discussion/236502",
  "author_name": "",
  "post_date": "2021-05-04T16:22:36.706988900Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is my first time working with neural networks, so please kindly correct me if I botch terminology or concepts. </p>\n<p>My understanding is that we’re borrowing a toolset developed largely for image recognition, where the goal is typically to recognize patterns regardless of the location in an image. But there is an asymmetry with audio spectrograms that doesn’t apply to generic image recognition. </p>\n<p>In audio, a horizontal (time) shift regarding where a target pattern is located is meaningless.</p>\n<p>In contrast, for bird sounds, a vertical (frequency) shift usually matters quite a bit. For example, the vertical placement on a spectrogram should be enough to differentiate a really high pitched Blackpoll Warbler trill from a lower pitched Worm-eating Warbler trill. The frequency distinction should hold up even if the soundscape was noisy and the bird was distant, so that the finer details of the trill aren’t apparent. </p>\n<p>This asymmetry seems like something that could be leveraged to minimize mismatches, but I’m not sure how to best do so in practice. Perhaps by tweaking max pooling windows? </p>\n<p>I assume this topic has been explored, but I’m having difficulty finding a treatment of it (or at least one that is intelligible to a beginner). </p>\n<p>In practice, what strategies might be used to maximize sensitivity to vertical pattern positioning and minimize sensitivity to horizontal pattern positioning for the benefit of bird sound identification? </p>",
  "messages": [
    {
      "id": "1293197",
      "postDate": "05/04/2021 16:22:36",
      "content": "<p>This is my first time working with neural networks, so please kindly correct me if I botch terminology or concepts. </p>\n<p>My understanding is that we’re borrowing a toolset developed largely for image recognition, where the goal is typically to recognize patterns regardless of the location in an image. But there is an asymmetry with audio spectrograms that doesn’t apply to generic image recognition. </p>\n<p>In audio, a horizontal (time) shift regarding where a target pattern is located is meaningless.</p>\n<p>In contrast, for bird sounds, a vertical (frequency) shift usually matters quite a bit. For example, the vertical placement on a spectrogram should be enough to differentiate a really high pitched Blackpoll Warbler trill from a lower pitched Worm-eating Warbler trill. The frequency distinction should hold up even if the soundscape was noisy and the bird was distant, so that the finer details of the trill aren’t apparent. </p>\n<p>This asymmetry seems like something that could be leveraged to minimize mismatches, but I’m not sure how to best do so in practice. Perhaps by tweaking max pooling windows? </p>\n<p>I assume this topic has been explored, but I’m having difficulty finding a treatment of it (or at least one that is intelligible to a beginner). </p>\n<p>In practice, what strategies might be used to maximize sensitivity to vertical pattern positioning and minimize sensitivity to horizontal pattern positioning for the benefit of bird sound identification? </p>",
      "rawMarkdown": "This is my first time working with neural networks, so please kindly correct me if I botch terminology or concepts. \n\nMy understanding is that we’re borrowing a toolset developed largely for image recognition, where the goal is typically to recognize patterns regardless of the location in an image. But there is an asymmetry with audio spectrograms that doesn’t apply to generic image recognition. \n\nIn audio, a horizontal (time) shift regarding where a target pattern is located is meaningless.\n\nIn contrast, for bird sounds, a vertical (frequency) shift usually matters quite a bit. For example, the vertical placement on a spectrogram should be enough to differentiate a really high pitched Blackpoll Warbler trill from a lower pitched Worm-eating Warbler trill. The frequency distinction should hold up even if the soundscape was noisy and the bird was distant, so that the finer details of the trill aren’t apparent. \n\nThis asymmetry seems like something that could be leveraged to minimize mismatches, but I’m not sure how to best do so in practice. Perhaps by tweaking max pooling windows? \n\nI assume this topic has been explored, but I’m having difficulty finding a treatment of it (or at least one that is intelligible to a beginner). \n\nIn practice, what strategies might be used to maximize sensitivity to vertical pattern positioning and minimize sensitivity to horizontal pattern positioning for the benefit of bird sound identification?",
      "votes": null
    },
    {
      "id": "1294078",
      "postDate": "05/05/2021 11:57:19",
      "content": "<p>I have not found an answer to my question. However, I did come across this article, which points out several additional differences between spectrograms and visual images, and ultimately implies that CNNs may not be the best approach for audio analysis in the long run: </p>\n<p><a href=\"https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\" target=\"_blank\">https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd</a></p>\n<p>Thoughts? Agreement? Disagreement? Practical advice/strategies for workarounds regarding those issues in the short term? </p>",
      "rawMarkdown": "I have not found an answer to my question. However, I did come across this article, which points out several additional differences between spectrograms and visual images, and ultimately implies that CNNs may not be the best approach for audio analysis in the long run: \n\nhttps://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\n\nThoughts? Agreement? Disagreement? Practical advice/strategies for workarounds regarding those issues in the short term?",
      "votes": null
    },
    {
      "id": "1295241",
      "postDate": "05/06/2021 10:08:50",
      "content": "<blockquote>\n  <p>I assume this topic has been explored, but I’m having difficulty finding a treatment of it (or at least one that is intelligible to a beginner). </p>\n</blockquote>\n<p>I gave you a way to find some : <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/234464\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/234464</a></p>",
      "rawMarkdown": "> I assume this topic has been explored, but I’m having difficulty finding a treatment of it (or at least one that is intelligible to a beginner). \n\nI gave you a way to find some : https://www.kaggle.com/c/birdclef-2021/discussion/234464",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1294078,
      "author_name": "jmreuter",
      "author_url": "",
      "post_date": "05/05/2021 11:57:19",
      "content": "<p>I have not found an answer to my question. However, I did come across this article, which points out several additional differences between spectrograms and visual images, and ultimately implies that CNNs may not be the best approach for audio analysis in the long run: </p>\n<p><a href=\"https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\" target=\"_blank\">https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd</a></p>\n<p>Thoughts? Agreement? Disagreement? Practical advice/strategies for workarounds regarding those issues in the short term? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1295241,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/06/2021 10:08:50",
          "content": "<blockquote>\n  <p>I assume this topic has been explored, but I’m having difficulty finding a treatment of it (or at least one that is intelligible to a beginner). </p>\n</blockquote>\n<p>I gave you a way to find some : <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/234464\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/234464</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1293197": "This is my first time working with neural networks, so please kindly correct me if I botch terminology or concepts. \n\nMy understanding is that we’re borrowing a toolset developed largely for image recognition, where the goal is typically to recognize patterns regardless of the location in an image. But there is an asymmetry with audio spectrograms that doesn’t apply to generic image recognition. \n\nIn audio, a horizontal (time) shift regarding where a target pattern is located is meaningless.\n\nIn contrast, for bird sounds, a vertical (frequency) shift usually matters quite a bit. For example, the vertical placement on a spectrogram should be enough to differentiate a really high pitched Blackpoll Warbler trill from a lower pitched Worm-eating Warbler trill. The frequency distinction should hold up even if the soundscape was noisy and the bird was distant, so that the finer details of the trill aren’t apparent. \n\nThis asymmetry seems like something that could be leveraged to minimize mismatches, but I’m not sure how to best do so in practice. Perhaps by tweaking max pooling windows? \n\nI assume this topic has been explored, but I’m having difficulty finding a treatment of it (or at least one that is intelligible to a beginner). \n\nIn practice, what strategies might be used to maximize sensitivity to vertical pattern positioning and minimize sensitivity to horizontal pattern positioning for the benefit of bird sound identification?",
    "1294078": "I have not found an answer to my question. However, I did come across this article, which points out several additional differences between spectrograms and visual images, and ultimately implies that CNNs may not be the best approach for audio analysis in the long run: \n\nhttps://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\n\nThoughts? Agreement? Disagreement? Practical advice/strategies for workarounds regarding those issues in the short term?",
    "1295241": "> I assume this topic has been explored, but I’m having difficulty finding a treatment of it (or at least one that is intelligible to a beginner). \n\nI gave you a way to find some : https://www.kaggle.com/c/birdclef-2021/discussion/234464"
  },
  "source": "meta"
}