{
  "id": 230733,
  "title": "Time localization / bird or no bird?",
  "url": "/competitions/birdclef-2021/discussion/230733",
  "author_name": "",
  "post_date": "2021-04-05T12:36:45.782149200Z",
  "votes": 25,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Within the short training samples, I'm interested in if/how people identify which parts contain the bird sound vs background noise.  Some ideas…</p>\n<ol>\n<li>It's likely that most samples have been clipped from a longer recording, so it's likely that the target bird can be found within the first 5s and the last 5s.  This is the very basic heuristic that I'm using at the moment and it's certainly right pretty often.</li>\n<li>Use a heuristic over the signal itself.  I've experimented with looking for an increase in <code>sum(fft_magnitudes)</code>.  Often this spots the birds, but frequently it also picks up background noise e.g. wind noise.  Armed with this, I'm planning to experiment with looking for increases in a target frequency range that aren't accompanied by increases across the spectrum.  I'd probably define the target range by hand, perhaps on a per-species basis.</li>\n<li>Use something very simple like #1 to train a model and then use the trained model to go back over all the training data to identify which parts do and don't actually contain the target bird.</li>\n</ol>\n<p>And I guess it's worth noting that I do this to (a) more accurately extract 5 second training samples from the short audio clips and (b) to get an additional source of <code>nobird</code> samples other than the few training soundscapes.  But maybe that isn't necessary?  Perhaps you find that you can just train on all 5 second windows for (a) and maybe you have some other solution for (b)?</p>\n<p>What techniques do you find to be effective for localizing where the bird is singing within the training samples?  Or do you somehow side-step this sub-problem altogether?</p>",
  "messages": [
    {
      "id": "1263456",
      "postDate": "04/05/2021 12:36:45",
      "content": "<p>Within the short training samples, I'm interested in if/how people identify which parts contain the bird sound vs background noise.  Some ideas…</p>\n<ol>\n<li>It's likely that most samples have been clipped from a longer recording, so it's likely that the target bird can be found within the first 5s and the last 5s.  This is the very basic heuristic that I'm using at the moment and it's certainly right pretty often.</li>\n<li>Use a heuristic over the signal itself.  I've experimented with looking for an increase in <code>sum(fft_magnitudes)</code>.  Often this spots the birds, but frequently it also picks up background noise e.g. wind noise.  Armed with this, I'm planning to experiment with looking for increases in a target frequency range that aren't accompanied by increases across the spectrum.  I'd probably define the target range by hand, perhaps on a per-species basis.</li>\n<li>Use something very simple like #1 to train a model and then use the trained model to go back over all the training data to identify which parts do and don't actually contain the target bird.</li>\n</ol>\n<p>And I guess it's worth noting that I do this to (a) more accurately extract 5 second training samples from the short audio clips and (b) to get an additional source of <code>nobird</code> samples other than the few training soundscapes.  But maybe that isn't necessary?  Perhaps you find that you can just train on all 5 second windows for (a) and maybe you have some other solution for (b)?</p>\n<p>What techniques do you find to be effective for localizing where the bird is singing within the training samples?  Or do you somehow side-step this sub-problem altogether?</p>",
      "rawMarkdown": "Within the short training samples, I'm interested in if/how people identify which parts contain the bird sound vs background noise.  Some ideas...\n\n1. It's likely that most samples have been clipped from a longer recording, so it's likely that the target bird can be found within the first 5s and the last 5s.  This is the very basic heuristic that I'm using at the moment and it's certainly right pretty often.\n2.  Use a heuristic over the signal itself.  I've experimented with looking for an increase in `sum(fft_magnitudes)`.  Often this spots the birds, but frequently it also picks up background noise e.g. wind noise.  Armed with this, I'm planning to experiment with looking for increases in a target frequency range that aren't accompanied by increases across the spectrum.  I'd probably define the target range by hand, perhaps on a per-species basis.\n3. Use something very simple like #1 to train a model and then use the trained model to go back over all the training data to identify which parts do and don't actually contain the target bird.\n\nAnd I guess it's worth noting that I do this to (a) more accurately extract 5 second training samples from the short audio clips and (b) to get an additional source of `nobird` samples other than the few training soundscapes.  But maybe that isn't necessary?  Perhaps you find that you can just train on all 5 second windows for (a) and maybe you have some other solution for (b)?\n\nWhat techniques do you find to be effective for localizing where the bird is singing within the training samples?  Or do you somehow side-step this sub-problem altogether?",
      "votes": null
    },
    {
      "id": "1265141",
      "postDate": "04/06/2021 16:41:03",
      "content": "<p>The first five seconds is, indeed, surprisingly effective.</p>\n<p>To get interesting audio from the middle of the file, you might be able to adapt methods from voice activity detection:<br>\n<a href=\"https://en.wikipedia.org/wiki/Voice_activity_detection\" target=\"_blank\">https://en.wikipedia.org/wiki/Voice_activity_detection</a></p>\n<p>I've had some good experience using <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.find_peaks_cwt.html\" target=\"_blank\">scipy.signal.find_peaks_cwt</a> on the summed spectral energy. (Possibly overkill, but here's a <a href=\"https://www.aclweb.org/anthology/O08-1016.pdf\" target=\"_blank\">paper on using wavelets for voice activity detection</a>.)</p>",
      "rawMarkdown": "The first five seconds is, indeed, surprisingly effective.\n\nTo get interesting audio from the middle of the file, you might be able to adapt methods from voice activity detection:\n[https://en.wikipedia.org/wiki/Voice_activity_detection](https://en.wikipedia.org/wiki/Voice_activity_detection)\n\nI've had some good experience using [scipy.signal.find_peaks_cwt](https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.find_peaks_cwt.html) on the summed spectral energy. (Possibly overkill, but here's a [paper on using wavelets for voice activity detection](https://www.aclweb.org/anthology/O08-1016.pdf).)",
      "votes": null
    },
    {
      "id": "1265442",
      "postDate": "04/06/2021 22:19:47",
      "content": "<blockquote>\n  <p>it's likely that the target bird can be found within the first 5s and the last 5s</p>\n</blockquote>\n<p>I made that assumption in last competition.</p>",
      "rawMarkdown": "> it's likely that the target bird can be found within the first 5s and the last 5s\n\nI made that assumption in last competition.",
      "votes": null
    },
    {
      "id": "1267818",
      "postDate": "04/08/2021 20:23:33",
      "content": "<p>I was thinking that one could have a nocall classifier and then just break recordings into chunks which you could use to determine if there's a bird or not in that section of recording. A more interesting model could take in the initial recording length and a representation of the recording in order to predict where the start and stop of the bird sound is; using a mixture model, you could predict multiple starting and stopping points too. In the end you would have a smaller but more relevant/refined dataset with which you could train… Please tell me your thoughts.  </p>",
      "rawMarkdown": "I was thinking that one could have a nocall classifier and then just break recordings into chunks which you could use to determine if there's a bird or not in that section of recording. A more interesting model could take in the initial recording length and a representation of the recording in order to predict where the start and stop of the bird sound is; using a mixture model, you could predict multiple starting and stopping points too. In the end you would have a smaller but more relevant/refined dataset with which you could train... Please tell me your thoughts.",
      "votes": null
    },
    {
      "id": "1274686",
      "postDate": "04/15/2021 14:15:41",
      "content": "<pre><code>It's likely that most samples have been clipped from a longer recording, ...\n</code></pre>\n<p>What's the reasoning behind this statement? </p>\n<p>Thanks for sharing the post though. :)</p>",
      "rawMarkdown": "```\nIt's likely that most samples have been clipped from a longer recording, ...\n```\n\nWhat's the reasoning behind this statement? \n\nThanks for sharing the post though. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1265141,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "04/06/2021 16:41:03",
      "content": "<p>The first five seconds is, indeed, surprisingly effective.</p>\n<p>To get interesting audio from the middle of the file, you might be able to adapt methods from voice activity detection:<br>\n<a href=\"https://en.wikipedia.org/wiki/Voice_activity_detection\" target=\"_blank\">https://en.wikipedia.org/wiki/Voice_activity_detection</a></p>\n<p>I've had some good experience using <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.find_peaks_cwt.html\" target=\"_blank\">scipy.signal.find_peaks_cwt</a> on the summed spectral energy. (Possibly overkill, but here's a <a href=\"https://www.aclweb.org/anthology/O08-1016.pdf\" target=\"_blank\">paper on using wavelets for voice activity detection</a>.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1265442,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/06/2021 22:19:47",
      "content": "<blockquote>\n  <p>it's likely that the target bird can be found within the first 5s and the last 5s</p>\n</blockquote>\n<p>I made that assumption in last competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1267818,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "04/08/2021 20:23:33",
      "content": "<p>I was thinking that one could have a nocall classifier and then just break recordings into chunks which you could use to determine if there's a bird or not in that section of recording. A more interesting model could take in the initial recording length and a representation of the recording in order to predict where the start and stop of the bird sound is; using a mixture model, you could predict multiple starting and stopping points too. In the end you would have a smaller but more relevant/refined dataset with which you could train… Please tell me your thoughts.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1274686,
      "author_name": "ayuraj",
      "author_url": "",
      "post_date": "04/15/2021 14:15:41",
      "content": "<pre><code>It's likely that most samples have been clipped from a longer recording, ...\n</code></pre>\n<p>What's the reasoning behind this statement? </p>\n<p>Thanks for sharing the post though. :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1263456": "Within the short training samples, I'm interested in if/how people identify which parts contain the bird sound vs background noise.  Some ideas...\n\n1. It's likely that most samples have been clipped from a longer recording, so it's likely that the target bird can be found within the first 5s and the last 5s.  This is the very basic heuristic that I'm using at the moment and it's certainly right pretty often.\n2.  Use a heuristic over the signal itself.  I've experimented with looking for an increase in `sum(fft_magnitudes)`.  Often this spots the birds, but frequently it also picks up background noise e.g. wind noise.  Armed with this, I'm planning to experiment with looking for increases in a target frequency range that aren't accompanied by increases across the spectrum.  I'd probably define the target range by hand, perhaps on a per-species basis.\n3. Use something very simple like #1 to train a model and then use the trained model to go back over all the training data to identify which parts do and don't actually contain the target bird.\n\nAnd I guess it's worth noting that I do this to (a) more accurately extract 5 second training samples from the short audio clips and (b) to get an additional source of `nobird` samples other than the few training soundscapes.  But maybe that isn't necessary?  Perhaps you find that you can just train on all 5 second windows for (a) and maybe you have some other solution for (b)?\n\nWhat techniques do you find to be effective for localizing where the bird is singing within the training samples?  Or do you somehow side-step this sub-problem altogether?",
    "1265141": "The first five seconds is, indeed, surprisingly effective.\n\nTo get interesting audio from the middle of the file, you might be able to adapt methods from voice activity detection:\n[https://en.wikipedia.org/wiki/Voice_activity_detection](https://en.wikipedia.org/wiki/Voice_activity_detection)\n\nI've had some good experience using [scipy.signal.find_peaks_cwt](https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.find_peaks_cwt.html) on the summed spectral energy. (Possibly overkill, but here's a [paper on using wavelets for voice activity detection](https://www.aclweb.org/anthology/O08-1016.pdf).)",
    "1265442": "> it's likely that the target bird can be found within the first 5s and the last 5s\n\nI made that assumption in last competition.",
    "1267818": "I was thinking that one could have a nocall classifier and then just break recordings into chunks which you could use to determine if there's a bird or not in that section of recording. A more interesting model could take in the initial recording length and a representation of the recording in order to predict where the start and stop of the bird sound is; using a mixture model, you could predict multiple starting and stopping points too. In the end you would have a smaller but more relevant/refined dataset with which you could train... Please tell me your thoughts.",
    "1274686": "```\nIt's likely that most samples have been clipped from a longer recording, ...\n```\n\nWhat's the reasoning behind this statement? \n\nThanks for sharing the post though. :)"
  },
  "source": "meta"
}