{
  "id": 315668,
  "title": "Keeping only relevant parts of the signals",
  "url": "/competitions/birdclef-2022/discussion/315668",
  "author_name": "",
  "post_date": "2022-03-29T09:46:24.266496400Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I am very new to sound projects, and I what is the best strategy to start this project.</p>\n<p>Some of the species are rare, and they are also not very talkative… And so in some of the audio samples, most of it is just background noise. <br>\nSo I was wondering if the best strategy is to feed a model with the full 5s sample, or try to isolate sounds of the audio prior to feeding them to a model.</p>\n<p><strong>Let's take the example of puaioh</strong></p>\n<p>This is one of the full audio:<br>\n<img src=\"https://i.ibb.co/mJw6G5S/ts-puaioh.png\" alt=\"\"></p>\n<p>As well as the associated constant Q transform<br>\n<img src=\"https://i.ibb.co/JmpVjQL/cqt.png\" alt=\"\"></p>\n<p>Given the lack of labels, instead of just using a sliding window to break the signal down, wouldn't it be more efficient to have a first stage specifically designed to isolate build crops around high-intensity sounds ?</p>\n<p><img src=\"https://i.ibb.co/mRCfM6j/ts-2.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/vJ3s5GZ/cqt-2.png\" alt=\"\"></p>\n<p>My intuition is that by filtering the signals based on the intensity, we should be able to remove large portions of noise, helping the models to generalize better. Is it actually the case?</p>",
  "messages": [
    {
      "id": "1738459",
      "postDate": "03/29/2022 09:46:24",
      "content": "<p>Hi all,</p>\n<p>I am very new to sound projects, and I what is the best strategy to start this project.</p>\n<p>Some of the species are rare, and they are also not very talkative… And so in some of the audio samples, most of it is just background noise. <br>\nSo I was wondering if the best strategy is to feed a model with the full 5s sample, or try to isolate sounds of the audio prior to feeding them to a model.</p>\n<p><strong>Let's take the example of puaioh</strong></p>\n<p>This is one of the full audio:<br>\n<img src=\"https://i.ibb.co/mJw6G5S/ts-puaioh.png\" alt=\"\"></p>\n<p>As well as the associated constant Q transform<br>\n<img src=\"https://i.ibb.co/JmpVjQL/cqt.png\" alt=\"\"></p>\n<p>Given the lack of labels, instead of just using a sliding window to break the signal down, wouldn't it be more efficient to have a first stage specifically designed to isolate build crops around high-intensity sounds ?</p>\n<p><img src=\"https://i.ibb.co/mRCfM6j/ts-2.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/vJ3s5GZ/cqt-2.png\" alt=\"\"></p>\n<p>My intuition is that by filtering the signals based on the intensity, we should be able to remove large portions of noise, helping the models to generalize better. Is it actually the case?</p>",
      "rawMarkdown": "Hi all,\n\nI am very new to sound projects, and I what is the best strategy to start this project.\n\nSome of the species are rare, and they are also not very talkative... And so in some of the audio samples, most of it is just background noise. \nSo I was wondering if the best strategy is to feed a model with the full 5s sample, or try to isolate sounds of the audio prior to feeding them to a model.\n\n**Let's take the example of puaioh**\n\nThis is one of the full audio:\n![](https://i.ibb.co/mJw6G5S/ts-puaioh.png)\n\nAs well as the associated constant Q transform\n![](https://i.ibb.co/JmpVjQL/cqt.png)\n\nGiven the lack of labels, instead of just using a sliding window to break the signal down, wouldn't it be more efficient to have a first stage specifically designed to isolate build crops around high-intensity sounds ?\n\n![](https://i.ibb.co/mRCfM6j/ts-2.png)\n![](https://i.ibb.co/vJ3s5GZ/cqt-2.png)\n\nMy intuition is that by filtering the signals based on the intensity, we should be able to remove large portions of noise, helping the models to generalize better. Is it actually the case?",
      "votes": null
    },
    {
      "id": "1738539",
      "postDate": "03/29/2022 11:06:50",
      "content": "<p>In previous competitions, a pre-trained bird chirping non-bird chirp classifier (freefield1010) was used for this in the preprocessing stage.</p>\n<p>Also, Google’s latest study deals with this, called MixIt.<br>\n<a href=\"https://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html\" target=\"_blank\">https://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html</a></p>\n<p>Intensity-based recognition can help with some of the raw audio file, as well as applying a low-pass filter, but not much.</p>",
      "rawMarkdown": "In previous competitions, a pre-trained bird chirping non-bird chirp classifier (freefield1010) was used for this in the preprocessing stage.\n\nAlso, Google’s latest study deals with this, called MixIt.\nhttps://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html\n\nIntensity-based recognition can help with some of the raw audio file, as well as applying a low-pass filter, but not much.",
      "votes": null
    },
    {
      "id": "1739531",
      "postDate": "03/30/2022 05:59:14",
      "content": "<p>I am also new to this field, but it seems to me that this method has been used for a long time[1].<br>\nIt is also used in papers published by the hosts of this competition[2].</p>\n<p>[1] <a href=\"https://watermark.silverchair.com/btl355.pdf\" target=\"_blank\">https://watermark.silverchair.com/btl355.pdf</a><br>\n[2] <a href=\"https://arxiv.org/abs/2110.03209\" target=\"_blank\">https://arxiv.org/abs/2110.03209</a></p>",
      "rawMarkdown": "I am also new to this field, but it seems to me that this method has been used for a long time[1].\nIt is also used in papers published by the hosts of this competition[2].\n\n[1] https://watermark.silverchair.com/btl355.pdf\n[2] https://arxiv.org/abs/2110.03209",
      "votes": null
    },
    {
      "id": "1739586",
      "postDate": "03/30/2022 07:06:42",
      "content": "<p>I wrote for the last two that it doesn’t improve much.  The first two improve a lot.  I hope what I wrote is not misunderstood.</p>",
      "rawMarkdown": "I wrote for the last two that it doesn’t improve much.  The first two improve a lot.  I hope what I wrote is not misunderstood.",
      "votes": null
    },
    {
      "id": "1739609",
      "postDate": "03/30/2022 07:30:05",
      "content": "<p>I see. Now I understand why most of the solutions in BirdCLEF 2020 adopt 2-stage approach. (I didn’t clearly understand why 1st stage is needed.)</p>\n<p>Thank you for pointing this out.</p>",
      "rawMarkdown": "I see. Now I understand why most of the solutions in BirdCLEF 2020 adopt 2-stage approach. (I didn’t clearly understand why 1st stage is needed.)\n\nThank you for pointing this out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1738539,
      "author_name": "ambrusattila",
      "author_url": "",
      "post_date": "03/29/2022 11:06:50",
      "content": "<p>In previous competitions, a pre-trained bird chirping non-bird chirp classifier (freefield1010) was used for this in the preprocessing stage.</p>\n<p>Also, Google’s latest study deals with this, called MixIt.<br>\n<a href=\"https://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html\" target=\"_blank\">https://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html</a></p>\n<p>Intensity-based recognition can help with some of the raw audio file, as well as applying a low-pass filter, but not much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1739531,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "03/30/2022 05:59:14",
      "content": "<p>I am also new to this field, but it seems to me that this method has been used for a long time[1].<br>\nIt is also used in papers published by the hosts of this competition[2].</p>\n<p>[1] <a href=\"https://watermark.silverchair.com/btl355.pdf\" target=\"_blank\">https://watermark.silverchair.com/btl355.pdf</a><br>\n[2] <a href=\"https://arxiv.org/abs/2110.03209\" target=\"_blank\">https://arxiv.org/abs/2110.03209</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1739586,
          "author_name": "ambrusattila",
          "author_url": "",
          "post_date": "03/30/2022 07:06:42",
          "content": "<p>I wrote for the last two that it doesn’t improve much.  The first two improve a lot.  I hope what I wrote is not misunderstood.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739609,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "03/30/2022 07:30:05",
          "content": "<p>I see. Now I understand why most of the solutions in BirdCLEF 2020 adopt 2-stage approach. (I didn’t clearly understand why 1st stage is needed.)</p>\n<p>Thank you for pointing this out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1738459": "Hi all,\n\nI am very new to sound projects, and I what is the best strategy to start this project.\n\nSome of the species are rare, and they are also not very talkative... And so in some of the audio samples, most of it is just background noise. \nSo I was wondering if the best strategy is to feed a model with the full 5s sample, or try to isolate sounds of the audio prior to feeding them to a model.\n\n**Let's take the example of puaioh**\n\nThis is one of the full audio:\n![](https://i.ibb.co/mJw6G5S/ts-puaioh.png)\n\nAs well as the associated constant Q transform\n![](https://i.ibb.co/JmpVjQL/cqt.png)\n\nGiven the lack of labels, instead of just using a sliding window to break the signal down, wouldn't it be more efficient to have a first stage specifically designed to isolate build crops around high-intensity sounds ?\n\n![](https://i.ibb.co/mRCfM6j/ts-2.png)\n![](https://i.ibb.co/vJ3s5GZ/cqt-2.png)\n\nMy intuition is that by filtering the signals based on the intensity, we should be able to remove large portions of noise, helping the models to generalize better. Is it actually the case?",
    "1738539": "In previous competitions, a pre-trained bird chirping non-bird chirp classifier (freefield1010) was used for this in the preprocessing stage.\n\nAlso, Google’s latest study deals with this, called MixIt.\nhttps://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html\n\nIntensity-based recognition can help with some of the raw audio file, as well as applying a low-pass filter, but not much.",
    "1739531": "I am also new to this field, but it seems to me that this method has been used for a long time[1].\nIt is also used in papers published by the hosts of this competition[2].\n\n[1] https://watermark.silverchair.com/btl355.pdf\n[2] https://arxiv.org/abs/2110.03209",
    "1739586": "I wrote for the last two that it doesn’t improve much.  The first two improve a lot.  I hope what I wrote is not misunderstood.",
    "1739609": "I see. Now I understand why most of the solutions in BirdCLEF 2020 adopt 2-stage approach. (I didn’t clearly understand why 1st stage is needed.)\n\nThank you for pointing this out."
  },
  "source": "meta"
}