{
  "id": 579393,
  "title": "Why do the audio files in this directory seem to contain only human speech?",
  "url": "/competitions/birdclef-2025/discussion/579393",
  "author_name": "",
  "post_date": "2025-05-17T05:33:21.272006100Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>train_audio =&gt; 1139490</p>\n<p>CSA36385.ogg<br>\nCSA36389.ogg</p>",
  "messages": [
    {
      "id": "3203659",
      "postDate": "05/17/2025 05:33:21",
      "content": "<p>train_audio =&gt; 1139490</p>\n<p>CSA36385.ogg<br>\nCSA36389.ogg</p>",
      "rawMarkdown": "train_audio => 1139490\n\nCSA36385.ogg\nCSA36389.ogg",
      "votes": null
    },
    {
      "id": "3203740",
      "postDate": "05/17/2025 08:51:39",
      "content": "<p>The target sound appears within the first 8 seconds.</p>",
      "rawMarkdown": "The target sound appears within the first 8 seconds.",
      "votes": null
    },
    {
      "id": "3203874",
      "postDate": "05/17/2025 12:25:30",
      "content": "<p>is there any page or source that mentions the target sound appearing within the first 8 seconds? If there isn’t such a rule, would it be correct to assume that the researcher has to manually identify and filter out human voices?</p>",
      "rawMarkdown": "is there any page or source that mentions the target sound appearing within the first 8 seconds? If there isn’t such a rule, would it be correct to assume that the researcher has to manually identify and filter out human voices?",
      "votes": null
    },
    {
      "id": "3205069",
      "postDate": "05/19/2025 09:10:37",
      "content": "<p>I listened to the samples and there is at least one CSA file or INat file where there is human voice in the first seconds of the sample. <br>\nBut the files you mentioned indeed contain the target sound in the first seconds. Maybe it is just quiet, so you didn't catch it, but it is obvious from spectrogram image for example.</p>\n<p>IMHO SileroVAD is the best approach, and then maybe augment with human voice.</p>",
      "rawMarkdown": "I listened to the samples and there is at least one CSA file or INat file where there is human voice in the first seconds of the sample. \nBut the files you mentioned indeed contain the target sound in the first seconds. Maybe it is just quiet, so you didn't catch it, but it is obvious from spectrogram image for example.\n\nIMHO SileroVAD is the best approach, and then maybe augment with human voice.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3203740,
      "author_name": "myso1987",
      "author_url": "",
      "post_date": "05/17/2025 08:51:39",
      "content": "<p>The target sound appears within the first 8 seconds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3203874,
          "author_name": "newbyy",
          "author_url": "",
          "post_date": "05/17/2025 12:25:30",
          "content": "<p>is there any page or source that mentions the target sound appearing within the first 8 seconds? If there isn’t such a rule, would it be correct to assume that the researcher has to manually identify and filter out human voices?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3205069,
              "author_name": "alexandergremyakov",
              "author_url": "",
              "post_date": "05/19/2025 09:10:37",
              "content": "<p>I listened to the samples and there is at least one CSA file or INat file where there is human voice in the first seconds of the sample. <br>\nBut the files you mentioned indeed contain the target sound in the first seconds. Maybe it is just quiet, so you didn't catch it, but it is obvious from spectrogram image for example.</p>\n<p>IMHO SileroVAD is the best approach, and then maybe augment with human voice.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3203659": "train_audio => 1139490\n\nCSA36385.ogg\nCSA36389.ogg",
    "3203740": "The target sound appears within the first 8 seconds.",
    "3203874": "is there any page or source that mentions the target sound appearing within the first 8 seconds? If there isn’t such a rule, would it be correct to assume that the researcher has to manually identify and filter out human voices?",
    "3205069": "I listened to the samples and there is at least one CSA file or INat file where there is human voice in the first seconds of the sample. \nBut the files you mentioned indeed contain the target sound in the first seconds. Maybe it is just quiet, so you didn't catch it, but it is obvious from spectrogram image for example.\n\nIMHO SileroVAD is the best approach, and then maybe augment with human voice."
  },
  "source": "meta"
}