{
  "id": 568886,
  "title": "Human voice in the recordings",
  "url": "/competitions/birdclef-2025/discussion/568886",
  "author_name": "Konstantin Dmitriev",
  "post_date": "2025-03-18T14:30:18.066000",
  "votes": 108,
  "comment_count": 37,
  "views": 0,
  "content": "<p>It was noticed by many participants, that the provided data often contains human voice. This is particularly  related to CSA dataset. I decided to explore this factor and maybe get some insights from the data.</p>\n<p>The details are presented <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\"><strong>in the notebook</strong></a>.</p>\n<h1>1. Exploration</h1>\n<h2>Is human voice presence common for all CSA recordings?</h2>\n<p>At first, I looked at the recordings by Fabio A. Sarria-S, as <a href=\"https://www.kaggle.com/martinapreusse\" target=\"_blank\">@martinapreusse</a> said, they contain voice. I have plotted some kind of dependence of the sound level on time, and a clear pattern appeared. The actual insect is separated from human description by short pauses.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fd89b7b930333d100c26feca344747b03%2F1.png?generation=1742306694933745&amp;alt=media\" alt=\"\"></p>\n<p>After a short analysis, it was clear, that <strong>almost all CSA recordings contain human voice</strong>.</p>\n<p>This could be easy to exclude human voice, but the recordings by other authors have different patterns. For example, Alexandra Buitrago-Cardona says something at the beginning of the recordings (and, sometimes, at the end or in the middle). </p>\n<h2>Do other collections contain human voice?</h2>\n<p>A short answer: <strong>yes</strong>!</p>\n<h2>Do train soundscapes contain human voice?</h2>\n<p>And again, the answer is <strong>yes</strong>!</p>\n<h1>2. Solution</h1>\n<p>I found a wonderful library <code>silero-vad</code> that can be used to detect human voice in the recordings. The use of the library is very simple:</p>\n<pre><code>model, (get_speech_timestamps, _, read_audio, _, _) = torch.hub.load(=, =)\nspeech_timestamps = get_speech_timestamps(torch.Tensor(wav), model)\n</code></pre>\n<p>The timestamps it provides are quite accurate. I have compared them with manual audio analysis. Also they correlate with the 'split-by-pause' method (if it is applicable). Below are some examples.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2F1ce810158bb88b54fb76dedc6e871843%2F2.png?generation=1742307610054378&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fa3f8814f8f3c270c4ae94ecea40a38af%2F3.png?generation=1742307619626563&amp;alt=media\" alt=\"\"></p>\n<h1>3. Results</h1>\n<p>I have run my  <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\"><strong>notebook</strong></a> on both train soundscapes and train audio data. The resulting <a href=\"https://www.kaggle.com/datasets/kdmitrie/bc25-separation-voice-from-data-by-silero-vad\" target=\"_blank\"><strong>dataset</strong></a> contains several files:</p>\n<ul>\n<li>*_voice_summary.txt - a list of files where human voice was found;</li>\n<li>*_voice_data.pkl - a dictionary saved with <code>pickle</code> library. Its keys are the names of files with human voice; each value is a list with exact locations of human voice.</li>\n</ul>",
  "messages": [
    {
      "id": 3153227,
      "postDate": "2025-03-18T14:30:18.067Z",
      "content": "<p>It was noticed by many participants, that the provided data often contains human voice. This is particularly  related to CSA dataset. I decided to explore this factor and maybe get some insights from the data.</p>\n<p>The details are presented <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\"><strong>in the notebook</strong></a>.</p>\n<h1>1. Exploration</h1>\n<h2>Is human voice presence common for all CSA recordings?</h2>\n<p>At first, I looked at the recordings by Fabio A. Sarria-S, as <a href=\"https://www.kaggle.com/martinapreusse\" target=\"_blank\">@martinapreusse</a> said, they contain voice. I have plotted some kind of dependence of the sound level on time, and a clear pattern appeared. The actual insect is separated from human description by short pauses.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fd89b7b930333d100c26feca344747b03%2F1.png?generation=1742306694933745&amp;alt=media\" alt=\"\"></p>\n<p>After a short analysis, it was clear, that <strong>almost all CSA recordings contain human voice</strong>.</p>\n<p>This could be easy to exclude human voice, but the recordings by other authors have different patterns. For example, Alexandra Buitrago-Cardona says something at the beginning of the recordings (and, sometimes, at the end or in the middle). </p>\n<h2>Do other collections contain human voice?</h2>\n<p>A short answer: <strong>yes</strong>!</p>\n<h2>Do train soundscapes contain human voice?</h2>\n<p>And again, the answer is <strong>yes</strong>!</p>\n<h1>2. Solution</h1>\n<p>I found a wonderful library <code>silero-vad</code> that can be used to detect human voice in the recordings. The use of the library is very simple:</p>\n<pre><code>model, (get_speech_timestamps, _, read_audio, _, _) = torch.hub.load(=, =)\nspeech_timestamps = get_speech_timestamps(torch.Tensor(wav), model)\n</code></pre>\n<p>The timestamps it provides are quite accurate. I have compared them with manual audio analysis. Also they correlate with the 'split-by-pause' method (if it is applicable). Below are some examples.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2F1ce810158bb88b54fb76dedc6e871843%2F2.png?generation=1742307610054378&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fa3f8814f8f3c270c4ae94ecea40a38af%2F3.png?generation=1742307619626563&amp;alt=media\" alt=\"\"></p>\n<h1>3. Results</h1>\n<p>I have run my  <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\"><strong>notebook</strong></a> on both train soundscapes and train audio data. The resulting <a href=\"https://www.kaggle.com/datasets/kdmitrie/bc25-separation-voice-from-data-by-silero-vad\" target=\"_blank\"><strong>dataset</strong></a> contains several files:</p>\n<ul>\n<li>*_voice_summary.txt - a list of files where human voice was found;</li>\n<li>*_voice_data.pkl - a dictionary saved with <code>pickle</code> library. Its keys are the names of files with human voice; each value is a list with exact locations of human voice.</li>\n</ul>",
      "rawMarkdown": "It was noticed by many participants, that the provided data often contains human voice. This is particularly  related to CSA dataset. I decided to explore this factor and maybe get some insights from the data.\n\nThe details are presented [**in the notebook**](https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data).\n# 1. Exploration\n## Is human voice presence common for all CSA recordings?\nAt first, I looked at the recordings by Fabio A. Sarria-S, as @martinapreusse said, they contain voice. I have plotted some kind of dependence of the sound level on time, and a clear pattern appeared. The actual insect is separated from human description by short pauses.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fd89b7b930333d100c26feca344747b03%2F1.png?generation=1742306694933745&alt=media)\n\nAfter a short analysis, it was clear, that **almost all CSA recordings contain human voice**.\n\nThis could be easy to exclude human voice, but the recordings by other authors have different patterns. For example, Alexandra Buitrago-Cardona says something at the beginning of the recordings (and, sometimes, at the end or in the middle). \n\n## Do other collections contain human voice?\nA short answer: **yes**!\n\n## Do train soundscapes contain human voice?\nAnd again, the answer is **yes**!\n\n# 2. Solution\nI found a wonderful library `silero-vad` that can be used to detect human voice in the recordings. The use of the library is very simple:\n```\nmodel, (get_speech_timestamps, _, read_audio, _, _) = torch.hub.load(repo_or_dir='snakers4/silero-vad', model='silero_vad')\nspeech_timestamps = get_speech_timestamps(torch.Tensor(wav), model)\n```\nThe timestamps it provides are quite accurate. I have compared them with manual audio analysis. Also they correlate with the 'split-by-pause' method (if it is applicable). Below are some examples.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2F1ce810158bb88b54fb76dedc6e871843%2F2.png?generation=1742307610054378&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fa3f8814f8f3c270c4ae94ecea40a38af%2F3.png?generation=1742307619626563&alt=media)\n\n# 3. Results\nI have run my  [**notebook**](https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data) on both train soundscapes and train audio data. The resulting [**dataset**](https://www.kaggle.com/datasets/kdmitrie/bc25-separation-voice-from-data-by-silero-vad) contains several files:\n- *_voice_summary.txt - a list of files where human voice was found;\n- *_voice_data.pkl - a dictionary saved with `pickle` library. Its keys are the names of files with human voice; each value is a list with exact locations of human voice.\n\n",
      "votes": 108
    },
    {
      "id": 3209121,
      "postDate": "2025-05-25T09:10:53.667Z",
      "content": "<p>It provide a clear understanding, very helpful! Thx!</p>",
      "rawMarkdown": "It provide a clear understanding, very helpful! Thx!",
      "votes": 1
    },
    {
      "id": 3154723,
      "postDate": "2025-03-20T10:20:06.513Z",
      "content": "<p>Hello, thank you so much for bringing up this idea.</p>\n<p>For all who reads this, I recommend you try <code>threshold</code> parameter. For example,<br>\n<code>speech_timestamps = get_speech_timestamps(torch.Tensor(segment), model, threshold=0.1)</code></p>\n<p><a href=\"https://github.com/snakers4/silero-vad/wiki/Quality-Metrics#threshold\" target=\"_blank\">https://github.com/snakers4/silero-vad/wiki/Quality-Metrics#threshold</a></p>\n<p>By default, the threshold is set to 0.5, but the model didn't work for <code>/kaggle/input/birdclef-2025/train_audio/1139490/CSA36385.ogg</code>. Once I've tweaked it to 0.4, it started working as expected.<br>\nHope this helps someone.</p>",
      "rawMarkdown": "Hello, thank you so much for bringing up this idea.\n\nFor all who reads this, I recommend you try `threshold` parameter. For example,\n`speech_timestamps = get_speech_timestamps(torch.Tensor(segment), model, threshold=0.1)`\n\nhttps://github.com/snakers4/silero-vad/wiki/Quality-Metrics#threshold\n\nBy default, the threshold is set to 0.5, but the model didn't work for `/kaggle/input/birdclef-2025/train_audio/1139490/CSA36385.ogg`. Once I've tweaked it to 0.4, it started working as expected.\nHope this helps someone.",
      "votes": 3,
      "replies": [
        {
          "id": 3154821,
          "postDate": "2025-03-20T12:35:32.647Z",
          "content": "<p><a href=\"https://www.kaggle.com/tomok1\" target=\"_blank\">@tomok1</a>, thank you very much for your suggestion! </p>\n<p>I have made the corrections in the code, introducing the threshold of 0.4. The new version will be ready in about 8-9 hours.</p>",
          "rawMarkdown": "@tomok1, thank you very much for your suggestion! \n\nI have made the corrections in the code, introducing the threshold of 0.4. The new version will be ready in about 8-9 hours.",
          "votes": 1
        },
        {
          "id": 3200997,
          "postDate": "2025-05-13T10:26:39.040Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3153738,
      "postDate": "2025-03-19T05:46:40.737Z",
      "content": "<p>Can we  find these voices remove  of human from  dataset  make the data clean for only bird voices.</p>",
      "rawMarkdown": "Can we  find these voices remove  of human from  dataset  make the data clean for only bird voices.",
      "votes": 3,
      "replies": [
        {
          "id": 3154168,
          "postDate": "2025-03-19T15:46:28.997Z",
          "content": "<p>This is a tricky question.  There are methods to identify if human voice in a recording segment - a few have been shared in notebooks or discussion.  So you could clean up the data - the issue, what happens if humans are talking in the test?   I am hoping to teach my model that humans are the 207th class in the data.</p>",
          "rawMarkdown": "This is a tricky question.  There are methods to identify if human voice in a recording segment - a few have been shared in notebooks or discussion.  So you could clean up the data - the issue, what happens if humans are talking in the test?   I am hoping to teach my model that humans are the 207th class in the data.",
          "votes": 8,
          "replies": [
            {
              "id": 3154220,
              "postDate": "2025-03-19T16:56:33.923Z",
              "content": "<p>My opinion is, humans do talk in the test, but not too much.</p>\n<p>In the <code>train_sounscapes</code>, there are 143 recordings that were confirmed to contain human voice. It is 1.5% of all soundscapes. However, the human voice inside <code>train_audio</code> is more dangerous, as the model can easily be misled by the clear and loud sound of speech instead of learning the weak insect sounds.</p>\n<p>I think, it's a good idea to remove human voice form <code>train_audio</code>, but leave it inside <code>train_soundscapes</code> to address the domain shift correctly.</p>",
              "rawMarkdown": "My opinion is, humans do talk in the test, but not too much.\n\nIn the `train_sounscapes`, there are 143 recordings that were confirmed to contain human voice. It is 1.5% of all soundscapes. However, the human voice inside `train_audio` is more dangerous, as the model can easily be misled by the clear and loud sound of speech instead of learning the weak insect sounds.\n\nI think, it's a good idea to remove human voice form `train_audio`, but leave it inside `train_soundscapes` to address the domain shift correctly.",
              "votes": 13
            },
            {
              "id": 3154446,
              "postDate": "2025-03-20T00:01:13.680Z",
              "content": "<p>Thanks.Could you please  tell me one more  thing i do not understand it we use train_audio for training data while prediction on we use test_audio.  What will be use of  train_soundscapes unlabeled data.</p>",
              "rawMarkdown": "Thanks.Could you please  tell me one more  thing i do not understand it we use train_audio for training data while prediction on we use test_audio.  What will be use of  train_soundscapes unlabeled data.",
              "votes": 1
            },
            {
              "id": 3154486,
              "postDate": "2025-03-20T02:30:28.773Z",
              "content": "<p>You can test that your model will process soundscapes by using train - I copy a couple over to the test folder on my local machine.</p>\n<p>You may also be able to label some of the audio and perhaps use it as additional training data.</p>",
              "rawMarkdown": "You can test that your model will process soundscapes by using train - I copy a couple over to the test folder on my local machine.\n\nYou may also be able to label some of the audio and perhaps use it as additional training data.",
              "votes": 1
            },
            {
              "id": 3154516,
              "postDate": "2025-03-20T03:04:02.007Z",
              "content": "<p>So they are not testing data but we can use them to see if our model is working fine to testing ready for submission process we can test with this direocty or labeled them use them for extra data for training am I get it right ?</p>",
              "rawMarkdown": "So they are not testing data but we can use them to see if our model is working fine to testing ready for submission process we can test with this direocty or labeled them use them for extra data for training am I get it right ?",
              "votes": 1
            },
            {
              "id": 3154524,
              "postDate": "2025-03-20T03:09:01.713Z",
              "content": "<p>As noted in the over view the train soundscapes are similar to the test.  If your model can do a good job of predicting than some of the unlabeled train soundscapes could make good training data to improve your model.  The training data we have been given is from all over the world, there is a decent chance that the train and test soundscapes have been recorded in a smaller number of places around the world.</p>\n<p>You need to have decent confidence that your model is doing a good job of labeling, but for sure my plan is to try to add as much of the train soundscapes as I can for further training data. </p>",
              "rawMarkdown": "As noted in the over view the train soundscapes are similar to the test.  If your model can do a good job of predicting than some of the unlabeled train soundscapes could make good training data to improve your model.  The training data we have been given is from all over the world, there is a decent chance that the train and test soundscapes have been recorded in a smaller number of places around the world.\n\nYou need to have decent confidence that your model is doing a good job of labeling, but for sure my plan is to try to add as much of the train soundscapes as I can for further training data. ",
              "votes": 2
            },
            {
              "id": 3154543,
              "postDate": "2025-03-20T03:38:47.900Z",
              "content": "<p>Thanks. I was able to submit my first submisison file. Thanks for you help. I really appreciate it.</p>",
              "rawMarkdown": "Thanks. I was able to submit my first submisison file. Thanks for you help. I really appreciate it.",
              "votes": 1
            },
            {
              "id": 3158620,
              "postDate": "2025-03-24T17:53:09.723Z",
              "content": "<p>I did a quick change were I removed any recording that had speech for training - bad idea as I did not end up with 206 classes :)</p>",
              "rawMarkdown": "I did a quick change were I removed any recording that had speech for training - bad idea as I did not end up with 206 classes :)",
              "votes": 1
            },
            {
              "id": 3162081,
              "postDate": "2025-03-28T19:23:05.867Z",
              "content": "<p>\"but rather a description of the heard species and recording environment.\"</p>\n<p>I hope that is not the case on the test environment as well!</p>\n<p>Otherwise this competition could end up picking the model that does the best human speech translation…</p>",
              "rawMarkdown": "\"but rather a description of the heard species and recording environment.\"\n\nI hope that is not the case on the test environment as well!\n\nOtherwise this competition could end up picking the model that does the best human speech translation...\n\n"
            }
          ]
        }
      ]
    },
    {
      "id": 3203446,
      "postDate": "2025-05-16T18:36:31.657Z",
      "content": "<p>Thanks to you, my LB improved! I would have never figured out the fix like this! </p>",
      "rawMarkdown": "Thanks to you, my LB improved! I would have never figured out the fix like this! ",
      "votes": 1
    },
    {
      "id": 3203384,
      "postDate": "2025-05-16T16:50:09.707Z",
      "content": "<p>Discarding frames that only contain human voices seems to really hurt my model’s performance. There are still some key points I haven’t figured out.</p>",
      "rawMarkdown": "Discarding frames that only contain human voices seems to really hurt my model’s performance. There are still some key points I haven’t figured out.",
      "votes": 1,
      "replies": [
        {
          "id": 3203511,
          "postDate": "2025-05-16T20:43:23.850Z",
          "content": "<p>What do your 5 second audio segments contain that are left over?</p>",
          "rawMarkdown": "What do your 5 second audio segments contain that are left over?"
        }
      ]
    },
    {
      "id": 3200773,
      "postDate": "2025-05-13T03:41:10.737Z",
      "content": "<p>Thank you for sharing the information. Thanks to this, I was able to significantly improve my LB score.</p>",
      "rawMarkdown": "Thank you for sharing the information. Thanks to this, I was able to significantly improve my LB score.",
      "replies": [
        {
          "id": 3200796,
          "postDate": "2025-05-13T04:32:22.637Z",
          "content": "<p>great! I removed the human voice and the lb dropped a lot. I will continue to explore</p>",
          "rawMarkdown": "great! I removed the human voice and the lb dropped a lot. I will continue to explore",
          "votes": 1
        }
      ]
    },
    {
      "id": 3171349,
      "postDate": "2025-04-05T15:37:48.313Z",
      "content": "<p>I noticed the problem but didn’t get as far as the OP in cracking the nut! 🥳 I was wondering however: Some of the CSA recordings, where there is a human voice, also have a synthetic voice reading out the label. I’m half convinced that that will create leakage somehow. Anybody had the same thought?</p>",
      "rawMarkdown": "I noticed the problem but didn’t get as far as the OP in cracking the nut! 🥳 I was wondering however: Some of the CSA recordings, where there is a human voice, also have a synthetic voice reading out the label. I’m half convinced that that will create leakage somehow. Anybody had the same thought?",
      "votes": 2,
      "replies": [
        {
          "id": 3172291,
          "postDate": "2025-04-06T17:02:03.357Z",
          "content": "<p>Aggreed. I posted a similar question but I did not get any replies from the organizers.</p>\n<p>If that is the case I am concerned if it is possible some of the top entries could be doing speech translation on the test set ( specially because you can access data way before/after the target region).</p>\n<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>  could you please clarify that any speech in the test set <em>does</em> not include the labels we are trying to identify? Thank you.</p>",
          "rawMarkdown": "Aggreed. I posted a similar question but I did not get any replies from the organizers.\n\nIf that is the case I am concerned if it is possible some of the top entries could be doing speech translation on the test set ( specially because you can access data way before/after the target region).\n\n\n\n@stefankahl  could you please clarify that any speech in the test set *does* not include the labels we are trying to identify? Thank you.\n",
          "votes": 1,
          "replies": [
            {
              "id": 3172311,
              "postDate": "2025-04-06T17:17:26.573Z",
              "content": "<p>I see what you mean. My understanding is that the test set would be soundscapes, and <em>naively</em> I’d assume that these are unlikely to include the <em>computer</em> voices reading out the labels clearly and distinctly. So I didn’t even thought about this possibility. Though I guess it is possible that a human voice in a test soundscape could mention one of the species recorded. </p>\n<p>My original question was meant like: If it isn’t advised to clean the train data of human voices altogether, because the test soundscapes are likely to contain some, would it still make sense to somehow detect and remove the generated voice?</p>",
              "rawMarkdown": "I see what you mean. My understanding is that the test set would be soundscapes, and *naively* I’d assume that these are unlikely to include the _computer_ voices reading out the labels clearly and distinctly. So I didn’t even thought about this possibility. Though I guess it is possible that a human voice in a test soundscape could mention one of the species recorded. \n\nMy original question was meant like: If it isn’t advised to clean the train data of human voices altogether, because the test soundscapes are likely to contain some, would it still make sense to somehow detect and remove the generated voice?"
            },
            {
              "id": 3204364,
              "postDate": "2025-05-18T07:30:17.320Z",
              "content": "<p>Replying to my own comment for posterity (and it has been mentioned elsewhere): by now I also have noticed that it isn't the label but the recording id that is being read out. I.e. no danger here</p>",
              "rawMarkdown": "Replying to my own comment for posterity (and it has been mentioned elsewhere): by now I also have noticed that it isn't the label but the recording id that is being read out. I.e. no danger here"
            }
          ]
        }
      ]
    },
    {
      "id": 3204607,
      "postDate": "2025-05-18T14:41:06.080Z",
      "content": "<p>This was really helpful. I did notice in some cases animals can be misclassified so folks should just take care to confirm if deciding to remove them outright. ruther1 for example was consistently marked as a human for me: <a href=\"https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1\" target=\"_blank\">https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1</a></p>\n<p>Really appreciate the work here -- I had tried other approaches but I found the one here the most consistent.</p>",
      "rawMarkdown": "This was really helpful. I did notice in some cases animals can be misclassified so folks should just take care to confirm if deciding to remove them outright. ruther1 for example was consistently marked as a human for me: https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1\n\nReally appreciate the work here -- I had tried other approaches but I found the one here the most consistent."
    },
    {
      "id": 3194321,
      "postDate": "2025-05-05T17:06:08.510Z",
      "content": "<p><strong>How do you handle the audio files listed in <code>train_voice_summary.txt</code> and <code>ss_voice_summary.txt</code></strong>?</p>\n<p>When applying data augmentation, do you split the audio into segments and then exclude only the segments that contain human voice? Or do you exclude the entire audio files listed in <code>train_voice_summary.txt</code> and <code>ss_voice_summary.txt</code> altogether?</p>",
      "rawMarkdown": "**How do you handle the audio files listed in `train_voice_summary.txt` and `ss_voice_summary.txt`**?\n\nWhen applying data augmentation, do you split the audio into segments and then exclude only the segments that contain human voice? Or do you exclude the entire audio files listed in `train_voice_summary.txt` and `ss_voice_summary.txt` altogether?\n\n",
      "replies": [
        {
          "id": 3194813,
          "postDate": "2025-05-06T10:21:43.967Z",
          "content": "<p>I exclude only the segments containing human voice, not the entire audios, because the number of records is quite small</p>",
          "rawMarkdown": "I exclude only the segments containing human voice, not the entire audios, because the number of records is quite small",
          "votes": 1
        }
      ]
    },
    {
      "id": 3173239,
      "postDate": "2025-04-07T17:16:25.770Z",
      "content": "<p>Great work very informative </p>",
      "rawMarkdown": "Great work very informative "
    },
    {
      "id": 3161992,
      "postDate": "2025-03-28T17:22:38.580Z",
      "content": "<p>This is really helpful! Thanks a lot!<br>\nSince there are human voices along with the bird sounds, rather than cutting them all, can we use LSTM model with attention mechanism and train it to focus only on patterns representing bird sounds?</p>",
      "rawMarkdown": "This is really helpful! Thanks a lot!\nSince there are human voices along with the bird sounds, rather than cutting them all, can we use LSTM model with attention mechanism and train it to focus only on patterns representing bird sounds?",
      "replies": [
        {
          "id": 3162780,
          "postDate": "2025-03-29T18:14:05.377Z",
          "content": "<p>That's what we are supposed to do (however, I didn't experiment with LSTM). I think, that the human voice is a problem, but only when dealing with the <code>train_audio</code>. Because we have such a voice in every species recording, it's hard for the NN to distinguish between the voice and the species call. In the soundscapes, on the contrary, the human voice can be considered as an additional augmentation.</p>",
          "rawMarkdown": "That's what we are supposed to do (however, I didn't experiment with LSTM). I think, that the human voice is a problem, but only when dealing with the `train_audio`. Because we have such a voice in every species recording, it's hard for the NN to distinguish between the voice and the species call. In the soundscapes, on the contrary, the human voice can be considered as an additional augmentation.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3159456,
      "postDate": "2025-03-25T16:08:39.403Z",
      "content": "<p>So, what to do with those <code>human-voice time segments</code>? Shall we discard them? if we are discarding them does that mean <code>test soundscapes</code> don't have <code>human-voice</code> in them?</p>",
      "rawMarkdown": "So, what to do with those `human-voice time segments`? Shall we discard them? if we are discarding them does that mean `test soundscapes` don't have `human-voice` in them?",
      "replies": [
        {
          "id": 3159888,
          "postDate": "2025-03-26T03:21:32.773Z",
          "content": "<p>I have been working on human-voice cleaning tasks recently, but I found that after cleaning the human-voice data, the PB score decreased significantly. I am also wondering if the test set contains human voice as well.</p>",
          "rawMarkdown": "I have been working on human-voice cleaning tasks recently, but I found that after cleaning the human-voice data, the PB score decreased significantly. I am also wondering if the test set contains human voice as well.",
          "votes": 1,
          "replies": [
            {
              "id": 3160891,
              "postDate": "2025-03-27T08:56:46.870Z",
              "content": "<p>I think, we should be careful with removing human voice if we do so with soundscapes (either test or train), because it may shift the statistical distribution of interferences or lead to data loss. But we can certainly remove the voice from train_audio, and particularly from CSA dataset where it is not a background noise but rather a description of the heard species and recording environment.</p>",
              "rawMarkdown": "I think, we should be careful with removing human voice if we do so with soundscapes (either test or train), because it may shift the statistical distribution of interferences or lead to data loss. But we can certainly remove the voice from train_audio, and particularly from CSA dataset where it is not a background noise but rather a description of the heard species and recording environment.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3157816,
      "postDate": "2025-03-23T21:37:34.213Z",
      "content": "<p>Wow thanks! I was trying other VAD but they were either too heavy or terrible performance. Nice to see some validation on this idea. I was staring to think it was a bad idea 🥴</p>",
      "rawMarkdown": "Wow thanks! I was trying other VAD but they were either too heavy or terrible performance. Nice to see some validation on this idea. I was staring to think it was a bad idea 🥴"
    },
    {
      "id": 3163546,
      "postDate": "2025-03-30T23:12:33.260Z",
      "content": "<p>This is very helpful. Thank you!</p>",
      "rawMarkdown": "This is very helpful. Thank you!"
    },
    {
      "id": 3160169,
      "postDate": "2025-03-26T12:55:46.907Z",
      "content": "<p>great work</p>",
      "rawMarkdown": "great work"
    },
    {
      "id": 3157669,
      "postDate": "2025-03-23T17:10:19.980Z",
      "content": "<p>Thank you for your reminder!</p>",
      "rawMarkdown": "Thank you for your reminder!"
    },
    {
      "id": 3177728,
      "postDate": "2025-04-13T07:29:44.657Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3209121,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-25T09:10:53.667000",
      "content": "<p>It provide a clear understanding, very helpful! Thx!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3154723,
      "author_name": "Tomoki Yokoyama",
      "author_url": "",
      "post_date": "2025-03-20T10:20:06.513000",
      "content": "<p>Hello, thank you so much for bringing up this idea.</p>\n<p>For all who reads this, I recommend you try <code>threshold</code> parameter. For example,<br>\n<code>speech_timestamps = get_speech_timestamps(torch.Tensor(segment), model, threshold=0.1)</code></p>\n<p><a href=\"https://github.com/snakers4/silero-vad/wiki/Quality-Metrics#threshold\" target=\"_blank\">https://github.com/snakers4/silero-vad/wiki/Quality-Metrics#threshold</a></p>\n<p>By default, the threshold is set to 0.5, but the model didn't work for <code>/kaggle/input/birdclef-2025/train_audio/1139490/CSA36385.ogg</code>. Once I've tweaked it to 0.4, it started working as expected.<br>\nHope this helps someone.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3154821,
          "author_name": "Konstantin Dmitriev",
          "author_url": "",
          "post_date": "2025-03-20T12:35:32.647000",
          "content": "<p><a href=\"https://www.kaggle.com/tomok1\" target=\"_blank\">@tomok1</a>, thank you very much for your suggestion! </p>\n<p>I have made the corrections in the code, introducing the threshold of 0.4. The new version will be ready in about 8-9 hours.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3200997,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-13T10:26:39.040000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3153738,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-19T05:46:40.737000",
      "content": "<p>Can we  find these voices remove  of human from  dataset  make the data clean for only bird voices.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3154168,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2025-03-19T15:46:28.997000",
          "content": "<p>This is a tricky question.  There are methods to identify if human voice in a recording segment - a few have been shared in notebooks or discussion.  So you could clean up the data - the issue, what happens if humans are talking in the test?   I am hoping to teach my model that humans are the 207th class in the data.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 3154220,
              "author_name": "Konstantin Dmitriev",
              "author_url": "",
              "post_date": "2025-03-19T16:56:33.923000",
              "content": "<p>My opinion is, humans do talk in the test, but not too much.</p>\n<p>In the <code>train_sounscapes</code>, there are 143 recordings that were confirmed to contain human voice. It is 1.5% of all soundscapes. However, the human voice inside <code>train_audio</code> is more dangerous, as the model can easily be misled by the clear and loud sound of speech instead of learning the weak insect sounds.</p>\n<p>I think, it's a good idea to remove human voice form <code>train_audio</code>, but leave it inside <code>train_soundscapes</code> to address the domain shift correctly.</p>",
              "votes": 13,
              "replies": []
            },
            {
              "id": 3154446,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-20T00:01:13.680000",
              "content": "<p>Thanks.Could you please  tell me one more  thing i do not understand it we use train_audio for training data while prediction on we use test_audio.  What will be use of  train_soundscapes unlabeled data.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3154486,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2025-03-20T02:30:28.773000",
              "content": "<p>You can test that your model will process soundscapes by using train - I copy a couple over to the test folder on my local machine.</p>\n<p>You may also be able to label some of the audio and perhaps use it as additional training data.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3154516,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-20T03:04:02.007000",
              "content": "<p>So they are not testing data but we can use them to see if our model is working fine to testing ready for submission process we can test with this direocty or labeled them use them for extra data for training am I get it right ?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3154524,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2025-03-20T03:09:01.713000",
              "content": "<p>As noted in the over view the train soundscapes are similar to the test.  If your model can do a good job of predicting than some of the unlabeled train soundscapes could make good training data to improve your model.  The training data we have been given is from all over the world, there is a decent chance that the train and test soundscapes have been recorded in a smaller number of places around the world.</p>\n<p>You need to have decent confidence that your model is doing a good job of labeling, but for sure my plan is to try to add as much of the train soundscapes as I can for further training data. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3154543,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-20T03:38:47.900000",
              "content": "<p>Thanks. I was able to submit my first submisison file. Thanks for you help. I really appreciate it.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3158620,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2025-03-24T17:53:09.723000",
              "content": "<p>I did a quick change were I removed any recording that had speech for training - bad idea as I did not end up with 206 classes :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3162081,
              "author_name": "IkaroSilva",
              "author_url": "",
              "post_date": "2025-03-28T19:23:05.867000",
              "content": "<p>\"but rather a description of the heard species and recording environment.\"</p>\n<p>I hope that is not the case on the test environment as well!</p>\n<p>Otherwise this competition could end up picking the model that does the best human speech translation…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3203446,
      "author_name": "Diganta",
      "author_url": "",
      "post_date": "2025-05-16T18:36:31.657000",
      "content": "<p>Thanks to you, my LB improved! I would have never figured out the fix like this! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3203384,
      "author_name": "Liu Zhongliang",
      "author_url": "",
      "post_date": "2025-05-16T16:50:09.707000",
      "content": "<p>Discarding frames that only contain human voices seems to really hurt my model’s performance. There are still some key points I haven’t figured out.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3203511,
          "author_name": "Kevin Willemse",
          "author_url": "",
          "post_date": "2025-05-16T20:43:23.850000",
          "content": "<p>What do your 5 second audio segments contain that are left over?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3200773,
      "author_name": "MYSO",
      "author_url": "",
      "post_date": "2025-05-13T03:41:10.737000",
      "content": "<p>Thank you for sharing the information. Thanks to this, I was able to significantly improve my LB score.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3200796,
          "author_name": "Baiph",
          "author_url": "",
          "post_date": "2025-05-13T04:32:22.637000",
          "content": "<p>great! I removed the human voice and the lb dropped a lot. I will continue to explore</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3171349,
      "author_name": "Yanik-Pascal Förster",
      "author_url": "",
      "post_date": "2025-04-05T15:37:48.313000",
      "content": "<p>I noticed the problem but didn’t get as far as the OP in cracking the nut! 🥳 I was wondering however: Some of the CSA recordings, where there is a human voice, also have a synthetic voice reading out the label. I’m half convinced that that will create leakage somehow. Anybody had the same thought?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3172291,
          "author_name": "IkaroSilva",
          "author_url": "",
          "post_date": "2025-04-06T17:02:03.357000",
          "content": "<p>Aggreed. I posted a similar question but I did not get any replies from the organizers.</p>\n<p>If that is the case I am concerned if it is possible some of the top entries could be doing speech translation on the test set ( specially because you can access data way before/after the target region).</p>\n<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>  could you please clarify that any speech in the test set <em>does</em> not include the labels we are trying to identify? Thank you.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3172311,
              "author_name": "Yanik-Pascal Förster",
              "author_url": "",
              "post_date": "2025-04-06T17:17:26.573000",
              "content": "<p>I see what you mean. My understanding is that the test set would be soundscapes, and <em>naively</em> I’d assume that these are unlikely to include the <em>computer</em> voices reading out the labels clearly and distinctly. So I didn’t even thought about this possibility. Though I guess it is possible that a human voice in a test soundscape could mention one of the species recorded. </p>\n<p>My original question was meant like: If it isn’t advised to clean the train data of human voices altogether, because the test soundscapes are likely to contain some, would it still make sense to somehow detect and remove the generated voice?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3204364,
              "author_name": "Yanik-Pascal Förster",
              "author_url": "",
              "post_date": "2025-05-18T07:30:17.320000",
              "content": "<p>Replying to my own comment for posterity (and it has been mentioned elsewhere): by now I also have noticed that it isn't the label but the recording id that is being read out. I.e. no danger here</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3204607,
      "author_name": "Nineso",
      "author_url": "",
      "post_date": "2025-05-18T14:41:06.080000",
      "content": "<p>This was really helpful. I did notice in some cases animals can be misclassified so folks should just take care to confirm if deciding to remove them outright. ruther1 for example was consistently marked as a human for me: <a href=\"https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1\" target=\"_blank\">https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1</a></p>\n<p>Really appreciate the work here -- I had tried other approaches but I found the one here the most consistent.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3194321,
      "author_name": "akacome",
      "author_url": "",
      "post_date": "2025-05-05T17:06:08.510000",
      "content": "<p><strong>How do you handle the audio files listed in <code>train_voice_summary.txt</code> and <code>ss_voice_summary.txt</code></strong>?</p>\n<p>When applying data augmentation, do you split the audio into segments and then exclude only the segments that contain human voice? Or do you exclude the entire audio files listed in <code>train_voice_summary.txt</code> and <code>ss_voice_summary.txt</code> altogether?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3194813,
          "author_name": "Konstantin Dmitriev",
          "author_url": "",
          "post_date": "2025-05-06T10:21:43.967000",
          "content": "<p>I exclude only the segments containing human voice, not the entire audios, because the number of records is quite small</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3173239,
      "author_name": "Amos Shehzad",
      "author_url": "",
      "post_date": "2025-04-07T17:16:25.770000",
      "content": "<p>Great work very informative </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3161992,
      "author_name": "Shivam",
      "author_url": "",
      "post_date": "2025-03-28T17:22:38.580000",
      "content": "<p>This is really helpful! Thanks a lot!<br>\nSince there are human voices along with the bird sounds, rather than cutting them all, can we use LSTM model with attention mechanism and train it to focus only on patterns representing bird sounds?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3162780,
          "author_name": "Konstantin Dmitriev",
          "author_url": "",
          "post_date": "2025-03-29T18:14:05.377000",
          "content": "<p>That's what we are supposed to do (however, I didn't experiment with LSTM). I think, that the human voice is a problem, but only when dealing with the <code>train_audio</code>. Because we have such a voice in every species recording, it's hard for the NN to distinguish between the voice and the species call. In the soundscapes, on the contrary, the human voice can be considered as an additional augmentation.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3159456,
      "author_name": "Hotson Honet",
      "author_url": "",
      "post_date": "2025-03-25T16:08:39.403000",
      "content": "<p>So, what to do with those <code>human-voice time segments</code>? Shall we discard them? if we are discarding them does that mean <code>test soundscapes</code> don't have <code>human-voice</code> in them?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3159888,
          "author_name": "noshakeplz",
          "author_url": "",
          "post_date": "2025-03-26T03:21:32.773000",
          "content": "<p>I have been working on human-voice cleaning tasks recently, but I found that after cleaning the human-voice data, the PB score decreased significantly. I am also wondering if the test set contains human voice as well.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3160891,
              "author_name": "Konstantin Dmitriev",
              "author_url": "",
              "post_date": "2025-03-27T08:56:46.870000",
              "content": "<p>I think, we should be careful with removing human voice if we do so with soundscapes (either test or train), because it may shift the statistical distribution of interferences or lead to data loss. But we can certainly remove the voice from train_audio, and particularly from CSA dataset where it is not a background noise but rather a description of the heard species and recording environment.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3157816,
      "author_name": "Diego Akel",
      "author_url": "",
      "post_date": "2025-03-23T21:37:34.213000",
      "content": "<p>Wow thanks! I was trying other VAD but they were either too heavy or terrible performance. Nice to see some validation on this idea. I was staring to think it was a bad idea 🥴</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3163546,
      "author_name": "Nayer Basim",
      "author_url": "",
      "post_date": "2025-03-30T23:12:33.260000",
      "content": "<p>This is very helpful. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3160169,
      "author_name": "gourav_gujariya",
      "author_url": "",
      "post_date": "2025-03-26T12:55:46.907000",
      "content": "<p>great work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3157669,
      "author_name": "想送你花.",
      "author_url": "",
      "post_date": "2025-03-23T17:10:19.980000",
      "content": "<p>Thank you for your reminder!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3177728,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-13T07:29:44.657000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3153227": "It was noticed by many participants, that the provided data often contains human voice. This is particularly  related to CSA dataset. I decided to explore this factor and maybe get some insights from the data.\n\nThe details are presented [**in the notebook**](https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data).\n# 1. Exploration\n## Is human voice presence common for all CSA recordings?\nAt first, I looked at the recordings by Fabio A. Sarria-S, as @martinapreusse said, they contain voice. I have plotted some kind of dependence of the sound level on time, and a clear pattern appeared. The actual insect is separated from human description by short pauses.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fd89b7b930333d100c26feca344747b03%2F1.png?generation=1742306694933745&alt=media)\n\nAfter a short analysis, it was clear, that **almost all CSA recordings contain human voice**.\n\nThis could be easy to exclude human voice, but the recordings by other authors have different patterns. For example, Alexandra Buitrago-Cardona says something at the beginning of the recordings (and, sometimes, at the end or in the middle). \n\n## Do other collections contain human voice?\nA short answer: **yes**!\n\n## Do train soundscapes contain human voice?\nAnd again, the answer is **yes**!\n\n# 2. Solution\nI found a wonderful library `silero-vad` that can be used to detect human voice in the recordings. The use of the library is very simple:\n```\nmodel, (get_speech_timestamps, _, read_audio, _, _) = torch.hub.load(repo_or_dir='snakers4/silero-vad', model='silero_vad')\nspeech_timestamps = get_speech_timestamps(torch.Tensor(wav), model)\n```\nThe timestamps it provides are quite accurate. I have compared them with manual audio analysis. Also they correlate with the 'split-by-pause' method (if it is applicable). Below are some examples.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2F1ce810158bb88b54fb76dedc6e871843%2F2.png?generation=1742307610054378&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fa3f8814f8f3c270c4ae94ecea40a38af%2F3.png?generation=1742307619626563&alt=media)\n\n# 3. Results\nI have run my  [**notebook**](https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data) on both train soundscapes and train audio data. The resulting [**dataset**](https://www.kaggle.com/datasets/kdmitrie/bc25-separation-voice-from-data-by-silero-vad) contains several files:\n- *_voice_summary.txt - a list of files where human voice was found;\n- *_voice_data.pkl - a dictionary saved with `pickle` library. Its keys are the names of files with human voice; each value is a list with exact locations of human voice.\n\n",
    "3209121": "It provide a clear understanding, very helpful! Thx!",
    "3154723": "Hello, thank you so much for bringing up this idea.\n\nFor all who reads this, I recommend you try `threshold` parameter. For example,\n`speech_timestamps = get_speech_timestamps(torch.Tensor(segment), model, threshold=0.1)`\n\nhttps://github.com/snakers4/silero-vad/wiki/Quality-Metrics#threshold\n\nBy default, the threshold is set to 0.5, but the model didn't work for `/kaggle/input/birdclef-2025/train_audio/1139490/CSA36385.ogg`. Once I've tweaked it to 0.4, it started working as expected.\nHope this helps someone.",
    "3153738": "Can we  find these voices remove  of human from  dataset  make the data clean for only bird voices.",
    "3203446": "Thanks to you, my LB improved! I would have never figured out the fix like this! ",
    "3203384": "Discarding frames that only contain human voices seems to really hurt my model’s performance. There are still some key points I haven’t figured out.",
    "3200773": "Thank you for sharing the information. Thanks to this, I was able to significantly improve my LB score.",
    "3171349": "I noticed the problem but didn’t get as far as the OP in cracking the nut! 🥳 I was wondering however: Some of the CSA recordings, where there is a human voice, also have a synthetic voice reading out the label. I’m half convinced that that will create leakage somehow. Anybody had the same thought?",
    "3204607": "This was really helpful. I did notice in some cases animals can be misclassified so folks should just take care to confirm if deciding to remove them outright. ruther1 for example was consistently marked as a human for me: https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1\n\nReally appreciate the work here -- I had tried other approaches but I found the one here the most consistent.",
    "3194321": "**How do you handle the audio files listed in `train_voice_summary.txt` and `ss_voice_summary.txt`**?\n\nWhen applying data augmentation, do you split the audio into segments and then exclude only the segments that contain human voice? Or do you exclude the entire audio files listed in `train_voice_summary.txt` and `ss_voice_summary.txt` altogether?\n\n",
    "3173239": "Great work very informative ",
    "3161992": "This is really helpful! Thanks a lot!\nSince there are human voices along with the bird sounds, rather than cutting them all, can we use LSTM model with attention mechanism and train it to focus only on patterns representing bird sounds?",
    "3159456": "So, what to do with those `human-voice time segments`? Shall we discard them? if we are discarding them does that mean `test soundscapes` don't have `human-voice` in them?",
    "3157816": "Wow thanks! I was trying other VAD but they were either too heavy or terrible performance. Nice to see some validation on this idea. I was staring to think it was a bad idea 🥴",
    "3163546": "This is very helpful. Thank you!",
    "3160169": "great work",
    "3157669": "Thank you for your reminder!",
    "3177728": ""
  }
}