{
  "id": 579407,
  "title": "Sampling Strategy",
  "url": "/competitions/birdclef-2025/discussion/579407",
  "author_name": "",
  "post_date": "2025-05-17T08:26:17.773327300Z",
  "votes": 16,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Based on my latest experiments.</p>\n<p>RMS Sampling &gt; Random Sampling &gt; First 5 seconds.</p>\n<p>Anyone tried any other strategies?</p>",
  "messages": [
    {
      "id": "3203723",
      "postDate": "05/17/2025 08:26:17",
      "content": "<p>Based on my latest experiments.</p>\n<p>RMS Sampling &gt; Random Sampling &gt; First 5 seconds.</p>\n<p>Anyone tried any other strategies?</p>",
      "rawMarkdown": "Based on my latest experiments.\n\nRMS Sampling > Random Sampling > First 5 seconds.\n\nAnyone tried any other strategies?",
      "votes": null
    },
    {
      "id": "3203803",
      "postDate": "05/17/2025 10:47:02",
      "content": "<p>Thanks for sharing! May I ask what's RMS Sampling?</p>",
      "rawMarkdown": "Thanks for sharing! May I ask what's RMS Sampling?",
      "votes": null
    },
    {
      "id": "3203804",
      "postDate": "05/17/2025 10:51:25",
      "content": "<p>it is root mean square based energy sampling: basically it highlights parts of audio containing high activity areas giving you a way to prevent the non-signal regions being passed into the model during training. </p>",
      "rawMarkdown": "it is root mean square based energy sampling: basically it highlights parts of audio containing high activity areas giving you a way to prevent the non-signal regions being passed into the model during training.",
      "votes": null
    },
    {
      "id": "3203849",
      "postDate": "05/17/2025 11:48:27",
      "content": "<p>Yeah, basically, we pick the chunk with highest RMS value in the audio file.</p>",
      "rawMarkdown": "Yeah, basically, we pick the chunk with highest RMS value in the audio file.",
      "votes": null
    },
    {
      "id": "3203856",
      "postDate": "05/17/2025 11:55:33",
      "content": "<p>I had previously considered using certain sound features for sampling, but I am worried about interference from human voices.</p>\n<p>Additionally, in the early stages of using the 5-fold cross-validation strategy, I predicted the out-of-fold audio for each model and selected the segment with the highest probability of the true label. However, this significantly lowered the lb.</p>",
      "rawMarkdown": "I had previously considered using certain sound features for sampling, but I am worried about interference from human voices.\n\nAdditionally, in the early stages of using the 5-fold cross-validation strategy, I predicted the out-of-fold audio for each model and selected the segment with the highest probability of the true label. However, this significantly lowered the lb.",
      "votes": null
    },
    {
      "id": "3203876",
      "postDate": "05/17/2025 12:30:06",
      "content": "<p>Nice Idea!!!</p>",
      "rawMarkdown": "Nice Idea!!!",
      "votes": null
    },
    {
      "id": "3203912",
      "postDate": "05/17/2025 13:29:19",
      "content": "<p>quick note and a word of caution pursuing the above approach, make sure you have logic in place to deal with human voices in the data or else the rms strategy will always focus on those areas of the audio defeating the purpose completely!</p>",
      "rawMarkdown": "quick note and a word of caution pursuing the above approach, make sure you have logic in place to deal with human voices in the data or else the rms strategy will always focus on those areas of the audio defeating the purpose completely!",
      "votes": null
    },
    {
      "id": "3203934",
      "postDate": "05/17/2025 14:11:49",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "3204238",
      "postDate": "05/18/2025 02:14:15",
      "content": "<p>What about center 5 seconds? I always try it, but I think I should try another. Center 5 seconds always have lower than half of human voice.</p>",
      "rawMarkdown": "What about center 5 seconds? I always try it, but I think I should try another. Center 5 seconds always have lower than half of human voice.",
      "votes": null
    },
    {
      "id": "3205091",
      "postDate": "05/19/2025 10:07:18",
      "content": "<p>Thank for sharing! Does RMS sampling means the sampling probability is depend on RMS power? Or just chunk the signal using the maximum RMS power value? To be specific, for a datapoint, does the chunk fixed in every epoch? Or the probability of chunks with large RMS power are larger than chunks with relatively low RMS power?</p>",
      "rawMarkdown": "Thank for sharing! Does RMS sampling means the sampling probability is depend on RMS power? Or just chunk the signal using the maximum RMS power value? To be specific, for a datapoint, does the chunk fixed in every epoch? Or the probability of chunks with large RMS power are larger than chunks with relatively low RMS power?",
      "votes": null
    },
    {
      "id": "3205110",
      "postDate": "05/19/2025 10:35:31",
      "content": "<p>In my experiment, for sed model, random &gt; rms; for cnn, rms &gt; random. Both have remove human voice clips</p>",
      "rawMarkdown": "In my experiment, for sed model, random > rms; for cnn, rms > random. Both have remove human voice clips",
      "votes": null
    },
    {
      "id": "3205111",
      "postDate": "05/19/2025 10:35:37",
      "content": "<p>I pick like this.</p>\n<pre><code> ():\n    duration_samples = duration_samples * \n     (waves) &lt;= duration_samples:\n        \n        repeat_count = (np.ceil(duration_samples / (waves)))\n        waves = np.tile(waves, repeat_count)[:duration_samples]\n         waves\n\n    stride = \n    max_rms = \n    max_rms_start = \n\n     start  (, (waves) - duration_samples + , stride):\n        window = waves[start:start + duration_samples]\n        rms = np.sqrt(np.mean(window ** ))\n         rms &gt; max_rms:\n            max_rms = rms\n            max_rms_start = start\n\n    wave = waves[max_rms_start:max_rms_start + duration_samples]\n     wave\n</code></pre>",
      "rawMarkdown": "I pick like this.\n```python\n\ndef pick_rms(self, waves, duration_samples):\n    duration_samples = duration_samples * 32000\n    if len(waves) <= duration_samples:\n        # Circular padding\n        repeat_count = int(np.ceil(duration_samples / len(waves)))\n        waves = np.tile(waves, repeat_count)[:duration_samples]\n        return waves\n\n    stride = 32000\n    max_rms = 0\n    max_rms_start = 0\n\n    for start in range(0, len(waves) - duration_samples + 1, stride):\n        window = waves[start:start + duration_samples]\n        rms = np.sqrt(np.mean(window ** 2))\n        if rms > max_rms:\n            max_rms = rms\n            max_rms_start = start\n\n    wave = waves[max_rms_start:max_rms_start + duration_samples]\n    return wave\n\n```",
      "votes": null
    },
    {
      "id": "3205261",
      "postDate": "05/19/2025 15:01:58",
      "content": "<p>I'd like to ask if these operations you are talking about are carried out during data processing, or throughout the entire process (I mean data processing, model training, inference process)</p>",
      "rawMarkdown": "I'd like to ask if these operations you are talking about are carried out during data processing, or throughout the entire process (I mean data processing, model training, inference process)",
      "votes": null
    },
    {
      "id": "3205271",
      "postDate": "05/19/2025 15:16:39",
      "content": "<p>I think it is just during data processing. </p>",
      "rawMarkdown": "I think it is just during data processing.",
      "votes": null
    },
    {
      "id": "3207751",
      "postDate": "05/23/2025 07:33:05",
      "content": "<p>Very helpful tips, appreciate!</p>",
      "rawMarkdown": "Very helpful tips, appreciate!",
      "votes": null
    },
    {
      "id": "3208607",
      "postDate": "05/24/2025 11:56:19",
      "content": "<p>Here's an attempt to remove speech. I may be throwing the baby out with the bath water but it should skip high rms clips with speech</p>\n<pre><code> librosa\n\n () -&gt; np.ndarray:\n    duration_samples = duration_samples * \n     (waves) &lt;= duration_samples:\n        \n        repeat_count = (np.ceil(duration_samples / (waves)))\n        waves = np.tile(waves, repeat_count)[:duration_samples]\n         waves\n\n    stride = \n    max_rms = \n    max_rms_start = \n\n    mel = librosa.feature.melspectrogram(\n        y=waves,\n        sr=,\n        fmin=,\n        fmax=,\n    )\n    mel_freqs = librosa.mel_frequencies(n_mels=mel.shape[], fmin=, fmax=)\n    speech_bins = np.where(mel_freqs &lt;= )[]\n    non_speech_bins = np.where(mel_freqs &gt; )[]\n\n     start  (, (waves) - duration_samples + , stride):\n        window = waves[start : start + duration_samples]\n        speech_power = np.(mel[speech_bins, start : start + duration_samples])\n        non_speech_power = np.(\n            mel[non_speech_bins, start : start + duration_samples]\n        )\n        rms = np.sqrt(np.mean(window**))\n         rms &gt; max_rms  speech_power / non_speech_power &lt; :\n            max_rms = rms\n            max_rms_start = start\n\n    wave = waves[max_rms_start : max_rms_start + duration_samples]\n     wave\n</code></pre>\n<p>I'm applying signal processing before the above function so a threshold of 1 may not be universal. Full disclosure: also not doing well in this one. </p>",
      "rawMarkdown": "Here's an attempt to remove speech. I may be throwing the baby out with the bath water but it should skip high rms clips with speech\n\n```python\nimport librosa\n\ndef pick_rms(waves: np.ndarray, duration_samples: int = 5) -> np.ndarray:\n    duration_samples = duration_samples * 32000\n    if len(waves) <= duration_samples:\n        # Circular padding\n        repeat_count = int(np.ceil(duration_samples / len(waves)))\n        waves = np.tile(waves, repeat_count)[:duration_samples]\n        return waves\n\n    stride = 32000\n    max_rms = 0\n    max_rms_start = 0\n\n    mel = librosa.feature.melspectrogram(\n        y=waves,\n        sr=32000,\n        fmin=10,\n        fmax=15500,\n    )\n    mel_freqs = librosa.mel_frequencies(n_mels=mel.shape[0], fmin=10, fmax=15500)\n    speech_bins = np.where(mel_freqs <= 3400)[0]\n    non_speech_bins = np.where(mel_freqs > 3400)[0]\n\n    for start in range(0, len(waves) - duration_samples + 1, stride):\n        window = waves[start : start + duration_samples]\n        speech_power = np.sum(mel[speech_bins, start : start + duration_samples])\n        non_speech_power = np.sum(\n            mel[non_speech_bins, start : start + duration_samples]\n        )\n        rms = np.sqrt(np.mean(window**2))\n        if rms > max_rms and speech_power / non_speech_power < 1:\n            max_rms = rms\n            max_rms_start = start\n\n    wave = waves[max_rms_start : max_rms_start + duration_samples]\n    return wave\n```\nI'm applying signal processing before the above function so a threshold of 1 may not be universal. Full disclosure: also not doing well in this one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3203803,
      "author_name": "xuanmingzhang777",
      "author_url": "",
      "post_date": "05/17/2025 10:47:02",
      "content": "<p>Thanks for sharing! May I ask what's RMS Sampling?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3203804,
          "author_name": "lavanbth99",
          "author_url": "",
          "post_date": "05/17/2025 10:51:25",
          "content": "<p>it is root mean square based energy sampling: basically it highlights parts of audio containing high activity areas giving you a way to prevent the non-signal regions being passed into the model during training. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3203876,
              "author_name": "xukongji",
              "author_url": "",
              "post_date": "05/17/2025 12:30:06",
              "content": "<p>Nice Idea!!!</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3203849,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "05/17/2025 11:48:27",
          "content": "<p>Yeah, basically, we pick the chunk with highest RMS value in the audio file.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3203912,
              "author_name": "lavanbth99",
              "author_url": "",
              "post_date": "05/17/2025 13:29:19",
              "content": "<p>quick note and a word of caution pursuing the above approach, make sure you have logic in place to deal with human voices in the data or else the rms strategy will always focus on those areas of the audio defeating the purpose completely!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3208607,
                  "author_name": "bgeier",
                  "author_url": "",
                  "post_date": "05/24/2025 11:56:19",
                  "content": "<p>Here's an attempt to remove speech. I may be throwing the baby out with the bath water but it should skip high rms clips with speech</p>\n<pre><code> librosa\n\n () -&gt; np.ndarray:\n    duration_samples = duration_samples * \n     (waves) &lt;= duration_samples:\n        \n        repeat_count = (np.ceil(duration_samples / (waves)))\n        waves = np.tile(waves, repeat_count)[:duration_samples]\n         waves\n\n    stride = \n    max_rms = \n    max_rms_start = \n\n    mel = librosa.feature.melspectrogram(\n        y=waves,\n        sr=,\n        fmin=,\n        fmax=,\n    )\n    mel_freqs = librosa.mel_frequencies(n_mels=mel.shape[], fmin=, fmax=)\n    speech_bins = np.where(mel_freqs &lt;= )[]\n    non_speech_bins = np.where(mel_freqs &gt; )[]\n\n     start  (, (waves) - duration_samples + , stride):\n        window = waves[start : start + duration_samples]\n        speech_power = np.(mel[speech_bins, start : start + duration_samples])\n        non_speech_power = np.(\n            mel[non_speech_bins, start : start + duration_samples]\n        )\n        rms = np.sqrt(np.mean(window**))\n         rms &gt; max_rms  speech_power / non_speech_power &lt; :\n            max_rms = rms\n            max_rms_start = start\n\n    wave = waves[max_rms_start : max_rms_start + duration_samples]\n     wave\n</code></pre>\n<p>I'm applying signal processing before the above function so a threshold of 1 may not be universal. Full disclosure: also not doing well in this one. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3203856,
      "author_name": "ryenhails",
      "author_url": "",
      "post_date": "05/17/2025 11:55:33",
      "content": "<p>I had previously considered using certain sound features for sampling, but I am worried about interference from human voices.</p>\n<p>Additionally, in the early stages of using the 5-fold cross-validation strategy, I predicted the out-of-fold audio for each model and selected the segment with the highest probability of the true label. However, this significantly lowered the lb.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3203934,
      "author_name": "yandsbnb666zhang",
      "author_url": "",
      "post_date": "05/17/2025 14:11:49",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3204238,
      "author_name": "kurisew",
      "author_url": "",
      "post_date": "05/18/2025 02:14:15",
      "content": "<p>What about center 5 seconds? I always try it, but I think I should try another. Center 5 seconds always have lower than half of human voice.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3205091,
      "author_name": "shtljw",
      "author_url": "",
      "post_date": "05/19/2025 10:07:18",
      "content": "<p>Thank for sharing! Does RMS sampling means the sampling probability is depend on RMS power? Or just chunk the signal using the maximum RMS power value? To be specific, for a datapoint, does the chunk fixed in every epoch? Or the probability of chunks with large RMS power are larger than chunks with relatively low RMS power?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3205111,
          "author_name": "salmanahmedtamu",
          "author_url": "",
          "post_date": "05/19/2025 10:35:37",
          "content": "<p>I pick like this.</p>\n<pre><code> ():\n    duration_samples = duration_samples * \n     (waves) &lt;= duration_samples:\n        \n        repeat_count = (np.ceil(duration_samples / (waves)))\n        waves = np.tile(waves, repeat_count)[:duration_samples]\n         waves\n\n    stride = \n    max_rms = \n    max_rms_start = \n\n     start  (, (waves) - duration_samples + , stride):\n        window = waves[start:start + duration_samples]\n        rms = np.sqrt(np.mean(window ** ))\n         rms &gt; max_rms:\n            max_rms = rms\n            max_rms_start = start\n\n    wave = waves[max_rms_start:max_rms_start + duration_samples]\n     wave\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 3207751,
              "author_name": "",
              "author_url": "",
              "post_date": "05/23/2025 07:33:05",
              "content": "<p>Very helpful tips, appreciate!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3205110,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "05/19/2025 10:35:31",
      "content": "<p>In my experiment, for sed model, random &gt; rms; for cnn, rms &gt; random. Both have remove human voice clips</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3205261,
      "author_name": "xiayuxuan",
      "author_url": "",
      "post_date": "05/19/2025 15:01:58",
      "content": "<p>I'd like to ask if these operations you are talking about are carried out during data processing, or throughout the entire process (I mean data processing, model training, inference process)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3205271,
          "author_name": "shtljw",
          "author_url": "",
          "post_date": "05/19/2025 15:16:39",
          "content": "<p>I think it is just during data processing. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3203723": "Based on my latest experiments.\n\nRMS Sampling > Random Sampling > First 5 seconds.\n\nAnyone tried any other strategies?",
    "3203803": "Thanks for sharing! May I ask what's RMS Sampling?",
    "3203804": "it is root mean square based energy sampling: basically it highlights parts of audio containing high activity areas giving you a way to prevent the non-signal regions being passed into the model during training.",
    "3203849": "Yeah, basically, we pick the chunk with highest RMS value in the audio file.",
    "3203856": "I had previously considered using certain sound features for sampling, but I am worried about interference from human voices.\n\nAdditionally, in the early stages of using the 5-fold cross-validation strategy, I predicted the out-of-fold audio for each model and selected the segment with the highest probability of the true label. However, this significantly lowered the lb.",
    "3203876": "Nice Idea!!!",
    "3203912": "quick note and a word of caution pursuing the above approach, make sure you have logic in place to deal with human voices in the data or else the rms strategy will always focus on those areas of the audio defeating the purpose completely!",
    "3203934": "Thanks for sharing!",
    "3204238": "What about center 5 seconds? I always try it, but I think I should try another. Center 5 seconds always have lower than half of human voice.",
    "3205091": "Thank for sharing! Does RMS sampling means the sampling probability is depend on RMS power? Or just chunk the signal using the maximum RMS power value? To be specific, for a datapoint, does the chunk fixed in every epoch? Or the probability of chunks with large RMS power are larger than chunks with relatively low RMS power?",
    "3205110": "In my experiment, for sed model, random > rms; for cnn, rms > random. Both have remove human voice clips",
    "3205111": "I pick like this.\n```python\n\ndef pick_rms(self, waves, duration_samples):\n    duration_samples = duration_samples * 32000\n    if len(waves) <= duration_samples:\n        # Circular padding\n        repeat_count = int(np.ceil(duration_samples / len(waves)))\n        waves = np.tile(waves, repeat_count)[:duration_samples]\n        return waves\n\n    stride = 32000\n    max_rms = 0\n    max_rms_start = 0\n\n    for start in range(0, len(waves) - duration_samples + 1, stride):\n        window = waves[start:start + duration_samples]\n        rms = np.sqrt(np.mean(window ** 2))\n        if rms > max_rms:\n            max_rms = rms\n            max_rms_start = start\n\n    wave = waves[max_rms_start:max_rms_start + duration_samples]\n    return wave\n\n```",
    "3205261": "I'd like to ask if these operations you are talking about are carried out during data processing, or throughout the entire process (I mean data processing, model training, inference process)",
    "3205271": "I think it is just during data processing.",
    "3207751": "Very helpful tips, appreciate!",
    "3208607": "Here's an attempt to remove speech. I may be throwing the baby out with the bath water but it should skip high rms clips with speech\n\n```python\nimport librosa\n\ndef pick_rms(waves: np.ndarray, duration_samples: int = 5) -> np.ndarray:\n    duration_samples = duration_samples * 32000\n    if len(waves) <= duration_samples:\n        # Circular padding\n        repeat_count = int(np.ceil(duration_samples / len(waves)))\n        waves = np.tile(waves, repeat_count)[:duration_samples]\n        return waves\n\n    stride = 32000\n    max_rms = 0\n    max_rms_start = 0\n\n    mel = librosa.feature.melspectrogram(\n        y=waves,\n        sr=32000,\n        fmin=10,\n        fmax=15500,\n    )\n    mel_freqs = librosa.mel_frequencies(n_mels=mel.shape[0], fmin=10, fmax=15500)\n    speech_bins = np.where(mel_freqs <= 3400)[0]\n    non_speech_bins = np.where(mel_freqs > 3400)[0]\n\n    for start in range(0, len(waves) - duration_samples + 1, stride):\n        window = waves[start : start + duration_samples]\n        speech_power = np.sum(mel[speech_bins, start : start + duration_samples])\n        non_speech_power = np.sum(\n            mel[non_speech_bins, start : start + duration_samples]\n        )\n        rms = np.sqrt(np.mean(window**2))\n        if rms > max_rms and speech_power / non_speech_power < 1:\n            max_rms = rms\n            max_rms_start = start\n\n    wave = waves[max_rms_start : max_rms_start + duration_samples]\n    return wave\n```\nI'm applying signal processing before the above function so a threshold of 1 may not be universal. Full disclosure: also not doing well in this one."
  },
  "source": "meta"
}