{
  "id": 498984,
  "title": "How to align?",
  "url": "/competitions/birdclef-2024/discussion/498984",
  "author_name": "",
  "post_date": "2024-04-30T08:04:57.885695100Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was wondering what was your approach to align transformed sound files?</p>\n<p>For example, imagine I sample a random 5s section from an audio file and then<br>\ndo a MEL spectrogram transformation. Each audio file will produce a tensor of different sizes: (1, p).</p>\n<p>How would you align those? Or maybe you first stack the waveforms and then apply the MEL spectrogram transformation?</p>\n<p>Please share your insights, thanks!</p>",
  "messages": [
    {
      "id": "2784346",
      "postDate": "04/30/2024 08:04:57",
      "content": "<p>I was wondering what was your approach to align transformed sound files?</p>\n<p>For example, imagine I sample a random 5s section from an audio file and then<br>\ndo a MEL spectrogram transformation. Each audio file will produce a tensor of different sizes: (1, p).</p>\n<p>How would you align those? Or maybe you first stack the waveforms and then apply the MEL spectrogram transformation?</p>\n<p>Please share your insights, thanks!</p>",
      "rawMarkdown": "I was wondering what was your approach to align transformed sound files?\n\nFor example, imagine I sample a random 5s section from an audio file and then\ndo a MEL spectrogram transformation. Each audio file will produce a tensor of different sizes: (1, p).\n\nHow would you align those? Or maybe you first stack the waveforms and then apply the MEL spectrogram transformation?\n\nPlease share your insights, thanks!",
      "votes": null
    },
    {
      "id": "2784377",
      "postDate": "04/30/2024 08:32:55",
      "content": "<p>Are you trying to deal with different-length waves and turn them into Mel spectra with a consistent resolution? The following code can provide you with some references.</p>\n<pre><code> ():\n\n    waveform_orig, sample_rate = librosa.load(filename, sr=, mono=)\n\n    wave_len = (waveform_orig)\n    waveform = np.concatenate([waveform_orig, waveform_orig, waveform_orig])\n\n    effective_length = sample_rate * period\n     (waveform) &lt; (period * sample_rate * ):\n        waveform = np.concatenate([waveform, waveform_orig])\n\n     start   :\n        start = start - (period - ) /  * sample_rate\n         start &lt; :\n            start += wave_len\n        start = (start)\n    :\n         wave_len &lt; effective_length:\n            start = np.random.randint(effective_length - wave_len)\n         wave_len &gt; effective_length:\n            start = np.random.randint(wave_len - effective_length)\n         wave_len == effective_length:\n            start = \n\n    waveform_seg = waveform[start: start + (effective_length)]\n\n     waveform_orig, waveform_seg, sample_rate, start\n</code></pre>",
      "rawMarkdown": "Are you trying to deal with different-length waves and turn them into Mel spectra with a consistent resolution? The following code can provide you with some references.\n\n```python\ndef load_wave_and_crop(filename, period, start=None):\n\n    waveform_orig, sample_rate = librosa.load(filename, sr=32000, mono=False)\n\n    wave_len = len(waveform_orig)\n    waveform = np.concatenate([waveform_orig, waveform_orig, waveform_orig])\n\n    effective_length = sample_rate * period\n    while len(waveform) < (period * sample_rate * 3):\n        waveform = np.concatenate([waveform, waveform_orig])\n    \n    if start is not None:\n        start = start - (period - 5) / 2 * sample_rate\n        while start < 0:\n            start += wave_len\n        start = int(start)\n    else:\n        if wave_len < effective_length:\n            start = np.random.randint(effective_length - wave_len)\n        elif wave_len > effective_length:\n            start = np.random.randint(wave_len - effective_length)\n        elif wave_len == effective_length:\n            start = 0\n\n    waveform_seg = waveform[start: start + int(effective_length)]\n\n    return waveform_orig, waveform_seg, sample_rate, start\n```",
      "votes": null
    },
    {
      "id": "2784443",
      "postDate": "04/30/2024 09:27:14",
      "content": "<p>Why it will produce tensors of different sizes if the input size is consistent (5 secs in your example)? You can either crop or pad if the input shape is different for the audios, such as padding to 5 secs if the input is 3 or cropping if the input is 7 secs.</p>",
      "rawMarkdown": "Why it will produce tensors of different sizes if the input size is consistent (5 secs in your example)? You can either crop or pad if the input shape is different for the audios, such as padding to 5 secs if the input is 3 or cropping if the input is 7 secs.",
      "votes": null
    },
    {
      "id": "2785048",
      "postDate": "04/30/2024 15:57:39",
      "content": "<p>Indeed, if I only do this (with missing padding for shorter samples): </p>\n<pre><code> torchaudio\nx, sample_rate = torchaudio.load(file_path)\n\nx = x[:, : * sample_rate]\n</code></pre>\n<p>the tensors can be stacked, no issue.</p>\n<p>However, if I add this step:</p>\n<pre><code> torchaudio.transforms  MelSpectrogram\nx = MelSpectrogram(**MEL_PARAMS)(x)\n</code></pre>\n<p>Then, I have the shape issue.</p>\n<p>I guess I will need to stack the tensors over the different <code>batch_size</code> and then do the transformation. </p>",
      "rawMarkdown": "Indeed, if I only do this (with missing padding for shorter samples): \n\n```python\nimport torchaudio\nx, sample_rate = torchaudio.load(file_path)\n# TODO: Need to pad...\nx = x[:, 0:5 * sample_rate]\n```\n\nthe tensors can be stacked, no issue.\n\nHowever, if I add this step:\n\n```python\nfrom torchaudio.transforms import MelSpectrogram\nx = MelSpectrogram(**MEL_PARAMS)(x)\n```\n\nThen, I have the shape issue.\n\nI guess I will need to stack the tensors over the different `batch_size` and then do the transformation.",
      "votes": null
    },
    {
      "id": "2785050",
      "postDate": "04/30/2024 15:59:02",
      "content": "<p>I almost have a similar way to do the waveform extraction. It is in the next step (see my other comment). Anyway, thanks for sharing your code snippet, it might help others!</p>",
      "rawMarkdown": "I almost have a similar way to do the waveform extraction. It is in the next step (see my other comment). Anyway, thanks for sharing your code snippet, it might help others!",
      "votes": null
    },
    {
      "id": "2785293",
      "postDate": "04/30/2024 17:51:31",
      "content": "<p>I see, then yeah there are multiple examples shared in the code section which can be useful further. </p>",
      "rawMarkdown": "I see, then yeah there are multiple examples shared in the code section which can be useful further.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2784377,
      "author_name": "zz1zzi",
      "author_url": "",
      "post_date": "04/30/2024 08:32:55",
      "content": "<p>Are you trying to deal with different-length waves and turn them into Mel spectra with a consistent resolution? The following code can provide you with some references.</p>\n<pre><code> ():\n\n    waveform_orig, sample_rate = librosa.load(filename, sr=, mono=)\n\n    wave_len = (waveform_orig)\n    waveform = np.concatenate([waveform_orig, waveform_orig, waveform_orig])\n\n    effective_length = sample_rate * period\n     (waveform) &lt; (period * sample_rate * ):\n        waveform = np.concatenate([waveform, waveform_orig])\n\n     start   :\n        start = start - (period - ) /  * sample_rate\n         start &lt; :\n            start += wave_len\n        start = (start)\n    :\n         wave_len &lt; effective_length:\n            start = np.random.randint(effective_length - wave_len)\n         wave_len &gt; effective_length:\n            start = np.random.randint(wave_len - effective_length)\n         wave_len == effective_length:\n            start = \n\n    waveform_seg = waveform[start: start + (effective_length)]\n\n     waveform_orig, waveform_seg, sample_rate, start\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2785050,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/30/2024 15:59:02",
          "content": "<p>I almost have a similar way to do the waveform extraction. It is in the next step (see my other comment). Anyway, thanks for sharing your code snippet, it might help others!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2784443,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "04/30/2024 09:27:14",
      "content": "<p>Why it will produce tensors of different sizes if the input size is consistent (5 secs in your example)? You can either crop or pad if the input shape is different for the audios, such as padding to 5 secs if the input is 3 or cropping if the input is 7 secs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2785048,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/30/2024 15:57:39",
          "content": "<p>Indeed, if I only do this (with missing padding for shorter samples): </p>\n<pre><code> torchaudio\nx, sample_rate = torchaudio.load(file_path)\n\nx = x[:, : * sample_rate]\n</code></pre>\n<p>the tensors can be stacked, no issue.</p>\n<p>However, if I add this step:</p>\n<pre><code> torchaudio.transforms  MelSpectrogram\nx = MelSpectrogram(**MEL_PARAMS)(x)\n</code></pre>\n<p>Then, I have the shape issue.</p>\n<p>I guess I will need to stack the tensors over the different <code>batch_size</code> and then do the transformation. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2785293,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "04/30/2024 17:51:31",
              "content": "<p>I see, then yeah there are multiple examples shared in the code section which can be useful further. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2784346": "I was wondering what was your approach to align transformed sound files?\n\nFor example, imagine I sample a random 5s section from an audio file and then\ndo a MEL spectrogram transformation. Each audio file will produce a tensor of different sizes: (1, p).\n\nHow would you align those? Or maybe you first stack the waveforms and then apply the MEL spectrogram transformation?\n\nPlease share your insights, thanks!",
    "2784377": "Are you trying to deal with different-length waves and turn them into Mel spectra with a consistent resolution? The following code can provide you with some references.\n\n```python\ndef load_wave_and_crop(filename, period, start=None):\n\n    waveform_orig, sample_rate = librosa.load(filename, sr=32000, mono=False)\n\n    wave_len = len(waveform_orig)\n    waveform = np.concatenate([waveform_orig, waveform_orig, waveform_orig])\n\n    effective_length = sample_rate * period\n    while len(waveform) < (period * sample_rate * 3):\n        waveform = np.concatenate([waveform, waveform_orig])\n    \n    if start is not None:\n        start = start - (period - 5) / 2 * sample_rate\n        while start < 0:\n            start += wave_len\n        start = int(start)\n    else:\n        if wave_len < effective_length:\n            start = np.random.randint(effective_length - wave_len)\n        elif wave_len > effective_length:\n            start = np.random.randint(wave_len - effective_length)\n        elif wave_len == effective_length:\n            start = 0\n\n    waveform_seg = waveform[start: start + int(effective_length)]\n\n    return waveform_orig, waveform_seg, sample_rate, start\n```",
    "2784443": "Why it will produce tensors of different sizes if the input size is consistent (5 secs in your example)? You can either crop or pad if the input shape is different for the audios, such as padding to 5 secs if the input is 3 or cropping if the input is 7 secs.",
    "2785048": "Indeed, if I only do this (with missing padding for shorter samples): \n\n```python\nimport torchaudio\nx, sample_rate = torchaudio.load(file_path)\n# TODO: Need to pad...\nx = x[:, 0:5 * sample_rate]\n```\n\nthe tensors can be stacked, no issue.\n\nHowever, if I add this step:\n\n```python\nfrom torchaudio.transforms import MelSpectrogram\nx = MelSpectrogram(**MEL_PARAMS)(x)\n```\n\nThen, I have the shape issue.\n\nI guess I will need to stack the tensors over the different `batch_size` and then do the transformation.",
    "2785050": "I almost have a similar way to do the waveform extraction. It is in the next step (see my other comment). Anyway, thanks for sharing your code snippet, it might help others!",
    "2785293": "I see, then yeah there are multiple examples shared in the code section which can be useful further."
  },
  "source": "meta"
}