{
  "id": 211585,
  "title": "How do I crop a spectrogram/MFCC",
  "url": "/competitions/rfcx-species-audio-detection/discussion/211585",
  "author_name": "",
  "post_date": "2021-01-15T17:59:02.682773300Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>After studying the data a bit I realised that a single audio file has multiple bird sounds as pointed out by <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> .<br>\nBut I am still confused as to how does one crop audio? Should I create an image of all the spectrograms and then use the image co-ordinates to make the crops? Can I do it directly on the spectrogram without converting the spectrogram to an image? If yes then how?</p>\n<p>Heres the bounding boxes I want to crop.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1604765%2F04415575fe06bdcc94041eed48acdf1f%2FScreenshot%20from%202021-01-15%2023-26-07.png?generation=1610733598342597&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1154573",
      "postDate": "01/15/2021 17:59:02",
      "content": "<p>After studying the data a bit I realised that a single audio file has multiple bird sounds as pointed out by <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> .<br>\nBut I am still confused as to how does one crop audio? Should I create an image of all the spectrograms and then use the image co-ordinates to make the crops? Can I do it directly on the spectrogram without converting the spectrogram to an image? If yes then how?</p>\n<p>Heres the bounding boxes I want to crop.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1604765%2F04415575fe06bdcc94041eed48acdf1f%2FScreenshot%20from%202021-01-15%2023-26-07.png?generation=1610733598342597&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "After studying the data a bit I realised that a single audio file has multiple bird sounds as pointed out by @kneroma .\nBut I am still confused as to how does one crop audio? Should I create an image of all the spectrograms and then use the image co-ordinates to make the crops? Can I do it directly on the spectrogram without converting the spectrogram to an image? If yes then how?\n\nHeres the bounding boxes I want to crop.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1604765%2F04415575fe06bdcc94041eed48acdf1f%2FScreenshot%20from%202021-01-15%2023-26-07.png?generation=1610733598342597&alt=media)",
      "votes": null
    },
    {
      "id": "1154651",
      "postDate": "01/15/2021 18:44:18",
      "content": "<p>The spectogram is a 2D numpy array that is easily subscriptable. For example, my code to isolate a section with a midpoint (in seconds) and a fixed width:</p>\n<pre><code>def sec2index(s, spec_len=5626):\n    return int(np.round(s*spec_len/60))\n\ndef make_slice(mid, width, spec_len):\n    end = mid + width\n    start = mid - width\n\n    if start &lt; 0:\n        end = end - start\n        start = 0\n\n    if end &gt; 60:\n        start = start - (end-60)\n        end = 60\n\n    start_idx = sec2index(start, spec_len)\n    end_idx = sec2index(end, spec_len)\n\n    return start_idx, end_idx\n\nmel_spec = librosa.feature.melspectrogram(audio_vec, n_fft=FFT, hop_length=HOP, sr=SR, fmin=F_MIN, fmax=F_MAX, power=POWER)\n\nspec_len = mel_spec.shape[1]\n\nstart, end = make_slice(t_mid, WINDOW_WIDTH, spec_len)\n\nmel_spec_slice = mel_spec[:, start:end]\n</code></pre>\n<p>I'm sure you can modify this code to work on the y-axis/frequency axis as well.</p>",
      "rawMarkdown": "The spectogram is a 2D numpy array that is easily subscriptable. For example, my code to isolate a section with a midpoint (in seconds) and a fixed width:\n\n```\ndef sec2index(s, spec_len=5626):\n    return int(np.round(s*spec_len/60))\n\ndef make_slice(mid, width, spec_len):\n    end = mid + width\n    start = mid - width\n    \n    if start < 0:\n        end = end - start\n        start = 0\n        \n    if end > 60:\n        start = start - (end-60)\n        end = 60\n        \n    start_idx = sec2index(start, spec_len)\n    end_idx = sec2index(end, spec_len)\n       \n    return start_idx, end_idx\n\nmel_spec = librosa.feature.melspectrogram(audio_vec, n_fft=FFT, hop_length=HOP, sr=SR, fmin=F_MIN, fmax=F_MAX, power=POWER)\n\nspec_len = mel_spec.shape[1]\n\nstart, end = make_slice(t_mid, WINDOW_WIDTH, spec_len)\n\nmel_spec_slice = mel_spec[:, start:end]\n```\n\nI'm sure you can modify this code to work on the y-axis/frequency axis as well.",
      "votes": null
    },
    {
      "id": "1156822",
      "postDate": "01/17/2021 12:37:43",
      "content": "<p>Thanks for comment <a href=\"https://www.kaggle.com/RNA\" target=\"_blank\">@RNA</a>, but for some reason the same logic doesnt really translate to the y axis :/</p>",
      "rawMarkdown": "Thanks for comment @RNA, but for some reason the same logic doesnt really translate to the y axis :/",
      "votes": null
    },
    {
      "id": "1156939",
      "postDate": "01/17/2021 14:38:01",
      "content": "<p>I solved it by limiting the librosa mel filter bank to the adequate freq limits, in case any one runs into the same problem</p>",
      "rawMarkdown": "I solved it by limiting the librosa mel filter bank to the adequate freq limits, in case any one runs into the same problem",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1154651,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "01/15/2021 18:44:18",
      "content": "<p>The spectogram is a 2D numpy array that is easily subscriptable. For example, my code to isolate a section with a midpoint (in seconds) and a fixed width:</p>\n<pre><code>def sec2index(s, spec_len=5626):\n    return int(np.round(s*spec_len/60))\n\ndef make_slice(mid, width, spec_len):\n    end = mid + width\n    start = mid - width\n\n    if start &lt; 0:\n        end = end - start\n        start = 0\n\n    if end &gt; 60:\n        start = start - (end-60)\n        end = 60\n\n    start_idx = sec2index(start, spec_len)\n    end_idx = sec2index(end, spec_len)\n\n    return start_idx, end_idx\n\nmel_spec = librosa.feature.melspectrogram(audio_vec, n_fft=FFT, hop_length=HOP, sr=SR, fmin=F_MIN, fmax=F_MAX, power=POWER)\n\nspec_len = mel_spec.shape[1]\n\nstart, end = make_slice(t_mid, WINDOW_WIDTH, spec_len)\n\nmel_spec_slice = mel_spec[:, start:end]\n</code></pre>\n<p>I'm sure you can modify this code to work on the y-axis/frequency axis as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156822,
      "author_name": "aryaman1999",
      "author_url": "",
      "post_date": "01/17/2021 12:37:43",
      "content": "<p>Thanks for comment <a href=\"https://www.kaggle.com/RNA\" target=\"_blank\">@RNA</a>, but for some reason the same logic doesnt really translate to the y axis :/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156939,
      "author_name": "aryaman1999",
      "author_url": "",
      "post_date": "01/17/2021 14:38:01",
      "content": "<p>I solved it by limiting the librosa mel filter bank to the adequate freq limits, in case any one runs into the same problem</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1154573": "After studying the data a bit I realised that a single audio file has multiple bird sounds as pointed out by @kneroma .\nBut I am still confused as to how does one crop audio? Should I create an image of all the spectrograms and then use the image co-ordinates to make the crops? Can I do it directly on the spectrogram without converting the spectrogram to an image? If yes then how?\n\nHeres the bounding boxes I want to crop.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1604765%2F04415575fe06bdcc94041eed48acdf1f%2FScreenshot%20from%202021-01-15%2023-26-07.png?generation=1610733598342597&alt=media)",
    "1154651": "The spectogram is a 2D numpy array that is easily subscriptable. For example, my code to isolate a section with a midpoint (in seconds) and a fixed width:\n\n```\ndef sec2index(s, spec_len=5626):\n    return int(np.round(s*spec_len/60))\n\ndef make_slice(mid, width, spec_len):\n    end = mid + width\n    start = mid - width\n    \n    if start < 0:\n        end = end - start\n        start = 0\n        \n    if end > 60:\n        start = start - (end-60)\n        end = 60\n        \n    start_idx = sec2index(start, spec_len)\n    end_idx = sec2index(end, spec_len)\n       \n    return start_idx, end_idx\n\nmel_spec = librosa.feature.melspectrogram(audio_vec, n_fft=FFT, hop_length=HOP, sr=SR, fmin=F_MIN, fmax=F_MAX, power=POWER)\n\nspec_len = mel_spec.shape[1]\n\nstart, end = make_slice(t_mid, WINDOW_WIDTH, spec_len)\n\nmel_spec_slice = mel_spec[:, start:end]\n```\n\nI'm sure you can modify this code to work on the y-axis/frequency axis as well.",
    "1156822": "Thanks for comment @RNA, but for some reason the same logic doesnt really translate to the y axis :/",
    "1156939": "I solved it by limiting the librosa mel filter bank to the adequate freq limits, in case any one runs into the same problem"
  },
  "source": "meta"
}