{
  "id": 212638,
  "title": "Best practices for creating spectrograms?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/212638",
  "author_name": "",
  "post_date": "2021-01-19T17:05:51.910525900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>What spectrogram parameters and augmentations have you guys found to work best?</p>\n<p>Currently I've been using using:<br>\n<code># create Mel spectrogram</code><br>\n<code>stft = librosa.feature.melspectrogram(loaded_audio, sr=sample_rate, power=power,\n                                          fmin=f_min, fmax=f_max, n_mels=shape[0])</code></p>\n<p><code>stft_to_db = librosa.core.amplitude_to_db(np.abs(stft))</code></p>\n<p>With </p>\n<p><code>shape = (256,512)</code> <br>\n<code>sample_rate = 48000</code> which is the native sampling rate of the audio file<br>\n<code>power=1.5</code><br>\n<code>f_min = round(df_in.f_min.min() * .75, 1)</code> -&gt; min freq reduced by 25%<br>\n<code>f_max = round(df_in.f_max.max() * 1.25, 1)</code> -&gt; max freq increased by 25%<br>\n<code>n_mels = 256</code> (x length of image size)</p>\n<p>Augmentations:</p>\n<ul>\n<li>Random offset of event plus or minus 3 seconds</li>\n<li>Adding random gaussian noise to spectrogram</li>\n<li>Random time masking</li>\n<li>Random frequency masking</li>\n</ul>\n<p>Example Spectrogram:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3436076%2Ffbe3decc95af953c4073fff60cec1b7b%2Ftest.png?generation=1611075880155421&amp;alt=media\" alt=\"\"></p>\n<p>Full code: <a href=\"https://www.kaggle.com/stegzz/readable-keras-resnet50-model\" target=\"_blank\">https://www.kaggle.com/stegzz/readable-keras-resnet50-model</a></p>",
  "messages": [
    {
      "id": "1160091",
      "postDate": "01/19/2021 17:05:51",
      "content": "<p>What spectrogram parameters and augmentations have you guys found to work best?</p>\n<p>Currently I've been using using:<br>\n<code># create Mel spectrogram</code><br>\n<code>stft = librosa.feature.melspectrogram(loaded_audio, sr=sample_rate, power=power,\n                                          fmin=f_min, fmax=f_max, n_mels=shape[0])</code></p>\n<p><code>stft_to_db = librosa.core.amplitude_to_db(np.abs(stft))</code></p>\n<p>With </p>\n<p><code>shape = (256,512)</code> <br>\n<code>sample_rate = 48000</code> which is the native sampling rate of the audio file<br>\n<code>power=1.5</code><br>\n<code>f_min = round(df_in.f_min.min() * .75, 1)</code> -&gt; min freq reduced by 25%<br>\n<code>f_max = round(df_in.f_max.max() * 1.25, 1)</code> -&gt; max freq increased by 25%<br>\n<code>n_mels = 256</code> (x length of image size)</p>\n<p>Augmentations:</p>\n<ul>\n<li>Random offset of event plus or minus 3 seconds</li>\n<li>Adding random gaussian noise to spectrogram</li>\n<li>Random time masking</li>\n<li>Random frequency masking</li>\n</ul>\n<p>Example Spectrogram:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3436076%2Ffbe3decc95af953c4073fff60cec1b7b%2Ftest.png?generation=1611075880155421&amp;alt=media\" alt=\"\"></p>\n<p>Full code: <a href=\"https://www.kaggle.com/stegzz/readable-keras-resnet50-model\" target=\"_blank\">https://www.kaggle.com/stegzz/readable-keras-resnet50-model</a></p>",
      "rawMarkdown": "What spectrogram parameters and augmentations have you guys found to work best?\n\nCurrently I've been using using:\n`# create Mel spectrogram`\n`stft = librosa.feature.melspectrogram(loaded_audio, sr=sample_rate, power=power,\n                                          fmin=f_min, fmax=f_max, n_mels=shape[0])`\n\n`stft_to_db = librosa.core.amplitude_to_db(np.abs(stft))`\n\n\nWith \n\n`shape = (256,512)` \n`sample_rate = 48000` which is the native sampling rate of the audio file\n`power=1.5`\n`f_min = round(df_in.f_min.min() * .75, 1)` -> min freq reduced by 25%\n`f_max = round(df_in.f_max.max() * 1.25, 1)` -> max freq increased by 25%\n`n_mels = 256` (x length of image size)\n\nAugmentations:\n- Random offset of event plus or minus 3 seconds\n- Adding random gaussian noise to spectrogram\n- Random time masking\n- Random frequency masking\n\nExample Spectrogram:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3436076%2Ffbe3decc95af953c4073fff60cec1b7b%2Ftest.png?generation=1611075880155421&alt=media)\n\nFull code: https://www.kaggle.com/stegzz/readable-keras-resnet50-model",
      "votes": null
    },
    {
      "id": "1160140",
      "postDate": "01/19/2021 17:31:53",
      "content": "<p>Have you tried adding mixup?</p>",
      "rawMarkdown": "Have you tried adding mixup?",
      "votes": null
    },
    {
      "id": "1160178",
      "postDate": "01/19/2021 18:06:32",
      "content": "<p>There is only one way to know: try various ways and see which settings lead to best model (better CV and LB).  </p>",
      "rawMarkdown": "There is only one way to know: try various ways and see which settings lead to best model (better CV and LB).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1160140,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "01/19/2021 17:31:53",
      "content": "<p>Have you tried adding mixup?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1160178,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/19/2021 18:06:32",
      "content": "<p>There is only one way to know: try various ways and see which settings lead to best model (better CV and LB).  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1160091": "What spectrogram parameters and augmentations have you guys found to work best?\n\nCurrently I've been using using:\n`# create Mel spectrogram`\n`stft = librosa.feature.melspectrogram(loaded_audio, sr=sample_rate, power=power,\n                                          fmin=f_min, fmax=f_max, n_mels=shape[0])`\n\n`stft_to_db = librosa.core.amplitude_to_db(np.abs(stft))`\n\n\nWith \n\n`shape = (256,512)` \n`sample_rate = 48000` which is the native sampling rate of the audio file\n`power=1.5`\n`f_min = round(df_in.f_min.min() * .75, 1)` -> min freq reduced by 25%\n`f_max = round(df_in.f_max.max() * 1.25, 1)` -> max freq increased by 25%\n`n_mels = 256` (x length of image size)\n\nAugmentations:\n- Random offset of event plus or minus 3 seconds\n- Adding random gaussian noise to spectrogram\n- Random time masking\n- Random frequency masking\n\nExample Spectrogram:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3436076%2Ffbe3decc95af953c4073fff60cec1b7b%2Ftest.png?generation=1611075880155421&alt=media)\n\nFull code: https://www.kaggle.com/stegzz/readable-keras-resnet50-model",
    "1160140": "Have you tried adding mixup?",
    "1160178": "There is only one way to know: try various ways and see which settings lead to best model (better CV and LB)."
  },
  "source": "meta"
}