{
  "id": 176184,
  "title": "top_db maximum value (librosa.split)",
  "url": "/competitions/birdsong-recognition/discussion/176184",
  "author_name": "",
  "post_date": "2020-08-20T19:36:00.232797800Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'm thinking about remixing each training recording by cutting the silence and noise out of it (and adding it back in later as random augmentation).</p>\n<p><code>librosa.split</code> looks really handy, but I was wondering if anyone has tried it and figured out the best value for the <code>top_db</code> argument. I don't want to leave noise in, but I also don't want to accidentally cut out signal either. </p>\n<p>What do you guys think?</p>",
  "messages": [
    {
      "id": "979375",
      "postDate": "08/20/2020 19:36:00",
      "content": "<p>I'm thinking about remixing each training recording by cutting the silence and noise out of it (and adding it back in later as random augmentation).</p>\n<p><code>librosa.split</code> looks really handy, but I was wondering if anyone has tried it and figured out the best value for the <code>top_db</code> argument. I don't want to leave noise in, but I also don't want to accidentally cut out signal either. </p>\n<p>What do you guys think?</p>",
      "rawMarkdown": "I'm thinking about remixing each training recording by cutting the silence and noise out of it (and adding it back in later as random augmentation).\n\n`librosa.split` looks really handy, but I was wondering if anyone has tried it and figured out the best value for the `top_db` argument. I don't want to leave noise in, but I also don't want to accidentally cut out signal either. \n\nWhat do you guys think?",
      "votes": null
    },
    {
      "id": "979409",
      "postDate": "08/20/2020 20:08:39",
      "content": "<p>Here is the distribution of DB for <code>train_audio/</code> if it can help. In one kernel they use <code>top_db=mean_db-std_db</code> to split clip in birds/noise. I've tried it and it cuts around 30% of clips to noise.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Faeaf310e49d0aa11a37d1ab5eeee2b76%2Fdb.png?generation=1597953889412560&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here is the distribution of DB for `train_audio/` if it can help. In one kernel they use `top_db=mean_db-std_db` to split clip in birds/noise. I've tried it and it cuts around 30% of clips to noise.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Faeaf310e49d0aa11a37d1ab5eeee2b76%2Fdb.png?generation=1597953889412560&alt=media)",
      "votes": null
    },
    {
      "id": "979420",
      "postDate": "08/20/2020 20:34:04",
      "content": "<p>That's really handy. If I may ask, how did you get the decibel distribution? Did you calculate it from the audio itself or get it from an outside source? I did not see that information in the <code>train.csv</code> file. The <code>volume</code> column did not have the expected information. </p>",
      "rawMarkdown": "That's really handy. If I may ask, how did you get the decibel distribution? Did you calculate it from the audio itself or get it from an outside source? I did not see that information in the `train.csv` file. The `volume` column did not have the expected information.",
      "votes": null
    },
    {
      "id": "979460",
      "postDate": "08/20/2020 21:20:12",
      "content": "<p>Compute for each audio content.</p>",
      "rawMarkdown": "Compute for each audio content.",
      "votes": null
    },
    {
      "id": "979517",
      "postDate": "08/20/2020 22:57:04",
      "content": "<p>you can obtain it from melspectrogram for example or spectrogram for that reason</p>",
      "rawMarkdown": "you can obtain it from melspectrogram for example or spectrogram for that reason",
      "votes": null
    },
    {
      "id": "979522",
      "postDate": "08/20/2020 23:07:20",
      "content": "<p>Yes, that makes sense. Just grab it from the STFTs of the signal. I just thought I might have missed the data existing somewhere else already. Thank you both. </p>",
      "rawMarkdown": "Yes, that makes sense. Just grab it from the STFTs of the signal. I just thought I might have missed the data existing somewhere else already. Thank you both.",
      "votes": null
    },
    {
      "id": "980597",
      "postDate": "08/21/2020 17:56:49",
      "content": "<p>Also notice that using <code>librosa.split</code> with <code>top_db=mean_db-std_db</code> does not really split birds songs / noise, it just splits based on db, if you listen on noise parts, there are a lot birds still singing in background. You can consider it to generate a \"compressed signal\" (to have more chance to get high quality birds songs).</p>",
      "rawMarkdown": "Also notice that using `librosa.split` with `top_db=mean_db-std_db` does not really split birds songs / noise, it just splits based on db, if you listen on noise parts, there are a lot birds still singing in background. You can consider it to generate a \"compressed signal\" (to have more chance to get high quality birds songs).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 979409,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "08/20/2020 20:08:39",
      "content": "<p>Here is the distribution of DB for <code>train_audio/</code> if it can help. In one kernel they use <code>top_db=mean_db-std_db</code> to split clip in birds/noise. I've tried it and it cuts around 30% of clips to noise.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Faeaf310e49d0aa11a37d1ab5eeee2b76%2Fdb.png?generation=1597953889412560&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 979420,
          "author_name": "wesleyneill",
          "author_url": "",
          "post_date": "08/20/2020 20:34:04",
          "content": "<p>That's really handy. If I may ask, how did you get the decibel distribution? Did you calculate it from the audio itself or get it from an outside source? I did not see that information in the <code>train.csv</code> file. The <code>volume</code> column did not have the expected information. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979460,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "08/20/2020 21:20:12",
          "content": "<p>Compute for each audio content.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979517,
          "author_name": "leodav",
          "author_url": "",
          "post_date": "08/20/2020 22:57:04",
          "content": "<p>you can obtain it from melspectrogram for example or spectrogram for that reason</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979522,
          "author_name": "wesleyneill",
          "author_url": "",
          "post_date": "08/20/2020 23:07:20",
          "content": "<p>Yes, that makes sense. Just grab it from the STFTs of the signal. I just thought I might have missed the data existing somewhere else already. Thank you both. </p>",
          "votes": null,
          "replies": [
            {
              "id": 980597,
              "author_name": "mpware",
              "author_url": "",
              "post_date": "08/21/2020 17:56:49",
              "content": "<p>Also notice that using <code>librosa.split</code> with <code>top_db=mean_db-std_db</code> does not really split birds songs / noise, it just splits based on db, if you listen on noise parts, there are a lot birds still singing in background. You can consider it to generate a \"compressed signal\" (to have more chance to get high quality birds songs).</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "979375": "I'm thinking about remixing each training recording by cutting the silence and noise out of it (and adding it back in later as random augmentation).\n\n`librosa.split` looks really handy, but I was wondering if anyone has tried it and figured out the best value for the `top_db` argument. I don't want to leave noise in, but I also don't want to accidentally cut out signal either. \n\nWhat do you guys think?",
    "979409": "Here is the distribution of DB for `train_audio/` if it can help. In one kernel they use `top_db=mean_db-std_db` to split clip in birds/noise. I've tried it and it cuts around 30% of clips to noise.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Faeaf310e49d0aa11a37d1ab5eeee2b76%2Fdb.png?generation=1597953889412560&alt=media)",
    "979420": "That's really handy. If I may ask, how did you get the decibel distribution? Did you calculate it from the audio itself or get it from an outside source? I did not see that information in the `train.csv` file. The `volume` column did not have the expected information.",
    "979460": "Compute for each audio content.",
    "979517": "you can obtain it from melspectrogram for example or spectrogram for that reason",
    "979522": "Yes, that makes sense. Just grab it from the STFTs of the signal. I just thought I might have missed the data existing somewhere else already. Thank you both.",
    "980597": "Also notice that using `librosa.split` with `top_db=mean_db-std_db` does not really split birds songs / noise, it just splits based on db, if you listen on noise parts, there are a lot birds still singing in background. You can consider it to generate a \"compressed signal\" (to have more chance to get high quality birds songs)."
  },
  "source": "meta"
}