{
  "id": 244065,
  "title": "How to Apply Audio-space Augmentation on (Mel)Spectrogram-space?",
  "url": "/competitions/birdclef-2021/discussion/244065",
  "author_name": "Alex Lau",
  "post_date": "2021-06-05T06:20:21.719000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Since this is the latest audio competition so I decide to ask the question here, in a hope that some participants may have come across the problem and found a solution:</p>\n<p><strong>Motivation</strong><br>\nComputing mel-spectrogram on the fly is too expensive in training (in my case my GPU utility is bounded by this step), so I would like to precompute mel-spectrogram of the original audio file instead. However, doing so make me give up the chance to apply augmentation on the original audio file. </p>\n<p><strong>Question</strong><br>\nMy questions are not meant to focus on augmentations usually applied on mel-spetrogram space, but the augmentations that usually applied on audio space (e.g. time stretching, pitch shifting, adding background noise). How could I apply these equivalent augmentations on mel-spectrogram space?  <br>\n<em>(Precomputing many melspectrogram of augmented audio files could be a work-around, but I am looking forward to a cleaner solution than this!)</em></p>",
  "messages": [
    {
      "id": 1336678,
      "postDate": "2021-06-05T06:20:21.720Z",
      "content": "<p>Since this is the latest audio competition so I decide to ask the question here, in a hope that some participants may have come across the problem and found a solution:</p>\n<p><strong>Motivation</strong><br>\nComputing mel-spectrogram on the fly is too expensive in training (in my case my GPU utility is bounded by this step), so I would like to precompute mel-spectrogram of the original audio file instead. However, doing so make me give up the chance to apply augmentation on the original audio file. </p>\n<p><strong>Question</strong><br>\nMy questions are not meant to focus on augmentations usually applied on mel-spetrogram space, but the augmentations that usually applied on audio space (e.g. time stretching, pitch shifting, adding background noise). How could I apply these equivalent augmentations on mel-spectrogram space?  <br>\n<em>(Precomputing many melspectrogram of augmented audio files could be a work-around, but I am looking forward to a cleaner solution than this!)</em></p>",
      "rawMarkdown": "Since this is the latest audio competition so I decide to ask the question here, in a hope that some participants may have come across the problem and found a solution:\n\n**Motivation**\nComputing mel-spectrogram on the fly is too expensive in training (in my case my GPU utility is bounded by this step), so I would like to precompute mel-spectrogram of the original audio file instead. However, doing so make me give up the chance to apply augmentation on the original audio file. \n\n**Question**\nMy questions are not meant to focus on augmentations usually applied on mel-spetrogram space, but the augmentations that usually applied on audio space (e.g. time stretching, pitch shifting, adding background noise). How could I apply these equivalent augmentations on mel-spectrogram space?  \n_(Precomputing many melspectrogram of augmented audio files could be a work-around, but I am looking forward to a cleaner solution than this!)_",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1336678": "Since this is the latest audio competition so I decide to ask the question here, in a hope that some participants may have come across the problem and found a solution:\n\n**Motivation**\nComputing mel-spectrogram on the fly is too expensive in training (in my case my GPU utility is bounded by this step), so I would like to precompute mel-spectrogram of the original audio file instead. However, doing so make me give up the chance to apply augmentation on the original audio file. \n\n**Question**\nMy questions are not meant to focus on augmentations usually applied on mel-spetrogram space, but the augmentations that usually applied on audio space (e.g. time stretching, pitch shifting, adding background noise). How could I apply these equivalent augmentations on mel-spectrogram space?  \n_(Precomputing many melspectrogram of augmented audio files could be a work-around, but I am looking forward to a cleaner solution than this!)_"
  }
}