{
  "id": 158908,
  "title": "Audio Pre-processing Techniques",
  "url": "/competitions/birdsong-recognition/discussion/158908",
  "author_name": "",
  "post_date": "2020-06-15T18:52:41.186702900Z",
  "votes": 42,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi All! </p>\n\n<p>While playing back a few audio files, I've realized, there is a huge amount of variance in the sound levels and there is some background wind/external noise. </p>\n\n<p>I'll keep updating this post as I find more techniques for pre-processing the audio, that might be helpful:</p>\n\n<ul>\n<li>SpecMix: A Simple Data Augmentation and Warm-up Pipeline to Leverage Clean and Noisy Set for Efficient Audio Tagging by @ebouteillon <a href=\"http://dcase.community/documents/challenge2019/technical_reports/DCASE2019_Bouteillon_27_t2.pdf\">Paper</a>, <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">Code</a></li>\n<li>A few data augmentation techniques using fastai audio <a href=\"https://github.com/zcaceres/fastai-audio/blob/master/DataAugmentation.ipynb\">Github</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=QEEBNF0aeeg\">Talk on audio classification</a></li>\n<li><a href=\"https://librosa.github.io/librosa/index.html\">Librosa package for music and audio pre-processing</a></li>\n</ul>\n\n<p>Good Luck!</p>",
  "messages": [
    {
      "id": "887591",
      "postDate": "06/15/2020 18:52:41",
      "content": "<p>Hi All! </p>\n\n<p>While playing back a few audio files, I've realized, there is a huge amount of variance in the sound levels and there is some background wind/external noise. </p>\n\n<p>I'll keep updating this post as I find more techniques for pre-processing the audio, that might be helpful:</p>\n\n<ul>\n<li>SpecMix: A Simple Data Augmentation and Warm-up Pipeline to Leverage Clean and Noisy Set for Efficient Audio Tagging by @ebouteillon <a href=\"http://dcase.community/documents/challenge2019/technical_reports/DCASE2019_Bouteillon_27_t2.pdf\">Paper</a>, <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">Code</a></li>\n<li>A few data augmentation techniques using fastai audio <a href=\"https://github.com/zcaceres/fastai-audio/blob/master/DataAugmentation.ipynb\">Github</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=QEEBNF0aeeg\">Talk on audio classification</a></li>\n<li><a href=\"https://librosa.github.io/librosa/index.html\">Librosa package for music and audio pre-processing</a></li>\n</ul>\n\n<p>Good Luck!</p>",
      "rawMarkdown": "Hi All! \n\nWhile playing back a few audio files, I've realized, there is a huge amount of variance in the sound levels and there is some background wind/external noise. \n\nI'll keep updating this post as I find more techniques for pre-processing the audio, that might be helpful:\n\n- SpecMix: A Simple Data Augmentation and Warm-up Pipeline to Leverage Clean and Noisy Set for Efficient Audio Tagging by @ebouteillon [Paper](http://dcase.community/documents/challenge2019/technical_reports/DCASE2019_Bouteillon_27_t2.pdf), [Code](https://github.com/ebouteillon/freesound-audio-tagging-2019)\n- A few data augmentation techniques using fastai audio [Github](https://github.com/zcaceres/fastai-audio/blob/master/DataAugmentation.ipynb)\n- [Talk on audio classification](https://www.youtube.com/watch?v=QEEBNF0aeeg)\n- [Librosa package for music and audio pre-processing](https://librosa.github.io/librosa/index.html)\n\nGood Luck!",
      "votes": null
    },
    {
      "id": "888037",
      "postDate": "06/16/2020 04:39:53",
      "content": "<p>How about Fourier transforms?</p>",
      "rawMarkdown": "How about Fourier transforms?",
      "votes": null
    },
    {
      "id": "890786",
      "postDate": "06/17/2020 17:42:33",
      "content": "<p>The main use-case / purpose of this competition is modelling for the domain mismatch between training and test data.\nThe train-audio files are clean recordings, while the test set contains noisy recordings made in the wild. I think using augmentation techniques that introduce background noise in the training data, as well as overlapping vocalizations, are going to be very important.</p>",
      "rawMarkdown": "The main use-case / purpose of this competition is modelling for the domain mismatch between training and test data.\nThe train-audio files are clean recordings, while the test set contains noisy recordings made in the wild. I think using augmentation techniques that introduce background noise in the training data, as well as overlapping vocalizations, are going to be very important.",
      "votes": null
    },
    {
      "id": "893987",
      "postDate": "06/20/2020 05:21:14",
      "content": "<p>nice info </p>",
      "rawMarkdown": "nice info",
      "votes": null
    },
    {
      "id": "954353",
      "postDate": "08/01/2020 16:49:24",
      "content": "<p>MFCCs could also be an approach I feel </p>",
      "rawMarkdown": "MFCCs could also be an approach I feel",
      "votes": null
    },
    {
      "id": "968521",
      "postDate": "08/13/2020 04:50:27",
      "content": "<p>Very helpful. <br>\n😊</p>",
      "rawMarkdown": "Very helpful. \n😊",
      "votes": null
    },
    {
      "id": "971038",
      "postDate": "08/15/2020 05:55:59",
      "content": "<p>Thanks for this :)</p>",
      "rawMarkdown": "Thanks for this :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 888037,
      "author_name": "nxrprime",
      "author_url": "",
      "post_date": "06/16/2020 04:39:53",
      "content": "<p>How about Fourier transforms?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 890786,
      "author_name": "dhruvrnaik",
      "author_url": "",
      "post_date": "06/17/2020 17:42:33",
      "content": "<p>The main use-case / purpose of this competition is modelling for the domain mismatch between training and test data.\nThe train-audio files are clean recordings, while the test set contains noisy recordings made in the wild. I think using augmentation techniques that introduce background noise in the training data, as well as overlapping vocalizations, are going to be very important.</p>",
      "votes": null,
      "replies": [
        {
          "id": 968521,
          "author_name": "palaksood97",
          "author_url": "",
          "post_date": "08/13/2020 04:50:27",
          "content": "<p>Very helpful. <br>\n😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 893987,
      "author_name": "muralidhar123",
      "author_url": "",
      "post_date": "06/20/2020 05:21:14",
      "content": "<p>nice info </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 954353,
      "author_name": "digvijayyadav",
      "author_url": "",
      "post_date": "08/01/2020 16:49:24",
      "content": "<p>MFCCs could also be an approach I feel </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971038,
      "author_name": "raoofnaushad",
      "author_url": "",
      "post_date": "08/15/2020 05:55:59",
      "content": "<p>Thanks for this :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "887591": "Hi All! \n\nWhile playing back a few audio files, I've realized, there is a huge amount of variance in the sound levels and there is some background wind/external noise. \n\nI'll keep updating this post as I find more techniques for pre-processing the audio, that might be helpful:\n\n- SpecMix: A Simple Data Augmentation and Warm-up Pipeline to Leverage Clean and Noisy Set for Efficient Audio Tagging by @ebouteillon [Paper](http://dcase.community/documents/challenge2019/technical_reports/DCASE2019_Bouteillon_27_t2.pdf), [Code](https://github.com/ebouteillon/freesound-audio-tagging-2019)\n- A few data augmentation techniques using fastai audio [Github](https://github.com/zcaceres/fastai-audio/blob/master/DataAugmentation.ipynb)\n- [Talk on audio classification](https://www.youtube.com/watch?v=QEEBNF0aeeg)\n- [Librosa package for music and audio pre-processing](https://librosa.github.io/librosa/index.html)\n\nGood Luck!",
    "888037": "How about Fourier transforms?",
    "890786": "The main use-case / purpose of this competition is modelling for the domain mismatch between training and test data.\nThe train-audio files are clean recordings, while the test set contains noisy recordings made in the wild. I think using augmentation techniques that introduce background noise in the training data, as well as overlapping vocalizations, are going to be very important.",
    "893987": "nice info",
    "954353": "MFCCs could also be an approach I feel",
    "968521": "Very helpful. \n😊",
    "971038": "Thanks for this :)"
  },
  "source": "meta"
}