{
  "id": 308497,
  "title": "Audio Preprocessing & Feature Engineering",
  "url": "/competitions/birdclef-2022/discussion/308497",
  "author_name": "Ravi Shah",
  "post_date": "2022-02-19T00:42:55.713000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/ravishah1/birdclef-2022-audio-feature-engineering-eda\" target=\"_blank\">https://www.kaggle.com/ravishah1/birdclef-2022-audio-feature-engineering-eda</a></p>\n<p>The notebook above demonstrates how to preprocess audio data using librosa. Hopefully this helps you get started with preprocessing so you can move onto modeling</p>\n<p>Here are the features it covers:</p>\n<p><strong>Amplitude signal</strong> <br>\nget the audio signal as a list of floating point values representing the amplitude of the signal at any given time. (Amplitude being the sound wave measured from its equilibrium position) </p>\n<p><strong>Fast-Fourier Transform</strong><br>\nused to create a power spectrum by computing the magnitude and frequency of a signal</p>\n<p><strong>Short-Time Fourier Transform</strong><br>\nused to create a spectrogram, then I convert the spectrogram from amplitude to decibels to calculate the log spectrogram</p>\n<p><strong>Mel Spectrogram</strong><br>\nvery popular image data that provides our models with sound information similar to what a human would perceive</p>\n<p><strong>Tempo</strong><br>\nthe speed which the sound is played as a tempo</p>\n<p><strong>Tempogram</strong><br>\na image representation of tempo information</p>\n<p><strong>mfcc delta</strong><br>\nMel-frequency cepstral coefficients at different deltas computed based on Savitsky-Golay filtering.</p>",
  "messages": [
    {
      "id": 1696604,
      "postDate": "2022-02-19T00:42:55.713Z",
      "content": "<p><a href=\"https://www.kaggle.com/ravishah1/birdclef-2022-audio-feature-engineering-eda\" target=\"_blank\">https://www.kaggle.com/ravishah1/birdclef-2022-audio-feature-engineering-eda</a></p>\n<p>The notebook above demonstrates how to preprocess audio data using librosa. Hopefully this helps you get started with preprocessing so you can move onto modeling</p>\n<p>Here are the features it covers:</p>\n<p><strong>Amplitude signal</strong> <br>\nget the audio signal as a list of floating point values representing the amplitude of the signal at any given time. (Amplitude being the sound wave measured from its equilibrium position) </p>\n<p><strong>Fast-Fourier Transform</strong><br>\nused to create a power spectrum by computing the magnitude and frequency of a signal</p>\n<p><strong>Short-Time Fourier Transform</strong><br>\nused to create a spectrogram, then I convert the spectrogram from amplitude to decibels to calculate the log spectrogram</p>\n<p><strong>Mel Spectrogram</strong><br>\nvery popular image data that provides our models with sound information similar to what a human would perceive</p>\n<p><strong>Tempo</strong><br>\nthe speed which the sound is played as a tempo</p>\n<p><strong>Tempogram</strong><br>\na image representation of tempo information</p>\n<p><strong>mfcc delta</strong><br>\nMel-frequency cepstral coefficients at different deltas computed based on Savitsky-Golay filtering.</p>",
      "rawMarkdown": "https://www.kaggle.com/ravishah1/birdclef-2022-audio-feature-engineering-eda\n\nThe notebook above demonstrates how to preprocess audio data using librosa. Hopefully this helps you get started with preprocessing so you can move onto modeling\n\nHere are the features it covers:\n\n**Amplitude signal** \nget the audio signal as a list of floating point values representing the amplitude of the signal at any given time. (Amplitude being the sound wave measured from its equilibrium position) \n\n**Fast-Fourier Transform**\nused to create a power spectrum by computing the magnitude and frequency of a signal\n\n**Short-Time Fourier Transform**\nused to create a spectrogram, then I convert the spectrogram from amplitude to decibels to calculate the log spectrogram\n\n**Mel Spectrogram**\nvery popular image data that provides our models with sound information similar to what a human would perceive\n\n**Tempo**\nthe speed which the sound is played as a tempo\n\n**Tempogram**\na image representation of tempo information\n\n**mfcc delta**\nMel-frequency cepstral coefficients at different deltas computed based on Savitsky-Golay filtering.",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1696604": "https://www.kaggle.com/ravishah1/birdclef-2022-audio-feature-engineering-eda\n\nThe notebook above demonstrates how to preprocess audio data using librosa. Hopefully this helps you get started with preprocessing so you can move onto modeling\n\nHere are the features it covers:\n\n**Amplitude signal** \nget the audio signal as a list of floating point values representing the amplitude of the signal at any given time. (Amplitude being the sound wave measured from its equilibrium position) \n\n**Fast-Fourier Transform**\nused to create a power spectrum by computing the magnitude and frequency of a signal\n\n**Short-Time Fourier Transform**\nused to create a spectrogram, then I convert the spectrogram from amplitude to decibels to calculate the log spectrogram\n\n**Mel Spectrogram**\nvery popular image data that provides our models with sound information similar to what a human would perceive\n\n**Tempo**\nthe speed which the sound is played as a tempo\n\n**Tempogram**\na image representation of tempo information\n\n**mfcc delta**\nMel-frequency cepstral coefficients at different deltas computed based on Savitsky-Golay filtering."
  }
}