{
  "id": 91166,
  "title": "Looking for some background knowledge",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/91166",
  "author_name": "",
  "post_date": "2019-05-01T15:47:37.208377400Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I haven't worked with signal processing before, hence it is quite a new thing for me. Although I understand how models work/will work, I am more interested in some background knowledge. I was looking through the documentation of <code>librosa</code> but didn't find much. Here are some of my doubts:\n1. What is <code>hop_length</code>, <code>fmin</code>, <code>fmax</code>,  <code>n_mel</code>, and <code>n_fft</code> means? \n2. What is their importance and how the values affect the signal overall? Is there any general guideline for these values or just a random search?\n3. What are the disadvantages of using <code>mel_spectogram</code>?</p>",
  "messages": [
    {
      "id": "525721",
      "postDate": "05/01/2019 15:47:37",
      "content": "<p>I haven't worked with signal processing before, hence it is quite a new thing for me. Although I understand how models work/will work, I am more interested in some background knowledge. I was looking through the documentation of <code>librosa</code> but didn't find much. Here are some of my doubts:\n1. What is <code>hop_length</code>, <code>fmin</code>, <code>fmax</code>,  <code>n_mel</code>, and <code>n_fft</code> means? \n2. What is their importance and how the values affect the signal overall? Is there any general guideline for these values or just a random search?\n3. What are the disadvantages of using <code>mel_spectogram</code>?</p>",
      "rawMarkdown": "I haven't worked with signal processing before, hence it is quite a new thing for me. Although I understand how models work/will work, I am more interested in some background knowledge. I was looking through the documentation of `librosa` but didn't find much. Here are some of my doubts:\n1. What is `hop_length`, `fmin`, `fmax`,  `n_mel`, and `n_fft` means? \n2. What is their importance and how the values affect the signal overall? Is there any general guideline for these values or just a random search?\n3. What are the disadvantages of using `mel_spectogram`?",
      "votes": null
    },
    {
      "id": "525771",
      "postDate": "05/01/2019 17:24:51",
      "content": "<p>Here's a blog post from someone who took part in last year's Freesound Kaggle challenge that might be helpful\n<a href=\"https://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89\">https://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89</a></p>\n\n<p>This is also a nice interactive way to learn some basic signal processing without getting drowned in theory\n<a href=\"https://jackschaedler.github.io/circles-sines-signals/dft_introduction.html\">https://jackschaedler.github.io/circles-sines-signals/dft_introduction.html</a></p>\n\n<p>Mel spectrograms are a good general-purpose audio feature and you could use the hyperparameters we have used in the Challenge Baseline as a sane starting point (see Input Features section of <a href=\"https://github.com/DCASE-REPO/dcase2019_task2_baseline/blob/master/README.md\">https://github.com/DCASE-REPO/dcase2019_task2_baseline/blob/master/README.md</a>)</p>",
      "rawMarkdown": "Here's a blog post from someone who took part in last year's Freesound Kaggle challenge that might be helpful\nhttps://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89\n\nThis is also a nice interactive way to learn some basic signal processing without getting drowned in theory\nhttps://jackschaedler.github.io/circles-sines-signals/dft_introduction.html\n\nMel spectrograms are a good general-purpose audio feature and you could use the hyperparameters we have used in the Challenge Baseline as a sane starting point (see Input Features section of https://github.com/DCASE-REPO/dcase2019_task2_baseline/blob/master/README.md)",
      "votes": null
    },
    {
      "id": "525788",
      "postDate": "05/01/2019 18:03:24",
      "content": "<p>Raw audio is some digital representation of air pressure. It goes to microphone similar like to human ear.\nHowever, we can suspect (or prove?) that humans hearing is focused on frequencies. The different representation of raw audio samples is spectrogram which is a visualization of Short Time Fourier Transform, STFT, frequency representation.\nWe could calculate all the frequencies in all the audio at once using Fast Fourier Transform, FFT. But it will not show the dynamics of audio, the changes in it, just all the present frequencies. So we can split the audio signal into some smaller parts and then calculate frequencies in consecutive parts. This is STFT.</p>\n\n<p>The questions here are: how do we split audio? Into some windows. </p>\n\n<p>The length of that window is *n_fft* (in samples we can convert it to milliseconds. ~2200 samples, sampling freq 44100, is around 50 milliseconds. Do you like this value? For speech, it seems too long because some sounds last less than that. Think about the best value for sound classification). </p>\n\n<p>If we split raw audion in equal windows, we can miss some information that is present in the border. So we can overlap the windows. *hop_length* is the value in samples, how big is the \"jump\" between the consecutive windows.</p>\n\n<p>We don't use raw spectrogram. We instead filter the frequencies using filters designed to mimic human hearing. these are mel filters. You can choose how many of them to create, setting *n_mel*. The signal has a sampling rate 44100, so the frequencies in it are between 0 and 22050. Using 128 mels means (not really because filters are not equal in freq sizes) that you represent 22050/128=172Hz on average. So one feature is around 172Hz in space between 0 and 22015 Hz. Or, you can set <em>fmin</em> and <em>fmax</em> to set the frequencies between which you will create mel filters.</p>\n\n<p>Maybe you will like this\n<a href=\"https://www.kaggle.com/davids1992/speech-representation-and-data-exploration\">https://www.kaggle.com/davids1992/speech-representation-and-data-exploration</a>\n<a href=\"https://www.kaggle.com/davids1992/audio-representation-what-it-s-all-about\">https://www.kaggle.com/davids1992/audio-representation-what-it-s-all-about</a></p>",
      "rawMarkdown": "Raw audio is some digital representation of air pressure. It goes to microphone similar like to human ear.\nHowever, we can suspect (or prove?) that humans hearing is focused on frequencies. The different representation of raw audio samples is spectrogram which is a visualization of Short Time Fourier Transform, STFT, frequency representation.\nWe could calculate all the frequencies in all the audio at once using Fast Fourier Transform, FFT. But it will not show the dynamics of audio, the changes in it, just all the present frequencies. So we can split the audio signal into some smaller parts and then calculate frequencies in consecutive parts. This is STFT.\n\nThe questions here are: how do we split audio? Into some windows. \n\nThe length of that window is *n_fft* (in samples we can convert it to milliseconds. ~2200 samples, sampling freq 44100, is around 50 milliseconds. Do you like this value? For speech, it seems too long because some sounds last less than that. Think about the best value for sound classification). \n\nIf we split raw audion in equal windows, we can miss some information that is present in the border. So we can overlap the windows. *hop_length* is the value in samples, how big is the \"jump\" between the consecutive windows.\n\nWe don't use raw spectrogram. We instead filter the frequencies using filters designed to mimic human hearing. these are mel filters. You can choose how many of them to create, setting *n_mel*. The signal has a sampling rate 44100, so the frequencies in it are between 0 and 22050. Using 128 mels means (not really because filters are not equal in freq sizes) that you represent 22050/128=172Hz on average. So one feature is around 172Hz in space between 0 and 22015 Hz. Or, you can set *fmin* and *fmax* to set the frequencies between which you will create mel filters.\n\n\nMaybe you will like this\nhttps://www.kaggle.com/davids1992/speech-representation-and-data-exploration\nhttps://www.kaggle.com/davids1992/audio-representation-what-it-s-all-about",
      "votes": null
    },
    {
      "id": "525990",
      "postDate": "05/02/2019 06:04:08",
      "content": "<p>Thanks <a href=\"/davids1992\">@davids1992</a> for the detailed info. This is quiet good for the start. Though I have one question. If you are using <code>Conv1D</code>, then it is better to pass raw audio to the convolution as it can figure out what to look for and in this case the input shape for conv would be <code>(nframes, 1)</code> but if you want to pass <code>spectogram</code> to <code>Conv1D</code>, you can do it as well but that won't be much useful as we already restricted the amount of information and the input shape in this case would be <code>(features_dim_of_Spectogram, timesteps)</code>, right?</p>",
      "rawMarkdown": "Thanks @davids1992 for the detailed info. This is quiet good for the start. Though I have one question. If you are using `Conv1D`, then it is better to pass raw audio to the convolution as it can figure out what to look for and in this case the input shape for conv would be `(nframes, 1)` but if you want to pass `spectogram` to `Conv1D`, you can do it as well but that won't be much useful as we already restricted the amount of information and the input shape in this case would be `(features_dim_of_Spectogram, timesteps)`, right?",
      "votes": null
    },
    {
      "id": "525991",
      "postDate": "05/02/2019 06:05:16",
      "content": "<p>Thanks <a href=\"/plakal\">@plakal</a> . I will look into it. The only weird thing is that you are still using <code>slim</code>, which I personally don't like as I more of a <code>keras</code> person but I will try to convert the code correspondingly</p>",
      "rawMarkdown": "Thanks @plakal . I will look into it. The only weird thing is that you are still using `slim`, which I personally don't like as I more of a `keras` person but I will try to convert the code correspondingly",
      "votes": null
    },
    {
      "id": "526006",
      "postDate": "05/02/2019 06:53:56",
      "content": "<p>You don't have to use the code. My suggestion was to use the hyperparameter values for feature extraction (STFT window/hop size, number of mel bands, etc) which you can then use in your librosa or keras implementation. </p>",
      "rawMarkdown": "You don't have to use the code. My suggestion was to use the hyperparameter values for feature extraction (STFT window/hop size, number of mel bands, etc) which you can then use in your librosa or keras implementation.",
      "votes": null
    },
    {
      "id": "526569",
      "postDate": "05/03/2019 09:42:41",
      "content": "<p>I don't have big experience with Convolutional Neural Networks. I think <em>Conv1D</em> can be useful with spectrograms as well as Conv2D. 1D should filter a few consecutive frames, not frequencies. It makes sense when you think about what the features represent.\nI don't think it's feasible to work directly on raw audio. Notice, however, that I'm doing pretty bad on this competition :)</p>",
      "rawMarkdown": "I don't have big experience with Convolutional Neural Networks. I think *Conv1D* can be useful with spectrograms as well as Conv2D. 1D should filter a few consecutive frames, not frequencies. It makes sense when you think about what the features represent.\nI don't think it's feasible to work directly on raw audio. Notice, however, that I'm doing pretty bad on this competition :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 525771,
      "author_name": "plakal",
      "author_url": "",
      "post_date": "05/01/2019 17:24:51",
      "content": "<p>Here's a blog post from someone who took part in last year's Freesound Kaggle challenge that might be helpful\n<a href=\"https://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89\">https://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89</a></p>\n\n<p>This is also a nice interactive way to learn some basic signal processing without getting drowned in theory\n<a href=\"https://jackschaedler.github.io/circles-sines-signals/dft_introduction.html\">https://jackschaedler.github.io/circles-sines-signals/dft_introduction.html</a></p>\n\n<p>Mel spectrograms are a good general-purpose audio feature and you could use the hyperparameters we have used in the Challenge Baseline as a sane starting point (see Input Features section of <a href=\"https://github.com/DCASE-REPO/dcase2019_task2_baseline/blob/master/README.md\">https://github.com/DCASE-REPO/dcase2019_task2_baseline/blob/master/README.md</a>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 525991,
          "author_name": "aakashnain",
          "author_url": "",
          "post_date": "05/02/2019 06:05:16",
          "content": "<p>Thanks <a href=\"/plakal\">@plakal</a> . I will look into it. The only weird thing is that you are still using <code>slim</code>, which I personally don't like as I more of a <code>keras</code> person but I will try to convert the code correspondingly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526006,
          "author_name": "plakal",
          "author_url": "",
          "post_date": "05/02/2019 06:53:56",
          "content": "<p>You don't have to use the code. My suggestion was to use the hyperparameter values for feature extraction (STFT window/hop size, number of mel bands, etc) which you can then use in your librosa or keras implementation. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 525788,
      "author_name": "davids1992",
      "author_url": "",
      "post_date": "05/01/2019 18:03:24",
      "content": "<p>Raw audio is some digital representation of air pressure. It goes to microphone similar like to human ear.\nHowever, we can suspect (or prove?) that humans hearing is focused on frequencies. The different representation of raw audio samples is spectrogram which is a visualization of Short Time Fourier Transform, STFT, frequency representation.\nWe could calculate all the frequencies in all the audio at once using Fast Fourier Transform, FFT. But it will not show the dynamics of audio, the changes in it, just all the present frequencies. So we can split the audio signal into some smaller parts and then calculate frequencies in consecutive parts. This is STFT.</p>\n\n<p>The questions here are: how do we split audio? Into some windows. </p>\n\n<p>The length of that window is *n_fft* (in samples we can convert it to milliseconds. ~2200 samples, sampling freq 44100, is around 50 milliseconds. Do you like this value? For speech, it seems too long because some sounds last less than that. Think about the best value for sound classification). </p>\n\n<p>If we split raw audion in equal windows, we can miss some information that is present in the border. So we can overlap the windows. *hop_length* is the value in samples, how big is the \"jump\" between the consecutive windows.</p>\n\n<p>We don't use raw spectrogram. We instead filter the frequencies using filters designed to mimic human hearing. these are mel filters. You can choose how many of them to create, setting *n_mel*. The signal has a sampling rate 44100, so the frequencies in it are between 0 and 22050. Using 128 mels means (not really because filters are not equal in freq sizes) that you represent 22050/128=172Hz on average. So one feature is around 172Hz in space between 0 and 22015 Hz. Or, you can set <em>fmin</em> and <em>fmax</em> to set the frequencies between which you will create mel filters.</p>\n\n<p>Maybe you will like this\n<a href=\"https://www.kaggle.com/davids1992/speech-representation-and-data-exploration\">https://www.kaggle.com/davids1992/speech-representation-and-data-exploration</a>\n<a href=\"https://www.kaggle.com/davids1992/audio-representation-what-it-s-all-about\">https://www.kaggle.com/davids1992/audio-representation-what-it-s-all-about</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 525990,
          "author_name": "aakashnain",
          "author_url": "",
          "post_date": "05/02/2019 06:04:08",
          "content": "<p>Thanks <a href=\"/davids1992\">@davids1992</a> for the detailed info. This is quiet good for the start. Though I have one question. If you are using <code>Conv1D</code>, then it is better to pass raw audio to the convolution as it can figure out what to look for and in this case the input shape for conv would be <code>(nframes, 1)</code> but if you want to pass <code>spectogram</code> to <code>Conv1D</code>, you can do it as well but that won't be much useful as we already restricted the amount of information and the input shape in this case would be <code>(features_dim_of_Spectogram, timesteps)</code>, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526569,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "05/03/2019 09:42:41",
          "content": "<p>I don't have big experience with Convolutional Neural Networks. I think <em>Conv1D</em> can be useful with spectrograms as well as Conv2D. 1D should filter a few consecutive frames, not frequencies. It makes sense when you think about what the features represent.\nI don't think it's feasible to work directly on raw audio. Notice, however, that I'm doing pretty bad on this competition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "525721": "I haven't worked with signal processing before, hence it is quite a new thing for me. Although I understand how models work/will work, I am more interested in some background knowledge. I was looking through the documentation of `librosa` but didn't find much. Here are some of my doubts:\n1. What is `hop_length`, `fmin`, `fmax`,  `n_mel`, and `n_fft` means? \n2. What is their importance and how the values affect the signal overall? Is there any general guideline for these values or just a random search?\n3. What are the disadvantages of using `mel_spectogram`?",
    "525771": "Here's a blog post from someone who took part in last year's Freesound Kaggle challenge that might be helpful\nhttps://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89\n\nThis is also a nice interactive way to learn some basic signal processing without getting drowned in theory\nhttps://jackschaedler.github.io/circles-sines-signals/dft_introduction.html\n\nMel spectrograms are a good general-purpose audio feature and you could use the hyperparameters we have used in the Challenge Baseline as a sane starting point (see Input Features section of https://github.com/DCASE-REPO/dcase2019_task2_baseline/blob/master/README.md)",
    "525788": "Raw audio is some digital representation of air pressure. It goes to microphone similar like to human ear.\nHowever, we can suspect (or prove?) that humans hearing is focused on frequencies. The different representation of raw audio samples is spectrogram which is a visualization of Short Time Fourier Transform, STFT, frequency representation.\nWe could calculate all the frequencies in all the audio at once using Fast Fourier Transform, FFT. But it will not show the dynamics of audio, the changes in it, just all the present frequencies. So we can split the audio signal into some smaller parts and then calculate frequencies in consecutive parts. This is STFT.\n\nThe questions here are: how do we split audio? Into some windows. \n\nThe length of that window is *n_fft* (in samples we can convert it to milliseconds. ~2200 samples, sampling freq 44100, is around 50 milliseconds. Do you like this value? For speech, it seems too long because some sounds last less than that. Think about the best value for sound classification). \n\nIf we split raw audion in equal windows, we can miss some information that is present in the border. So we can overlap the windows. *hop_length* is the value in samples, how big is the \"jump\" between the consecutive windows.\n\nWe don't use raw spectrogram. We instead filter the frequencies using filters designed to mimic human hearing. these are mel filters. You can choose how many of them to create, setting *n_mel*. The signal has a sampling rate 44100, so the frequencies in it are between 0 and 22050. Using 128 mels means (not really because filters are not equal in freq sizes) that you represent 22050/128=172Hz on average. So one feature is around 172Hz in space between 0 and 22015 Hz. Or, you can set *fmin* and *fmax* to set the frequencies between which you will create mel filters.\n\n\nMaybe you will like this\nhttps://www.kaggle.com/davids1992/speech-representation-and-data-exploration\nhttps://www.kaggle.com/davids1992/audio-representation-what-it-s-all-about",
    "525990": "Thanks @davids1992 for the detailed info. This is quiet good for the start. Though I have one question. If you are using `Conv1D`, then it is better to pass raw audio to the convolution as it can figure out what to look for and in this case the input shape for conv would be `(nframes, 1)` but if you want to pass `spectogram` to `Conv1D`, you can do it as well but that won't be much useful as we already restricted the amount of information and the input shape in this case would be `(features_dim_of_Spectogram, timesteps)`, right?",
    "525991": "Thanks @plakal . I will look into it. The only weird thing is that you are still using `slim`, which I personally don't like as I more of a `keras` person but I will try to convert the code correspondingly",
    "526006": "You don't have to use the code. My suggestion was to use the hyperparameter values for feature extraction (STFT window/hop size, number of mel bands, etc) which you can then use in your librosa or keras implementation.",
    "526569": "I don't have big experience with Convolutional Neural Networks. I think *Conv1D* can be useful with spectrograms as well as Conv2D. 1D should filter a few consecutive frames, not frequencies. It makes sense when you think about what the features represent.\nI don't think it's feasible to work directly on raw audio. Notice, however, that I'm doing pretty bad on this competition :)"
  },
  "source": "meta"
}