{
  "id": 197827,
  "title": "Let's talk about some fundamentals before modelling",
  "url": "/competitions/rfcx-species-audio-detection/discussion/197827",
  "author_name": "NAIN",
  "post_date": "2020-11-18T08:17:06.320000",
  "votes": 18,
  "comment_count": 8,
  "views": 0,
  "content": "<p>There have been many audio competitions in the past as well but I didn't get a chance to get fully involved in them. Although Computer Vision is my main focus, speech is definitely something that takes the second spot on my interest list. Given that I am learning about audio, I have certain questions related to the fundamentals of speech processing, which I am sure many Kagglers would be aware of. That's one of the reasons I am posting these questions here. </p>\n<ol>\n<li><p>Let's say we have a 5 second audio clip. Spectrograms and MelSpectrograms are fairly common ways of processing the audio. But how do you decide the number of <code>timesteps</code> a spectrogram should represent? How much of the audio signal does a single timestep in a sepctrogram should cover? </p></li>\n<li><p>I know that things like <code>nfft</code>, <code>hop_length</code>, <code>n_mels</code> are hyperparameters but what's the intuition behind increasing/decreasing the values of these hyperparameters? For example, why would I want to increase/decrease my <code>hop_length</code>?</p></li>\n</ol>",
  "messages": [
    {
      "id": 1082782,
      "postDate": "2020-11-18T08:17:06.320Z",
      "content": "<p>There have been many audio competitions in the past as well but I didn't get a chance to get fully involved in them. Although Computer Vision is my main focus, speech is definitely something that takes the second spot on my interest list. Given that I am learning about audio, I have certain questions related to the fundamentals of speech processing, which I am sure many Kagglers would be aware of. That's one of the reasons I am posting these questions here. </p>\n<ol>\n<li><p>Let's say we have a 5 second audio clip. Spectrograms and MelSpectrograms are fairly common ways of processing the audio. But how do you decide the number of <code>timesteps</code> a spectrogram should represent? How much of the audio signal does a single timestep in a sepctrogram should cover? </p></li>\n<li><p>I know that things like <code>nfft</code>, <code>hop_length</code>, <code>n_mels</code> are hyperparameters but what's the intuition behind increasing/decreasing the values of these hyperparameters? For example, why would I want to increase/decrease my <code>hop_length</code>?</p></li>\n</ol>",
      "rawMarkdown": "There have been many audio competitions in the past as well but I didn't get a chance to get fully involved in them. Although Computer Vision is my main focus, speech is definitely something that takes the second spot on my interest list. Given that I am learning about audio, I have certain questions related to the fundamentals of speech processing, which I am sure many Kagglers would be aware of. That's one of the reasons I am posting these questions here. \n\n1. Let's say we have a 5 second audio clip. Spectrograms and MelSpectrograms are fairly common ways of processing the audio. But how do you decide the number of `timesteps` a spectrogram should represent? How much of the audio signal does a single timestep in a sepctrogram should cover? \n\n2. I know that things like `nfft`, `hop_length`, `n_mels` are hyperparameters but what's the intuition behind increasing/decreasing the values of these hyperparameters? For example, why would I want to increase/decrease my `hop_length`?\n ",
      "votes": 18
    },
    {
      "id": 1095981,
      "postDate": "2020-11-30T05:44:27.730Z",
      "content": "<p>Here's a response to your first question:</p>\n<p>Timesteps are important when looking at the overall **pattern ** in the sound. There's a really good <a href=\"https://research.atspotify.com/making-sense-of-music-by-extracting-and-analyzing-individual-instruments-in-a-song/\" target=\"_blank\">Spotify Research</a> piece which mentions the importance of time steps when looking at the sound of vocals versus bass guitar.</p>\n<p>Vocals typically move fast in songs, since singers change their pitch, sing many words, etc. In this case, you'd want your timesteps fine-grained to sufficiently capture what's changing in the vocal track.</p>\n<p>On the other hand, a bass guitar typically plays slow, long notes. You don't want to sample too fast, because not much would be changing from one timestep to another. Imagine taking a picture, then streeeeeeetching out the width. With something like a CNN, your algorithm relies on a neighborhood of pixels. By including many more width pixels than required, the CNN won't find solutions as easily since the X direction doesn't offer much differentiation. It's the same with audio.</p>\n<p>Essentially, you want to pick a timestep appropriate to see the pattern in the sound, without going overboard. If you're not noticing much difference along the Time axis in your spectrogram, maybe reduce your timestep.</p>",
      "rawMarkdown": "Here's a response to your first question:\n\nTimesteps are important when looking at the overall **pattern ** in the sound. There's a really good [Spotify Research](https://research.atspotify.com/making-sense-of-music-by-extracting-and-analyzing-individual-instruments-in-a-song/) piece which mentions the importance of time steps when looking at the sound of vocals versus bass guitar.\n\nVocals typically move fast in songs, since singers change their pitch, sing many words, etc. In this case, you'd want your timesteps fine-grained to sufficiently capture what's changing in the vocal track.\n\nOn the other hand, a bass guitar typically plays slow, long notes. You don't want to sample too fast, because not much would be changing from one timestep to another. Imagine taking a picture, then streeeeeeetching out the width. With something like a CNN, your algorithm relies on a neighborhood of pixels. By including many more width pixels than required, the CNN won't find solutions as easily since the X direction doesn't offer much differentiation. It's the same with audio.\n\nEssentially, you want to pick a timestep appropriate to see the pattern in the sound, without going overboard. If you're not noticing much difference along the Time axis in your spectrogram, maybe reduce your timestep.",
      "votes": 11,
      "replies": [
        {
          "id": 1097625,
          "postDate": "2020-12-01T06:20:26.433Z",
          "content": "<p>Thank you. This makes sense. And thanks for the spotify research piece. </p>",
          "rawMarkdown": "Thank you. This makes sense. And thanks for the spotify research piece. "
        }
      ]
    },
    {
      "id": 1141094,
      "postDate": "2021-01-06T13:50:06.127Z",
      "content": "<p>One way is to think of SFFT as convolution.</p>\n<p>if time is the x axis and frequency the y axis then:</p>\n<p>nftt is the proportional to the width of your output, n_mels is proportional to the height of the output, and hop_length is stride on x axis.</p>",
      "rawMarkdown": "One way is to think of SFFT as convolution.\n\nif time is the x axis and frequency the y axis then:\n\nnftt is the proportional to the width of your output, n_mels is proportional to the height of the output, and hop_length is stride on x axis.",
      "votes": 6,
      "replies": [
        {
          "id": 1166311,
          "postDate": "2021-01-23T14:32:10.653Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>",
          "rawMarkdown": "Thanks @cpmpml ",
          "votes": 1
        },
        {
          "id": 1167561,
          "postDate": "2021-01-24T11:03:19.113Z",
          "content": "<p>No problem.  Will you enter this competition?</p>",
          "rawMarkdown": "No problem.  Will you enter this competition?"
        }
      ]
    },
    {
      "id": 1168052,
      "postDate": "2021-01-24T16:58:34.013Z",
      "content": "<p><a href=\"https://www.kaggle.com/aakashnain\" target=\"_blank\">@aakashnain</a>  when I was writing my notebook for a <a href=\"https://www.kaggle.com/dimitreoliveira/rainforest-audio-classification-tensorflow-starter\" target=\"_blank\">beginner on TF for audio</a>, I came across <a href=\"https://www.coursera.org/lecture/audio-signal-processing/stft-2-tjEQe\" target=\"_blank\">this video</a>, I am no expert in the field, but it seems to be a good explanation about STFT, I am not sure yet if I will get the chance to really put my hands in this competition, but I hope so.</p>",
      "rawMarkdown": "@aakashnain  when I was writing my notebook for a [beginner on TF for audio](https://www.kaggle.com/dimitreoliveira/rainforest-audio-classification-tensorflow-starter), I came across [this video](https://www.coursera.org/lecture/audio-signal-processing/stft-2-tjEQe), I am no expert in the field, but it seems to be a good explanation about STFT, I am not sure yet if I will get the chance to really put my hands in this competition, but I hope so.",
      "votes": 2,
      "replies": [
        {
          "id": 1197963,
          "postDate": "2021-02-12T15:09:33.803Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> </p>",
          "rawMarkdown": "Thanks @dimitreoliveira ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1170371,
      "postDate": "2021-01-26T07:08:59.343Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1095981,
      "author_name": "Mike A",
      "author_url": "",
      "post_date": "2020-11-30T05:44:27.730000",
      "content": "<p>Here's a response to your first question:</p>\n<p>Timesteps are important when looking at the overall **pattern ** in the sound. There's a really good <a href=\"https://research.atspotify.com/making-sense-of-music-by-extracting-and-analyzing-individual-instruments-in-a-song/\" target=\"_blank\">Spotify Research</a> piece which mentions the importance of time steps when looking at the sound of vocals versus bass guitar.</p>\n<p>Vocals typically move fast in songs, since singers change their pitch, sing many words, etc. In this case, you'd want your timesteps fine-grained to sufficiently capture what's changing in the vocal track.</p>\n<p>On the other hand, a bass guitar typically plays slow, long notes. You don't want to sample too fast, because not much would be changing from one timestep to another. Imagine taking a picture, then streeeeeeetching out the width. With something like a CNN, your algorithm relies on a neighborhood of pixels. By including many more width pixels than required, the CNN won't find solutions as easily since the X direction doesn't offer much differentiation. It's the same with audio.</p>\n<p>Essentially, you want to pick a timestep appropriate to see the pattern in the sound, without going overboard. If you're not noticing much difference along the Time axis in your spectrogram, maybe reduce your timestep.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 1097625,
          "author_name": "NAIN",
          "author_url": "",
          "post_date": "2020-12-01T06:20:26.433000",
          "content": "<p>Thank you. This makes sense. And thanks for the spotify research piece. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1141094,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-01-06T13:50:06.127000",
      "content": "<p>One way is to think of SFFT as convolution.</p>\n<p>if time is the x axis and frequency the y axis then:</p>\n<p>nftt is the proportional to the width of your output, n_mels is proportional to the height of the output, and hop_length is stride on x axis.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1166311,
          "author_name": "NAIN",
          "author_url": "",
          "post_date": "2021-01-23T14:32:10.653000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1167561,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-01-24T11:03:19.113000",
          "content": "<p>No problem.  Will you enter this competition?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1168052,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2021-01-24T16:58:34.013000",
      "content": "<p><a href=\"https://www.kaggle.com/aakashnain\" target=\"_blank\">@aakashnain</a>  when I was writing my notebook for a <a href=\"https://www.kaggle.com/dimitreoliveira/rainforest-audio-classification-tensorflow-starter\" target=\"_blank\">beginner on TF for audio</a>, I came across <a href=\"https://www.coursera.org/lecture/audio-signal-processing/stft-2-tjEQe\" target=\"_blank\">this video</a>, I am no expert in the field, but it seems to be a good explanation about STFT, I am not sure yet if I will get the chance to really put my hands in this competition, but I hope so.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1197963,
          "author_name": "NAIN",
          "author_url": "",
          "post_date": "2021-02-12T15:09:33.803000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1170371,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-26T07:08:59.343000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1082782": "There have been many audio competitions in the past as well but I didn't get a chance to get fully involved in them. Although Computer Vision is my main focus, speech is definitely something that takes the second spot on my interest list. Given that I am learning about audio, I have certain questions related to the fundamentals of speech processing, which I am sure many Kagglers would be aware of. That's one of the reasons I am posting these questions here. \n\n1. Let's say we have a 5 second audio clip. Spectrograms and MelSpectrograms are fairly common ways of processing the audio. But how do you decide the number of `timesteps` a spectrogram should represent? How much of the audio signal does a single timestep in a sepctrogram should cover? \n\n2. I know that things like `nfft`, `hop_length`, `n_mels` are hyperparameters but what's the intuition behind increasing/decreasing the values of these hyperparameters? For example, why would I want to increase/decrease my `hop_length`?\n ",
    "1095981": "Here's a response to your first question:\n\nTimesteps are important when looking at the overall **pattern ** in the sound. There's a really good [Spotify Research](https://research.atspotify.com/making-sense-of-music-by-extracting-and-analyzing-individual-instruments-in-a-song/) piece which mentions the importance of time steps when looking at the sound of vocals versus bass guitar.\n\nVocals typically move fast in songs, since singers change their pitch, sing many words, etc. In this case, you'd want your timesteps fine-grained to sufficiently capture what's changing in the vocal track.\n\nOn the other hand, a bass guitar typically plays slow, long notes. You don't want to sample too fast, because not much would be changing from one timestep to another. Imagine taking a picture, then streeeeeeetching out the width. With something like a CNN, your algorithm relies on a neighborhood of pixels. By including many more width pixels than required, the CNN won't find solutions as easily since the X direction doesn't offer much differentiation. It's the same with audio.\n\nEssentially, you want to pick a timestep appropriate to see the pattern in the sound, without going overboard. If you're not noticing much difference along the Time axis in your spectrogram, maybe reduce your timestep.",
    "1141094": "One way is to think of SFFT as convolution.\n\nif time is the x axis and frequency the y axis then:\n\nnftt is the proportional to the width of your output, n_mels is proportional to the height of the output, and hop_length is stride on x axis.",
    "1168052": "@aakashnain  when I was writing my notebook for a [beginner on TF for audio](https://www.kaggle.com/dimitreoliveira/rainforest-audio-classification-tensorflow-starter), I came across [this video](https://www.coursera.org/lecture/audio-signal-processing/stft-2-tjEQe), I am no expert in the field, but it seems to be a good explanation about STFT, I am not sure yet if I will get the chance to really put my hands in this competition, but I hope so.",
    "1170371": ""
  }
}