{
  "id": 201129,
  "title": "Analyzing and processing the features in frequency domain only, but why?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/201129",
  "author_name": "Ultron",
  "post_date": "2020-12-03T09:48:25.947000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Let's understand what does competition demand? It simply asks us to find the probability of the presence of audio of given target species in the test audio recordings. Right? We don't have to find out exactly where the audio of a species locates in a sample, we just need to find out the probability that whether it's there or not. </p>\n<p>Now, let's become a bit domain-specific. An audio signal can contain one or more frequency components. But these effects or detection of these frequency components in time domain representation is obfuscated. As there remains a traditional time-frequency trade-off, the effects of frequency components aren't quite considerable in time domain representation. But our target is to find whether a particular frequency or range of frequencies is/are present in a given audio sample or not. That's why we all are using frequency domain analysis,  spectral components, Fourier transforms, etc. </p>\n<p>Hope it makes sense, correct my intuitions if it's wrong. Thanks.</p>\n<p>Here is my detailed notebook: <a href=\"https://www.kaggle.com/mrutyunjaybiswal/rainforest-audio-detection-detailsexplained-eda\" target=\"_blank\">https://www.kaggle.com/mrutyunjaybiswal/rainforest-audio-detection-detailsexplained-eda</a></p>",
  "messages": [
    {
      "id": 1100741,
      "postDate": "2020-12-03T09:48:25.947Z",
      "content": "<p>Let's understand what does competition demand? It simply asks us to find the probability of the presence of audio of given target species in the test audio recordings. Right? We don't have to find out exactly where the audio of a species locates in a sample, we just need to find out the probability that whether it's there or not. </p>\n<p>Now, let's become a bit domain-specific. An audio signal can contain one or more frequency components. But these effects or detection of these frequency components in time domain representation is obfuscated. As there remains a traditional time-frequency trade-off, the effects of frequency components aren't quite considerable in time domain representation. But our target is to find whether a particular frequency or range of frequencies is/are present in a given audio sample or not. That's why we all are using frequency domain analysis,  spectral components, Fourier transforms, etc. </p>\n<p>Hope it makes sense, correct my intuitions if it's wrong. Thanks.</p>\n<p>Here is my detailed notebook: <a href=\"https://www.kaggle.com/mrutyunjaybiswal/rainforest-audio-detection-detailsexplained-eda\" target=\"_blank\">https://www.kaggle.com/mrutyunjaybiswal/rainforest-audio-detection-detailsexplained-eda</a></p>",
      "rawMarkdown": "Let's understand what does competition demand? It simply asks us to find the probability of the presence of audio of given target species in the test audio recordings. Right? We don't have to find out exactly where the audio of a species locates in a sample, we just need to find out the probability that whether it's there or not. \n\nNow, let's become a bit domain-specific. An audio signal can contain one or more frequency components. But these effects or detection of these frequency components in time domain representation is obfuscated. As there remains a traditional time-frequency trade-off, the effects of frequency components aren't quite considerable in time domain representation. But our target is to find whether a particular frequency or range of frequencies is/are present in a given audio sample or not. That's why we all are using frequency domain analysis,  spectral components, Fourier transforms, etc. \n\nHope it makes sense, correct my intuitions if it's wrong. Thanks.\n\nHere is my detailed notebook: https://www.kaggle.com/mrutyunjaybiswal/rainforest-audio-detection-detailsexplained-eda",
      "votes": 6
    },
    {
      "id": 1101720,
      "postDate": "2020-12-04T07:08:33.813Z",
      "content": "<p>Yes you are right. When we analyze the audio signal in frequency space, signals can be decomposed by frequency band. And it help detect the target signal. The Convolution layer with the Fully Connected layer work effectively to get the various energy constant of the each frequency band. <br>\nOn the other hands, If we analyze the audio signal in the time space, it is difficult to detect a target signal. Because there is background noise. And also target signal information is loss during down sampling.<br>\nAdditionally, the operation like integral and differentiation can be expressed by convolution kernel in frequency domain. I think this is the most reason why we analysis the audio signal on frequency domain. </p>\n<p>However, I think Fourier transform is not suffice to express all of the audio signal information. Because we can only use finite frequency band information. So, I think creating features without using spectrogram is also important to this project. I will try to select feature manually.</p>",
      "rawMarkdown": "Yes you are right. When we analyze the audio signal in frequency space, signals can be decomposed by frequency band. And it help detect the target signal. The Convolution layer with the Fully Connected layer work effectively to get the various energy constant of the each frequency band. \nOn the other hands, If we analyze the audio signal in the time space, it is difficult to detect a target signal. Because there is background noise. And also target signal information is loss during down sampling.\nAdditionally, the operation like integral and differentiation can be expressed by convolution kernel in frequency domain. I think this is the most reason why we analysis the audio signal on frequency domain. \n\nHowever, I think Fourier transform is not suffice to express all of the audio signal information. Because we can only use finite frequency band information. So, I think creating features without using spectrogram is also important to this project. I will try to select feature manually.",
      "votes": 1,
      "replies": [
        {
          "id": 1102310,
          "postDate": "2020-12-04T19:31:12.020Z",
          "content": "<p>I would think that all of the information required can be found in a spectrogram, depending on what sampling rate you use for the time-step component. </p>\n<p>I agree a single FFT will lose a lot of information, such as rhythm, but a spectrogram (since it has the time-component) should be able to capture rhythm.</p>\n<p>I think taking the sum of all frequencies per time step in the spectrogram would yield a similar plot to the original time series, albeit at a reduced sampling rate and on a different scale. I may try this myself to see how it looks, it's an interesting idea!</p>",
          "rawMarkdown": "I would think that all of the information required can be found in a spectrogram, depending on what sampling rate you use for the time-step component. \n\nI agree a single FFT will lose a lot of information, such as rhythm, but a spectrogram (since it has the time-component) should be able to capture rhythm.\n\nI think taking the sum of all frequencies per time step in the spectrogram would yield a similar plot to the original time series, albeit at a reduced sampling rate and on a different scale. I may try this myself to see how it looks, it's an interesting idea!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1101723,
      "postDate": "2020-12-04T07:12:56.497Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1101720,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2020-12-04T07:08:33.813000",
      "content": "<p>Yes you are right. When we analyze the audio signal in frequency space, signals can be decomposed by frequency band. And it help detect the target signal. The Convolution layer with the Fully Connected layer work effectively to get the various energy constant of the each frequency band. <br>\nOn the other hands, If we analyze the audio signal in the time space, it is difficult to detect a target signal. Because there is background noise. And also target signal information is loss during down sampling.<br>\nAdditionally, the operation like integral and differentiation can be expressed by convolution kernel in frequency domain. I think this is the most reason why we analysis the audio signal on frequency domain. </p>\n<p>However, I think Fourier transform is not suffice to express all of the audio signal information. Because we can only use finite frequency band information. So, I think creating features without using spectrogram is also important to this project. I will try to select feature manually.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1102310,
          "author_name": "Mike A",
          "author_url": "",
          "post_date": "2020-12-04T19:31:12.020000",
          "content": "<p>I would think that all of the information required can be found in a spectrogram, depending on what sampling rate you use for the time-step component. </p>\n<p>I agree a single FFT will lose a lot of information, such as rhythm, but a spectrogram (since it has the time-component) should be able to capture rhythm.</p>\n<p>I think taking the sum of all frequencies per time step in the spectrogram would yield a similar plot to the original time series, albeit at a reduced sampling rate and on a different scale. I may try this myself to see how it looks, it's an interesting idea!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1101723,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-04T07:12:56.497000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1100741": "Let's understand what does competition demand? It simply asks us to find the probability of the presence of audio of given target species in the test audio recordings. Right? We don't have to find out exactly where the audio of a species locates in a sample, we just need to find out the probability that whether it's there or not. \n\nNow, let's become a bit domain-specific. An audio signal can contain one or more frequency components. But these effects or detection of these frequency components in time domain representation is obfuscated. As there remains a traditional time-frequency trade-off, the effects of frequency components aren't quite considerable in time domain representation. But our target is to find whether a particular frequency or range of frequencies is/are present in a given audio sample or not. That's why we all are using frequency domain analysis,  spectral components, Fourier transforms, etc. \n\nHope it makes sense, correct my intuitions if it's wrong. Thanks.\n\nHere is my detailed notebook: https://www.kaggle.com/mrutyunjaybiswal/rainforest-audio-detection-detailsexplained-eda",
    "1101720": "Yes you are right. When we analyze the audio signal in frequency space, signals can be decomposed by frequency band. And it help detect the target signal. The Convolution layer with the Fully Connected layer work effectively to get the various energy constant of the each frequency band. \nOn the other hands, If we analyze the audio signal in the time space, it is difficult to detect a target signal. Because there is background noise. And also target signal information is loss during down sampling.\nAdditionally, the operation like integral and differentiation can be expressed by convolution kernel in frequency domain. I think this is the most reason why we analysis the audio signal on frequency domain. \n\nHowever, I think Fourier transform is not suffice to express all of the audio signal information. Because we can only use finite frequency band information. So, I think creating features without using spectrogram is also important to this project. I will try to select feature manually.",
    "1101723": ""
  }
}