{
  "id": 199048,
  "title": "Beginner Help",
  "url": "/competitions/rfcx-species-audio-detection/discussion/199048",
  "author_name": "",
  "post_date": "2020-11-24T06:36:28.402297500Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello,<br>\nI am new to audio competitions so it would be helpful if anyone suggested any notebook in kaggle or external source from where I can get the idea to get started with this sort of  competition and understand basics concepts used in the kernels</p>",
  "messages": [
    {
      "id": "1089029",
      "postDate": "11/24/2020 06:36:28",
      "content": "<p>Hello,<br>\nI am new to audio competitions so it would be helpful if anyone suggested any notebook in kaggle or external source from where I can get the idea to get started with this sort of  competition and understand basics concepts used in the kernels</p>",
      "rawMarkdown": "Hello,\nI am new to audio competitions so it would be helpful if anyone suggested any notebook in kaggle or external source from where I can get the idea to get started with this sort of  competition and understand basics concepts used in the kernels",
      "votes": null
    },
    {
      "id": "1089094",
      "postDate": "11/24/2020 07:49:03",
      "content": "<p>Hi, if you are new to any kind of sequence related problem, I would suggest going through any lectures. There are plenty available on youtube. Then chose a framework and understand how to do sequence related sampling.</p>",
      "rawMarkdown": "Hi, if you are new to any kind of sequence related problem, I would suggest going through any lectures. There are plenty available on youtube. Then chose a framework and understand how to do sequence related sampling.",
      "votes": null
    },
    {
      "id": "1089420",
      "postDate": "11/24/2020 13:34:13",
      "content": "<p>One common approach to working with audio is to create a spectrogram image of the signal (e.g. Mel Spectrogram) and then use an image model like ResNet to train on the images. The problem is then similar to something like MNIST.</p>\n<p>Essentially the Mel Spectrogram consists of 3 ingredients:</p>\n<ul>\n<li>FFT (The Fast Fourier Transform aka the Discrete Fourier Transform) this will tell you the strength of each frequency in a signal but it will not tell you how the frequencies changes over time.</li>\n<li>STFT (The Short Time Fourier Transform) The FFT is applied multiple times to small chunks of the signal along the time axis in a sliding window fashion and will tell you how the strength of the frequencies change over time. </li>\n<li>Mel Scale Transform. Transforms frequencies from Hertz to Mels. The Mel Scale is meant to approximate how humans perceive differences in pitches. <a href=\"https://en.wikipedia.org/wiki/Mel_scale\" target=\"_blank\">https://en.wikipedia.org/wiki/Mel_scale</a> </li>\n</ul>\n<p>Mel Spectrogram = STFT + Mel Scale Transform.</p>\n<p>Kernels using Mel Spectrogram + Image model approach:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/doanquanvietnamca/training-baseline-resnest-rfcx-audio-detection\" target=\"_blank\">https://www.kaggle.com/doanquanvietnamca/training-baseline-resnest-rfcx-audio-detection</a></li>\n<li><a href=\"https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection\" target=\"_blank\">https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection</a></li>\n<li><a href=\"https://www.kaggle.com/jackvial/pytorch-lightning-starter\" target=\"_blank\">https://www.kaggle.com/jackvial/pytorch-lightning-starter</a></li>\n</ul>\n<p>This articles might also help in building an understanding of the Mel Spectrogram <a href=\"https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\" target=\"_blank\">https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0</a></p>\n<p>Another common approach is to take the FFT (as described above) of the signal and use this as the features to a model like Logistic Regression, Random Forest, SVM:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast\" target=\"_blank\">https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast</a></li>\n<li><a href=\"https://www.kaggle.com/tunguz/rainforest-rapids-baseline\" target=\"_blank\">https://www.kaggle.com/tunguz/rainforest-rapids-baseline</a></li>\n</ul>",
      "rawMarkdown": "One common approach to working with audio is to create a spectrogram image of the signal (e.g. Mel Spectrogram) and then use an image model like ResNet to train on the images. The problem is then similar to something like MNIST.\n\nEssentially the Mel Spectrogram consists of 3 ingredients:\n- FFT (The Fast Fourier Transform aka the Discrete Fourier Transform) this will tell you the strength of each frequency in a signal but it will not tell you how the frequencies changes over time.\n- STFT (The Short Time Fourier Transform) The FFT is applied multiple times to small chunks of the signal along the time axis in a sliding window fashion and will tell you how the strength of the frequencies change over time. \n- Mel Scale Transform. Transforms frequencies from Hertz to Mels. The Mel Scale is meant to approximate how humans perceive differences in pitches. https://en.wikipedia.org/wiki/Mel_scale \n \nMel Spectrogram = STFT + Mel Scale Transform.\n\nKernels using Mel Spectrogram + Image model approach:\n- https://www.kaggle.com/doanquanvietnamca/training-baseline-resnest-rfcx-audio-detection\n- https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection\n- https://www.kaggle.com/jackvial/pytorch-lightning-starter\n\nThis articles might also help in building an understanding of the Mel Spectrogram https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\n\nAnother common approach is to take the FFT (as described above) of the signal and use this as the features to a model like Logistic Regression, Random Forest, SVM:\n- https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast\n- https://www.kaggle.com/tunguz/rainforest-rapids-baseline",
      "votes": null
    },
    {
      "id": "1102089",
      "postDate": "12/04/2020 15:20:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nur988\" target=\"_blank\">@nur988</a> </p>\n<p>I suggest you look at earlier sound competitions and see their solutions as well . it will give you some sense of what has worked<br>\nI have complied the best solution of the last sound competition: Cornell Birdcall Identification in link below:<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197873\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197873</a></p>",
      "rawMarkdown": "Hi @nur988 \n\nI suggest you look at earlier sound competitions and see their solutions as well . it will give you some sense of what has worked\nI have complied the best solution of the last sound competition: Cornell Birdcall Identification in link below:\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197873",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1089094,
      "author_name": "archanghosh",
      "author_url": "",
      "post_date": "11/24/2020 07:49:03",
      "content": "<p>Hi, if you are new to any kind of sequence related problem, I would suggest going through any lectures. There are plenty available on youtube. Then chose a framework and understand how to do sequence related sampling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1089420,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "11/24/2020 13:34:13",
      "content": "<p>One common approach to working with audio is to create a spectrogram image of the signal (e.g. Mel Spectrogram) and then use an image model like ResNet to train on the images. The problem is then similar to something like MNIST.</p>\n<p>Essentially the Mel Spectrogram consists of 3 ingredients:</p>\n<ul>\n<li>FFT (The Fast Fourier Transform aka the Discrete Fourier Transform) this will tell you the strength of each frequency in a signal but it will not tell you how the frequencies changes over time.</li>\n<li>STFT (The Short Time Fourier Transform) The FFT is applied multiple times to small chunks of the signal along the time axis in a sliding window fashion and will tell you how the strength of the frequencies change over time. </li>\n<li>Mel Scale Transform. Transforms frequencies from Hertz to Mels. The Mel Scale is meant to approximate how humans perceive differences in pitches. <a href=\"https://en.wikipedia.org/wiki/Mel_scale\" target=\"_blank\">https://en.wikipedia.org/wiki/Mel_scale</a> </li>\n</ul>\n<p>Mel Spectrogram = STFT + Mel Scale Transform.</p>\n<p>Kernels using Mel Spectrogram + Image model approach:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/doanquanvietnamca/training-baseline-resnest-rfcx-audio-detection\" target=\"_blank\">https://www.kaggle.com/doanquanvietnamca/training-baseline-resnest-rfcx-audio-detection</a></li>\n<li><a href=\"https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection\" target=\"_blank\">https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection</a></li>\n<li><a href=\"https://www.kaggle.com/jackvial/pytorch-lightning-starter\" target=\"_blank\">https://www.kaggle.com/jackvial/pytorch-lightning-starter</a></li>\n</ul>\n<p>This articles might also help in building an understanding of the Mel Spectrogram <a href=\"https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\" target=\"_blank\">https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0</a></p>\n<p>Another common approach is to take the FFT (as described above) of the signal and use this as the features to a model like Logistic Regression, Random Forest, SVM:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast\" target=\"_blank\">https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast</a></li>\n<li><a href=\"https://www.kaggle.com/tunguz/rainforest-rapids-baseline\" target=\"_blank\">https://www.kaggle.com/tunguz/rainforest-rapids-baseline</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1102089,
      "author_name": "kmldas",
      "author_url": "",
      "post_date": "12/04/2020 15:20:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nur988\" target=\"_blank\">@nur988</a> </p>\n<p>I suggest you look at earlier sound competitions and see their solutions as well . it will give you some sense of what has worked<br>\nI have complied the best solution of the last sound competition: Cornell Birdcall Identification in link below:<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197873\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197873</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1089029": "Hello,\nI am new to audio competitions so it would be helpful if anyone suggested any notebook in kaggle or external source from where I can get the idea to get started with this sort of  competition and understand basics concepts used in the kernels",
    "1089094": "Hi, if you are new to any kind of sequence related problem, I would suggest going through any lectures. There are plenty available on youtube. Then chose a framework and understand how to do sequence related sampling.",
    "1089420": "One common approach to working with audio is to create a spectrogram image of the signal (e.g. Mel Spectrogram) and then use an image model like ResNet to train on the images. The problem is then similar to something like MNIST.\n\nEssentially the Mel Spectrogram consists of 3 ingredients:\n- FFT (The Fast Fourier Transform aka the Discrete Fourier Transform) this will tell you the strength of each frequency in a signal but it will not tell you how the frequencies changes over time.\n- STFT (The Short Time Fourier Transform) The FFT is applied multiple times to small chunks of the signal along the time axis in a sliding window fashion and will tell you how the strength of the frequencies change over time. \n- Mel Scale Transform. Transforms frequencies from Hertz to Mels. The Mel Scale is meant to approximate how humans perceive differences in pitches. https://en.wikipedia.org/wiki/Mel_scale \n \nMel Spectrogram = STFT + Mel Scale Transform.\n\nKernels using Mel Spectrogram + Image model approach:\n- https://www.kaggle.com/doanquanvietnamca/training-baseline-resnest-rfcx-audio-detection\n- https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection\n- https://www.kaggle.com/jackvial/pytorch-lightning-starter\n\nThis articles might also help in building an understanding of the Mel Spectrogram https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0\n\nAnother common approach is to take the FFT (as described above) of the signal and use this as the features to a model like Logistic Regression, Random Forest, SVM:\n- https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast\n- https://www.kaggle.com/tunguz/rainforest-rapids-baseline",
    "1102089": "Hi @nur988 \n\nI suggest you look at earlier sound competitions and see their solutions as well . it will give you some sense of what has worked\nI have complied the best solution of the last sound competition: Cornell Birdcall Identification in link below:\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197873"
  },
  "source": "meta"
}