{
  "id": 172573,
  "title": "Audio Feature Extraction: Best Practices",
  "url": "/competitions/birdsong-recognition/discussion/172573",
  "author_name": "",
  "post_date": "2020-08-05T15:51:13.992037600Z",
  "votes": 47,
  "comment_count": 15,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>I have recently come round a few interesting papers to outline some of the best practices in feature extraction from the audio files.</p>\n<p>Although these papers related to areas other then birdcall identification/classification, their discoveries could be useful in course of this competition.</p>\n<p>I am currently building my pipelines around the processes summarized in the papers shared. Hopefully such an expertise would be useful for other fellow Kagglers trying to contribute to the birdcall identification.</p>\n<h1>Most Useful Audio Features</h1>\n<p>As per the nice article of \"How I Understood: What features to consider while training audio files?\" ( <a href=\"https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b\" target=\"_blank\">https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b</a> ), many audio identification ML projects/experiments benefit from grabbing the features below</p>\n<ul>\n<li>MFCC (Mel-Frequency Cepstral Coefficients)</li>\n<li>Zero-crossing rate</li>\n<li>Energy</li>\n<li>Spectral roll-off</li>\n<li>Spectral flux</li>\n<li>Spectral entropy</li>\n<li>Chroma features (chromatogram), with Chroma vector and Chroma deviation considered to be the most important ones within this group</li>\n<li>Pitch</li>\n</ul>\n<p>The above-mentioned article contains the code examples on how to extract the above-mentioned features in Python, using <em>librosa</em></p>\n<p><em>Note:</em> some theory on MFCC is explained in <a href=\"https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\" target=\"_blank\">https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd</a></p>\n<h1>Practical Case Study</h1>\n<p>There has been undertaken the research to analyze and classify real coronavirus patient cough sounds. </p>\n<p>The research results have been published in \"Coronavirus: Using Machine Learning to Triage COVID-19 Patients\" (<a href=\"https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4)\" target=\"_blank\">https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4)</a>.</p>\n<p>The major take-aways from the project are summarized below</p>\n<h2>Feature Selection</h2>\n<p>Some of the features extracted from the cough samples were</p>\n<ul>\n<li>Chroma-based features which calculates how much of each chromatic pitch class (C, C♯, D, D♯, E, F, F♯, G, G♯, A, A♯, B) exists in the signal.</li>\n<li>Mel-Frequency Cepstral Coefficients (MFCCs). Pratheeksha Nair’s post named ‘The dummy’s guide to MFCC’ notes that any sound generated by a human is determined by the shape of their vocal tract (including the tongue, teeth, etc). If this shape can be determined correctly, any sound produced can be accurately represented.</li>\n<li>Zero-crossing rate measures the number of times the amplitude of a signal passes through a value of zero in a given time frame.</li>\n<li>Spectral Centroid is an indicator of the “brightness” of a given sound, representing the spectral centre of gravity. If you were to take the spectrum, make a wooden block out of it and try to balance it on your finger (across the X-axis), the spectral centroid would be the frequency that your finger “touches” when it successfully balances.</li>\n<li>Spectral Roll-Off is the frequency below which is contained 99% of the energy of the spectrum.</li>\n<li>The Root Mean Squared (RMS) of the waveform that corresponds to its loudness.</li>\n</ul>\n<h2>ML/DL</h2>\n<p>Once the researches had successfully extracted the features from the underlying audio data, they began by fitting a simple Neural Network with a Keras and Tensorflow backend.</p>\n<p>They highly relied on the recommendations of <em>comet.ml</em>'s team on how to apply machine learning and deep learning methods to audio analysis (see <a href=\"https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc\" target=\"_blank\">https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc</a>)</p>\n<h1>References</h1>\n<ul>\n<li>How I Understood: What features to consider while training audio files? - <a href=\"https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b\" target=\"_blank\">https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b</a></li>\n<li>Coronavirus: Using Machine Learning to Triage COVID-19 Patients - <a href=\"https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4\" target=\"_blank\">https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4</a></li>\n<li>The dummy’s guide to MFCC - <a href=\"https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\" target=\"_blank\">https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd</a></li>\n<li>How to apply machine learning and deep learning methods to audio analysis - <a href=\"https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc\" target=\"_blank\">https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc</a></li>\n</ul>",
  "messages": [
    {
      "id": "959465",
      "postDate": "08/05/2020 15:51:13",
      "content": "<h1>Introduction</h1>\n<p>I have recently come round a few interesting papers to outline some of the best practices in feature extraction from the audio files.</p>\n<p>Although these papers related to areas other then birdcall identification/classification, their discoveries could be useful in course of this competition.</p>\n<p>I am currently building my pipelines around the processes summarized in the papers shared. Hopefully such an expertise would be useful for other fellow Kagglers trying to contribute to the birdcall identification.</p>\n<h1>Most Useful Audio Features</h1>\n<p>As per the nice article of \"How I Understood: What features to consider while training audio files?\" ( <a href=\"https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b\" target=\"_blank\">https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b</a> ), many audio identification ML projects/experiments benefit from grabbing the features below</p>\n<ul>\n<li>MFCC (Mel-Frequency Cepstral Coefficients)</li>\n<li>Zero-crossing rate</li>\n<li>Energy</li>\n<li>Spectral roll-off</li>\n<li>Spectral flux</li>\n<li>Spectral entropy</li>\n<li>Chroma features (chromatogram), with Chroma vector and Chroma deviation considered to be the most important ones within this group</li>\n<li>Pitch</li>\n</ul>\n<p>The above-mentioned article contains the code examples on how to extract the above-mentioned features in Python, using <em>librosa</em></p>\n<p><em>Note:</em> some theory on MFCC is explained in <a href=\"https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\" target=\"_blank\">https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd</a></p>\n<h1>Practical Case Study</h1>\n<p>There has been undertaken the research to analyze and classify real coronavirus patient cough sounds. </p>\n<p>The research results have been published in \"Coronavirus: Using Machine Learning to Triage COVID-19 Patients\" (<a href=\"https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4)\" target=\"_blank\">https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4)</a>.</p>\n<p>The major take-aways from the project are summarized below</p>\n<h2>Feature Selection</h2>\n<p>Some of the features extracted from the cough samples were</p>\n<ul>\n<li>Chroma-based features which calculates how much of each chromatic pitch class (C, C♯, D, D♯, E, F, F♯, G, G♯, A, A♯, B) exists in the signal.</li>\n<li>Mel-Frequency Cepstral Coefficients (MFCCs). Pratheeksha Nair’s post named ‘The dummy’s guide to MFCC’ notes that any sound generated by a human is determined by the shape of their vocal tract (including the tongue, teeth, etc). If this shape can be determined correctly, any sound produced can be accurately represented.</li>\n<li>Zero-crossing rate measures the number of times the amplitude of a signal passes through a value of zero in a given time frame.</li>\n<li>Spectral Centroid is an indicator of the “brightness” of a given sound, representing the spectral centre of gravity. If you were to take the spectrum, make a wooden block out of it and try to balance it on your finger (across the X-axis), the spectral centroid would be the frequency that your finger “touches” when it successfully balances.</li>\n<li>Spectral Roll-Off is the frequency below which is contained 99% of the energy of the spectrum.</li>\n<li>The Root Mean Squared (RMS) of the waveform that corresponds to its loudness.</li>\n</ul>\n<h2>ML/DL</h2>\n<p>Once the researches had successfully extracted the features from the underlying audio data, they began by fitting a simple Neural Network with a Keras and Tensorflow backend.</p>\n<p>They highly relied on the recommendations of <em>comet.ml</em>'s team on how to apply machine learning and deep learning methods to audio analysis (see <a href=\"https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc\" target=\"_blank\">https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc</a>)</p>\n<h1>References</h1>\n<ul>\n<li>How I Understood: What features to consider while training audio files? - <a href=\"https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b\" target=\"_blank\">https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b</a></li>\n<li>Coronavirus: Using Machine Learning to Triage COVID-19 Patients - <a href=\"https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4\" target=\"_blank\">https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4</a></li>\n<li>The dummy’s guide to MFCC - <a href=\"https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\" target=\"_blank\">https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd</a></li>\n<li>How to apply machine learning and deep learning methods to audio analysis - <a href=\"https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc\" target=\"_blank\">https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc</a></li>\n</ul>",
      "rawMarkdown": "# Introduction\n\nI have recently come round a few interesting papers to outline some of the best practices in feature extraction from the audio files.\n\nAlthough these papers related to areas other then birdcall identification/classification, their discoveries could be useful in course of this competition.\n\nI am currently building my pipelines around the processes summarized in the papers shared. Hopefully such an expertise would be useful for other fellow Kagglers trying to contribute to the birdcall identification.\n\n# Most Useful Audio Features\n\nAs per the nice article of \"How I Understood: What features to consider while training audio files?\" ( https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b ), many audio identification ML projects/experiments benefit from grabbing the features below\n\n- MFCC (Mel-Frequency Cepstral Coefficients)\n- Zero-crossing rate\n- Energy\n- Spectral roll-off\n- Spectral flux\n- Spectral entropy\n- Chroma features (chromatogram), with Chroma vector and Chroma deviation considered to be the most important ones within this group\n- Pitch\n\nThe above-mentioned article contains the code examples on how to extract the above-mentioned features in Python, using *librosa*\n\n*Note:* some theory on MFCC is explained in https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\n\n# Practical Case Study\n\nThere has been undertaken the research to analyze and classify real coronavirus patient cough sounds. \n\nThe research results have been published in \"Coronavirus: Using Machine Learning to Triage COVID-19 Patients\" (https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4).\n\nThe major take-aways from the project are summarized below\n\n## Feature Selection\n\nSome of the features extracted from the cough samples were\n\n- Chroma-based features which calculates how much of each chromatic pitch class (C, C♯, D, D♯, E, F, F♯, G, G♯, A, A♯, B) exists in the signal.\n- Mel-Frequency Cepstral Coefficients (MFCCs). Pratheeksha Nair’s post named ‘The dummy’s guide to MFCC’ notes that any sound generated by a human is determined by the shape of their vocal tract (including the tongue, teeth, etc). If this shape can be determined correctly, any sound produced can be accurately represented.\n- Zero-crossing rate measures the number of times the amplitude of a signal passes through a value of zero in a given time frame.\n- Spectral Centroid is an indicator of the “brightness” of a given sound, representing the spectral centre of gravity. If you were to take the spectrum, make a wooden block out of it and try to balance it on your finger (across the X-axis), the spectral centroid would be the frequency that your finger “touches” when it successfully balances.\n- Spectral Roll-Off is the frequency below which is contained 99% of the energy of the spectrum.\n- The Root Mean Squared (RMS) of the waveform that corresponds to its loudness.\n\n## ML/DL \n\nOnce the researches had successfully extracted the features from the underlying audio data, they began by fitting a simple Neural Network with a Keras and Tensorflow backend.\n\nThey highly relied on the recommendations of *comet.ml*'s team on how to apply machine learning and deep learning methods to audio analysis (see https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc)\n\n# References\n\n- How I Understood: What features to consider while training audio files? - https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b\n- Coronavirus: Using Machine Learning to Triage COVID-19 Patients - https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4\n- The dummy’s guide to MFCC - https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\n- How to apply machine learning and deep learning methods to audio analysis - https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc",
      "votes": null
    },
    {
      "id": "959485",
      "postDate": "08/05/2020 16:09:19",
      "content": "<p>Really nice summary for feature selection on audio files. Thanks for sharing. Kudos 👍 </p>",
      "rawMarkdown": "Really nice summary for feature selection on audio files. Thanks for sharing. Kudos 👍",
      "votes": null
    },
    {
      "id": "959892",
      "postDate": "08/06/2020 02:29:27",
      "content": "<p>Nice one </p>",
      "rawMarkdown": "Nice one",
      "votes": null
    },
    {
      "id": "960098",
      "postDate": "08/06/2020 06:19:18",
      "content": "<p>thanks <a href=\"/gvyshnya\">@gvyshnya</a> for this info, i was looking for something like this.</p>",
      "rawMarkdown": "thanks @gvyshnya for this info, i was looking for something like this.",
      "votes": null
    },
    {
      "id": "960430",
      "postDate": "08/06/2020 11:53:06",
      "content": "<p>Thanks for sharing. It's very useful !</p>",
      "rawMarkdown": "Thanks for sharing. It's very useful !",
      "votes": null
    },
    {
      "id": "962490",
      "postDate": "08/08/2020 06:57:55",
      "content": "<p>super</p>",
      "rawMarkdown": "super",
      "votes": null
    },
    {
      "id": "962513",
      "postDate": "08/08/2020 07:25:31",
      "content": "<p>Very nice blog to jump start aural analysis!!!</p>",
      "rawMarkdown": "Very nice blog to jump start aural analysis!!!",
      "votes": null
    },
    {
      "id": "962571",
      "postDate": "08/08/2020 08:45:07",
      "content": "<p>Great work.  Upvoted for you!</p>",
      "rawMarkdown": "Great work.  Upvoted for you!",
      "votes": null
    },
    {
      "id": "963474",
      "postDate": "08/09/2020 04:12:05",
      "content": "<p>Thank you very much for that post!</p>",
      "rawMarkdown": "Thank you very much for that post!",
      "votes": null
    },
    {
      "id": "964754",
      "postDate": "08/10/2020 06:29:40",
      "content": "<p>I have recently worked in this domain and found that for Audio Feature Extraction there are many libraries like Librosa, pyAudioAnalysis, etc. I personally find pyAudioAnalysis and Librosa very useful. </p>\n\n<p>pyaudioAnalysis has a total number of 34 short-term features implemented and the feature details are given here- <a href=\"https://github.com/tyiannak/pyAudioAnalysis/wiki/3.-Feature-Extraction\">https://github.com/tyiannak/pyAudioAnalysis/wiki/3.-Feature-Extraction</a>  </p>\n\n<p>pyAudioAnalysis also comes with already implemented classifications and regression models too and the details can be found here-  <a href=\"https://github.com/tyiannak/pyAudioAnalysis/wiki/4.-Classification-and-Regression\">https://github.com/tyiannak/pyAudioAnalysis/wiki/4.-Classification-and-Regression</a></p>",
      "rawMarkdown": "I have recently worked in this domain and found that for Audio Feature Extraction there are many libraries like Librosa, pyAudioAnalysis, etc. I personally find pyAudioAnalysis and Librosa very useful. \n\npyaudioAnalysis has a total number of 34 short-term features implemented and the feature details are given here- https://github.com/tyiannak/pyAudioAnalysis/wiki/3.-Feature-Extraction  \n\npyAudioAnalysis also comes with already implemented classifications and regression models too and the details can be found here-  https://github.com/tyiannak/pyAudioAnalysis/wiki/4.-Classification-and-Regression",
      "votes": null
    },
    {
      "id": "966493",
      "postDate": "08/11/2020 13:10:24",
      "content": "<p>Thanks. It's very helpful.</p>",
      "rawMarkdown": "Thanks. It's very helpful.",
      "votes": null
    },
    {
      "id": "985448",
      "postDate": "08/25/2020 18:21:05",
      "content": "<p>Thank you for sharing these informative articles this can give me some approach I need to start!!</p>",
      "rawMarkdown": "Thank you for sharing these informative articles this can give me some approach I need to start!!",
      "votes": null
    },
    {
      "id": "1501915",
      "postDate": "09/03/2021 16:29:23",
      "content": "<p>Your notebooks give me so much insight and resources to keep on learning. Thank you for sharing</p>",
      "rawMarkdown": "Your notebooks give me so much insight and resources to keep on learning. Thank you for sharing",
      "votes": null
    },
    {
      "id": "1502078",
      "postDate": "09/03/2021 19:33:43",
      "content": "<p><a href=\"https://www.kaggle.com/juanpasutti\" target=\"_blank\">@juanpasutti</a> : Hey Juan, you are welcome!</p>",
      "rawMarkdown": "juanpasutti : Hey Juan, you are welcome!",
      "votes": null
    },
    {
      "id": "3170326",
      "postDate": "04/04/2025 15:33:25",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/gvyshnya\" target=\"_blank\">@gvyshnya</a> , as a complete beginner to audio signal processing with previous ML knowledge it was very helpful to me! </p>",
      "rawMarkdown": "Thank you very much @gvyshnya , as a complete beginner to audio signal processing with previous ML knowledge it was very helpful to me!",
      "votes": null
    },
    {
      "id": "3249992",
      "postDate": "07/17/2025 13:46:54",
      "content": "<p><a href=\"https://www.kaggle.com/nakulkrishnakumar\" target=\"_blank\">@nakulkrishnakumar</a> : my pleasure!</p>",
      "rawMarkdown": "nakulkrishnakumar : my pleasure!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 963474,
      "author_name": "leongurtler",
      "author_url": "",
      "post_date": "08/09/2020 04:12:05",
      "content": "<p>Thank you very much for that post!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 966493,
      "author_name": "shadowburning",
      "author_url": "",
      "post_date": "08/11/2020 13:10:24",
      "content": "<p>Thanks. It's very helpful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 985448,
      "author_name": "digvijayyadav",
      "author_url": "",
      "post_date": "08/25/2020 18:21:05",
      "content": "<p>Thank you for sharing these informative articles this can give me some approach I need to start!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1501915,
      "author_name": "juanpasutti",
      "author_url": "",
      "post_date": "09/03/2021 16:29:23",
      "content": "<p>Your notebooks give me so much insight and resources to keep on learning. Thank you for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 1502078,
          "author_name": "gvyshnya",
          "author_url": "",
          "post_date": "09/03/2021 19:33:43",
          "content": "<p><a href=\"https://www.kaggle.com/juanpasutti\" target=\"_blank\">@juanpasutti</a> : Hey Juan, you are welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3170326,
      "author_name": "nakulkrishnakumar",
      "author_url": "",
      "post_date": "04/04/2025 15:33:25",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/gvyshnya\" target=\"_blank\">@gvyshnya</a> , as a complete beginner to audio signal processing with previous ML knowledge it was very helpful to me! </p>",
      "votes": null,
      "replies": [
        {
          "id": 3249992,
          "author_name": "gvyshnya",
          "author_url": "",
          "post_date": "07/17/2025 13:46:54",
          "content": "<p><a href=\"https://www.kaggle.com/nakulkrishnakumar\" target=\"_blank\">@nakulkrishnakumar</a> : my pleasure!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 959485,
      "author_name": "narendrageek",
      "author_url": "",
      "post_date": "08/05/2020 16:09:19",
      "content": "<p>Really nice summary for feature selection on audio files. Thanks for sharing. Kudos 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 959892,
      "author_name": "sudarshanpatil",
      "author_url": "",
      "post_date": "08/06/2020 02:29:27",
      "content": "<p>Nice one </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 960098,
      "author_name": "saurah403",
      "author_url": "",
      "post_date": "08/06/2020 06:19:18",
      "content": "<p>thanks <a href=\"/gvyshnya\">@gvyshnya</a> for this info, i was looking for something like this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 960430,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "08/06/2020 11:53:06",
      "content": "<p>Thanks for sharing. It's very useful !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962490,
      "author_name": "kushagra7744",
      "author_url": "",
      "post_date": "08/08/2020 06:57:55",
      "content": "<p>super</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962513,
      "author_name": "aayuver",
      "author_url": "",
      "post_date": "08/08/2020 07:25:31",
      "content": "<p>Very nice blog to jump start aural analysis!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962571,
      "author_name": "",
      "author_url": "",
      "post_date": "08/08/2020 08:45:07",
      "content": "<p>Great work.  Upvoted for you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 964754,
      "author_name": "palaksood97",
      "author_url": "",
      "post_date": "08/10/2020 06:29:40",
      "content": "<p>I have recently worked in this domain and found that for Audio Feature Extraction there are many libraries like Librosa, pyAudioAnalysis, etc. I personally find pyAudioAnalysis and Librosa very useful. </p>\n\n<p>pyaudioAnalysis has a total number of 34 short-term features implemented and the feature details are given here- <a href=\"https://github.com/tyiannak/pyAudioAnalysis/wiki/3.-Feature-Extraction\">https://github.com/tyiannak/pyAudioAnalysis/wiki/3.-Feature-Extraction</a>  </p>\n\n<p>pyAudioAnalysis also comes with already implemented classifications and regression models too and the details can be found here-  <a href=\"https://github.com/tyiannak/pyAudioAnalysis/wiki/4.-Classification-and-Regression\">https://github.com/tyiannak/pyAudioAnalysis/wiki/4.-Classification-and-Regression</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "959465": "# Introduction\n\nI have recently come round a few interesting papers to outline some of the best practices in feature extraction from the audio files.\n\nAlthough these papers related to areas other then birdcall identification/classification, their discoveries could be useful in course of this competition.\n\nI am currently building my pipelines around the processes summarized in the papers shared. Hopefully such an expertise would be useful for other fellow Kagglers trying to contribute to the birdcall identification.\n\n# Most Useful Audio Features\n\nAs per the nice article of \"How I Understood: What features to consider while training audio files?\" ( https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b ), many audio identification ML projects/experiments benefit from grabbing the features below\n\n- MFCC (Mel-Frequency Cepstral Coefficients)\n- Zero-crossing rate\n- Energy\n- Spectral roll-off\n- Spectral flux\n- Spectral entropy\n- Chroma features (chromatogram), with Chroma vector and Chroma deviation considered to be the most important ones within this group\n- Pitch\n\nThe above-mentioned article contains the code examples on how to extract the above-mentioned features in Python, using *librosa*\n\n*Note:* some theory on MFCC is explained in https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\n\n# Practical Case Study\n\nThere has been undertaken the research to analyze and classify real coronavirus patient cough sounds. \n\nThe research results have been published in \"Coronavirus: Using Machine Learning to Triage COVID-19 Patients\" (https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4).\n\nThe major take-aways from the project are summarized below\n\n## Feature Selection\n\nSome of the features extracted from the cough samples were\n\n- Chroma-based features which calculates how much of each chromatic pitch class (C, C♯, D, D♯, E, F, F♯, G, G♯, A, A♯, B) exists in the signal.\n- Mel-Frequency Cepstral Coefficients (MFCCs). Pratheeksha Nair’s post named ‘The dummy’s guide to MFCC’ notes that any sound generated by a human is determined by the shape of their vocal tract (including the tongue, teeth, etc). If this shape can be determined correctly, any sound produced can be accurately represented.\n- Zero-crossing rate measures the number of times the amplitude of a signal passes through a value of zero in a given time frame.\n- Spectral Centroid is an indicator of the “brightness” of a given sound, representing the spectral centre of gravity. If you were to take the spectrum, make a wooden block out of it and try to balance it on your finger (across the X-axis), the spectral centroid would be the frequency that your finger “touches” when it successfully balances.\n- Spectral Roll-Off is the frequency below which is contained 99% of the energy of the spectrum.\n- The Root Mean Squared (RMS) of the waveform that corresponds to its loudness.\n\n## ML/DL \n\nOnce the researches had successfully extracted the features from the underlying audio data, they began by fitting a simple Neural Network with a Keras and Tensorflow backend.\n\nThey highly relied on the recommendations of *comet.ml*'s team on how to apply machine learning and deep learning methods to audio analysis (see https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc)\n\n# References\n\n- How I Understood: What features to consider while training audio files? - https://towardsdatascience.com/how-i-understood-what-features-to-consider-while-training-audio-files-eedfb6e9002b\n- Coronavirus: Using Machine Learning to Triage COVID-19 Patients - https://towardsdatascience.com/coronavirus-using-machine-learning-to-triage-covid-19-patients-980e62489fd4\n- The dummy’s guide to MFCC - https://medium.com/prathena/the-dummys-guide-to-mfcc-aceab2450fd\n- How to apply machine learning and deep learning methods to audio analysis - https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc",
    "959485": "Really nice summary for feature selection on audio files. Thanks for sharing. Kudos 👍",
    "959892": "Nice one",
    "960098": "thanks @gvyshnya for this info, i was looking for something like this.",
    "960430": "Thanks for sharing. It's very useful !",
    "962490": "super",
    "962513": "Very nice blog to jump start aural analysis!!!",
    "962571": "Great work.  Upvoted for you!",
    "963474": "Thank you very much for that post!",
    "964754": "I have recently worked in this domain and found that for Audio Feature Extraction there are many libraries like Librosa, pyAudioAnalysis, etc. I personally find pyAudioAnalysis and Librosa very useful. \n\npyaudioAnalysis has a total number of 34 short-term features implemented and the feature details are given here- https://github.com/tyiannak/pyAudioAnalysis/wiki/3.-Feature-Extraction  \n\npyAudioAnalysis also comes with already implemented classifications and regression models too and the details can be found here-  https://github.com/tyiannak/pyAudioAnalysis/wiki/4.-Classification-and-Regression",
    "966493": "Thanks. It's very helpful.",
    "985448": "Thank you for sharing these informative articles this can give me some approach I need to start!!",
    "1501915": "Your notebooks give me so much insight and resources to keep on learning. Thank you for sharing",
    "1502078": "juanpasutti : Hey Juan, you are welcome!",
    "3170326": "Thank you very much @gvyshnya , as a complete beginner to audio signal processing with previous ML knowledge it was very helpful to me!",
    "3249992": "nakulkrishnakumar : my pleasure!"
  },
  "source": "meta"
}