{
  "id": 179662,
  "title": "Parallelized Audio Feature Extraction Study",
  "url": "/competitions/birdsong-recognition/discussion/179662",
  "author_name": "Georgii Vyshnia",
  "post_date": "2020-09-02T15:33:42.326000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Preface: Curse of Speed Limitations</h1>\n<p>Python has been the standard de facto for the majority of industrial ML/AI solutions. Its extreme programmer-friendliness, along with the convenience and wide range of libraries for data analytics, data processing, data engineering, system integration and ML made it the language of choice of many experts, vendors and corporate Data Science teams.</p>\n<p>At the same time, it isn’t the fastest programming language around (although doing much better than R, its eternal AI/ML rival). Some of Python’s speed limitations are due to its default implementation, cPython, being single-threaded. That is, cPython doesn’t use more than one hardware thread at a time.</p>\n<p>While you can use threading module built into Python to speed things up, threading only gives you concurrency, not parallelism. It’s good for running multiple tasks that aren’t CPU-dependent but does nothing to speed up multiple tasks that each require a full CPU. </p>\n<p>Python does include a native way to run a Python workload across multiple CPUs. The <em>multiprocessing</em> module spins up multiple copies of the Python interpreter, each on a separate core, and provides primitives for splitting tasks across cores. For years and years, multiprocessing outperformed a bunch of third-party Python libraries invented to facilitate the same. Python’s core maintenance team was good at keeping the focus on the operability and usability of this part of the native built-in framework.</p>\n<p>However, the advent of distributed multi-node computation technologies had an unexpected impact on Python-based multiprocessing technologies. A number of new frameworks arrived that started to be on a par with the standard multiprocessing module, even if run in a single-node configuration.</p>\n<p>Among such newcomers, you can point our two outstanding leaders, namely</p>\n<ul>\n<li>Ray ( <a href=\"https://ray.io/\" target=\"_blank\">https://ray.io/</a> )</li>\n<li>Dask ( <a href=\"https://dask.org/\" target=\"_blank\">https://dask.org/</a> )</li>\n</ul>\n<p>In the sections below, we will compare the multiprocessing capabilities and performance/speed of three implementation of audio feature extractions from the bird call tracks.</p>\n<h1>Audio Feature Extraction Flow</h1>\n<p>As a part of the effort to classify bird calls (songs) in Cornell Birdcall Identification competition, there is a need to extract audio features from the digital audio records of the bird songs. In this project, librosa library is used for audio feature extraction. Since librosa-based audio feature computations are CPU-intensive, parallelizing such computations is the key architectural pattern to build a high-performance data transformation toolset to process the audio files.</p>\n<p>From the conceptual standpoint, the audio feature extraction flow was implemented as follows</p>\n<ol>\n<li>The training set was split by the species</li>\n<li>Every per-species batch of audio files with the bird songs have been processed with librosa as follows</li>\n</ol>\n<ul>\n<li>Features extracted/calculated per the list below</li>\n<li>Every feature that librosa extracts as np.array is further processed to calculate mean across the numeric values in np.array (actually it is everything except BPM variable)</li>\n<li>The final results saved as a pandas dataframe</li>\n<li>The dataframes serialized as CSV files for future use by ML pipeline tools down the road</li>\n</ul>\n<p>Activities per step 2 have been parallelized using three technologies compared in this research. </p>\n<p>When working with Dask and Ray, the output in a Pandas dataframe was consciously preferred over the native Dask and Ray counterparts of Pandsas to maximize the reuse of audio feature extraction code across all of the solutions. </p>\n<p>Data files for just three species ('American Avocet', 'American Bittern', and 'American Crow') have been used in order to run the calculation within the reasonable time frame (as well as not to hit the Kaggle timeout issues when running the experiment on Kaggle).</p>\n<p>The list of audio features extracted with librosa is below</p>\n<ul>\n<li>Spectrogram (decibel-scaled)</li>\n<li>Mel Spectrograms (decibel-scaled)</li>\n<li>Zero-crossing rate</li>\n<li>Harmonics (harmonics are characteristics that represent the sound color)</li>\n<li>Perceptual shock wave (it represents the sound rhythm and emotion of a soundtrack)</li>\n<li>Spectral Centroid (it indicates where the ”center of mass” for a sound is located and is calculated as the weighted mean of the frequencies present in the sound)</li>\n<li>Chroma Frequencies (Chroma features). These are an interesting and powerful representation for music audio in which the entire spectrum is projected onto 12 bins representing the 12 distinct semitones (or chromas) of the musical octave.</li>\n<li>Tempo BPM (beats per minute). Dynamic programming beat tracker.</li>\n<li>Spectral Rolloff. This is a measure of the shape of the signal. It represents the frequency below which a specified percentage of the total spectral energy (e.g. 85 %) lies.</li>\n<li>Spectral flux.</li>\n<li>Spectral Bandwidth. The spectral bandwidth is defined as the width of the band of light at one-half the peak maximum (or full width at half maximum (FWHM) and is represented by the two verticalred lines and λSB on the wavelength axis (in fact, we measure three spectral bandwidth features in this experiment)</li>\n<li>MFCC features (20 MFCC coefficients, 20 MFCC delta coefficients, and 20 MFCC accelerate coefficients)</li>\n</ul>\n<p>Total of 87 audio features has been extracted.</p>\n<p><strong>Note:</strong> the audio feature extraction flow has been designed to calculate tabular features for each audio file provided in the training set. Such an approach will be helpful to try both DL and classical ML algorithms during the feature importance analysis and ML predictions. However, you can find alternative approaches to audio feature extraction where the entire np.arrays are saved as files for futher processing (you can see examples of such a flow in <a href=\"https://github.com/Cocoxili/DCASE2018Task2/blob/master/data_transform.py\" target=\"_blank\">https://github.com/Cocoxili/DCASE2018Task2/blob/master/data_transform.py</a> etc., for instance)</p>\n<h1>Summary and Comparison</h1>\n<p>When running the experiment on a local notebook (Intel Core i7-8750H, 2 CPU @ 2,20 Gh, 8 virtual CPUs over hyper-v, 16 GB RAM, Windows 10, Python 3.7 / Anaconda), the comparative performance has been measured as follows</p>\n<ul>\n<li>Classical multiprocessing with parallelizing across 6 CPUs  - 6 min 17 s</li>\n<li>Parallelized processing with Dask - 7 min 21 s</li>\n</ul>\n<p>It was not possible to run a Ray-based feature extraction workflow on a Windows platform due to the limitations explained in <a href=\"https://github.com/ray-project/ray/issues/631\" target=\"_blank\">https://github.com/ray-project/ray/issues/631</a></p>\n<p>When running the experiment on Kaggle kernels, the indicative times of the parallelized audio features extraction were displayed as follows</p>\n<ul>\n<li>Classical multiprocessing with parallelizing across 6 CPUs (<a href=\"https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-mp/\" target=\"_blank\">https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-mp/</a> ) – 11 min 14 s</li>\n<li>Parallelized processing with Dask (<a href=\"https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-dask/\" target=\"_blank\">https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-dask/</a> ) – 12 min 55 s</li>\n<li>Parallelized processing with Ray (<a href=\"https://www.kaggle.com/gvyshnya/parallel-audio-extraction-with-ray/\" target=\"_blank\">https://www.kaggle.com/gvyshnya/parallel-audio-extraction-with-ray/</a> ) – 30 min 18 s</li>\n</ul>\n<p>So the classical <em>multiprocessing</em> performed a little better then <em>Dask</em> on both the local notebook hardware and Kaggle Kernels. </p>\n<p>Ray did not perform well on Kaggle Kernels (probably, due to the limitations in its setup in the docker image for  Kaggle Python notebooks), despite the promise it shared per other computational experiments (see <a href=\"https://towardsdatascience.com/10x-faster-parallel-python-without-python-multiprocessing-e5017c93cce1\" target=\"_blank\">https://towardsdatascience.com/10x-faster-parallel-python-without-python-multiprocessing-e5017c93cce1</a> ).</p>\n<p><strong>Note:</strong> Dask in its default settings has spinned 4 parallel processes when running in Kaggle. When running on the local computer, it was able to run 8 processes in parallel.</p>",
  "messages": [
    {
      "id": 995584,
      "postDate": "2020-09-02T15:33:42.327Z",
      "content": "<h1>Preface: Curse of Speed Limitations</h1>\n<p>Python has been the standard de facto for the majority of industrial ML/AI solutions. Its extreme programmer-friendliness, along with the convenience and wide range of libraries for data analytics, data processing, data engineering, system integration and ML made it the language of choice of many experts, vendors and corporate Data Science teams.</p>\n<p>At the same time, it isn’t the fastest programming language around (although doing much better than R, its eternal AI/ML rival). Some of Python’s speed limitations are due to its default implementation, cPython, being single-threaded. That is, cPython doesn’t use more than one hardware thread at a time.</p>\n<p>While you can use threading module built into Python to speed things up, threading only gives you concurrency, not parallelism. It’s good for running multiple tasks that aren’t CPU-dependent but does nothing to speed up multiple tasks that each require a full CPU. </p>\n<p>Python does include a native way to run a Python workload across multiple CPUs. The <em>multiprocessing</em> module spins up multiple copies of the Python interpreter, each on a separate core, and provides primitives for splitting tasks across cores. For years and years, multiprocessing outperformed a bunch of third-party Python libraries invented to facilitate the same. Python’s core maintenance team was good at keeping the focus on the operability and usability of this part of the native built-in framework.</p>\n<p>However, the advent of distributed multi-node computation technologies had an unexpected impact on Python-based multiprocessing technologies. A number of new frameworks arrived that started to be on a par with the standard multiprocessing module, even if run in a single-node configuration.</p>\n<p>Among such newcomers, you can point our two outstanding leaders, namely</p>\n<ul>\n<li>Ray ( <a href=\"https://ray.io/\" target=\"_blank\">https://ray.io/</a> )</li>\n<li>Dask ( <a href=\"https://dask.org/\" target=\"_blank\">https://dask.org/</a> )</li>\n</ul>\n<p>In the sections below, we will compare the multiprocessing capabilities and performance/speed of three implementation of audio feature extractions from the bird call tracks.</p>\n<h1>Audio Feature Extraction Flow</h1>\n<p>As a part of the effort to classify bird calls (songs) in Cornell Birdcall Identification competition, there is a need to extract audio features from the digital audio records of the bird songs. In this project, librosa library is used for audio feature extraction. Since librosa-based audio feature computations are CPU-intensive, parallelizing such computations is the key architectural pattern to build a high-performance data transformation toolset to process the audio files.</p>\n<p>From the conceptual standpoint, the audio feature extraction flow was implemented as follows</p>\n<ol>\n<li>The training set was split by the species</li>\n<li>Every per-species batch of audio files with the bird songs have been processed with librosa as follows</li>\n</ol>\n<ul>\n<li>Features extracted/calculated per the list below</li>\n<li>Every feature that librosa extracts as np.array is further processed to calculate mean across the numeric values in np.array (actually it is everything except BPM variable)</li>\n<li>The final results saved as a pandas dataframe</li>\n<li>The dataframes serialized as CSV files for future use by ML pipeline tools down the road</li>\n</ul>\n<p>Activities per step 2 have been parallelized using three technologies compared in this research. </p>\n<p>When working with Dask and Ray, the output in a Pandas dataframe was consciously preferred over the native Dask and Ray counterparts of Pandsas to maximize the reuse of audio feature extraction code across all of the solutions. </p>\n<p>Data files for just three species ('American Avocet', 'American Bittern', and 'American Crow') have been used in order to run the calculation within the reasonable time frame (as well as not to hit the Kaggle timeout issues when running the experiment on Kaggle).</p>\n<p>The list of audio features extracted with librosa is below</p>\n<ul>\n<li>Spectrogram (decibel-scaled)</li>\n<li>Mel Spectrograms (decibel-scaled)</li>\n<li>Zero-crossing rate</li>\n<li>Harmonics (harmonics are characteristics that represent the sound color)</li>\n<li>Perceptual shock wave (it represents the sound rhythm and emotion of a soundtrack)</li>\n<li>Spectral Centroid (it indicates where the ”center of mass” for a sound is located and is calculated as the weighted mean of the frequencies present in the sound)</li>\n<li>Chroma Frequencies (Chroma features). These are an interesting and powerful representation for music audio in which the entire spectrum is projected onto 12 bins representing the 12 distinct semitones (or chromas) of the musical octave.</li>\n<li>Tempo BPM (beats per minute). Dynamic programming beat tracker.</li>\n<li>Spectral Rolloff. This is a measure of the shape of the signal. It represents the frequency below which a specified percentage of the total spectral energy (e.g. 85 %) lies.</li>\n<li>Spectral flux.</li>\n<li>Spectral Bandwidth. The spectral bandwidth is defined as the width of the band of light at one-half the peak maximum (or full width at half maximum (FWHM) and is represented by the two verticalred lines and λSB on the wavelength axis (in fact, we measure three spectral bandwidth features in this experiment)</li>\n<li>MFCC features (20 MFCC coefficients, 20 MFCC delta coefficients, and 20 MFCC accelerate coefficients)</li>\n</ul>\n<p>Total of 87 audio features has been extracted.</p>\n<p><strong>Note:</strong> the audio feature extraction flow has been designed to calculate tabular features for each audio file provided in the training set. Such an approach will be helpful to try both DL and classical ML algorithms during the feature importance analysis and ML predictions. However, you can find alternative approaches to audio feature extraction where the entire np.arrays are saved as files for futher processing (you can see examples of such a flow in <a href=\"https://github.com/Cocoxili/DCASE2018Task2/blob/master/data_transform.py\" target=\"_blank\">https://github.com/Cocoxili/DCASE2018Task2/blob/master/data_transform.py</a> etc., for instance)</p>\n<h1>Summary and Comparison</h1>\n<p>When running the experiment on a local notebook (Intel Core i7-8750H, 2 CPU @ 2,20 Gh, 8 virtual CPUs over hyper-v, 16 GB RAM, Windows 10, Python 3.7 / Anaconda), the comparative performance has been measured as follows</p>\n<ul>\n<li>Classical multiprocessing with parallelizing across 6 CPUs  - 6 min 17 s</li>\n<li>Parallelized processing with Dask - 7 min 21 s</li>\n</ul>\n<p>It was not possible to run a Ray-based feature extraction workflow on a Windows platform due to the limitations explained in <a href=\"https://github.com/ray-project/ray/issues/631\" target=\"_blank\">https://github.com/ray-project/ray/issues/631</a></p>\n<p>When running the experiment on Kaggle kernels, the indicative times of the parallelized audio features extraction were displayed as follows</p>\n<ul>\n<li>Classical multiprocessing with parallelizing across 6 CPUs (<a href=\"https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-mp/\" target=\"_blank\">https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-mp/</a> ) – 11 min 14 s</li>\n<li>Parallelized processing with Dask (<a href=\"https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-dask/\" target=\"_blank\">https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-dask/</a> ) – 12 min 55 s</li>\n<li>Parallelized processing with Ray (<a href=\"https://www.kaggle.com/gvyshnya/parallel-audio-extraction-with-ray/\" target=\"_blank\">https://www.kaggle.com/gvyshnya/parallel-audio-extraction-with-ray/</a> ) – 30 min 18 s</li>\n</ul>\n<p>So the classical <em>multiprocessing</em> performed a little better then <em>Dask</em> on both the local notebook hardware and Kaggle Kernels. </p>\n<p>Ray did not perform well on Kaggle Kernels (probably, due to the limitations in its setup in the docker image for  Kaggle Python notebooks), despite the promise it shared per other computational experiments (see <a href=\"https://towardsdatascience.com/10x-faster-parallel-python-without-python-multiprocessing-e5017c93cce1\" target=\"_blank\">https://towardsdatascience.com/10x-faster-parallel-python-without-python-multiprocessing-e5017c93cce1</a> ).</p>\n<p><strong>Note:</strong> Dask in its default settings has spinned 4 parallel processes when running in Kaggle. When running on the local computer, it was able to run 8 processes in parallel.</p>",
      "rawMarkdown": "# Preface: Curse of Speed Limitations\n\nPython has been the standard de facto for the majority of industrial ML/AI solutions. Its extreme programmer-friendliness, along with the convenience and wide range of libraries for data analytics, data processing, data engineering, system integration and ML made it the language of choice of many experts, vendors and corporate Data Science teams.\n\nAt the same time, it isn’t the fastest programming language around (although doing much better than R, its eternal AI/ML rival). Some of Python’s speed limitations are due to its default implementation, cPython, being single-threaded. That is, cPython doesn’t use more than one hardware thread at a time.\n\nWhile you can use threading module built into Python to speed things up, threading only gives you concurrency, not parallelism. It’s good for running multiple tasks that aren’t CPU-dependent but does nothing to speed up multiple tasks that each require a full CPU. \n\nPython does include a native way to run a Python workload across multiple CPUs. The *multiprocessing* module spins up multiple copies of the Python interpreter, each on a separate core, and provides primitives for splitting tasks across cores. For years and years, multiprocessing outperformed a bunch of third-party Python libraries invented to facilitate the same. Python’s core maintenance team was good at keeping the focus on the operability and usability of this part of the native built-in framework.\n\nHowever, the advent of distributed multi-node computation technologies had an unexpected impact on Python-based multiprocessing technologies. A number of new frameworks arrived that started to be on a par with the standard multiprocessing module, even if run in a single-node configuration.\n\nAmong such newcomers, you can point our two outstanding leaders, namely\n\n* Ray ( https://ray.io/ )\n* Dask ( https://dask.org/ )\n\nIn the sections below, we will compare the multiprocessing capabilities and performance/speed of three implementation of audio feature extractions from the bird call tracks.\n\n# Audio Feature Extraction Flow\n\nAs a part of the effort to classify bird calls (songs) in Cornell Birdcall Identification competition, there is a need to extract audio features from the digital audio records of the bird songs. In this project, librosa library is used for audio feature extraction. Since librosa-based audio feature computations are CPU-intensive, parallelizing such computations is the key architectural pattern to build a high-performance data transformation toolset to process the audio files.\n\nFrom the conceptual standpoint, the audio feature extraction flow was implemented as follows\n\n1. The training set was split by the species\n2. Every per-species batch of audio files with the bird songs have been processed with librosa as follows\n- Features extracted/calculated per the list below\n- Every feature that librosa extracts as np.array is further processed to calculate mean across the numeric values in np.array (actually it is everything except BPM variable)\n- The final results saved as a pandas dataframe\n- The dataframes serialized as CSV files for future use by ML pipeline tools down the road\n\nActivities per step 2 have been parallelized using three technologies compared in this research. \n\nWhen working with Dask and Ray, the output in a Pandas dataframe was consciously preferred over the native Dask and Ray counterparts of Pandsas to maximize the reuse of audio feature extraction code across all of the solutions. \n\nData files for just three species ('American Avocet', 'American Bittern', and 'American Crow') have been used in order to run the calculation within the reasonable time frame (as well as not to hit the Kaggle timeout issues when running the experiment on Kaggle).\n\nThe list of audio features extracted with librosa is below\n\n- Spectrogram (decibel-scaled)\n- Mel Spectrograms (decibel-scaled)\n- Zero-crossing rate\n- Harmonics (harmonics are characteristics that represent the sound color)\n- Perceptual shock wave (it represents the sound rhythm and emotion of a soundtrack)\n- Spectral Centroid (it indicates where the ”center of mass” for a sound is located and is calculated as the weighted mean of the frequencies present in the sound)\n- Chroma Frequencies (Chroma features). These are an interesting and powerful representation for music audio in which the entire spectrum is projected onto 12 bins representing the 12 distinct semitones (or chromas) of the musical octave.\n- Tempo BPM (beats per minute). Dynamic programming beat tracker.\n- Spectral Rolloff. This is a measure of the shape of the signal. It represents the frequency below which a specified percentage of the total spectral energy (e.g. 85 %) lies.\n- Spectral flux.\n- Spectral Bandwidth. The spectral bandwidth is defined as the width of the band of light at one-half the peak maximum (or full width at half maximum (FWHM) and is represented by the two verticalred lines and λSB on the wavelength axis (in fact, we measure three spectral bandwidth features in this experiment)\n- MFCC features (20 MFCC coefficients, 20 MFCC delta coefficients, and 20 MFCC accelerate coefficients)\n\nTotal of 87 audio features has been extracted.\n\n**Note:** the audio feature extraction flow has been designed to calculate tabular features for each audio file provided in the training set. Such an approach will be helpful to try both DL and classical ML algorithms during the feature importance analysis and ML predictions. However, you can find alternative approaches to audio feature extraction where the entire np.arrays are saved as files for futher processing (you can see examples of such a flow in https://github.com/Cocoxili/DCASE2018Task2/blob/master/data_transform.py etc., for instance)\n\n# Summary and Comparison\n\nWhen running the experiment on a local notebook (Intel Core i7-8750H, 2 CPU @ 2,20 Gh, 8 virtual CPUs over hyper-v, 16 GB RAM, Windows 10, Python 3.7 / Anaconda), the comparative performance has been measured as follows\n\n- Classical multiprocessing with parallelizing across 6 CPUs  - 6 min 17 s\n- Parallelized processing with Dask - 7 min 21 s\n\nIt was not possible to run a Ray-based feature extraction workflow on a Windows platform due to the limitations explained in https://github.com/ray-project/ray/issues/631\n\nWhen running the experiment on Kaggle kernels, the indicative times of the parallelized audio features extraction were displayed as follows\n\n- Classical multiprocessing with parallelizing across 6 CPUs (https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-mp/ ) – 11 min 14 s\n- Parallelized processing with Dask (https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-dask/ ) – 12 min 55 s\n- Parallelized processing with Ray (https://www.kaggle.com/gvyshnya/parallel-audio-extraction-with-ray/ ) – 30 min 18 s\n\nSo the classical *multiprocessing* performed a little better then *Dask* on both the local notebook hardware and Kaggle Kernels. \n\nRay did not perform well on Kaggle Kernels (probably, due to the limitations in its setup in the docker image for  Kaggle Python notebooks), despite the promise it shared per other computational experiments (see https://towardsdatascience.com/10x-faster-parallel-python-without-python-multiprocessing-e5017c93cce1 ).\n\n**Note:** Dask in its default settings has spinned 4 parallel processes when running in Kaggle. When running on the local computer, it was able to run 8 processes in parallel.\n\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "995584": "# Preface: Curse of Speed Limitations\n\nPython has been the standard de facto for the majority of industrial ML/AI solutions. Its extreme programmer-friendliness, along with the convenience and wide range of libraries for data analytics, data processing, data engineering, system integration and ML made it the language of choice of many experts, vendors and corporate Data Science teams.\n\nAt the same time, it isn’t the fastest programming language around (although doing much better than R, its eternal AI/ML rival). Some of Python’s speed limitations are due to its default implementation, cPython, being single-threaded. That is, cPython doesn’t use more than one hardware thread at a time.\n\nWhile you can use threading module built into Python to speed things up, threading only gives you concurrency, not parallelism. It’s good for running multiple tasks that aren’t CPU-dependent but does nothing to speed up multiple tasks that each require a full CPU. \n\nPython does include a native way to run a Python workload across multiple CPUs. The *multiprocessing* module spins up multiple copies of the Python interpreter, each on a separate core, and provides primitives for splitting tasks across cores. For years and years, multiprocessing outperformed a bunch of third-party Python libraries invented to facilitate the same. Python’s core maintenance team was good at keeping the focus on the operability and usability of this part of the native built-in framework.\n\nHowever, the advent of distributed multi-node computation technologies had an unexpected impact on Python-based multiprocessing technologies. A number of new frameworks arrived that started to be on a par with the standard multiprocessing module, even if run in a single-node configuration.\n\nAmong such newcomers, you can point our two outstanding leaders, namely\n\n* Ray ( https://ray.io/ )\n* Dask ( https://dask.org/ )\n\nIn the sections below, we will compare the multiprocessing capabilities and performance/speed of three implementation of audio feature extractions from the bird call tracks.\n\n# Audio Feature Extraction Flow\n\nAs a part of the effort to classify bird calls (songs) in Cornell Birdcall Identification competition, there is a need to extract audio features from the digital audio records of the bird songs. In this project, librosa library is used for audio feature extraction. Since librosa-based audio feature computations are CPU-intensive, parallelizing such computations is the key architectural pattern to build a high-performance data transformation toolset to process the audio files.\n\nFrom the conceptual standpoint, the audio feature extraction flow was implemented as follows\n\n1. The training set was split by the species\n2. Every per-species batch of audio files with the bird songs have been processed with librosa as follows\n- Features extracted/calculated per the list below\n- Every feature that librosa extracts as np.array is further processed to calculate mean across the numeric values in np.array (actually it is everything except BPM variable)\n- The final results saved as a pandas dataframe\n- The dataframes serialized as CSV files for future use by ML pipeline tools down the road\n\nActivities per step 2 have been parallelized using three technologies compared in this research. \n\nWhen working with Dask and Ray, the output in a Pandas dataframe was consciously preferred over the native Dask and Ray counterparts of Pandsas to maximize the reuse of audio feature extraction code across all of the solutions. \n\nData files for just three species ('American Avocet', 'American Bittern', and 'American Crow') have been used in order to run the calculation within the reasonable time frame (as well as not to hit the Kaggle timeout issues when running the experiment on Kaggle).\n\nThe list of audio features extracted with librosa is below\n\n- Spectrogram (decibel-scaled)\n- Mel Spectrograms (decibel-scaled)\n- Zero-crossing rate\n- Harmonics (harmonics are characteristics that represent the sound color)\n- Perceptual shock wave (it represents the sound rhythm and emotion of a soundtrack)\n- Spectral Centroid (it indicates where the ”center of mass” for a sound is located and is calculated as the weighted mean of the frequencies present in the sound)\n- Chroma Frequencies (Chroma features). These are an interesting and powerful representation for music audio in which the entire spectrum is projected onto 12 bins representing the 12 distinct semitones (or chromas) of the musical octave.\n- Tempo BPM (beats per minute). Dynamic programming beat tracker.\n- Spectral Rolloff. This is a measure of the shape of the signal. It represents the frequency below which a specified percentage of the total spectral energy (e.g. 85 %) lies.\n- Spectral flux.\n- Spectral Bandwidth. The spectral bandwidth is defined as the width of the band of light at one-half the peak maximum (or full width at half maximum (FWHM) and is represented by the two verticalred lines and λSB on the wavelength axis (in fact, we measure three spectral bandwidth features in this experiment)\n- MFCC features (20 MFCC coefficients, 20 MFCC delta coefficients, and 20 MFCC accelerate coefficients)\n\nTotal of 87 audio features has been extracted.\n\n**Note:** the audio feature extraction flow has been designed to calculate tabular features for each audio file provided in the training set. Such an approach will be helpful to try both DL and classical ML algorithms during the feature importance analysis and ML predictions. However, you can find alternative approaches to audio feature extraction where the entire np.arrays are saved as files for futher processing (you can see examples of such a flow in https://github.com/Cocoxili/DCASE2018Task2/blob/master/data_transform.py etc., for instance)\n\n# Summary and Comparison\n\nWhen running the experiment on a local notebook (Intel Core i7-8750H, 2 CPU @ 2,20 Gh, 8 virtual CPUs over hyper-v, 16 GB RAM, Windows 10, Python 3.7 / Anaconda), the comparative performance has been measured as follows\n\n- Classical multiprocessing with parallelizing across 6 CPUs  - 6 min 17 s\n- Parallelized processing with Dask - 7 min 21 s\n\nIt was not possible to run a Ray-based feature extraction workflow on a Windows platform due to the limitations explained in https://github.com/ray-project/ray/issues/631\n\nWhen running the experiment on Kaggle kernels, the indicative times of the parallelized audio features extraction were displayed as follows\n\n- Classical multiprocessing with parallelizing across 6 CPUs (https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-mp/ ) – 11 min 14 s\n- Parallelized processing with Dask (https://www.kaggle.com/gvyshnya/parallel-audio-feature-extraction-with-dask/ ) – 12 min 55 s\n- Parallelized processing with Ray (https://www.kaggle.com/gvyshnya/parallel-audio-extraction-with-ray/ ) – 30 min 18 s\n\nSo the classical *multiprocessing* performed a little better then *Dask* on both the local notebook hardware and Kaggle Kernels. \n\nRay did not perform well on Kaggle Kernels (probably, due to the limitations in its setup in the docker image for  Kaggle Python notebooks), despite the promise it shared per other computational experiments (see https://towardsdatascience.com/10x-faster-parallel-python-without-python-multiprocessing-e5017c93cce1 ).\n\n**Note:** Dask in its default settings has spinned 4 parallel processes when running in Kaggle. When running on the local computer, it was able to run 8 processes in parallel.\n\n"
  }
}