{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":8900,"databundleVersionId":862232,"sourceType":"competition"},{"sourceId":6391,"sourceType":"datasetVersion","datasetId":4114}],"dockerImageVersionId":29271,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Part 2 - Extracting Audio Features\n\nEu Jin Lok\n<br>Kernel post for __[Speech Accent Archive](https://www.kaggle.com/rtatman/speech-accent-archive)__ on Kaggle\n<br>20 January 2019\n\n\n## Introduction \n#### To understand the various features we can extract from audio, and use it to predict gender and accents\n\n\nIn part 2 of this series, I'll introduce the various types of features from audio and we can see how these features differ when we compare between genders and accent types. I'm hoping one of these features will be distinctive enough that we could use it for predictive modelling in the next part. \n\nI would also like to acknowledge 2 people whom I've learnt alot from: \n- __[Jayesh Saita](https://towardsdatascience.com/ok-google-how-to-do-speech-recognition-f77b5d7cbe0b)__, whose blog I really liked and helped me understand audio features and MFCCs in the early stages of my journey   \n- __[Zafarullah Mahmood](https://www.kaggle.com/fizzbuzz/beginner-s-guide-to-audio-data)__, whose Kaggle Kernel basically became the canvas for all audio work that I do at work and outside work. Please do check out his kernels whom I've drawn lots of inspiration from! \n\nThis part 2 kernel will cover quite a few items so I'll provide a brief agenda: \n- [Core concepts in Audio](#core)\n- [Time domain features](#time) \n    - [1. Audio wave](#time)\n- [Frequency domain features](#mfcc)\n    - [2. MFCC](#mfcc)\n    - [3. Log Mel-spectogram](#melspec)\n    - [4. Harmonic-percussive source separation (HPSS)](#hpss)\n    - [5. Chroma](#chroma)\n- [Final thoughts](#final)\n\nAgain, thanks to the awesome Kaggle community, and the broader data science community. Without further ado, lets begin!","metadata":{"_uuid":"a8f691d6-bf3c-4375-a295-baaeb01bb61b","_cell_guid":"f4459c84-dd7e-4911-8943-3c2b84cc8666","nbpresent":{"id":"42344751-df85-4fa9-af08-b4bf50b41e05"},"trusted":true}},{"cell_type":"code","source":"import pandas as pd       \nimport os \nimport math \nimport numpy as np\nimport matplotlib.pyplot as plt  \nimport IPython.display as ipd  # To play sound in the notebook\nimport librosa\nimport librosa.display\n# !apt-get -y install ffmpeg\n!conda install -c conda-forge -y ffmpeg\n# !pip install ffmpeg\n# os.chdir(\"/kaggle/input/freesound-audio-tagging/audio_train\")\n#os.getcwd()\nos.chdir(\"/kaggle/input/speech-accent-archive/recordings\")\n# print(os.listdir(\"/kaggle/input/freesound-audio-tagging/audio_train/audio_train/\"))\n# print(os.listdir(\"recordings\"))","metadata":{"_uuid":"99e9ca45-c54e-4c36-92af-4d402adb2db6","_cell_guid":"d3895d77-89ac-4d49-ad77-e276bef0c96c","collapsed":false,"nbpresent":{"id":"6b931974-2fc8-4fdc-9f32-a93d21ec08ef"},"_kg_hide-output":false,"_kg_hide-input":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:10:04.461269Z","iopub.execute_input":"2023-12-15T03:10:04.461603Z","iopub.status.idle":"2023-12-15T03:14:31.657562Z","shell.execute_reply.started":"2023-12-15T03:10:04.461537Z","shell.execute_reply":"2023-12-15T03:14:31.656499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"------------------------------\nAfter loading the libraries and setting our directory path, we're going to use the two Kentucky accent speakers that we heard from __[Part 1](https://www.kaggle.com/ejlok1/part-1-data-exploration-deep-learning-series)__. Here's another quick replay of how they sound if you've forgotten:","metadata":{"_uuid":"0bd0ba43-846b-43ed-bd05-c24e60d22cb3","_cell_guid":"1c36e067-1652-4c8f-a91e-c1d80341189d","trusted":true}},{"cell_type":"code","source":"# Play female from Kentucky\nfname_f = 'recordings/' + 'english385.mp3'   \nipd.Audio(fname_f)","metadata":{"_uuid":"4502615a-9b86-41b0-bafe-89fff8c7e159","_cell_guid":"9678f87a-11eb-411e-9f73-9ea4a9461d79","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:31.66011Z","iopub.execute_input":"2023-12-15T03:14:31.660406Z","iopub.status.idle":"2023-12-15T03:14:31.698478Z","shell.execute_reply.started":"2023-12-15T03:14:31.660326Z","shell.execute_reply":"2023-12-15T03:14:31.697443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Play male from Kentucky\nfname_m = 'recordings/' + 'english381.mp3'\nipd.Audio(fname_m)","metadata":{"_uuid":"d4fc3dae-908c-4610-b0ea-9781b0989859","_cell_guid":"bcf06a2b-dc6c-4e81-9b60-6edae906ecf9","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:31.699988Z","iopub.execute_input":"2023-12-15T03:14:31.700292Z","iopub.status.idle":"2023-12-15T03:14:32.968374Z","shell.execute_reply.started":"2023-12-15T03:14:31.700241Z","shell.execute_reply":"2023-12-15T03:14:32.967629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"core\"></a>\n##  Core concepts in Audio","metadata":{"_uuid":"4abb2929-6214-44fc-883d-4e81752fad05","_cell_guid":"28ea2b6e-d1ac-48b2-9718-f91c127ff62a","trusted":true}},{"cell_type":"markdown","source":"Before diving straight to the meaty bits, I'm going to quickly introduce a few key concepts to anyone who's new to audio \n- Duration\n- sampling rate\n- Amplitude \n- Frequency\n\n\n__Duration__ is the length of the audio call in terms of time. \n\n__Sampling rate__ is the number of samples of audio per second, measured in Hz / KHz. This is similar to resolution in images, where the higher the resolution (or more pixels), the clearer it is. A full sampling rate is 44100 Hz (44.1 KHz) but, you don't always need to have it in at 'High Fidelity' format. A more reasonable sampling rate is 22050 Hz (22 KHz), because that is the audible sound to a human. \n\n__Amplitude__ is the fluctuation of the soud wave. The shorter and more frequent the waves are, the higher the pitch or frequency. Plotting the audio by time against amplitute is probably the most intuitive way of understanding the audio. However, its not the only way to represent the data or used as feature. Another equally good way of doing this is to look at it by the frequency domain, which is a nice segway to our final concept. \n\n__Frequency__, the best way of understanding it is through visualising it. Imagining the audio in terms of time is probably the most intutive way of thinking about it. But frequency, although not as intuitive, is actually much more efficient as signal in the frequency domain requires much less computational space for storage. Below is a nice visualisation of how to differentiate the Time vs Frequency domain","metadata":{"_uuid":"ba4eaa66-7a58-4774-aedb-46d335c067f9","_cell_guid":"3a2dbd84-5023-497a-bf0e-380cce57fa81","trusted":true}},{"cell_type":"markdown","source":"![Audio%20wave.png](attachment:Audio%20wave.png)\nThe time vs frequency domain sourced from __[here](https://docs.google.com/presentation/d/1zzgNu_HbKL2iPkHS8-qhtDV20QfWt9lC3ZwPVZo8Rw0/pub?start=false&loop=false&delayms=3000&slide=id.g5a7a9806e_0_84)__","metadata":{"_uuid":"4adad233-f7a6-4552-99f6-95eaeff7e0b5","_cell_guid":"8a7702a5-a5a8-4c2f-a6cf-a5e40fdcb8e1","trusted":true}},{"cell_type":"markdown","source":"So what we're going to do is use the same audio file but plot the 2 different sampling rates to see how they differ, one at the full 44100 Hz and the other at 6000 Hz. We're also going to down sample them and chop the audio file at 5 seconds just so the differences are more perceptible","metadata":{"_uuid":"449da528-8b93-4200-9352-2ed6419fb327","_cell_guid":"8f5d8619-b619-42f8-8995-0b99dbcb50d1","trusted":true}},{"cell_type":"markdown","source":"<a id=\"time\"></a>\n## 1. Audio wave\n\nAudio wave is our first feature. This is of course the first most common form of the data, and is the only time domain feature that I am aware off. Lets use this opportunity to plot the two audio files again at their most native format, which is the wave form. To start off, I'm going to plot them in 3 different sampling rates, at 44kHz, 6kHz and 1000kHz and see how they differ, using the female version of the audio (_fname = english385.mp3_)\n\nNote that I'm only taking the first 5 seconds of audio for illustration purposes.","metadata":{"_uuid":"bd48b090-accc-4ec8-ac0a-e646d3c05e60","_cell_guid":"f50dc068-f97b-446d-b73b-98cacc9e99d6","trusted":true}},{"cell_type":"code","source":"# The full 'high fidelity' sampling rate of 44k \nSAMPLE_RATE = 44100\nfname_f = 'recordings/' + 'english385.mp3' \ny, sr = librosa.load(fname_f, sr=SAMPLE_RATE, duration = 5)\n\nplt.figure(figsize=(12, 3))\nplt.figure()\nlibrosa.display.waveplot(y, sr=sr)\nplt.title('Audio sampled at 44100 hrz')","metadata":{"_uuid":"8bf294d0-72fb-4e1d-b124-1cdc953b0f34","_cell_guid":"17bf076a-2d7a-4690-ab4d-6d2f6cedf761","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:32.969843Z","iopub.execute_input":"2023-12-15T03:14:32.970112Z","iopub.status.idle":"2023-12-15T03:14:33.456993Z","shell.execute_reply.started":"2023-12-15T03:14:32.970061Z","shell.execute_reply":"2023-12-15T03:14:33.455902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The very 'low fidelity' sampling rate of 6k \nSAMPLE_RATE = 6000\nfname_f = 'recordings/' + 'english385.mp3' \ny, sr = librosa.load(fname_f, sr=SAMPLE_RATE, duration = 5)\n\nplt.figure(figsize=(12, 3))\nplt.figure()\nlibrosa.display.waveplot(y, sr=sr)\nplt.title('Audio sampled at 6000 hrz')","metadata":{"_uuid":"daed13ba-1032-493c-91c4-6584f77eaf93","_cell_guid":"245b0a0b-52da-4878-8ec6-3170dc2daa64","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:33.461905Z","iopub.execute_input":"2023-12-15T03:14:33.462565Z","iopub.status.idle":"2023-12-15T03:14:34.883429Z","shell.execute_reply.started":"2023-12-15T03:14:33.46249Z","shell.execute_reply":"2023-12-15T03:14:34.882062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The very very 'low fidelity' sampling rate of 1k \nSAMPLE_RATE = 1000\nfname_f = 'recordings/' + 'english385.mp3' \ny, sr = librosa.core.load(fname_f, sr=SAMPLE_RATE, duration = 5)\n\nplt.figure(figsize=(12, 3))\nplt.figure()\nlibrosa.display.waveplot(y, sr=sr)\nplt.title('Audio sampled at 1000 hrz')","metadata":{"_uuid":"f7ae32ec-b5c8-42b1-a26c-e1be661c380c","_cell_guid":"5ac46bb6-38c4-46bf-8b56-4b2034e1bb1a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:34.886119Z","iopub.execute_input":"2023-12-15T03:14:34.886569Z","iopub.status.idle":"2023-12-15T03:14:35.58504Z","shell.execute_reply.started":"2023-12-15T03:14:34.886483Z","shell.execute_reply":"2023-12-15T03:14:35.584207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note the minor differences between the 44100 Hz and the 6000 Hz? Things start to change once you go down to sampling rate of 1000 Hz, and audio becomes more blurry. Generally 21500 Hz sampling rate is the most common as previously mentioned. And the difference with 44 KHz is very minor. For now, we're going to use a sampling rate of 6000 Hz, to compare the male and female speakers from Kentucky. \n\nSince we've plotted the female version, lets hear the male version now (_fname = english381.mp3_)... at 6000 Hz and chopped at 5 seconds. You can see has some distinctive pattern over the female equivalent","metadata":{"_uuid":"a0781b5d-7067-4133-8ede-a3f102214f80","_cell_guid":"3b7605a0-146c-4d84-b4e6-a0cc299af090","trusted":true}},{"cell_type":"code","source":"# The 'low fidelity' sampling rate of 6k \nSAMPLE_RATE = 6000\nfname_m = 'recordings/' + 'english381.mp3' \ny, sr = librosa.load(fname_m, sr=SAMPLE_RATE, duration = 5)\n\nplt.figure(figsize=(12, 3))\nplt.figure()\nlibrosa.display.waveplot(y, sr=sr)\nplt.title('Audio sampled at 6000 hrz')","metadata":{"_uuid":"8d43fddc-67fc-45fd-9acd-5b37f62a7806","_cell_guid":"ad730b50-fdfb-49a0-83e8-196592cf6abc","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:35.586528Z","iopub.execute_input":"2023-12-15T03:14:35.586972Z","iopub.status.idle":"2023-12-15T03:14:36.375556Z","shell.execute_reply.started":"2023-12-15T03:14:35.586921Z","shell.execute_reply":"2023-12-15T03:14:36.374625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There's some distinctive differences in the wave pattern between male and female but we don't really know what information it is capturing. It could be capturing the supposedly higher pitch voice that female has, or it could just be the different accent pronoucation between the 2 speakers. And that's why we use machine learning right? \n\nI suppose if there's one thing that's clear between the 2 audio file is that, the male speaker seems to have a more consistent pitch overtime versus the female counter part, whom we can clearly see an obvious huge spike in amplitude.","metadata":{"_uuid":"56e13ad4-e9c2-49e5-a733-113a167c58c7","_cell_guid":"62867453-9c19-4ab4-b1bf-ed73cfc948d4","trusted":true}},{"cell_type":"markdown","source":"<a id=\"mfcc\"></a>\n## 2. MFCC\nNow moving into the Frequency domain feature, this is my favourite which is MFCC, short for Mel-Frequency Cepstral Coefficient. MFCC is a sentence, is a \"representation\" of the vocal tract that produces the sound. Think of it like an x-ray of your mouth. The processing steps to get MFCC is quite lengthy but  in my simple mind, that is how I understand it, or explain to my wife. A more indepth and technical explanation can be found __[here](http://practicalcryptography.com/miscellaneous/machine-learning/guide-mel-frequency-cepstral-coefficients-mfccs/)__. \n\nSo we're now going to compare males and female voices via the MFCC coefficient plots","metadata":{"_uuid":"7044ec97-bab7-49d3-93ad-3de0e1ede711","_cell_guid":"856cb0b1-7c7e-4b0f-8e28-e31025847cf6","trusted":true}},{"cell_type":"code","source":"# MFCC for female \nSAMPLE_RATE = 22050\nfname_f = 'recordings/' + 'english385.mp3'  \ny, sr = librosa.load(fname_f, sr=SAMPLE_RATE, duration = 5) # Chop audio at 5 secs... \nmfcc = librosa.feature.mfcc(y=y, sr=SAMPLE_RATE, n_mfcc = 5) # 5 MFCC components\n\nplt.figure(figsize=(12, 6))\nplt.subplot(3,1,1)\nlibrosa.display.specshow(mfcc)\nplt.ylabel('MFCC')\nplt.colorbar()","metadata":{"_uuid":"ac139930-d4f1-4347-b688-57ae7d3725f1","_cell_guid":"78ea22c7-3546-488f-a90d-c0f3d238a450","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:36.377644Z","iopub.execute_input":"2023-12-15T03:14:36.378092Z","iopub.status.idle":"2023-12-15T03:14:36.967856Z","shell.execute_reply.started":"2023-12-15T03:14:36.378013Z","shell.execute_reply":"2023-12-15T03:14:36.966766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# MFCC for male  \nSAMPLE_RATE = 22050\nfname_m = 'recordings/' + 'english381.mp3'  \ny, sr = librosa.load(fname_m, sr=SAMPLE_RATE, duration = 5)\nmfcc = librosa.feature.mfcc(y=y, sr=SAMPLE_RATE, n_mfcc = 5)\n\nplt.figure(figsize=(12, 6))\nplt.subplot(3,1,1)\nlibrosa.display.specshow(mfcc)\nplt.ylabel('MFCC')\nplt.colorbar()","metadata":{"_uuid":"bf24404f-5e56-475b-bd96-ce3adc8f56a2","_cell_guid":"201fd715-8c6a-4a1e-8185-a098d1a4045b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:36.96968Z","iopub.execute_input":"2023-12-15T03:14:36.97017Z","iopub.status.idle":"2023-12-15T03:14:37.553288Z","shell.execute_reply.started":"2023-12-15T03:14:36.970113Z","shell.execute_reply":"2023-12-15T03:14:37.55209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notice the difference? Seems like there could be something here. Generally in practice one would use not just the entire duration of the audio, but also set a higher number of MFCC components, usually between 20 and 40. Here I used 5 MFCC components just so its easier to visualise the difference. We will use the optimal settings later in the actual modelling component. In my limited experience so far with audio, I have gotten the most success in using MFCC as the feature choice for supervise learning","metadata":{"_uuid":"aecc5b05-343a-48ca-91ce-3d365d04dec6","_cell_guid":"e466339b-078d-4e5c-826d-aa9063c78742","trusted":true}},{"cell_type":"markdown","source":"<a id=\"melspec\"></a>\n## 3. Log Mel-spectogram\n\nOther than MFCC, the next most popular, if not the most popular audio feature is the Mel-spectogram. A time by frequency representation of the audio wave form is called a spectogram. That spectogram is then mapped to the Mel-scales thus giving us the Mel-spectogram. But because human perception of sound intensity is logarithmic in nature, the log form of the Mel-spectogram is the better one in theory. Thou I admit I haven't research this in alot of depth. \n\nLike MFCC, Log Mel-spectogram are well known to be discriminative features in audio. So once again, lets see how the Log Mel-spectogram varies between male and female.","metadata":{"_uuid":"e82984e8-93b3-47aa-89ef-fc5f23add116","_cell_guid":"1f640e4a-635e-49cc-b2cd-ebb65b1bc3a7","trusted":true}},{"cell_type":"code","source":"# Log Mel-spectogram for female \nSAMPLE_RATE = 22050\nfname_f = 'recordings/' + 'english385.mp3'  \ny, sr = librosa.load(fname_f, sr=SAMPLE_RATE, duration = 5) # Chop audio at 5 secs... \nmelspec = librosa.feature.melspectrogram(y, sr=sr, n_mels=128)\n\n# Convert to log scale (dB). We'll use the peak power (max) as reference.\nlog_S = librosa.amplitude_to_db(melspec)\n\n# Display the log mel spectrogram\nplt.figure(figsize=(12,4))\nlibrosa.display.specshow(log_S, sr=sr, x_axis='time', y_axis='mel')\nplt.title('Log mel spectrogram for female')\nplt.colorbar(format='%+02.0f dB')\nplt.tight_layout()","metadata":{"_uuid":"f8f3bb33-3f49-4ef2-b404-2067bbb07a2e","_cell_guid":"06144c55-44cb-4dce-93bf-83a70f5956cd","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:37.555381Z","iopub.execute_input":"2023-12-15T03:14:37.556009Z","iopub.status.idle":"2023-12-15T03:14:39.32958Z","shell.execute_reply.started":"2023-12-15T03:14:37.555932Z","shell.execute_reply":"2023-12-15T03:14:39.328292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Log Mel-spectogram for male \nSAMPLE_RATE = 22050\nfname_m = 'recordings/' + 'english381.mp3'  \ny, sr = librosa.load(fname_m, sr=SAMPLE_RATE, duration = 5) # Chop audio at 5 secs... \nmelspec = librosa.feature.melspectrogram(y, sr=sr, n_mels=128)\n\n# Convert to log scale (dB). We'll use the peak power (max) as reference.\nlog_S = librosa.amplitude_to_db(melspec)\n\n# Display the log mel spectrogram\nplt.figure(figsize=(12,4))\nlibrosa.display.specshow(log_S, sr=sr, x_axis='time', y_axis='mel')\nplt.title('Log mel spectrogram for male')\nplt.colorbar(format='%+02.0f dB')\nplt.tight_layout()","metadata":{"_uuid":"3b8854d1-69e3-4e7c-ac9f-45b0a026c5be","_cell_guid":"300a8084-0ec7-40a7-af2d-a217d7cc4573","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:39.331915Z","iopub.execute_input":"2023-12-15T03:14:39.33267Z","iopub.status.idle":"2023-12-15T03:14:40.692827Z","shell.execute_reply.started":"2023-12-15T03:14:39.332592Z","shell.execute_reply":"2023-12-15T03:14:40.691494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The difference may not be as obvious but is definitely noticable after careful inspection. In particular, the female version  reaches a higher frequency compared to the male counterpart. \n\nI've also heard about a Mel Power spectogram as another feature, thou I haven't had any experience using it, hence I really can't comment how good it is compared to a Log Mel spectogram. Perhaps someone else who's a real expert can elaborate further. Leave the comments down below!","metadata":{"_uuid":"5dee94c5-cde6-407e-8efc-c4c3b91c65c6","_cell_guid":"1d6da86a-65f9-40c6-8285-134d89f76cc1","trusted":true}},{"cell_type":"markdown","source":"<a id=\"hpss\"></a>\n## 4. Harmonic-percussive source separation\n\nIts a mouthful to pronounce, but it's really easy to remember and I rarely use its short form of HPSS. Despite the complex description, it is essentially quite literal in what it means. This feature seperates the harmonic and the percussive source of an audio file. There's a really good notebook that explains how this works __[here](https://musicinformationretrieval.com/hpss.html)__\n\nLets take one of the audio files and __seperate its harmonic and perussive source__ (see told you its easy to remember), then listen to it too see how different they are...","metadata":{"_uuid":"73f7562e-627e-4bcb-888e-372f7fc7cd0b","_cell_guid":"cd3361c3-6bbf-4b0a-ba92-66672483bba5","trusted":true}},{"cell_type":"code","source":"SAMPLE_RATE = 22050\nfname_f = 'recordings/' + 'english385.mp3'  \ny, sr = librosa.load(fname_f, sr=SAMPLE_RATE, duration = 5) \ny_harmonic, y_percussive = librosa.effects.hpss(y)\n\nipd.Audio(y_harmonic, rate=sr)","metadata":{"_uuid":"0ab2cbb9-b40f-41ce-956b-f2539cf69e75","_cell_guid":"412e5b6c-f2e7-4636-b19f-f7d7874c85d9","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:40.694885Z","iopub.execute_input":"2023-12-15T03:14:40.695323Z","iopub.status.idle":"2023-12-15T03:14:41.972041Z","shell.execute_reply.started":"2023-12-15T03:14:40.695247Z","shell.execute_reply":"2023-12-15T03:14:41.971255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ipd.Audio(y_percussive, rate=sr)","metadata":{"_uuid":"1ba1b853-a7aa-421b-8aae-2af9009bf771","_cell_guid":"16e228c3-b065-49c3-8a01-a335ddbabc95","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:41.973404Z","iopub.execute_input":"2023-12-15T03:14:41.973696Z","iopub.status.idle":"2023-12-15T03:14:42.030399Z","shell.execute_reply.started":"2023-12-15T03:14:41.973642Z","shell.execute_reply":"2023-12-15T03:14:42.029666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's a pity that playing these audio files aren't the best way to really get a feel for how different the harmonic and percussive source sounds like. Generally its really obvious for music as it seperates out the beats from drums. In either case, maybe converting the source to a Mel spectogram may shed abit more light...","metadata":{"_uuid":"0eb66dd2-4e12-49a2-b7ab-5a0d12274521","_cell_guid":"765de0b7-0846-405f-ae55-8aa0bdb01b96","trusted":true}},{"cell_type":"code","source":"# harmonic \nmelspec = librosa.feature.melspectrogram(y_harmonic, sr=sr, n_mels=128)\nlog_h = librosa.amplitude_to_db(melspec)\n\n# percussive\nmelspec = librosa.feature.melspectrogram(y_percussive, sr=sr, n_mels=128)\nlog_p = librosa.amplitude_to_db(melspec)\n\n# Display the log mel spectrogram of both harmonic and percussive\nplt.figure(figsize=(12,6))\n\nplt.subplot(2,1,1)\nlibrosa.display.specshow(log_h, sr=sr, x_axis='time', y_axis='mel')\nplt.title('Log mel spectrogram for female harmonic')\nplt.colorbar(format='%+02.0f dB')\n\nplt.subplot(2,1,2)\nlibrosa.display.specshow(log_p, sr=sr, x_axis='time', y_axis='mel')\nplt.title('Log mel spectrogram for female percussive')\nplt.colorbar(format='%+02.0f dB')","metadata":{"_uuid":"d79406cf-c3e3-4e2e-be60-db0eea91cdf2","_cell_guid":"dc8efffb-6a5d-4c12-a2e7-12cfb735df5f","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:42.03163Z","iopub.execute_input":"2023-12-15T03:14:42.031893Z","iopub.status.idle":"2023-12-15T03:14:44.36376Z","shell.execute_reply.started":"2023-12-15T03:14:42.031843Z","shell.execute_reply":"2023-12-15T03:14:44.362889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here's another alternative way to go about getting HPSS which is the decompose route. Output's the same but just a shortcut","metadata":{"_uuid":"f7186bd8-22df-4dfa-8627-53e7aee6e36c","_cell_guid":"12e7e6e0-0c5d-4912-a6a6-00b33d1b3294","trusted":true}},{"cell_type":"code","source":"# Lets use this one for the male \nSAMPLE_RATE = 22050\nfname_f = 'recordings/' + 'english381.mp3'  \ny, sr = librosa.load(fname_f, sr=SAMPLE_RATE, duration = 5)\nX = librosa.stft(y)\nH, P = librosa.decompose.hpss(X)  # Both Harmonic and Percussive as spectogram \nHmag = librosa.amplitude_to_db(H) # Get log mel-spectogram \nPmag = librosa.amplitude_to_db(P)\n\n# Have a listen to male harmonic \nh = librosa.istft(H)\nipd.Audio(h, rate=SAMPLE_RATE)","metadata":{"_uuid":"51f2e35f-2f1a-4bf8-812e-3064c76106d2","_cell_guid":"309351c8-0e0a-4515-ab34-b9061d687292","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:44.36516Z","iopub.execute_input":"2023-12-15T03:14:44.365586Z","iopub.status.idle":"2023-12-15T03:14:44.991994Z","shell.execute_reply.started":"2023-12-15T03:14:44.365484Z","shell.execute_reply":"2023-12-15T03:14:44.991225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Have a listen to male percussive \np = librosa.istft(P)\nipd.Audio(p, rate=SAMPLE_RATE)","metadata":{"_uuid":"f5b58e83-0bbd-4c8d-8cc3-2a51421a4577","_cell_guid":"1faf96c7-3002-4c27-817a-5725619100e8","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:44.993543Z","iopub.execute_input":"2023-12-15T03:14:44.993829Z","iopub.status.idle":"2023-12-15T03:14:45.057023Z","shell.execute_reply.started":"2023-12-15T03:14:44.993773Z","shell.execute_reply":"2023-12-15T03:14:45.056319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the log mel spectrogram of both harmonic and percussive\nplt.figure(figsize=(12,6))\n\nplt.subplot(2,1,1)\nlibrosa.display.specshow(Hmag, sr=sr, x_axis='time', y_axis='mel')\nplt.title('Log mel spectrogram for male harmonic')\nplt.colorbar(format='%+02.0f dB')\n\nplt.subplot(2,1,2)\nlibrosa.display.specshow(Pmag, sr=sr, x_axis='time', y_axis='mel')\nplt.title('Log mel spectrogram for male percussive')\nplt.colorbar(format='%+02.0f dB')","metadata":{"_uuid":"0a958d05-a2f9-4f32-8c23-e1a256109eb5","_cell_guid":"0dea38ee-a9a3-4a34-b859-69f81c250e66","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:45.058204Z","iopub.execute_input":"2023-12-15T03:14:45.058455Z","iopub.status.idle":"2023-12-15T03:14:48.015541Z","shell.execute_reply.started":"2023-12-15T03:14:45.058411Z","shell.execute_reply":"2023-12-15T03:14:48.014465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notice how the HPSS for both female and male are quite diffferent? It sort of looks like the male version is abit rough? I'm no expert in interpreting the outputs of the HPSS but safe to say it will serve as a good feature for the machine learning model","metadata":{"_uuid":"0334a620-7515-40a6-92c0-1ae275653b4f","_cell_guid":"0bfdb273-26fa-472e-8a7f-46c30954184d","trusted":true}},{"cell_type":"markdown","source":"<a id=\"chroma\"></a>\n## 5. Chroma\n\nThis feature is still pretty new to me. But the way I understand it is that it classifies the pitches into 12 distinct classes called 'pitch profiles. Often used in music but again, I haven't done much research on this feature and have never tried it on a deep learning architecture before, at least not yet. I'll aim to give this a try in this series! \n\nBut before we do, lets compare how female and male differs in Chroma","metadata":{"_uuid":"ebb5f3c2-97ec-4115-9149-18da2d4126cc","_cell_guid":"68d0f650-dcd4-4cf7-a36e-d477116592d6","trusted":true}},{"cell_type":"code","source":"SAMPLE_RATE = 22050\nfname_f = 'recordings/' + 'english381.mp3'  \ny, sr = librosa.load(fname_f, sr=SAMPLE_RATE, duration = 5)\nC = librosa.feature.chroma_cqt(y=y, sr=sr)\n\n# Make a new figure\nplt.figure(figsize=(12,4))\n# To make sure that the colors span the full range of chroma values, set vmin and vmax\nlibrosa.display.specshow(C, sr=sr, x_axis='time', y_axis='chroma', vmin=0, vmax=1)\nplt.title('Chromagram')\nplt.colorbar()\nplt.tight_layout()","metadata":{"_uuid":"035903d8-6804-459f-b08f-5ec2cdfd6212","_cell_guid":"814a4842-062c-4b41-9c0a-f0a9dec2b87a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:48.017544Z","iopub.execute_input":"2023-12-15T03:14:48.018234Z","iopub.status.idle":"2023-12-15T03:14:49.097022Z","shell.execute_reply.started":"2023-12-15T03:14:48.018159Z","shell.execute_reply":"2023-12-15T03:14:49.095888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"SAMPLE_RATE = 22050\nfname_f = 'recordings/' + 'english385.mp3'  \ny, sr = librosa.load(fname_f, sr=SAMPLE_RATE, duration = 5)\nC = librosa.feature.chroma_cqt(y=y, sr=sr)\n\n# Make a new figure\nplt.figure(figsize=(12,4))\nlibrosa.display.specshow(C, sr=sr, x_axis='time', y_axis='chroma', vmin=0, vmax=1)\nplt.title('Chromagram')\nplt.colorbar()\nplt.tight_layout()","metadata":{"_uuid":"8c64f00b-8e51-4d73-af8e-45c30f163eb8","_cell_guid":"b0334637-138e-4029-8eef-a559ace2704a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-12-15T03:14:49.099238Z","iopub.execute_input":"2023-12-15T03:14:49.099991Z","iopub.status.idle":"2023-12-15T03:14:50.103329Z","shell.execute_reply.started":"2023-12-15T03:14:49.099906Z","shell.execute_reply":"2023-12-15T03:14:50.102177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Quite surprisingly, Chroma seems to be the most discriminative feature between male and female, or at least visually. We should definitely try and use Chroma for our deep learning.","metadata":{"_uuid":"745a7516-ded4-4c97-b66e-41be4ff6d39d","_cell_guid":"13a8f6c1-755c-446a-b4ec-3999418c11e7","trusted":true}},{"cell_type":"markdown","source":"------------------------------\n<a id=\"final\"></a>\n## Final thoughts\n\nAnd there you have it! The 5 features you can extract from audio. Note that there are more features you can extract such as Beat Tracking, spectogram, spectral informations and etc. But I'm not familiar with them. So I'll let the experts elaborate on those other features.\n\nNow moving on, part 3 we will go into the actual modelling. First up, I will show how you could use the traditional machine learning models first such as Gradient Boosting and Random Forest, before jumping into deep learning. Then at Part 4 we'll go into Deep Learning. \n\nStay Tuned!","metadata":{"_uuid":"7d084903-db30-4615-9e66-4bfcf307f1e2","_cell_guid":"e09603e3-ae83-4b06-9cd0-ac8059a39950","trusted":true}}]}