{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Importing packages","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom tqdm.notebook import tqdm\n\n## For audio\nfrom scipy import fft      \nimport librosa\nimport IPython.display as ipd\n\ncolor_pal = plt.rcParams[\"axes.prop_cycle\"].by_key()[\"color\"]","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-17T05:50:11.388486Z","iopub.execute_input":"2024-04-17T05:50:11.389184Z","iopub.status.idle":"2024-04-17T05:50:12.419516Z","shell.execute_reply.started":"2024-04-17T05:50:11.389158Z","shell.execute_reply":"2024-04-17T05:50:12.418749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Sample audio File and other configs","metadata":{}},{"cell_type":"code","source":"audiofile = '/kaggle/input/birdclef-2024/train_audio/asbfly/XC134896.ogg'","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:14.379482Z","iopub.execute_input":"2024-04-17T05:50:14.379885Z","iopub.status.idle":"2024-04-17T05:50:14.385455Z","shell.execute_reply.started":"2024-04-17T05:50:14.379862Z","shell.execute_reply":"2024-04-17T05:50:14.384792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Reading file and ploting it\n- An audio file is actually a time series data of numbers \n- depicting the amplitude at each time step, \n- the **size of the time series = time_in_sec x sample_rate)**. \n- Sample rate is the points to sample for each second of data.","metadata":{}},{"cell_type":"markdown","source":"## a. Audio is time series data","metadata":{}},{"cell_type":"code","source":"audio_data, sample_rate = librosa.load(audiofile, sr=32000)\nprint(audio_data.shape, sample_rate)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:15.188119Z","iopub.execute_input":"2024-04-17T05:50:15.188427Z","iopub.status.idle":"2024-04-17T05:50:23.756945Z","shell.execute_reply.started":"2024-04-17T05:50:15.188400Z","shell.execute_reply":"2024-04-17T05:50:23.756004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ipd.Audio(audiofile)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:23.758284Z","iopub.execute_input":"2024-04-17T05:50:23.758646Z","iopub.status.idle":"2024-04-17T05:50:23.786461Z","shell.execute_reply.started":"2024-04-17T05:50:23.758625Z","shell.execute_reply":"2024-04-17T05:50:23.785692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Here, audio_data.shape = (875207,);\n- For sample_rate = 32000, audio_data.shape = (875207,) ~= sample_rate * time_of_audio\n- As the sample_rate=32000 so the max frequency we can sample is sample_rate/2=16KHz by Nyquist criteria.","metadata":{"execution":{"iopub.status.busy":"2024-04-16T08:15:24.729569Z","iopub.execute_input":"2024-04-16T08:15:24.731092Z","iopub.status.idle":"2024-04-16T08:15:24.745286Z","shell.execute_reply.started":"2024-04-16T08:15:24.731046Z","shell.execute_reply":"2024-04-16T08:15:24.743470Z"}}},{"cell_type":"markdown","source":"## b. ploting audio data as timeseries","metadata":{"execution":{"iopub.status.busy":"2024-04-16T08:11:24.505597Z","iopub.execute_input":"2024-04-16T08:11:24.506062Z","iopub.status.idle":"2024-04-16T08:11:24.511636Z","shell.execute_reply.started":"2024-04-16T08:11:24.506029Z","shell.execute_reply":"2024-04-16T08:11:24.510231Z"}}},{"cell_type":"markdown","source":"**Trimming Audio( first 5 seconds)**","metadata":{}},{"cell_type":"code","source":"audio_data = audio_data[:5*32000]","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:23.787561Z","iopub.execute_input":"2024-04-17T05:50:23.787795Z","iopub.status.idle":"2024-04-17T05:50:23.791562Z","shell.execute_reply.started":"2024-04-17T05:50:23.787748Z","shell.execute_reply":"2024-04-17T05:50:23.790657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.Series(audio_data).plot(figsize=(10, 5),\n                          lw=1,\n                          title='Raw Audio',\n                         color=color_pal[0])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:23.793221Z","iopub.execute_input":"2024-04-17T05:50:23.793630Z","iopub.status.idle":"2024-04-17T05:50:24.062386Z","shell.execute_reply.started":"2024-04-17T05:50:23.793608Z","shell.execute_reply":"2024-04-17T05:50:24.061613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Sliced Audio**","metadata":{}},{"cell_type":"code","source":"pd.Series(audio_data[60000:60100]).plot(figsize=(10, 5),\n                          lw=1,\n                          title='Sliced Audio',\n                         color=color_pal[1])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:24.063459Z","iopub.execute_input":"2024-04-17T05:50:24.063657Z","iopub.status.idle":"2024-04-17T05:50:24.236280Z","shell.execute_reply.started":"2024-04-17T05:50:24.063638Z","shell.execute_reply":"2024-04-17T05:50:24.235625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Interpretting audio data","metadata":{}},{"cell_type":"markdown","source":"Audio data as raw(time-series format) is not useful. \nSome of the important features of audio data is :\n1. Frequency\n2. Amplitude of those frequency or power ∝ (Amplitude)²\n\n- Now, the audio signal is made of different frequency(think of sine wave of different frequencies) overlapped over one another \n- For getting the frequency present in signal we need to apply a fourier transform. \n- And FFT is an faster algorithm for calculating fourier transform.\n- If ``x`` is a 1d array, then the `fft` is equivalent to ::\n\n    ``y[k] = np.sum(x * np.exp(-2j * np.pi * k * np.arange(n)/n))``","metadata":{}},{"cell_type":"markdown","source":"## a. Calculating FFT:","metadata":{}},{"cell_type":"code","source":"D = fft.fft(audio_data)\nS_db = librosa.amplitude_to_db(np.abs(D), ref=np.max)\naudio_data.shape, D.shape, S_db.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:24.237088Z","iopub.execute_input":"2024-04-17T05:50:24.237712Z","iopub.status.idle":"2024-04-17T05:50:24.255527Z","shell.execute_reply.started":"2024-04-17T05:50:24.237690Z","shell.execute_reply":"2024-04-17T05:50:24.254614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Understanding output Shape:\n- It ouputs a complex number for each frequency.\n- So, here there are 160000 discrete frequency and amplitude( as a complex number).\n- If at ith index the value of S_db is high(`v` ), that mean s the ith frequency is present in audio with `v`  intensity.\n- Hence you can compare relative strength of each frequency in the audio.","metadata":{}},{"cell_type":"code","source":"x = audio_data\nn = len(audio_data)\nk = 100\nnp.sum(x * np.exp(-2j * np.pi * k * np.arange(n)/n)), D[k]","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:24.257561Z","iopub.execute_input":"2024-04-17T05:50:24.258040Z","iopub.status.idle":"2024-04-17T05:50:24.269736Z","shell.execute_reply.started":"2024-04-17T05:50:24.258013Z","shell.execute_reply":"2024-04-17T05:50:24.269126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notice: both are same (apart from the precision)","metadata":{}},{"cell_type":"code","source":"N = audio_data.shape[0]\nyf = fft.rfft(audio_data)\nxf = fft.rfftfreq(N, 1 / 32000)\n\nplt.plot(xf, np.abs(yf), lw=1, color=color_pal[2])\nplt.title('Power Spectrogram of audio')\nplt.xlabel(\"Frequency(Hz)\")\nplt.ylabel(\"Amplitude\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:24.270860Z","iopub.execute_input":"2024-04-17T05:50:24.271329Z","iopub.status.idle":"2024-04-17T05:50:24.450940Z","shell.execute_reply.started":"2024-04-17T05:50:24.271303Z","shell.execute_reply":"2024-04-17T05:50:24.450161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, these Frequencies are prominent in this audio. (5KHz-8.5KHz). This will come later in other spectrograms.\n","metadata":{}},{"cell_type":"code","source":"plt.plot(xf[5000:8000], np.abs(yf[5000:8000]), lw=1, color=color_pal[2])\nplt.title('Sliced power Spectrogram of audio')\nplt.xlabel(\"Frequency(Hz)\")\nplt.ylabel(\"Amplitude\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:24.451612Z","iopub.execute_input":"2024-04-17T05:50:24.451829Z","iopub.status.idle":"2024-04-17T05:50:24.645524Z","shell.execute_reply.started":"2024-04-17T05:50:24.451809Z","shell.execute_reply":"2024-04-17T05:50:24.644912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## b. Calculating STFT","metadata":{}},{"cell_type":"markdown","source":"- FFT( a DFT algo) gives you the overall frequency components for the entire time series. So, it won’t tell you about how the frequency appears to be changing over time. \n- If you wanted to get a sense, for example, where the high and low pitches were in music, you could apply the STFT.\n- Short-time Fourier transform (STFT) is a method of taking a “window” that slides along the time series and performing the DFT on the time dependent segment. \n- This would give you a DFT that changes with time. This would not be readily apparent from just applying the DFT to the entire time series, as it gives one set of components that isn’t time dependent.","metadata":{}},{"cell_type":"markdown","source":"Important params for STFT:\n1. n_fft: The length of window for which DFT (using FFT algo) has to be calculated. The algo is accurate when this is power of 2. Generally for speech, it should be 23 milliseconds length. for 32kHz, n_fft = 32*23 = 736. Nearest power of 2 comes 512 which corresponds to 16 milliseconds.\n2. hop_length: The gap between start of current window to previous window. Think of this as stride in 1D convolution.\nThe next 2 are related to window \n3. win_length: The length of window , defaults to n_fft\n4. window: a function used to smoothening the signal (in the window), usually a `hann` or `hanning` function.","metadata":{}},{"cell_type":"code","source":"n_fft = 512\nD = librosa.stft(audio_data,n_fft=n_fft, hop_length=n_fft//4)\nS_db = librosa.amplitude_to_db(np.abs(D), ref=np.max)\nD.shape, S_db.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:24.648153Z","iopub.execute_input":"2024-04-17T05:50:24.649682Z","iopub.status.idle":"2024-04-17T05:50:24.663493Z","shell.execute_reply.started":"2024-04-17T05:50:24.649659Z","shell.execute_reply":"2024-04-17T05:50:24.662797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Understanding output Shape:\n- rows(denoting frequency) = 1 + n_fft/2 = 1 + 512/2 = 257 , i.e. 257 frequency components for each window. window_size = 512 so 0 to 512/2 frequncy would be there. \n- cols(denoting time) = no. of FFT calculated = audio_data.shape[0]/hop_length ~= 1251\n\nFrequency mapping can be calculated from `fft.rfftfreq(512, 1 / 32000)` window_length=512=n_ftt\n","metadata":{}},{"cell_type":"markdown","source":"## c. Mel Frequency","metadata":{}},{"cell_type":"markdown","source":"The mel scale is a scale of pitches judged by listeners to be equal in distance one from another. The reference point between this scale and normal frequency measurement is defined by equating a 1000 Hz tone, 40 dB above the listener's threshold, with a pitch of 1000 mels. Below about 500 Hz the mel and hertz scales coincide; above that, larger and larger INTERVALs are judged by listeners to produce equal pitch increments. \n\nIn simple words, 2 frequency 100Hz apart can be distinguised well when they are in lower range(say 2000Hz-3000Hz) than in higher range(say 18000Hz-19000Hz). Thats the precise need for Mel-scale. As the frequency grows the gap in Mel-range becomes larger and larger. You can say it as more descrete representation of audio data.","metadata":{}},{"cell_type":"code","source":"S = librosa.feature.melspectrogram(y=audio_data,\n                                   sr=32000,\n                                   n_mels=128*2,\n                                   n_fft=n_fft, \n                                   hop_length=n_fft//4)\nS_db_mel = librosa.amplitude_to_db(S, ref=np.max)\nS.shape, S_db_mel.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:24.664256Z","iopub.execute_input":"2024-04-17T05:50:24.664455Z","iopub.status.idle":"2024-04-17T05:50:25.693240Z","shell.execute_reply.started":"2024-04-17T05:50:24.664437Z","shell.execute_reply":"2024-04-17T05:50:25.692260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mel_f = librosa.filters.mel(sr=32000, n_fft=n_fft, n_mels=128)\nmel_f.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:25.694120Z","iopub.execute_input":"2024-04-17T05:50:25.695028Z","iopub.status.idle":"2024-04-17T05:50:25.703049Z","shell.execute_reply.started":"2024-04-17T05:50:25.695003Z","shell.execute_reply":"2024-04-17T05:50:25.702372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## c. Spectrogram","metadata":{}},{"cell_type":"markdown","source":"A spectogram gives us the way to represent the power of each frequency at a given time.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 5))\nimg = librosa.display.specshow(S_db,\n                                n_fft =512,\n                                hop_length=512//4,\n                                sr=32000,\n                                x_axis='s',\n                                y_axis='hz',\n                                ax=ax)\nax.set_title('Spectrogram  of audio_data', fontsize=20)\nfig.colorbar(img, ax=ax, format=f'%0.2f')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:25.704673Z","iopub.execute_input":"2024-04-17T05:50:25.704931Z","iopub.status.idle":"2024-04-17T05:50:26.196346Z","shell.execute_reply.started":"2024-04-17T05:50:25.704909Z","shell.execute_reply":"2024-04-17T05:50:26.195703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## d. Mel-Spectrogram","metadata":{}},{"cell_type":"markdown","source":"It is a spectrogram in Mel Scale:\n\n`mel_f.dot(S**power) #power=2, S=(stft of audio data)**2`\n\n- Esentially it converts the spectogram from frequncy domain to mel- frequncy domain.\n- For doing so, it calculated mel_f tranfomatation matrix of shape (n_mels, S.shape[0])\n- This is done through: `librosa.filters.mel(sr=32000, n_fft=n_fft, n_mels=128)`","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 5))\n# Plot the mel spectogram\nimg = librosa.display.specshow(S_db_mel,\n                              n_fft =512,\n                                hop_length=512//4,\n                                sr=32000,\n                                x_axis='s',\n                                y_axis='mel',\n                              ax=ax)\nax.set_title('Mel Spectrogram of audio_data', fontsize=20)\nfig.colorbar(img, ax=ax, format=f'%0.2f')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:26.425492Z","iopub.execute_input":"2024-04-17T05:50:26.426457Z","iopub.status.idle":"2024-04-17T05:50:26.866652Z","shell.execute_reply.started":"2024-04-17T05:50:26.426415Z","shell.execute_reply":"2024-04-17T05:50:26.865871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This filters lot of noises for us.","metadata":{}},{"cell_type":"markdown","source":"# 5. More Features","metadata":{}},{"cell_type":"markdown","source":"A DFT, STFT, MEL-freqency and spectograms can be features of Audio as described above. \n\nSo, you captured all frequency, their power in some specified windows(say 16ms). Also found it in Mel scale.\nBut Frequency and amplitude in itself, are not much useful either. \n\nA characterstic of audio is also defined by a changing amplitude/frequency. Enter **`MFCC`**","metadata":{}},{"cell_type":"code","source":"D = librosa.feature.mfcc(y=audio_data, sr=32000, n_fft=n_fft, hop_length=n_fft//4)\nS_db_mfcc = librosa.amplitude_to_db(np.abs(D), ref=np.max)\nD.shape, S_db_mfcc.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:28.538542Z","iopub.execute_input":"2024-04-17T05:50:28.538853Z","iopub.status.idle":"2024-04-17T05:50:28.558069Z","shell.execute_reply.started":"2024-04-17T05:50:28.538832Z","shell.execute_reply":"2024-04-17T05:50:28.557350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 5))\n# Plot the mel spectogram\nimg = librosa.display.specshow(S_db_mfcc,\n                              n_fft =512,\n                                hop_length=512//4,\n                                sr=32000,\n                                x_axis='s',\n                              ax=ax)\nax.set_title('MFCC of audio_data', fontsize=20)\nfig.colorbar(img, ax=ax, format=f'%0.2f')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T05:50:28.816658Z","iopub.execute_input":"2024-04-17T05:50:28.816966Z","iopub.status.idle":"2024-04-17T05:50:29.034131Z","shell.execute_reply.started":"2024-04-17T05:50:28.816945Z","shell.execute_reply":"2024-04-17T05:50:29.033521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we can calculate the delta deltas of MFCC features created earlier.","metadata":{}},{"cell_type":"code","source":"mfcc = D\nmfcc_delta = librosa.feature.delta(mfcc)\nmfcc_delta2 = librosa.feature.delta(mfcc, order=2)   \nmfccs = (np.vstack(( mfcc,  mfcc_delta,  mfcc_delta2)))\nS_db_mfcc_w_deltas = librosa.amplitude_to_db(np.abs(mfccs), ref=np.max)\nmfccs.shape, S_db_mfcc_w_deltas.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-17T06:04:43.371314Z","iopub.execute_input":"2024-04-17T06:04:43.371635Z","iopub.status.idle":"2024-04-17T06:04:43.382109Z","shell.execute_reply.started":"2024-04-17T06:04:43.371614Z","shell.execute_reply":"2024-04-17T06:04:43.381389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 5))\n# Plot the mel spectogram\nimg = librosa.display.specshow(S_db_mfcc_w_deltas,\n                              n_fft =512,\n                                hop_length=512//4,\n                                sr=32000,\n                                x_axis='s',\n                              ax=ax)\nax.set_title('MFCC with Deltas of audio_data', fontsize=20)\nfig.colorbar(img, ax=ax, format=f'%0.2f')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T06:05:38.093695Z","iopub.execute_input":"2024-04-17T06:05:38.094005Z","iopub.status.idle":"2024-04-17T06:05:38.506550Z","shell.execute_reply.started":"2024-04-17T06:05:38.093984Z","shell.execute_reply":"2024-04-17T06:05:38.505791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. The road ahead 🛣️","metadata":{}},{"cell_type":"markdown","source":"Now that, we have intensity at (frequency, time) i.e. a 2D representation of data. A good choic eis to use a CNN.\nWhy? Because a CNN filter can capture spatial info (change along x, y axis), here:\n1. Change along x is change in a particular frequency along time or the temporal change in a frequency's amplitude.\n2. Change along y is change in a neigbouring frequencies' intensity at a particular time or the intensity change in a frequency domain.\n3. Other complex features.","metadata":{}}]}