{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Audio data is an important type of data that is increasingly being used in deep learning applications, such as speech recognition, audio classification, and sound event detection. OpenAI's **Whisper model** (https://openai.com/research/whisper), which uses a Transformer sequence-to-sequence architecture to tackle a range of speech processing tasks, from multilingual speech recognition to voice activity detection is an example of exciting progress being made in the field of deep learning and audio processing.\n\nHowever, raw audio data is not directly usable by deep learning models, and therefore it needs to be processed into a format that can be effectively used for training.\n\nThe following sections provide a step-by-step guide on how to handle audio data. It is an introduction to audio data processing for deep learning applications, and can be useful for beginners looking to work with audio data in their projects.\n\nSome of the content used here was sourced from **Valerio Velardo - The Sound of AI** (https://www.youtube.com/@ValerioVelardoTheSoundofAI).","metadata":{}},{"cell_type":"markdown","source":"# Import necessary modules","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport IPython.display as ipd\nimport librosa.display as lid\n\nimport torch\nimport librosa\nimport torchaudio\nimport torchvision.transforms as T\nimport torchvision.transforms.functional as F\n\nfrom matplotlib import pyplot as plt\nfrom torchaudio.transforms import MelSpectrogram, AmplitudeToDB\n","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:46:46.364989Z","iopub.execute_input":"2023-03-12T14:46:46.365538Z","iopub.status.idle":"2023-03-12T14:46:50.571222Z","shell.execute_reply.started":"2023-03-12T14:46:46.365490Z","shell.execute_reply":"2023-03-12T14:46:50.569471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta_df = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\nmeta_df = meta_df[['primary_label', 'filename']]\nmeta_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:46:50.573953Z","iopub.execute_input":"2023-03-12T14:46:50.574558Z","iopub.status.idle":"2023-03-12T14:46:50.737960Z","shell.execute_reply.started":"2023-03-12T14:46:50.574522Z","shell.execute_reply":"2023-03-12T14:46:50.736541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TRAIN_ROOT_PATH = \"/kaggle/input/birdclef-2023/train_audio\"\nTEST_ROOT_PATH = \"/kaggle/input/birdclef-2023/test_soundscapes\"\n\nfile_names = list(meta_df['filename'])\nfile_path = os.path.join(TRAIN_ROOT_PATH, file_names[10])\nprint(file_path)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:46:50.739767Z","iopub.execute_input":"2023-03-12T14:46:50.741307Z","iopub.status.idle":"2023-03-12T14:46:50.755897Z","shell.execute_reply.started":"2023-03-12T14:46:50.741250Z","shell.execute_reply":"2023-03-12T14:46:50.754075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the audio file","metadata":{}},{"cell_type":"markdown","source":"The **torchaudio.load** function returns a tuple containing the signal tensor and the sample rate of the audio file. The signal tensor contains the actual audio data and is represented as a 2D tensor, where the first dimension represents the **n_channel** and the second dimension represents the **n_sample**. The **sample_rate** indicates how many samples per second were used to record the audio. This information is crucial for further processing of the audio signal.\n\nn_channel refer to the number of independent audio signals that are recorded or played back simultaneously. **Mono audio**, also known as monophonic or single-channel audio, consists of a single audio signal that is reproduced through a single speaker or headphone. Mono audio is typically used for voice recordings, as well as in situations where sound quality is less important, such as in telephone systems or AM radio. **Stereo audio**, on the other hand, uses two channels to create a more immersive listening experience. In stereo audio, two independent audio signals are recorded or played back simultaneously, with each signal representing a different channel. These channels are typically designated as the left channel and the right channel, and they are mixed together to create a sense of space and depth in the audio.","metadata":{}},{"cell_type":"code","source":"signal, sr = torchaudio.load(file_path)\nn_channel, n_sample = signal.shape\n\nprint(f'n_channel = {n_channel}')\nprint(f'n_sample = {n_sample}')\nprint(f'sample_rate = {sr}')","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:07.236137Z","iopub.execute_input":"2023-03-12T14:48:07.236634Z","iopub.status.idle":"2023-03-12T14:48:07.289466Z","shell.execute_reply.started":"2023-03-12T14:48:07.236570Z","shell.execute_reply":"2023-03-12T14:48:07.287854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Display the audio file","metadata":{}},{"cell_type":"code","source":"display(ipd.Audio(signal, rate=sr))","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:07.828739Z","iopub.execute_input":"2023-03-12T14:48:07.829174Z","iopub.status.idle":"2023-03-12T14:48:07.871741Z","shell.execute_reply.started":"2023-03-12T14:48:07.829136Z","shell.execute_reply":"2023-03-12T14:48:07.870321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Display the signal","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(15,3))\nlid.waveshow(signal.numpy(), sr=sr, ax=ax)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:08.407550Z","iopub.execute_input":"2023-03-12T14:48:08.408116Z","iopub.status.idle":"2023-03-12T14:48:08.799266Z","shell.execute_reply.started":"2023-03-12T14:48:08.408076Z","shell.execute_reply":"2023-03-12T14:48:08.797896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Resample the signal if necessary","metadata":{}},{"cell_type":"markdown","source":"Audio resampling refers to the process of changing the sample rate of an audio signal. This is necessary when the sample rate of an audio file is not compatible with the system or software being used to process or play the audio. Here we resample the audio signal to a **target_sample_rate** (tsr) for the sake of uniformity.","metadata":{}},{"cell_type":"code","source":"tsr = 32_000\n\nif sr != tsr:\n    print(\"Resampling the signal..\")\n    signal = torchaudio.functional.resample(signal, sr, tsr)\n    \nprint(signal.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:10.171355Z","iopub.execute_input":"2023-03-12T14:48:10.171791Z","iopub.status.idle":"2023-03-12T14:48:10.179321Z","shell.execute_reply.started":"2023-03-12T14:48:10.171751Z","shell.execute_reply":"2023-03-12T14:48:10.178023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mix down the channel if necessary","metadata":{}},{"cell_type":"markdown","source":"Mix down of channels refers to the process of combining multiple audio channels into a single stereo or mono track. This is necessary when dealing with audio files that have multiple channels, such as stereo or surround sound files. We do this by taking mean of the channels using **torch.mean** function.","metadata":{}},{"cell_type":"code","source":"# signal = torch.rand(3, 640000)\n\nif n_channel > 1:\n    print(\"Mixing down..\")\n    signal = torch.mean(signal, dim=0, keepdim=True)\n    \nprint(signal.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:12.453069Z","iopub.execute_input":"2023-03-12T14:48:12.453509Z","iopub.status.idle":"2023-03-12T14:48:12.461110Z","shell.execute_reply.started":"2023-03-12T14:48:12.453465Z","shell.execute_reply":"2023-03-12T14:48:12.459644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Extract Mel Spectrogram","metadata":{}},{"cell_type":"markdown","source":"Here's a step-by-step process for converting a 1-dimensional audio signal with 64,000 samples into a mel spectrogram:\n\n**Windowing**: Apply a window function, such as a Hann window, to the audio signal to reduce the impact of edge effects and improve the frequency resolution of the spectrogram.\n\n**Short-Time Fourier Transform** (STFT): Apply the STFT to the windowed audio signal to obtain a complex-valued spectrogram. The STFT is computed by dividing the signal into overlapping frames of a fixed length (**win_length/n_fft**), applying a Fourier transform to each frame, and stacking the resulting spectra to form a 2-dimensional spectrogram.\n\n**Magnitude Spectrum**: Take the magnitude of the complex spectrogram to obtain a real-valued spectrogram.\n\n**Mel-scale Filterbank**: Apply a mel-scale filterbank to the magnitude spectrogram to obtain a mel spectrogram. The filterbank consists of a set of triangular filters (**n_mels**) that are spaced uniformly on the mel scale, which is a non-linear scale that approximates the human auditory system's perception of frequency.\n\n**Logarithmic Compression**: Apply a logarithmic compression to the mel spectrogram to compress the dynamic range and make it more suitable for analysis. This is typically done using a logarithmic function, such as the decibel scale","metadata":{}},{"cell_type":"markdown","source":"# But before that","metadata":{}},{"cell_type":"markdown","source":"# Crop/Pad signal if necessary","metadata":{}},{"cell_type":"markdown","source":"Signals of varying size will have different numbers of windows and, therefore, different numbers of spectra in the resulting spectrogram. The number of spectra in the spectrogram depends on the length of the signal, the length of the window, and the amount of overlap between adjacent windows. In general, longer signals will have more spectra, while shorter signals will have fewer spectra.\n\nWhen processing signals of varying length, it's common practice to set a maximum length and truncate longer signals or pad shorter signals to match that length, as discussed earlier. This ensures that all signals are represented by the same number of spectra, which is necessary for feeding them into a machine learning model that expects inputs of fixed size.","metadata":{}},{"cell_type":"code","source":"duration = 20\ndesired_n_sample = duration * tsr\n\nif n_sample > desired_n_sample:\n    print(\"Cropping the waveform..\")\n    signal = signal[:,:desired_n_sample]\n\nprint(signal.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:15.154461Z","iopub.execute_input":"2023-03-12T14:48:15.154953Z","iopub.status.idle":"2023-03-12T14:48:15.162154Z","shell.execute_reply.started":"2023-03-12T14:48:15.154911Z","shell.execute_reply":"2023-03-12T14:48:15.160871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#signal = torch.rand(1, 62000)\n\nif n_sample < desired_n_sample:\n    print(\"Padding the waveform..\")\n    padding = desired_n_sample - n_sample\n    signal = torch.nn.functional.pad(signal, (0, padding))\n\nprint(signal.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:15.732089Z","iopub.execute_input":"2023-03-12T14:48:15.732503Z","iopub.status.idle":"2023-03-12T14:48:15.746787Z","shell.execute_reply.started":"2023-03-12T14:48:15.732467Z","shell.execute_reply":"2023-03-12T14:48:15.745740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The MelSpectrogram transform from the torchaudio.transforms is used to compute the Mel Spectrogram of an audio signal.","metadata":{}},{"cell_type":"code","source":"n_fft=1024\nwin_length=1024\nn_mels=64\nhop_length=512\n\nmelspectrogram = MelSpectrogram(\n    sample_rate=tsr, \n    n_fft=n_fft,\n    n_mels=n_mels, \n    hop_length=hop_length,\n    win_length=win_length\n)\n\nmel = melspectrogram(signal)\nprint(mel.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:18.535378Z","iopub.execute_input":"2023-03-12T14:48:18.535854Z","iopub.status.idle":"2023-03-12T14:48:18.561304Z","shell.execute_reply.started":"2023-03-12T14:48:18.535812Z","shell.execute_reply":"2023-03-12T14:48:18.559872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The AudioToDecibel transform from the torchaudio.transforms is used to convert the amplitude values of an audio signal from the linear scale to the decibel (dB) scale. The dB scale is a logarithmic scale that is commonly used in acoustics to express the relative difference in sound pressure levels.","metadata":{}},{"cell_type":"code","source":"amplitude2db = AmplitudeToDB(\n    stype='power',\n    top_db=None\n)\n\nmel = amplitude2db(mel)\nprint(mel.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:19.908196Z","iopub.execute_input":"2023-03-12T14:48:19.909141Z","iopub.status.idle":"2023-03-12T14:48:19.917752Z","shell.execute_reply.started":"2023-03-12T14:48:19.909088Z","shell.execute_reply":"2023-03-12T14:48:19.916076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Display Mel Spectrogram","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(15, 3))\nlid.specshow(\n    data=mel[0,:].numpy(), \n    sr=tsr, \n    hop_length=hop_length,\n    cmap = 'coolwarm',\n    n_fft=n_fft,\n    ax=ax\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:21.268284Z","iopub.execute_input":"2023-03-12T14:48:21.268863Z","iopub.status.idle":"2023-03-12T14:48:21.553289Z","shell.execute_reply.started":"2023-03-12T14:48:21.268805Z","shell.execute_reply":"2023-03-12T14:48:21.552086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Refactor Mel Spectrogram","metadata":{}},{"cell_type":"markdown","source":"Stack the extracted 2D Mel Spectrogram thrice over one another to make it resemble the dimensions of an image with RGB channels.","metadata":{}},{"cell_type":"code","source":"mel = torch.concat([mel for _ in range(3)])\nprint(mel.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:27.431008Z","iopub.execute_input":"2023-03-12T14:48:27.431777Z","iopub.status.idle":"2023-03-12T14:48:27.438454Z","shell.execute_reply.started":"2023-03-12T14:48:27.431735Z","shell.execute_reply":"2023-03-12T14:48:27.437146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Resize the image to a desired dimension.","metadata":{}},{"cell_type":"code","source":"mel = F.resize(mel, (64,64))\nprint(mel.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:28.942999Z","iopub.execute_input":"2023-03-12T14:48:28.943462Z","iopub.status.idle":"2023-03-12T14:48:28.951006Z","shell.execute_reply.started":"2023-03-12T14:48:28.943420Z","shell.execute_reply":"2023-03-12T14:48:28.949519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"normalize the image.","metadata":{}},{"cell_type":"code","source":"# mean = torch.mean(mel, dim=(1,2))\n# std = torch.std(mel, dim=(1,2))\n# print(f'mean={mean}, std={std}')\n\n# norm = T.Normalize(mean, std)\n# mel = norm(mel)\n\nmel = torch.nn.functional.normalize(mel)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:55.797864Z","iopub.execute_input":"2023-03-12T14:48:55.799096Z","iopub.status.idle":"2023-03-12T14:48:55.804959Z","shell.execute_reply.started":"2023-03-12T14:48:55.799046Z","shell.execute_reply":"2023-03-12T14:48:55.803280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(5,5))\nlibrosa.display.specshow(mel[0,:].numpy(), y_axis='mel', x_axis='time', ax=ax)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T14:48:56.978618Z","iopub.execute_input":"2023-03-12T14:48:56.979682Z","iopub.status.idle":"2023-03-12T14:48:57.203483Z","shell.execute_reply.started":"2023-03-12T14:48:56.979637Z","shell.execute_reply":"2023-03-12T14:48:57.201925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}