{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Bengali.AI Speech Recognition: From Audio Signals to Mel Spectrogram\n\n<img src=\"https://tds-images.thedailystar.net/sites/default/files/styles/very_big_1/public/feature/images/amar_ekushey_8.jpg?itok=dt-VJGk_\" width=\"500\" height=\"300\" />\n\n\n# Introduction\n\nBeing a Bengali, I'm very thrilled to participate in this competition and excited to learn many new things along the way. I am very new to working with audio signals. I'll share all my learning as I'll progress through this competition. I hope I'll be able to contribute something meaningful to open-source speech recognition methods for Bengali at the end of this competition.\n\nThe goal of this notebook is to\n- Get overview of the data provided for this competition\n- Understand the fundamentals of audio signals and their representations.\n- Fourier transform and why is this important.\n- Explore the concept of spectrograms as visual representations of audio frequency content.\n- Mel Spectrogram","metadata":{}},{"cell_type":"markdown","source":"### Imports","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pylab as plt\nimport seaborn as sns\n\nfrom glob import glob\n\nimport librosa\nimport librosa.display\nimport IPython.display as ipd\n\nfrom itertools import cycle\n\nsns.set_theme(style = \"white\", palette = None)\ncolor_pal = plt.rcParams[\"axes.prop_cycle\"].by_key()[\"color\"]\ncolor_cycle = cycle(plt.rcParams[\"axes.prop_cycle\"].by_key()[\"color\"])","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:06.987829Z","iopub.execute_input":"2023-08-15T11:17:06.988259Z","iopub.status.idle":"2023-08-15T11:17:06.996745Z","shell.execute_reply.started":"2023-08-15T11:17:06.988223Z","shell.execute_reply":"2023-08-15T11:17:06.995827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# About the dataset\n- The competition dataset comprises about 1200 hours of recordings of Bengali speech.\n- The full test set contains about 20 hours of speech in almost 8000 MP3 audio files. All of the files in the test set are encoded at a sample rate of 32k, a bit rate of 48k, in one channel. The test set annotations were normalized with the bnUnicodeNormalizer.\n- `train/` The training set, comprising several thousand recordings in MP3 format.\n- `test/` The test set, comprising spontaneous speech recordings from eighteen domains, seventeen of which are out-of-distribution with respect to the training set. There may be domains in the private test set that are not in the public test set.\n- `examples/` An example recording for each test set domain. You may find these example recordings helpful for creating models robust to domain variation. These are representative recordings and none of them are present in the test set.\n- `train.csv` Sentence labels for the training set.\n    - id A unique identifier for this instance. Corresponds to the file {id}.mp3 in train/.\n    - sentence A plain-text transcription of the recording. Your goal is to predict these sentences for each recording in the test set.\n    - split Whether train or valid. The annotations in the valid split have been manually reviewed and corrected, while the annotations in the train split have only been algorithmically cleaned. The valid samples will generally have higher quality annotations than the train samples, but are otherwise drawn from the same distribution.","metadata":{}},{"cell_type":"markdown","source":"### Lets listen one of the audio files present in the examples folder.\n","metadata":{}},{"cell_type":"code","source":"#Play audio file\nexamples = glob(\"../input/bengaliai-speech/examples/*.wav\")\nipd.Audio(examples[6])","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:07.657034Z","iopub.execute_input":"2023-08-15T11:17:07.657701Z","iopub.status.idle":"2023-08-15T11:17:07.991391Z","shell.execute_reply.started":"2023-08-15T11:17:07.657653Z","shell.execute_reply":"2023-08-15T11:17:07.989698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above audio is recitation of a famous bengali poem বিদ্রোহী(The Rebel) by Kazi Nazrul Islam\n\nবল বীর —<br>\nবল উন্নত মম শির!<br>\nশির     নেহারি আমারি, নত-শির ওই শিখর হিমাদ্রীর!<br>\nবল বীর —<br>\nবল     মহাবিশ্বের মহাকাশ ফাড়ি’<br>\nচন্দ্র সূর্য্য গ্রহ তারা ছাড়ি’<br>\nভূলোক দ্যুলোক গোলক ভেদিয়া,<br>\nখোদার আসন ‘আরশ’ ছেদিয়া<br>\nউঠিয়াছি চির-বিস্ময় আমি বিশ্ব-বিধাত্রীর!<br>\nমম     ললাটে রুদ্র-ভগবান জ্বলে রাজ-রাজটীকা দীপ্ত জয়শ্রীর!<br>\nবল বীর —<br>\nআমি চির-উন্নত শির!<br>\n\nTranslation:\n\nProclaim, Hero,<br>\nproclaim: I raise my head high!<br>\nBefore me bows down the Himalayan peaks!<br>\n\nProclaim, Hero,<br>\nproclaim: rending through the sky,<br>\nsurpassing the moon, the sun,<br>\nthe planets, the stars,<br>\npiercing through the earth,<br>\nthe heavens, the cosmos<br>\nand the Almighty's throne,<br>\nhave I risen, the eternal wonder<br>\nof the Creator of the universe.<br>\nThe furious Shiva shines on my forehead<br>\nlike a royal medallion of victory!<br>\n\nProclaim, Hero,<br>\nproclaim: My head is ever held high!","metadata":{"execution":{"iopub.status.busy":"2023-08-14T20:20:41.552205Z","iopub.execute_input":"2023-08-14T20:20:41.553194Z","iopub.status.idle":"2023-08-14T20:20:41.568346Z","shell.execute_reply.started":"2023-08-14T20:20:41.553160Z","shell.execute_reply":"2023-08-14T20:20:41.566928Z"}}},{"cell_type":"markdown","source":"### Let's read train.csv\ntrian.csv contains the ids of the audio files and its transcripts.\n- We have 963636 records.\n- out of which 29588 records are labeled as valid. The valid split have been manually reviewed and corrected. So, these samples will generally have higher quality annotations than the other samples","metadata":{}},{"cell_type":"code","source":"# read the train data\ntrain = pd.read_csv(\"../input/bengaliai-speech/train.csv\")\nprint(train.head())\nprint(f\"Shape: {train.shape}\")\nprint(train.split.value_counts())","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:08.264566Z","iopub.execute_input":"2023-08-15T11:17:08.265783Z","iopub.status.idle":"2023-08-15T11:17:12.712222Z","shell.execute_reply.started":"2023-08-15T11:17:08.265724Z","shell.execute_reply":"2023-08-15T11:17:12.710975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reading in Audio Data\nWe have been provided `mp3` files for this competition. Other types of audio files can be: `wav`, `m4a`, `flac`, `ogg`","metadata":{}},{"cell_type":"code","source":"audio_files = glob(\"../input/bengaliai-speech/train_mp3s/*.mp3\")\nlen(audio_files)","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:12.715092Z","iopub.execute_input":"2023-08-15T11:17:12.715572Z","iopub.status.idle":"2023-08-15T11:17:18.880206Z","shell.execute_reply.started":"2023-08-15T11:17:12.715526Z","shell.execute_reply":"2023-08-15T11:17:18.879023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Play the first audio file\nipd.Audio(audio_files[0])","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:18.881784Z","iopub.execute_input":"2023-08-15T11:17:18.882729Z","iopub.status.idle":"2023-08-15T11:17:18.893125Z","shell.execute_reply.started":"2023-08-15T11:17:18.882688Z","shell.execute_reply":"2023-08-15T11:17:18.891999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Before we load the audio files we need to know few basic concepts\nA signal is a variation in a certain quantity over time. For audio, the quantity that varies is air pressure. How do we capture this information digitally? We can take samples of the air pressure over time. The rate at which we sample the data is called Sample Rate.\n<img src=\"https://jvbalen.github.io/figs/wavenet.gif\" />\n#### Frequency (Hz)\n- Frequency describes the differences of wave lengths.\n- We interperate frequency has high and low pitches.\n\n<img src = \"https://cdn.britannica.com/83/194283-004-37696A2F.jpg\" />\n\n#### Intensity (db / power)\n- Intensity describes the amplitude (height) of the wave.\n\n<img src = \"https://ars.els-cdn.com/content/image/3-s2.0-B9780124722804500162-f13-15-9780124722804.gif\" />\n\n#### Sample Rate\n- Sample rate is specific to how the computer reads in the audio file.\n- In digital audio, sound is represented as a series of discrete samples, where each sample is a snapshot of the audio signal at a specific point in time.\n- Sample rate is how many \"snapshots\" of sound are taken every second to turn it into digital data.\n- Think of it as the \"resolution\" of the audio.  Higher sample rate can potentially capture more detail in the audio signal.\n\n<img src = \"https://audiosolace.com/wp-content/uploads/2021/06/image-4-1024x512.png\" />","metadata":{}},{"cell_type":"code","source":"# Length of the audio\nprint(f\"Length of the audio in seconds: {librosa.get_duration(path=audio_files[0])}\")\n# load the 1st audio file. It returns the audio time series and sample rate\ny, sr = librosa.load(audio_files[0])\nprint(f'y: {y[:10]}')\nprint(f'shape: {y.shape}')\nprint(f'sample rate: {sr}')\nprint(f\"Length of the audio calculated from audio time series: {librosa.get_duration(y=y, sr = sr)}\")","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:18.896596Z","iopub.execute_input":"2023-08-15T11:17:18.897248Z","iopub.status.idle":"2023-08-15T11:17:18.914139Z","shell.execute_reply.started":"2023-08-15T11:17:18.897216Z","shell.execute_reply":"2023-08-15T11:17:18.912967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above code reads the audio signal and returns the audio as a time series - which basically contains the amplitudes of the samples. By default, sample rate = 22050 is provided while loading the audio file, but you can provide the desired value while loading it.\n\nWe can also calculate the lenght of the audio from the total number of samples and sample rate.\n\nAudio Length (seconds) = Total Number of Samples / Sample Rate\n                       = 65886/22050 = 2.9880272108843537\n","metadata":{}},{"cell_type":"code","source":"pd.Series(y).plot(figsize = (10,5),\n                  lw = 1,\n                  title = \"Raw Audio Example\",\n                  color = color_pal[0])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:18.915491Z","iopub.execute_input":"2023-08-15T11:17:18.915914Z","iopub.status.idle":"2023-08-15T11:17:19.373041Z","shell.execute_reply.started":"2023-08-15T11:17:18.915863Z","shell.execute_reply":"2023-08-15T11:17:19.371926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#trimming leading/lagging silence\ny_trimmed, _ = librosa.effects.trim(y, top_db = 20)\n\npd.Series(y_trimmed).plot(figsize = (10,5),\n                  lw = 1,\n                  title = \"Raw Audio Trimmed Example\",\n                  color = color_pal[1])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:19.374777Z","iopub.execute_input":"2023-08-15T11:17:19.375521Z","iopub.status.idle":"2023-08-15T11:17:19.843992Z","shell.execute_reply.started":"2023-08-15T11:17:19.375476Z","shell.execute_reply":"2023-08-15T11:17:19.842841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#zoomed in view\npd.Series(y[15000:15500]).plot(figsize = (10,5),\n                  lw = 1,\n                  title = \"Raw Audio Zoomed In Example\",\n                  color = color_pal[2])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:19.845661Z","iopub.execute_input":"2023-08-15T11:17:19.846804Z","iopub.status.idle":"2023-08-15T11:17:20.244703Z","shell.execute_reply.started":"2023-08-15T11:17:19.846713Z","shell.execute_reply":"2023-08-15T11:17:20.243780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Fourier Transform\nAn audio signal esspecially speech audio is comprised of several single-frequency sound waves. <br><br>\nRefer the above zoomed in plot of our audio example - it's not pure sinusoidal wave. It's a mix of different pure frequencies and we only see the final sum. So, how do we take a signal like this and decompose it to pure frequencies that make it up?\n<br><br>\n**Fourier Transform** is like a magical tool that allows us to decompose a signal into its individual frequencies and the frequency's amplitude.<br>\n In other words, it converts the signal from the time domain into the frequency domain. The result is called a **spectrum**.\n <br>\n<img src = \"https://miro.medium.com/v2/resize:fit:720/format:webp/1*xTYCtcx_7otHVu-uToI9dA.png\" />\n\n<br>\n\nThe **fast Fourier transform (FFT)** is an algorithm that can efficiently compute the Fourier transform. It is widely used in signal processing and we will use this algorithm on a windowed segment of our example audio.","metadata":{}},{"cell_type":"code","source":"# Default FFT window size\nn_fft = 2048 # FFT window size\nhop_length = 512 # number audio of frames between STFT columns (looks like a good default)\n\n# Short-time Fourier transform (STFT)\nD = np.abs(librosa.stft(y, n_fft = n_fft, hop_length = hop_length))\n\nprint('Shape of D object:', np.shape(D))","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:20.245900Z","iopub.execute_input":"2023-08-15T11:17:20.246737Z","iopub.status.idle":"2023-08-15T11:17:20.260392Z","shell.execute_reply.started":"2023-08-15T11:17:20.246686Z","shell.execute_reply":"2023-08-15T11:17:20.258827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (16, 6))\nplt.plot(D)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:32.207495Z","iopub.execute_input":"2023-08-15T11:17:32.207985Z","iopub.status.idle":"2023-08-15T11:17:32.753776Z","shell.execute_reply.started":"2023-08-15T11:17:32.207938Z","shell.execute_reply":"2023-08-15T11:17:32.752910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_fft = 2048\nft = np.abs(librosa.stft(y_trimmed[:n_fft], hop_length = n_fft+1))\nplt.plot(ft);\nplt.title('Spectrum');\nplt.xlabel('Frequency Bin');\nplt.ylabel('Amplitude');","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:41.858560Z","iopub.execute_input":"2023-08-15T11:17:41.859166Z","iopub.status.idle":"2023-08-15T11:17:42.287591Z","shell.execute_reply.started":"2023-08-15T11:17:41.859132Z","shell.execute_reply":"2023-08-15T11:17:42.286454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Spectrogram\nThe fast Fourier transform is a powerful tool that helps us understand the different frequencies in a signal. Imagine if the frequencies in a signal change as time goes by - like in music or speech. We call these signals **\"non-periodic.\"** To make sense of them, we need a way to show how their frequencies change over time. You might be thinking, \"Can't we analyze different parts of the signal with FFT?\" Absolutely! That's exactly what the **short-time Fourier transform** does. It breaks the signal into overlapping sections, uses FFT on each one, and creates something called a **spectrogram**. It might sound complex, but a visual example will make it clearer.\n<br>\n<img src =\"https://www.mathworks.com/help/dsp/ref/stft_output.png\" />\n<br>\nImagine a spectrogram like a stack of FFTs, showing how the loudness of a sound changes over time at different frequencies. The y-axis is converted to a log scale, and the color dimension is converted to decibels (you can think of this as the log scale of the amplitude).\n<br>\n<img src = \"https://source-separation.github.io/tutorial/_images/stft_process.png\" />","metadata":{}},{"cell_type":"code","source":"# Default FFT window size\nn_fft = 2048 # FFT window size\nhop_length = 512 # number audio of frames between STFT columns (looks like a good default)\n\nD = librosa.stft(y, n_fft=n_fft, hop_length=hop_length)\nS_db = librosa.amplitude_to_db(np.abs(D), ref = np.max)\nS_db.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:42.826481Z","iopub.execute_input":"2023-08-15T11:17:42.827634Z","iopub.status.idle":"2023-08-15T11:17:42.846314Z","shell.execute_reply.started":"2023-08-15T11:17:42.827584Z","shell.execute_reply":"2023-08-15T11:17:42.845327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the transformed audio\nfig, ax = plt.subplots(figsize = (10, 5))\nimg = librosa.display.specshow(S_db,\n                               x_axis= 'time',\n                               y_axis= 'log',\n                               ax=ax)\nax.set_title(\"Spectogram Example\", fontsize = 15)\nfig.colorbar(img, ax=ax, format=f\"%0.2f dB\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T11:17:44.063514Z","iopub.execute_input":"2023-08-15T11:17:44.063903Z","iopub.status.idle":"2023-08-15T11:17:44.690327Z","shell.execute_reply.started":"2023-08-15T11:17:44.063871Z","shell.execute_reply":"2023-08-15T11:17:44.689104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### How to read?<br>\n- Time goes from left to right.\n- Pitch(frequency) goes from bottom to top.\n- Color shows the amplitude of that particular frequency i.e. how loud different parts are. \n- Patterns and shapes help you see different sounds, like vowels in speech or musical notes. \n- It's like a colorful timeline that lets you \"see\" sound events over time.","metadata":{}},{"cell_type":"markdown","source":"# Mel Spectogram","metadata":{}},{"cell_type":"markdown","source":"**Humans perceive frequency logarithmically**: Research has found that humans don't hear frequencies on a linear scale. We're better at noticing differences in lower frequencies than higher ones. For instance, we can easily hear the difference between 500 and 1000 Hz, but it's tough for us to distinguish between 10,000 and 10,500 Hz, even though the gap between the two pairs is the same.\n<br><br>\nIn 1937, Stevens, Volkmann, and Newmann proposed a unit of pitch such that equal distances in pitch sounded equally distant to the listener. This is called the **mel scale.**\n<br>\n<img src = \"https://www.researchgate.net/publication/259479391/figure/fig3/AS:667809986125842@1536229716748/The-mel-scale-used-to-map-the-linear-frequency-scale-to-a-logarithmic-one.png\" />\n<br><br>\nA **Mel Spectrogram** is a spectrogram where the frequencies are converted to the mel scale.","metadata":{}},{"cell_type":"code","source":"S = librosa.feature.melspectrogram(y=y,\n                                   sr=sr,\n                                   n_mels= 128)\nS_db_mel = librosa.amplitude_to_db(S, ref = np.max)\nS_db_mel.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-15T09:43:43.781927Z","iopub.execute_input":"2023-08-15T09:43:43.782406Z","iopub.status.idle":"2023-08-15T09:43:43.806833Z","shell.execute_reply.started":"2023-08-15T09:43:43.782362Z","shell.execute_reply":"2023-08-15T09:43:43.805256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If you see the shape of the output spectrum is (128, 129).\n- Number of rows is equal to number of mel bands we have provided i.e. 128.\n- Number of columns is equal to the number of frames the audio signal is partitioned into i.e. Number of samples / hop length considering sample length is fixed during the whole process.\n- So for our example, Number of Frames = Number of samples / hop length = 65886/512 = 128.68359375 = 129","metadata":{}},{"cell_type":"code","source":"#Plot the mel spectogram\nfig, ax = plt.subplots(figsize = (10, 5))\nimg = librosa.display.specshow(S_db_mel,\n                               x_axis= 'time',\n                               y_axis= 'log',\n                               ax=ax)\nax.set_title(\"Mel Spectogram Example\", fontsize =  15)\nfig.colorbar(img, ax=ax, format=f\"%0.2f dB\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T08:59:42.590503Z","iopub.execute_input":"2023-08-15T08:59:42.590932Z","iopub.status.idle":"2023-08-15T08:59:43.154323Z","shell.execute_reply.started":"2023-08-15T08:59:42.590900Z","shell.execute_reply":"2023-08-15T08:59:43.153277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The mel spectrogram** is a valuable tool in speech recognition because it focuses on key aspects of speech, reduces noise impact, and simplifies data for machine learning. It extracts phonetic features, aligns with human hearing, and works well with deep learning models. This makes it essential for accurate and efficient speech analysis and recognition.","metadata":{}},{"cell_type":"markdown","source":"# Credits and Acknowledgements:\n- https://www.kaggle.com/code/robikscube/working-with-audio-in-python --> Working with audio data in python\n- https://medium.com/analytics-vidhya/understanding-the-mel-spectrogram-fca2afa2ce53 --> For understanding Fourier Transform and Mel Spectrogram in a simple and easy way\n- https://www.youtube.com/watch?v=spUNpyF58BY&ab_channel=3Blue1Brown --> For better understanding of Fourier Transform\n- https://github.com/musikalkemist/AudioSignalProcessingForML/blob/master/17%20-%20Mel%20Spectrogram%20Explained%20Easily/Mel%20Spectrograms%20Explained%20Easily.pdf--> For understanding Mel Spectrogram in easy way","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}