{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is Spectrograms?\n\nA spectrogram is a visual representation of the spectrum of frequencies of a signal as it varies with time. In audio machine learning problems, spectrograms are computed for preprocessing audio signals.\n\n# Why Spectrograms?\n\nYes, you can train your machine learning models just using raw audio. But spectrograms tend to be more easier to work with as they contain richer features than raw audio clips.\n\nThere are many types of spectrogram. Few of the popular ones are listed below.\n1. [STFT (Short Time Fourier Transform)](https://en.wikipedia.org/wiki/Short-time_Fourier_transform#:~:text=The%20Short%2Dtime%20Fourier%20transform,as%20it%20changes%20over%20time.)\n2. [Mel Spectrogram](https://medium.com/analytics-vidhya/understanding-the-mel-spectrogram-fca2afa2ce53)\n3. [MFCC (Mel-frequency cepstrum)](https://en.wikipedia.org/wiki/Mel-frequency_cepstrum#:~:text=Mel%2Dfrequency%20cepstral%20coefficients%20(MFCCs,%2Da%2Dspectrum%22).)\n\nLuckily, implementations of the algorithm for computing them can be found in <code>librosa</code>!","metadata":{}},{"cell_type":"code","source":"import warnings\n\nimport librosa\nimport librosa.display\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport IPython.display as ipd\n\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-02T12:35:10.483193Z","iopub.execute_input":"2022-07-02T12:35:10.483791Z","iopub.status.idle":"2022-07-02T12:35:13.198290Z","shell.execute_reply.started":"2022-07-02T12:35:10.483686Z","shell.execute_reply":"2022-07-02T12:35:13.196655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# librosa.load() takes the path and returns a numpy array and the sample rate of the audio.\n# You can display the audio using IPython\n\nAUDIO_PATH = '../input/dlsprint/train_files/common_voice_bn_30614352.mp3'\naudio, sr = librosa.load(AUDIO_PATH)\n\nprint('Shape of the audio: ', audio.shape)\nprint('Sample rate of the audio: ', sr)\n\nipd.display(ipd.Audio(data=audio, rate=sr))","metadata":{"execution":{"iopub.status.busy":"2022-07-02T12:35:13.200230Z","iopub.execute_input":"2022-07-02T12:35:13.200677Z","iopub.status.idle":"2022-07-02T12:35:15.524452Z","shell.execute_reply.started":"2022-07-02T12:35:13.200627Z","shell.execute_reply":"2022-07-02T12:35:15.522925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to librosa, we can compute our choice of spectrogram in a single line.","metadata":{}},{"cell_type":"code","source":"stft_feature = np.abs(librosa.stft(y=audio)) ** 2\nprint('Shape of STFT features: ', stft_feature.shape)\n\nms_feature = librosa.feature.melspectrogram(y=audio, sr=sr)\nprint('Shape of Mel Spectrogram features: ', ms_feature.shape)\n\nmfcc_feature = librosa.feature.mfcc(y=audio, sr=sr)\nprint('Shape of MFCC features: ', mfcc_feature.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-02T12:35:15.526123Z","iopub.execute_input":"2022-07-02T12:35:15.526505Z","iopub.status.idle":"2022-07-02T12:35:15.586161Z","shell.execute_reply.started":"2022-07-02T12:35:15.526470Z","shell.execute_reply":"2022-07-02T12:35:15.584829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can notice two things here.\n\n1. A 1D numpy array has been transformed into 2D numpy array.\n2. Mel Spectrogram and MFCC helps in reducing the size of the input.\n\nWe can visualize the spectrograms using <code>librosa</code>.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\nms_dB = librosa.power_to_db(ms_feature, ref=np.max)\nimg = librosa.display.specshow(ms_dB, x_axis='time', y_axis='mel', sr=sr, ax=ax)\nfig.colorbar(img, ax=ax, format='%+2.0f dB')\nax.set(title='Mel-frequency spectrogram');","metadata":{"execution":{"iopub.status.busy":"2022-07-02T12:35:15.589647Z","iopub.execute_input":"2022-07-02T12:35:15.591060Z","iopub.status.idle":"2022-07-02T12:35:16.010445Z","shell.execute_reply.started":"2022-07-02T12:35:15.591002Z","shell.execute_reply":"2022-07-02T12:35:16.009105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Spectrograms and CNN\n\nAs spectrograms are 2D matrix, we can treat them like an image. That's a great thing! Because now we can use state of the art computer vision models (like ResNet) for our audio task. I guess now you all have a reason to love spectrograms.\n\nHope you have learned something new. Best of luck comrade! 😊","metadata":{}}]}