{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Simple Audio emotion exploration and classification ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"This is a notebook that explores the audio data in Audio Emotions dataset and shows a simple model to classify these images. \nI learned most of the code from a dude in youtube, his channel is - Valerio Velardo - The Sound of AI. Go check him out if you are interested in audio classification and sound AI stuff. \n\nThis notebook is mainly just to give some basic information about the data set. I did not go in depth in comparing the emotions and datasets, but mostly just gave you some stuff to get started more easily if you are planning to make audio classifier yourself!\n\nI am new to data science and classification, so this is just basic and im still learning and this notebook might not be up to some people's standards. \nIf you have any ideas and advice on how to upgrade it, feel free to hit me up.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Few imports:","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import librosa\nfrom librosa import display\nimport os\nimport glob \nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport time","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Loading files\nFirst we need to convert all the files from .wav to something we can use in our neural network.\nWe can make a list (lst) that contains the arrays of MFCC (40 dimensional vector that contains the information about audio files that we can use in training)\n\nThis does take a while, so if you are planning to run this, feel free to take a break and make yourself a coffe while this loads (Around 900seconds)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '/kaggle/input/audio-emotions/Emotions'\nlst = []\ni = -2\nstart_time = time.time()\n\nfor subdir, dirs, files in os.walk(path):\n  i=i+1\n  print(subdir)\n  print(i)\n  for file in files:\n\n        #Load librosa array, obtain mfcss, add them to array and then to list.\n        X, sample_rate = librosa.load(os.path.join(subdir,file), res_type='kaiser_fast')\n        mfccs = np.mean(librosa.feature.mfcc(y=X, sr=sample_rate, n_fft=4096, hop_length=256, n_mfcc=40).T,axis=0) \n        arr = mfccs, i\n        lst.append(arr) #Here we append the MFCCs to our list.\n\nprint(\"--- Data loaded. Loading time: %s seconds ---\" % (time.time() - start_time))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Lets assign paths to some variables to paths of files from different datasets, so we can compare them and explore further.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Lets explore the Neutral path. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Add paths and get signals.\n\nfile1='/kaggle/input/audio-emotions/Emotions/Neutral/03-02-01-01-02-02-20.wav'\nsignal1, sample_rate = librosa.load(file1, sr=22050)\n\nfile2='/kaggle/input/audio-emotions/Emotions/Neutral/1007_WSI_NEU_XX.wav'\nsignal2, sample_rate = librosa.load(file2, sr=22050)\n\nfile3='/kaggle/input/audio-emotions/Emotions/Neutral/n01.wav'\nsignal3, sample_rate = librosa.load(file3, sr=22050)\n\nfile4='/kaggle/input/audio-emotions/Emotions/Neutral/YAF_vote_neutral.wav'\nsignal4, sample_rate = librosa.load(file4, sr=22050)\n\nemotion='Neutral'\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Exploration\n\nLet's draw the simplest of visualisations WaveForms!\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(15,8))\n# WAVEFORM\n# display waveform\nplt.subplot(2, 2, 1)\nlibrosa.display.waveplot(signal1,sample_rate, alpha=0.4)\nplt.xlabel(\"Time (s)\")\nplt.ylabel(\"Amplitude\")\nplt.title(\"RAVDESS Waveform \"+emotion)\n\nplt.subplot(2, 2, 2)\nlibrosa.display.waveplot(signal2,sample_rate, alpha=0.4)\nplt.xlabel(\"Time (s)\")\nplt.ylabel(\"Amplitude\")\nplt.title(\"CREMA-D Waveform \"+emotion)\n\n\nfig = plt.figure(figsize=(15,8))\nplt.subplot(2, 2, 3)\nlibrosa.display.waveplot(signal3,sample_rate, alpha=0.4)\nplt.xlabel(\"Time (s)\")\nplt.ylabel(\"Amplitude\")\nplt.title(\"SAVEE Waveform \"+emotion)\n\nplt.subplot(2, 2, 4)\nlibrosa.display.waveplot(signal4,sample_rate, alpha=0.4)\nplt.xlabel(\"Time (s)\")\nplt.ylabel(\"Amplitude\")\nplt.title(\"TESS Waveform \"+emotion)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here we can see, that the aplitude difference is quite visible, witch is probably due to different recording environments, text and individual voice characterstics.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Now lets draw the power spectrums by performing Fourier transformations!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# FFT -> power spectrum\n# perform Fourier transform\nfft1 = np.fft.fft(signal1)\nfft2 = np.fft.fft(signal2)\nfft3 = np.fft.fft(signal3)\nfft4 = np.fft.fft(signal4)\n\n# calculate abs values on complex numbers to get magnitude\nspectrum1 = np.abs(fft1)\nspectrum2 = np.abs(fft2)\nspectrum3 = np.abs(fft3)\nspectrum4 = np.abs(fft4)\n\n# create frequency variable\nf1 = np.linspace(0, sample_rate, len(spectrum1))\nf2 = np.linspace(0, sample_rate, len(spectrum2))\nf3 = np.linspace(0, sample_rate, len(spectrum3))\nf4 = np.linspace(0, sample_rate, len(spectrum4))\n\n# take half of the spectrum and frequency\nleft_spectrum1 = spectrum1[:int(len(spectrum1)/2)]\nleft_f1 = f1[:int(len(spectrum1)/2)]\n# take half of the spectrum and frequency\nleft_spectrum2 = spectrum2[:int(len(spectrum2)/2)]\nleft_f2 = f2[:int(len(spectrum2)/2)]\n# take half of the spectrum and frequency\nleft_spectrum3 = spectrum3[:int(len(spectrum3)/2)]\nleft_f3 = f3[:int(len(spectrum3)/2)]\n# take half of the spectrum and frequency\nleft_spectrum4 = spectrum4[:int(len(spectrum4)/2)]\nleft_f4 = f4[:int(len(spectrum4)/2)]\n\nfig = plt.figure(figsize=(8,10))\nplt.subplot(2, 2, 1)\n# plot spectrum\nplt.plot(left_f1, left_spectrum1, alpha=0.4)\nplt.xlabel(\"Frequency\")\nplt.ylabel(\"Magnitude\")\nplt.title(\"RAVDESS  Power spectrum \"+emotion)\n\nplt.subplot(2, 2,2)\n# plot spectrum\nplt.plot(left_f2, left_spectrum2, alpha=0.4)\nplt.xlabel(\"Frequency\")\nplt.ylabel(\"Magnitude\")\nplt.title(\"CREMA-D Power spectrum \"+emotion)\n\n\nfig = plt.figure(figsize=(8,10))\n\nplt.subplot(2, 2, 3)\nplt.plot(left_f3, left_spectrum3, alpha=0.4)\nplt.xlabel(\"Frequency\")\nplt.ylabel(\"Magnitude\")\nplt.title(\"SAVEE Power spectrum \"+emotion)\n\nplt.subplot(2, 2, 4)\nplt.plot(left_f4, left_spectrum4, alpha=0.4)\nplt.xlabel(\"Frequency\")\nplt.ylabel(\"Magnitude\")\nplt.title(\"TESS Power spectrum \"+emotion)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here we can see that between the data sets the power spectrums are very different.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Lets draw the spectograms of theese files ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# STFT -> spectrogram\nhop_length =256 # in num. of samples\nn_fft = 4096 # window in num. of samples\n\n# calculate duration hop length and window in seconds\nhop_length_duration = float(hop_length)/sample_rate\nn_fft_duration = float(n_fft)/sample_rate\n\nprint(\"STFT hop length duration is: {}s\".format(hop_length_duration))\nprint(\"STFT window duration is: {}s\".format(n_fft_duration))\n\n# perform stft\nstft1 = librosa.stft(signal1, n_fft=n_fft, hop_length=hop_length)\nstft2 = librosa.stft(signal2, n_fft=n_fft, hop_length=hop_length)\nstft3 = librosa.stft(signal3, n_fft=n_fft, hop_length=hop_length)\nstft4 = librosa.stft(signal4, n_fft=n_fft, hop_length=hop_length)\n\n# calculate abs values on complex numbers to get magnitude\nspectrogram1 = np.abs(stft1)\nspectrogram2 = np.abs(stft2)\nspectrogram3 = np.abs(stft3)\nspectrogram4 = np.abs(stft4)\n\n\n# display spectrogram\n\n\nfig = plt.figure(figsize=(8,6))\nplt.subplot(2, 2, 1)\nlibrosa.display.specshow(spectrogram1, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"Frequency\")\nplt.colorbar()\nplt.title(\"RAVDESS Spectrogram \"+emotion)\n\nplt.subplot(2, 2,2)\nlibrosa.display.specshow(spectrogram2, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"Frequency\")\nplt.colorbar()\nplt.title(\"CREMA-D Spectrogram \"+emotion)\n\n\nfig = plt.figure(figsize=(8,6))\nplt.subplot(2, 2, 3)\nlibrosa.display.specshow(spectrogram3, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"Frequency\")\nplt.colorbar()\nplt.title(\"SAVEE Spectrogram \"+emotion)\n\nplt.subplot(2, 2, 4)\nlibrosa.display.specshow(spectrogram4, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"Frequency\")\nplt.colorbar()\nplt.title(\"TESS Spectrogram \"+emotion)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This doesn't show much does it. Lets try to apply logarhithm to cast amplitude to decibels!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"log_spectrogram1 = librosa.amplitude_to_db(spectrogram1)\nlog_spectrogram2 = librosa.amplitude_to_db(spectrogram2)\nlog_spectrogram3 = librosa.amplitude_to_db(spectrogram3)\nlog_spectrogram4 = librosa.amplitude_to_db(spectrogram4)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Lets display Spectogramms now!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(8,6))\nplt.subplot(2, 2, 1)\nlibrosa.display.specshow(log_spectrogram1, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"Frequency\")\nplt.colorbar(format=\"%+2.0f dB\")\nplt.title(\"RAVDESS Spectogramm (dB) \"+emotion)\n\nplt.subplot(2, 2,2)\nlibrosa.display.specshow(log_spectrogram2, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"Frequency\")\nplt.colorbar(format=\"%+2.0f dB\")\nplt.title(\"CREMA-D Spectogramm (dB) \"+emotion)\n\n\nfig = plt.figure(figsize=(8,6))\nplt.subplot(2, 2, 3)\nlibrosa.display.specshow(log_spectrogram3, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"Frequency\")\nplt.colorbar(format=\"%+2.0f dB\")\nplt.title(\"SAVEE Spectogramm (dB) \"+emotion)\n\nplt.subplot(2, 2, 4)\nlibrosa.display.specshow(log_spectrogram4, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"Frequency\")\nplt.colorbar(format=\"%+2.0f dB\")\nplt.title(\"TESS Spectogramm (dB) \"+emotion)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Ah, much better. I've seen some classificators that use theese images as input to convolutional neural networks, so that might be a fun thing to try!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Now, lets actually see how the MFCC look! ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# MFCCs\n# extract 13 MFCCs\nMFCCs1 = librosa.feature.mfcc(signal1, sample_rate, n_fft=4096, hop_length=256, n_mfcc=40)\nMFCCs2 = librosa.feature.mfcc(signal2, sample_rate, n_fft=4096, hop_length=256, n_mfcc=40)\nMFCCs3 = librosa.feature.mfcc(signal3, sample_rate, n_fft=4096, hop_length=256, n_mfcc=40)\nMFCCs4 = librosa.feature.mfcc(signal4, sample_rate, n_fft=4096, hop_length=256, n_mfcc=40)\n\n# display MFCCs\nhop_length=256\n\n\nfig = plt.figure(figsize=(8,6))\nplt.subplot(2, 2, 1)\nlibrosa.display.specshow(MFCCs1, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"MFCC coefficients\")\nplt.colorbar()\nplt.title(\"RAVDESS MFCCs \"+emotion)\n\nplt.subplot(2, 2,2)\nlibrosa.display.specshow(MFCCs2, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"MFCC coefficients\")\nplt.colorbar()\nplt.title(\"CREMA-D MFCCs \"+emotion)\n\n\nfig = plt.figure(figsize=(8,6))\nplt.subplot(2, 2, 3)\nlibrosa.display.specshow(MFCCs3, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"MFCC coefficients\")\nplt.colorbar()\nplt.title(\"SAVEE MFCCs \"+emotion)\n\n\nplt.subplot(2, 2, 4)\nlibrosa.display.specshow(MFCCs4, sr=sample_rate, hop_length=hop_length)\nplt.xlabel(\"Time\")\nplt.ylabel(\"MFCC coefficients\")\nplt.colorbar()\nplt.title(\"TESS MFCCs \"+emotion)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So that is basically the information that we are going to be inputing in our neural network!\nIf you don't know what MFCCS are here's a wiki page - https://en.wikipedia.org/wiki/Mel-frequency_cepstrum#:~:text=Mel%2Dfrequency%20cepstral%20coefficients%20(MFCCs,%2Da%2Dspectrum%22).","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Let's build the classificator!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"I din't add data augmentation, but that could be a fun thing to add in future versions, to see how results change!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Lets make the X and y and lets split them into test and train!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Creating X and y: zip makes a list of all the first elements, and a list of all the second elements.\nX, y = zip(*lst)\nimport numpy as np\nX = np.asarray(X)\ny = np.asarray(y)\n\nfrom sklearn.model_selection import train_test_split\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.1, random_state=42)\n\n#As always we need to expand the dimensions, so we can input the data to NN.\nx_traincnn = np.expand_dims(X_train, axis=2) \nx_testcnn = np.expand_dims(X_test, axis=2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Some imports! Not sure if we need all of them, but i did need them at some point so maybe they stayed. If they are not needed, let me know, and ill remove them!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import keras\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom keras.preprocessing import sequence\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Embedding\nfrom keras.utils import to_categorical\nfrom keras.layers import Input, Flatten, Dropout, Activation\nfrom keras.layers import Conv1D, MaxPooling1D\nfrom keras.models import Model\nfrom keras.callbacks import ModelCheckpoint\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's build a simple Artificial neural network with some regularization and dropouts! Im still new to neural networks so this is definetly a sub-optimal solution! Feel free to change all of this.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Simple model\n\nmodel = keras.Sequential([\n\n        # input layer\n        keras.layers.Flatten(input_shape=(40, 1)),\n\n        # 1st dense layer\n        keras.layers.Dense(256, activation='relu', kernel_regularizer=keras.regularizers.l2(0.1)),\n        keras.layers.Dropout(0.5),\n\n        # 2nd dense layer\n        keras.layers.Dense(256, activation='relu', kernel_regularizer=keras.regularizers.l2(0.01)),\n        keras.layers.Dropout(0.5),\n\n        # 3rd dense layer\n        keras.layers.Dense(128, activation='relu', kernel_regularizer=keras.regularizers.l2(0.001)),\n        keras.layers.Dropout(0.3),\n\n        # output layer\n        keras.layers.Dense(10, activation='softmax')\n    ])\n\n    # compile model\noptimiser = keras.optimizers.Adam(learning_rate=0.00001)\nmodel.compile(optimizer=optimiser,\n                  loss='sparse_categorical_crossentropy',\n                  metrics=['accuracy'])\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Lets fit and visualize our training data!\nThis does take a while.","execution_count":null},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"cnnhistory=model.fit(x_traincnn, y_train, batch_size=64, epochs=700, validation_data=(x_testcnn, y_test))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nplt.plot(cnnhistory.history['loss'])\nplt.plot(cnnhistory.history['val_loss'])\nplt.title('Loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['Train', 'Test'], loc='upper left')\nplt.show()\n\n\n\nplt.plot(cnnhistory.history['accuracy'])\nplt.plot(cnnhistory.history['val_accuracy'])\nplt.title('Accuracy')\nplt.ylabel('acc')\nplt.xlabel('epoch')\nplt.legend(['Train', 'Test'], loc='upper left')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Results!\n\nLets see how we did.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Lets get a classification report and lets see in more detail which emotions are easier and which are harder to classify.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.metrics import classification_report\n\npredictions = model.predict_classes(x_testcnn)\ny_test = y_test.astype(int)\nreport = classification_report(y_test, predictions)\nprint(report)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's make a classification matrix to see more in depth about classifications.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"\nfrom sklearn.metrics import confusion_matrix\nmatrix = confusion_matrix(y_test, predictions)\nprint (matrix)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"That looks, too boring. Lets get seaborn heatmap to help us with that!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import seaborn as sns\nplt.figure(figsize=(10,8))\nsns.heatmap(matrix, annot=True, fmt=\"d\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here we can make a lot of different conclusions, like 2nd class is often classified as 0 and so on.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"If you want you can also save the model!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"model_name = 'EmotionClassificationModel.h5'\nsave_dir = '/kaggle/working'\n# Save model and weights\nif not os.path.isdir(save_dir):\n    os.makedirs(save_dir)\nmodel_path = os.path.join(save_dir, model_name)\nmodel.save(model_path)\nprint('Saved trained model at %s ' % model_path)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}