{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on July 17, 2023. By Marília Prata, mpwolke","metadata":{"_kg_hide-input":false}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport plotly.express as px\nimport plotly.graph_objs as go\nimport itertools\n\nfrom IPython.display import Markdown, display\nimport matplotlib.pyplot as plt\nfrom plotly.subplots import make_subplots\nfrom wordcloud import WordCloud\nimport re\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=Warning)","metadata":{"execution":{"iopub.status.busy":"2023-07-17T22:44:37.834863Z","iopub.execute_input":"2023-07-17T22:44:37.835228Z","iopub.status.idle":"2023-07-17T22:44:39.705884Z","shell.execute_reply.started":"2023-07-17T22:44:37.835201Z","shell.execute_reply":"2023-07-17T22:44:39.705196Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"@article{rakib2023oodspeech,\n\ntitle={{OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking},\n\nauthor={Rakib, Fazle Rabbi and Dip, Souhardya Saha and Alam, Samiul and Tasnim, Nazia and Shihab, Md Istiak Hossain and Ansary, Md Nazmuddoha and Hossen, Syed Mobassir and Meghla, Marsia Haque and Mamun, Mamunur and Sadeque, Farig and others},\n\njournal={Proc. Interspeech 2023},\n\nyear={2023} }","metadata":{}},{"cell_type":"markdown","source":"![](https://youngbutterfly.in/wp-content/uploads/2021/07/Learn-and-Speak-Bengali-Language.png)youngbutterfly","metadata":{}},{"cell_type":"markdown","source":"#Bengali Language - Bangla\n\n\"Bengali generally known by its endonym Bangla, is an Indo-Aryan language native to the Bengal region of South Asia. With approximately 235 million native speakers and another 40 million as second language speakers, Bengali is the sixth most spoken native language and the seventh most spoken language by the total number of speakers in the world. Bengali is the fifth most spoken Indo-European language.\"\n\nhttps://en.wikipedia.org/wiki/Bengali_language","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"../input/bengaliai-speech/train.csv\")\ntrain.head()\n#test = pd.read_csv(\"../input/chaii-hindi-and-tamil-question-answering/test.csv\")\n#sample = pd.read_csv(\"../input/chaii-hindi-and-tamil-question-answering/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-07-17T22:46:13.044427Z","iopub.execute_input":"2023-07-17T22:46:13.044748Z","iopub.status.idle":"2023-07-17T22:46:17.112871Z","shell.execute_reply.started":"2023-07-17T22:46:13.044722Z","shell.execute_reply":"2023-07-17T22:46:17.111760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by Lucas Abrahão https://www.kaggle.com/lucasabrahao/trabalho-manufatura-an-lise-de-dados-no-brasil\n\ntrain[\"split\"].value_counts().plot.barh(figsize = (8,6),color=['red', 'green'], title='BengaliAI Speech Recognition');","metadata":{"execution":{"iopub.status.busy":"2023-07-18T01:26:34.828611Z","iopub.execute_input":"2023-07-18T01:26:34.829656Z","iopub.status.idle":"2023-07-18T01:26:35.037511Z","shell.execute_reply.started":"2023-07-18T01:26:34.829618Z","shell.execute_reply":"2023-07-18T01:26:35.036240Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = pd.read_csv(\"../input/bengaliai-speech/sample_submission.csv\")\nsample.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-17T22:52:02.972236Z","iopub.execute_input":"2023-07-17T22:52:02.972600Z","iopub.status.idle":"2023-07-17T22:52:02.984129Z","shell.execute_reply.started":"2023-07-17T22:52:02.972560Z","shell.execute_reply":"2023-07-17T22:52:02.983086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#OOD-Speech (Out-of Distribution Speech)\n\nOOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking\n\nAuthors: Fazle Rabbi Rakib, Souhardya Saha Dip, Samiul Alam, Nazia Tasnim, Md. Istiak Hossain Shihab, Md. Nazmuddoha Ansary, Syed Mobassir Hossen, Marsia Haque Meghla, Mamunur Mamun, Farig Sadeque, Sayma Sultana Chowdhury, Tahsin Reasat, Asif Sushmit, Ahmed Imtiaz Humayun\n\n\"The authors presented OOD-Speech, the first out-of-distribution (OOD) benchmarking dataset for Bengali automatic speech recognition (ASR). Being one of the most spoken languages globally, Bengali portrays large diversity in dialects and prosodic features, which demands ASR frameworks to be robust towards distribution shifts.\"\n\n\"For example, islamic religious sermons in Bengali are delivered with a tonality that is significantly different from regular speech. Their training dataset is collected via massively online crowdsourcing campaigns which resulted in 1177.94 hours collected and curated from 22,645 native Bengali speakers from South Asia.\"\n\n\"Their test dataset comprises 23.03 hours of speech collected and manually annotated from 17 different sources, e.g., Bengali TV drama, Audiobook, Talk show, Online class, and Islamic sermons to name a few. OOD-Speech is jointly the largest publicly available speech dataset, as well as the first out-of-distribution ASR benchmarking dataset for Bengali.\"\n\nhttps://arxiv.org/abs/2305.09688 arXiv:2305.09688","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install langdetect # Language Detection\n!pip install bnlp_toolkit # For Bangla Word Cloud\n!wget https://www.omicronlab.com/download/fonts/kalpurush.ttf # Bangla Font For the Word Cloud","metadata":{"execution":{"iopub.status.busy":"2023-07-17T22:57:52.710719Z","iopub.execute_input":"2023-07-17T22:57:52.711150Z","iopub.status.idle":"2023-07-17T22:58:17.237072Z","shell.execute_reply.started":"2023-07-17T22:57:52.711114Z","shell.execute_reply":"2023-07-17T22:58:17.235480Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def printmd(string):\n  display(Markdown(string))\n\nfrom langdetect import detect\nimport unicodedata\nimport html","metadata":{"execution":{"iopub.status.busy":"2023-07-17T22:58:30.893178Z","iopub.execute_input":"2023-07-17T22:58:30.893509Z","iopub.status.idle":"2023-07-17T22:58:30.909385Z","shell.execute_reply.started":"2023-07-17T22:58:30.893483Z","shell.execute_reply":"2023-07-17T22:58:30.908442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by Tanjima Nasreen Jenia  https://www.kaggle.com/tanjimanasreenjenia/bangladeshi-restaurants-analytics/notebook\n\nfrom bnlp.corpus import stopwords, punctuations\nregex = r\"[\\u0980-\\u09FF]+\" \ndata = train\n\nwc = WordCloud(background_color = 'green',\n                      height =2000,\n                      width = 2000,\n                      colormap = \"Reds\",\n                      font_path=\"./kalpurush.ttf\"\n                     ).generate(str(train[\"sentence\"]))                                                                                                                        \n\nplt.rcParams['figure.figsize'] = (12,12)\nplt.imshow(wc, interpolation=\"bilinear\")\nplt.axis('off')\nplt.show()\nresult = wc.to_file(\"Bangla_word_cloud.png\")\nprintmd(\"BengaliAI Sentences\")","metadata":{"execution":{"iopub.status.busy":"2023-07-17T23:01:44.101404Z","iopub.execute_input":"2023-07-17T23:01:44.101748Z","iopub.status.idle":"2023-07-17T23:01:51.364627Z","shell.execute_reply.started":"2023-07-17T23:01:44.101720Z","shell.execute_reply":"2023-07-17T23:01:51.363712Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by Tanjima Nasreen Jenia  https://www.kaggle.com/tanjimanasreenjenia/bangladeshi-restaurants-analytics/notebook\n\nfrom bnlp.corpus import stopwords, punctuations\nregex = r\"[\\u0980-\\u09FF]+\" \ndata = sample\n\nwc = WordCloud(background_color = 'white',               \n                      height =2000,\n                      width = 2000,               \n                      font_path=\"./kalpurush.ttf\",\n                      colormap ='Purples'\n                     ).generate(str(sample[\"sentence\"]))                                                                                                                        \n\nplt.rcParams['figure.figsize'] = (12,12)\nplt.imshow(wc, interpolation=\"bilinear\")\nplt.axis('off')\nplt.show()\nresult = wc.to_file(\"Bangla_word_cloud.png\")\nprintmd(\"Bengali Speech\")","metadata":{"execution":{"iopub.status.busy":"2023-07-17T23:04:36.690812Z","iopub.execute_input":"2023-07-17T23:04:36.691169Z","iopub.status.idle":"2023-07-17T23:04:42.951283Z","shell.execute_reply.started":"2023-07-17T23:04:36.691144Z","shell.execute_reply":"2023-07-17T23:04:42.950224Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"! pip install pydub","metadata":{"execution":{"iopub.status.busy":"2023-07-17T23:16:17.221292Z","iopub.execute_input":"2023-07-17T23:16:17.221651Z","iopub.status.idle":"2023-07-17T23:16:26.061692Z","shell.execute_reply.started":"2023-07-17T23:16:17.221625Z","shell.execute_reply":"2023-07-17T23:16:26.060570Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Slang Profanity (mp3) I choose 2 because it's not in order","metadata":{}},{"cell_type":"code","source":"bengali = []\ncwd = '/kaggle/input/bengaliai-speech/examples/'\n\nfor dirname, _, filenames in os.walk('/kaggle/input/bengaliai-speech/examples/'):\n    for filename in filenames:\n        #print(os.path.join(dirname, filename))\n        bengali.append(filename)\n#data = pd.read_csv('/kaggle/input/us-election-2020-presidential-debates/us_election_2020_1st_presidential_debate.csv')\nbengali.pop(2)#Original is 0 zero though the 1st mp3 is on the 13th line They are Not in order here.\n#I chose Slang Profanity manualy ","metadata":{"execution":{"iopub.status.busy":"2023-07-17T23:29:41.437012Z","iopub.execute_input":"2023-07-17T23:29:41.437393Z","iopub.status.idle":"2023-07-17T23:29:41.446637Z","shell.execute_reply.started":"2023-07-17T23:29:41.437363Z","shell.execute_reply":"2023-07-17T23:29:41.445530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pydub import AudioSegment\nimport IPython\n\n# We will listen to this file:\n# 213_1p5_Pr_mc_AKGC417L.wav\nfile = '/kaggle/input/bengaliai-speech/Stage Drama Jatra.wav'\nprint(cwd+bengali[13])#Original is zero\nIPython.display.Audio(cwd+bengali[13])#Original is 2","metadata":{"execution":{"iopub.status.busy":"2023-07-17T23:55:34.074051Z","iopub.execute_input":"2023-07-17T23:55:34.074418Z","iopub.status.idle":"2023-07-17T23:55:34.152432Z","shell.execute_reply.started":"2023-07-17T23:55:34.074392Z","shell.execute_reply":"2023-07-17T23:55:34.151002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://www.kaggle.com/rakibilly/extract-audio-starter\nimport subprocess\nimport glob\nimport os\nfrom pathlib import Path\nimport shutil\nfrom zipfile import ZipFile","metadata":{"execution":{"iopub.status.busy":"2023-07-17T23:46:08.256791Z","iopub.execute_input":"2023-07-17T23:46:08.257150Z","iopub.status.idle":"2023-07-17T23:46:08.261852Z","shell.execute_reply.started":"2023-07-17T23:46:08.257123Z","shell.execute_reply":"2023-07-17T23:46:08.260972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Added ffmpeg Static build though it didn't help me in any step.","metadata":{}},{"cell_type":"code","source":"! tar xvf ../input/ffmpeg-static-build/ffmpeg-git-amd64-static.tar.xz","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:00:13.817610Z","iopub.execute_input":"2023-07-18T00:00:13.817947Z","iopub.status.idle":"2023-07-18T00:00:17.510231Z","shell.execute_reply.started":"2023-07-18T00:00:13.817921Z","shell.execute_reply":"2023-07-18T00:00:17.509226Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert MP3s to WAV for easy conversion to numpy arrays:\noutput_format = 'wav'  # can also use aac, wav, etc\noutput_dir = Path(f\"{output_format}s\")\nPath(output_dir).mkdir(exist_ok=True, parents=True)\n\n#Only do first 50 because notebook memory limitations...\nfor song in songs[:50]:\n    file = cwd+song\n    file_name = song.replace(\".mp3\",\"\")\n    command = f\"../working/ffmpeg-git-20191209-amd64-static/ffmpeg -i {file} -ab 192000 -ac 2 -ar 44100 -vn {output_dir/file_name}.{output_format}\"\n    subprocess.call(command, shell=True)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:00:44.168340Z","iopub.execute_input":"2023-07-18T00:00:44.168703Z","iopub.status.idle":"2023-07-18T00:00:44.174802Z","shell.execute_reply.started":"2023-07-18T00:00:44.168674Z","shell.execute_reply":"2023-07-18T00:00:44.174085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from scipy.io.wavfile import read, write\n#a = read(\"adios.wav\")\nwavs = []\nnp_arrays = []\nfor dirname, _, filenames in os.walk('/kaggle/working/wavs/'):\n    for filename in filenames:\n        wav_file = dirname+filename\n        #print(wav_file)\n        wavs.append(wav_file)\n        try:\n            fs, io_file = read(wav_file)\n        except ValueError:\n            continue\n        data = np.array(io_file,dtype=float)\n        wav_info= {\n            'name': filename,\n            'fs' : fs,\n            'left': data[:,0],\n            'right': data[:,1]\n        }\n        \n        np_arrays.append(wav_info)\n\nprint(\"Succesfully converted: \"+str(len(np_arrays)))","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:01:09.626592Z","iopub.execute_input":"2023-07-18T00:01:09.626928Z","iopub.status.idle":"2023-07-18T00:01:09.635132Z","shell.execute_reply.started":"2023-07-18T00:01:09.626902Z","shell.execute_reply":"2023-07-18T00:01:09.633857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#There were only 2 MP3 files the others were already wav files, so ZERO converted above ","metadata":{}},{"cell_type":"code","source":"# Other  \nimport librosa\nimport librosa.display\nimport json\nimport tensorflow as tf\nfrom matplotlib.pyplot import specgram\nimport glob \nimport os\nfrom tqdm import tqdm\nimport pickle\nimport IPython.display as ipd  # To play sound in the notebook","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-07-18T00:07:24.599299Z","iopub.execute_input":"2023-07-18T00:07:24.599662Z","iopub.status.idle":"2023-07-18T00:07:32.120647Z","shell.execute_reply.started":"2023-07-18T00:07:24.599636Z","shell.execute_reply":"2023-07-18T00:07:32.119341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Poem Recital","metadata":{}},{"cell_type":"code","source":"#By Eu Jin Lok https://www.kaggle.com/ejlok1/audio-emotion-part-5-data-augmentation/notebook\n\n# Use one audio file in previous parts again\nfname = '/kaggle/input/bengaliai-speech/examples/Poem Recital.wav'  \ndata, sampling_rate = librosa.load(fname)\nplt.figure(figsize=(15, 5))\nlibrosa.display.waveshow(data, sr=sampling_rate)#librosa.display' has no attribute 'waveplot'\n\n# Paly it again to refresh our memory\nipd.Audio(data, rate=sampling_rate)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:11:16.340583Z","iopub.execute_input":"2023-07-18T00:11:16.340935Z","iopub.status.idle":"2023-07-18T00:11:17.029892Z","shell.execute_reply.started":"2023-07-18T00:11:16.340907Z","shell.execute_reply":"2023-07-18T00:11:17.028887Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"general_path = '../input/bengaliai-speech/examples'","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:21:31.013385Z","iopub.execute_input":"2023-07-18T00:21:31.013736Z","iopub.status.idle":"2023-07-18T00:21:31.019062Z","shell.execute_reply.started":"2023-07-18T00:21:31.013706Z","shell.execute_reply":"2023-07-18T00:21:31.017804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Importing 1 file\ny, sr = librosa.load(f'{general_path}/Poem Recital.wav')\n\nprint('y:', y, '\\n')\nprint('y shape:', np.shape(y), '\\n')\nprint('Sample Rate (KHz):', sr, '\\n')\n\n# Verify length of the audio\nprint('Check Len of Audio:', 661794/22050)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:22:16.592220Z","iopub.execute_input":"2023-07-18T00:22:16.592565Z","iopub.status.idle":"2023-07-18T00:22:16.669836Z","shell.execute_reply.started":"2023-07-18T00:22:16.592540Z","shell.execute_reply":"2023-07-18T00:22:16.668470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Trim leading and trailing silence from an audio signal (silence before and after the actual audio)\naudio_file, _ = librosa.effects.trim(y)\n\n# the result is an numpy ndarray\nprint('Audio File:', audio_file, '\\n')\nprint('Audio File shape:', np.shape(audio_file))","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:23:03.668575Z","iopub.execute_input":"2023-07-18T00:23:03.668902Z","iopub.status.idle":"2023-07-18T00:23:04.922951Z","shell.execute_reply.started":"2023-07-18T00:23:03.668877Z","shell.execute_reply":"2023-07-18T00:23:04.921742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#2D Representation: Sound Waves","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\nplt.figure(figsize = (16, 6))\nlibrosa.display.waveshow(y = audio_file, sr = sr, color = \"#A300F9\");#Waveplot is waveshow\nplt.title(\"Bengali Poem Recital\", fontsize = 23);","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:24:11.556243Z","iopub.execute_input":"2023-07-18T00:24:11.556582Z","iopub.status.idle":"2023-07-18T00:24:12.255913Z","shell.execute_reply.started":"2023-07-18T00:24:11.556552Z","shell.execute_reply":"2023-07-18T00:24:12.254309Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Fourier Transform\n\n\"Function that gets a signal in the time domain as input, and outputs its decomposition into frequencies\nTransform both the y-axis (frequency) to log scale, and the “color” axis (amplitude) to Decibels, which is approx. the log scale of amplitudes.\"","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Default FFT window size\nn_fft = 2048 # FFT window size\nhop_length = 512 # number audio of frames between STFT columns (looks like a good default)\n\n# Short-time Fourier transform (STFT)\nD = np.abs(librosa.stft(audio_file, n_fft = n_fft, hop_length = hop_length))\n\nprint('Shape of D object:', np.shape(D))","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:24:54.797420Z","iopub.execute_input":"2023-07-18T00:24:54.797805Z","iopub.status.idle":"2023-07-18T00:24:54.840482Z","shell.execute_reply.started":"2023-07-18T00:24:54.797780Z","shell.execute_reply":"2023-07-18T00:24:54.839194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (16, 6))\nplt.plot(D);","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:25:10.143282Z","iopub.execute_input":"2023-07-18T00:25:10.143640Z","iopub.status.idle":"2023-07-18T00:25:13.468196Z","shell.execute_reply.started":"2023-07-18T00:25:10.143612Z","shell.execute_reply":"2023-07-18T00:25:13.467381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Spectrogram","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Convert an amplitude spectrogram to Decibels-scaled spectrogram.\nDB = librosa.amplitude_to_db(D, ref = np.max)\n\n# Creating the Spectogram\nplt.figure(figsize = (16, 6))\nlibrosa.display.specshow(DB, sr = sr, hop_length = hop_length, x_axis = 'time', y_axis = 'log',\n                        cmap = 'cool')\nplt.colorbar();","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:25:33.251746Z","iopub.execute_input":"2023-07-18T00:25:33.252084Z","iopub.status.idle":"2023-07-18T00:25:35.068270Z","shell.execute_reply.started":"2023-07-18T00:25:33.252059Z","shell.execute_reply":"2023-07-18T00:25:35.067545Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install librosa==0.9.2","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:44:23.287233Z","iopub.execute_input":"2023-07-18T00:44:23.287612Z","iopub.status.idle":"2023-07-18T00:44:33.215724Z","shell.execute_reply.started":"2023-07-18T00:44:23.287584Z","shell.execute_reply":"2023-07-18T00:44:33.214578Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Mel Spectrogram \n\n\"The Mel Scale, mathematically speaking, is the result of some non-linear transformation of the frequency scale. The Mel Spectrogram is a normal Spectrogram, but with a Mel Scale on the y axis.\"","metadata":{}},{"cell_type":"code","source":"#https://librosa.org/doc/main/generated/librosa.feature.melspectrogram.html\n\ny, sr = librosa.load(f'{general_path}/Puthi Literature.wav')\ny, _ = librosa.effects.trim(y)\n\n\nS = librosa.feature.melspectrogram(y=y, sr=sr)#y=y fixed passes 0 argument error\nS_DB = librosa.amplitude_to_db(S, ref=np.max)\nplt.figure(figsize = (16, 6))\nlibrosa.display.specshow(S_DB, sr=sr, hop_length=hop_length, x_axis = 'time', y_axis = 'log',\n                        cmap = 'cool');\nplt.colorbar();\nplt.title(\"Puthi Literature Mel Spectrogram\", fontsize = 23);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-18T01:12:32.758951Z","iopub.execute_input":"2023-07-18T01:12:32.759324Z","iopub.status.idle":"2023-07-18T01:12:33.475504Z","shell.execute_reply.started":"2023-07-18T01:12:32.759296Z","shell.execute_reply":"2023-07-18T01:12:33.474101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Classical Mel Spectrogram","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\ny, sr = librosa.load(f'{general_path}/Bangladeshi TV Drama.wav')\ny, _ = librosa.effects.trim(y)\n\n\nS = librosa.feature.melspectrogram(y=y, sr=sr)\nS_DB = librosa.amplitude_to_db(S, ref=np.max)\nplt.figure(figsize = (16, 6))\nlibrosa.display.specshow(S_DB, sr=sr, hop_length=hop_length, x_axis = 'time', y_axis = 'log',\n                        cmap = 'cool');\nplt.colorbar();\nplt.title(\"Classical Mel Spectrogram\", fontsize = 23);","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:55:27.172412Z","iopub.execute_input":"2023-07-18T00:55:27.172762Z","iopub.status.idle":"2023-07-18T00:55:28.047835Z","shell.execute_reply.started":"2023-07-18T00:55:27.172737Z","shell.execute_reply":"2023-07-18T00:55:28.046356Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Audio Features\n\nZero Crossing Rate: the rate at which the signal changes from positive to negative or back.","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Total zero_crossings in our 1 song\nzero_crossings = librosa.zero_crossings(audio_file, pad=False)\nprint(sum(zero_crossings))","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:36:08.426822Z","iopub.execute_input":"2023-07-18T00:36:08.427175Z","iopub.status.idle":"2023-07-18T00:36:08.507541Z","shell.execute_reply.started":"2023-07-18T00:36:08.427148Z","shell.execute_reply":"2023-07-18T00:36:08.506353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Harmonics and Perceptrual\n\n\"Harmonics are characteristichs that human years can't distinguish (represents the sound color)\nPerceptrual understanding shock wave represents the sound rhythm and emotion\"","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\ny_harm, y_perc = librosa.effects.hpss(audio_file)\n\nplt.figure(figsize = (16, 6))\nplt.plot(y_harm, color = '#A300F9');\nplt.plot(y_perc, color = '#FFB100');","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:36:58.271719Z","iopub.execute_input":"2023-07-18T00:36:58.272096Z","iopub.status.idle":"2023-07-18T00:37:02.612597Z","shell.execute_reply.started":"2023-07-18T00:36:58.272068Z","shell.execute_reply":"2023-07-18T00:37:02.611607Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Tempo BMP (beats per minute)\n\nDynamic programming beat tracker.","metadata":{}},{"cell_type":"code","source":"tempo, _ = librosa.beat.beat_track(y=y, sr = sr)\ntempo","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:57:06.650713Z","iopub.execute_input":"2023-07-18T00:57:06.651072Z","iopub.status.idle":"2023-07-18T00:57:06.914961Z","shell.execute_reply.started":"2023-07-18T00:57:06.651023Z","shell.execute_reply":"2023-07-18T00:57:06.914054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Spectral Centroid\n\n\"Indicates where the ”centre of mass” for a sound is located and is calculated as the weighted mean of the frequencies present in the sound.\"","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Calculate the Spectral Centroids\nspectral_centroids = librosa.feature.spectral_centroid(y=y, sr=sr)[0]#Changed audio_file to y=y\n\n# Shape is a vector\nprint('Centroids:', spectral_centroids, '\\n')\nprint('Shape of Spectral Centroids:', spectral_centroids.shape, '\\n')\n\n# Computing the time variable for visualization\nframes = range(len(spectral_centroids))\n\n# Converts frame counts to time (seconds)\nt = librosa.frames_to_time(frames)\n\nprint('frames:', frames, '\\n')\nprint('t:', t)\n\n# Function that normalizes the Sound Data\ndef normalize(x, axis=0):\n    return sklearn.preprocessing.minmax_scale(x, axis=axis)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T00:58:34.644297Z","iopub.execute_input":"2023-07-18T00:58:34.644623Z","iopub.status.idle":"2023-07-18T00:58:34.691453Z","shell.execute_reply.started":"2023-07-18T00:58:34.644600Z","shell.execute_reply":"2023-07-18T00:58:34.690393Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n#Plotting the Spectral Centroid along the waveform\nplt.figure(figsize = (16, 6))\nlibrosa.display.waveshow(y=y, sr=sr, alpha=0.4, color = '#A300F9');\nplt.plot(t, normalize(spectral_centroids), color='#FFB100');","metadata":{"execution":{"iopub.status.busy":"2023-07-18T01:00:03.079491Z","iopub.execute_input":"2023-07-18T01:00:03.079850Z","iopub.status.idle":"2023-07-18T01:00:03.780676Z","shell.execute_reply.started":"2023-07-18T01:00:03.079824Z","shell.execute_reply":"2023-07-18T01:00:03.779476Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Spectral Rolloff\n\n\"It's a measure of the shape of the signal. It represents the frequency below which a specified percentage of the total spectral energy, e.g. 85%, lies\"","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Spectral RollOff Vector\nspectral_rolloff = librosa.feature.spectral_rolloff(y=y, sr=sr)[0]\n\n# The plot\nplt.figure(figsize = (16, 6))\nlibrosa.display.waveshow(audio_file, sr=sr, alpha=0.4, color = '#A300F9');\nplt.plot(t, normalize(spectral_rolloff), color='#FFB100');","metadata":{"execution":{"iopub.status.busy":"2023-07-18T01:00:41.551710Z","iopub.execute_input":"2023-07-18T01:00:41.552078Z","iopub.status.idle":"2023-07-18T01:00:42.307277Z","shell.execute_reply.started":"2023-07-18T01:00:41.552024Z","shell.execute_reply":"2023-07-18T01:00:42.306049Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Mel-Frequency Cepstral Coefficients:\n\nThe Mel frequency cepstral coefficients (MFCCs) of a signal are a small set of features (usually about 10–20) which concisely describe the overall shape of a spectral envelope. It models the characteristics of the human voice.","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\nimport sklearn\nmfccs = librosa.feature.mfcc(y=y, sr=sr)\nprint('mfccs shape:', mfccs.shape)\n\n#Displaying  the MFCCs:\nplt.figure(figsize = (16, 6))\nlibrosa.display.specshow(mfccs, sr=sr, x_axis='time', cmap = 'cool');","metadata":{"execution":{"iopub.status.busy":"2023-07-18T01:01:57.882268Z","iopub.execute_input":"2023-07-18T01:01:57.882602Z","iopub.status.idle":"2023-07-18T01:01:58.227833Z","shell.execute_reply.started":"2023-07-18T01:01:57.882575Z","shell.execute_reply":"2023-07-18T01:01:58.226982Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Scaling","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Perform Feature Scaling\nmfccs = sklearn.preprocessing.scale(mfccs, axis=1)\nprint('Mean:', mfccs.mean(), '\\n')\nprint('Var:', mfccs.var())\n\nplt.figure(figsize = (16, 6))\nlibrosa.display.specshow(mfccs, sr=sr, x_axis='time', cmap = 'cool');","metadata":{"execution":{"iopub.status.busy":"2023-07-18T01:03:07.466094Z","iopub.execute_input":"2023-07-18T01:03:07.466447Z","iopub.status.idle":"2023-07-18T01:03:07.743203Z","shell.execute_reply.started":"2023-07-18T01:03:07.466418Z","shell.execute_reply":"2023-07-18T01:03:07.742392Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Chroma Frequencies\n\n\"Chroma features are an interesting and powerful representation for music audio in which the entire spectrum is projected onto 12 bins representing the 12 distinct semitones (or chroma) of the musical octave.\"","metadata":{}},{"cell_type":"code","source":"#Andrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\n# Increase or decrease hop_length to change how granular you want your data to be\nhop_length = 5000\n\n# Chromogram\nchromagram = librosa.feature.chroma_stft(y=y, sr=sr, hop_length=hop_length)#y=y fix error\nprint('Chromogram shape:', chromagram.shape)\n\nplt.figure(figsize=(16, 6))\nlibrosa.display.specshow(chromagram, x_axis='time', y_axis='chroma', hop_length=hop_length, cmap='coolwarm');","metadata":{"execution":{"iopub.status.busy":"2023-07-18T01:04:06.129575Z","iopub.execute_input":"2023-07-18T01:04:06.129934Z","iopub.status.idle":"2023-07-18T01:04:06.578249Z","shell.execute_reply.started":"2023-07-18T01:04:06.129906Z","shell.execute_reply":"2023-07-18T01:04:06.577328Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#I didn't learn Bangla but I learned that y=y\n\nTo avoid many Errors \"takes 0 positional arguments but 1 were given\". I had to add y=y (original was just y). Example: (y=y, sr=sr) ","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements:\n\nAndrada Olteanu https://www.kaggle.com/code/andradaolteanu/work-w-audio-data-visualise-classify-recommend\n\nTanjima Nasreen Jenia  https://www.kaggle.com/tanjimanasreenjenia/bangladeshi-restaurants-analytics/notebook\n\nEu Jin Lok https://www.kaggle.com/ejlok1/audio-emotion-part-5-data-augmentation/notebook\n\nLucas Abrahão https://www.kaggle.com/lucasabrahao/trabalho-manufatura-an-lise-de-dados-no-brasil","metadata":{}}]}