{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"> Since the main challenge of this competition of this competition is ```Recognize Bengali speech from out-of-distribution audio recordings```,we need to evaluate how our model is **performing on Out-of-distribuition data**. But we don't have OOD data, except the audios provides in ```/kaggle/input/bengaliai-speech/test_mp3s``` directory. \n\n> They have provided **17 audio samples for 17 domains**. \n\n> **I have hand-annotated these 17 OOD audios**. You can use these annotations to check how your model is performing. ","metadata":{}},{"cell_type":"markdown","source":"> Disclaimer : **These annotations might not be the exact match with the actual annotations**. Some of the audios were too difficult to understand, some of the annotations might have my biased judgement on some words and some words might have different annotation protocol in the actual test set. So any numerical evaluation using this annotaions might not be accurate. However it might give you an idea on how your model is actually performing.","metadata":{}},{"cell_type":"markdown","source":"# Insights","metadata":{}},{"cell_type":"markdown","source":"> **TLDR** : I have some insights on the OOD Test data. Let me share those first. Then we'll explore some of the audios.\n\n> **The main challenges in the test audios/OOD audios :**\n```\n1. Most of the audios are dialogues where two or more persons are speaker.Some of them contain overlapping sounds. It'll be challenging to predict these words.\n2. There are audios of different regions/accents. This'll be challenging for the model to understand different regional accents\n3. Some Audios contain background music, challenging to focus to the text part.\n4. Some domain contain same words with different accents.\n5. The speed of the speaker is also a vital issue. Some audios are hard to keep track of\n6. Presence of Name, Mix of english words/numbers. These will be hard to detect\n7. Different annotation possibility: Since this a crowd-sourced data, different users might have had different annotations for the same words/sentences\n```","metadata":{}},{"cell_type":"markdown","source":"We'll look into the audios now and from hearing the audios, you'll understand what challenges I'm talking about. ","metadata":{}},{"cell_type":"markdown","source":"# Explore the audios and annotations","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport soundfile as sf\nfrom pydub import AudioSegment","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:00.808795Z","iopub.execute_input":"2023-07-21T03:22:00.810374Z","iopub.status.idle":"2023-07-21T03:22:00.890005Z","shell.execute_reply.started":"2023-07-21T03:22:00.810284Z","shell.execute_reply":"2023-07-21T03:22:00.888681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv(\"/kaggle/input/ood-example-audios-hand-annotations/annoated.csv\",sep=\"\\t\")\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:00.892095Z","iopub.execute_input":"2023-07-21T03:22:00.892724Z","iopub.status.idle":"2023-07-21T03:22:00.935829Z","shell.execute_reply.started":"2023-07-21T03:22:00.892686Z","shell.execute_reply":"2023-07-21T03:22:00.933315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# There are audios of different regions/accents","metadata":{}},{"cell_type":"code","source":"path = \"/kaggle/input/bengaliai-speech/examples/\"\nfile = \"Cartoon.wav\"\n\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:00.969265Z","iopub.execute_input":"2023-07-21T03:22:00.970622Z","iopub.status.idle":"2023-07-21T03:22:02.978602Z","shell.execute_reply.started":"2023-07-21T03:22:00.970578Z","shell.execute_reply":"2023-07-21T03:22:02.977490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is not standard bangla, rather having a regional accent.","metadata":{}},{"cell_type":"markdown","source":"# Some Audios contain background music","metadata":{}},{"cell_type":"code","source":"path = \"/kaggle/input/bengaliai-speech/examples/\"\nfile = \"Indian TV Drama.wav\"\n\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:02.980504Z","iopub.execute_input":"2023-07-21T03:22:02.980982Z","iopub.status.idle":"2023-07-21T03:22:04.204497Z","shell.execute_reply.started":"2023-07-21T03:22:02.980942Z","shell.execute_reply":"2023-07-21T03:22:04.203262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In some parts, it's difficult to seperate the words from the background musing","metadata":{}},{"cell_type":"code","source":"file = \"Movie.wav\"\n\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:04.206332Z","iopub.execute_input":"2023-07-21T03:22:04.208280Z","iopub.status.idle":"2023-07-21T03:22:05.469464Z","shell.execute_reply.started":"2023-07-21T03:22:04.208240Z","shell.execute_reply":"2023-07-21T03:22:05.468370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Some of them contain overlapping sounds","metadata":{}},{"cell_type":"markdown","source":"An example of this is the Bengali Advertisement.wav file. Let's hear it","metadata":{}},{"cell_type":"code","source":"file = \"Bengali Advertisement.wav\"\n\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:05.472787Z","iopub.execute_input":"2023-07-21T03:22:05.473612Z","iopub.status.idle":"2023-07-21T03:22:07.216374Z","shell.execute_reply.started":"2023-07-21T03:22:05.473569Z","shell.execute_reply":"2023-07-21T03:22:07.215086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are multiple speakers in this audio.","metadata":{}},{"cell_type":"markdown","source":"# Some domain contain same words/sentence with different accents.","metadata":{}},{"cell_type":"code","source":"file = \"Puthi Literature.wav\"\n\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:07.217977Z","iopub.execute_input":"2023-07-21T03:22:07.219040Z","iopub.status.idle":"2023-07-21T03:22:08.302847Z","shell.execute_reply.started":"2023-07-21T03:22:07.218995Z","shell.execute_reply":"2023-07-21T03:22:08.301437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The speed of the speaker is also a vital issue","metadata":{}},{"cell_type":"code","source":"file = \"Waz Islamic Sermon.wav\"\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:08.304671Z","iopub.execute_input":"2023-07-21T03:22:08.305073Z","iopub.status.idle":"2023-07-21T03:22:09.352014Z","shell.execute_reply.started":"2023-07-21T03:22:08.305040Z","shell.execute_reply":"2023-07-21T03:22:09.350755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"file = \"Debate.wav\"\n\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:09.353669Z","iopub.execute_input":"2023-07-21T03:22:09.354556Z","iopub.status.idle":"2023-07-21T03:22:10.252471Z","shell.execute_reply.started":"2023-07-21T03:22:09.354521Z","shell.execute_reply":"2023-07-21T03:22:10.251092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Presence of Name, Mix of english words/numbers","metadata":{}},{"cell_type":"code","source":"file = \"Audiobook.wav\"\n\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:10.254833Z","iopub.execute_input":"2023-07-21T03:22:10.255307Z","iopub.status.idle":"2023-07-21T03:22:11.419953Z","shell.execute_reply.started":"2023-07-21T03:22:10.255270Z","shell.execute_reply":"2023-07-21T03:22:11.418671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> এক হাজার,কয়েকহাজার > Number\n\n> ডিনারে,ডলার,কয়েন -> English words","metadata":{}},{"cell_type":"code","source":"file = \"Debate.wav\"\n\nprint(file)\ndisplay(AudioSegment.from_file(path+file))\ndf[df['file']==file].sentence.tolist()[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-21T03:22:11.422617Z","iopub.execute_input":"2023-07-21T03:22:11.423406Z","iopub.status.idle":"2023-07-21T03:22:12.243240Z","shell.execute_reply.started":"2023-07-21T03:22:11.423362Z","shell.execute_reply":"2023-07-21T03:22:12.241947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lot of names in this audio.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}