{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <center style=\"font-family: consolas; font-size: 28px; font-weight: bold;\"> Bengali.AI Speech Recognition : Detailed EDA, Normalizer and WER</center>\n<p><center style=\"color:#949494; font-family: consolas; font-size: 20px;\"> Let's contribute to the Bangla Language Together🤗 </center></p>","metadata":{}},{"cell_type":"markdown","source":"Version History : \n\n> * **Version 2** : Includes audio length distribuition\n> * **Version 3** : Includes Intent distribuition of the text sentences","metadata":{}},{"cell_type":"code","source":"from pydub import AudioSegment\npath = \"/kaggle/input/bengaliai-speech/train_mp3s/83efbd035221.mp3\"\ndisplay(AudioSegment.from_file(path))\npath = \"/kaggle/input/bengaliai-speech/train_mp3s/910ec4c6e1b9.mp3\"\ndisplay(AudioSegment.from_file(path))","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:31.361121Z","iopub.execute_input":"2023-08-26T16:22:31.361525Z","iopub.status.idle":"2023-08-26T16:22:32.762248Z","shell.execute_reply.started":"2023-08-26T16:22:31.361490Z","shell.execute_reply":"2023-08-26T16:22:32.760243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><div class=\"alert alert-block alert-info\" style=\"margin: 2em; line-height: 1.7em; font-family: Verdana;\">\n    <b style=\"font-size: 18px; color:green\"> &nbsp; DO YOU KNOW WHAT THE ABOVE AUDIOS ARE SAYING?</b><br><br><b style=\"font-size: 18px; color:green\">\"স্বাগতম, কেমন আছেেন\" -> Welcome, how are you?</b><br>\n</div></center>","metadata":{}},{"cell_type":"markdown","source":"<center> <b style=\"font-size: 18px; color:purple\">Hello everyone!</b></center>\n<br>\nIf you didn't understand what the the speakers in the previous audios were saying, well you're in the right place. This is what this competition is all about. To make machines understand Bangla Audios, and generate texts from them !\n\nLet's get started.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\"> ✍️Table of Contents✍️</center></div>\n\n1.  [Imports](#1)\n1.  [What is this competition about?](#2)\n1.  [Peeking into the dataset](#3)\n1.  [EDA on text sentences ](#4)\n1. [Text Normalization](#5)\n1. [Evaluation Metric](#6)\n1. [Audio EDA](#7)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\"> ⬇️Imports⬇️</center>\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport soundfile as sf\nfrom pydub import AudioSegment\nimport IPython.display as ipd\nfrom collections import Counter\nimport os","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:32.764339Z","iopub.execute_input":"2023-08-26T16:22:32.764750Z","iopub.status.idle":"2023-08-26T16:22:32.788209Z","shell.execute_reply.started":"2023-08-26T16:22:32.764716Z","shell.execute_reply":"2023-08-26T16:22:32.786978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\"> What is this competition about?😕</center>\n","metadata":{}},{"cell_type":"markdown","source":"\n\nWell, to answer in short this is an Automatic Speech Recognition(ASR) task. We have to develop models for Bengali ASR that can generate sentence level transcription from a given audio. This competition proivdes the largest dataset on Bengali ASR domain, which also introduces data samples from several domains and challenge you to build models that can perform well on Out of Domain Data too.\n\nThis is what the organizers have to say  : \n```\nThe goal of this competition is to recognize Bengali speech from out-of-distribution audio recordings. You will build a model trained on the first Massively Crowdsourced (MaCro) Bengali speech dataset with 1,200 hours of data from ~24,000 people from India and Bangladesh. The test set contains samples from 17 different domains that are not present in training.\n\nYour efforts could improve Bengali speech recognition using the first Bengali out-of-distribution speech recognition dataset. In addition, your submission will be among the first open-source speech recognition methods for Bengali\n```","metadata":{}},{"cell_type":"markdown","source":"Let's hope this competition will bring ground breaking results on the Bengali ASR Domain 🤗","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\">  Peeking into the dataset 👀</center>\n","metadata":{}},{"cell_type":"markdown","source":"So what do we have here? Well,","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/input/bengaliai-speech","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:32.789812Z","iopub.execute_input":"2023-08-26T16:22:32.790214Z","iopub.status.idle":"2023-08-26T16:22:33.913701Z","shell.execute_reply.started":"2023-08-26T16:22:32.790177Z","shell.execute_reply":"2023-08-26T16:22:33.912316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*  **train/** The training set, comprising several thousand recordings in MP3 format.\n* **test/** The test set, comprising spontaneous speech recordings from eighteen domains, seventeen of which are out-of-distribution with respect to the training set. There may be domains in the private test set that are not in the public test set.\n* **examples/** An example recording for each test set domain. You may find these example recordings helpful for creating models robust to domain variation. These are representative recordings and none of them are present in the test set.\n* **train.csv** Sentence labels for the training set.\n\n    \n* **sample_submission.csv** A sample submission file in the correct format. See the Evaluation page for more details.","metadata":{}},{"cell_type":"markdown","source":"Let's load the train.csv","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv(\"/kaggle/input/bengaliai-speech/train.csv\")\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:33.919954Z","iopub.execute_input":"2023-08-26T16:22:33.921229Z","iopub.status.idle":"2023-08-26T16:22:41.341091Z","shell.execute_reply.started":"2023-08-26T16:22:33.921151Z","shell.execute_reply":"2023-08-26T16:22:41.339164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **id** A unique identifier for this instance. Corresponds to the file {id}.mp3 in    train/.\n* **sentence** A plain-text transcription of the recording. Your goal is to predict these sentences for each recording in the test set.\n* **split** Whether train or valid. The annotations in the valid split have been manually reviewed and corrected, while the annotations in the train split have only been algorithmically cleaned. The valid samples will generally have higher quality annotations than the train samples, but are otherwise drawn from the same distribution.","metadata":{}},{"cell_type":"code","source":"df.split.unique()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:41.343376Z","iopub.execute_input":"2023-08-26T16:22:41.343968Z","iopub.status.idle":"2023-08-26T16:22:41.446789Z","shell.execute_reply.started":"2023-08-26T16:22:41.343915Z","shell.execute_reply":"2023-08-26T16:22:41.445312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How many train and how many validation samples?","metadata":{}},{"cell_type":"code","source":"n_train_samples = sum(df[\"split\"]==\"train\")\nn_valid_samples = sum(df[\"split\"]==\"valid\")\nprint(f\"Total training samples : \",n_train_samples)\nprint(f\"Total validation samples : \",n_valid_samples)\nprint(\"Validation/Train ratio : \",n_valid_samples/n_train_samples)","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:41.448900Z","iopub.execute_input":"2023-08-26T16:22:41.449387Z","iopub.status.idle":"2023-08-26T16:22:42.066395Z","shell.execute_reply.started":"2023-08-26T16:22:41.449346Z","shell.execute_reply":"2023-08-26T16:22:42.065098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Okay so validation set is very small compared to training set. We might need to use additional data from the train set as validation set.","metadata":{}},{"cell_type":"code","source":"plt.bar([\"train\",\"valid\"],[n_train_samples,n_valid_samples],color = ['blue', 'yellow'])","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:42.067865Z","iopub.execute_input":"2023-08-26T16:22:42.068304Z","iopub.status.idle":"2023-08-26T16:22:42.380300Z","shell.execute_reply.started":"2023-08-26T16:22:42.068260Z","shell.execute_reply":"2023-08-26T16:22:42.379344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's hear some audio files and see the corresponding text","metadata":{}},{"cell_type":"code","source":"root_path = \"/kaggle/input/bengaliai-speech/train_mp3s\"\n\nfor idx in range(1,len(df),99999):\n    \n    mp3_path = os.path.join(root_path,df['id'].iloc[idx])+ \".mp3\"\n    text = df['sentence'].iloc[idx]\n    display(AudioSegment.from_file(mp3_path))\n    print(\"Original transcription : \",text)\n","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:42.381896Z","iopub.execute_input":"2023-08-26T16:22:42.382463Z","iopub.status.idle":"2023-08-26T16:22:45.779845Z","shell.execute_reply.started":"2023-08-26T16:22:42.382424Z","shell.execute_reply":"2023-08-26T16:22:45.778414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From these audios it's quite evident that the audios has a lot of varieties, in domains, in audio quality and other forms. Feel free to change the idxs and listen to different audios","metadata":{}},{"cell_type":"markdown","source":"For the test set, we'll have data from 17 domains. They have provided a sample for each of them. Let's hear some of them.","metadata":{}},{"cell_type":"code","source":"mp3_path = \"/kaggle/input/bengaliai-speech/examples/Cartoon.wav\"\nprint(\"Cartoon Audio\")\ndisplay(AudioSegment.from_file(mp3_path))\n\nmp3_path = \"/kaggle/input/bengaliai-speech/examples/Poem Recital.wav\"\nprint(\"Poem recital \")\ndisplay(AudioSegment.from_file(mp3_path))\n","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:45.781628Z","iopub.execute_input":"2023-08-26T16:22:45.782220Z","iopub.status.idle":"2023-08-26T16:22:48.110875Z","shell.execute_reply.started":"2023-08-26T16:22:45.782169Z","shell.execute_reply":"2023-08-26T16:22:48.109528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This Cartoon is probably from the popular Cartoon show \"Gopal Bhar\" and the poem was written by the National Poet of Bangladesh, \"Kazi Nazrul Islam\"","metadata":{}},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\"> 🧐EDA on text sentences🧐</center>\n","metadata":{}},{"cell_type":"code","source":"print(df.shape)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:48.117781Z","iopub.execute_input":"2023-08-26T16:22:48.118908Z","iopub.status.idle":"2023-08-26T16:22:48.135984Z","shell.execute_reply.started":"2023-08-26T16:22:48.118846Z","shell.execute_reply":"2023-08-26T16:22:48.134572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A total of 963636 sentences. But how many unique sentences?","metadata":{}},{"cell_type":"code","source":"print(\"Total sentences :\",len(df))\nprint(\"Total unique sentences : \",df.sentence.nunique())\nprint(\"Percentage of unique sentences ; \",df.sentence.nunique()/len(df))","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:48.137670Z","iopub.execute_input":"2023-08-26T16:22:48.138678Z","iopub.status.idle":"2023-08-26T16:22:50.358699Z","shell.execute_reply.started":"2023-08-26T16:22:48.138634Z","shell.execute_reply":"2023-08-26T16:22:50.356969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's interesting. 50% of the overall sentences are actually unique. So multiple of them has been repeated in the dataset. Let's look at the most frequent 10 sentences","metadata":{}},{"cell_type":"code","source":"print(\"Most frequent sentences in the dataset \\n\")\ndf.sentence.value_counts()[:10]","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:50.360390Z","iopub.execute_input":"2023-08-26T16:22:50.360827Z","iopub.status.idle":"2023-08-26T16:22:51.634958Z","shell.execute_reply.started":"2023-08-26T16:22:50.360788Z","shell.execute_reply":"2023-08-26T16:22:51.633502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Let's look at the sentence length distribuition**","metadata":{}},{"cell_type":"code","source":"x = df.sentence.apply(lambda x: len(x))","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:51.636805Z","iopub.execute_input":"2023-08-26T16:22:51.637237Z","iopub.status.idle":"2023-08-26T16:22:52.511643Z","shell.execute_reply.started":"2023-08-26T16:22:51.637202Z","shell.execute_reply":"2023-08-26T16:22:52.510377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.xlabel(\"Sentence Length\")\nplt.ylabel(\"Frequency\")\nplt.title(\"Sentence Length Distribuition\")\nplt.hist(x)","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:52.513626Z","iopub.execute_input":"2023-08-26T16:22:52.514174Z","iopub.status.idle":"2023-08-26T16:22:52.926258Z","shell.execute_reply.started":"2023-08-26T16:22:52.514122Z","shell.execute_reply":"2023-08-26T16:22:52.924816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the sentences have length<=100\n\nLet's look at the overall vocabulary size and the most frequent words\n","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\nvocab = {}\nfor sen in tqdm(df.sentence):\n    for j in sen.split(\" \"):\n        try:\n            vocab[j]+=1\n        except:\n            vocab[j]=1\nprint(\"Total words in vocabulary : \",len(vocab))","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:52.928330Z","iopub.execute_input":"2023-08-26T16:22:52.928993Z","iopub.status.idle":"2023-08-26T16:22:59.899837Z","shell.execute_reply.started":"2023-08-26T16:22:52.928945Z","shell.execute_reply":"2023-08-26T16:22:59.898464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So total 235761 words in vocabulary. That's huge. Let's look at the most frequent ones","metadata":{}},{"cell_type":"code","source":"sorted_vocab = sorted(vocab.items(),key = lambda kv:kv[1],reverse=True)\nsorted_vocab[:30]","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:22:59.901693Z","iopub.execute_input":"2023-08-26T16:22:59.902220Z","iopub.status.idle":"2023-08-26T16:23:00.018997Z","shell.execute_reply.started":"2023-08-26T16:22:59.902168Z","shell.execute_reply":"2023-08-26T16:23:00.017270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**So mostly bengali pronouns and common verbs.**","metadata":{}},{"cell_type":"markdown","source":"This is to remind you that this is the vocabulary of **train+validation**. You might wanna check how much **out of vocabulary** words does the validation set have. Let's check it.","metadata":{}},{"cell_type":"code","source":"def vocabulary(df):\n    vocab = {}\n    for sen in tqdm(df.sentence):\n        for j in sen.split(\" \"):\n            try:\n                vocab[j]+=1\n            except:\n                vocab[j]=1\n    return vocab","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:00.020867Z","iopub.execute_input":"2023-08-26T16:23:00.021461Z","iopub.status.idle":"2023-08-26T16:23:00.030717Z","shell.execute_reply.started":"2023-08-26T16:23:00.021402Z","shell.execute_reply":"2023-08-26T16:23:00.028991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = df[df[\"split\"]==\"train\"]\nval = df[df[\"split\"]==\"valid\"]\n\ntrain_vocab = vocabulary(train)\nval_vocab = vocabulary(val)\nprint(\"Total words in train vocabulary : \",len(train_vocab))\nprint(\"Total words in validation vocabulary : \",len(val_vocab))\n\ntrain_words = set([key for key,value in train_vocab.items()])\nvocab_words = set([key for key,value in val_vocab.items()])\n\nprint(\"Total Out of vocabulary words : \",len(vocab_words-train_words))\n","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:00.032572Z","iopub.execute_input":"2023-08-26T16:23:00.033006Z","iopub.status.idle":"2023-08-26T16:23:06.561363Z","shell.execute_reply.started":"2023-08-26T16:23:00.032970Z","shell.execute_reply":"2023-08-26T16:23:06.560021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Okay not so much words.\n\nNow let's look at the characters present.","metadata":{}},{"cell_type":"code","source":"chars = {}\nfor sen in tqdm(df.sentence):\n    for j in sen:\n        try:\n            chars[j]+=1\n        except:\n            chars[j]=1\nchars","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:06.562828Z","iopub.execute_input":"2023-08-26T16:23:06.563249Z","iopub.status.idle":"2023-08-26T16:23:26.051824Z","shell.execute_reply.started":"2023-08-26T16:23:06.563212Z","shell.execute_reply":"2023-08-26T16:23:26.050449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Total characters :\",len(chars))","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:26.053610Z","iopub.execute_input":"2023-08-26T16:23:26.054426Z","iopub.status.idle":"2023-08-26T16:23:26.060038Z","shell.execute_reply.started":"2023-08-26T16:23:26.054379Z","shell.execute_reply":"2023-08-26T16:23:26.059018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sorted_chars = sorted(chars.items(),key = lambda x:x[1],reverse = True)\nprint(sorted_chars[:10])\nsorted_chars = dict(sorted_chars)\nplt.figure(figsize=(15,8))\nplt.title('Character Occurance')\nplt.bar(sorted_chars.keys(),sorted_chars.values())\nplt.xlabel('Unique characters')\nplt.ylabel('Sample Count')\nplt.grid(axis='y')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:26.061633Z","iopub.execute_input":"2023-08-26T16:23:26.062071Z","iopub.status.idle":"2023-08-26T16:23:27.370481Z","shell.execute_reply.started":"2023-08-26T16:23:26.062017Z","shell.execute_reply":"2023-08-26T16:23:27.369136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\"> 🔍Text Normalization 🔍</center>\n","metadata":{}},{"cell_type":"markdown","source":"At this point you might ask what is text normalization and why do we need it here? Well, this is one of the challenges in bangla. It has so manyyy characters whereas English has only 26 characters. As a result a word could have different forms.\n\nFor example, look at this pair :\n\n**নিয়ে == নিয়ে**\n\n**হয়ে == হয়ে**\n\n**কথায় == কথায়**\n\n**চিড়ে == চিড়ে**\n\n\nDon't they look absolutely same? Well let's see.","metadata":{}},{"cell_type":"code","source":"print(\"নিয়ে\" == \"নিয়ে\")\nprint(\"হয়ে\" == \"হয়ে\")\nprint(\"চিড়ে\"==\"চিড়ে\")\n","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:27.372348Z","iopub.execute_input":"2023-08-26T16:23:27.372782Z","iopub.status.idle":"2023-08-26T16:23:27.380397Z","shell.execute_reply.started":"2023-08-26T16:23:27.372744Z","shell.execute_reply":"2023-08-26T16:23:27.378805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Wait what? Why the computer is saying they aren't equal? They look absolutely identitcal.**","metadata":{}},{"cell_type":"code","source":"print(list(\"নিয়ে\"))\nprint(list(\"নিয়ে\"))","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:27.382265Z","iopub.execute_input":"2023-08-26T16:23:27.382729Z","iopub.status.idle":"2023-08-26T16:23:27.395632Z","shell.execute_reply.started":"2023-08-26T16:23:27.382686Z","shell.execute_reply":"2023-08-26T16:23:27.394129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(list(\"চিড়ে\"))\nprint(list(\"চিড়ে\"))","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:27.398475Z","iopub.execute_input":"2023-08-26T16:23:27.399563Z","iopub.status.idle":"2023-08-26T16:23:27.410988Z","shell.execute_reply.started":"2023-08-26T16:23:27.399500Z","shell.execute_reply":"2023-08-26T16:23:27.409594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Do you get it now? Although they look identical, they both have unequal amount of characters.\n\n**\"নিয়ে\"** has an extra character **\"'়'\"**\n\nWell here's the thing. The character \"য়\" you are seeing can be written in two ways. \n\n**\"য়\" -> 'য়' (First word)**\n\n**\"য়\" ->  'য'+ '়'**\n\n**This '়' symbol is called Nukta, the characters having dots under them (\"য়\",'ড়' etc.) can have it , it some notation they might not. So if the dataset contain both forms, the metric will count them as unequal words and thus add an error.**","metadata":{}},{"cell_type":"markdown","source":"Also think, you're confused seeing these words. The model will also get confused seeing two representations of the same word.\n\nHow to avoid this problem? Well normalizer to the rescue. We want to bring all the words in a canonical form. Normalizer will do that for us.\n\nWe'll be using [this normalizer](https://github.com/csebuetnlp/normalizer) for this Notebook. There are several others for bangla.\n\nLet's first install it","metadata":{}},{"cell_type":"code","source":"!pip install git+https://github.com/csebuetnlp/normalizer","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:27.412603Z","iopub.execute_input":"2023-08-26T16:23:27.413010Z","iopub.status.idle":"2023-08-26T16:23:54.033474Z","shell.execute_reply.started":"2023-08-26T16:23:27.412973Z","shell.execute_reply":"2023-08-26T16:23:54.032093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How to use? We can use it on the whole sentence","metadata":{}},{"cell_type":"code","source":"from normalizer import normalize\nsentence = \"এমনকি উকুন ঘরবাড়ি ও খাদ্য-সম্ভারের উপর ছড়িয়ে পড়তে লাগল।\"\nnormalized = normalize(sentence)\nprint(normalized)\nprint(normalize==sentence)","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:54.035396Z","iopub.execute_input":"2023-08-26T16:23:54.035959Z","iopub.status.idle":"2023-08-26T16:23:54.286593Z","shell.execute_reply.started":"2023-08-26T16:23:54.035895Z","shell.execute_reply":"2023-08-26T16:23:54.284890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Or we can use it on word level","metadata":{}},{"cell_type":"code","source":"normalize(\"নিয়ে\")==normalize(\"নিয়ে\")","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:23:54.288689Z","iopub.execute_input":"2023-08-26T16:23:54.289109Z","iopub.status.idle":"2023-08-26T16:23:54.298507Z","shell.execute_reply.started":"2023-08-26T16:23:54.289073Z","shell.execute_reply":"2023-08-26T16:23:54.297105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"See! The problem is now gone.","metadata":{}},{"cell_type":"markdown","source":"\n<a id=\"8\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\"> Intent of texts</center>\n","metadata":{}},{"cell_type":"markdown","source":"We can find out the intents of the text sentences. This might help out to find the most suitable punctuations for each sentence. Let's first see which type of punctuations are present in the sentences","metadata":{}},{"cell_type":"code","source":"base_puncs = ['!','?','।']\n\npuncs = {}\n\nfor i in tqdm(train.sentence.tolist()):\n    last = i[-1]\n    \n    if last not in base_puncs:\n        last = \"others\"\n        \n    try:\n        puncs[last]+=1\n    except:\n        puncs[last] = 1\n","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:37:22.764420Z","iopub.execute_input":"2023-08-26T16:37:22.764935Z","iopub.status.idle":"2023-08-26T16:37:24.022374Z","shell.execute_reply.started":"2023-08-26T16:37:22.764897Z","shell.execute_reply":"2023-08-26T16:37:24.020968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"puncs","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:37:25.644896Z","iopub.execute_input":"2023-08-26T16:37:25.645389Z","iopub.status.idle":"2023-08-26T16:37:25.654750Z","shell.execute_reply.started":"2023-08-26T16:37:25.645353Z","shell.execute_reply":"2023-08-26T16:37:25.653124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This can be categorized in 4 types:\n> **'।'** -> **statements**\n\n> **'!'** --> Exclamation\n\n> **?** --> Question\n\n> **Others** -> Neutral/Wrong annotation/Incompleted sentence","metadata":{}},{"cell_type":"markdown","source":"**Let's look at the validation set**","metadata":{}},{"cell_type":"code","source":"base_puncs = ['!','?','।']\n\npuncs = {}\n\nfor i in tqdm(val.sentence.tolist()):\n    last = i[-1]\n    \n    if last not in base_puncs:\n        last = \"others\"\n        \n    try:\n        puncs[last]+=1\n    except:\n        puncs[last] = 1\npuncs\n","metadata":{"execution":{"iopub.status.busy":"2023-08-26T16:35:38.146684Z","iopub.execute_input":"2023-08-26T16:35:38.147230Z","iopub.status.idle":"2023-08-26T16:35:38.222636Z","shell.execute_reply.started":"2023-08-26T16:35:38.147191Z","shell.execute_reply":"2023-08-26T16:35:38.221169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\"> 🖋️ Evaluation Metric : WER 🖋️</center>","metadata":{}},{"cell_type":"markdown","source":"WER stands for Word Error Rate. Word error rate is a common metric of the performance of a speech recognition or machine translation system. (from wiki) \n\n ~~~\n It measures the percentage of words that are incorrectly recognized or transcribed compared to a reference or ground truth.\n\nThe formula for calculating WER is as follows:\n\nWER = (S + D + I) / N\n\nwhere:\nS = the number of word substitutions (words in the reference that are replaced with incorrect words in the recognition)\nD = the number of word deletions (words in the reference that are not present in the recognition)\nI = the number of word insertions (additional words inserted in the recognition that are not present in the reference)\nN = the total number of words in the reference\n\nTo calculate WER, the ASR or OCR system's output is compared to a known and correct reference text. Any differences in words, including substitutions, deletions, and insertions, are counted and divided by the total number of words in the reference.\n\nWER is typically expressed as a percentage, where lower values indicate better performance. For example, a WER of 10% means that 10% of the words in the recognized output are incorrect or differ from the reference. (From my good friend ChatGPT)\n~~~","metadata":{}},{"cell_type":"markdown","source":"Now how to calculate? We'll use jiwer to calculate WER","metadata":{}},{"cell_type":"code","source":"df['sentence'].tolist()[:20]","metadata":{"execution":{"iopub.status.busy":"2023-07-18T17:13:07.291383Z","iopub.execute_input":"2023-07-18T17:13:07.291809Z","iopub.status.idle":"2023-07-18T17:13:07.348123Z","shell.execute_reply.started":"2023-07-18T17:13:07.291771Z","shell.execute_reply":"2023-07-18T17:13:07.347040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install jiwer","metadata":{"execution":{"iopub.status.busy":"2023-07-18T17:10:38.825466Z","iopub.execute_input":"2023-07-18T17:10:38.826412Z","iopub.status.idle":"2023-07-18T17:10:52.285629Z","shell.execute_reply.started":"2023-07-18T17:10:38.826365Z","shell.execute_reply":"2023-07-18T17:10:52.283970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from jiwer import wer\n\nground_truth =  'আজ পর্যন্ত কেউই কার্তিকেয়ানকে সন্দেহ কিংবা প্রশ্নবিদ্ধ করেনি।'\nprediction = normalize(ground_truth)\nprint(\"Ground Truth : \",ground_truth)\nprint(\"prediction : \",prediction)\n\nprint(\"Word Error Rate : \",wer(ground_truth,prediction))\n","metadata":{"execution":{"iopub.status.busy":"2023-07-18T17:14:20.473081Z","iopub.execute_input":"2023-07-18T17:14:20.473456Z","iopub.status.idle":"2023-07-18T17:14:20.479857Z","shell.execute_reply.started":"2023-07-18T17:14:20.473428Z","shell.execute_reply":"2023-07-18T17:14:20.479019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ground_truth =  'ক্যাপ্টেন, আমরা ব্লাস্টারটাকে ঠিক ধরে ফেলবো।'\nprediction = 'ক্যাপ্টেন, আমরা ভ্লাচ স্টার টাকে ঠিক ধরিতর্ভব।'\nprint(\"Ground Truth : \",ground_truth)\nprint(\"prediction : \",prediction)\n\nprint(\"Word Error Rate : \",wer(ground_truth,prediction))","metadata":{"execution":{"iopub.status.busy":"2023-07-18T17:16:11.716505Z","iopub.execute_input":"2023-07-18T17:16:11.716937Z","iopub.status.idle":"2023-07-18T17:16:11.724687Z","shell.execute_reply.started":"2023-07-18T17:16:11.716896Z","shell.execute_reply":"2023-07-18T17:16:11.723373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:purple\"> ⬇️Audio EDA⬇️</center>\n","metadata":{}},{"cell_type":"markdown","source":"Let's first look at the audio durations and sampling rates. We'll use mutagen since it's faster.","metadata":{}},{"cell_type":"code","source":"!pip install mutagen","metadata":{"execution":{"iopub.status.busy":"2023-07-25T21:19:48.941039Z","iopub.execute_input":"2023-07-25T21:19:48.941844Z","iopub.status.idle":"2023-07-25T21:20:16.648684Z","shell.execute_reply.started":"2023-07-25T21:19:48.941774Z","shell.execute_reply":"2023-07-25T21:20:16.647081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport glob\nimport concurrent.futures\nfrom mutagen.mp3 import MP3\n\ndef process_mp3_file(mp3_file):\n    try:\n        audio = MP3(mp3_file)\n        duration = audio.info.length\n        sample_rate = audio.info.sample_rate\n        return mp3_file, duration, sample_rate\n    except Exception as e:\n        print(f\"Error processing {mp3_file}: {e}\")\n        return mp3_file, None, None\n\ndef process_mp3_files_in_batch(mp3_files):\n    durs = []\n    srs = []\n    with concurrent.futures.ThreadPoolExecutor() as executor:\n        futures = [executor.submit(process_mp3_file, mp3_file) for mp3_file in mp3_files]\n        for future in concurrent.futures.as_completed(futures):\n            result = future.result()\n            durs.append(result[1])\n            srs.append(result[2])\n    return durs,srs\n\ndef durations(mp3_files_directory):\n    # Get a list of all MP3 files in the directory\n    mp3_files = glob.glob(os.path.join(mp3_files_directory, \"*.mp3\"))\n\n    # Split the list of files into batches for processing\n    batch_size = 1000\n    file_batches = [mp3_files[i:i + batch_size] for i in tqdm(range(0, len(mp3_files), batch_size))]\n\n    durs = []\n    srs = []\n    for batch in tqdm(file_batches):\n        dur,sr = process_mp3_files_in_batch(batch)\n        durs.extend(dur)\n        srs.extend(sr)\n    return durs,srs","metadata":{"execution":{"iopub.status.busy":"2023-07-25T21:30:48.318598Z","iopub.execute_input":"2023-07-25T21:30:48.319761Z","iopub.status.idle":"2023-07-25T21:30:48.333868Z","shell.execute_reply.started":"2023-07-25T21:30:48.319702Z","shell.execute_reply":"2023-07-25T21:30:48.332478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_path = \"/kaggle/input/bengaliai-speech/train_mp3s\"\ndurs,srs = durations(train_path)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T21:30:49.685229Z","iopub.execute_input":"2023-07-25T21:30:49.685767Z","iopub.status.idle":"2023-07-25T21:31:13.511068Z","shell.execute_reply.started":"2023-07-25T21:30:49.685703Z","shell.execute_reply":"2023-07-25T21:31:13.509409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.xlabel(\"Audio Length in seconds\")\nplt.ylabel(\"Frequency\")\nplt.title(\"Audio Length Distribuition\")\nplt.hist(durs)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T21:29:07.069454Z","iopub.execute_input":"2023-07-25T21:29:07.069971Z","iopub.status.idle":"2023-07-25T21:29:07.485779Z","shell.execute_reply.started":"2023-07-25T21:29:07.069916Z","shell.execute_reply":"2023-07-25T21:29:07.484597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nice! Staircase.","metadata":{}},{"cell_type":"markdown","source":"For sampling rate :","metadata":{}},{"cell_type":"code","source":"set(srs)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T21:27:49.770029Z","iopub.execute_input":"2023-07-25T21:27:49.770499Z","iopub.status.idle":"2023-07-25T21:27:49.778639Z","shell.execute_reply.started":"2023-07-25T21:27:49.770463Z","shell.execute_reply":"2023-07-25T21:27:49.777789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So only one value. All sampled at 32k.","metadata":{}},{"cell_type":"code","source":"plt.xlabel(\"Sampling Rate\")\nplt.ylabel(\"Frequency\")\nplt.title(\"Sampling rate Distribuition\")\nplt.hist(srs)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T21:27:23.413041Z","iopub.execute_input":"2023-07-25T21:27:23.413493Z","iopub.status.idle":"2023-07-25T21:27:23.816537Z","shell.execute_reply.started":"2023-07-25T21:27:23.413459Z","shell.execute_reply":"2023-07-25T21:27:23.815158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#To DO : Audio length, sampling rate analysis","metadata":{},"execution_count":null,"outputs":[]}]}