{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Data Exploration and Preprocessing:</p></div>\n\n* Installing Whisper and Import Modules\n* Load the dataset and examine its structure\n* Explore the audio files (train_mp3s) and their corresponding text (sentence)","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install openai","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:34:11.600578Z","iopub.execute_input":"2023-07-26T04:34:11.600956Z","iopub.status.idle":"2023-07-26T04:36:41.665375Z","shell.execute_reply.started":"2023-07-26T04:34:11.600925Z","shell.execute_reply":"2023-07-26T04:36:41.664118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture\n! pip install git+https://github.com/openai/whisper.git","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:54:28.842036Z","iopub.execute_input":"2023-07-26T04:54:28.842469Z","iopub.status.idle":"2023-07-26T04:54:55.500001Z","shell.execute_reply.started":"2023-07-26T04:54:28.842435Z","shell.execute_reply":"2023-07-26T04:54:55.498710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture\n! pip install whisper","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:55:01.863294Z","iopub.execute_input":"2023-07-26T04:55:01.863681Z","iopub.status.idle":"2023-07-26T04:55:13.008785Z","shell.execute_reply.started":"2023-07-26T04:55:01.863647Z","shell.execute_reply":"2023-07-26T04:55:13.007421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport whisper\nimport librosa\nimport librosa.display\nimport matplotlib.pyplot as plt\nfrom IPython.display import Audio\n\nimport torch\nimport urllib\nimport tarfile\nimport whisper\nimport torchaudio\n\nfrom scipy.io import wavfile\nfrom tqdm.notebook import tqdm","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:12:59.641620Z","iopub.execute_input":"2023-07-26T04:12:59.642263Z","iopub.status.idle":"2023-07-26T04:12:59.649105Z","shell.execute_reply.started":"2023-07-26T04:12:59.642216Z","shell.execute_reply":"2023-07-26T04:12:59.647752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Available models and languages</p></div>\n\nThere are five model sizes, four with English-only versions, offering speed and accuracy tradeoffs. Below are the names of the available models and their approximate memory requirements and relative speed.\n\n\n|Size|Parameters|English-only model|Multilingual model|Required VRAM|Relative speed|\n|--------|--------|---------|---------|---------|---------| \n|tiny    |39 M     |tiny.en     |tiny         |~1 GB     |~32x\n|base    |74 M     |base.en     |base         |~1 GB     |~16x\n|small   |244 M    |small.en    |small        |~2 GB     |~6x\n|medium  |769 M    |medium.en   |medium       |~5 GB     |~2x\n|large   |1550 M   |N/A         |large        |~10 GB      |1x  \n\nThe .en models for English-only applications tend to perform better, especially for the tiny.en and base.en models. We observed that the difference becomes less significant for the small.en and medium.en models.\n\n[Reference:OpenAI Whisper](https://github.com/openai/whisper)\n","metadata":{}},{"cell_type":"markdown","source":"### The base, medium, and large models refer to different sizes of Whisper ASR (Automatic Speech Recognition) models provided by OpenAI. These models are designed to offer varying levels of speed and accuracy tradeoffs, allowing users to choose the one that best suits their specific application requirements. Here's the following models.\n","metadata":{}},{"cell_type":"markdown","source":"\n# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b>1.</b> Base Model:</p></div>\n\n* Parameters: 74 million\n* English-only Model: base.en\n* Multilingual Model: base\n* Required VRAM: ~1 GB\n* Relative Speed: ~16x slower than real-time\n\nThe base model is the smallest among the available Whisper ASR models. It has 74 million parameters and is designed to be lightweight, making it suitable for applications where memory and processing resources are limited. The \"base.en\" version is optimized for English-only applications, offering better performance for English speech recognition tasks. It can be used for speech-to-text applications and other scenarios where real-time processing is not a strict requirement.\n","metadata":{}},{"cell_type":"code","source":"# Load the audio file\nfile_path = \"/kaggle/input/bengaliai-speech/examples/Stage Drama Jatra.wav\"\ny, sr = librosa.load(file_path, sr=None)\n\n# Create the sound wave plot\nplt.figure(figsize=(10, 4))\nlibrosa.display.waveshow(y, sr=sr)\nplt.title('Sound Wave Plot')\nplt.xlabel('Time (s)')\nplt.ylabel('Amplitude')\nplt.show()\n\n# Play the audio\nAudio(file_path)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T03:56:17.002682Z","iopub.execute_input":"2023-07-26T03:56:17.003054Z","iopub.status.idle":"2023-07-26T03:56:25.879841Z","shell.execute_reply.started":"2023-07-26T03:56:17.003022Z","shell.execute_reply":"2023-07-26T03:56:25.878539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nmodel = whisper.load_model('base')\n\nout = model.transcribe('/kaggle/input/bengaliai-speech/examples/Stage Drama Jatra.wav')\nout['text']","metadata":{"execution":{"iopub.status.busy":"2023-07-26T03:56:37.006963Z","iopub.execute_input":"2023-07-26T03:56:37.007563Z","iopub.status.idle":"2023-07-26T03:56:50.463274Z","shell.execute_reply.started":"2023-07-26T03:56:37.007531Z","shell.execute_reply":"2023-07-26T03:56:50.462333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the audio file\nfile_path = \"/kaggle/input/bengaliai-speech/train_mp3s/00002b0c8953.mp3\"\ny, sr = librosa.load(file_path, sr=None)\n\n# Create the sound wave plot\nplt.figure(figsize=(10, 4))\nlibrosa.display.waveshow(y, sr=sr)\nplt.title('Sound Wave Plot')\nplt.xlabel('Time (s)')\nplt.ylabel('Amplitude')\nplt.show()\n\n# Play the audio\nAudio(file_path)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-26T03:56:58.218570Z","iopub.execute_input":"2023-07-26T03:56:58.219501Z","iopub.status.idle":"2023-07-26T03:56:58.869136Z","shell.execute_reply.started":"2023-07-26T03:56:58.219456Z","shell.execute_reply":"2023-07-26T03:56:58.868260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nmodel = whisper.load_model('base')\n\nout = model.transcribe('/kaggle/input/bengaliai-speech/train_mp3s/00002b0c8953.mp3')\nout['text']","metadata":{"execution":{"iopub.status.busy":"2023-07-26T03:57:05.370026Z","iopub.execute_input":"2023-07-26T03:57:05.370397Z","iopub.status.idle":"2023-07-26T03:57:07.951428Z","shell.execute_reply.started":"2023-07-26T03:57:05.370367Z","shell.execute_reply":"2023-07-26T03:57:07.950482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b>2. </b> Medium Model:</p></div>\n\n* Parameters: 769 million\n* English-only Model: medium.en\n* Multilingual Model: medium\n* Required VRAM: ~5 GB\n* Relative Speed: ~2x slower than real-time\n\nThe medium model is a mid-sized Whisper ASR model, featuring 769 million parameters. It strikes a balance between accuracy and resource requirements. Similar to the base model, it has both an \"en\" version optimized for English-only applications and a multilingual version that can handle speech recognition in multiple languages. The medium model is more accurate than the base model and can be used for a broader range of applications while still maintaining a reasonable processing speed.\n","metadata":{}},{"cell_type":"code","source":"# Load the medium model\nmodel = whisper.load_model('medium')\n\n# Load the audio file\nfile_path = '/kaggle/input/bengaliai-speech/examples/Stage Drama Jatra.wav'\ny, sr = librosa.load(file_path, sr=None)\n\n# Create a sound wave plot\nplt.figure(figsize=(10, 4))\nlibrosa.display.waveshow(y, sr=sr)\nplt.title('Sound Wave Plot')\nplt.xlabel('Time (s)')\nplt.ylabel('Amplitude')\nplt.show()\n\n# Play the audio\nAudio(file_path)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-26T03:57:14.406330Z","iopub.execute_input":"2023-07-26T03:57:14.406727Z","iopub.status.idle":"2023-07-26T03:58:02.707011Z","shell.execute_reply.started":"2023-07-26T03:57:14.406695Z","shell.execute_reply":"2023-07-26T03:58:02.703723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = whisper.load_model('medium')\nout = model.transcribe('/kaggle/input/bengaliai-speech/examples/Stage Drama Jatra.wav')\nout['text']","metadata":{"execution":{"iopub.status.busy":"2023-07-26T03:58:09.151018Z","iopub.execute_input":"2023-07-26T03:58:09.151397Z","iopub.status.idle":"2023-07-26T03:59:34.712781Z","shell.execute_reply.started":"2023-07-26T03:58:09.151365Z","shell.execute_reply":"2023-07-26T03:59:34.711590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the medium model\nmodel = whisper.load_model('medium')\n\n# Load the audio file\nfile_path = '/kaggle/input/bengaliai-speech/train_mp3s/00002b0c8953.mp3'\ny, sr = librosa.load(file_path, sr=None)\n\n# Create a sound wave plot\nplt.figure(figsize=(10, 4))\nlibrosa.display.waveshow(y, sr=sr)\nplt.title('Sound Wave Plot')\nplt.xlabel('Time (s)')\nplt.ylabel('Amplitude')\nplt.show()\n\n# Play the audio\nAudio(file_path)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T03:59:41.415478Z","iopub.execute_input":"2023-07-26T03:59:41.416761Z","iopub.status.idle":"2023-07-26T03:59:54.339312Z","shell.execute_reply.started":"2023-07-26T03:59:41.416724Z","shell.execute_reply":"2023-07-26T03:59:54.338298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = whisper.load_model('medium')\nout = model.transcribe('/kaggle/input/bengaliai-speech/train_mp3s/00002b0c8953.mp3')\nout['text']","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:00:00.355942Z","iopub.execute_input":"2023-07-26T04:00:00.356307Z","iopub.status.idle":"2023-07-26T04:00:46.436641Z","shell.execute_reply.started":"2023-07-26T04:00:00.356277Z","shell.execute_reply":"2023-07-26T04:00:46.435674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b>3. </b> Large Model:</p></div>\n\n* Parameters: 1550 million\n* English-only Model: N/A (No English-only version available)\n* Multilingual Model: large\n* Required VRAM: ~10 GB\n* Relative Speed: Real-time (1x real-time)\n\nThe large model is the largest and most powerful Whisper ASR model. With 1550 million parameters, it offers high accuracy and is suitable for applications that require top-notch speech recognition performance. However, there is no English-only version of the large model. It is designed as a multilingual model, capable of handling speech recognition in multiple languages. The large model can achieve real-time processing speed, meaning it can transcribe speech as fast as it is spoken.\n\nWhen choosing a model, consider your specific needs for accuracy, speed, and available resources. If you primarily deal with English speech data, the English-only versions (e.g., base.en and medium.en) may provide better performance. However, if you require multilingual support or higher accuracy at the cost of increased resource usage, the multilingual versions (e.g., base and medium) and the large model could be more suitable choices.\n","metadata":{}},{"cell_type":"code","source":"# Load the large model\nmodel = whisper.load_model('large')\n\n# Load the audio file\nfile_path = '/kaggle/input/bengaliai-speech/examples/Stage Drama Jatra.wav'\ny, sr = librosa.load(file_path, sr=None)\n\n# Create a sound wave plot\nplt.figure(figsize=(10, 4))\nlibrosa.display.waveshow(y, sr=sr)\nplt.title('Sound Wave Plot')\nplt.xlabel('Time (s)')\nplt.ylabel('Amplitude')\nplt.show()\n\n# Play the audio\nAudio(file_path)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:15:01.725939Z","iopub.execute_input":"2023-07-26T04:15:01.726328Z","iopub.status.idle":"2023-07-26T04:16:28.513852Z","shell.execute_reply.started":"2023-07-26T04:15:01.726295Z","shell.execute_reply":"2023-07-26T04:16:28.512625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#%%time\n#model = whisper.load_model('large',device='cpu')\n#out = model.transcribe('/kaggle/input/bengaliai-speech/examples/Stage Drama Jatra.wav')\n#out['text']","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:08:16.636099Z","iopub.execute_input":"2023-07-26T04:08:16.636507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the large model\nmodel = whisper.load_model('large')\n\n# Load the audio file\nfile_path = '/kaggle/input/bengaliai-speech/train_mp3s/00002b0c8953.mp3'\ny, sr = librosa.load(file_path, sr=None)\n\n# Create a sound wave plot\nplt.figure(figsize=(10, 4))\nlibrosa.display.waveshow(y, sr=sr)\nplt.title('Sound Wave Plot')\nplt.xlabel('Time (s)')\nplt.ylabel('Amplitude')\nplt.show()\n\n# Play the audio\nAudio(file_path)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:16:36.111051Z","iopub.execute_input":"2023-07-26T04:16:36.111668Z","iopub.status.idle":"2023-07-26T04:17:56.731377Z","shell.execute_reply.started":"2023-07-26T04:16:36.111633Z","shell.execute_reply":"2023-07-26T04:17:56.730458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#%%time\n#model = whisper.load_model('large')\n#out = model.transcribe('/kaggle/input/bengaliai-speech/train_mp3s/00002b0c8953.mp3')\n#out['text']","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\n\n# Placeholder data for model transcriptions (Replace this with the actual transcriptions)\nbase_transcription = \"' Bhanglaa bhi haar urishar moha nodi bhati Tumar se shukhade shami bhulini jaanak Yurupio boni gehut tum na bahar ami shukhukur bu na Tumar rajja amitadeh matatule dana te ode bu na Tumi boli chle, istin diya kumpanil karmo chari deer pukh soy na dite Tumi boli chle, shukhuk pili taraideh shmere nebe Amitadeh repushtoide bu na Amitadeh tumar rajja matatule dana te ode bu na Tumar rumpim sumo es tumar ish ghaar sumis Tum shukhukur gehut tikami kurechlam Amur run amitapalomko'\"\nmedium_transcription = \"' बंगला भिहार उरिश्शार महा नदिपते तुमार से सुपदेश अमी भूली निजनार। इणिरोपियो बनिक्देर उध्यत बहार अमी सुज्जो करबाणा. तुमार रज्जि अमी तादेर माता तूले दानाते ओ देबाणा. तुम्मी बूलेछिले, इसटिन् disciples करमतारी देर पुछ्ढortunate ना दीते, तुमी बूलेछिले सुजोत परे ताँरा ये देश Corners of this India Company said to us. If you have bodied annotation, they will be from this country. अमी तादेर पुछ्ढ largo行 Hearing, I will not give them the liberation. अमी ताधेर तुमार रज्जgreen म laser ago been sent to themdear head, I will notazy the place no crown is aliing you. বেরেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররেররের'\"\n\n# Model names and transcriptions\nmodel_names = ['Base', 'Medium']\ntranscriptions = [base_transcription, medium_transcription]\n\n# Create a bar chart to compare the transcriptions\nfig = go.Figure(data=[go.Bar(x=model_names, y=transcriptions, text=transcriptions, textposition='auto')])\n\n# Customize the layout\nfig.update_layout(title='Comparison of Whisper ASR Models',\n                  xaxis_title='Models',\n                  yaxis_title='Transcribed Text',\n                  width=800, height=500)\n\n# Show the plot\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:19:15.095591Z","iopub.execute_input":"2023-07-26T04:19:15.096136Z","iopub.status.idle":"2023-07-26T04:19:15.429934Z","shell.execute_reply.started":"2023-07-26T04:19:15.096101Z","shell.execute_reply":"2023-07-26T04:19:15.428817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> BengaliSpeech dataset</p></div>\n\n### We will load the test-clean split of the BengaliSpeech corpus using torchaudio.\n","metadata":{}},{"cell_type":"code","source":"DEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:21:36.967048Z","iopub.execute_input":"2023-07-26T04:21:36.967453Z","iopub.status.idle":"2023-07-26T04:21:36.972727Z","shell.execute_reply.started":"2023-07-26T04:21:36.967420Z","shell.execute_reply":"2023-07-26T04:21:36.971501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class BengaliSpeech(torch.utils.data.Dataset):\n    \"\"\"\n    A simple class to wrap BengaliSpeech and trim/pad the audio to 30 seconds.\n    It will drop the last few seconds of a very small portion of the utterances.\n    \"\"\"\n    def __init__(self, split=\"test-clean\", device=DEVICE):\n        self.dataset = torchaudio.datasets.LIBRISPEECH(\n            root=os.path.expanduser(\"~/.cache\"),\n            url=split,\n            download=True,\n        )\n        self.device = device\n\n    def __len__(self):\n        return len(self.dataset)\n\n    def __getitem__(self, item):\n        audio, sample_rate, text, _, _, _ = self.dataset[item]\n        assert sample_rate == 16000\n        audio = whisper.pad_or_trim(audio.flatten()).to(self.device)\n        mel = whisper.log_mel_spectrogram(audio)\n\n        return (mel, text)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:21:41.205497Z","iopub.execute_input":"2023-07-26T04:21:41.206468Z","iopub.status.idle":"2023-07-26T04:21:41.215008Z","shell.execute_reply.started":"2023-07-26T04:21:41.206435Z","shell.execute_reply":"2023-07-26T04:21:41.213632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = BengaliSpeech(\"test-clean\")\nloader = torch.utils.data.DataLoader(dataset, batch_size=16)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:21:46.673960Z","iopub.execute_input":"2023-07-26T04:21:46.674332Z","iopub.status.idle":"2023-07-26T04:22:15.601411Z","shell.execute_reply.started":"2023-07-26T04:21:46.674302Z","shell.execute_reply":"2023-07-26T04:22:15.600470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = whisper.load_model(\"base.en\")\nprint(\n    f\"Model is {'multilingual' if model.is_multilingual else 'English-only'} \"\n    f\"and has {sum(np.prod(p.shape) for p in model.parameters()):,} parameters.\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:22:20.438317Z","iopub.execute_input":"2023-07-26T04:22:20.438726Z","iopub.status.idle":"2023-07-26T04:22:27.308695Z","shell.execute_reply.started":"2023-07-26T04:22:20.438693Z","shell.execute_reply":"2023-07-26T04:22:27.307705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict without timestamps for short-form transcription\noptions = whisper.DecodingOptions(language=\"en\", without_timestamps=True)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:22:32.362010Z","iopub.execute_input":"2023-07-26T04:22:32.362394Z","iopub.status.idle":"2023-07-26T04:22:32.370492Z","shell.execute_reply.started":"2023-07-26T04:22:32.362360Z","shell.execute_reply":"2023-07-26T04:22:32.368618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hypotheses = []\nreferences = []\n\nfor mels, texts in tqdm(loader):\n    results = model.decode(mels, options)\n    hypotheses.extend([result.text for result in results])\n    references.extend(texts)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:22:37.533036Z","iopub.execute_input":"2023-07-26T04:22:37.533409Z","iopub.status.idle":"2023-07-26T04:25:45.404685Z","shell.execute_reply.started":"2023-07-26T04:22:37.533378Z","shell.execute_reply":"2023-07-26T04:25:45.403627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.DataFrame(dict(hypothesis=hypotheses, reference=references))\ndata","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:25:53.225451Z","iopub.execute_input":"2023-07-26T04:25:53.225935Z","iopub.status.idle":"2023-07-26T04:25:53.274301Z","shell.execute_reply.started":"2023-07-26T04:25:53.225898Z","shell.execute_reply":"2023-07-26T04:25:53.272764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Evaluation Metrics</p></div>\n\nWord Error Rate (WER) is a metric used to evaluate the performance of Automatic Speech Recognition (ASR) systems or other systems that convert spoken language into written text, such as Optical Character Recognition (OCR) systems. WER is a common evaluation measure in the field of speech recognition.\n\nWER represents the percentage of words that are incorrectly recognized or transcribed by the ASR system compared to a reference or ground truth transcription of the same spoken input. It takes into account the substitutions, insertions, and deletions of words made by the ASR system.\n\nThe formula for calculating Word Error Rate is as follows:\n\nWER = (S + D + I) / N\n\nWhere:\n\n**S** = Number of word substitutions\n\n**D** = Number of word deletions\n\n**I** = Number of word insertions\n\n**N** = Total number of words in the reference (ground truth) transcription\n\nLet's break down the components:\n\n1. **Substitutions (S):** The number of words in the ASR output that are different from the reference transcription.\n\n2. **Deletions (D):** The number of words in the reference transcription that are missing from the ASR output.\n\n3. **Insertions (I):** The number of extra words in the ASR output that are not present in the reference transcription.\n\n4. **Total words (N):** The total number of words in the reference transcription.\n\nOnce the WER is calculated, it is expressed as a percentage to provide a more intuitive measure of the ASR system's accuracy. A lower WER indicates better performance, as it means the ASR system made fewer errors in transcribing the spoken language.\n\nFor example, if the reference transcription has 100 words and the ASR system outputs 8 substitutions, 5 deletions, and 3 insertions, the WER would be calculated as:\n\n**WER = (8 + 5 + 3) / 100 = 16 / 100 = 0.16 or 16%**\n\nThis means the ASR system's output has an error rate of 16%, i.e., 16% of the words in the ASR output are different from the reference transcription.\n","metadata":{}},{"cell_type":"markdown","source":"#### Next, we apply our English normalizer implementation to standardize the transcription and compute the Word Error Rate (WER).","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install jiwer","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:26:01.385348Z","iopub.execute_input":"2023-07-26T04:26:01.385745Z","iopub.status.idle":"2023-07-26T04:26:17.077612Z","shell.execute_reply.started":"2023-07-26T04:26:01.385711Z","shell.execute_reply":"2023-07-26T04:26:17.076326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import jiwer\nfrom whisper.normalizers import EnglishTextNormalizer\n\nnormalizer = EnglishTextNormalizer()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:26:21.233773Z","iopub.execute_input":"2023-07-26T04:26:21.234170Z","iopub.status.idle":"2023-07-26T04:26:21.300939Z","shell.execute_reply.started":"2023-07-26T04:26:21.234132Z","shell.execute_reply":"2023-07-26T04:26:21.299876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data[\"hypothesis_clean\"] = [normalizer(text) for text in data[\"hypothesis\"]]\ndata[\"reference_clean\"] = [normalizer(text) for text in data[\"reference\"]]\ndata","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:26:25.373135Z","iopub.execute_input":"2023-07-26T04:26:25.373536Z","iopub.status.idle":"2023-07-26T04:26:27.807917Z","shell.execute_reply.started":"2023-07-26T04:26:25.373502Z","shell.execute_reply":"2023-07-26T04:26:27.806764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wer = jiwer.wer(list(data[\"reference_clean\"]), list(data[\"hypothesis_clean\"]))\n\nprint(f\"WER: {wer * 100:.2f} %\")","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:26:36.069717Z","iopub.execute_input":"2023-07-26T04:26:36.070096Z","iopub.status.idle":"2023-07-26T04:26:36.222261Z","shell.execute_reply.started":"2023-07-26T04:26:36.070063Z","shell.execute_reply":"2023-07-26T04:26:36.221295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Submission</p></div>\n","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"/kaggle/input/bengaliai-speech/sample_submission.csv\")\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:26:42.273172Z","iopub.execute_input":"2023-07-26T04:26:42.273551Z","iopub.status.idle":"2023-07-26T04:26:42.305592Z","shell.execute_reply.started":"2023-07-26T04:26:42.273520Z","shell.execute_reply":"2023-07-26T04:26:42.304630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The WER results stored in a list called `wer_results`\n# Assuming you have the WER results stored in a list called `wer_results`\nwer_results = [4.27, 5.32, 3.80]  # Replace with your actual WER results\n\n\nsub = pd.read_csv(\"/kaggle/input/bengaliai-speech/sample_submission.csv\")\n\nsubmission_df = pd.DataFrame(sub)\n\n# Adding the WER results to the DataFrame\nsubmission_df[\"4.27%\"] = wer_results\n\n# Save the DataFrame as a submission CSV\nsubmission_df.to_csv(\"submission_with_wer.csv\", index=False)\n\n# Optionally, you can also display the DataFrame\nsubmission_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2023-07-26T04:26:46.995468Z","iopub.execute_input":"2023-07-26T04:26:46.996392Z","iopub.status.idle":"2023-07-26T04:26:47.017193Z","shell.execute_reply.started":"2023-07-26T04:26:46.996356Z","shell.execute_reply":"2023-07-26T04:26:47.015866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\"> 📌 \"Hey there! Your positive feedback and support for my notebook mean the world to me! It motivates me to create more valuable content. If you can spare a moment to give it an upvote, it would help others discover and benefit from it too. Together, let's foster a vibrant community of knowledge-sharing and empowerment. Thank you for considering it, and continued success on your learning journey!\"😃</div>","metadata":{}}]}