{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#EE82EE;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Dataset Description:</p></div>\n\nThe competition dataset comprises about 1200 hours of recordings of Bengali speech. Your goal is to transcribe recordings of speech that is out-of-distribution with respect to the training set.\n\nNote that this is a Code Competition, in which the actual test set is hidden. In this public version, we give some sample data in the correct format to help you author your solutions. The full test set contains about 20 hours of speech in almost 8000 MP3 audio files. All of the files in the test set are encoded at a sample rate of 32k, a bit rate of 48k, in one channel.\n\nDetails on the dataset are available in the dataset paper: https://arxiv.org/abs/2305.09688","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#EE82EE;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Files and Field Descriptions:</p></div>\n\n\n- **train/ The training set**, comprising several thousand recordings in MP3 format.\n- **test/ The test set**, comprising spontaneous speech recordings from eighteen domains, seventeen of which are out-of-distribution with respect to the training set. There may be domains in the private test set that are not in the public test set.\n- **examples/** An example recording for each test set domain. You may find these example recordings helpful for creating models robust to domain variation. These are representative recordings and none of them are present in the test set.\n- **train.csv** Sentence labels for the training set.\n  * `id` A unique identifier for this instance. Corresponds to the file `{id}.mp3` in train/.\n  * `sentence` A plain-text transcription of the recording. Your goal is to predict these sentences for each recording in the test set.\n  * `split` Whether `train` or `valid`. The annotations in the `valid` split have been manually reviewed and corrected, while the annotations in the `train` split have only been algorithmically cleaned. The `valid` samples will generally have higher quality annotations than the `train` samples, but are otherwise drawn from the same distribution.\n- **sample_submission.csv** A sample submission file in the correct format. See the Evaluation page for more details.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#EE82EE;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Data Exploration and Preprocessing:</p></div>\n\n* Installing Whisper and Import Modules \n* Load the dataset and examine its structure\n* Explore the audio files (train_mp3s) and their corresponding text (sentence)\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFA500;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:150%;letter-spacing:0.5px;margin:0\"><b> </b> Installing Whisper</p></div>","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install -U openai-whisper","metadata":{"execution":{"iopub.status.busy":"2023-07-25T03:55:51.179442Z","iopub.execute_input":"2023-07-25T03:55:51.179840Z","iopub.status.idle":"2023-07-25T03:56:03.276762Z","shell.execute_reply.started":"2023-07-25T03:55:51.179807Z","shell.execute_reply":"2023-07-25T03:56:03.275422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture\n! pip install git+https://github.com/openai/whisper.git\n! pip install jiwer","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:19:37.122146Z","iopub.execute_input":"2023-07-25T04:19:37.122588Z","iopub.status.idle":"2023-07-25T04:20:14.467073Z","shell.execute_reply.started":"2023-07-25T04:19:37.122548Z","shell.execute_reply":"2023-07-25T04:20:14.465672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#%%capture\n#!pip install git+https://github.com/openai/whisper.git -q\n#!sudo apt update && sudo apt install ffmpeg","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:00:39.041786Z","iopub.execute_input":"2023-07-25T04:00:39.042555Z","iopub.status.idle":"2023-07-25T04:01:08.434956Z","shell.execute_reply.started":"2023-07-25T04:00:39.042516Z","shell.execute_reply":"2023-07-25T04:01:08.433578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture\n!pip install --upgrade torch pandas whisper torchaudio tensorflow tensorflow-io","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:03:01.695803Z","iopub.execute_input":"2023-07-25T04:03:01.696205Z","iopub.status.idle":"2023-07-25T04:03:14.649959Z","shell.execute_reply.started":"2023-07-25T04:03:01.696170Z","shell.execute_reply":"2023-07-25T04:03:14.648518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import io\nimport os\nimport numpy as np\n\nimport torch\nimport pandas as pd\nimport urllib\nimport tarfile\nimport whisper\nimport torchaudio\n\nimport librosa\nimport librosa.display\nimport matplotlib.pyplot as plt\n\n\nfrom scipy.io import wavfile\nfrom tqdm.notebook import tqdm","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:01:27.728229Z","iopub.execute_input":"2023-07-25T04:01:27.728645Z","iopub.status.idle":"2023-07-25T04:01:27.735415Z","shell.execute_reply.started":"2023-07-25T04:01:27.728609Z","shell.execute_reply":"2023-07-25T04:01:27.734241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset and examine its structure\ndataset = pd.read_csv(\"/kaggle/input/bengaliai-speech/train.csv\")\nprint(dataset.shape)\ndataset.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:01:43.040620Z","iopub.execute_input":"2023-07-25T04:01:43.041002Z","iopub.status.idle":"2023-07-25T04:01:46.568408Z","shell.execute_reply.started":"2023-07-25T04:01:43.040974Z","shell.execute_reply":"2023-07-25T04:01:46.567424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Explore the Audio Files (train_mp3s) and Their Corresponding Text (sentence)\nsample_audio_file = \"train_mp3s/000005f3362c.mp3\"\nsample_sentence = dataset[dataset[\"id\"] == \"000005f3362c\"][\"sentence\"].values[0]\n","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:01:55.845394Z","iopub.execute_input":"2023-07-25T04:01:55.845773Z","iopub.status.idle":"2023-07-25T04:01:56.006016Z","shell.execute_reply.started":"2023-07-25T04:01:55.845743Z","shell.execute_reply":"2023-07-25T04:01:56.004965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import IPython.display as ipd\n#ipd.Audio(sample_audio_file)\n\nprint(\"Sample Audio File:\", sample_audio_file)\nprint(\"Sample Sentence:\", sample_sentence)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:02:04.308123Z","iopub.execute_input":"2023-07-25T04:02:04.308541Z","iopub.status.idle":"2023-07-25T04:02:04.315111Z","shell.execute_reply.started":"2023-07-25T04:02:04.308506Z","shell.execute_reply":"2023-07-25T04:02:04.314014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#EE82EE;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Data Visualization:</p></div>\n\n* Visualize audio waveforms and spectrograms of different classes\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Harmonic and percussive components with transparency</p></div>\n\n","metadata":{}},{"cell_type":"code","source":"from IPython.display import Audio\nAudio(\"/kaggle/input/bengaliai-speech/examples/Waz Islamic Sermon.wav\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:02:10.753429Z","iopub.execute_input":"2023-07-25T04:02:10.754457Z","iopub.status.idle":"2023-07-25T04:02:11.030904Z","shell.execute_reply.started":"2023-07-25T04:02:10.754415Z","shell.execute_reply":"2023-07-25T04:02:11.029578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import librosa.display\n\nfile_path = '/kaggle/input/bengaliai-speech/examples/Waz Islamic Sermon.wav'\ny, sr = librosa.load(file_path, duration=15)\n\n# Create a 3x1 subplot grid\nfig, ax = plt.subplots(nrows=3, ncols=1, sharex=True)\n\n# Plot the waveform\nlibrosa.display.waveshow(y, sr=sr, alpha=0.5, ax=ax[0])\nax[0].set(title='Waveform')\n\n# Decompose the signal into harmonic and percussive components\ny_harm, y_perc = librosa.effects.hpss(y)\n\n# Plot the harmonic component\nlibrosa.display.waveshow(y_harm, sr=sr, alpha=0.5, ax=ax[1], label='Harmonic')\nax[1].set(title='Harmonic Component')\nax[1].legend()\n\n# Plot the percussive component\nlibrosa.display.waveshow(y_perc, sr=sr, color='r', alpha=0.5, ax=ax[2], label='Percussive')\nax[2].set(title='Percussive Component')\nax[2].legend()\n\n# Adjust layout and display the plots\nplt.tight_layout()\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-25T04:02:18.185403Z","iopub.execute_input":"2023-07-25T04:02:18.185806Z","iopub.status.idle":"2023-07-25T04:02:30.157659Z","shell.execute_reply.started":"2023-07-25T04:02:18.185776Z","shell.execute_reply":"2023-07-25T04:02:30.156291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Plotting a transposed wave along with a self-similarity matrix</p></div>\n","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplot_mosaic(\"hSSS;hSSS;hSSS;.vvv\")\ny, sr = librosa.load(librosa.ex('trumpet'))\nchroma = librosa.feature.chroma_cqt(y=y, sr=sr)\nsim = librosa.segment.recurrence_matrix(chroma, mode='affinity')\nlibrosa.display.specshow(sim, ax=ax['S'], sr=sr,\n                         x_axis='time', y_axis='time',\n                         auto_aspect=False)\nax['S'].label_outer()\nax['S'].sharex(ax['v'])\nax['S'].sharey(ax['h'])\nax['S'].set(title='Self-similarity')\nlibrosa.display.waveshow(y, ax=ax['v'])\nax['v'].label_outer()\nax['v'].set(title='transpose=False')\nlibrosa.display.waveshow(y, ax=ax['h'], transpose=True)\nax['h'].label_outer()\nax['h'].set(title='transpose=True')","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:03:25.692091Z","iopub.execute_input":"2023-07-25T04:03:25.692530Z","iopub.status.idle":"2023-07-25T04:03:27.642040Z","shell.execute_reply.started":"2023-07-25T04:03:25.692493Z","shell.execute_reply":"2023-07-25T04:03:27.641075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Tempogram - Onset autocorrelation</p></div>","metadata":{}},{"cell_type":"code","source":"# Compute local onset autocorrelation\nfile_path = '/kaggle/input/bengaliai-speech/examples/Waz Islamic Sermon.wav'\ny, sr = librosa.load(file_path, duration=30)\n\nhop_length = 512\noenv = librosa.onset.onset_strength(y=y, sr=sr, hop_length=hop_length)\ntempogram = librosa.feature.tempogram(onset_envelope=oenv, sr=sr,\n                                      hop_length=hop_length)\n# Compute global onset autocorrelation\nac_global = librosa.autocorrelate(oenv, max_size=tempogram.shape[0])\nac_global = librosa.util.normalize(ac_global)\n# Estimate the global tempo for display purposes\ntempo = librosa.feature.tempo(onset_envelope=oenv, sr=sr,\n                              hop_length=hop_length)[0]","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:03:33.525679Z","iopub.execute_input":"2023-07-25T04:03:33.526060Z","iopub.status.idle":"2023-07-25T04:03:33.906101Z","shell.execute_reply.started":"2023-07-25T04:03:33.526029Z","shell.execute_reply":"2023-07-25T04:03:33.905041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfig, ax = plt.subplots(nrows=4, figsize=(10, 10))\ntimes = librosa.times_like(oenv, sr=sr, hop_length=hop_length)\nax[0].plot(times, oenv, label='Onset strength')\nax[0].label_outer()\nax[0].legend(frameon=True)\nlibrosa.display.specshow(tempogram, sr=sr, hop_length=hop_length,\n                         x_axis='time', y_axis='tempo', cmap='magma',\n                         ax=ax[1])\nax[1].axhline(tempo, color='w', linestyle='--', alpha=1,\n            label='Estimated tempo={:g}'.format(tempo))\nax[1].legend(loc='upper right')\nax[1].set(title='Tempogram')\nx = np.linspace(0, tempogram.shape[0] * float(hop_length) / sr,\n                num=tempogram.shape[0])\nax[2].plot(x, np.mean(tempogram, axis=1), label='Mean local autocorrelation')\nax[2].plot(x, ac_global, '--', alpha=0.75, label='Global autocorrelation')\nax[2].set(xlabel='Lag (seconds)')\nax[2].legend(frameon=True)\nfreqs = librosa.tempo_frequencies(tempogram.shape[0], hop_length=hop_length, sr=sr)\nax[3].semilogx(freqs[1:], np.mean(tempogram[1:], axis=1),\n             label='Mean local autocorrelation', base=2)\nax[3].semilogx(freqs[1:], ac_global[1:], '--', alpha=0.75,\n             label='Global autocorrelation', base=2)\nax[3].axvline(tempo, color='black', linestyle='--', alpha=.8,\n            label='Estimated tempo={:g}'.format(tempo))\nax[3].legend(frameon=True)\nax[3].set(xlabel='BPM')\nax[3].grid(True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-25T04:03:39.615843Z","iopub.execute_input":"2023-07-25T04:03:39.616247Z","iopub.status.idle":"2023-07-25T04:03:41.499771Z","shell.execute_reply.started":"2023-07-25T04:03:39.616212Z","shell.execute_reply":"2023-07-25T04:03:41.498658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compute local onset autocorrelation\nfile_path = '/kaggle/input/bengaliai-speech/examples/Waz Islamic Sermon.wav'\ny, sr = librosa.load(file_path, duration=30)\n\nhop_length = 512\noenv = librosa.onset.onset_strength(y=y, sr=sr, hop_length=hop_length)\ntempogram = librosa.feature.fourier_tempogram(onset_envelope=oenv, sr=sr,\n                                              hop_length=hop_length)\n# Compute the auto-correlation tempogram, unnormalized to make comparison easier\nac_tempogram = librosa.feature.tempogram(onset_envelope=oenv, sr=sr,\n                                         hop_length=hop_length, norm=None)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:03:48.975815Z","iopub.execute_input":"2023-07-25T04:03:48.976206Z","iopub.status.idle":"2023-07-25T04:03:49.178949Z","shell.execute_reply.started":"2023-07-25T04:03:48.976174Z","shell.execute_reply":"2023-07-25T04:03:49.177554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(nrows=3, sharex=True)\nax[0].plot(librosa.times_like(oenv), oenv, label='Onset strength')\nax[0].legend(frameon=True)\nax[0].label_outer()\nlibrosa.display.specshow(np.abs(tempogram), sr=sr, hop_length=hop_length,\n                         x_axis='time', y_axis='fourier_tempo', cmap='magma',\n                         ax=ax[1])\nax[1].set(title='Fourier tempogram')\nax[1].label_outer()\nlibrosa.display.specshow(ac_tempogram, sr=sr, hop_length=hop_length,\n                         x_axis='time', y_axis='tempo', cmap='magma',\n                         ax=ax[2])\nax[2].set(title='Autocorrelation tempogram')","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:03:54.491443Z","iopub.execute_input":"2023-07-25T04:03:54.491850Z","iopub.status.idle":"2023-07-25T04:03:55.747366Z","shell.execute_reply.started":"2023-07-25T04:03:54.491811Z","shell.execute_reply":"2023-07-25T04:03:55.746408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Tempogram ratio features</p></div>\n\nCompute tempogram ratio features using the default factors for a waltz (3/4 time)","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ntempogram = librosa.feature.tempogram(y=y, sr=sr)\ntgr = librosa.feature.tempogram_ratio(tg=tempogram, sr=sr)\nfig, ax = plt.subplots(nrows=2, sharex=True)\nlibrosa.display.specshow(tempogram, x_axis='time', y_axis='tempo',\n                         ax=ax[0])\nlibrosa.display.specshow(tgr, x_axis='time', ax=ax[1])\nax[0].label_outer()\nax[0].set(title=\"Tempogram\")\nax[1].set(title=\"Tempogram ratio\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:04:02.181657Z","iopub.execute_input":"2023-07-25T04:04:02.182057Z","iopub.status.idle":"2023-07-25T04:04:05.583047Z","shell.execute_reply.started":"2023-07-25T04:04:02.182026Z","shell.execute_reply":"2023-07-25T04:04:05.581897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Construct a standard onset function</p></div>\n\nConstruct a standard onset function,Median aggregation, and custom mel options","metadata":{}},{"cell_type":"code","source":"onset_env = librosa.onset.onset_strength(y=y, sr=sr)\n\n# Create a new subplot for the onset plot\nfig, ax = plt.subplots(nrows=3, sharex=True)\n\n# Plot the spectrogram on the first axis (ax[0])\nD = np.abs(librosa.stft(y))\ntimes = librosa.times_like(D)\nlibrosa.display.specshow(librosa.amplitude_to_db(D, ref=np.max),\n                         y_axis='log', x_axis='time', ax=ax[0])\nax[0].set(title='Power spectrogram')\nax[0].label_outer()\n\n# Plot the onset strength on the second axis (ax[1])\nax[1].plot(times, 2 + onset_env / onset_env.max(), alpha=0.8, label='Mean (mel)')\nax[1].set(title='Onset Strength/Mean')\nax[1].label_outer()\n\n# Plot the onset strength on the second axis (ax[2])\nax[2].plot(times, 1 + onset_env / onset_env.max(), alpha=0.8, label='Median (custom mel)')\nax[2].set(title='Onset Strength/Median')\nax[2].label_outer()\n\n# Add legends to the plots\nax[0].legend()\nax[1].legend()\nax[2].legend()\n\n# Adjust layout and display the plots\nplt.tight_layout()\nplt.show()\n\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-25T04:04:11.400270Z","iopub.execute_input":"2023-07-25T04:04:11.401361Z","iopub.status.idle":"2023-07-25T04:11:38.195944Z","shell.execute_reply.started":"2023-07-25T04:04:11.401323Z","shell.execute_reply":"2023-07-25T04:11:38.195057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#EE82EE;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Whisper</p></div>\n\n**Whisper** is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification.[Reference](https://github.com/openai/whisper)","metadata":{}},{"cell_type":"markdown","source":"## **Approach:**\n\n![image](https://raw.githubusercontent.com/openai/whisper/main/approach.png)\n\n[Image Source and Ref.](https://github.com/openai/whisper)\n\nA Transformer sequence-to-sequence model is trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection. These tasks are jointly represented as a sequence of tokens to be predicted by the decoder, allowing a single model to replace many stages of a traditional speech-processing pipeline. The multitask training format uses a set of special tokens that serve as task specifiers or classification targets.\n","metadata":{}},{"cell_type":"code","source":"import whisper\n\nmodel = whisper.load_model(\"base\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:11:51.796561Z","iopub.execute_input":"2023-07-25T04:11:51.796988Z","iopub.status.idle":"2023-07-25T04:11:57.062746Z","shell.execute_reply.started":"2023-07-25T04:11:51.796953Z","shell.execute_reply":"2023-07-25T04:11:57.061635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.device","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:12:00.066285Z","iopub.execute_input":"2023-07-25T04:12:00.067264Z","iopub.status.idle":"2023-07-25T04:12:00.078432Z","shell.execute_reply.started":"2023-07-25T04:12:00.067211Z","shell.execute_reply":"2023-07-25T04:12:00.077434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.display import Audio\nAudio(\"/kaggle/input/bengaliai-speech/train_mp3s/00023fd6ed9e.mp3\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:12:04.367280Z","iopub.execute_input":"2023-07-25T04:12:04.368329Z","iopub.status.idle":"2023-07-25T04:12:04.382424Z","shell.execute_reply.started":"2023-07-25T04:12:04.368292Z","shell.execute_reply":"2023-07-25T04:12:04.381459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Audio(\"/kaggle/input/bengaliai-speech/train_mp3s/000057f5f4a3.mp3\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:12:08.572278Z","iopub.execute_input":"2023-07-25T04:12:08.572684Z","iopub.status.idle":"2023-07-25T04:12:08.584821Z","shell.execute_reply.started":"2023-07-25T04:12:08.572652Z","shell.execute_reply":"2023-07-25T04:12:08.583862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\n#from numba import jit\n#!whisper \"/kaggle/input/bengaliai-speech/train_mp3s/000057f5f4a3.mp3\" --model medium.en","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:28:39.480266Z","iopub.execute_input":"2023-07-25T04:28:39.481358Z","iopub.status.idle":"2023-07-25T04:29:03.310970Z","shell.execute_reply.started":"2023-07-25T04:28:39.481319Z","shell.execute_reply":"2023-07-25T04:29:03.309562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#EE82EE;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Multilingual-ASR</p></div>\n\n**Multilingual Automatic Speech Recognition** (Multilingual ASR) is a technology that enables the recognition of speech in multiple languages using a single system. ASR systems are designed to convert spoken language into written text, and traditional ASR systems are typically trained and optimized for specific languages. However, with the increasing need for handling multilingual data and supporting multiple languages, researchers and developers have been working on creating ASR systems that can handle multiple languages effectively.\n\nThe main **objective of Multilingual ASR** is to build a single ASR system that can accurately transcribe speech in multiple languages without requiring separate models for each language. This is particularly useful in scenarios where speech data involves a mixture of languages or when resources and computational power are limited.","metadata":{}},{"cell_type":"code","source":"import io\nimport os\nimport numpy as np\n\ntry:\n    import tensorflow  \nexcept ImportError:\n    pass\n\nimport torch\nimport pandas as pd\nimport urllib\nimport tarfile\nimport whisper\nimport torchaudio\n\nfrom scipy.io import wavfile\nfrom tqdm.notebook import tqdm\n\n\npd.options.display.max_rows = 100\npd.options.display.max_colwidth = 1000\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-25T04:30:13.849002Z","iopub.execute_input":"2023-07-25T04:30:13.849483Z","iopub.status.idle":"2023-07-25T04:30:21.284871Z","shell.execute_reply.started":"2023-07-25T04:30:13.849444Z","shell.execute_reply":"2023-07-25T04:30:21.283839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Choose the Bengali dataset language for download. Keep in mind that the transcription and translation performance can significantly differ based on the selected language.","metadata":{}},{"cell_type":"code","source":"import ipywidgets as widgets\n\nlanguages = {\"af_za\": \"Afrikaans\", \"am_et\": \"Amharic\", \"ar_eg\": \"Arabic\", \"as_in\": \"Assamese\", \"az_az\": \"Azerbaijani\", \"be_by\": \"Belarusian\", \"bg_bg\": \"Bulgarian\", \"bn_in\": \"Bengali\", \"bs_ba\": \"Bosnian\", \"ca_es\": \"Catalan\", \"cmn_hans_cn\": \"Chinese\", \"cs_cz\": \"Czech\", \"cy_gb\": \"Welsh\", \"da_dk\": \"Danish\", \"de_de\": \"German\", \"el_gr\": \"Greek\", \"en_us\": \"English\", \"es_419\": \"Spanish\", \"et_ee\": \"Estonian\", \"fa_ir\": \"Persian\", \"fi_fi\": \"Finnish\", \"fil_ph\": \"Tagalog\", \"fr_fr\": \"French\", \"gl_es\": \"Galician\", \"gu_in\": \"Gujarati\", \"ha_ng\": \"Hausa\", \"he_il\": \"Hebrew\", \"hi_in\": \"Hindi\", \"hr_hr\": \"Croatian\", \"hu_hu\": \"Hungarian\", \"hy_am\": \"Armenian\", \"id_id\": \"Indonesian\", \"is_is\": \"Icelandic\", \"it_it\": \"Italian\", \"ja_jp\": \"Japanese\", \"jv_id\": \"Javanese\", \"ka_ge\": \"Georgian\", \"kk_kz\": \"Kazakh\", \"km_kh\": \"Khmer\", \"kn_in\": \"Kannada\", \"ko_kr\": \"Korean\", \"lb_lu\": \"Luxembourgish\", \"ln_cd\": \"Lingala\", \"lo_la\": \"Lao\", \"lt_lt\": \"Lithuanian\", \"lv_lv\": \"Latvian\", \"mi_nz\": \"Maori\", \"mk_mk\": \"Macedonian\", \"ml_in\": \"Malayalam\", \"mn_mn\": \"Mongolian\", \"mr_in\": \"Marathi\", \"ms_my\": \"Malay\", \"mt_mt\": \"Maltese\", \"my_mm\": \"Myanmar\", \"nb_no\": \"Norwegian\", \"ne_np\": \"Nepali\", \"nl_nl\": \"Dutch\", \"oc_fr\": \"Occitan\", \"pa_in\": \"Punjabi\", \"pl_pl\": \"Polish\", \"ps_af\": \"Pashto\", \"pt_br\": \"Portuguese\", \"ro_ro\": \"Romanian\", \"ru_ru\": \"Russian\", \"sd_in\": \"Sindhi\", \"sk_sk\": \"Slovak\", \"sl_si\": \"Slovenian\", \"sn_zw\": \"Shona\", \"so_so\": \"Somali\", \"sr_rs\": \"Serbian\", \"sv_se\": \"Swedish\", \"sw_ke\": \"Swahili\", \"ta_in\": \"Tamil\", \"te_in\": \"Telugu\", \"tg_tj\": \"Tajik\", \"th_th\": \"Thai\", \"tr_tr\": \"Turkish\", \"uk_ua\": \"Ukrainian\", \"ur_pk\": \"Urdu\", \"uz_uz\": \"Uzbek\", \"vi_vn\": \"Vietnamese\", \"yo_ng\": \"Yoruba\"}\nselection = widgets.Dropdown(\n    options=[(\"Select language\", None), (\"----------\", None)] + sorted([(f\"{v} ({k})\", k) for k, v in languages.items()]),\n    value=\"bn_in\",\n    description='Language:',\n    disabled=False,\n)\n\nselection","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:30:35.190429Z","iopub.execute_input":"2023-07-25T04:30:35.191448Z","iopub.status.idle":"2023-07-25T04:30:35.219232Z","shell.execute_reply.started":"2023-07-25T04:30:35.191406Z","shell.execute_reply":"2023-07-25T04:30:35.217271Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lang = selection.value\nlanguage = languages[lang]\n\nassert lang is not None, \"Please select a language\"\nprint(f\"Selected language: {language} ({lang})\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:30:40.217003Z","iopub.execute_input":"2023-07-25T04:30:40.217520Z","iopub.status.idle":"2023-07-25T04:30:40.225070Z","shell.execute_reply.started":"2023-07-25T04:30:40.217477Z","shell.execute_reply":"2023-07-25T04:30:40.223982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def download(url: str, target_path: str):\n    with urllib.request.urlopen(url) as source, open(target_path, \"wb\") as output:\n        with tqdm(total=int(source.info().get(\"Content-Length\")), ncols=80, unit='iB', unit_scale=True, unit_divisor=1024) as loop:\n            while True:\n                buffer = source.read(8192)\n                if not buffer:\n                    break\n\n                output.write(buffer)\n                loop.update(len(buffer))\n\n\nclass Bengali(torch.utils.data.Dataset):\n    \"\"\"\n    A simple class to wrap Fleurs and subsample a portion of the dataset as needed.\n    \"\"\"\n    def __init__(self, lang, split=\"test\", subsample_rate=1, device=DEVICE):\n        url = f\"https://storage.googleapis.com/xtreme_translations/FLEURS102/{lang}.tar.gz\"\n        tar_path = os.path.expanduser(f\"~/.cache/fleurs/{lang}.tgz\")\n        os.makedirs(os.path.dirname(tar_path), exist_ok=True)\n\n        if not os.path.exists(tar_path):\n            download(url, tar_path)\n\n        all_audio = {}\n        with tarfile.open(tar_path, \"r:gz\") as tar:\n            for member in tar.getmembers():\n                name = member.name\n                if name.endswith(f\"{split}.tsv\"):\n                    labels = pd.read_table(tar.extractfile(member), names=(\"id\", \"file_name\", \"raw_transcription\", \"transcription\", \"_\", \"num_samples\", \"gender\"))\n\n                if f\"/{split}/\" in name and name.endswith(\".wav\"):\n                    audio_bytes = tar.extractfile(member).read()\n                    all_audio[os.path.basename(name)] = wavfile.read(io.BytesIO(audio_bytes))[1]                    \n\n        self.labels = labels.to_dict(\"records\")[::subsample_rate]\n        self.all_audio = all_audio\n        self.device = device\n\n    def __len__(self):\n        return len(self.labels)\n\n    def __getitem__(self, item):\n        record = self.labels[item]\n        audio = torch.from_numpy(self.all_audio[record[\"file_name\"]].copy())\n        text = record[\"transcription\"]\n        \n        return (audio, text)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:30:44.442591Z","iopub.execute_input":"2023-07-25T04:30:44.442981Z","iopub.status.idle":"2023-07-25T04:30:44.458092Z","shell.execute_reply.started":"2023-07-25T04:30:44.442951Z","shell.execute_reply":"2023-07-25T04:30:44.456943Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = Bengali(lang, subsample_rate=10)  ","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:30:53.923744Z","iopub.execute_input":"2023-07-25T04:30:53.924209Z","iopub.status.idle":"2023-07-25T04:32:43.255032Z","shell.execute_reply.started":"2023-07-25T04:30:53.924167Z","shell.execute_reply":"2023-07-25T04:32:43.253914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Inference</p></div>\n\nInference will be performed on the dataset using a medium Whisper model. The process involves transcribing and translating utterances and may take a few minutes to finish.","metadata":{}},{"cell_type":"code","source":"model = whisper.load_model(\"medium\")\nprint(\n    f\"Model is {'multilingual' if model.is_multilingual else 'English-only'} \"\n    f\"and has {sum(np.prod(p.shape) for p in model.parameters()):,} parameters.\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:32:43.257485Z","iopub.execute_input":"2023-07-25T04:32:43.257900Z","iopub.status.idle":"2023-07-25T04:33:40.401268Z","shell.execute_reply.started":"2023-07-25T04:32:43.257854Z","shell.execute_reply":"2023-07-25T04:33:40.400050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"options = dict(language=language, beam_size=5, best_of=5)\ntranscribe_options = dict(task=\"transcribe\", **options)\ntranslate_options = dict(task=\"translate\", **options)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:34:39.458974Z","iopub.execute_input":"2023-07-25T04:34:39.459445Z","iopub.status.idle":"2023-07-25T04:34:39.465132Z","shell.execute_reply.started":"2023-07-25T04:34:39.459401Z","shell.execute_reply":"2023-07-25T04:34:39.464153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"references = []\ntranscriptions = []\ntranslations = []\n\nfor audio, text in tqdm(dataset):\n    transcription = model.transcribe(audio, **transcribe_options)[\"text\"]\n    translation = model.transcribe(audio, **translate_options)[\"text\"]\n    \n    transcriptions.append(transcription)\n    translations.append(translation)\n    references.append(text)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T04:34:45.757182Z","iopub.execute_input":"2023-07-25T04:34:45.757599Z","iopub.status.idle":"2023-07-25T05:57:16.823621Z","shell.execute_reply.started":"2023-07-25T04:34:45.757565Z","shell.execute_reply":"2023-07-25T05:57:16.822518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.DataFrame(dict(reference=references, transcription=transcriptions, translation=translations))\ndata","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:57:33.857243Z","iopub.execute_input":"2023-07-25T05:57:33.858149Z","iopub.status.idle":"2023-07-25T05:57:33.886459Z","shell.execute_reply.started":"2023-07-25T05:57:33.858107Z","shell.execute_reply":"2023-07-25T05:57:33.885422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Word-level timestamps</p></div>\n\nHere, we utilize attention weights to establish word-level timestamps, enabling a more detailed alignment. Employing heuristics and dynamic time warping (DTW), we determine the correlation between the audio and transcript to achieve accurate alignments.","metadata":{}},{"cell_type":"code","source":"%%capture\n! pip install dtw-python","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:58:09.006136Z","iopub.execute_input":"2023-07-25T05:58:09.006594Z","iopub.status.idle":"2023-07-25T05:58:21.872367Z","shell.execute_reply.started":"2023-07-25T05:58:09.006556Z","shell.execute_reply":"2023-07-25T05:58:21.870898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import string\nimport matplotlib.pyplot as plt\nimport matplotlib.font_manager as fm\nimport matplotlib.ticker as ticker\n\nfrom IPython.display import display, HTML\nfrom whisper.tokenizer import get_tokenizer\nfrom dtw import dtw\nfrom scipy.ndimage import median_filter\n\n%matplotlib inline\n%config InlineBackend.figure_format = \"retina\"","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:58:28.161784Z","iopub.execute_input":"2023-07-25T05:58:28.162196Z","iopub.status.idle":"2023-07-25T05:58:28.190988Z","shell.execute_reply.started":"2023-07-25T05:58:28.162163Z","shell.execute_reply":"2023-07-25T05:58:28.189967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AUDIO_SAMPLES_PER_TOKEN = whisper.audio.HOP_LENGTH * 2\nAUDIO_TIME_PER_TOKEN = AUDIO_SAMPLES_PER_TOKEN / whisper.audio.SAMPLE_RATE\n\nmedfilt_width = 7\nqk_scale = 1.0\n\ntokenizer = get_tokenizer(model.is_multilingual, language=languages[lang])\n","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:58:34.214672Z","iopub.execute_input":"2023-07-25T05:58:34.215073Z","iopub.status.idle":"2023-07-25T05:58:34.223176Z","shell.execute_reply.started":"2023-07-25T05:58:34.215041Z","shell.execute_reply":"2023-07-25T05:58:34.222095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if languages[lang] in {\"Japanese\",\"Bengali\",\"English\"}:\n    font = \"GoNotoCJKCore.ttf\"\nelse:\n    font = \"GoNotoCurrent.ttf\"\n\nfont_release = \"https://github.com/satbyy/go-noto-universal/releases/download/v5.2\"\nif not os.path.exists(font):\n    download(f\"{font_release}/{font}\", font)\n\nprop = fm.FontProperties(fname=font)\nprops = {'fontproperties': prop}","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:58:39.524283Z","iopub.execute_input":"2023-07-25T05:58:39.525338Z","iopub.status.idle":"2023-07-25T05:58:40.264784Z","shell.execute_reply.started":"2023-07-25T05:58:39.525297Z","shell.execute_reply":"2023-07-25T05:58:40.263654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def split_tokens_on_unicode(tokens: torch.Tensor):\n    words = []\n    word_tokens = []\n    current_tokens = []\n    \n    for token in tokens.tolist():\n        current_tokens.append(token)\n        decoded = tokenizer.decode_with_timestamps(current_tokens)\n        if \"\\ufffd\" not in decoded:\n            words.append(decoded)\n            word_tokens.append(current_tokens)\n            current_tokens = []\n    \n    return words, word_tokens","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:58:45.206847Z","iopub.execute_input":"2023-07-25T05:58:45.207240Z","iopub.status.idle":"2023-07-25T05:58:45.215742Z","shell.execute_reply.started":"2023-07-25T05:58:45.207206Z","shell.execute_reply":"2023-07-25T05:58:45.213573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def split_tokens_on_spaces(tokens: torch.Tensor):\n    subwords, subword_tokens_list = split_tokens_on_unicode(tokens)\n    words = []\n    word_tokens = []\n    \n    for subword, subword_tokens in zip(subwords, subword_tokens_list):\n        special = subword_tokens[0] >= tokenizer.eot\n        with_space = subword.startswith(\" \")\n        punctuation = subword.strip() in string.punctuation\n        if special or with_space or punctuation:\n            words.append(subword)\n            word_tokens.append(subword_tokens)\n        else:\n            words[-1] = words[-1] + subword\n            word_tokens[-1].extend(subword_tokens)\n    \n    return words, word_tokens","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:58:49.556333Z","iopub.execute_input":"2023-07-25T05:58:49.557353Z","iopub.status.idle":"2023-07-25T05:58:49.565830Z","shell.execute_reply.started":"2023-07-25T05:58:49.557316Z","shell.execute_reply":"2023-07-25T05:58:49.563636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if languages[lang] in {\" Japanese\",\" Bengali\",\"English\", \"Thai\"}:\n    \n    split_tokens = split_tokens_on_unicode\nelse:\n    split_tokens = split_tokens_on_spaces","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:58:54.179462Z","iopub.execute_input":"2023-07-25T05:58:54.179882Z","iopub.status.idle":"2023-07-25T05:58:54.185769Z","shell.execute_reply.started":"2023-07-25T05:58:54.179844Z","shell.execute_reply":"2023-07-25T05:58:54.184421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# install hooks on the cross attention layers to retrieve the attention weights\nQKs = [None] * model.dims.n_text_layer\n\nfor i, block in enumerate(model.decoder.blocks):\n    block.cross_attn.register_forward_hook(\n        lambda _, ins, outs, index=i: QKs.__setitem__(index, outs[-1])\n    )\n","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:58:59.984635Z","iopub.execute_input":"2023-07-25T05:58:59.985644Z","iopub.status.idle":"2023-07-25T05:58:59.992537Z","shell.execute_reply.started":"2023-07-25T05:58:59.985593Z","shell.execute_reply":"2023-07-25T05:58:59.991321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.options.display.max_rows = 100\npd.options.display.max_colwidth = 1000\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:59:04.993726Z","iopub.execute_input":"2023-07-25T05:59:04.994125Z","iopub.status.idle":"2023-07-25T05:59:05.000252Z","shell.execute_reply.started":"2023-07-25T05:59:04.994093Z","shell.execute_reply":"2023-07-25T05:59:04.999263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mel = whisper.log_mel_spectrogram(whisper.pad_or_trim(audio))  # Remove .cuda()\n","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:59:08.885003Z","iopub.execute_input":"2023-07-25T05:59:08.885412Z","iopub.status.idle":"2023-07-25T05:59:08.906543Z","shell.execute_reply.started":"2023-07-25T05:59:08.885358Z","shell.execute_reply":"2023-07-25T05:59:08.905522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.font_manager as font_manager\n\n# Replace 'your_font_path.ttf' with the path to the font file that supports Bengali and Telugu.\nplt.rcParams['font.family'] = 'your_font_name'\n","metadata":{"execution":{"iopub.status.busy":"2023-07-25T05:59:13.619395Z","iopub.execute_input":"2023-07-25T05:59:13.619858Z","iopub.status.idle":"2023-07-25T05:59:13.625190Z","shell.execute_reply.started":"2023-07-25T05:59:13.619827Z","shell.execute_reply":"2023-07-25T05:59:13.624217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# for the first 8 examples in the dataset\nfor (audio, label), transcription in zip(dataset, transcriptions[:8]):\n    print(transcription)\n  \n    duration = len(audio)\n    mel = whisper.log_mel_spectrogram(whisper.pad_or_trim(audio)).cuda()\n    tokens = torch.tensor(\n        [\n            *tokenizer.sot_sequence,\n            tokenizer.timestamp_begin,\n        ] + tokenizer.encode(transcription) + [\n            tokenizer.timestamp_begin + duration // AUDIO_SAMPLES_PER_TOKEN,\n            tokenizer.eot,\n        ]\n    ).cuda()\n    with torch.no_grad():\n        logits = model(mel.unsqueeze(0), tokens.unsqueeze(0))\n\n    weights = torch.cat(QKs)  # layers * heads * tokens * frames    \n    weights = weights[:, :, :, : duration // AUDIO_SAMPLES_PER_TOKEN].cpu()\n    weights = median_filter(weights, (1, 1, 1, medfilt_width))\n    weights = torch.tensor(weights * qk_scale).softmax(dim=-1)\n    \n    w = weights / weights.norm(dim=-2, keepdim=True)\n    matrix = w[-6:].mean(axis=(0, 1))\n\n    alignment = dtw(-matrix.double().numpy())\n\n    jumps = np.pad(np.diff(alignment.index1s), (1, 0), constant_values=1).astype(bool)\n    jump_times = alignment.index2s[jumps] * AUDIO_TIME_PER_TOKEN\n    words, word_tokens = split_tokens(tokens)\n\n    # display the normalized attention weights and the alignment\n    plt.figure(figsize=(8, 8))\n    plt.imshow(matrix, aspect=\"auto\")\n    plt.plot(alignment.index2s, alignment.index1s, color=\"red\")\n\n    xticks = np.arange(0, matrix.shape[1], 1 / AUDIO_TIME_PER_TOKEN)\n    xticklabels = (xticks * AUDIO_TIME_PER_TOKEN).round().astype(np.int32) \n    plt.xticks(xticks, xticklabels)\n    plt.xlabel(\"Time (s)\")\n    \n    # display tokens and words as tick labels\n    ylims = plt.gca().get_ylim()\n\n    ax = plt.gca()\n    ax.tick_params('both', length=0, width=0, which='minor', pad=6)\n\n    ax.yaxis.set_ticks_position(\"left\")\n    ax.yaxis.set_label_position(\"left\")\n    ax.invert_yaxis()\n    ax.set_ylim(ylims)\n\n    major_ticks = [-0.5]\n    minor_ticks = []\n    current_y = 0\n    \n    for word, word_token in zip(words, word_tokens):\n        minor_ticks.append(current_y + len(word_token) / 2 - 0.5)\n        current_y += len(word_token)\n        major_ticks.append(current_y - 0.5)\n        \n    ax.yaxis.set_minor_locator(ticker.FixedLocator(minor_ticks))\n    #ax.yaxis.set_minor_formatter(ticker.FixedFormatter(words))\n    ax.set_yticks(major_ticks)\n    ax.yaxis.set_major_formatter(ticker.NullFormatter())\n    \n    for label in ax.get_yminorticklabels():\n        label.set_fontproperties(prop)\n\n    plt.ylabel(\"Words\")\n    plt.show()\n\n    # display the word-level timestamps in a table\n    word_boundaries = np.pad(np.cumsum([len(t) for t in word_tokens[:-1]]), (1, 0))\n    begin_times = jump_times[word_boundaries[:-1]]\n    end_times = jump_times[word_boundaries[1:]]\n\n    data = [\n        dict(word=word, begin=begin, end=end)\n        for word, begin, end in zip(words[:-1], begin_times, end_times)\n        if not word.startswith(\"<|\") and word.strip() not in \".,!?、。\"\n    ]\n\n    display(pd.DataFrame(data))\n    display(HTML(\"<hr>\"))\n    \n# Finally, show the plot\nplt.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-25T05:59:17.323231Z","iopub.execute_input":"2023-07-25T05:59:17.324021Z","iopub.status.idle":"2023-07-25T05:59:57.180056Z","shell.execute_reply.started":"2023-07-25T05:59:17.323983Z","shell.execute_reply":"2023-07-25T05:59:57.178950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#FFFF00;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> BengaliSpeech dataset</p></div> \n\nWe will load the test-clean split of the BengaliSpeech corpus using torchaudio.","metadata":{}},{"cell_type":"code","source":"class BengaliSpeech(torch.utils.data.Dataset):\n    \"\"\"\n    A simple class to wrap BengaliSpeech and trim/pad the audio to 30 seconds.\n    It will drop the last few seconds of a very small portion of the utterances.\n    \"\"\"\n    def __init__(self, split=\"test-clean\", device=DEVICE):\n        self.dataset = torchaudio.datasets.LIBRISPEECH(\n            root=os.path.expanduser(\"~/.cache\"),\n            url=split,\n            download=True,\n        )\n        self.device = device\n\n    def __len__(self):\n        return len(self.dataset)\n\n    def __getitem__(self, item):\n        audio, sample_rate, text, _, _, _ = self.dataset[item]\n        assert sample_rate == 16000\n        audio = whisper.pad_or_trim(audio.flatten()).to(self.device)\n        mel = whisper.log_mel_spectrogram(audio)\n        \n        return (mel, text)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:00:36.264032Z","iopub.execute_input":"2023-07-25T06:00:36.265074Z","iopub.status.idle":"2023-07-25T06:00:36.273698Z","shell.execute_reply.started":"2023-07-25T06:00:36.265033Z","shell.execute_reply":"2023-07-25T06:00:36.272286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = BengaliSpeech(\"test-clean\")\nloader = torch.utils.data.DataLoader(dataset, batch_size=16)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:00:44.021726Z","iopub.execute_input":"2023-07-25T06:00:44.022224Z","iopub.status.idle":"2023-07-25T06:01:05.684843Z","shell.execute_reply.started":"2023-07-25T06:00:44.022185Z","shell.execute_reply":"2023-07-25T06:01:05.683795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = whisper.load_model(\"base.en\")\nprint(\n    f\"Model is {'multilingual' if model.is_multilingual else 'English-only'} \"\n    f\"and has {sum(np.prod(p.shape) for p in model.parameters()):,} parameters.\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:01:15.272355Z","iopub.execute_input":"2023-07-25T06:01:15.273064Z","iopub.status.idle":"2023-07-25T06:01:19.230567Z","shell.execute_reply.started":"2023-07-25T06:01:15.273030Z","shell.execute_reply":"2023-07-25T06:01:19.229444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict without timestamps for short-form transcription\noptions = whisper.DecodingOptions(language=\"en\", without_timestamps=True)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:01:24.241368Z","iopub.execute_input":"2023-07-25T06:01:24.242040Z","iopub.status.idle":"2023-07-25T06:01:24.246897Z","shell.execute_reply.started":"2023-07-25T06:01:24.242001Z","shell.execute_reply":"2023-07-25T06:01:24.245828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hypotheses = []\nreferences = []\n\nfor mels, texts in tqdm(loader):\n    results = model.decode(mels, options)\n    hypotheses.extend([result.text for result in results])\n    references.extend(texts)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:01:29.317611Z","iopub.execute_input":"2023-07-25T06:01:29.318316Z","iopub.status.idle":"2023-07-25T06:04:27.833405Z","shell.execute_reply.started":"2023-07-25T06:01:29.318277Z","shell.execute_reply":"2023-07-25T06:04:27.832268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.DataFrame(dict(hypothesis=hypotheses, reference=references))\ndata","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:04:32.873737Z","iopub.execute_input":"2023-07-25T06:04:32.874141Z","iopub.status.idle":"2023-07-25T06:04:32.891781Z","shell.execute_reply.started":"2023-07-25T06:04:32.874109Z","shell.execute_reply":"2023-07-25T06:04:32.890598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#EE82EE;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Evaluation Metrics</p></div>\n\nWord Error Rate (WER) is a metric used to evaluate the performance of Automatic Speech Recognition (ASR) systems or other systems that convert spoken language into written text, such as Optical Character Recognition (OCR) systems. WER is a common evaluation measure in the field of speech recognition.\n\nWER represents the percentage of words that are incorrectly recognized or transcribed by the ASR system compared to a reference or ground truth transcription of the same spoken input. It takes into account the substitutions, insertions, and deletions of words made by the ASR system.\n\nThe formula for calculating Word Error Rate is as follows:\n\nWER = (S + D + I) / N\n\nWhere:\n\n**S** = Number of word substitutions\n\n**D** = Number of word deletions\n\n**I** = Number of word insertions\n\n**N** = Total number of words in the reference (ground truth) transcription\n\nLet's break down the components:\n\n1. **Substitutions (S):** The number of words in the ASR output that are different from the reference transcription.\n\n2. **Deletions (D):** The number of words in the reference transcription that are missing from the ASR output.\n\n3. **Insertions (I):** The number of extra words in the ASR output that are not present in the reference transcription.\n\n4. **Total words (N):** The total number of words in the reference transcription.\n\nOnce the WER is calculated, it is expressed as a percentage to provide a more intuitive measure of the ASR system's accuracy. A lower WER indicates better performance, as it means the ASR system made fewer errors in transcribing the spoken language.\n\nFor example, if the reference transcription has 100 words and the ASR system outputs 8 substitutions, 5 deletions, and 3 insertions, the WER would be calculated as:\n\n**WER = (8 + 5 + 3) / 100 = 16 / 100 = 0.16 or 16%**\n\nThis means the ASR system's output has an error rate of 16%, i.e., 16% of the words in the ASR output are different from the reference transcription.\n","metadata":{}},{"cell_type":"markdown","source":"Next, we apply our English normalizer implementation to standardize the transcription and compute the Word Error Rate (WER).","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install jiwer","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:04:55.860037Z","iopub.execute_input":"2023-07-25T06:04:55.860462Z","iopub.status.idle":"2023-07-25T06:05:08.473628Z","shell.execute_reply.started":"2023-07-25T06:04:55.860428Z","shell.execute_reply":"2023-07-25T06:05:08.472094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import jiwer\nfrom whisper.normalizers import EnglishTextNormalizer\n\nnormalizer = EnglishTextNormalizer()","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:05:12.193727Z","iopub.execute_input":"2023-07-25T06:05:12.194163Z","iopub.status.idle":"2023-07-25T06:05:12.336755Z","shell.execute_reply.started":"2023-07-25T06:05:12.194129Z","shell.execute_reply":"2023-07-25T06:05:12.335703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data[\"hypothesis_clean\"] = [normalizer(text) for text in data[\"hypothesis\"]]\ndata[\"reference_clean\"] = [normalizer(text) for text in data[\"reference\"]]\ndata","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:05:17.601851Z","iopub.execute_input":"2023-07-25T06:05:17.602254Z","iopub.status.idle":"2023-07-25T06:05:20.088529Z","shell.execute_reply.started":"2023-07-25T06:05:17.602223Z","shell.execute_reply":"2023-07-25T06:05:20.087519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wer = jiwer.wer(list(data[\"reference_clean\"]), list(data[\"hypothesis_clean\"]))\n\nprint(f\"WER: {wer * 100:.2f} %\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:05:45.253668Z","iopub.execute_input":"2023-07-25T06:05:45.254087Z","iopub.status.idle":"2023-07-25T06:05:45.414231Z","shell.execute_reply.started":"2023-07-25T06:05:45.254053Z","shell.execute_reply":"2023-07-25T06:05:45.413193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:black;display:inline-block;border-radius:5px;background-color:#EE82EE;font-family:cursive;overflow:hidden\"><p style=\"padding:15px;color:black;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Submission</p></div>\n\n","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"/kaggle/input/bengaliai-speech/sample_submission.csv\")\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:05:49.313980Z","iopub.execute_input":"2023-07-25T06:05:49.314758Z","iopub.status.idle":"2023-07-25T06:05:49.332470Z","shell.execute_reply.started":"2023-07-25T06:05:49.314712Z","shell.execute_reply":"2023-07-25T06:05:49.331415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# The WER results stored in a list called `wer_results`\n# Assuming you have the WER results stored in a list called `wer_results`\nwer_results = [4.27, 5.32, 3.80]  # Replace with your actual WER results\n\n\nsub = pd.read_csv(\"/kaggle/input/bengaliai-speech/sample_submission.csv\")\n\nsubmission_df = pd.DataFrame(sub)\n\n# Adding the WER results to the DataFrame\nsubmission_df[\"4.27%\"] = wer_results\n\n# Save the DataFrame as a submission CSV\nsubmission_df.to_csv(\"submission_with_wer.csv\", index=False)\n\n# Optionally, you can also display the DataFrame\nsubmission_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2023-07-25T06:05:53.571493Z","iopub.execute_input":"2023-07-25T06:05:53.571930Z","iopub.status.idle":"2023-07-25T06:05:53.602086Z","shell.execute_reply.started":"2023-07-25T06:05:53.571900Z","shell.execute_reply":"2023-07-25T06:05:53.601035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reference:\n\n[Website:Librosa](https://librosa.org/)\n\n[OpenAI Whisper - MultiLingual AI Speech Recognition Live App Tutorial](https://www.youtube.com/watch?v=ywIyc8l1K1Q&t=237s)\n\n[Acknowledgement: jong wook kim](https://github.com/openai/whisper/blob/main/notebooks/LibriSpeech.ipynb)\n\n[Acknowledgement: Umar Farooqi](https://github.com/openai/whisper/blob/main/notebooks/Multilingual_ASR.ipynb)\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\"> 📌 \"Hey there! Your positive feedback and support for my notebook mean the world to me! It motivates me to create more valuable content. If you can spare a moment to give it an upvote, it would help others discover and benefit from it too. Together, let's foster a vibrant community of knowledge-sharing and empowerment. Thank you for considering it, and continued success on your learning journey!\"😃</div>","metadata":{}}]}