{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"> In this notebook, we'll see how to submit a notebook in this competition without internet connection. We'll be using the baseline pretrained NeMo Conformer model by Bengali.AI (The competition hosts).\n\nSome of code snippets were taken from [hengck23](https://www.kaggle.com/hengck23) [notebook](https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer). Kudos to him!\n\nLet's get started.","metadata":{"execution":{"iopub.status.busy":"2023-07-20T16:59:33.973984Z","iopub.execute_input":"2023-07-20T16:59:33.974796Z","iopub.status.idle":"2023-07-20T16:59:36.599802Z","shell.execute_reply.started":"2023-07-20T16:59:33.974763Z","shell.execute_reply":"2023-07-20T16:59:36.598481Z"}}},{"cell_type":"markdown","source":"# Install the packages","metadata":{}},{"cell_type":"markdown","source":"> Well since we don't have internet connection in this notebook, we can't just use !pip install to install pakcages. We'll install the packages from wheels. Please see [this notebook](https://www.kaggle.com/mbmmurad/how-to-submit-code-w-o-internet-creating-wheels/edit) to understand how to create wheels.","metadata":{}},{"cell_type":"markdown","source":"``` Install ffmpeg : ```","metadata":{}},{"cell_type":"code","source":"!pip install --no-index --no-deps /kaggle/input/bengaliai-conformer-packages/ffmpeg/*.whl","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"``` Install sox ```","metadata":{}},{"cell_type":"code","source":"!pip install --no-index --no-deps /kaggle/input/bengaliai-conformer-packages/sox/*.whl","metadata":{"execution":{"iopub.status.busy":"2023-07-20T16:59:36.603983Z","iopub.execute_input":"2023-07-20T16:59:36.604438Z","iopub.status.idle":"2023-07-20T16:59:45.229833Z","shell.execute_reply.started":"2023-07-20T16:59:36.604290Z","shell.execute_reply":"2023-07-20T16:59:45.228441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"``` Install NeMo ```","metadata":{}},{"cell_type":"code","source":"!cp -a /kaggle/input/nemo-pckgs/NeMo /kaggle/working/","metadata":{"execution":{"iopub.status.busy":"2023-07-20T16:59:45.232734Z","iopub.execute_input":"2023-07-20T16:59:45.233170Z","iopub.status.idle":"2023-07-20T17:00:26.568229Z","shell.execute_reply.started":"2023-07-20T16:59:45.233097Z","shell.execute_reply":"2023-07-20T17:00:26.566826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cd /kaggle/working/NeMo","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:00:26.571483Z","iopub.execute_input":"2023-07-20T17:00:26.571972Z","iopub.status.idle":"2023-07-20T17:00:26.579183Z","shell.execute_reply.started":"2023-07-20T17:00:26.571915Z","shell.execute_reply":"2023-07-20T17:00:26.578263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torchmetrics\ntorchmetrics.__version__","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:00:26.580636Z","iopub.execute_input":"2023-07-20T17:00:26.581200Z","iopub.status.idle":"2023-07-20T17:00:48.199397Z","shell.execute_reply.started":"2023-07-20T17:00:26.581167Z","shell.execute_reply":"2023-07-20T17:00:48.198281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture tst\n!pip install --no-index --no-deps /kaggle/input/nemo-pckgs/NeMo/*.whl","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:00:48.201346Z","iopub.execute_input":"2023-07-20T17:00:48.202099Z","iopub.status.idle":"2023-07-20T17:04:38.795460Z","shell.execute_reply.started":"2023-07-20T17:00:48.202059Z","shell.execute_reply":"2023-07-20T17:04:38.793838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sox import Transformer","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:05:31.505393Z","iopub.execute_input":"2023-07-20T17:05:31.505813Z","iopub.status.idle":"2023-07-20T17:05:31.510466Z","shell.execute_reply.started":"2023-07-20T17:05:31.505784Z","shell.execute_reply":"2023-07-20T17:05:31.509241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pwd","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:05:38.200237Z","iopub.execute_input":"2023-07-20T17:05:38.200949Z","iopub.status.idle":"2023-07-20T17:05:38.207467Z","shell.execute_reply.started":"2023-07-20T17:05:38.200914Z","shell.execute_reply":"2023-07-20T17:05:38.206555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"from pydub import AudioSegment\nimport nemo\nprint('nemo', nemo.__version__)\nimport librosa\nimport numpy as np\nimport pandas as pd\n#import jiwer\nimport nemo.collections.asr as nemo_asr","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:05:40.498602Z","iopub.execute_input":"2023-07-20T17:05:40.499309Z","iopub.status.idle":"2023-07-20T17:05:46.722115Z","shell.execute_reply.started":"2023-07-20T17:05:40.499258Z","shell.execute_reply":"2023-07-20T17:05:46.721047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cd /kaggle/working/","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:05:54.906773Z","iopub.execute_input":"2023-07-20T17:05:54.908014Z","iopub.status.idle":"2023-07-20T17:05:54.916077Z","shell.execute_reply.started":"2023-07-20T17:05:54.907976Z","shell.execute_reply":"2023-07-20T17:05:54.914983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> NeMo requires the model to be sampled at **16000 sample rate**. But test audios are sampled at **32k**. So we'll use AudioSegment to convert mp3s to wavs and then soundfile to export the audio to **16k sample rate** and **16 bit PCM** format. This bit was taken from this notebook","metadata":{}},{"cell_type":"code","source":"import soundfile as sf\n\nif 1:\n    mp3_dir = f'/kaggle/input/bengaliai-speech/train_mp3s'\n    id = '000005f3362c'\n    mp3_file  = f'{mp3_dir}/{id}.mp3'    \n    \n    sound = AudioSegment.from_mp3(mp3_file)\n    sound.export('temp.wav', format=\"wav\") \n    y, sr = librosa.load('temp.wav')\n    y = librosa.resample(y, orig_sr=sr, target_sr=16000)\n    y = librosa.to_mono(y)\n    sf.write('temp.wav', y, 16000, 'PCM_16')","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:06:26.058104Z","iopub.execute_input":"2023-07-20T17:06:26.058502Z","iopub.status.idle":"2023-07-20T17:06:26.288239Z","shell.execute_reply.started":"2023-07-20T17:06:26.058469Z","shell.execute_reply":"2023-07-20T17:06:26.286738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading the model","metadata":{}},{"cell_type":"markdown","source":"> Another problem is NeMo requires to download the **tokenizer from huggingface**. But we don't have internet connection, so we can't download the model from huggingface. So we'll have to load the model from local directory. \n\n> It requires the **base-bert-case** model. We've downloaded it and added it here as a dataset. We'll copy it to the /kaggle/working directory","metadata":{}},{"cell_type":"code","source":"!cp -a /kaggle/input/bert-base-uncased-bengaliai/bert-base-cased /kaggle/working/","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:39:09.808379Z","iopub.execute_input":"2023-07-20T17:39:09.808861Z","iopub.status.idle":"2023-07-20T17:39:37.727189Z","shell.execute_reply.started":"2023-07-20T17:39:09.808820Z","shell.execute_reply":"2023-07-20T17:39:37.725613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pwd","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:29:12.744815Z","iopub.execute_input":"2023-07-20T17:29:12.745285Z","iopub.status.idle":"2023-07-20T17:29:12.755643Z","shell.execute_reply.started":"2023-07-20T17:29:12.745244Z","shell.execute_reply":"2023-07-20T17:29:12.754334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Now there's a file that imports the tokenizer : /kaggle/working/NeMo/nemo/collections/common/tokenizers/huggingface/auto_tokenizer.py\n\n\nThis one tries to download the file from huggingface. We make a slight change here . This file uses this : \n\n```\nAUTOTOKENIZER.from_pretrained(\n                    pretrained_model_name_or_path=pretrained_model_name,\n                    vocab_file=vocab_file,\n                    merges_file=merges_file,\n                    use_fast=use_fast)\n```\n\nWe'll set the argument ```local_files_only=True``` to load the model locally. \n\n\n``` \nAUTOTOKENIZER.from_pretrained(\n                    pretrained_model_name_or_path=pretrained_model_name,\n                    vocab_file=vocab_file,\n                    merges_file=merges_file,\n                    use_fast=use_fast,\n                    local_files_only=True,\n                )\n```","metadata":{}},{"cell_type":"markdown","source":"I have changed this file and added it in the /kaggle/input/nemo-toolkit dataset. So we'll remove the previous \"auto_tokenizer.py\" and replace it with our changed version.","metadata":{}},{"cell_type":"code","source":"!rm /kaggle/working/NeMo/nemo/collections/common/tokenizers/huggingface/auto_tokenizer.py","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:29:26.632821Z","iopub.execute_input":"2023-07-20T17:29:26.634056Z","iopub.status.idle":"2023-07-20T17:29:27.662847Z","shell.execute_reply.started":"2023-07-20T17:29:26.634007Z","shell.execute_reply":"2023-07-20T17:29:27.661261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!cp /kaggle/input/nemo-toolkit/auto_tokenizer.py /kaggle/working/NeMo/nemo/collections/common/tokenizers/huggingface/","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:29:48.479555Z","iopub.execute_input":"2023-07-20T17:29:48.480583Z","iopub.status.idle":"2023-07-20T17:29:49.513564Z","shell.execute_reply.started":"2023-07-20T17:29:48.480540Z","shell.execute_reply":"2023-07-20T17:29:49.512132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls /kaggle/working/NeMo/nemo/collections/common/tokenizers/huggingface/","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:29:58.833495Z","iopub.execute_input":"2023-07-20T17:29:58.833920Z","iopub.status.idle":"2023-07-20T17:29:59.874499Z","shell.execute_reply.started":"2023-07-20T17:29:58.833884Z","shell.execute_reply":"2023-07-20T17:29:59.873221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#########################################################################\nprint(\"Initialize model loading\")\ncheckpoint_file = \"/kaggle/input/baseline-bengaliai-banglaconformer/BanglaConformer/Conformer-CTC-BPE.nemo\"\nasr_model = nemo_asr.models.EncDecCTCModelBPE.restore_from(restore_path=checkpoint_file)\n#asr_model.cuda()\nprint(\"Model Loaded\")","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:39:47.342998Z","iopub.execute_input":"2023-07-20T17:39:47.344412Z","iopub.status.idle":"2023-07-20T17:41:00.724638Z","shell.execute_reply.started":"2023-07-20T17:39:47.344359Z","shell.execute_reply":"2023-07-20T17:41:00.723507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Niceeee! We have loaded the model successfully. Now let's infer on the test data.","metadata":{}},{"cell_type":"code","source":"#import sys\n#sys.path.append('/kaggle/input/my-nemo/NeMo-main')\nmode='submit'\nif mode=='debug':\n    mp3_dir = f'/kaggle/input/bengaliai-speech/train_mp3s'\n    valid_df = pd.read_csv('/kaggle/input/bengaliai-speech/train.csv')\n    valid_df = valid_df[:25]\nif mode=='submit':\n    mp3_dir = f'/kaggle/input/bengaliai-speech/test_mp3s'\n    valid_df = pd.read_csv('/kaggle/input/bengaliai-speech/sample_submission.csv')\n    \nprint(len(valid_df))\nprint(valid_df['id'][:5].tolist())","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:41:09.485779Z","iopub.execute_input":"2023-07-20T17:41:09.486220Z","iopub.status.idle":"2023-07-20T17:41:09.513627Z","shell.execute_reply.started":"2023-07-20T17:41:09.486170Z","shell.execute_reply":"2023-07-20T17:41:09.512497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\nimport os\nfiles = os.listdir(\"/kaggle/input/bengaliai-speech/test_mp3s\")\nbase_path = \"/kaggle/input/bengaliai-speech/test_mp3s\"\nids = []\npreds = []\nfor file in files:\n    ids.append(file.split(\".\")[0])\n    mp3_file  = f'{base_path}/{file}'\n    \n    \n    sound = AudioSegment.from_mp3(mp3_file)\n    sound.export('temp.wav', format=\"wav\") \n    y, sr = librosa.load('temp.wav')\n    y = librosa.resample(y, orig_sr=sr, target_sr=16000)\n    y = librosa.to_mono(y)\n    sf.write('temp.wav', y, 16000, 'PCM_16')\n    \n    \n    p = asr_model.transcribe(paths2audio_files=['temp.wav', ], batch_size=1)[0] \n    p = p[:-2]+p[-1] ##?\n    preds.append(p)\n    \n    if mode=='debug':\n        print(p)\n        print(d.sentence)\n        print('')\n\nif mode=='debug':\n    score = jiwer.wer(valid_df['sentence'].to_list(), predict)\n    print('jiwer', score)\n    \nprint('')","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:44:58.284463Z","iopub.execute_input":"2023-07-20T17:44:58.284880Z","iopub.status.idle":"2023-07-20T17:44:59.298130Z","shell.execute_reply.started":"2023-07-20T17:44:58.284848Z","shell.execute_reply":"2023-07-20T17:44:59.296987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Removing the unnecessary folders","metadata":{}},{"cell_type":"code","source":"!rm -rf /kaggle/working/NeMo\n!rm -rf /kaggle/working/bert-base-cased\n!rm /kaggle/working/temp.wav","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating Submission.csv","metadata":{}},{"cell_type":"code","source":"df = pd.DataFrame({\"id\":ids,\"sentence\":preds})\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:45:39.295699Z","iopub.execute_input":"2023-07-20T17:45:39.296118Z","iopub.status.idle":"2023-07-20T17:45:39.309849Z","shell.execute_reply.started":"2023-07-20T17:45:39.296082Z","shell.execute_reply":"2023-07-20T17:45:39.308740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Normalize","metadata":{}},{"cell_type":"markdown","source":"Before submitting we'll normalize the sentences using [bnunicodenormalizer](https://github.com/mnansary/bnUnicodeNormalizer)\n> Install it from wheels","metadata":{}},{"cell_type":"code","source":"!cp -r /kaggle/input/python-packages2 ./","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!tar xvfz ./python-packages2/normalizer.tgz\n!pip install ./normalizer/bnunicodenormalizer-0.0.24.tar.gz -f ./ --no-index","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from bnunicodenormalizer import Normalizer \nbnorm = Normalizer()\ndef normalize(sen):\n    _words = [bnorm(word)['normalized']  for word in sen.split()]\n    return \" \".join([word for word in _words if word is not None])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.sentence = df.sentence.apply(lambda x:normalize(x))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv(\"submission.csv\",index=False)\nprint('SUBMIT DONE !!!!!!!')","metadata":{"execution":{"iopub.status.busy":"2023-07-20T17:45:51.331745Z","iopub.execute_input":"2023-07-20T17:45:51.332576Z","iopub.status.idle":"2023-07-20T17:45:51.340571Z","shell.execute_reply.started":"2023-07-20T17:45:51.332534Z","shell.execute_reply":"2023-07-20T17:45:51.339386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All set. Now let's hit save and run all and then submit the output to the competition. Hope you enjoyed this!","metadata":{}}]}