{
  "id": 425339,
  "title": "[lb0.68] how to setup baseline NeMO conformer-CTC baseline offline without internet",
  "url": "/competitions/bengaliai-speech/discussion/425339",
  "author_name": "hengck23",
  "post_date": "2023-07-18T09:14:11.121000",
  "votes": 19,
  "comment_count": 10,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F664cba6234f067a9e307b493b16e27ae%2FSelection_999(2736).png?generation=1689671608172327&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://huggingface.co/bengaliAI/BanglaConformer\" target=\"_blank\">https://huggingface.co/bengaliAI/BanglaConformer</a></p>\n<p>It takes me one full days to clear all the issues and finally it works.<br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb0-68-1hr-nemo-baseline-conformer-w-o-internet?scriptVersionId=137334272\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-68-1hr-nemo-baseline-conformer-w-o-internet?scriptVersionId=137334272</a></p>\n<p>here are some useful notes for those who are trying ….</p>",
  "messages": [
    {
      "id": 2349262,
      "postDate": "2023-07-18T09:14:11.120Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F664cba6234f067a9e307b493b16e27ae%2FSelection_999(2736).png?generation=1689671608172327&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://huggingface.co/bengaliAI/BanglaConformer\" target=\"_blank\">https://huggingface.co/bengaliAI/BanglaConformer</a></p>\n<p>It takes me one full days to clear all the issues and finally it works.<br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb0-68-1hr-nemo-baseline-conformer-w-o-internet?scriptVersionId=137334272\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-68-1hr-nemo-baseline-conformer-w-o-internet?scriptVersionId=137334272</a></p>\n<p>here are some useful notes for those who are trying ….</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F664cba6234f067a9e307b493b16e27ae%2FSelection_999(2736).png?generation=1689671608172327&alt=media)\n\nhttps://huggingface.co/bengaliAI/BanglaConformer\n\nIt takes me one full days to clear all the issues and finally it works.\nhttps://www.kaggle.com/code/hengck23/lb0-68-1hr-nemo-baseline-conformer-w-o-internet?scriptVersionId=137334272\n\nhere are some useful notes for those who are trying ....",
      "votes": 18
    },
    {
      "id": 2349724,
      "postDate": "2023-07-18T16:44:43.177Z",
      "content": "<p>seem that a quick baseline is pytorch-lightning + nemo + conformer-CTC<br>\n<a href=\"https://github.com/NVIDIA/NeMo/blob/main/tutorials/asr/ASR_with_NeMo.ipynb\" target=\"_blank\">https://github.com/NVIDIA/NeMo/blob/main/tutorials/asr/ASR_with_NeMo.ipynb</a></p>\n<p>just 3 line of code to start training<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcaf16ac18694ea8af0b44c4e5fcd4a3c%2FSelection_999(2758).png?generation=1689851826527020&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "seem that a quick baseline is pytorch-lightning + nemo + conformer-CTC\nhttps://github.com/NVIDIA/NeMo/blob/main/tutorials/asr/ASR_with_NeMo.ipynb\n\njust 3 line of code to start training\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcaf16ac18694ea8af0b44c4e5fcd4a3c%2FSelection_999(2758).png?generation=1689851826527020&alt=media)",
      "votes": 1
    },
    {
      "id": 2351320,
      "postDate": "2023-07-20T03:51:02.303Z",
      "content": "<p>useful tools:</p>\n<ol>\n<li>download folder from kaggle/working</li>\n</ol>\n<pre><code>import subprocess\n IPython.display import FileLink, display\n\ndef download_file(path, download_zip_file_name):\n    os.chdir()\n    zip_name = f\n    command = f\n    ()\n    result = subprocess.(command, =, =, =)\n     result.returncode != 0:\n        ()\n        (result.stderr)\n        return\n    display(FileLink(f))\n 0:   \n    download_file(, )\n</code></pre>",
      "rawMarkdown": "useful tools:\n\n1. download folder from kaggle/working\n\n```\n\nimport subprocess\nfrom IPython.display import FileLink, display\n\ndef download_file(path, download_zip_file_name):\n    os.chdir('/kaggle/working/')\n    zip_name = f\"/kaggle/working/{download_zip_file_name}\"\n    command = f\"zip {zip_name} {path} -r\"\n    print('zip ...')\n    result = subprocess.run(command, shell=True, capture_output=True, text=True)\n    if result.returncode != 0:\n        print(\"Unable to run zip command!\")\n        print(result.stderr)\n        return\n    display(FileLink(f'{download_zip_file_name}'))\nif 0:   \n    download_file('/kaggle/working/hf_cache', 'hf_cache.out.zip')\n\n```",
      "replies": [
        {
          "id": 2351322,
          "postDate": "2023-07-20T03:55:57.033Z",
          "content": "<ol>\n<li>Most of the wheel files are be download via \"pip download \" commands.<br>\nBut some wheels are not found on pip repo (e.g. for python 3.10) and they are compiled on the local kenrel notebook.<br>\nFor such, copy them from the temp cache folder: </li>\n</ol>\n<pre><code>#e.g.\n\nCollecting youtokentome\n  Downloading youtokentome-..tar.gz ( kB)\n     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ s eta ::\n  Preparing metadata (setup.py) ... done\nRequirement already satisfied: Click&gt;= in condapython3./site-packages ( youtokentome) (.)\nBuilding wheels  collected packages: youtokentome\n  Building wheel  youtokentome (setup.py) ... done\n  Created wheel  youtokentome: filename=youtokentome-.-cp310-cp310-linux_x86_64.whl = sha256=a4d14bc764849f0a2b73d19ec0dbd601303ef9632d5f62fead0c90c33875ae\n  Stored in directory: .cachewheelsd2ba45f43f30bed2fe413efa760bc726b8b660ed9c2900c\nSuccessfully built youtokentome\nInstalling collected packages: youtokentome\nSuccessfully installed youtokentome-.\n</code></pre>\n<p>then</p>\n<pre><code>!cp -r .cachewheelsd2ba45f43f30bed2fe413efa760bc726b8b660ed9c2900c working/\n\ndownload  : youtokentome-.-cp310-cp310-linux_x86_64.whl \n</code></pre>",
          "rawMarkdown": "2.  Most of the wheel files are be download via \"pip download <package name>\" commands.\nBut some wheels are not found on pip repo (e.g. for python 3.10) and they are compiled on the local kenrel notebook.\nFor such, copy them from the temp cache folder: \n\n```\n#e.g.\n\nCollecting youtokentome\n  Downloading youtokentome-1.0.6.tar.gz (86 kB)\n     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 86.7/86.7 kB 5.1 MB/s eta 0:00:00\n  Preparing metadata (setup.py) ... done\nRequirement already satisfied: Click>=7.0 in /opt/conda/lib/python3.10/site-packages (from youtokentome) (8.1.3)\nBuilding wheels for collected packages: youtokentome\n  Building wheel for youtokentome (setup.py) ... done\n  Created wheel for youtokentome: filename=youtokentome-1.0.6-cp310-cp310-linux_x86_64.whl size=177151 sha256=88a4d14bc764849f0a2b73d19ec0dbd601303ef9632d5f62fead0c90c33875ae\n  Stored in directory: /root/.cache/pip/wheels/df/85/f8/301d2ba45f43f30bed2fe413efa760bc726b8b660ed9c2900c\nSuccessfully built youtokentome\nInstalling collected packages: youtokentome\nSuccessfully installed youtokentome-1.0.6\n\n```\n\nthen\n\n```\n!cp -r /root/.cache/pip/wheels/df/85/f8/301d2ba45f43f30bed2fe413efa760bc726b8b660ed9c2900c /kaggle/working/\n\ndownload this file: youtokentome-1.0.6-cp310-cp310-linux_x86_64.whl \n```",
          "replies": [
            {
              "id": 2351326,
              "postDate": "2023-07-20T04:00:18.493Z",
              "content": "<p>apt get install of sox </p>\n<ul>\n<li>you can repalce sox transformer with the following code:</li>\n</ul>\n<pre><code>\n\n     sox import Transformer\n    tfm = Transformer()\n    tfm.rate(=16000)\n    tfm.channels(=1)\n\n    mp3_dir = f\n    id = \n    mp3_file  = f   \n    tfm.build(=mp3_file, =)\n</code></pre>\n<pre><code>#   sox\n\n librosa\n soundfile  sf\n pydub\n pydub  AudioSegment\n\n    sound = AudioSegment.from_mp3(mp3_file)\n    sound.export(, =\"wav\") \n    y, sr = librosa.()\n    y = librosa.resample(y, orig_sr=sr, target_sr=)\n    y = librosa.to_mono(y)\n    sf.(, y, , ) #PCM_16\n</code></pre>\n<p>using librosa, WER degrades by 0.01</p>",
              "rawMarkdown": "apt get install of sox \n- you can repalce sox transformer with the following code:\n\n```\n#using sox\n\n    from sox import Transformer\n    tfm = Transformer()\n    tfm.rate(samplerate=16000)\n    tfm.channels(n_channels=1)\n\n    mp3_dir = f'/kaggle/input/bengaliai-speech/train_mp3s'\n    id = '000005f3362c'\n    mp3_file  = f'{mp3_dir}/{id}.mp3'   \n    tfm.build(input_filepath=mp3_file, output_filepath='temp.wav')\n\n```\n\n```\n# not using sox\n\nimport librosa\nimport soundfile as sf\nimport pydub\nfrom pydub import AudioSegment\n\n    sound = AudioSegment.from_mp3(mp3_file)\n    sound.export('temp.wav', format=\"wav\") \n    y, sr = librosa.load('temp.wav')\n    y = librosa.resample(y, orig_sr=sr, target_sr=16000)\n    y = librosa.to_mono(y)\n    sf.write('temp.wav', y, 16000, 'PCM_16') #PCM_16\n```\n\nusing librosa, WER degrades by 0.01"
            },
            {
              "id": 2351329,
              "postDate": "2023-07-20T04:03:03.157Z",
              "content": "<p>by if you reall want to apt-get install sox without internet, copy the depency tree from an interneted connected notebook</p>\n<pre><code>  \n upgraded,  newly installed,  to remove and  not upgraded.\n to get  kB of archives.\n this operation,  kB of additional disk space will be used.\n: http://archive.ubuntu.com/ubuntu focal/main amd64 libmagic-mgc amd64 :.-\n: http://archive.ubuntu.com/ubuntu focal/main amd64 libmagic1 amd64 :.-\n: http://archive.ubuntu.com/ubuntu focal/universe amd64 libid3tag0 amd64 ..b-\n: http://archive.ubuntu.com/ubuntu focal/universe amd64 libmad0 amd64 ..b-ubuntu1\n: http://archive.ubuntu.com/ubuntu focal/universe amd64 libopencore-amrnb0 amd64 ..-\n: http://archive.ubuntu.com/ubuntu focal/universe amd64 libopencore-amrwb0 amd64 ..-\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox3 amd64 ..+git20190427-+deb11u2build\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-alsa amd64 ..+git20190427-+deb11u2build\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-base amd64 ..+git20190427-+deb11u2build\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-mp3 amd64 ..+git20190427-+deb11u2build\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 sox amd64 ..+git20190427-+deb11u2build\n</code></pre>\n<p>download these deb files and upload to your dataset. use !dpkg -i to install. installation order is important !!!</p>\n<pre><code>e.g.\n\n\n\n\n\n\n      \n \n</code></pre>",
              "rawMarkdown": "by if you reall want to apt-get install sox without internet, copy the depency tree from an interneted connected notebook\n\n```\n  sox\n0 upgraded, 11 newly installed, 0 to remove and 6 not upgraded.\nNeed to get 917 kB of archives.\nAfter this operation, 8035 kB of additional disk space will be used.\nGet:1 http://archive.ubuntu.com/ubuntu focal/main amd64 libmagic-mgc amd64 1:5.38-4 [218 kB]\nGet:2 http://archive.ubuntu.com/ubuntu focal/main amd64 libmagic1 amd64 1:5.38-4 [75.9 kB]\nGet:3 http://archive.ubuntu.com/ubuntu focal/universe amd64 libid3tag0 amd64 0.15.1b-14 [31.3 kB]\nGet:4 http://archive.ubuntu.com/ubuntu focal/universe amd64 libmad0 amd64 0.15.1b-10ubuntu1 [63.1 kB]\nGet:5 http://archive.ubuntu.com/ubuntu focal/universe amd64 libopencore-amrnb0 amd64 0.1.5-1 [94.8 kB]\nGet:6 http://archive.ubuntu.com/ubuntu focal/universe amd64 libopencore-amrwb0 amd64 0.1.5-1 [49.1 kB]\nGet:7 http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox3 amd64 14.4.2+git20190427-2+deb11u2build0.20.04.1 [225 kB]\nGet:8 http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-alsa amd64 14.4.2+git20190427-2+deb11u2build0.20.04.1 [10.5 kB]\nGet:9 http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-base amd64 14.4.2+git20190427-2+deb11u2build0.20.04.1 [31.4 kB]\nGet:10 http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-mp3 amd64 14.4.2+git20190427-2+deb11u2build0.20.04.1 [15.9 kB]\nGet:11 http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 sox amd64 14.4.2+git20190427-2+deb11u2build0.20.04.1 [102 kB]\n\n\n```\n\ndownload these deb files and upload to your dataset. use !dpkg -i to install. installation order is important !!!\n```\ne.g.\n# for d in [\n#     '/kaggle/input/sox-deb/libmagic1_13a5.38-4_amd64.deb',\n#     '/kaggle/input/sox-deb/libsox3_14.4.2git20190427-2deb11u2build0.20.04.1_amd64.deb',\n#     '/kaggle/input/sox-deb/libsox-fmt-base_14.4.2git20190427-2deb11u2build0.20.04.1_amd64.deb',\n#     '/kaggle/input/sox-deb/libsox-fmt-alsa_14.4.2git20190427-2deb11u2build0.20.04.1_amd64.deb',\n#     '/kaggle/input/sox-deb/sox_14.4.2git20190427-2deb11u2build0.20.04.1_amd64.deb',\n# ]:      \n#     !dpkg -i $d \n\n```"
            }
          ]
        }
      ]
    },
    {
      "id": 2351208,
      "postDate": "2023-07-19T23:07:36.790Z",
      "content": "<p>there is something strange about Nemo code:</p>\n<pre><code>anaconda3.nemopython3.nemoasrxins/mixins.py\nline \n\n            self.tokenizer = tokenizers.AutoTokenizer(\n                pretrained_model_name=,\n                vocab_file=self.vocab_path,\n</code></pre>\n<p>tokenizers.AutoTokenizer is actually huggingface auto-tokenizer.</p>\n<p>pretrained_model_name is defaulted. you cannot change to local path. Hence it always connect to internet to look for tokonizer, even if i have overide it in configure file.</p>",
      "rawMarkdown": "there is something strange about Nemo code:\n\n```\n/.../anaconda3.10/envs/nemo/lib/python3.8/site-packages/nemo/collections/asr/parts/mixins/mixins.py\nline 157\n\n            self.tokenizer = tokenizers.AutoTokenizer(\n                pretrained_model_name='bert-base-cased',\n                vocab_file=self.vocab_path,\n\n```\n\n tokenizers.AutoTokenizer is actually huggingface auto-tokenizer.\n\npretrained\\_model\\_name is defaulted. you cannot change to local path. Hence it always connect to internet to look for tokonizer, even if i have overide it in configure file.",
      "replies": [
        {
          "id": 2351288,
          "postDate": "2023-07-20T03:00:18.800Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2351310,
              "postDate": "2023-07-20T03:40:22.317Z",
              "content": "<p>i solved it by copying files from an internet-connected notebook to a non-connected notebook (via kaggle dataset).<br>\nThen I set hf cache director as follows:</p>\n<pre><code># to copy bert-base-cased\n!cp -r   \n \n.environ[]=\n.environ[]=\n.environ[]=\n</code></pre>\n<pre><code>\n\n\n\n</code></pre>",
              "rawMarkdown": "i solved it by copying files from an internet-connected notebook to a non-connected notebook (via kaggle dataset).\nThen I set hf cache director as follows:\n\n\n```\n# to copy bert-base-cased\n!cp -r '/kaggle/input/my-nemo/hf_cache' '/kaggle/working/' \nimport os\nos.environ['PYTORCH_TRANSFORMERS_CACHE']='/kaggle/working/hf_cache/hub'\nos.environ['HUGGINGFACE_HUB_CACHE']='/kaggle/working/hf_cache/hub'\nos.environ['HF_HOME']='/kaggle/working/hf_cache'\n\n```\n\n```\n# https://stackoverflow.com/questions/63312859/how-to-change-huggingface-transformers-default-cache-directory\n# https://github.com/huggingface/transformers/blob/6112b1c6442aaf7affd2b0676a1cd4eee30c45cf/src/transformers/utils/hub.py#L100\n# https://huggingface.co/docs/huggingface_hub/guides/manage-cache\n# https://huggingface.co/docs/huggingface_hub/package_reference/environment_variables\n\n```"
            },
            {
              "id": 2351379,
              "postDate": "2023-07-20T05:08:24.553Z",
              "content": "<p>the setting up of offline hf cache tool me the longest time to debug becuase i didn't realise that kaggle dataset disallow symbolic link, empty files and folders.</p>\n<p>so i have to recreate them.</p>\n<pre><code>\n\n\n\n\n\n!cp -r '/kaggle/input/my-nemo/hf_cache' '/kaggle/working/' \n!mkdir '/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots'\n!mkdir '/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/cc56f1d4bb1f5c76a55d11f846e0' \n!touch '/kaggle/working/hf_cache/hub/models--bert-base-cased/.no_exist/cc56f1d4bb1f5c76a55d11f846e0/added_tokens.json'\n!touch '/kaggle/working/hf_cache/hub/models--bert-base-cased/.no_exist/cc56f1d4bb1f5c76a55d11f846e0/special_tokens_map.json'\n\n\nfor source_file, symbolic_link in [\n    ['6be4f921afb3fdfd2ae79d', 'config.json'],\n    ['e3c6d456fbf01a9a6cd01a1be1a3ed22', 'tokenizer_config.json'],\n    ['2ea941cc79a6f3dcaef4f67dad62af04', 'vocab.txt'],\n]:\n    source_file = \\\n        f'/kaggle/working/hf_cache/hub/models--bert-base-cased/blobs/{source_file}'\n    symbolic_link = \\\n        f'/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/cc56f1d4bb1f5c76a55d11f846e0/{symbolic_link}'\n    !ln -s $source_file $symbolic_link\n\nos.chdir('/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/cc56f1d4bb1f5c76a55d11f846e0')\n!pwd\n!ls -all\n</code></pre>\n<p>if local cached bert-base-cased is set up correctly, you will see this</p>\n<pre><code>asr_model = nemo_asr.models.EncDecCTCModelBPE.restore_from(\n    =checkpoint_file,\n    #=cfg\n)\n</code></pre>\n<pre><code>Using eos_token,      yet.\nUsing bos_token,      yet.\n[NeMo I  :: mixins:] Tokenizer AutoTokenizer initialized   tokens\n</code></pre>",
              "rawMarkdown": "the setting up of offline hf cache tool me the longest time to debug becuase i didn't realise that kaggle dataset disallow symbolic link, empty files and folders.\n\nso i have to recreate them.\n\n```\n\n# to copy bert-base-cased\n# https://stackoverflow.com/questions/63312859/how-to-change-huggingface-transformers-default-cache-directory\n# https://github.com/huggingface/transformers/blob/6112b1c6442aaf7affd2b0676a1cd4eee30c45cf/src/transformers/utils/hub.py#L100\n# https://huggingface.co/docs/huggingface_hub/guides/manage-cache\n# https://huggingface.co/docs/huggingface_hub/package_reference/environment_variables\n\n!cp -r '/kaggle/input/my-nemo/hf_cache' '/kaggle/working/' \n!mkdir '/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots'\n!mkdir '/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/5532cc56f74641d4bb33641f5c76a55d11f846e0' \n!touch '/kaggle/working/hf_cache/hub/models--bert-base-cased/.no_exist/5532cc56f74641d4bb33641f5c76a55d11f846e0/added_tokens.json'\n!touch '/kaggle/working/hf_cache/hub/models--bert-base-cased/.no_exist/5532cc56f74641d4bb33641f5c76a55d11f846e0/special_tokens_map.json'\n\n# i think kaggle cannot upload empty folder and symbolic links ????\nfor source_file, symbolic_link in [\n    ['107460496b431545e4f921afb3fd5486fd2ae79d', 'config.json'],\n    ['e3c6d456fb2616f01a9a6cd01a1be1a36353ed22', 'tokenizer_config.json'],\n    ['2ea941cc79a6f3d7985ca6991ef4f67dad62af04', 'vocab.txt'],\n]:\n    source_file = \\\n        f'/kaggle/working/hf_cache/hub/models--bert-base-cased/blobs/{source_file}'\n    symbolic_link = \\\n        f'/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/5532cc56f74641d4bb33641f5c76a55d11f846e0/{symbolic_link}'\n    !ln -s $source_file $symbolic_link\n\nos.chdir('/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/5532cc56f74641d4bb33641f5c76a55d11f846e0')\n!pwd\n!ls -all\n\n\n```\n\nif local cached bert-base-cased is set up correctly, you will see this\n```\nasr_model = nemo_asr.models.EncDecCTCModelBPE.restore_from(\n\trestore_path=checkpoint_file,\n\t#override_config_path=cfg\n)\n\n\n```\n\n```\nUsing eos_token, but it is not set yet.\nUsing bos_token, but it is not set yet.\n[NeMo I 2023-07-20 05:07:47 mixins:170] Tokenizer AutoTokenizer initialized with 32000 tokens\n\n```",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2350041,
      "postDate": "2023-07-18T23:30:27.967Z",
      "content": "<p>see <a href=\"https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer?scriptVersionId=137210579\" target=\"_blank\">https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer?scriptVersionId=137210579</a></p>\n<p>local validation WER is 0.26</p>\n<pre><code>    mp3_dir = f\n\n    train_file = f\n    train_df = pd.read_csv(train_file)\n    fold_df = train_df[train_df[] == ].reset_index(=)\n    d = fold_df.iloc[0]\n\n    tfm = Transformer()\n    tfm.rate(=16000)\n    tfm.channels(=1)\n\n    checkpoint_file = \\\n        \n    asr_model = nemo_asr.models.EncDecCTCModelBPE.restore_from(=checkpoint_file)\n    (asr_model)\n\n     t, d  fold_df.iterrows():\n        mp3_file  = f\n        tfm.build(=mp3_file, =)\n        t = asr_model.transcribe(paths2audio_files=[, ], =1)[0]\n        (t)\n        (d.sentence)\n        ()\n</code></pre>\n<p>results</p>\n<pre><code>output_file: .wav already   will be overwritten  build\nTranscribing: %|██████████| / [:&lt;:,  it/s]\noutput_file: .wav already   will be overwritten  build\nতিনি এবং তার মা তাদের পৈতৃক বাড়িতে থেকে প্রতিবেশীদের দ্বারা অনেক তিরস্কার সহ্য করেন ।\nতিনি এবং তাঁর মা তাদের পৈতৃক বাড়িতে থেকে প্রতিবেশীদের দ্বারা অনেক তিরস্কার সহ্য করেন।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nভিত্তিবাস রামায়ণ বহির্ভূত অনেক গল্প এই অনুবাদ গ্রহণ করেছিলেন ।\nকৃত্তিবাস রামায়ণ-বহির্ভূত অনেক গল্প এই অনুবাদে গ্রহণ করেছিলেন।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nতিনি তার সুশৃঙ্খল সামরিক বাহিনী এবং সুগঠিত শাসন কাঠামোর মাধ্যমে একটি দক্ষ শাসনব্যবস্থা প্রতিষ্ঠিত করেন ।\nতিনি তার সুশৃঙ্খল সামরিক বাহিনী এবং সুগঠিত শাসন কাঠামোর মাধ্যমে একটি দক্ষ শাসন ব্যবস্থা প্রতিষ্ঠিত করেন।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nতিনি বিজয় সাম্রাজ্যের বিরুদ্ধে এবং বিজ মুসলিম প্রতিবেশীদের বিরুদ্ধে যুদ্ধ করেছিলেন ।\nতিনি বিজয়নগর সাম্রাজ্যের বিরুদ্ধে এবং বিজাপুরের মুসলিম প্রতিবেশীদের বিরুদ্ধেও যুদ্ধ করেছিলেন।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nএটি মূলত একটি মরুময় অঞ্চল ।\nএটি মূলত একটি মরুময় অঞ্চল।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nসড়কটি বিহার পশ্চিমবঙ্গ সীমান্ত অতিক্রম করে পশ্চিমবঙ্গ রাজ্যে প্রবেশ করে উত্তর দিনাজপুর জেলা হয়ে ।\nসড়কটি বিহার-পশ্চিমবঙ্গ সীমান্ত অতিক্রম করে পশ্চিমবঙ্গ রাজ্যে প্রবেশ করে উত্তর দিনাজপুর জেলা হয়ে।\n</code></pre>",
      "rawMarkdown": "see https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer?scriptVersionId=137210579\n\nlocal validation WER is 0.26\n\n```\n\tmp3_dir = f'{root_dir}/data/bengaliai-speech/train_mp3s'\n\n\ttrain_file = f'{root_dir}/data/bengaliai-speech/train.csv'\n\ttrain_df = pd.read_csv(train_file)\n\tfold_df = train_df[train_df['split'] == 'valid'].reset_index(drop=True)\n\td = fold_df.iloc[0]\n\n\ttfm = Transformer()\n\ttfm.rate(samplerate=16000)\n\ttfm.channels(n_channels=1)\n\n\tcheckpoint_file = \\\n\t\t'/home/titanx/hengck/share1/kaggle/2022/bengali-asr/data/other/bangla-conformer/Conformer-CTC-BPE.nemo'\n\tasr_model = nemo_asr.models.EncDecCTCModelBPE.restore_from(restore_path=checkpoint_file)\n\tprint(asr_model)\n\n\tfor t, d in fold_df.iterrows():\n\t\tmp3_file  = f'{mp3_dir}/{d[\"id\"]}.mp3'\n\t\ttfm.build(input_filepath=mp3_file, output_filepath='temp.wav')\n\t\tt = asr_model.transcribe(paths2audio_files=['temp.wav', ], batch_size=1)[0]\n\t\tprint(t)\n\t\tprint(d.sentence)\n\t\tprint('')\n\n```\n\nresults\n\n```\noutput_file: temp.wav already exists and will be overwritten on build\nTranscribing: 100%|██████████| 1/1 [00:00<00:00,  1.12it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nতিনি এবং তার মা তাদের পৈতৃক বাড়িতে থেকে প্রতিবেশীদের দ্বারা অনেক তিরস্কার সহ্য করেন ।\nতিনি এবং তাঁর মা তাদের পৈতৃক বাড়িতে থেকে প্রতিবেশীদের দ্বারা অনেক তিরস্কার সহ্য করেন।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 16.44it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nভিত্তিবাস রামায়ণ বহির্ভূত অনেক গল্প এই অনুবাদ গ্রহণ করেছিলেন ।\nকৃত্তিবাস রামায়ণ-বহির্ভূত অনেক গল্প এই অনুবাদে গ্রহণ করেছিলেন।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 15.31it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nতিনি তার সুশৃঙ্খল সামরিক বাহিনী এবং সুগঠিত শাসন কাঠামোর মাধ্যমে একটি দক্ষ শাসনব্যবস্থা প্রতিষ্ঠিত করেন ।\nতিনি তার সুশৃঙ্খল সামরিক বাহিনী এবং সুগঠিত শাসন কাঠামোর মাধ্যমে একটি দক্ষ শাসন ব্যবস্থা প্রতিষ্ঠিত করেন।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 15.50it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nতিনি বিজয় সাম্রাজ্যের বিরুদ্ধে এবং বিজ মুসলিম প্রতিবেশীদের বিরুদ্ধে যুদ্ধ করেছিলেন ।\nতিনি বিজয়নগর সাম্রাজ্যের বিরুদ্ধে এবং বিজাপুরের মুসলিম প্রতিবেশীদের বিরুদ্ধেও যুদ্ধ করেছিলেন।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 18.29it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nএটি মূলত একটি মরুময় অঞ্চল ।\nএটি মূলত একটি মরুময় অঞ্চল।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 15.23it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nসড়কটি বিহার পশ্চিমবঙ্গ সীমান্ত অতিক্রম করে পশ্চিমবঙ্গ রাজ্যে প্রবেশ করে উত্তর দিনাজপুর জেলা হয়ে ।\nসড়কটি বিহার-পশ্চিমবঙ্গ সীমান্ত অতিক্রম করে পশ্চিমবঙ্গ রাজ্যে প্রবেশ করে উত্তর দিনাজপুর জেলা হয়ে।\n\n```"
    }
  ],
  "comments": [
    {
      "id": 2349724,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-18T16:44:43.177000",
      "content": "<p>seem that a quick baseline is pytorch-lightning + nemo + conformer-CTC<br>\n<a href=\"https://github.com/NVIDIA/NeMo/blob/main/tutorials/asr/ASR_with_NeMo.ipynb\" target=\"_blank\">https://github.com/NVIDIA/NeMo/blob/main/tutorials/asr/ASR_with_NeMo.ipynb</a></p>\n<p>just 3 line of code to start training<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcaf16ac18694ea8af0b44c4e5fcd4a3c%2FSelection_999(2758).png?generation=1689851826527020&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2351320,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-20T03:51:02.303000",
      "content": "<p>useful tools:</p>\n<ol>\n<li>download folder from kaggle/working</li>\n</ol>\n<pre><code>import subprocess\n IPython.display import FileLink, display\n\ndef download_file(path, download_zip_file_name):\n    os.chdir()\n    zip_name = f\n    command = f\n    ()\n    result = subprocess.(command, =, =, =)\n     result.returncode != 0:\n        ()\n        (result.stderr)\n        return\n    display(FileLink(f))\n 0:   \n    download_file(, )\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2351322,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-07-20T03:55:57.033000",
          "content": "<ol>\n<li>Most of the wheel files are be download via \"pip download \" commands.<br>\nBut some wheels are not found on pip repo (e.g. for python 3.10) and they are compiled on the local kenrel notebook.<br>\nFor such, copy them from the temp cache folder: </li>\n</ol>\n<pre><code>#e.g.\n\nCollecting youtokentome\n  Downloading youtokentome-..tar.gz ( kB)\n     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ s eta ::\n  Preparing metadata (setup.py) ... done\nRequirement already satisfied: Click&gt;= in condapython3./site-packages ( youtokentome) (.)\nBuilding wheels  collected packages: youtokentome\n  Building wheel  youtokentome (setup.py) ... done\n  Created wheel  youtokentome: filename=youtokentome-.-cp310-cp310-linux_x86_64.whl = sha256=a4d14bc764849f0a2b73d19ec0dbd601303ef9632d5f62fead0c90c33875ae\n  Stored in directory: .cachewheelsd2ba45f43f30bed2fe413efa760bc726b8b660ed9c2900c\nSuccessfully built youtokentome\nInstalling collected packages: youtokentome\nSuccessfully installed youtokentome-.\n</code></pre>\n<p>then</p>\n<pre><code>!cp -r .cachewheelsd2ba45f43f30bed2fe413efa760bc726b8b660ed9c2900c working/\n\ndownload  : youtokentome-.-cp310-cp310-linux_x86_64.whl \n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2351326,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-07-20T04:00:18.493000",
              "content": "<p>apt get install of sox </p>\n<ul>\n<li>you can repalce sox transformer with the following code:</li>\n</ul>\n<pre><code>\n\n     sox import Transformer\n    tfm = Transformer()\n    tfm.rate(=16000)\n    tfm.channels(=1)\n\n    mp3_dir = f\n    id = \n    mp3_file  = f   \n    tfm.build(=mp3_file, =)\n</code></pre>\n<pre><code>#   sox\n\n librosa\n soundfile  sf\n pydub\n pydub  AudioSegment\n\n    sound = AudioSegment.from_mp3(mp3_file)\n    sound.export(, =\"wav\") \n    y, sr = librosa.()\n    y = librosa.resample(y, orig_sr=sr, target_sr=)\n    y = librosa.to_mono(y)\n    sf.(, y, , ) #PCM_16\n</code></pre>\n<p>using librosa, WER degrades by 0.01</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2351329,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-07-20T04:03:03.157000",
              "content": "<p>by if you reall want to apt-get install sox without internet, copy the depency tree from an interneted connected notebook</p>\n<pre><code>  \n upgraded,  newly installed,  to remove and  not upgraded.\n to get  kB of archives.\n this operation,  kB of additional disk space will be used.\n: http://archive.ubuntu.com/ubuntu focal/main amd64 libmagic-mgc amd64 :.-\n: http://archive.ubuntu.com/ubuntu focal/main amd64 libmagic1 amd64 :.-\n: http://archive.ubuntu.com/ubuntu focal/universe amd64 libid3tag0 amd64 ..b-\n: http://archive.ubuntu.com/ubuntu focal/universe amd64 libmad0 amd64 ..b-ubuntu1\n: http://archive.ubuntu.com/ubuntu focal/universe amd64 libopencore-amrnb0 amd64 ..-\n: http://archive.ubuntu.com/ubuntu focal/universe amd64 libopencore-amrwb0 amd64 ..-\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox3 amd64 ..+git20190427-+deb11u2build\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-alsa amd64 ..+git20190427-+deb11u2build\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-base amd64 ..+git20190427-+deb11u2build\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 libsox-fmt-mp3 amd64 ..+git20190427-+deb11u2build\n: http://archive.ubuntu.com/ubuntu focal-updates/universe amd64 sox amd64 ..+git20190427-+deb11u2build\n</code></pre>\n<p>download these deb files and upload to your dataset. use !dpkg -i to install. installation order is important !!!</p>\n<pre><code>e.g.\n\n\n\n\n\n\n      \n \n</code></pre>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2351208,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-19T23:07:36.790000",
      "content": "<p>there is something strange about Nemo code:</p>\n<pre><code>anaconda3.nemopython3.nemoasrxins/mixins.py\nline \n\n            self.tokenizer = tokenizers.AutoTokenizer(\n                pretrained_model_name=,\n                vocab_file=self.vocab_path,\n</code></pre>\n<p>tokenizers.AutoTokenizer is actually huggingface auto-tokenizer.</p>\n<p>pretrained_model_name is defaulted. you cannot change to local path. Hence it always connect to internet to look for tokonizer, even if i have overide it in configure file.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2351288,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-07-20T03:00:18.800000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2351310,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-07-20T03:40:22.317000",
              "content": "<p>i solved it by copying files from an internet-connected notebook to a non-connected notebook (via kaggle dataset).<br>\nThen I set hf cache director as follows:</p>\n<pre><code># to copy bert-base-cased\n!cp -r   \n \n.environ[]=\n.environ[]=\n.environ[]=\n</code></pre>\n<pre><code>\n\n\n\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2351379,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-07-20T05:08:24.553000",
              "content": "<p>the setting up of offline hf cache tool me the longest time to debug becuase i didn't realise that kaggle dataset disallow symbolic link, empty files and folders.</p>\n<p>so i have to recreate them.</p>\n<pre><code>\n\n\n\n\n\n!cp -r '/kaggle/input/my-nemo/hf_cache' '/kaggle/working/' \n!mkdir '/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots'\n!mkdir '/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/cc56f1d4bb1f5c76a55d11f846e0' \n!touch '/kaggle/working/hf_cache/hub/models--bert-base-cased/.no_exist/cc56f1d4bb1f5c76a55d11f846e0/added_tokens.json'\n!touch '/kaggle/working/hf_cache/hub/models--bert-base-cased/.no_exist/cc56f1d4bb1f5c76a55d11f846e0/special_tokens_map.json'\n\n\nfor source_file, symbolic_link in [\n    ['6be4f921afb3fdfd2ae79d', 'config.json'],\n    ['e3c6d456fbf01a9a6cd01a1be1a3ed22', 'tokenizer_config.json'],\n    ['2ea941cc79a6f3dcaef4f67dad62af04', 'vocab.txt'],\n]:\n    source_file = \\\n        f'/kaggle/working/hf_cache/hub/models--bert-base-cased/blobs/{source_file}'\n    symbolic_link = \\\n        f'/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/cc56f1d4bb1f5c76a55d11f846e0/{symbolic_link}'\n    !ln -s $source_file $symbolic_link\n\nos.chdir('/kaggle/working/hf_cache/hub/models--bert-base-cased/snapshots/cc56f1d4bb1f5c76a55d11f846e0')\n!pwd\n!ls -all\n</code></pre>\n<p>if local cached bert-base-cased is set up correctly, you will see this</p>\n<pre><code>asr_model = nemo_asr.models.EncDecCTCModelBPE.restore_from(\n    =checkpoint_file,\n    #=cfg\n)\n</code></pre>\n<pre><code>Using eos_token,      yet.\nUsing bos_token,      yet.\n[NeMo I  :: mixins:] Tokenizer AutoTokenizer initialized   tokens\n</code></pre>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2350041,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-18T23:30:27.967000",
      "content": "<p>see <a href=\"https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer?scriptVersionId=137210579\" target=\"_blank\">https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer?scriptVersionId=137210579</a></p>\n<p>local validation WER is 0.26</p>\n<pre><code>    mp3_dir = f\n\n    train_file = f\n    train_df = pd.read_csv(train_file)\n    fold_df = train_df[train_df[] == ].reset_index(=)\n    d = fold_df.iloc[0]\n\n    tfm = Transformer()\n    tfm.rate(=16000)\n    tfm.channels(=1)\n\n    checkpoint_file = \\\n        \n    asr_model = nemo_asr.models.EncDecCTCModelBPE.restore_from(=checkpoint_file)\n    (asr_model)\n\n     t, d  fold_df.iterrows():\n        mp3_file  = f\n        tfm.build(=mp3_file, =)\n        t = asr_model.transcribe(paths2audio_files=[, ], =1)[0]\n        (t)\n        (d.sentence)\n        ()\n</code></pre>\n<p>results</p>\n<pre><code>output_file: .wav already   will be overwritten  build\nTranscribing: %|██████████| / [:&lt;:,  it/s]\noutput_file: .wav already   will be overwritten  build\nতিনি এবং তার মা তাদের পৈতৃক বাড়িতে থেকে প্রতিবেশীদের দ্বারা অনেক তিরস্কার সহ্য করেন ।\nতিনি এবং তাঁর মা তাদের পৈতৃক বাড়িতে থেকে প্রতিবেশীদের দ্বারা অনেক তিরস্কার সহ্য করেন।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nভিত্তিবাস রামায়ণ বহির্ভূত অনেক গল্প এই অনুবাদ গ্রহণ করেছিলেন ।\nকৃত্তিবাস রামায়ণ-বহির্ভূত অনেক গল্প এই অনুবাদে গ্রহণ করেছিলেন।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nতিনি তার সুশৃঙ্খল সামরিক বাহিনী এবং সুগঠিত শাসন কাঠামোর মাধ্যমে একটি দক্ষ শাসনব্যবস্থা প্রতিষ্ঠিত করেন ।\nতিনি তার সুশৃঙ্খল সামরিক বাহিনী এবং সুগঠিত শাসন কাঠামোর মাধ্যমে একটি দক্ষ শাসন ব্যবস্থা প্রতিষ্ঠিত করেন।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nতিনি বিজয় সাম্রাজ্যের বিরুদ্ধে এবং বিজ মুসলিম প্রতিবেশীদের বিরুদ্ধে যুদ্ধ করেছিলেন ।\nতিনি বিজয়নগর সাম্রাজ্যের বিরুদ্ধে এবং বিজাপুরের মুসলিম প্রতিবেশীদের বিরুদ্ধেও যুদ্ধ করেছিলেন।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nএটি মূলত একটি মরুময় অঞ্চল ।\nএটি মূলত একটি মরুময় অঞ্চল।\n\nTranscribing: %|██████████| / [:&lt;:, it/s]\noutput_file: .wav already   will be overwritten  build\nসড়কটি বিহার পশ্চিমবঙ্গ সীমান্ত অতিক্রম করে পশ্চিমবঙ্গ রাজ্যে প্রবেশ করে উত্তর দিনাজপুর জেলা হয়ে ।\nসড়কটি বিহার-পশ্চিমবঙ্গ সীমান্ত অতিক্রম করে পশ্চিমবঙ্গ রাজ্যে প্রবেশ করে উত্তর দিনাজপুর জেলা হয়ে।\n</code></pre>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2349262": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F664cba6234f067a9e307b493b16e27ae%2FSelection_999(2736).png?generation=1689671608172327&alt=media)\n\nhttps://huggingface.co/bengaliAI/BanglaConformer\n\nIt takes me one full days to clear all the issues and finally it works.\nhttps://www.kaggle.com/code/hengck23/lb0-68-1hr-nemo-baseline-conformer-w-o-internet?scriptVersionId=137334272\n\nhere are some useful notes for those who are trying ....",
    "2349724": "seem that a quick baseline is pytorch-lightning + nemo + conformer-CTC\nhttps://github.com/NVIDIA/NeMo/blob/main/tutorials/asr/ASR_with_NeMo.ipynb\n\njust 3 line of code to start training\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcaf16ac18694ea8af0b44c4e5fcd4a3c%2FSelection_999(2758).png?generation=1689851826527020&alt=media)",
    "2351320": "useful tools:\n\n1. download folder from kaggle/working\n\n```\n\nimport subprocess\nfrom IPython.display import FileLink, display\n\ndef download_file(path, download_zip_file_name):\n    os.chdir('/kaggle/working/')\n    zip_name = f\"/kaggle/working/{download_zip_file_name}\"\n    command = f\"zip {zip_name} {path} -r\"\n    print('zip ...')\n    result = subprocess.run(command, shell=True, capture_output=True, text=True)\n    if result.returncode != 0:\n        print(\"Unable to run zip command!\")\n        print(result.stderr)\n        return\n    display(FileLink(f'{download_zip_file_name}'))\nif 0:   \n    download_file('/kaggle/working/hf_cache', 'hf_cache.out.zip')\n\n```",
    "2351208": "there is something strange about Nemo code:\n\n```\n/.../anaconda3.10/envs/nemo/lib/python3.8/site-packages/nemo/collections/asr/parts/mixins/mixins.py\nline 157\n\n            self.tokenizer = tokenizers.AutoTokenizer(\n                pretrained_model_name='bert-base-cased',\n                vocab_file=self.vocab_path,\n\n```\n\n tokenizers.AutoTokenizer is actually huggingface auto-tokenizer.\n\npretrained\\_model\\_name is defaulted. you cannot change to local path. Hence it always connect to internet to look for tokonizer, even if i have overide it in configure file.",
    "2350041": "see https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer?scriptVersionId=137210579\n\nlocal validation WER is 0.26\n\n```\n\tmp3_dir = f'{root_dir}/data/bengaliai-speech/train_mp3s'\n\n\ttrain_file = f'{root_dir}/data/bengaliai-speech/train.csv'\n\ttrain_df = pd.read_csv(train_file)\n\tfold_df = train_df[train_df['split'] == 'valid'].reset_index(drop=True)\n\td = fold_df.iloc[0]\n\n\ttfm = Transformer()\n\ttfm.rate(samplerate=16000)\n\ttfm.channels(n_channels=1)\n\n\tcheckpoint_file = \\\n\t\t'/home/titanx/hengck/share1/kaggle/2022/bengali-asr/data/other/bangla-conformer/Conformer-CTC-BPE.nemo'\n\tasr_model = nemo_asr.models.EncDecCTCModelBPE.restore_from(restore_path=checkpoint_file)\n\tprint(asr_model)\n\n\tfor t, d in fold_df.iterrows():\n\t\tmp3_file  = f'{mp3_dir}/{d[\"id\"]}.mp3'\n\t\ttfm.build(input_filepath=mp3_file, output_filepath='temp.wav')\n\t\tt = asr_model.transcribe(paths2audio_files=['temp.wav', ], batch_size=1)[0]\n\t\tprint(t)\n\t\tprint(d.sentence)\n\t\tprint('')\n\n```\n\nresults\n\n```\noutput_file: temp.wav already exists and will be overwritten on build\nTranscribing: 100%|██████████| 1/1 [00:00<00:00,  1.12it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nতিনি এবং তার মা তাদের পৈতৃক বাড়িতে থেকে প্রতিবেশীদের দ্বারা অনেক তিরস্কার সহ্য করেন ।\nতিনি এবং তাঁর মা তাদের পৈতৃক বাড়িতে থেকে প্রতিবেশীদের দ্বারা অনেক তিরস্কার সহ্য করেন।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 16.44it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nভিত্তিবাস রামায়ণ বহির্ভূত অনেক গল্প এই অনুবাদ গ্রহণ করেছিলেন ।\nকৃত্তিবাস রামায়ণ-বহির্ভূত অনেক গল্প এই অনুবাদে গ্রহণ করেছিলেন।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 15.31it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nতিনি তার সুশৃঙ্খল সামরিক বাহিনী এবং সুগঠিত শাসন কাঠামোর মাধ্যমে একটি দক্ষ শাসনব্যবস্থা প্রতিষ্ঠিত করেন ।\nতিনি তার সুশৃঙ্খল সামরিক বাহিনী এবং সুগঠিত শাসন কাঠামোর মাধ্যমে একটি দক্ষ শাসন ব্যবস্থা প্রতিষ্ঠিত করেন।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 15.50it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nতিনি বিজয় সাম্রাজ্যের বিরুদ্ধে এবং বিজ মুসলিম প্রতিবেশীদের বিরুদ্ধে যুদ্ধ করেছিলেন ।\nতিনি বিজয়নগর সাম্রাজ্যের বিরুদ্ধে এবং বিজাপুরের মুসলিম প্রতিবেশীদের বিরুদ্ধেও যুদ্ধ করেছিলেন।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 18.29it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nএটি মূলত একটি মরুময় অঞ্চল ।\nএটি মূলত একটি মরুময় অঞ্চল।\n\nTranscribing: 100%|██████████| 1/1 [00:00<00:00, 15.23it/s]\noutput_file: temp.wav already exists and will be overwritten on build\nসড়কটি বিহার পশ্চিমবঙ্গ সীমান্ত অতিক্রম করে পশ্চিমবঙ্গ রাজ্যে প্রবেশ করে উত্তর দিনাজপুর জেলা হয়ে ।\nসড়কটি বিহার-পশ্চিমবঙ্গ সীমান্ত অতিক্রম করে পশ্চিমবঙ্গ রাজ্যে প্রবেশ করে উত্তর দিনাজপুর জেলা হয়ে।\n\n```"
  }
}