{
  "id": 443778,
  "title": "How to use hugging face language models with pyctcdecode",
  "url": "/competitions/bengaliai-speech/discussion/443778",
  "author_name": "Antar",
  "post_date": "2023-09-28T18:31:37.778000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello, I'm newbie and want to use this model as language model:<br>\n<a href=\"https://huggingface.co/arijitx/wav2vec2-large-xlsr-bengali\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-large-xlsr-bengali</a> as in this notebook <a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline/notebook\" target=\"_blank\">https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline/notebook</a></p>\n<p>I use pyctcdecode but I can't find the 5gram.bin (used in the notebook) to pass to <br>\n<code>\ndecoder = build_ctcdecoder(\n    labels,\n    kenlm_model_path=\".cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin\",  # either .arpa or .bin file\n)\n</code></p>\n<p>when I pass the downloaded .bin model to kenlm_model_path it gives me this error:<br>\n<code>Cannot read model '/home/asr_r02/.cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin' (lm/read_arpa.cc:65 in void lm::ReadARPACounts(util::FilePiece&amp;, std::vector&lt;long unsigned int&gt;&amp;) threw FormatLoadException. first non-empty line was \"PK\u0003\u0004)</code><br>\nplease hlep</p>",
  "messages": [
    {
      "id": 2460382,
      "postDate": "2023-09-28T18:31:37.780Z",
      "content": "<p>Hello, I'm newbie and want to use this model as language model:<br>\n<a href=\"https://huggingface.co/arijitx/wav2vec2-large-xlsr-bengali\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-large-xlsr-bengali</a> as in this notebook <a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline/notebook\" target=\"_blank\">https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline/notebook</a></p>\n<p>I use pyctcdecode but I can't find the 5gram.bin (used in the notebook) to pass to <br>\n<code>\ndecoder = build_ctcdecoder(\n    labels,\n    kenlm_model_path=\".cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin\",  # either .arpa or .bin file\n)\n</code></p>\n<p>when I pass the downloaded .bin model to kenlm_model_path it gives me this error:<br>\n<code>Cannot read model '/home/asr_r02/.cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin' (lm/read_arpa.cc:65 in void lm::ReadARPACounts(util::FilePiece&amp;, std::vector&lt;long unsigned int&gt;&amp;) threw FormatLoadException. first non-empty line was \"PK\u0003\u0004)</code><br>\nplease hlep</p>",
      "rawMarkdown": "Hello, I'm newbie and want to use this model as language model:\nhttps://huggingface.co/arijitx/wav2vec2-large-xlsr-bengali as in this notebook https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline/notebook\n\nI use pyctcdecode but I can't find the 5gram.bin (used in the notebook) to pass to \n`\ndecoder = build_ctcdecoder(\n    labels,\n    kenlm_model_path=\".cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin\",  # either .arpa or .bin file\n)\n`\n\nwhen I pass the downloaded .bin model to kenlm_model_path it gives me this error:\n`Cannot read model '/home/asr_r02/.cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin' (lm/read_arpa.cc:65 in void lm::ReadARPACounts(util::FilePiece&, std::vector<long unsigned int>&) threw FormatLoadException. first non-empty line was \"PK\u0003\u0004)`\nplease hlep",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2460382": "Hello, I'm newbie and want to use this model as language model:\nhttps://huggingface.co/arijitx/wav2vec2-large-xlsr-bengali as in this notebook https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline/notebook\n\nI use pyctcdecode but I can't find the 5gram.bin (used in the notebook) to pass to \n`\ndecoder = build_ctcdecoder(\n    labels,\n    kenlm_model_path=\".cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin\",  # either .arpa or .bin file\n)\n`\n\nwhen I pass the downloaded .bin model to kenlm_model_path it gives me this error:\n`Cannot read model '/home/asr_r02/.cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin' (lm/read_arpa.cc:65 in void lm::ReadARPACounts(util::FilePiece&, std::vector<long unsigned int>&) threw FormatLoadException. first non-empty line was \"PK\u0003\u0004)`\nplease hlep"
  }
}