{
  "id": 427872,
  "title": "[SOLVED] Problem when inferring",
  "url": "/competitions/bengaliai-speech/discussion/427872",
  "author_name": "",
  "post_date": "2023-07-30T05:21:50.235730100Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I tried YellowKing's Model, and it reported this problem. However, I referred to the original code of Yellowking's, and used their original processor when training, thus the problem is now solved.</p>\n<h4>Problem:</h4>\n<p>Here is my code for inference:</p>\n<pre><code> os\n numpy  np\n tqdm.auto  tqdm\n glob  glob\n transformers  AutoFeatureExtractor, pipeline\n pandas  pd\n librosa\n IPython\n datasets  load_metric\n tqdm.auto  tqdm\n torch.utils.data  Dataset, DataLoader\n torch\n gc\n wave\n scipy.io  wavfile\n scipy.signal  sps\n pyctcdecode\n\ntqdm.pandas()\n warnings\nwarnings.filterwarnings()\nos.environ[] = \n\nBATCH_SIZE = \nTEST_DIRECTORY = \npaths = glob(os.path.join(TEST_DIRECTORY,))\n :\n    my_model_name = \n    processor_name = \n transformers  Wav2Vec2ProcessorWithLM\n\nprocessor = Wav2Vec2ProcessorWithLM.from_pretrained(CFG.processor_name)\n transformers  Wav2Vec2ForCTC\n\nmodel = Wav2Vec2ForCTC.from_pretrained(CFG.my_model_name)\nmy_asrLM = pipeline(, model=model ,feature_extractor =processor.feature_extractor, tokenizer= processor.tokenizer,decoder=processor.decoder ,device=)\n ():\n     ():\n        self.paths = paths\n     ():\n         (self.paths)\n     ():\n        speech, sr = librosa.load(self.paths[idx], sr=processor.feature_extractor.sampling_rate) \n\n         speech\ndataset = AudioDataset(paths)\ndevice = \n ():\n    \n    \n    lengths = torch.tensor([ t.shape[]  t  batch ])\n    \n    batch = [ torch.Tensor(t)  t  batch ]\n    batch = torch.nn.utils.rnn.pad_sequence(batch)\n    \n    mask = (batch != )\n     batch, lengths, mask\ndataloader = DataLoader(dataset, batch_size=, shuffle=, num_workers=, collate_fn=collate_fn_padd)\n ():\n    speech, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n\n    my_LM_prediction = my_asrLM(speech)\n    (my_LM_prediction[])\n     my_LM_prediction[]\n ():\n    \n    results = []\n     path  audio_paths:\n        pred = \n        pred = infer(path)\n\n\n\n\n         (pred)==:\n            pred = \n        results.append(pred)\n\n     results\n\naudio_paths=[audio_path  audio_path  tqdm(paths)]\n\n\n\n\nsentences=[]\n idx  tqdm((,(audio_paths),BATCH_SIZE)):\n    batch_paths=audio_paths[idx:idx+BATCH_SIZE]\n    sentence=batch_infer(batch_paths)\n    (sentence)\n    sentences+=sentence\n\n\n\n\n\n</code></pre>\n<p>I used my own trained model in <code>/kaggle/input/run-004-wav2vec2-15data-constant-lr2e-6-1ksteps</code> and Yellowking's processor, then when running the core inference code</p>\n<pre><code>pred = infer(path)\n</code></pre>\n<p>the notebook met this problem:</p>\n<pre><code>---------------------------------------------------------------------------\nValueError                                Traceback (most recent  )\n/tmp/ipykernel_25/. in \n        idx in tqdm((,(audio_paths),BATCH_SIZE)):\n           batch_paths=audio_paths[idx:idx+BATCH_SIZE]\n---&gt;      sentence=batch_infer(batch_paths)\n          (sentence)\n          sentences+=sentence\n\n/tmp/ipykernel_25/. in batch_infer(audio_paths, batch_size)\n           path in audio_paths:\n              pred = \n---&gt;          pred = infer(path)\n      #         :\n      #             pred = infer(path)\n\n/tmp/ipykernel_25/. in infer(audio_path)\n           speech, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n       \n----&gt;      my_LM_prediction = my_asrLM(speech)\n           (my_LM_prediction[])\n            my_LM_prediction[]\n\n//conda/lib/./site-packages/transformers/pipelines/automatic_speech_recognition. in __call__(self, inputs, **kwargs)\n                             `.(chunk[]  chunk in output[])`.\n             \n--&gt;           super().__call__(inputs, **kwargs)\n     \n         def _sanitize_parameters(self, **kwargs):\n\n//conda/lib/./site-packages/transformers/pipelines/base. in __call__(self, inputs, num_workers, batch_size, *, **kwargs)\n                 self.iterate(inputs, preprocess_params, forward_params, postprocess_params)\n            :\n-&gt;               self.run_single(inputs, preprocess_params, forward_params, postprocess_params)\n    \n        def run_multi(self, inputs, preprocess_params, forward_params, postprocess_params):\n\n//conda/lib/./site-packages/transformers/pipelines/base. in run_single(self, inputs, preprocess_params, forward_params, postprocess_params)\n                model_outputs = self.forward(model_inputs, **forward_params)\n                all_outputs.(model_outputs)\n-&gt;          outputs = self.postprocess(all_outputs, **postprocess_params)\n             outputs\n    \n\n//conda/lib/./site-packages/transformers/pipelines/automatic_speech_recognition. in postprocess(self, model_outputs, decoder_kwargs, return_timestamps)\n                  decoder_kwargs  None:\n                     decoder_kwargs = {}\n--&gt;              beams = self.decoder.decode_beams(, **decoder_kwargs)\n                 text = beams[][]\n                  return_timestamps:\n\n//conda/lib/./site-packages/pyctcdecode/decoder. in decode_beams(self, logits, beam_width, beam_prune_logp, token_min_logp, prune_history, hotwords, hotword_weight, lm_start_state)\n                 List of beams of  OUTPUT_BEAM with various meta information\n             \n--&gt;          self._check_logits_dimension(logits)\n             # prepare hotword \n             hotword_scorer = HotwordScorer.build_scorer(hotwords, weight=hotword_weight)\n\n//conda/lib/./site-packages/pyctcdecode/decoder. in _check_logits_dimension(self, logits)\n                 raise ValueError(\n                     \n--&gt;                   % (logits.shape, (self._idx2vocab))\n                 )\n     \n\nValueError: Input logits shape  (, ), but vocabulary  size . Need logits of shape: (time, vocabulary)\n</code></pre>\n<p>I used Yellowking's Wav2Vec2 model as base model and trained it many steps. The validation WER seemed normal (approx. 0.4, 0.5) How could this happen? And how can I fix this? Thanks a lot!</p>",
  "messages": [
    {
      "id": "2365239",
      "postDate": "07/30/2023 05:21:50",
      "content": "<p>I tried YellowKing's Model, and it reported this problem. However, I referred to the original code of Yellowking's, and used their original processor when training, thus the problem is now solved.</p>\n<h4>Problem:</h4>\n<p>Here is my code for inference:</p>\n<pre><code> os\n numpy  np\n tqdm.auto  tqdm\n glob  glob\n transformers  AutoFeatureExtractor, pipeline\n pandas  pd\n librosa\n IPython\n datasets  load_metric\n tqdm.auto  tqdm\n torch.utils.data  Dataset, DataLoader\n torch\n gc\n wave\n scipy.io  wavfile\n scipy.signal  sps\n pyctcdecode\n\ntqdm.pandas()\n warnings\nwarnings.filterwarnings()\nos.environ[] = \n\nBATCH_SIZE = \nTEST_DIRECTORY = \npaths = glob(os.path.join(TEST_DIRECTORY,))\n :\n    my_model_name = \n    processor_name = \n transformers  Wav2Vec2ProcessorWithLM\n\nprocessor = Wav2Vec2ProcessorWithLM.from_pretrained(CFG.processor_name)\n transformers  Wav2Vec2ForCTC\n\nmodel = Wav2Vec2ForCTC.from_pretrained(CFG.my_model_name)\nmy_asrLM = pipeline(, model=model ,feature_extractor =processor.feature_extractor, tokenizer= processor.tokenizer,decoder=processor.decoder ,device=)\n ():\n     ():\n        self.paths = paths\n     ():\n         (self.paths)\n     ():\n        speech, sr = librosa.load(self.paths[idx], sr=processor.feature_extractor.sampling_rate) \n\n         speech\ndataset = AudioDataset(paths)\ndevice = \n ():\n    \n    \n    lengths = torch.tensor([ t.shape[]  t  batch ])\n    \n    batch = [ torch.Tensor(t)  t  batch ]\n    batch = torch.nn.utils.rnn.pad_sequence(batch)\n    \n    mask = (batch != )\n     batch, lengths, mask\ndataloader = DataLoader(dataset, batch_size=, shuffle=, num_workers=, collate_fn=collate_fn_padd)\n ():\n    speech, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n\n    my_LM_prediction = my_asrLM(speech)\n    (my_LM_prediction[])\n     my_LM_prediction[]\n ():\n    \n    results = []\n     path  audio_paths:\n        pred = \n        pred = infer(path)\n\n\n\n\n         (pred)==:\n            pred = \n        results.append(pred)\n\n     results\n\naudio_paths=[audio_path  audio_path  tqdm(paths)]\n\n\n\n\nsentences=[]\n idx  tqdm((,(audio_paths),BATCH_SIZE)):\n    batch_paths=audio_paths[idx:idx+BATCH_SIZE]\n    sentence=batch_infer(batch_paths)\n    (sentence)\n    sentences+=sentence\n\n\n\n\n\n</code></pre>\n<p>I used my own trained model in <code>/kaggle/input/run-004-wav2vec2-15data-constant-lr2e-6-1ksteps</code> and Yellowking's processor, then when running the core inference code</p>\n<pre><code>pred = infer(path)\n</code></pre>\n<p>the notebook met this problem:</p>\n<pre><code>---------------------------------------------------------------------------\nValueError                                Traceback (most recent  )\n/tmp/ipykernel_25/. in \n        idx in tqdm((,(audio_paths),BATCH_SIZE)):\n           batch_paths=audio_paths[idx:idx+BATCH_SIZE]\n---&gt;      sentence=batch_infer(batch_paths)\n          (sentence)\n          sentences+=sentence\n\n/tmp/ipykernel_25/. in batch_infer(audio_paths, batch_size)\n           path in audio_paths:\n              pred = \n---&gt;          pred = infer(path)\n      #         :\n      #             pred = infer(path)\n\n/tmp/ipykernel_25/. in infer(audio_path)\n           speech, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n       \n----&gt;      my_LM_prediction = my_asrLM(speech)\n           (my_LM_prediction[])\n            my_LM_prediction[]\n\n//conda/lib/./site-packages/transformers/pipelines/automatic_speech_recognition. in __call__(self, inputs, **kwargs)\n                             `.(chunk[]  chunk in output[])`.\n             \n--&gt;           super().__call__(inputs, **kwargs)\n     \n         def _sanitize_parameters(self, **kwargs):\n\n//conda/lib/./site-packages/transformers/pipelines/base. in __call__(self, inputs, num_workers, batch_size, *, **kwargs)\n                 self.iterate(inputs, preprocess_params, forward_params, postprocess_params)\n            :\n-&gt;               self.run_single(inputs, preprocess_params, forward_params, postprocess_params)\n    \n        def run_multi(self, inputs, preprocess_params, forward_params, postprocess_params):\n\n//conda/lib/./site-packages/transformers/pipelines/base. in run_single(self, inputs, preprocess_params, forward_params, postprocess_params)\n                model_outputs = self.forward(model_inputs, **forward_params)\n                all_outputs.(model_outputs)\n-&gt;          outputs = self.postprocess(all_outputs, **postprocess_params)\n             outputs\n    \n\n//conda/lib/./site-packages/transformers/pipelines/automatic_speech_recognition. in postprocess(self, model_outputs, decoder_kwargs, return_timestamps)\n                  decoder_kwargs  None:\n                     decoder_kwargs = {}\n--&gt;              beams = self.decoder.decode_beams(, **decoder_kwargs)\n                 text = beams[][]\n                  return_timestamps:\n\n//conda/lib/./site-packages/pyctcdecode/decoder. in decode_beams(self, logits, beam_width, beam_prune_logp, token_min_logp, prune_history, hotwords, hotword_weight, lm_start_state)\n                 List of beams of  OUTPUT_BEAM with various meta information\n             \n--&gt;          self._check_logits_dimension(logits)\n             # prepare hotword \n             hotword_scorer = HotwordScorer.build_scorer(hotwords, weight=hotword_weight)\n\n//conda/lib/./site-packages/pyctcdecode/decoder. in _check_logits_dimension(self, logits)\n                 raise ValueError(\n                     \n--&gt;                   % (logits.shape, (self._idx2vocab))\n                 )\n     \n\nValueError: Input logits shape  (, ), but vocabulary  size . Need logits of shape: (time, vocabulary)\n</code></pre>\n<p>I used Yellowking's Wav2Vec2 model as base model and trained it many steps. The validation WER seemed normal (approx. 0.4, 0.5) How could this happen? And how can I fix this? Thanks a lot!</p>",
      "rawMarkdown": "I tried YellowKing's Model, and it reported this problem. However, I referred to the original code of Yellowking's, and used their original processor when training, thus the problem is now solved.\n\n#### Problem:\n\nHere is my code for inference:\n```py\nimport os\nimport numpy as np\nfrom tqdm.auto import tqdm\nfrom glob import glob\nfrom transformers import AutoFeatureExtractor, pipeline\nimport pandas as pd\nimport librosa\nimport IPython\nfrom datasets import load_metric\nfrom tqdm.auto import tqdm\nfrom torch.utils.data import Dataset, DataLoader\nimport torch\nimport gc\nimport wave\nfrom scipy.io import wavfile\nimport scipy.signal as sps\nimport pyctcdecode\n\ntqdm.pandas()\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nos.environ[\"WANDB_DISABLED\"] = \"true\"\n# CHANGE ACCORDINGLY\nBATCH_SIZE = 1\nTEST_DIRECTORY = '/kaggle/input/bengaliai-speech/test_mp3s'\npaths = glob(os.path.join(TEST_DIRECTORY,'*.mp3'))\nclass CFG:\n    my_model_name = '/kaggle/input/run-004-wav2vec2-15data-constant-lr2e-6-1ksteps'\n    processor_name = '/kaggle/input/yellowking-dlsprint-model/YellowKing_processor'\nfrom transformers import Wav2Vec2ProcessorWithLM\n\nprocessor = Wav2Vec2ProcessorWithLM.from_pretrained(CFG.processor_name)\nfrom transformers import Wav2Vec2ForCTC\n\nmodel = Wav2Vec2ForCTC.from_pretrained(CFG.my_model_name)\nmy_asrLM = pipeline(\"automatic-speech-recognition\", model=model ,feature_extractor =processor.feature_extractor, tokenizer= processor.tokenizer,decoder=processor.decoder ,device=0)\nclass AudioDataset(Dataset):\n    def __init__(self, paths):\n        self.paths = paths\n    def __len__(self):\n        return len(self.paths)\n    def __getitem__(self,idx):\n        speech, sr = librosa.load(self.paths[idx], sr=processor.feature_extractor.sampling_rate) \n#         print(speech.shape)\n        return speech\ndataset = AudioDataset(paths)\ndevice = 'cuda:0'\ndef collate_fn_padd(batch):\n    '''\n    Padds batch of variable length\n\n    note: it converts things ToTensor manually here since the ToTensor transform\n    assume it takes in images rather than arbitrary tensors.\n    '''\n    ## get sequence lengths\n    lengths = torch.tensor([ t.shape[0] for t in batch ])\n    ## padd\n    batch = [ torch.Tensor(t) for t in batch ]\n    batch = torch.nn.utils.rnn.pad_sequence(batch)\n    ## compute mask\n    mask = (batch != 0)\n    return batch, lengths, mask\ndataloader = DataLoader(dataset, batch_size=1, shuffle=False, num_workers=4, collate_fn=collate_fn_padd)\ndef infer(audio_path):\n    speech, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n\n    my_LM_prediction = my_asrLM(speech)\n    print(my_LM_prediction['text'])\n    return my_LM_prediction['text']\ndef batch_infer(audio_paths, batch_size=BATCH_SIZE):\n    '''\n    infers on a batch of audio\n    args:\n      audio_paths  : list of path to audio files <list of string>\n    returns:\n      bangla predicted texts <list of string>\n    '''\n    results = []\n    for path in audio_paths:\n        pred = \"\"\n        pred = infer(path)\n#         try:\n#             pred = infer(path)\n#         except:\n#             pred = \"এ\"\n        if len(pred)==0:\n            pred = \"এ\"\n        results.append(pred)\n    \n    return results\n# preds_all = []\naudio_paths=[audio_path for audio_path in tqdm(paths)]\n# files = os.listdir(\"/kaggle/input/bengaliai-speech/test_mp3s\")\n# paths = []\n# for i in files:\n#     paths.append(i.split(\".\")[0])\nsentences=[]\nfor idx in tqdm(range(0,len(audio_paths),BATCH_SIZE)):\n    batch_paths=audio_paths[idx:idx+BATCH_SIZE]\n    sentence=batch_infer(batch_paths)\n    print(sentence)\n    sentences+=sentence\n\n# for batch, lengths, mask in dataloader:\n#     print(lengths)\n#     preds = my_asrLM(list(batch.numpy().transpose()))\n#     preds_all+=preds\n```\n\nI used my own trained model in `/kaggle/input/run-004-wav2vec2-15data-constant-lr2e-6-1ksteps` and Yellowking's processor, then when running the core inference code\n```py\npred = infer(path)\n```\nthe notebook met this problem:\n```\n---------------------------------------------------------------------------\nValueError                                Traceback (most recent call last)\n/tmp/ipykernel_25/1626861118.py in <module>\n      8 for idx in tqdm(range(0,len(audio_paths),BATCH_SIZE)):\n      9     batch_paths=audio_paths[idx:idx+BATCH_SIZE]\n---> 10     sentence=batch_infer(batch_paths)\n     11     print(sentence)\n     12     sentences+=sentence\n\n/tmp/ipykernel_25/2639051409.py in batch_infer(audio_paths, batch_size)\n     10     for path in audio_paths:\n     11         pred = \"\"\n---> 12         pred = infer(path)\n     13 #         try:\n     14 #             pred = infer(path)\n\n/tmp/ipykernel_25/4007737660.py in infer(audio_path)\n      2     speech, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n      3 \n----> 4     my_LM_prediction = my_asrLM(speech)\n      5     print(my_LM_prediction['text'])\n      6     return my_LM_prediction['text']\n\n/opt/conda/lib/python3.7/site-packages/transformers/pipelines/automatic_speech_recognition.py in __call__(self, inputs, **kwargs)\n    180                         `\"\".join(chunk[\"text\"] for chunk in output[\"chunks\"])`.\n    181         \"\"\"\n--> 182         return super().__call__(inputs, **kwargs)\n    183 \n    184     def _sanitize_parameters(self, **kwargs):\n\n/opt/conda/lib/python3.7/site-packages/transformers/pipelines/base.py in __call__(self, inputs, num_workers, batch_size, *args, **kwargs)\n   1041             return self.iterate(inputs, preprocess_params, forward_params, postprocess_params)\n   1042         else:\n-> 1043             return self.run_single(inputs, preprocess_params, forward_params, postprocess_params)\n   1044 \n   1045     def run_multi(self, inputs, preprocess_params, forward_params, postprocess_params):\n\n/opt/conda/lib/python3.7/site-packages/transformers/pipelines/base.py in run_single(self, inputs, preprocess_params, forward_params, postprocess_params)\n   1065             model_outputs = self.forward(model_inputs, **forward_params)\n   1066             all_outputs.append(model_outputs)\n-> 1067         outputs = self.postprocess(all_outputs, **postprocess_params)\n   1068         return outputs\n   1069 \n\n/opt/conda/lib/python3.7/site-packages/transformers/pipelines/automatic_speech_recognition.py in postprocess(self, model_outputs, decoder_kwargs, return_timestamps)\n    363             if decoder_kwargs is None:\n    364                 decoder_kwargs = {}\n--> 365             beams = self.decoder.decode_beams(items, **decoder_kwargs)\n    366             text = beams[0][0]\n    367             if return_timestamps:\n\n/opt/conda/lib/python3.7/site-packages/pyctcdecode/decoder.py in decode_beams(self, logits, beam_width, beam_prune_logp, token_min_logp, prune_history, hotwords, hotword_weight, lm_start_state)\n    547             List of beams of type OUTPUT_BEAM with various meta information\n    548         \"\"\"\n--> 549         self._check_logits_dimension(logits)\n    550         # prepare hotword input\n    551         hotword_scorer = HotwordScorer.build_scorer(hotwords, weight=hotword_weight)\n\n/opt/conda/lib/python3.7/site-packages/pyctcdecode/decoder.py in _check_logits_dimension(self, logits)\n    271             raise ValueError(\n    272                 \"Input logits shape is %s, but vocabulary is size %s. \"\n--> 273                 \"Need logits of shape: (time, vocabulary)\" % (logits.shape, len(self._idx2vocab))\n    274             )\n    275 \n\nValueError: Input logits shape is (377, 32), but vocabulary is size 112. Need logits of shape: (time, vocabulary)\n```\n\nI used Yellowking's Wav2Vec2 model as base model and trained it many steps. The validation WER seemed normal (approx. 0.4, 0.5) How could this happen? And how can I fix this? Thanks a lot!",
      "votes": null
    },
    {
      "id": "2365371",
      "postDate": "07/30/2023 07:30:28",
      "content": "<p>I wonder if the problem is from that I used LM in my inference. However, when training, I didn't use LM.</p>",
      "rawMarkdown": "I wonder if the problem is from that I used LM in my inference. However, when training, I didn't use LM.",
      "votes": null
    },
    {
      "id": "2365901",
      "postDate": "07/30/2023 14:44:57",
      "content": "<p>During training you don't need LM </p>",
      "rawMarkdown": "During training you don't need LM",
      "votes": null
    },
    {
      "id": "2368605",
      "postDate": "08/01/2023 09:02:30",
      "content": "<p>The problem here might be due to vocabulary mismatch. I see you used your own model with yellowking processor. What was the vocab size when you trained your model? Was it the same as yellowking's? I suggest creating your own processor and use it with your model. The LM can also give a rise to this problem. Make sure you normalize the text before building the LM as well. Should solve the problem making sure of these.</p>\n<p>Currently, the processor you used in this code has 112 characters but your model is giving the output in a 32 sized vocab. That's why the problem has risen.</p>",
      "rawMarkdown": "The problem here might be due to vocabulary mismatch. I see you used your own model with yellowking processor. What was the vocab size when you trained your model? Was it the same as yellowking's? I suggest creating your own processor and use it with your model. The LM can also give a rise to this problem. Make sure you normalize the text before building the LM as well. Should solve the problem making sure of these.\n\nCurrently, the processor you used in this code has 112 characters but your model is giving the output in a 32 sized vocab. That's why the problem has risen.",
      "votes": null
    },
    {
      "id": "2369014",
      "postDate": "08/01/2023 13:50:51",
      "content": "<p>I tried YellowKing's Model, and it reported this problem. However, I referred to the original code of Yellowking's, and used their original processor when training, thus the problem is now solved.</p>",
      "rawMarkdown": "I tried YellowKing's Model, and it reported this problem. However, I referred to the original code of Yellowking's, and used their original processor when training, thus the problem is now solved.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2365371,
      "author_name": "nisshokuitsuki",
      "author_url": "",
      "post_date": "07/30/2023 07:30:28",
      "content": "<p>I wonder if the problem is from that I used LM in my inference. However, when training, I didn't use LM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2365901,
          "author_name": "arunodhayan",
          "author_url": "",
          "post_date": "07/30/2023 14:44:57",
          "content": "<p>During training you don't need LM </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2368605,
      "author_name": "mahfuzulkabirsourav",
      "author_url": "",
      "post_date": "08/01/2023 09:02:30",
      "content": "<p>The problem here might be due to vocabulary mismatch. I see you used your own model with yellowking processor. What was the vocab size when you trained your model? Was it the same as yellowking's? I suggest creating your own processor and use it with your model. The LM can also give a rise to this problem. Make sure you normalize the text before building the LM as well. Should solve the problem making sure of these.</p>\n<p>Currently, the processor you used in this code has 112 characters but your model is giving the output in a 32 sized vocab. That's why the problem has risen.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2369014,
          "author_name": "nisshokuitsuki",
          "author_url": "",
          "post_date": "08/01/2023 13:50:51",
          "content": "<p>I tried YellowKing's Model, and it reported this problem. However, I referred to the original code of Yellowking's, and used their original processor when training, thus the problem is now solved.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2365239": "I tried YellowKing's Model, and it reported this problem. However, I referred to the original code of Yellowking's, and used their original processor when training, thus the problem is now solved.\n\n#### Problem:\n\nHere is my code for inference:\n```py\nimport os\nimport numpy as np\nfrom tqdm.auto import tqdm\nfrom glob import glob\nfrom transformers import AutoFeatureExtractor, pipeline\nimport pandas as pd\nimport librosa\nimport IPython\nfrom datasets import load_metric\nfrom tqdm.auto import tqdm\nfrom torch.utils.data import Dataset, DataLoader\nimport torch\nimport gc\nimport wave\nfrom scipy.io import wavfile\nimport scipy.signal as sps\nimport pyctcdecode\n\ntqdm.pandas()\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nos.environ[\"WANDB_DISABLED\"] = \"true\"\n# CHANGE ACCORDINGLY\nBATCH_SIZE = 1\nTEST_DIRECTORY = '/kaggle/input/bengaliai-speech/test_mp3s'\npaths = glob(os.path.join(TEST_DIRECTORY,'*.mp3'))\nclass CFG:\n    my_model_name = '/kaggle/input/run-004-wav2vec2-15data-constant-lr2e-6-1ksteps'\n    processor_name = '/kaggle/input/yellowking-dlsprint-model/YellowKing_processor'\nfrom transformers import Wav2Vec2ProcessorWithLM\n\nprocessor = Wav2Vec2ProcessorWithLM.from_pretrained(CFG.processor_name)\nfrom transformers import Wav2Vec2ForCTC\n\nmodel = Wav2Vec2ForCTC.from_pretrained(CFG.my_model_name)\nmy_asrLM = pipeline(\"automatic-speech-recognition\", model=model ,feature_extractor =processor.feature_extractor, tokenizer= processor.tokenizer,decoder=processor.decoder ,device=0)\nclass AudioDataset(Dataset):\n    def __init__(self, paths):\n        self.paths = paths\n    def __len__(self):\n        return len(self.paths)\n    def __getitem__(self,idx):\n        speech, sr = librosa.load(self.paths[idx], sr=processor.feature_extractor.sampling_rate) \n#         print(speech.shape)\n        return speech\ndataset = AudioDataset(paths)\ndevice = 'cuda:0'\ndef collate_fn_padd(batch):\n    '''\n    Padds batch of variable length\n\n    note: it converts things ToTensor manually here since the ToTensor transform\n    assume it takes in images rather than arbitrary tensors.\n    '''\n    ## get sequence lengths\n    lengths = torch.tensor([ t.shape[0] for t in batch ])\n    ## padd\n    batch = [ torch.Tensor(t) for t in batch ]\n    batch = torch.nn.utils.rnn.pad_sequence(batch)\n    ## compute mask\n    mask = (batch != 0)\n    return batch, lengths, mask\ndataloader = DataLoader(dataset, batch_size=1, shuffle=False, num_workers=4, collate_fn=collate_fn_padd)\ndef infer(audio_path):\n    speech, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n\n    my_LM_prediction = my_asrLM(speech)\n    print(my_LM_prediction['text'])\n    return my_LM_prediction['text']\ndef batch_infer(audio_paths, batch_size=BATCH_SIZE):\n    '''\n    infers on a batch of audio\n    args:\n      audio_paths  : list of path to audio files <list of string>\n    returns:\n      bangla predicted texts <list of string>\n    '''\n    results = []\n    for path in audio_paths:\n        pred = \"\"\n        pred = infer(path)\n#         try:\n#             pred = infer(path)\n#         except:\n#             pred = \"এ\"\n        if len(pred)==0:\n            pred = \"এ\"\n        results.append(pred)\n    \n    return results\n# preds_all = []\naudio_paths=[audio_path for audio_path in tqdm(paths)]\n# files = os.listdir(\"/kaggle/input/bengaliai-speech/test_mp3s\")\n# paths = []\n# for i in files:\n#     paths.append(i.split(\".\")[0])\nsentences=[]\nfor idx in tqdm(range(0,len(audio_paths),BATCH_SIZE)):\n    batch_paths=audio_paths[idx:idx+BATCH_SIZE]\n    sentence=batch_infer(batch_paths)\n    print(sentence)\n    sentences+=sentence\n\n# for batch, lengths, mask in dataloader:\n#     print(lengths)\n#     preds = my_asrLM(list(batch.numpy().transpose()))\n#     preds_all+=preds\n```\n\nI used my own trained model in `/kaggle/input/run-004-wav2vec2-15data-constant-lr2e-6-1ksteps` and Yellowking's processor, then when running the core inference code\n```py\npred = infer(path)\n```\nthe notebook met this problem:\n```\n---------------------------------------------------------------------------\nValueError                                Traceback (most recent call last)\n/tmp/ipykernel_25/1626861118.py in <module>\n      8 for idx in tqdm(range(0,len(audio_paths),BATCH_SIZE)):\n      9     batch_paths=audio_paths[idx:idx+BATCH_SIZE]\n---> 10     sentence=batch_infer(batch_paths)\n     11     print(sentence)\n     12     sentences+=sentence\n\n/tmp/ipykernel_25/2639051409.py in batch_infer(audio_paths, batch_size)\n     10     for path in audio_paths:\n     11         pred = \"\"\n---> 12         pred = infer(path)\n     13 #         try:\n     14 #             pred = infer(path)\n\n/tmp/ipykernel_25/4007737660.py in infer(audio_path)\n      2     speech, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n      3 \n----> 4     my_LM_prediction = my_asrLM(speech)\n      5     print(my_LM_prediction['text'])\n      6     return my_LM_prediction['text']\n\n/opt/conda/lib/python3.7/site-packages/transformers/pipelines/automatic_speech_recognition.py in __call__(self, inputs, **kwargs)\n    180                         `\"\".join(chunk[\"text\"] for chunk in output[\"chunks\"])`.\n    181         \"\"\"\n--> 182         return super().__call__(inputs, **kwargs)\n    183 \n    184     def _sanitize_parameters(self, **kwargs):\n\n/opt/conda/lib/python3.7/site-packages/transformers/pipelines/base.py in __call__(self, inputs, num_workers, batch_size, *args, **kwargs)\n   1041             return self.iterate(inputs, preprocess_params, forward_params, postprocess_params)\n   1042         else:\n-> 1043             return self.run_single(inputs, preprocess_params, forward_params, postprocess_params)\n   1044 \n   1045     def run_multi(self, inputs, preprocess_params, forward_params, postprocess_params):\n\n/opt/conda/lib/python3.7/site-packages/transformers/pipelines/base.py in run_single(self, inputs, preprocess_params, forward_params, postprocess_params)\n   1065             model_outputs = self.forward(model_inputs, **forward_params)\n   1066             all_outputs.append(model_outputs)\n-> 1067         outputs = self.postprocess(all_outputs, **postprocess_params)\n   1068         return outputs\n   1069 \n\n/opt/conda/lib/python3.7/site-packages/transformers/pipelines/automatic_speech_recognition.py in postprocess(self, model_outputs, decoder_kwargs, return_timestamps)\n    363             if decoder_kwargs is None:\n    364                 decoder_kwargs = {}\n--> 365             beams = self.decoder.decode_beams(items, **decoder_kwargs)\n    366             text = beams[0][0]\n    367             if return_timestamps:\n\n/opt/conda/lib/python3.7/site-packages/pyctcdecode/decoder.py in decode_beams(self, logits, beam_width, beam_prune_logp, token_min_logp, prune_history, hotwords, hotword_weight, lm_start_state)\n    547             List of beams of type OUTPUT_BEAM with various meta information\n    548         \"\"\"\n--> 549         self._check_logits_dimension(logits)\n    550         # prepare hotword input\n    551         hotword_scorer = HotwordScorer.build_scorer(hotwords, weight=hotword_weight)\n\n/opt/conda/lib/python3.7/site-packages/pyctcdecode/decoder.py in _check_logits_dimension(self, logits)\n    271             raise ValueError(\n    272                 \"Input logits shape is %s, but vocabulary is size %s. \"\n--> 273                 \"Need logits of shape: (time, vocabulary)\" % (logits.shape, len(self._idx2vocab))\n    274             )\n    275 \n\nValueError: Input logits shape is (377, 32), but vocabulary is size 112. Need logits of shape: (time, vocabulary)\n```\n\nI used Yellowking's Wav2Vec2 model as base model and trained it many steps. The validation WER seemed normal (approx. 0.4, 0.5) How could this happen? And how can I fix this? Thanks a lot!",
    "2365371": "I wonder if the problem is from that I used LM in my inference. However, when training, I didn't use LM.",
    "2365901": "During training you don't need LM",
    "2368605": "The problem here might be due to vocabulary mismatch. I see you used your own model with yellowking processor. What was the vocab size when you trained your model? Was it the same as yellowking's? I suggest creating your own processor and use it with your model. The LM can also give a rise to this problem. Make sure you normalize the text before building the LM as well. Should solve the problem making sure of these.\n\nCurrently, the processor you used in this code has 112 characters but your model is giving the output in a 32 sized vocab. That's why the problem has risen.",
    "2369014": "I tried YellowKing's Model, and it reported this problem. However, I referred to the original code of Yellowking's, and used their original processor when training, thus the problem is now solved."
  },
  "source": "meta"
}