{
  "id": 448006,
  "title": "5th place solution - ensembling works",
  "url": "/competitions/bengaliai-speech/discussion/448006",
  "author_name": "Benedikt Droste",
  "post_date": "2023-10-18T05:22:17.734000",
  "votes": 34,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thank you for organizing this competition so well. Whenever questions arose, the hosts responded promptly and clearly. I really enjoyed working on this problem. Thank you very much <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> </p>\n<p><strong>Modeling</strong><br>\nLike most of the participants, I started with the YellowKing pipeline. I changed the backbone to IndicWav2Vec and reinitialized the CTC layer. With this method and the public Dari/normalizing pipeline, it was already possible to achieve 0.48x scores without a Language Model. However, there was a big problem: <strong>the local WER score did not match the public lb anymore if the model was trained for too long</strong>. At a learning rate of 3e-5, local overfitting started at 40k steps (batch size 16). I suspect that the model started learning the relationship of erroneous audio/annotation pairs.</p>\n<p>Due to the noisy annotation, I excluded all samples with a mos-score of &gt;2.0 in the subsequent training stage, reaching a score of 0.472. After re-labeling with the new model, I removed all samples with a WER above 0.5 for my model and a WER above 0.5 for YellowKing. <strong>This resolved the local overfitting issue</strong>, ensuring that local improvements correlated with better public lb scores. I trained the model over 210k steps (bs 16) with a learning rate of 8e-5, achieving a score of 0.452 without a language model on top.</p>\n<p>Additionally, I trained a model with a <a href=\"https://huggingface.co/facebook/wav2vec2-xls-r-1b\" target=\"_blank\">larger backbone</a>. Unfortunately, due to lack of time, I could not fully exploit the potential, but after 135k steps I could already achieve a score of 0.454. The learning rate for the larger model had to be much smaller (1e-5).</p>\n<p><strong>Key Takeaways:</strong></p>\n<ul>\n<li>Quality filtering of data significantly enhanced training performance.</li>\n</ul>\n<p><strong>Ensembling</strong><br>\nI tried early to average the model logits of different finetuned models. However, it doesn´t work because the predictions may not be aligned. This is even the case for models of the same architecture. In <a href=\"https://arxiv.org/pdf/2206.05518.pdf\" target=\"_blank\">this paper</a> a procedure is described, where <strong>features of different ASR models are concatenated</strong>, followed by the addition of a transformer encoder and a CTC layer. The weights of the asr models are freezed and only the transformer encoder and the CTC layer are trained. I applied this approach to my finetuned models:</p>\n<ul>\n<li>Extract embeddings of the finetuned models</li>\n<li>Concatenate the features</li>\n<li>Concatenated features are passed into the Transformer encoder and then processed in the CTC layer</li>\n<li>A few steps (7k steps for batchsize 8) on the same training data were enough to increase the performance of the overall model from 0.355 to 0.344 (I justed tested it in the full pipeline, so with language model, punctuation and normalizing)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fadb7e22f32c07808c16ed5ece66f5032%2FBild1.png?generation=1697605346189304&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p><strong>Key takeways:</strong></p>\n<ul>\n<li>Ensembling asr model outputs by concatenating the last hidden states and adding a transformer encoder with CTC layer improves performance a lot</li>\n</ul>\n<p><strong>Potentials:</strong></p>\n<ul>\n<li>I think this is where the <strong>greatest further potential lies</strong> in my solution. I didn't have enough time, but it would be interesting to see how Whisper or HUBERT in the ensemble would have improved the performance. I think because of the different architectures it should be even better because of the larger diversity in the ensemble.</li>\n<li>My finetuned model with the 1b backbone is undertrained. There is still potential for better performance if the model is trained for a longer period.</li>\n<li>I have hardly experimented with the parameters and training routine of the added transformer encoder. There is room for improvement as well.</li>\n</ul>\n<p><strong>Language Model</strong><br>\nI have thought a lot about the language model part. I had thought about interpreting the whole thing as a seq2seq problem and taking a pretrained transformer as the language model. In the end I just used a standard KenLM model. For this I used datasets which were shared in the forums (e.g. IndicCorp, SLR, Comp data).</p>\n<p>In preprocessing, I radically <strong>limited to the character set of the training data</strong>, as the<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/433469#2410416\" target=\"_blank\"> hosts said </a> there are no new characters in the test data. I replaced numbers, other languages, and other characters with special tokens. This reduced the number of unigrams significantly which was a bottleneck in the Kaggle environment. I ended up with <strong>a 0 0 0 2 2 pruned 5-ngram</strong>. Adding more data didn´t increase the performance at a certain point.</p>\n<p>Much more important for the Kaggle workflow was the following fact: If you use the standard Wav2Vec2ProcessorWithLM of Huggingface and then remove the model, the <a href=\"https://github.com/kensho-technologies/pyctcdecode/pull/111\" target=\"_blank\">RAM is not released</a>. I therefore disassembled the pipeline and first computed all logits, stored them locally, and then read them back in for the decoding process. Afterwards I removed the language model cleanly with decoder.cleanup() before deleting the decoder. Without this measure, I would not have been able to use the ensemble model because of out of memory errors.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fd077b0fb9e9042395af9401e0b509c65%2FBild2.png?generation=1697605762499198&amp;alt=media\" alt=\"\"></p>\n<p><strong>Key Takeways:</strong></p>\n<ul>\n<li>The language model is important: precise preprocessing manages size while excessive data may not necessarily improve the LM's performance. Corrupted data can even degrade the performance of the language model.</li>\n</ul>\n<p><strong>Punctuation</strong><br>\nI tried different punctuation models and ensembles. In the end, I used a <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">XLM-Roberta-Large model</a>. It improved the score by ~0.02. I could improve the performance even more by forcing the model to output a | or ? after the last word (+0.002).</p>\n<p><strong>Code</strong><br>\nCleaned up inference code can be found here:<br>\n<a href=\"https://www.kaggle.com/code/benbla/5th-place-solution/notebook\" target=\"_blank\">Kaggle Inference</a><br>\nCleaned up training code can be found here:<br>\n<a href=\"https://github.com/bd317/BengaliAI_Speech_Recognition_5th_solution\" target=\"_blank\">Github training</a></p>",
  "messages": [
    {
      "id": 2486677,
      "postDate": "2023-10-18T05:22:17.733Z",
      "content": "<p>Thank you for organizing this competition so well. Whenever questions arose, the hosts responded promptly and clearly. I really enjoyed working on this problem. Thank you very much <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> </p>\n<p><strong>Modeling</strong><br>\nLike most of the participants, I started with the YellowKing pipeline. I changed the backbone to IndicWav2Vec and reinitialized the CTC layer. With this method and the public Dari/normalizing pipeline, it was already possible to achieve 0.48x scores without a Language Model. However, there was a big problem: <strong>the local WER score did not match the public lb anymore if the model was trained for too long</strong>. At a learning rate of 3e-5, local overfitting started at 40k steps (batch size 16). I suspect that the model started learning the relationship of erroneous audio/annotation pairs.</p>\n<p>Due to the noisy annotation, I excluded all samples with a mos-score of &gt;2.0 in the subsequent training stage, reaching a score of 0.472. After re-labeling with the new model, I removed all samples with a WER above 0.5 for my model and a WER above 0.5 for YellowKing. <strong>This resolved the local overfitting issue</strong>, ensuring that local improvements correlated with better public lb scores. I trained the model over 210k steps (bs 16) with a learning rate of 8e-5, achieving a score of 0.452 without a language model on top.</p>\n<p>Additionally, I trained a model with a <a href=\"https://huggingface.co/facebook/wav2vec2-xls-r-1b\" target=\"_blank\">larger backbone</a>. Unfortunately, due to lack of time, I could not fully exploit the potential, but after 135k steps I could already achieve a score of 0.454. The learning rate for the larger model had to be much smaller (1e-5).</p>\n<p><strong>Key Takeaways:</strong></p>\n<ul>\n<li>Quality filtering of data significantly enhanced training performance.</li>\n</ul>\n<p><strong>Ensembling</strong><br>\nI tried early to average the model logits of different finetuned models. However, it doesn´t work because the predictions may not be aligned. This is even the case for models of the same architecture. In <a href=\"https://arxiv.org/pdf/2206.05518.pdf\" target=\"_blank\">this paper</a> a procedure is described, where <strong>features of different ASR models are concatenated</strong>, followed by the addition of a transformer encoder and a CTC layer. The weights of the asr models are freezed and only the transformer encoder and the CTC layer are trained. I applied this approach to my finetuned models:</p>\n<ul>\n<li>Extract embeddings of the finetuned models</li>\n<li>Concatenate the features</li>\n<li>Concatenated features are passed into the Transformer encoder and then processed in the CTC layer</li>\n<li>A few steps (7k steps for batchsize 8) on the same training data were enough to increase the performance of the overall model from 0.355 to 0.344 (I justed tested it in the full pipeline, so with language model, punctuation and normalizing)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fadb7e22f32c07808c16ed5ece66f5032%2FBild1.png?generation=1697605346189304&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p><strong>Key takeways:</strong></p>\n<ul>\n<li>Ensembling asr model outputs by concatenating the last hidden states and adding a transformer encoder with CTC layer improves performance a lot</li>\n</ul>\n<p><strong>Potentials:</strong></p>\n<ul>\n<li>I think this is where the <strong>greatest further potential lies</strong> in my solution. I didn't have enough time, but it would be interesting to see how Whisper or HUBERT in the ensemble would have improved the performance. I think because of the different architectures it should be even better because of the larger diversity in the ensemble.</li>\n<li>My finetuned model with the 1b backbone is undertrained. There is still potential for better performance if the model is trained for a longer period.</li>\n<li>I have hardly experimented with the parameters and training routine of the added transformer encoder. There is room for improvement as well.</li>\n</ul>\n<p><strong>Language Model</strong><br>\nI have thought a lot about the language model part. I had thought about interpreting the whole thing as a seq2seq problem and taking a pretrained transformer as the language model. In the end I just used a standard KenLM model. For this I used datasets which were shared in the forums (e.g. IndicCorp, SLR, Comp data).</p>\n<p>In preprocessing, I radically <strong>limited to the character set of the training data</strong>, as the<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/433469#2410416\" target=\"_blank\"> hosts said </a> there are no new characters in the test data. I replaced numbers, other languages, and other characters with special tokens. This reduced the number of unigrams significantly which was a bottleneck in the Kaggle environment. I ended up with <strong>a 0 0 0 2 2 pruned 5-ngram</strong>. Adding more data didn´t increase the performance at a certain point.</p>\n<p>Much more important for the Kaggle workflow was the following fact: If you use the standard Wav2Vec2ProcessorWithLM of Huggingface and then remove the model, the <a href=\"https://github.com/kensho-technologies/pyctcdecode/pull/111\" target=\"_blank\">RAM is not released</a>. I therefore disassembled the pipeline and first computed all logits, stored them locally, and then read them back in for the decoding process. Afterwards I removed the language model cleanly with decoder.cleanup() before deleting the decoder. Without this measure, I would not have been able to use the ensemble model because of out of memory errors.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fd077b0fb9e9042395af9401e0b509c65%2FBild2.png?generation=1697605762499198&amp;alt=media\" alt=\"\"></p>\n<p><strong>Key Takeways:</strong></p>\n<ul>\n<li>The language model is important: precise preprocessing manages size while excessive data may not necessarily improve the LM's performance. Corrupted data can even degrade the performance of the language model.</li>\n</ul>\n<p><strong>Punctuation</strong><br>\nI tried different punctuation models and ensembles. In the end, I used a <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">XLM-Roberta-Large model</a>. It improved the score by ~0.02. I could improve the performance even more by forcing the model to output a | or ? after the last word (+0.002).</p>\n<p><strong>Code</strong><br>\nCleaned up inference code can be found here:<br>\n<a href=\"https://www.kaggle.com/code/benbla/5th-place-solution/notebook\" target=\"_blank\">Kaggle Inference</a><br>\nCleaned up training code can be found here:<br>\n<a href=\"https://github.com/bd317/BengaliAI_Speech_Recognition_5th_solution\" target=\"_blank\">Github training</a></p>",
      "rawMarkdown": "Thank you for organizing this competition so well. Whenever questions arose, the hosts responded promptly and clearly. I really enjoyed working on this problem. Thank you very much @imtiazprio \n\n**Modeling**\nLike most of the participants, I started with the YellowKing pipeline. I changed the backbone to IndicWav2Vec and reinitialized the CTC layer. With this method and the public Dari/normalizing pipeline, it was already possible to achieve 0.48x scores without a Language Model. However, there was a big problem: **the local WER score did not match the public lb anymore if the model was trained for too long**. At a learning rate of 3e-5, local overfitting started at 40k steps (batch size 16). I suspect that the model started learning the relationship of erroneous audio/annotation pairs.\n\nDue to the noisy annotation, I excluded all samples with a mos-score of >2.0 in the subsequent training stage, reaching a score of 0.472. After re-labeling with the new model, I removed all samples with a WER above 0.5 for my model and a WER above 0.5 for YellowKing. **This resolved the local overfitting issue**, ensuring that local improvements correlated with better public lb scores. I trained the model over 210k steps (bs 16) with a learning rate of 8e-5, achieving a score of 0.452 without a language model on top.\n\nAdditionally, I trained a model with a [larger backbone](https://huggingface.co/facebook/wav2vec2-xls-r-1b). Unfortunately, due to lack of time, I could not fully exploit the potential, but after 135k steps I could already achieve a score of 0.454. The learning rate for the larger model had to be much smaller (1e-5).\n\n**Key Takeaways:**\n- Quality filtering of data significantly enhanced training performance.\n\n**Ensembling**\nI tried early to average the model logits of different finetuned models. However, it doesn´t work because the predictions may not be aligned. This is even the case for models of the same architecture. In [this paper](https://arxiv.org/pdf/2206.05518.pdf) a procedure is described, where **features of different ASR models are concatenated**, followed by the addition of a transformer encoder and a CTC layer. The weights of the asr models are freezed and only the transformer encoder and the CTC layer are trained. I applied this approach to my finetuned models:\n- Extract embeddings of the finetuned models\n- Concatenate the features\n- Concatenated features are passed into the Transformer encoder and then processed in the CTC layer\n- A few steps (7k steps for batchsize 8) on the same training data were enough to increase the performance of the overall model from 0.355 to 0.344 (I justed tested it in the full pipeline, so with language model, punctuation and normalizing)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fadb7e22f32c07808c16ed5ece66f5032%2FBild1.png?generation=1697605346189304&alt=media)\n\n**Key takeways:**\n- Ensembling asr model outputs by concatenating the last hidden states and adding a transformer encoder with CTC layer improves performance a lot\n\n**Potentials:**\n- I think this is where the **greatest further potential lies** in my solution. I didn't have enough time, but it would be interesting to see how Whisper or HUBERT in the ensemble would have improved the performance. I think because of the different architectures it should be even better because of the larger diversity in the ensemble.\n- My finetuned model with the 1b backbone is undertrained. There is still potential for better performance if the model is trained for a longer period.\n- I have hardly experimented with the parameters and training routine of the added transformer encoder. There is room for improvement as well.\n\n**Language Model**\nI have thought a lot about the language model part. I had thought about interpreting the whole thing as a seq2seq problem and taking a pretrained transformer as the language model. In the end I just used a standard KenLM model. For this I used datasets which were shared in the forums (e.g. IndicCorp, SLR, Comp data).\n\nIn preprocessing, I radically **limited to the character set of the training data**, as the[ hosts said ](https://www.kaggle.com/competitions/bengaliai-speech/discussion/433469#2410416) there are no new characters in the test data. I replaced numbers, other languages, and other characters with special tokens. This reduced the number of unigrams significantly which was a bottleneck in the Kaggle environment. I ended up with **a 0 0 0 2 2 pruned 5-ngram**. Adding more data didn´t increase the performance at a certain point.\n\nMuch more important for the Kaggle workflow was the following fact: If you use the standard Wav2Vec2ProcessorWithLM of Huggingface and then remove the model, the [RAM is not released](https://github.com/kensho-technologies/pyctcdecode/pull/111). I therefore disassembled the pipeline and first computed all logits, stored them locally, and then read them back in for the decoding process. Afterwards I removed the language model cleanly with decoder.cleanup() before deleting the decoder. Without this measure, I would not have been able to use the ensemble model because of out of memory errors.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fd077b0fb9e9042395af9401e0b509c65%2FBild2.png?generation=1697605762499198&alt=media)\n\n**Key Takeways:**\n- The language model is important: precise preprocessing manages size while excessive data may not necessarily improve the LM's performance. Corrupted data can even degrade the performance of the language model.\n\n**Punctuation**\nI tried different punctuation models and ensembles. In the end, I used a [XLM-Roberta-Large model](https://github.com/xashru/punctuation-restoration). It improved the score by ~0.02. I could improve the performance even more by forcing the model to output a | or ? after the last word (+0.002).\n\n**Code**\nCleaned up inference code can be found here:\n[Kaggle Inference](https://www.kaggle.com/code/benbla/5th-place-solution/notebook)\nCleaned up training code can be found here:\n[Github training](https://github.com/bd317/BengaliAI_Speech_Recognition_5th_solution)",
      "votes": 34
    },
    {
      "id": 2488104,
      "postDate": "2023-10-19T03:27:27.620Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> on your amazing achievement. I had some questions regarding building a n-gram Kenlm language model to be used with Beam Search. </p>\n<p>Let's say I have trained a wav2vec2 STT model using a custom Bengali vocabulary set of 70 Bengali unicodes and special tokens which is different from the vocabulary (vocab.json) of the pre-trained wav2vec2 STT model like <a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali/tree/main\" target=\"_blank\">ai4bharat/indicwav2vec_v1_bengali</a>. How should I train a custom n-gram Kenlm model? As i have found that open source Kenlm models, for example: <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model</a> degrade the inference result when used with wav2vec2 STT models with custom vocabulary. Any suggestions regarding this problem? </p>\n<p>Apologies for any ignorant questions. </p>\n<p>Thanks. </p>",
      "rawMarkdown": "Congratulations @benbla on your amazing achievement. I had some questions regarding building a n-gram Kenlm language model to be used with Beam Search. \n\nLet's say I have trained a wav2vec2 STT model using a custom Bengali vocabulary set of 70 Bengali unicodes and special tokens which is different from the vocabulary (vocab.json) of the pre-trained wav2vec2 STT model like [ai4bharat/indicwav2vec_v1_bengali](https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali/tree/main). How should I train a custom n-gram Kenlm model? As i have found that open source Kenlm models, for example: https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model degrade the inference result when used with wav2vec2 STT models with custom vocabulary. Any suggestions regarding this problem? \n\nApologies for any ignorant questions. \n\nThanks. ",
      "votes": 1,
      "replies": [
        {
          "id": 2488288,
          "postDate": "2023-10-19T07:15:07.130Z",
          "content": "<p>It is important to normalize the texts the same way you normalize the training data for the audio model. </p>\n<p>Therefore you should preprocess and normalize the texts before training your KenLM model. If you use the <strong>bnunicodenormalizer</strong> for the training annotations, you should use it for the language model text corpus as well. </p>\n<p>I think the problem with the linked KenLM model is not related to the custom vocabulary, but rather that the texts for the model were normalized differently or not at all.</p>",
          "rawMarkdown": "It is important to normalize the texts the same way you normalize the training data for the audio model. \n\nTherefore you should preprocess and normalize the texts before training your KenLM model. If you use the **bnunicodenormalizer** for the training annotations, you should use it for the language model text corpus as well. \n\nI think the problem with the linked KenLM model is not related to the custom vocabulary, but rather that the texts for the model were normalized differently or not at all.",
          "votes": 1,
          "replies": [
            {
              "id": 2490186,
              "postDate": "2023-10-20T14:05:41.087Z",
              "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a>. Could you share the steps or possibly any tutorials you followed to train your custom KenlM model?</p>",
              "rawMarkdown": "Thanks for the answer @benbla. Could you share the steps or possibly any tutorials you followed to train your custom KenlM model?"
            }
          ]
        }
      ]
    },
    {
      "id": 2491219,
      "postDate": "2023-10-21T12:33:47.037Z",
      "content": "<p>Inference code is cleaned up and published:<br>\n<a href=\"https://www.kaggle.com/code/benbla/5th-place-solution/notebook\" target=\"_blank\">https://www.kaggle.com/code/benbla/5th-place-solution/notebook</a></p>",
      "rawMarkdown": "Inference code is cleaned up and published:\nhttps://www.kaggle.com/code/benbla/5th-place-solution/notebook"
    },
    {
      "id": 2486763,
      "postDate": "2023-10-18T06:49:46.503Z",
      "content": "<p>Congratulations on your achievement. I have some questions.</p>\n<ol>\n<li>Did you split the output words/letters into phonemes?</li>\n<li>If yes how did you get the phonemes from individual letters in Bengali?</li>\n</ol>\n<p>Apologies for any ignorant questions. 😀</p>\n<p>Thanks</p>",
      "rawMarkdown": "Congratulations on your achievement. I have some questions.\n\n1. Did you split the output words/letters into phonemes?\n2. If yes how did you get the phonemes from individual letters in Bengali?\n\nApologies for any ignorant questions. 😀\n\nThanks",
      "replies": [
        {
          "id": 2488297,
          "postDate": "2023-10-19T07:23:31.233Z",
          "content": "<p>You receive as raw output of a Wav2Vec2 model a matrix with the logits. The number of lines is proportional to the length of the audio file and the number of columns is proportional to the length of the vocabulary. If we assume greedy decoding as the simplest method, we use argmax to check for each line which character has the highest probability. This means that the output is not autoregressive at the character level. The CTC algorithm then handles pauses and consecutive letters and creates entire words. This means that in this case you have a string of words that can be further processed into sentences by adding punctuation marks.</p>\n<p>Therefore, there is no need to manually put the output together or to split the output again for further processing. I hope, it answered your question.</p>",
          "rawMarkdown": "You receive as raw output of a Wav2Vec2 model a matrix with the logits. The number of lines is proportional to the length of the audio file and the number of columns is proportional to the length of the vocabulary. If we assume greedy decoding as the simplest method, we use argmax to check for each line which character has the highest probability. This means that the output is not autoregressive at the character level. The CTC algorithm then handles pauses and consecutive letters and creates entire words. This means that in this case you have a string of words that can be further processed into sentences by adding punctuation marks.\n\nTherefore, there is no need to manually put the output together or to split the output again for further processing. I hope, it answered your question."
        }
      ]
    },
    {
      "id": 2493370,
      "postDate": "2023-10-23T11:05:40.367Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2488104,
      "author_name": "Fahim Shahriar Khan",
      "author_url": "",
      "post_date": "2023-10-19T03:27:27.620000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> on your amazing achievement. I had some questions regarding building a n-gram Kenlm language model to be used with Beam Search. </p>\n<p>Let's say I have trained a wav2vec2 STT model using a custom Bengali vocabulary set of 70 Bengali unicodes and special tokens which is different from the vocabulary (vocab.json) of the pre-trained wav2vec2 STT model like <a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali/tree/main\" target=\"_blank\">ai4bharat/indicwav2vec_v1_bengali</a>. How should I train a custom n-gram Kenlm model? As i have found that open source Kenlm models, for example: <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model</a> degrade the inference result when used with wav2vec2 STT models with custom vocabulary. Any suggestions regarding this problem? </p>\n<p>Apologies for any ignorant questions. </p>\n<p>Thanks. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2488288,
          "author_name": "Benedikt Droste",
          "author_url": "",
          "post_date": "2023-10-19T07:15:07.130000",
          "content": "<p>It is important to normalize the texts the same way you normalize the training data for the audio model. </p>\n<p>Therefore you should preprocess and normalize the texts before training your KenLM model. If you use the <strong>bnunicodenormalizer</strong> for the training annotations, you should use it for the language model text corpus as well. </p>\n<p>I think the problem with the linked KenLM model is not related to the custom vocabulary, but rather that the texts for the model were normalized differently or not at all.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2490186,
              "author_name": "Fahim Shahriar Khan",
              "author_url": "",
              "post_date": "2023-10-20T14:05:41.087000",
              "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a>. Could you share the steps or possibly any tutorials you followed to train your custom KenlM model?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2491219,
      "author_name": "Benedikt Droste",
      "author_url": "",
      "post_date": "2023-10-21T12:33:47.037000",
      "content": "<p>Inference code is cleaned up and published:<br>\n<a href=\"https://www.kaggle.com/code/benbla/5th-place-solution/notebook\" target=\"_blank\">https://www.kaggle.com/code/benbla/5th-place-solution/notebook</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2486763,
      "author_name": "Abhisek Dash",
      "author_url": "",
      "post_date": "2023-10-18T06:49:46.503000",
      "content": "<p>Congratulations on your achievement. I have some questions.</p>\n<ol>\n<li>Did you split the output words/letters into phonemes?</li>\n<li>If yes how did you get the phonemes from individual letters in Bengali?</li>\n</ol>\n<p>Apologies for any ignorant questions. 😀</p>\n<p>Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2488297,
          "author_name": "Benedikt Droste",
          "author_url": "",
          "post_date": "2023-10-19T07:23:31.233000",
          "content": "<p>You receive as raw output of a Wav2Vec2 model a matrix with the logits. The number of lines is proportional to the length of the audio file and the number of columns is proportional to the length of the vocabulary. If we assume greedy decoding as the simplest method, we use argmax to check for each line which character has the highest probability. This means that the output is not autoregressive at the character level. The CTC algorithm then handles pauses and consecutive letters and creates entire words. This means that in this case you have a string of words that can be further processed into sentences by adding punctuation marks.</p>\n<p>Therefore, there is no need to manually put the output together or to split the output again for further processing. I hope, it answered your question.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2493370,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-23T11:05:40.367000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486677": "Thank you for organizing this competition so well. Whenever questions arose, the hosts responded promptly and clearly. I really enjoyed working on this problem. Thank you very much @imtiazprio \n\n**Modeling**\nLike most of the participants, I started with the YellowKing pipeline. I changed the backbone to IndicWav2Vec and reinitialized the CTC layer. With this method and the public Dari/normalizing pipeline, it was already possible to achieve 0.48x scores without a Language Model. However, there was a big problem: **the local WER score did not match the public lb anymore if the model was trained for too long**. At a learning rate of 3e-5, local overfitting started at 40k steps (batch size 16). I suspect that the model started learning the relationship of erroneous audio/annotation pairs.\n\nDue to the noisy annotation, I excluded all samples with a mos-score of >2.0 in the subsequent training stage, reaching a score of 0.472. After re-labeling with the new model, I removed all samples with a WER above 0.5 for my model and a WER above 0.5 for YellowKing. **This resolved the local overfitting issue**, ensuring that local improvements correlated with better public lb scores. I trained the model over 210k steps (bs 16) with a learning rate of 8e-5, achieving a score of 0.452 without a language model on top.\n\nAdditionally, I trained a model with a [larger backbone](https://huggingface.co/facebook/wav2vec2-xls-r-1b). Unfortunately, due to lack of time, I could not fully exploit the potential, but after 135k steps I could already achieve a score of 0.454. The learning rate for the larger model had to be much smaller (1e-5).\n\n**Key Takeaways:**\n- Quality filtering of data significantly enhanced training performance.\n\n**Ensembling**\nI tried early to average the model logits of different finetuned models. However, it doesn´t work because the predictions may not be aligned. This is even the case for models of the same architecture. In [this paper](https://arxiv.org/pdf/2206.05518.pdf) a procedure is described, where **features of different ASR models are concatenated**, followed by the addition of a transformer encoder and a CTC layer. The weights of the asr models are freezed and only the transformer encoder and the CTC layer are trained. I applied this approach to my finetuned models:\n- Extract embeddings of the finetuned models\n- Concatenate the features\n- Concatenated features are passed into the Transformer encoder and then processed in the CTC layer\n- A few steps (7k steps for batchsize 8) on the same training data were enough to increase the performance of the overall model from 0.355 to 0.344 (I justed tested it in the full pipeline, so with language model, punctuation and normalizing)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fadb7e22f32c07808c16ed5ece66f5032%2FBild1.png?generation=1697605346189304&alt=media)\n\n**Key takeways:**\n- Ensembling asr model outputs by concatenating the last hidden states and adding a transformer encoder with CTC layer improves performance a lot\n\n**Potentials:**\n- I think this is where the **greatest further potential lies** in my solution. I didn't have enough time, but it would be interesting to see how Whisper or HUBERT in the ensemble would have improved the performance. I think because of the different architectures it should be even better because of the larger diversity in the ensemble.\n- My finetuned model with the 1b backbone is undertrained. There is still potential for better performance if the model is trained for a longer period.\n- I have hardly experimented with the parameters and training routine of the added transformer encoder. There is room for improvement as well.\n\n**Language Model**\nI have thought a lot about the language model part. I had thought about interpreting the whole thing as a seq2seq problem and taking a pretrained transformer as the language model. In the end I just used a standard KenLM model. For this I used datasets which were shared in the forums (e.g. IndicCorp, SLR, Comp data).\n\nIn preprocessing, I radically **limited to the character set of the training data**, as the[ hosts said ](https://www.kaggle.com/competitions/bengaliai-speech/discussion/433469#2410416) there are no new characters in the test data. I replaced numbers, other languages, and other characters with special tokens. This reduced the number of unigrams significantly which was a bottleneck in the Kaggle environment. I ended up with **a 0 0 0 2 2 pruned 5-ngram**. Adding more data didn´t increase the performance at a certain point.\n\nMuch more important for the Kaggle workflow was the following fact: If you use the standard Wav2Vec2ProcessorWithLM of Huggingface and then remove the model, the [RAM is not released](https://github.com/kensho-technologies/pyctcdecode/pull/111). I therefore disassembled the pipeline and first computed all logits, stored them locally, and then read them back in for the decoding process. Afterwards I removed the language model cleanly with decoder.cleanup() before deleting the decoder. Without this measure, I would not have been able to use the ensemble model because of out of memory errors.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fd077b0fb9e9042395af9401e0b509c65%2FBild2.png?generation=1697605762499198&alt=media)\n\n**Key Takeways:**\n- The language model is important: precise preprocessing manages size while excessive data may not necessarily improve the LM's performance. Corrupted data can even degrade the performance of the language model.\n\n**Punctuation**\nI tried different punctuation models and ensembles. In the end, I used a [XLM-Roberta-Large model](https://github.com/xashru/punctuation-restoration). It improved the score by ~0.02. I could improve the performance even more by forcing the model to output a | or ? after the last word (+0.002).\n\n**Code**\nCleaned up inference code can be found here:\n[Kaggle Inference](https://www.kaggle.com/code/benbla/5th-place-solution/notebook)\nCleaned up training code can be found here:\n[Github training](https://github.com/bd317/BengaliAI_Speech_Recognition_5th_solution)",
    "2488104": "Congratulations @benbla on your amazing achievement. I had some questions regarding building a n-gram Kenlm language model to be used with Beam Search. \n\nLet's say I have trained a wav2vec2 STT model using a custom Bengali vocabulary set of 70 Bengali unicodes and special tokens which is different from the vocabulary (vocab.json) of the pre-trained wav2vec2 STT model like [ai4bharat/indicwav2vec_v1_bengali](https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali/tree/main). How should I train a custom n-gram Kenlm model? As i have found that open source Kenlm models, for example: https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model degrade the inference result when used with wav2vec2 STT models with custom vocabulary. Any suggestions regarding this problem? \n\nApologies for any ignorant questions. \n\nThanks. ",
    "2491219": "Inference code is cleaned up and published:\nhttps://www.kaggle.com/code/benbla/5th-place-solution/notebook",
    "2486763": "Congratulations on your achievement. I have some questions.\n\n1. Did you split the output words/letters into phonemes?\n2. If yes how did you get the phonemes from individual letters in Bengali?\n\nApologies for any ignorant questions. 😀\n\nThanks",
    "2493370": ""
  }
}