{
  "id": 448343,
  "title": "23nd Solution - My first competition medal",
  "url": "/competitions/bengaliai-speech/writeups/hubert-23nd-solution-my-first-competition-medal",
  "author_name": "",
  "post_date": "2023-10-20T15:30:03.367Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I would like to express my gratitude to the organizers of this competition. As a relative newcomer to Automatic Speech Recognition (ASR), I'm thrilled to have secured a medal in this event. I am immensely satisfied with the outcome, and I extend my thanks to both the competition participants and its hosts.</p>\n<p><strong>The solution consists of 2 components:</strong></p>\n<ol>\n<li>ASR model</li>\n<li>Ngram KenLM model</li>\n</ol>\n<p><strong>1. ASR model</strong><br>\nInitially, I experimented with numerous Hugging Face (HF) models, training them on approximately 20,000 data samples. The models I tested included:</p>\n<ul>\n<li>LegolasTheElf/Wav2Vec2_XLSR_Bengali_1b</li>\n<li>arijitx/wav2vec2-large-xlsr-bengali</li>\n<li>arijitx/wav2vec2-xls-r-300m-bengali</li>\n<li>tanmoyio/wav2vec2-large-xlsr-bengali</li>\n<li>ai4bharat/indicwav2vec_v1_bengali</li>\n<li>Umong/wav2vec2-large-mms-1b-bengali</li>\n<li>kabir5297/Wav2Vec2-90k-Bengali</li>\n<li>tanmoyio/wav2vec2-large-xlsr-bengali</li>\n<li>bayartsogt/bengali-2023-0016</li>\n<li>shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice</li>\n<li>jonatasgrosman/wav2vec2-large-xlsr-53-english</li>\n<li>wav2vec2-xls-r-2b - training from scratch</li>\n<li>wav2vec2-xls-r-1b - training from scratch</li>\n</ul>\n<p>Eventually, I settled on the \"shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\" model. You can find all the fine-tuned models in my HF repository <a href=\"url\" target=\"_blank\">https://huggingface.co/Aspik101</a><br>\nThe final fine-tuned model can be accessed: <a href=\"url\" target=\"_blank\">https://huggingface.co/Aspik101/shahruk10_checkpoint-360_2444</a></p>\n<p>The final model underwent a two-step training process. In the first step, the model was trained on filtered data, where the Word Error Rate (WER) was less than 80, sourced from the yellowking model, MADASR dataset, and google/fleurs. This training was executed based on a streaming mode and took approximately 20 hours.</p>\n<pre><code>trainer = Trainer(\n        model=model,\n        data_collator=daata_collator,\n        args=training_args,\n        compute_metrics=compute_metrics,\n        train_dataset=IterableWrapper(train),\n        eval_dataset=IterableWrapper(val),\n        tokenizer=processor.feature_extractor)\n</code></pre>\n<p>Loading data via streaming on Hugging Face can be useful when processing large data sets or when you want to feed data to a model while it is running rather than loading the entire data set into memory at once.</p>\n<p>At this stage, the best result I managed to get was around 0.424 on the leaderboard</p>\n<p>In the next step, the model was trained using data from the competition but from the <strong>examples folder!</strong> Thanks to this, I managed to get results of <strong>0.404.</strong> with KenLM model.</p>\n<p><strong>2. Ngram KenLM model</strong><br>\n5-gram kenlm model trained on competition dataset external Bengali corpus IndicCorp V1+V2.</p>\n<p><strong>What doesn't work</strong></p>\n<ul>\n<li>Punctuation model <a href=\"url\" target=\"_blank\">https://huggingface.co/1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase</a> On the test data set, my calculations show that I should get about 0.02 WER, unfortunately it did not work on the final data set</li>\n<li>Attempting to use a large language model to correct typos, such as Qwen 7b, from <a href=\"url\" target=\"_blank\">https://huggingface.co/Qwen/Qwen-7B-Chat</a> showed potential but did not yield successful results in my case.</li>\n<li>I experimented with the \"indicparser\" as follows:</li>\n</ul>\n<pre><code> indicparser  graphemeParserg\ngp=graphemeParser()\n</code></pre>\n<ul>\n<li>Exploring the stacking of logits from various wav2vec models.</li>\n<li>Generating transcriptions using the Conformer model available <a href=\"url\" target=\"_blank\">https://huggingface.co/bengaliAI/BanglaConformer</a>, while considering only punctuation and combining it with the final wav2vec model.</li>\n<li>And many other ideas that I have already forgotten</li>\n</ul>\n<p>In conclusion, I want to reiterate my appreciation to the competition organizers and participants. This journey has been an incredible learning experience for me, and I'm excited to continue exploring the world of ASR.</p>",
  "messages": [
    {
      "id": "2488465",
      "postDate": "10/19/2023 09:16:45",
      "content": "<p>I would like to express my gratitude to the organizers of this competition. As a relative newcomer to Automatic Speech Recognition (ASR), I'm thrilled to have secured a medal in this event. I am immensely satisfied with the outcome, and I extend my thanks to both the competition participants and its hosts.</p>\n<p><strong>The solution consists of 2 components:</strong></p>\n<ol>\n<li>ASR model</li>\n<li>Ngram KenLM model</li>\n</ol>\n<p><strong>1. ASR model</strong><br>\nInitially, I experimented with numerous Hugging Face (HF) models, training them on approximately 20,000 data samples. The models I tested included:</p>\n<ul>\n<li>LegolasTheElf/Wav2Vec2_XLSR_Bengali_1b</li>\n<li>arijitx/wav2vec2-large-xlsr-bengali</li>\n<li>arijitx/wav2vec2-xls-r-300m-bengali</li>\n<li>tanmoyio/wav2vec2-large-xlsr-bengali</li>\n<li>ai4bharat/indicwav2vec_v1_bengali</li>\n<li>Umong/wav2vec2-large-mms-1b-bengali</li>\n<li>kabir5297/Wav2Vec2-90k-Bengali</li>\n<li>tanmoyio/wav2vec2-large-xlsr-bengali</li>\n<li>bayartsogt/bengali-2023-0016</li>\n<li>shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice</li>\n<li>jonatasgrosman/wav2vec2-large-xlsr-53-english</li>\n<li>wav2vec2-xls-r-2b - training from scratch</li>\n<li>wav2vec2-xls-r-1b - training from scratch</li>\n</ul>\n<p>Eventually, I settled on the \"shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\" model. You can find all the fine-tuned models in my HF repository <a href=\"url\" target=\"_blank\">https://huggingface.co/Aspik101</a><br>\nThe final fine-tuned model can be accessed: <a href=\"url\" target=\"_blank\">https://huggingface.co/Aspik101/shahruk10_checkpoint-360_2444</a></p>\n<p>The final model underwent a two-step training process. In the first step, the model was trained on filtered data, where the Word Error Rate (WER) was less than 80, sourced from the yellowking model, MADASR dataset, and google/fleurs. This training was executed based on a streaming mode and took approximately 20 hours.</p>\n<pre><code>trainer = Trainer(\n        model=model,\n        data_collator=daata_collator,\n        args=training_args,\n        compute_metrics=compute_metrics,\n        train_dataset=IterableWrapper(train),\n        eval_dataset=IterableWrapper(val),\n        tokenizer=processor.feature_extractor)\n</code></pre>\n<p>Loading data via streaming on Hugging Face can be useful when processing large data sets or when you want to feed data to a model while it is running rather than loading the entire data set into memory at once.</p>\n<p>At this stage, the best result I managed to get was around 0.424 on the leaderboard</p>\n<p>In the next step, the model was trained using data from the competition but from the <strong>examples folder!</strong> Thanks to this, I managed to get results of <strong>0.404.</strong> with KenLM model.</p>\n<p><strong>2. Ngram KenLM model</strong><br>\n5-gram kenlm model trained on competition dataset external Bengali corpus IndicCorp V1+V2.</p>\n<p><strong>What doesn't work</strong></p>\n<ul>\n<li>Punctuation model <a href=\"url\" target=\"_blank\">https://huggingface.co/1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase</a> On the test data set, my calculations show that I should get about 0.02 WER, unfortunately it did not work on the final data set</li>\n<li>Attempting to use a large language model to correct typos, such as Qwen 7b, from <a href=\"url\" target=\"_blank\">https://huggingface.co/Qwen/Qwen-7B-Chat</a> showed potential but did not yield successful results in my case.</li>\n<li>I experimented with the \"indicparser\" as follows:</li>\n</ul>\n<pre><code> indicparser  graphemeParserg\ngp=graphemeParser()\n</code></pre>\n<ul>\n<li>Exploring the stacking of logits from various wav2vec models.</li>\n<li>Generating transcriptions using the Conformer model available <a href=\"url\" target=\"_blank\">https://huggingface.co/bengaliAI/BanglaConformer</a>, while considering only punctuation and combining it with the final wav2vec model.</li>\n<li>And many other ideas that I have already forgotten</li>\n</ul>\n<p>In conclusion, I want to reiterate my appreciation to the competition organizers and participants. This journey has been an incredible learning experience for me, and I'm excited to continue exploring the world of ASR.</p>",
      "rawMarkdown": "I would like to express my gratitude to the organizers of this competition. As a relative newcomer to Automatic Speech Recognition (ASR), I'm thrilled to have secured a medal in this event. I am immensely satisfied with the outcome, and I extend my thanks to both the competition participants and its hosts.\n\n\n**The solution consists of 2 components:**\n\n1.  ASR model\n2. Ngram KenLM model\n\n\n**1. ASR model**\nInitially, I experimented with numerous Hugging Face (HF) models, training them on approximately 20,000 data samples. The models I tested included:\n- LegolasTheElf/Wav2Vec2_XLSR_Bengali_1b\n- arijitx/wav2vec2-large-xlsr-bengali\n- arijitx/wav2vec2-xls-r-300m-bengali\n- tanmoyio/wav2vec2-large-xlsr-bengali\n- ai4bharat/indicwav2vec_v1_bengali\n- Umong/wav2vec2-large-mms-1b-bengali\n- kabir5297/Wav2Vec2-90k-Bengali\n- tanmoyio/wav2vec2-large-xlsr-bengali\n- bayartsogt/bengali-2023-0016\n- shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\n- jonatasgrosman/wav2vec2-large-xlsr-53-english\n- wav2vec2-xls-r-2b - training from scratch\n- wav2vec2-xls-r-1b - training from scratch\n\nEventually, I settled on the \"shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\" model. You can find all the fine-tuned models in my HF repository [https://huggingface.co/Aspik101](url)\nThe final fine-tuned model can be accessed: [https://huggingface.co/Aspik101/shahruk10_checkpoint-360_2444](url)\n\nThe final model underwent a two-step training process. In the first step, the model was trained on filtered data, where the Word Error Rate (WER) was less than 80, sourced from the yellowking model, MADASR dataset, and google/fleurs. This training was executed based on a streaming mode and took approximately 20 hours.\n\n```python\ntrainer = Trainer(\n        model=model,\n        data_collator=daata_collator,\n        args=training_args,\n        compute_metrics=compute_metrics,\n        train_dataset=IterableWrapper(train),\n        eval_dataset=IterableWrapper(val),\n        tokenizer=processor.feature_extractor)\n```\nLoading data via streaming on Hugging Face can be useful when processing large data sets or when you want to feed data to a model while it is running rather than loading the entire data set into memory at once.\n\nAt this stage, the best result I managed to get was around 0.424 on the leaderboard\n\nIn the next step, the model was trained using data from the competition but from the **examples folder!** Thanks to this, I managed to get results of **0.404.** with KenLM model.\n\n\n**2. Ngram KenLM model**\n5-gram kenlm model trained on competition dataset external Bengali corpus IndicCorp V1+V2.\n\n\n**What doesn't work**\n- Punctuation model [https://huggingface.co/1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase](url) On the test data set, my calculations show that I should get about 0.02 WER, unfortunately it did not work on the final data set\n- Attempting to use a large language model to correct typos, such as Qwen 7b, from [https://huggingface.co/Qwen/Qwen-7B-Chat](url) showed potential but did not yield successful results in my case.\n- I experimented with the \"indicparser\" as follows:\n```python\nfrom indicparser import graphemeParserg\ngp=graphemeParser(\"bangla\")\n```\n- Exploring the stacking of logits from various wav2vec models.\n- Generating transcriptions using the Conformer model available [https://huggingface.co/bengaliAI/BanglaConformer](url), while considering only punctuation and combining it with the final wav2vec model.\n- And many other ideas that I have already forgotten\n\nIn conclusion, I want to reiterate my appreciation to the competition organizers and participants. This journey has been an incredible learning experience for me, and I'm excited to continue exploring the world of ASR.",
      "votes": null
    },
    {
      "id": "2488747",
      "postDate": "10/19/2023 13:13:37",
      "content": "<p>Hearty congratulations to you <a href=\"https://www.kaggle.com/hubert101\" target=\"_blank\">@hubert101</a> for the fantastic result! I am sure that this will be a start for many more medals for you!</p>",
      "rawMarkdown": "Hearty congratulations to you @hubert101 for the fantastic result! I am sure that this will be a start for many more medals for you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2488747,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/19/2023 13:13:37",
      "content": "<p>Hearty congratulations to you <a href=\"https://www.kaggle.com/hubert101\" target=\"_blank\">@hubert101</a> for the fantastic result! I am sure that this will be a start for many more medals for you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2488465": "I would like to express my gratitude to the organizers of this competition. As a relative newcomer to Automatic Speech Recognition (ASR), I'm thrilled to have secured a medal in this event. I am immensely satisfied with the outcome, and I extend my thanks to both the competition participants and its hosts.\n\n\n**The solution consists of 2 components:**\n\n1.  ASR model\n2. Ngram KenLM model\n\n\n**1. ASR model**\nInitially, I experimented with numerous Hugging Face (HF) models, training them on approximately 20,000 data samples. The models I tested included:\n- LegolasTheElf/Wav2Vec2_XLSR_Bengali_1b\n- arijitx/wav2vec2-large-xlsr-bengali\n- arijitx/wav2vec2-xls-r-300m-bengali\n- tanmoyio/wav2vec2-large-xlsr-bengali\n- ai4bharat/indicwav2vec_v1_bengali\n- Umong/wav2vec2-large-mms-1b-bengali\n- kabir5297/Wav2Vec2-90k-Bengali\n- tanmoyio/wav2vec2-large-xlsr-bengali\n- bayartsogt/bengali-2023-0016\n- shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\n- jonatasgrosman/wav2vec2-large-xlsr-53-english\n- wav2vec2-xls-r-2b - training from scratch\n- wav2vec2-xls-r-1b - training from scratch\n\nEventually, I settled on the \"shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\" model. You can find all the fine-tuned models in my HF repository [https://huggingface.co/Aspik101](url)\nThe final fine-tuned model can be accessed: [https://huggingface.co/Aspik101/shahruk10_checkpoint-360_2444](url)\n\nThe final model underwent a two-step training process. In the first step, the model was trained on filtered data, where the Word Error Rate (WER) was less than 80, sourced from the yellowking model, MADASR dataset, and google/fleurs. This training was executed based on a streaming mode and took approximately 20 hours.\n\n```python\ntrainer = Trainer(\n        model=model,\n        data_collator=daata_collator,\n        args=training_args,\n        compute_metrics=compute_metrics,\n        train_dataset=IterableWrapper(train),\n        eval_dataset=IterableWrapper(val),\n        tokenizer=processor.feature_extractor)\n```\nLoading data via streaming on Hugging Face can be useful when processing large data sets or when you want to feed data to a model while it is running rather than loading the entire data set into memory at once.\n\nAt this stage, the best result I managed to get was around 0.424 on the leaderboard\n\nIn the next step, the model was trained using data from the competition but from the **examples folder!** Thanks to this, I managed to get results of **0.404.** with KenLM model.\n\n\n**2. Ngram KenLM model**\n5-gram kenlm model trained on competition dataset external Bengali corpus IndicCorp V1+V2.\n\n\n**What doesn't work**\n- Punctuation model [https://huggingface.co/1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase](url) On the test data set, my calculations show that I should get about 0.02 WER, unfortunately it did not work on the final data set\n- Attempting to use a large language model to correct typos, such as Qwen 7b, from [https://huggingface.co/Qwen/Qwen-7B-Chat](url) showed potential but did not yield successful results in my case.\n- I experimented with the \"indicparser\" as follows:\n```python\nfrom indicparser import graphemeParserg\ngp=graphemeParser(\"bangla\")\n```\n- Exploring the stacking of logits from various wav2vec models.\n- Generating transcriptions using the Conformer model available [https://huggingface.co/bengaliAI/BanglaConformer](url), while considering only punctuation and combining it with the final wav2vec model.\n- And many other ideas that I have already forgotten\n\nIn conclusion, I want to reiterate my appreciation to the competition organizers and participants. This journey has been an incredible learning experience for me, and I'm excited to continue exploring the world of ASR.",
    "2488747": "Hearty congratulations to you @hubert101 for the fantastic result! I am sure that this will be a start for many more medals for you!"
  },
  "source": "meta"
}