{
  "id": 447965,
  "title": "14th Place Solution for the Bengali.AI Speech Recognition Competition",
  "url": "/competitions/bengaliai-speech/discussion/447965",
  "author_name": "neilus",
  "post_date": "2023-10-18T01:06:11.711000",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks to the organizers and Kaggle staff for holding the competition, and congratulations to the winners!</p>\n<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/overview\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/data</a></p>\n<h1>Overview of the Approach</h1>\n<p>My approach addressed two main challenges:</p>\n<p><strong>Challenge:</strong></p>\n<ol>\n<li>The need for robust speech recognition capable of handling diverse speakers.</li>\n<li>The requirement to restore punctuation in transcriptions.</li>\n</ol>\n<p><strong>Approach:</strong></p>\n<ol>\n<li>Fine-tuning the <code>indicwav2vec_v1_bengali</code> model using competition data.</li>\n<li>Leveraging the Punctuation Restoration tool from <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a>.</li>\n</ol>\n<p>This approach led to a Public Leaderboard score of 0.38.</p>\n<p>I began with the foundation provided by <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>’s <a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\" target=\"_blank\">notebook</a>.</p>\n<h1>Details of the submission</h1>\n<p><strong>Diverse Speaker Recognition:</strong><br>\nThe test audio data comes mostly from YouTube, which means that the speakers' identities are often unknown. To create a versatile model, I trained it on diverse audio data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F04724b5e0fab1e4b550271153025dcd3%2Fdomain.png?generation=1697591004781091&amp;alt=media\" alt=\"\"><br>\n(dataset paper: <a href=\"https://arxiv.org/abs/2305.09688\" target=\"_blank\">https://arxiv.org/abs/2305.09688</a> C.1.1. Data Scraping Roadmap &amp; Prerequisites)</p>\n<p><strong>Punctuation Restoration:</strong><br>\nIt's known that the labels are normalized, which implies that punctuation is preserved, as mentioned <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110\" target=\"_blank\">here</a>.<br>\nPredicting punctuation during transcription is difficult, and leaving it out would result in word errors.</p>\n<p>By restoring punctuation, Word Error Rate (WER) can be reduced:</p>\n<ul>\n<li>(label) hello. how are you?</li>\n<li>(predict) hello how are you<br>\n→ wer: 0.5</li>\n<li>(restore) hello. how are you.<br>\n→ wer: 0.25</li>\n</ul>\n<p>While the public notebook appends periods at the end of sentences, the test data has an average of 34.42 words per sample and a Macro Train/Validation set with averages of 8.42/9.21. This suggests that multiple sentences may exist in one audio file, making it necessary to restore punctuation at points other than sentence endings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F1aa92dc382aaefb5ee7797fdc42a5c94%2Fwps.png?generation=1697591051306745&amp;alt=media\" alt=\"\"><br>\n(dataset paper: <a href=\"https://arxiv.org/abs/2305.09688\" target=\"_blank\">https://arxiv.org/abs/2305.09688</a> Table 1: OOD-Speech Dataset Statistics)</p>\n<h2>Models</h2>\n<ol>\n<li>Wav2vec2CTC model<ul>\n<li>Fine-tuned Wav2vec2 <code>ai4bharat/indicwav2vec_v1_bengali</code></li></ul></li>\n<li>Language model<ul>\n<li>KenLM <code>arijitx/wav2vec2-xls-r-300m-bengali</code></li></ul></li>\n<li>Punctuation Restore model<ul>\n<li>XLM-RoBERTa-large from <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a></li></ul></li>\n</ol>\n<h2>Training</h2>\n<p><a href=\"https://github.com/Neilsaw/kaggle_Bengali.AI_ASR_16th_solution\" target=\"_blank\">Github</a></p>\n<p>Fine-tuning Wav2vec2CTC with Transformers involved choosing datasets based on Yellowking’s WER, CER, and MOS_PRED metrics, as outlined in this <a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">notebook</a>. Two dataset splits were used:</p>\n<ol>\n<li><strong>Easy Data:</strong> audio samples where inference was straightforward.<ul>\n<li>YKG WER &lt; 0.1</li></ul></li>\n<li><strong>Hard Data:</strong> audio samples where the character content was correct but the WER was high.<ul>\n<li>0.3 &lt; WER &lt; 1.5</li>\n<li>CER &lt; 0.15</li>\n<li>MOS_PRED &gt; 3</li></ul></li>\n</ol>\n<p>The first dataset helped the model adapt to a variety of voices, while the second dataset allowed it to handle audio with higher WER.</p>\n<p>Additionally, the inclusion of white noise during training led to a slight improvement in the Leaderboard score by 0.001.</p>\n<h2>Validate</h2>\n<p>I only used LB for Validate.<br>\nI couldn't rely on local cross-validation because the domain shift between the training data and the test data was too significant.</p>\n<h2>Inference</h2>\n<p><a href=\"https://www.kaggle.com/code/neilus/16th-solution/notebook\" target=\"_blank\">inference notebook</a></p>\n<p>improved things from public notebook</p>\n<ul>\n<li><p>Punctuation restoration (-0.023)</p>\n<ul>\n<li>for using <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">xashru/punctuation-restoration</a>, we need to change transformers==2.11.0.</li>\n<li>so after LM inference, pip install transformers==2.11.0 and execute punctuation-restoration  on command line for reset import packages.</li></ul></li>\n<li><p>using unigrams.txt for KenLM ( -0.005)</p></li>\n</ul>\n<pre><code> ( / , encoding=)  :\n    unigram_list = [t.()  t  f.().().()]\n\ndecoder = pyctcdecode.(\n    (sorted_vocab_dict.()),\n    ( / ),\n    unigram_list,\n)\n</code></pre>\n<ul>\n<li>beam width 1500 (-0.002 ~ -0.001)</li>\n</ul>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public</th>\n<th>Improvement</th>\n<th></th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.471</td>\n<td></td>\n<td></td>\n<td>0.564</td>\n</tr>\n<tr>\n<td>Training (Easy Data)</td>\n<td>0.425</td>\n<td>-0.046</td>\n<td></td>\n<td>0.508</td>\n</tr>\n<tr>\n<td>Beam width 1500</td>\n<td>0.423</td>\n<td>-0.002</td>\n<td></td>\n<td>0.506</td>\n</tr>\n<tr>\n<td>Punctuation restoration</td>\n<td>0.400</td>\n<td>-0.023</td>\n<td></td>\n<td>0.488</td>\n</tr>\n<tr>\n<td>Using unigrams.txt</td>\n<td>0.395</td>\n<td>-0.005</td>\n<td></td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>Training (Hard Data)</td>\n<td>0.380</td>\n<td>-0.015</td>\n<td></td>\n<td>0.458</td>\n</tr>\n</tbody>\n</table>\n<h2>Things didn't work for me</h2>\n<ul>\n<li>NER (Named Entity Recognition)<ul>\n<li>For restore \"-\" to NER, but using NER for ASR output sentences occur a lot of  False detection and LB down.</li></ul></li>\n<li>create own LM<ul>\n<li>since i only used MaCro Train data, may be too small sentence.</li></ul></li>\n<li>fine tuned punctuation model<ul>\n<li>same reason LM. only used MaCro Train data.</li></ul></li>\n<li>denoiser (<a href=\"https://github.com/facebookresearch/denoiser\" target=\"_blank\">https://github.com/facebookresearch/denoiser</a>)<ul>\n<li>LB and CV down. so I didn`t use.</li></ul></li>\n<li>add various noise while training<ul>\n<li>only White noise was work.</li></ul></li>\n</ul>\n<h1>Conclusion</h1>\n<p>At the beginning of the competition, I tried adding noise to adapt to different domains, but it didn't improve the LB. From this, I thought that there were other challenges to address besides noise.</p>\n<p>This competition challenged participants to achieve generalization and deal with label noise (punctuation). It required addressing the question of how much generalization is necessary and how to handle punctuation noise effectively. </p>\n<p>I am grateful for the opportunity to learn from this competition.</p>\n<h1>Source</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\" target=\"_blank\">https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110</a></li>\n<li><a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a></li>\n</ul>",
  "messages": [
    {
      "id": 2486486,
      "postDate": "2023-10-18T01:06:11.710Z",
      "content": "<p>Thanks to the organizers and Kaggle staff for holding the competition, and congratulations to the winners!</p>\n<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/overview\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/data</a></p>\n<h1>Overview of the Approach</h1>\n<p>My approach addressed two main challenges:</p>\n<p><strong>Challenge:</strong></p>\n<ol>\n<li>The need for robust speech recognition capable of handling diverse speakers.</li>\n<li>The requirement to restore punctuation in transcriptions.</li>\n</ol>\n<p><strong>Approach:</strong></p>\n<ol>\n<li>Fine-tuning the <code>indicwav2vec_v1_bengali</code> model using competition data.</li>\n<li>Leveraging the Punctuation Restoration tool from <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a>.</li>\n</ol>\n<p>This approach led to a Public Leaderboard score of 0.38.</p>\n<p>I began with the foundation provided by <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>’s <a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\" target=\"_blank\">notebook</a>.</p>\n<h1>Details of the submission</h1>\n<p><strong>Diverse Speaker Recognition:</strong><br>\nThe test audio data comes mostly from YouTube, which means that the speakers' identities are often unknown. To create a versatile model, I trained it on diverse audio data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F04724b5e0fab1e4b550271153025dcd3%2Fdomain.png?generation=1697591004781091&amp;alt=media\" alt=\"\"><br>\n(dataset paper: <a href=\"https://arxiv.org/abs/2305.09688\" target=\"_blank\">https://arxiv.org/abs/2305.09688</a> C.1.1. Data Scraping Roadmap &amp; Prerequisites)</p>\n<p><strong>Punctuation Restoration:</strong><br>\nIt's known that the labels are normalized, which implies that punctuation is preserved, as mentioned <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110\" target=\"_blank\">here</a>.<br>\nPredicting punctuation during transcription is difficult, and leaving it out would result in word errors.</p>\n<p>By restoring punctuation, Word Error Rate (WER) can be reduced:</p>\n<ul>\n<li>(label) hello. how are you?</li>\n<li>(predict) hello how are you<br>\n→ wer: 0.5</li>\n<li>(restore) hello. how are you.<br>\n→ wer: 0.25</li>\n</ul>\n<p>While the public notebook appends periods at the end of sentences, the test data has an average of 34.42 words per sample and a Macro Train/Validation set with averages of 8.42/9.21. This suggests that multiple sentences may exist in one audio file, making it necessary to restore punctuation at points other than sentence endings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F1aa92dc382aaefb5ee7797fdc42a5c94%2Fwps.png?generation=1697591051306745&amp;alt=media\" alt=\"\"><br>\n(dataset paper: <a href=\"https://arxiv.org/abs/2305.09688\" target=\"_blank\">https://arxiv.org/abs/2305.09688</a> Table 1: OOD-Speech Dataset Statistics)</p>\n<h2>Models</h2>\n<ol>\n<li>Wav2vec2CTC model<ul>\n<li>Fine-tuned Wav2vec2 <code>ai4bharat/indicwav2vec_v1_bengali</code></li></ul></li>\n<li>Language model<ul>\n<li>KenLM <code>arijitx/wav2vec2-xls-r-300m-bengali</code></li></ul></li>\n<li>Punctuation Restore model<ul>\n<li>XLM-RoBERTa-large from <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a></li></ul></li>\n</ol>\n<h2>Training</h2>\n<p><a href=\"https://github.com/Neilsaw/kaggle_Bengali.AI_ASR_16th_solution\" target=\"_blank\">Github</a></p>\n<p>Fine-tuning Wav2vec2CTC with Transformers involved choosing datasets based on Yellowking’s WER, CER, and MOS_PRED metrics, as outlined in this <a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">notebook</a>. Two dataset splits were used:</p>\n<ol>\n<li><strong>Easy Data:</strong> audio samples where inference was straightforward.<ul>\n<li>YKG WER &lt; 0.1</li></ul></li>\n<li><strong>Hard Data:</strong> audio samples where the character content was correct but the WER was high.<ul>\n<li>0.3 &lt; WER &lt; 1.5</li>\n<li>CER &lt; 0.15</li>\n<li>MOS_PRED &gt; 3</li></ul></li>\n</ol>\n<p>The first dataset helped the model adapt to a variety of voices, while the second dataset allowed it to handle audio with higher WER.</p>\n<p>Additionally, the inclusion of white noise during training led to a slight improvement in the Leaderboard score by 0.001.</p>\n<h2>Validate</h2>\n<p>I only used LB for Validate.<br>\nI couldn't rely on local cross-validation because the domain shift between the training data and the test data was too significant.</p>\n<h2>Inference</h2>\n<p><a href=\"https://www.kaggle.com/code/neilus/16th-solution/notebook\" target=\"_blank\">inference notebook</a></p>\n<p>improved things from public notebook</p>\n<ul>\n<li><p>Punctuation restoration (-0.023)</p>\n<ul>\n<li>for using <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">xashru/punctuation-restoration</a>, we need to change transformers==2.11.0.</li>\n<li>so after LM inference, pip install transformers==2.11.0 and execute punctuation-restoration  on command line for reset import packages.</li></ul></li>\n<li><p>using unigrams.txt for KenLM ( -0.005)</p></li>\n</ul>\n<pre><code> ( / , encoding=)  :\n    unigram_list = [t.()  t  f.().().()]\n\ndecoder = pyctcdecode.(\n    (sorted_vocab_dict.()),\n    ( / ),\n    unigram_list,\n)\n</code></pre>\n<ul>\n<li>beam width 1500 (-0.002 ~ -0.001)</li>\n</ul>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public</th>\n<th>Improvement</th>\n<th></th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.471</td>\n<td></td>\n<td></td>\n<td>0.564</td>\n</tr>\n<tr>\n<td>Training (Easy Data)</td>\n<td>0.425</td>\n<td>-0.046</td>\n<td></td>\n<td>0.508</td>\n</tr>\n<tr>\n<td>Beam width 1500</td>\n<td>0.423</td>\n<td>-0.002</td>\n<td></td>\n<td>0.506</td>\n</tr>\n<tr>\n<td>Punctuation restoration</td>\n<td>0.400</td>\n<td>-0.023</td>\n<td></td>\n<td>0.488</td>\n</tr>\n<tr>\n<td>Using unigrams.txt</td>\n<td>0.395</td>\n<td>-0.005</td>\n<td></td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>Training (Hard Data)</td>\n<td>0.380</td>\n<td>-0.015</td>\n<td></td>\n<td>0.458</td>\n</tr>\n</tbody>\n</table>\n<h2>Things didn't work for me</h2>\n<ul>\n<li>NER (Named Entity Recognition)<ul>\n<li>For restore \"-\" to NER, but using NER for ASR output sentences occur a lot of  False detection and LB down.</li></ul></li>\n<li>create own LM<ul>\n<li>since i only used MaCro Train data, may be too small sentence.</li></ul></li>\n<li>fine tuned punctuation model<ul>\n<li>same reason LM. only used MaCro Train data.</li></ul></li>\n<li>denoiser (<a href=\"https://github.com/facebookresearch/denoiser\" target=\"_blank\">https://github.com/facebookresearch/denoiser</a>)<ul>\n<li>LB and CV down. so I didn`t use.</li></ul></li>\n<li>add various noise while training<ul>\n<li>only White noise was work.</li></ul></li>\n</ul>\n<h1>Conclusion</h1>\n<p>At the beginning of the competition, I tried adding noise to adapt to different domains, but it didn't improve the LB. From this, I thought that there were other challenges to address besides noise.</p>\n<p>This competition challenged participants to achieve generalization and deal with label noise (punctuation). It required addressing the question of how much generalization is necessary and how to handle punctuation noise effectively. </p>\n<p>I am grateful for the opportunity to learn from this competition.</p>\n<h1>Source</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\" target=\"_blank\">https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110</a></li>\n<li><a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a></li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers and Kaggle staff for holding the competition, and congratulations to the winners!\n\n# Context\nBusiness context: https://www.kaggle.com/competitions/bengaliai-speech/overview\nData context: https://www.kaggle.com/competitions/bengaliai-speech/data\n\n# Overview of the Approach\n\nMy approach addressed two main challenges:\n\n**Challenge:**\n\n1. The need for robust speech recognition capable of handling diverse speakers.\n2. The requirement to restore punctuation in transcriptions.\n\n**Approach:**\n\n1. Fine-tuning the `indicwav2vec_v1_bengali` model using competition data.\n2. Leveraging the Punctuation Restoration tool from https://github.com/xashru/punctuation-restoration.\n\nThis approach led to a Public Leaderboard score of 0.38.\n\nI began with the foundation provided by @ttahara’s [notebook](https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline).\n\n# Details of the submission\n\n**Diverse Speaker Recognition:**\nThe test audio data comes mostly from YouTube, which means that the speakers' identities are often unknown. To create a versatile model, I trained it on diverse audio data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F04724b5e0fab1e4b550271153025dcd3%2Fdomain.png?generation=1697591004781091&alt=media)\n(dataset paper: https://arxiv.org/abs/2305.09688 C.1.1. Data Scraping Roadmap & Prerequisites)\n\n**Punctuation Restoration:**\nIt's known that the labels are normalized, which implies that punctuation is preserved, as mentioned [here](https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110).\nPredicting punctuation during transcription is difficult, and leaving it out would result in word errors.\n\nBy restoring punctuation, Word Error Rate (WER) can be reduced:\n\n- (label) hello. how are you?\n- (predict) hello how are you\n→ wer: 0.5\n- (restore) hello. how are you.\n→ wer: 0.25\n\nWhile the public notebook appends periods at the end of sentences, the test data has an average of 34.42 words per sample and a Macro Train/Validation set with averages of 8.42/9.21. This suggests that multiple sentences may exist in one audio file, making it necessary to restore punctuation at points other than sentence endings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F1aa92dc382aaefb5ee7797fdc42a5c94%2Fwps.png?generation=1697591051306745&alt=media)\n(dataset paper: https://arxiv.org/abs/2305.09688 Table 1: OOD-Speech Dataset Statistics)\n\n## Models\n\n1. Wav2vec2CTC model\n    - Fine-tuned Wav2vec2 `ai4bharat/indicwav2vec_v1_bengali`\n2. Language model\n    - KenLM `arijitx/wav2vec2-xls-r-300m-bengali`\n3. Punctuation Restore model\n    - XLM-RoBERTa-large from https://github.com/xashru/punctuation-restoration\n\n## Training\n[Github](https://github.com/Neilsaw/kaggle_Bengali.AI_ASR_16th_solution)\n\nFine-tuning Wav2vec2CTC with Transformers involved choosing datasets based on Yellowking’s WER, CER, and MOS_PRED metrics, as outlined in this [notebook](https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda). Two dataset splits were used:\n\n1. **Easy Data:** audio samples where inference was straightforward.\n    - YKG WER < 0.1\n2. **Hard Data:** audio samples where the character content was correct but the WER was high.\n    - 0.3 < WER < 1.5\n    - CER < 0.15\n    - MOS_PRED > 3\n\nThe first dataset helped the model adapt to a variety of voices, while the second dataset allowed it to handle audio with higher WER.\n\nAdditionally, the inclusion of white noise during training led to a slight improvement in the Leaderboard score by 0.001.\n\n## Validate\nI only used LB for Validate.\nI couldn't rely on local cross-validation because the domain shift between the training data and the test data was too significant.\n\n## Inference\n\n[inference notebook](https://www.kaggle.com/code/neilus/16th-solution/notebook)\n\nimproved things from public notebook\n\n- Punctuation restoration (-0.023)\n  - for using [xashru/punctuation-restoration](https://github.com/xashru/punctuation-restoration), we need to change transformers==2.11.0.\n  - so after LM inference, pip install transformers==2.11.0 and execute punctuation-restoration  on command line for reset import packages.\n\n- using unigrams.txt for KenLM ( -0.005)\n\n```jsx\nwith open(LM_PATH / \"unigrams.txt\", encoding=\"utf-8\") as f:\n    unigram_list = [t.lower() for t in f.read().strip().split(\"\\n\")]\n\ndecoder = pyctcdecode.build_ctcdecoder(\n    list(sorted_vocab_dict.keys()),\n    str(LM_PATH / \"5gram.bin\"),\n    unigram_list,\n)\n```\n\n- beam width 1500 (-0.002 ~ -0.001)\n\n\n## Results\n\n|  | Public| Improvement ||Private|\n| --- | --- | --- | --- | --- |\n| Baseline | 0.471 |  ||0.564|\n| Training (Easy Data) | 0.425 | -0.046 ||0.508|\n| Beam width 1500 | 0.423 | -0.002 ||0.506|\n| Punctuation restoration | 0.400 | -0.023 ||0.488|\n| Using unigrams.txt | 0.395 | -0.005 ||0.48|\n| Training (Hard Data) | 0.380 | -0.015 ||0.458|\n\n## Things didn't work for me \n- NER (Named Entity Recognition)\n  - For restore \"-\" to NER, but using NER for ASR output sentences occur a lot of  False detection and LB down.\n- create own LM\n  - since i only used MaCro Train data, may be too small sentence.\n- fine tuned punctuation model\n  - same reason LM. only used MaCro Train data.\n- denoiser (https://github.com/facebookresearch/denoiser)\n  - LB and CV down. so I didn`t use.\n- add various noise while training\n  - only White noise was work.\n\n# Conclusion\n\nAt the beginning of the competition, I tried adding noise to adapt to different domains, but it didn't improve the LB. From this, I thought that there were other challenges to address besides noise.\n\nThis competition challenged participants to achieve generalization and deal with label noise (punctuation). It required addressing the question of how much generalization is necessary and how to handle punctuation noise effectively. \n\nI am grateful for the opportunity to learn from this competition.\n\n# Source\n- https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\n- https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110\n- https://github.com/xashru/punctuation-restoration",
      "votes": 15
    },
    {
      "id": 2486493,
      "postDate": "2023-10-18T01:16:11.610Z",
      "content": "<p>I found good result after changing beam width 512 -&gt; 1024. But you got more improvement with 1500! Kudos to you for your achievement :) </p>",
      "rawMarkdown": "I found good result after changing beam width 512 -> 1024. But you got more improvement with 1500! Kudos to you for your achievement :) ",
      "votes": 1,
      "replies": [
        {
          "id": 2486792,
          "postDate": "2023-10-18T07:04:37.710Z",
          "content": "<p>Thank you!👍</p>",
          "rawMarkdown": "Thank you!👍"
        }
      ]
    },
    {
      "id": 2487171,
      "postDate": "2023-10-18T12:00:47.813Z",
      "content": "<p>Has xlm-roberta-large been fine-tuned?</p>",
      "rawMarkdown": "Has xlm-roberta-large been fine-tuned?",
      "replies": [
        {
          "id": 2487194,
          "postDate": "2023-10-18T12:26:08.367Z",
          "content": "<p>Hasn't. using pre-trained model from github repo.</p>\n<p>I tried fine-tune with MacRo Train sentence.<br>\nbut LB became worse (0.400 -&gt; 0.405) so I gave up.</p>",
          "rawMarkdown": "Hasn't. using pre-trained model from github repo.\n\nI tried fine-tune with MacRo Train sentence.\nbut LB became worse (0.400 -> 0.405) so I gave up."
        }
      ]
    },
    {
      "id": 2487854,
      "postDate": "2023-10-18T19:53:37.623Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2487168,
      "postDate": "2023-10-18T12:00:17.497Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2486493,
      "author_name": "AIFahim",
      "author_url": "",
      "post_date": "2023-10-18T01:16:11.610000",
      "content": "<p>I found good result after changing beam width 512 -&gt; 1024. But you got more improvement with 1500! Kudos to you for your achievement :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2486792,
          "author_name": "neilus",
          "author_url": "",
          "post_date": "2023-10-18T07:04:37.710000",
          "content": "<p>Thank you!👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2487171,
      "author_name": "answer-qzd2",
      "author_url": "",
      "post_date": "2023-10-18T12:00:47.813000",
      "content": "<p>Has xlm-roberta-large been fine-tuned?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2487194,
          "author_name": "neilus",
          "author_url": "",
          "post_date": "2023-10-18T12:26:08.367000",
          "content": "<p>Hasn't. using pre-trained model from github repo.</p>\n<p>I tried fine-tune with MacRo Train sentence.<br>\nbut LB became worse (0.400 -&gt; 0.405) so I gave up.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2487854,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:53:37.623000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2487168,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T12:00:17.497000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486486": "Thanks to the organizers and Kaggle staff for holding the competition, and congratulations to the winners!\n\n# Context\nBusiness context: https://www.kaggle.com/competitions/bengaliai-speech/overview\nData context: https://www.kaggle.com/competitions/bengaliai-speech/data\n\n# Overview of the Approach\n\nMy approach addressed two main challenges:\n\n**Challenge:**\n\n1. The need for robust speech recognition capable of handling diverse speakers.\n2. The requirement to restore punctuation in transcriptions.\n\n**Approach:**\n\n1. Fine-tuning the `indicwav2vec_v1_bengali` model using competition data.\n2. Leveraging the Punctuation Restoration tool from https://github.com/xashru/punctuation-restoration.\n\nThis approach led to a Public Leaderboard score of 0.38.\n\nI began with the foundation provided by @ttahara’s [notebook](https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline).\n\n# Details of the submission\n\n**Diverse Speaker Recognition:**\nThe test audio data comes mostly from YouTube, which means that the speakers' identities are often unknown. To create a versatile model, I trained it on diverse audio data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F04724b5e0fab1e4b550271153025dcd3%2Fdomain.png?generation=1697591004781091&alt=media)\n(dataset paper: https://arxiv.org/abs/2305.09688 C.1.1. Data Scraping Roadmap & Prerequisites)\n\n**Punctuation Restoration:**\nIt's known that the labels are normalized, which implies that punctuation is preserved, as mentioned [here](https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110).\nPredicting punctuation during transcription is difficult, and leaving it out would result in word errors.\n\nBy restoring punctuation, Word Error Rate (WER) can be reduced:\n\n- (label) hello. how are you?\n- (predict) hello how are you\n→ wer: 0.5\n- (restore) hello. how are you.\n→ wer: 0.25\n\nWhile the public notebook appends periods at the end of sentences, the test data has an average of 34.42 words per sample and a Macro Train/Validation set with averages of 8.42/9.21. This suggests that multiple sentences may exist in one audio file, making it necessary to restore punctuation at points other than sentence endings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F1aa92dc382aaefb5ee7797fdc42a5c94%2Fwps.png?generation=1697591051306745&alt=media)\n(dataset paper: https://arxiv.org/abs/2305.09688 Table 1: OOD-Speech Dataset Statistics)\n\n## Models\n\n1. Wav2vec2CTC model\n    - Fine-tuned Wav2vec2 `ai4bharat/indicwav2vec_v1_bengali`\n2. Language model\n    - KenLM `arijitx/wav2vec2-xls-r-300m-bengali`\n3. Punctuation Restore model\n    - XLM-RoBERTa-large from https://github.com/xashru/punctuation-restoration\n\n## Training\n[Github](https://github.com/Neilsaw/kaggle_Bengali.AI_ASR_16th_solution)\n\nFine-tuning Wav2vec2CTC with Transformers involved choosing datasets based on Yellowking’s WER, CER, and MOS_PRED metrics, as outlined in this [notebook](https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda). Two dataset splits were used:\n\n1. **Easy Data:** audio samples where inference was straightforward.\n    - YKG WER < 0.1\n2. **Hard Data:** audio samples where the character content was correct but the WER was high.\n    - 0.3 < WER < 1.5\n    - CER < 0.15\n    - MOS_PRED > 3\n\nThe first dataset helped the model adapt to a variety of voices, while the second dataset allowed it to handle audio with higher WER.\n\nAdditionally, the inclusion of white noise during training led to a slight improvement in the Leaderboard score by 0.001.\n\n## Validate\nI only used LB for Validate.\nI couldn't rely on local cross-validation because the domain shift between the training data and the test data was too significant.\n\n## Inference\n\n[inference notebook](https://www.kaggle.com/code/neilus/16th-solution/notebook)\n\nimproved things from public notebook\n\n- Punctuation restoration (-0.023)\n  - for using [xashru/punctuation-restoration](https://github.com/xashru/punctuation-restoration), we need to change transformers==2.11.0.\n  - so after LM inference, pip install transformers==2.11.0 and execute punctuation-restoration  on command line for reset import packages.\n\n- using unigrams.txt for KenLM ( -0.005)\n\n```jsx\nwith open(LM_PATH / \"unigrams.txt\", encoding=\"utf-8\") as f:\n    unigram_list = [t.lower() for t in f.read().strip().split(\"\\n\")]\n\ndecoder = pyctcdecode.build_ctcdecoder(\n    list(sorted_vocab_dict.keys()),\n    str(LM_PATH / \"5gram.bin\"),\n    unigram_list,\n)\n```\n\n- beam width 1500 (-0.002 ~ -0.001)\n\n\n## Results\n\n|  | Public| Improvement ||Private|\n| --- | --- | --- | --- | --- |\n| Baseline | 0.471 |  ||0.564|\n| Training (Easy Data) | 0.425 | -0.046 ||0.508|\n| Beam width 1500 | 0.423 | -0.002 ||0.506|\n| Punctuation restoration | 0.400 | -0.023 ||0.488|\n| Using unigrams.txt | 0.395 | -0.005 ||0.48|\n| Training (Hard Data) | 0.380 | -0.015 ||0.458|\n\n## Things didn't work for me \n- NER (Named Entity Recognition)\n  - For restore \"-\" to NER, but using NER for ASR output sentences occur a lot of  False detection and LB down.\n- create own LM\n  - since i only used MaCro Train data, may be too small sentence.\n- fine tuned punctuation model\n  - same reason LM. only used MaCro Train data.\n- denoiser (https://github.com/facebookresearch/denoiser)\n  - LB and CV down. so I didn`t use.\n- add various noise while training\n  - only White noise was work.\n\n# Conclusion\n\nAt the beginning of the competition, I tried adding noise to adapt to different domains, but it didn't improve the LB. From this, I thought that there were other challenges to address besides noise.\n\nThis competition challenged participants to achieve generalization and deal with label noise (punctuation). It required addressing the question of how much generalization is necessary and how to handle punctuation noise effectively. \n\nI am grateful for the opportunity to learn from this competition.\n\n# Source\n- https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\n- https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305#2400110\n- https://github.com/xashru/punctuation-restoration",
    "2486493": "I found good result after changing beam width 512 -> 1024. But you got more improvement with 1500! Kudos to you for your achievement :) ",
    "2487171": "Has xlm-roberta-large been fine-tuned?",
    "2487854": "",
    "2487168": ""
  }
}