{
  "id": 448119,
  "title": "24th Place Solution for the Bengali.AI Speech Recognition Competition",
  "url": "/competitions/bengaliai-speech/discussion/448119",
  "author_name": "Mutian Hong",
  "post_date": "2023-10-18T14:48:39.845000",
  "votes": 10,
  "comment_count": 8,
  "views": 0,
  "content": "<h2>First of all, we want to thank the organizers who held this wonderful competition. This is my first kaggle competition, I fell satisfied with the final outcome. Thanks my teammate <a href=\"https://www.kaggle.com/nisshokuitsuki\" target=\"_blank\">@nisshokuitsuki</a>  , who worked with me during this three months journel. Thanks everyone who contributed to the discussion and the notebooks, your works gave us a lot insperiations. </h2>\n<h1>Model</h1>\n<p>The model we used is</p>\n<pre><code>\n</code></pre>\n<p>We finetuned it further on this model <a href=\"https://www.kaggle.com/datasets/nischaydnk/bengali-wav2vec2-finetuned\" target=\"_blank\">bengali_wav2vec2_finetuned (kaggle.com)</a><br>\nSince the data quality of the competition dataset is low, we used the train split of <a href=\"https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv\" target=\"_blank\">Common Voice 13 | Bengali (Normalized) (kaggle.com)</a>to train. We trained about 10 epochs with the 20000 data in the dataset, and splited about 700 data for validation. <br>\nThe data here is normalized and removed punctuations.<br>\nHere is the training arguments:</p>\n<pre><code>training_args = TrainingArguments(\n&amp;nbsp; &amp;nbsp; group_by_length=,\n&amp;nbsp; &amp;nbsp; weight_decay=,\n&amp;nbsp; &amp;nbsp; num_train_epochs=,\n&amp;nbsp; &amp;nbsp; fp16=,\n&amp;nbsp; &amp;nbsp; learning_rate=,\n&amp;nbsp; &amp;nbsp; warmup_steps=,\n)\n</code></pre>\n<p>And we used cosine optimizer.</p>\n<p>The best model got local wer 0.15, and Improved the Public Score from <em>0.445-&gt;0.434</em></p>\n<p>However, my teammate also trained a model with the same data for just 60 steps( different args), and got the same score. (even 0.001 better on the private LB). How interesting and confusing.</p>\n<h1>Language Model</h1>\n<p>We trained a 6gram with <a href=\"https://www.kaggle.com/datasets/mbmmurad/lm-no-punc\" target=\"_blank\">lm_no_punc (kaggle.com)</a>. <br>\nMention that there is a error in the LM provided in <a href=\"https://www.kaggle.com/datasets/sameen53/yellowking-dlsprint-model\" target=\"_blank\">YellowKing_DLSprint_Model (kaggle.com)</a>. </p>\n<pre><code>There is no terminator   arpa . So we need    &lt;/s&gt;   .\n</code></pre>\n<p>Thanks to this notebook <a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\" target=\"_blank\">Build an n-gram with KenLM | MaCro | Kaggle</a>, we are able to realize this. <br>\nAdding the  improved the LB from <em>0.445-&gt;0.422</em><br>\nAnd building an 6gram with <a href=\"https://www.kaggle.com/datasets/mbmmurad/lm-no-punc\" target=\"_blank\">lm_no_punc (kaggle.com)</a>.Improved about <em>0.001</em></p>\n<h1>Punctuation Restoration</h1>\n<p>Punctuation really matters. Thanks to this post <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305\" target=\"_blank\">Bengali.AI Speech Recognition | Kaggle</a>, we are able to realize this. And an response under this post showed us a way to restore the punctuation:<br>\n<a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">xashru/punctuation-restoration: Punctuation Restoration using Transformer Models for High-and Low-Resource Languages (github.com)</a> First we trained a model with the dataset provided in this repositorie. It can restore 3 punctuations : </p>\n<pre><code>{: , : , : }\n</code></pre>\n<p>This improved the LB from <em>0.422-&gt;0.400</em><br>\nCombining with the finetuned model, we have <em>0.400-&gt;0.397</em><br>\nAfterwards, we thought that 3 punctuations might be not enouth. So we made a dataset with <a href=\"https://huggingface.co/datasets/oscar\" target=\"_blank\">oscar · Datasets at Hugging Face</a>, filterd datas that have only have bengali words. We chose 7<br>\npunctuations: </p>\n<pre><code>{: , : , : , : , : , : , : }\n</code></pre>\n<p>We trained 6 epochs with the default pharams.<br>\nAnd we have <em>0.393-&gt;0.387</em></p>\n<h1>Model ensemble</h1>\n<p>We simply ensembled our model like this: </p>\n<pre><code>            y = model_1(x).logits* + model_2(x).logits* + model_3(x).logits*\n</code></pre>\n<p>Wait, This works???<br>\nYes, thouth the predictions may not be aligned, But since the three models are trained on same datasets, the no-aligning problem is paritially solved. <br>\nThis improved our performance about <em>0.001</em></p>\n<h1>Decoder pharams selection</h1>\n<p>There are three main pharams for the decoder:</p>\n<pre><code>alpha: weight  language model during shallow fusion\nbeta: weight   score adjustment  during scoring\nbeam_width: determines    candidate output sequences retained    step.\n</code></pre>\n<p>To find the best pharams, we used optuna <a href=\"https://www.kaggle.com/code/snnclsr/0-444-optimize-decoding-parameters-with-optuna\" target=\"_blank\">[0.444] Optimize Decoding Parameters with Optuna | Kaggle</a> to search the best pharams. We searched the pharams with the example datas in the dataset, which is ood data, brought us better LB score.<br>\nThe final decoder pharams are:</p>\n<pre><code>{'alpha': , 'beta': , 'beam_width': }\n</code></pre>\n<h1>What doesn't work for us</h1>\n<ul>\n<li>Data augmentation. We added background noise downloaded from <a href=\"https://pixabay.com/sound-effects/search/noise/\" target=\"_blank\">https://pixabay.com/sound-effects/search/noise/</a> and also pitch shift , time stretch etc. But The LB got worse (<em>0.397-&gt;0.415</em>). Every experiment of data augmentation takes too much time and I fells to tired to do more experiments. Maybe I could write some codes to do it automatically</li>\n<li>Denoise model. We tried three denoising models: <br>\n    UVR: Notebook Run Out of time<br>\n    <a href=\"https://github.com/facebookresearch/denoiser\" target=\"_blank\">facebookresearch/denoiser</a>: decreased about 0.01<br>\n    <a href=\"https://github.com/NVIDIA/CleanUNet/blob/main/exp/DNS-large-high/checkpoint/pretrained.pkl\" target=\"_blank\">CleanUNet</a>: decreased about 0.005<br>\nWe thought that denoising harms the features and makes some short syllables unrecornizable.</li>\n<li>Train a bigger LM with more data. We used 15G normalized benglai data to build an kenlm, the score got worse. We still haven't found the cause of the problem. </li>\n</ul>",
  "messages": [
    {
      "id": 2487380,
      "postDate": "2023-10-18T14:48:39.847Z",
      "content": "<h2>First of all, we want to thank the organizers who held this wonderful competition. This is my first kaggle competition, I fell satisfied with the final outcome. Thanks my teammate <a href=\"https://www.kaggle.com/nisshokuitsuki\" target=\"_blank\">@nisshokuitsuki</a>  , who worked with me during this three months journel. Thanks everyone who contributed to the discussion and the notebooks, your works gave us a lot insperiations. </h2>\n<h1>Model</h1>\n<p>The model we used is</p>\n<pre><code>\n</code></pre>\n<p>We finetuned it further on this model <a href=\"https://www.kaggle.com/datasets/nischaydnk/bengali-wav2vec2-finetuned\" target=\"_blank\">bengali_wav2vec2_finetuned (kaggle.com)</a><br>\nSince the data quality of the competition dataset is low, we used the train split of <a href=\"https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv\" target=\"_blank\">Common Voice 13 | Bengali (Normalized) (kaggle.com)</a>to train. We trained about 10 epochs with the 20000 data in the dataset, and splited about 700 data for validation. <br>\nThe data here is normalized and removed punctuations.<br>\nHere is the training arguments:</p>\n<pre><code>training_args = TrainingArguments(\n&amp;nbsp; &amp;nbsp; group_by_length=,\n&amp;nbsp; &amp;nbsp; weight_decay=,\n&amp;nbsp; &amp;nbsp; num_train_epochs=,\n&amp;nbsp; &amp;nbsp; fp16=,\n&amp;nbsp; &amp;nbsp; learning_rate=,\n&amp;nbsp; &amp;nbsp; warmup_steps=,\n)\n</code></pre>\n<p>And we used cosine optimizer.</p>\n<p>The best model got local wer 0.15, and Improved the Public Score from <em>0.445-&gt;0.434</em></p>\n<p>However, my teammate also trained a model with the same data for just 60 steps( different args), and got the same score. (even 0.001 better on the private LB). How interesting and confusing.</p>\n<h1>Language Model</h1>\n<p>We trained a 6gram with <a href=\"https://www.kaggle.com/datasets/mbmmurad/lm-no-punc\" target=\"_blank\">lm_no_punc (kaggle.com)</a>. <br>\nMention that there is a error in the LM provided in <a href=\"https://www.kaggle.com/datasets/sameen53/yellowking-dlsprint-model\" target=\"_blank\">YellowKing_DLSprint_Model (kaggle.com)</a>. </p>\n<pre><code>There is no terminator   arpa . So we need    &lt;/s&gt;   .\n</code></pre>\n<p>Thanks to this notebook <a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\" target=\"_blank\">Build an n-gram with KenLM | MaCro | Kaggle</a>, we are able to realize this. <br>\nAdding the  improved the LB from <em>0.445-&gt;0.422</em><br>\nAnd building an 6gram with <a href=\"https://www.kaggle.com/datasets/mbmmurad/lm-no-punc\" target=\"_blank\">lm_no_punc (kaggle.com)</a>.Improved about <em>0.001</em></p>\n<h1>Punctuation Restoration</h1>\n<p>Punctuation really matters. Thanks to this post <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305\" target=\"_blank\">Bengali.AI Speech Recognition | Kaggle</a>, we are able to realize this. And an response under this post showed us a way to restore the punctuation:<br>\n<a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">xashru/punctuation-restoration: Punctuation Restoration using Transformer Models for High-and Low-Resource Languages (github.com)</a> First we trained a model with the dataset provided in this repositorie. It can restore 3 punctuations : </p>\n<pre><code>{: , : , : }\n</code></pre>\n<p>This improved the LB from <em>0.422-&gt;0.400</em><br>\nCombining with the finetuned model, we have <em>0.400-&gt;0.397</em><br>\nAfterwards, we thought that 3 punctuations might be not enouth. So we made a dataset with <a href=\"https://huggingface.co/datasets/oscar\" target=\"_blank\">oscar · Datasets at Hugging Face</a>, filterd datas that have only have bengali words. We chose 7<br>\npunctuations: </p>\n<pre><code>{: , : , : , : , : , : , : }\n</code></pre>\n<p>We trained 6 epochs with the default pharams.<br>\nAnd we have <em>0.393-&gt;0.387</em></p>\n<h1>Model ensemble</h1>\n<p>We simply ensembled our model like this: </p>\n<pre><code>            y = model_1(x).logits* + model_2(x).logits* + model_3(x).logits*\n</code></pre>\n<p>Wait, This works???<br>\nYes, thouth the predictions may not be aligned, But since the three models are trained on same datasets, the no-aligning problem is paritially solved. <br>\nThis improved our performance about <em>0.001</em></p>\n<h1>Decoder pharams selection</h1>\n<p>There are three main pharams for the decoder:</p>\n<pre><code>alpha: weight  language model during shallow fusion\nbeta: weight   score adjustment  during scoring\nbeam_width: determines    candidate output sequences retained    step.\n</code></pre>\n<p>To find the best pharams, we used optuna <a href=\"https://www.kaggle.com/code/snnclsr/0-444-optimize-decoding-parameters-with-optuna\" target=\"_blank\">[0.444] Optimize Decoding Parameters with Optuna | Kaggle</a> to search the best pharams. We searched the pharams with the example datas in the dataset, which is ood data, brought us better LB score.<br>\nThe final decoder pharams are:</p>\n<pre><code>{'alpha': , 'beta': , 'beam_width': }\n</code></pre>\n<h1>What doesn't work for us</h1>\n<ul>\n<li>Data augmentation. We added background noise downloaded from <a href=\"https://pixabay.com/sound-effects/search/noise/\" target=\"_blank\">https://pixabay.com/sound-effects/search/noise/</a> and also pitch shift , time stretch etc. But The LB got worse (<em>0.397-&gt;0.415</em>). Every experiment of data augmentation takes too much time and I fells to tired to do more experiments. Maybe I could write some codes to do it automatically</li>\n<li>Denoise model. We tried three denoising models: <br>\n    UVR: Notebook Run Out of time<br>\n    <a href=\"https://github.com/facebookresearch/denoiser\" target=\"_blank\">facebookresearch/denoiser</a>: decreased about 0.01<br>\n    <a href=\"https://github.com/NVIDIA/CleanUNet/blob/main/exp/DNS-large-high/checkpoint/pretrained.pkl\" target=\"_blank\">CleanUNet</a>: decreased about 0.005<br>\nWe thought that denoising harms the features and makes some short syllables unrecornizable.</li>\n<li>Train a bigger LM with more data. We used 15G normalized benglai data to build an kenlm, the score got worse. We still haven't found the cause of the problem. </li>\n</ul>",
      "rawMarkdown": "First of all, we want to thank the organizers who held this wonderful competition. This is my first kaggle competition, I fell satisfied with the final outcome. Thanks my teammate @nisshokuitsuki  , who worked with me during this three months journel. Thanks everyone who contributed to the discussion and the notebooks, your works gave us a lot insperiations. \n---\n# Model\nThe model we used is\n```\nWav2Vec2ForCTC\n```\nWe finetuned it further on this model [bengali_wav2vec2_finetuned (kaggle.com)](https://www.kaggle.com/datasets/nischaydnk/bengali-wav2vec2-finetuned)\nSince the data quality of the competition dataset is low, we used the train split of [Common Voice 13 | Bengali (Normalized) (kaggle.com)](https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv)to train. We trained about 10 epochs with the 20000 data in the dataset, and splited about 700 data for validation. \nThe data here is normalized and removed punctuations.\nHere is the training arguments:\n```python\ntraining_args = TrainingArguments(\n    group_by_length=False,\n    weight_decay=0.01,\n    num_train_epochs=10,\n    fp16=True,\n    learning_rate=4e-5,\n    warmup_steps=600,\n)\n```\nAnd we used cosine optimizer.\n\nThe best model got local wer 0.15, and Improved the Public Score from *0.445->0.434*\n\nHowever, my teammate also trained a model with the same data for just 60 steps( different args), and got the same score. (even 0.001 better on the private LB). How interesting and confusing.\n\n# Language Model\nWe trained a 6gram with [lm_no_punc (kaggle.com)](https://www.kaggle.com/datasets/mbmmurad/lm-no-punc). \nMention that there is a error in the LM provided in [YellowKing_DLSprint_Model (kaggle.com)](https://www.kaggle.com/datasets/sameen53/yellowking-dlsprint-model). \n```\nThere is no terminator in the arpa file. So we need to add an </s> into the file.\n```\nThanks to this notebook [Build an n-gram with KenLM | MaCro | Kaggle](https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro), we are able to realize this. \nAdding the </s> improved the LB from *0.445->0.422*\nAnd building an 6gram with [lm_no_punc (kaggle.com)](https://www.kaggle.com/datasets/mbmmurad/lm-no-punc).Improved about *0.001*\n# Punctuation Restoration\nPunctuation really matters. Thanks to this post [Bengali.AI Speech Recognition | Kaggle](https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305), we are able to realize this. And an response under this post showed us a way to restore the punctuation:\n[xashru/punctuation-restoration: Punctuation Restoration using Transformer Models for High-and Low-Resource Languages (github.com)](https://github.com/xashru/punctuation-restoration) First we trained a model with the dataset provided in this repositorie. It can restore 3 punctuations : \n```python\n{1: ',', 2: '।', 3: '?'}\n```\n\nThis improved the LB from *0.422->0.400*\nCombining with the finetuned model, we have *0.400->0.397*\nAfterwards, we thought that 3 punctuations might be not enouth. So we made a dataset with [oscar · Datasets at Hugging Face](https://huggingface.co/datasets/oscar), filterd datas that have only have bengali words. We chose 7\npunctuations: \n```python\n{1: ',', 2: '।', 3: '?', 4: '!', 5: '-', 6: '\"', 7: ':'}\n```\nWe trained 6 epochs with the default pharams.\nAnd we have *0.393->0.387*\n\n# Model ensemble\nWe simply ensembled our model like this: \n```python\n            y = model_1(x).logits*0.7 + model_2(x).logits*0.2 + model_3(x).logits*0.1\n```\nWait, This works???\nYes, thouth the predictions may not be aligned, But since the three models are trained on same datasets, the no-aligning problem is paritially solved. \nThis improved our performance about *0.001*\n\n# Decoder pharams selection\nThere are three main pharams for the decoder:\n```\nalpha: weight for language model during shallow fusion\nbeta: weight for length score adjustment of during scoring\nbeam_width: determines the number of candidate output sequences retained at each time step.\n```\nTo find the best pharams, we used optuna [[0.444] Optimize Decoding Parameters with Optuna | Kaggle](https://www.kaggle.com/code/snnclsr/0-444-optimize-decoding-parameters-with-optuna) to search the best pharams. We searched the pharams with the example datas in the dataset, which is ood data, brought us better LB score.\nThe final decoder pharams are:\n```\n{'alpha': 0.46570704474381447, 'beta': 0.8635977171858652, 'beam_width': 768}\n```\n\n\n# What doesn't work for us\n- Data augmentation. We added background noise downloaded from https://pixabay.com/sound-effects/search/noise/ and also pitch shift , time stretch etc. But The LB got worse (*0.397->0.415*). Every experiment of data augmentation takes too much time and I fells to tired to do more experiments. Maybe I could write some codes to do it automatically\n- Denoise model. We tried three denoising models: \n\t\tUVR: Notebook Run Out of time\n\t\t[facebookresearch/denoiser](https://github.com/facebookresearch/denoiser): decreased about 0.01\n\t\t[CleanUNet](https://github.com/NVIDIA/CleanUNet/blob/main/exp/DNS-large-high/checkpoint/pretrained.pkl): decreased about 0.005\n\tWe thought that denoising harms the features and makes some short syllables unrecornizable.\n- Train a bigger LM with more data. We used 15G normalized benglai data to build an kenlm, the score got worse. We still haven't found the cause of the problem. \n",
      "votes": 10
    },
    {
      "id": 2487449,
      "postDate": "2023-10-18T15:43:27.847Z",
      "content": "<p>I can confirm that way of ensembling somehow works :)<br>\nThank you for sharing</p>",
      "rawMarkdown": "I can confirm that way of ensembling somehow works :)\nThank you for sharing",
      "votes": 2,
      "replies": [
        {
          "id": 2488155,
          "postDate": "2023-10-19T04:30:29.533Z",
          "content": "<p>Did you used DTW to align logits? You said before in previous ensemble discussion thread?</p>",
          "rawMarkdown": "Did you used DTW to align logits? You said before in previous ensemble discussion thread?",
          "replies": [
            {
              "id": 2488891,
              "postDate": "2023-10-19T15:10:19.800Z",
              "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> Sadly no, I didn't have time. I found a publication that supports it though.  If you're interested, I can point you to this paper: <a href=\"https://ieeexplore.ieee.org/document/9687953\" target=\"_blank\">https://ieeexplore.ieee.org/document/9687953</a>  <br>\n\"Warped Ensembles: A Novel Technique for Improving CTC Based End-to-End Speech Recognition\" <br>\nThey mentioned the use of DTW there.  </p>",
              "rawMarkdown": "@aifahim Sadly no, I didn't have time. I found a publication that supports it though.  If you're interested, I can point you to this paper: https://ieeexplore.ieee.org/document/9687953  \n\"Warped Ensembles: A Novel Technique for Improving CTC Based End-to-End Speech Recognition\" \nThey mentioned the use of DTW there.  ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3106246,
      "postDate": "2025-01-24T15:33:31.577Z",
      "content": "<p>牛但是得凑够十个字才能发送</p>",
      "rawMarkdown": "牛但是得凑够十个字才能发送"
    },
    {
      "id": 2488159,
      "postDate": "2023-10-19T04:35:46.877Z",
      "content": "<p>Congratulation for your great achievement. Still didn't get you ensemble part! Can you share your pseudocode how you do this? </p>",
      "rawMarkdown": "Congratulation for your great achievement. Still didn't get you ensemble part! Can you share your pseudocode how you do this? ",
      "replies": [
        {
          "id": 2488514,
          "postDate": "2023-10-19T09:47:32.540Z",
          "content": "<p>Yes, here is our main inference code:</p>\n<pre><code> torch.no_grad():\n     batch  tqdm(test_loader):\n        x = batch[]\n        x = x.to(device, non_blocking=)\n\n         torch.cuda.amp.autocast():\n            y = model_1(x).logits* + model_2(x).logits* + model_3(x).logits*\n        y = y.detach().cpu().numpy()\n\n         l  y:  \n            sentence = processor_with_lm.decode(l, **best_params).text\n            pred_sentence_list.append(sentence)\n</code></pre>\n<p>However, it is the wrong way to do the ensemble. Because the format of the logits of wav2vec2 model is not aligned. There are other ways to do this, like (<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/448074#2487921\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/448074#2487921</a>) this solution.</p>",
          "rawMarkdown": "Yes, here is our main inference code:\n```python\nwith torch.no_grad():\n    for batch in tqdm(test_loader):\n        x = batch[\"input_values\"]\n        x = x.to(device, non_blocking=True)\n#         x = x.squeeze(0)\n        with torch.cuda.amp.autocast(True):\n            y = model_1(x).logits*0.7 + model_2(x).logits*0.2 + model_3(x).logits*0.1\n        y = y.detach().cpu().numpy()\n        \n        for l in y:  \n            sentence = processor_with_lm.decode(l, **best_params).text\n            pred_sentence_list.append(sentence)\n```\nHowever, it is the wrong way to do the ensemble. Because the format of the logits of wav2vec2 model is not aligned. There are other ways to do this, like (https://www.kaggle.com/competitions/bengaliai-speech/discussion/448074#2487921) this solution.",
          "votes": 1,
          "replies": [
            {
              "id": 2488980,
              "postDate": "2023-10-19T15:58:58.650Z",
              "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/hongori\" target=\"_blank\">@hongori</a> 😃</p>",
              "rawMarkdown": "Thank you so much @hongori 😃"
            }
          ]
        }
      ]
    },
    {
      "id": 2487858,
      "postDate": "2023-10-18T19:57:10.077Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2487449,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2023-10-18T15:43:27.847000",
      "content": "<p>I can confirm that way of ensembling somehow works :)<br>\nThank you for sharing</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2488155,
          "author_name": "AIFahim",
          "author_url": "",
          "post_date": "2023-10-19T04:30:29.533000",
          "content": "<p>Did you used DTW to align logits? You said before in previous ensemble discussion thread?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2488891,
              "author_name": "yukiya",
              "author_url": "",
              "post_date": "2023-10-19T15:10:19.800000",
              "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> Sadly no, I didn't have time. I found a publication that supports it though.  If you're interested, I can point you to this paper: <a href=\"https://ieeexplore.ieee.org/document/9687953\" target=\"_blank\">https://ieeexplore.ieee.org/document/9687953</a>  <br>\n\"Warped Ensembles: A Novel Technique for Improving CTC Based End-to-End Speech Recognition\" <br>\nThey mentioned the use of DTW there.  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3106246,
      "author_name": "Donald Armstrong",
      "author_url": "",
      "post_date": "2025-01-24T15:33:31.577000",
      "content": "<p>牛但是得凑够十个字才能发送</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2488159,
      "author_name": "AIFahim",
      "author_url": "",
      "post_date": "2023-10-19T04:35:46.877000",
      "content": "<p>Congratulation for your great achievement. Still didn't get you ensemble part! Can you share your pseudocode how you do this? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2488514,
          "author_name": "Mutian Hong",
          "author_url": "",
          "post_date": "2023-10-19T09:47:32.540000",
          "content": "<p>Yes, here is our main inference code:</p>\n<pre><code> torch.no_grad():\n     batch  tqdm(test_loader):\n        x = batch[]\n        x = x.to(device, non_blocking=)\n\n         torch.cuda.amp.autocast():\n            y = model_1(x).logits* + model_2(x).logits* + model_3(x).logits*\n        y = y.detach().cpu().numpy()\n\n         l  y:  \n            sentence = processor_with_lm.decode(l, **best_params).text\n            pred_sentence_list.append(sentence)\n</code></pre>\n<p>However, it is the wrong way to do the ensemble. Because the format of the logits of wav2vec2 model is not aligned. There are other ways to do this, like (<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/448074#2487921\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/448074#2487921</a>) this solution.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2488980,
              "author_name": "AIFahim",
              "author_url": "",
              "post_date": "2023-10-19T15:58:58.650000",
              "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/hongori\" target=\"_blank\">@hongori</a> 😃</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2487858,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:57:10.077000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2487380": "First of all, we want to thank the organizers who held this wonderful competition. This is my first kaggle competition, I fell satisfied with the final outcome. Thanks my teammate @nisshokuitsuki  , who worked with me during this three months journel. Thanks everyone who contributed to the discussion and the notebooks, your works gave us a lot insperiations. \n---\n# Model\nThe model we used is\n```\nWav2Vec2ForCTC\n```\nWe finetuned it further on this model [bengali_wav2vec2_finetuned (kaggle.com)](https://www.kaggle.com/datasets/nischaydnk/bengali-wav2vec2-finetuned)\nSince the data quality of the competition dataset is low, we used the train split of [Common Voice 13 | Bengali (Normalized) (kaggle.com)](https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv)to train. We trained about 10 epochs with the 20000 data in the dataset, and splited about 700 data for validation. \nThe data here is normalized and removed punctuations.\nHere is the training arguments:\n```python\ntraining_args = TrainingArguments(\n    group_by_length=False,\n    weight_decay=0.01,\n    num_train_epochs=10,\n    fp16=True,\n    learning_rate=4e-5,\n    warmup_steps=600,\n)\n```\nAnd we used cosine optimizer.\n\nThe best model got local wer 0.15, and Improved the Public Score from *0.445->0.434*\n\nHowever, my teammate also trained a model with the same data for just 60 steps( different args), and got the same score. (even 0.001 better on the private LB). How interesting and confusing.\n\n# Language Model\nWe trained a 6gram with [lm_no_punc (kaggle.com)](https://www.kaggle.com/datasets/mbmmurad/lm-no-punc). \nMention that there is a error in the LM provided in [YellowKing_DLSprint_Model (kaggle.com)](https://www.kaggle.com/datasets/sameen53/yellowking-dlsprint-model). \n```\nThere is no terminator in the arpa file. So we need to add an </s> into the file.\n```\nThanks to this notebook [Build an n-gram with KenLM | MaCro | Kaggle](https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro), we are able to realize this. \nAdding the </s> improved the LB from *0.445->0.422*\nAnd building an 6gram with [lm_no_punc (kaggle.com)](https://www.kaggle.com/datasets/mbmmurad/lm-no-punc).Improved about *0.001*\n# Punctuation Restoration\nPunctuation really matters. Thanks to this post [Bengali.AI Speech Recognition | Kaggle](https://www.kaggle.com/competitions/bengaliai-speech/discussion/432305), we are able to realize this. And an response under this post showed us a way to restore the punctuation:\n[xashru/punctuation-restoration: Punctuation Restoration using Transformer Models for High-and Low-Resource Languages (github.com)](https://github.com/xashru/punctuation-restoration) First we trained a model with the dataset provided in this repositorie. It can restore 3 punctuations : \n```python\n{1: ',', 2: '।', 3: '?'}\n```\n\nThis improved the LB from *0.422->0.400*\nCombining with the finetuned model, we have *0.400->0.397*\nAfterwards, we thought that 3 punctuations might be not enouth. So we made a dataset with [oscar · Datasets at Hugging Face](https://huggingface.co/datasets/oscar), filterd datas that have only have bengali words. We chose 7\npunctuations: \n```python\n{1: ',', 2: '।', 3: '?', 4: '!', 5: '-', 6: '\"', 7: ':'}\n```\nWe trained 6 epochs with the default pharams.\nAnd we have *0.393->0.387*\n\n# Model ensemble\nWe simply ensembled our model like this: \n```python\n            y = model_1(x).logits*0.7 + model_2(x).logits*0.2 + model_3(x).logits*0.1\n```\nWait, This works???\nYes, thouth the predictions may not be aligned, But since the three models are trained on same datasets, the no-aligning problem is paritially solved. \nThis improved our performance about *0.001*\n\n# Decoder pharams selection\nThere are three main pharams for the decoder:\n```\nalpha: weight for language model during shallow fusion\nbeta: weight for length score adjustment of during scoring\nbeam_width: determines the number of candidate output sequences retained at each time step.\n```\nTo find the best pharams, we used optuna [[0.444] Optimize Decoding Parameters with Optuna | Kaggle](https://www.kaggle.com/code/snnclsr/0-444-optimize-decoding-parameters-with-optuna) to search the best pharams. We searched the pharams with the example datas in the dataset, which is ood data, brought us better LB score.\nThe final decoder pharams are:\n```\n{'alpha': 0.46570704474381447, 'beta': 0.8635977171858652, 'beam_width': 768}\n```\n\n\n# What doesn't work for us\n- Data augmentation. We added background noise downloaded from https://pixabay.com/sound-effects/search/noise/ and also pitch shift , time stretch etc. But The LB got worse (*0.397->0.415*). Every experiment of data augmentation takes too much time and I fells to tired to do more experiments. Maybe I could write some codes to do it automatically\n- Denoise model. We tried three denoising models: \n\t\tUVR: Notebook Run Out of time\n\t\t[facebookresearch/denoiser](https://github.com/facebookresearch/denoiser): decreased about 0.01\n\t\t[CleanUNet](https://github.com/NVIDIA/CleanUNet/blob/main/exp/DNS-large-high/checkpoint/pretrained.pkl): decreased about 0.005\n\tWe thought that denoising harms the features and makes some short syllables unrecornizable.\n- Train a bigger LM with more data. We used 15G normalized benglai data to build an kenlm, the score got worse. We still haven't found the cause of the problem. \n",
    "2487449": "I can confirm that way of ensembling somehow works :)\nThank you for sharing",
    "3106246": "牛但是得凑够十个字才能发送",
    "2488159": "Congratulation for your great achievement. Still didn't get you ensemble part! Can you share your pseudocode how you do this? ",
    "2487858": ""
  }
}