{
  "id": 447995,
  "title": "4th place solution",
  "url": "/competitions/bengaliai-speech/discussion/447995",
  "author_name": "hanx",
  "post_date": "2023-10-18T04:11:13.212000",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Many thanks to the organizers for this interesting competition. Speech recognition is a very interesting direction and I did learn a lot from the discussion and public code of the many contestants in this competition. The past three months have been stressful but rewarding. I will try to make my solution clear in my broken English</p>\n<h2>Summary</h2>\n<p>My solution is relatively simple, using a wav2vec2 1b model as the pretrain model and training a Wav2Vec2ForCTC model. During the post-processing stage, a 6-gram language model is trained using KenLM, followed by further post-processing of normalization and dari on the output results.</p>\n<h2>Wav2Vec2ForCTC Training</h2>\n<p>Specifically, I use <a href=\"https://huggingface.co/facebook/wav2vec2-xls-r-1b\" target=\"_blank\">facebook/wav2vec2-xls-r-1b</a> as the pretrain model. Training this model requires three stages, with different random seeds and  consistent data augmentation and parameters in each stage:</p>\n<ul>\n<li>Optimizer: AdamW (weight_decay: 0.05; betas: (0.9, 0.999))</li>\n<li>Scheduler:  Modified linear_warmup_cosine scheduler (init_lr: 1e-5; min_lr: 5e-6; warmup_start_lr: 1e-6; warmup_steps: 1000; max_epoch: 120; iters_per_epoch: 1000)</li>\n</ul>\n<pre><code> :\n     ():\n        self.optimizer = optimizer\n\n        self.max_epoch = max_epoch\n        self.min_lr = min_lr\n\n        self.init_lr = init_lr\n        self.warmup_steps = warmup_steps\n        self.iters_per_epoch = iters_per_epoch\n        self.warmup_start_lr = warmup_start_lr  warmup_start_lr &gt;=   init_lr\n        self.max_iters = max_epoch * iters_per_epoch\n\n     ():\n        \n        total_steps = cur_epoch * self.iters_per_epoch + cur_step\n         total_steps &lt; self.warmup_steps:\n            warmup_lr_schedule(\n                step=cur_step,\n                optimizer=self.optimizer,\n                max_step=self.warmup_steps,\n                init_lr=self.warmup_start_lr,\n                max_lr=self.init_lr,\n            )\n         total_steps &lt;= self.max_iters // :\n            cosine_lr_schedule(\n                epoch=total_steps,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // ,\n                init_lr=self.init_lr,\n                min_lr=self.min_lr,\n            )\n         total_steps &lt;= self.max_iters // :\n            cosine_lr_schedule(\n                epoch=self.max_iters // ,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // ,\n                init_lr=self.init_lr,\n                min_lr=self.min_lr,\n            )\n        :  \n            cosine_lr_schedule(\n                epoch=total_steps - self.max_iters // ,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // ,\n                init_lr=self.min_lr,\n                min_lr=,\n            )\n\n\n ():\n    \n    lr = (init_lr - min_lr) *  * (\n             + math.cos(math.pi * epoch / max_epoch)\n    ) + min_lr\n     param_group  optimizer.param_groups:\n        param_group[] = lr\n</code></pre>\n<ul>\n<li>base DataAugmentation (denode as <code>base_aug</code>): </li>\n</ul>\n<pre><code> ():\n    trans = Compose(\n        [\n            TimeStretch(min_rate=, max_rate=, p=, leave_length_unchanged=),\n            Gain(min_gain_in_db=-, max_gain_in_db=, p=),\n            PitchShift(min_semitones=-, max_semitones=, p=),\n            OneOf(\n                [\n                    \n                    AddBackgroundNoise(sounds_path=musan_dir, min_snr_in_db=, max_snr_in_db=,\n                                       noise_transform=PolarityInversion(), p=),\n                    AddGaussianNoise(min_amplitude=, max_amplitude=, p=),\n                ]  musan_dir     [\n                    AddGaussianNoise(min_amplitude=, max_amplitude=, p=), ],\n                p=,\n            ),\n        ]\n    )\n     trans\n</code></pre>\n<ul>\n<li><p>composite DataAugmentation (denode as <code>comp_aug</code>)。I use three comp_augs：</p>\n<ol>\n<li>split an audio wave evenly into 3 segments and perform base_aug on each segment;</li>\n<li>randomly select two speeches from the dataset, perform base_aug on each of them separately, and then concatenate them together;</li>\n<li>combine the above two data augmentation methods.</li></ol></li>\n<li><p>dataset<br>\nFor Wav2Vec2ForCTC training, I didn't use external data, I filtered the competition dataset in the following steps:</p>\n<ol>\n<li>train a model based on <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">arijitx/wav2vec2-xls-r-300m-bengali</a> </li>\n<li>use the model above to inference the whole dataset, sort all sample scores from small to large and retain the top 70% of the data</li></ol></li>\n</ul>\n<h2>KenLM training</h2>\n<ul>\n<li><p>dataset:<br>\nI use IndicCorpv1 and IndicCorpv2 as corpus. After cleaning, the two corpus are combined without reduplicates.</p></li>\n<li><p>corpus cleaning:<br>\nI clean each sentence in the corpus using the below code:</p></li>\n</ul>\n<pre><code>chars_to_ignore = re.()\nlong_space_to_ignore = re.()\nbnorm = Normalizer()\n\n ():\n    \n    text = re.sub(chars_to_ignore, , text)\n    \n    text = re.sub(long_space_to_ignore, , text).strip()\n\n     text\n\n ():\n    sentence = normalize(sentence)\n    sentence = fix_text(sentence)\n    words = sentence.split()\n    :\n        all_words = [bnorm(word)[]  word  words]\n        all_words = [_  _  all_words  _]\n         (all_words) &lt; :\n             \n         .join(all_words).strip()\n     TypeError:\n         \n</code></pre>\n<ul>\n<li>6gram language model is trained.</li>\n</ul>\n<h2>Tried but not work</h2>\n<ol>\n<li>use external data, like openslr, shrutilipi</li>\n<li>use deepfilternet to denose the audio</li>\n<li>use larger model (wav2vec2-xls-r-2b)</li>\n<li>use whisper-small</li>\n<li>use more corpus to train an language model (BanglaLM)</li>\n<li>train a spelling error correction model</li>\n<li>train the fourth stage model</li>\n<li>…<br>\nThere's a lot more, off the top of my head</li>\n</ol>\n<h2>Update</h2>\n<ul>\n<li>I just found the 7gram-language kenlm model is chosen to be the final submission version, not the 6gram version, even though their scores are the same.</li>\n<li>inference code is made public: <a href=\"https://www.kaggle.com/code/hanx2smile/4th-place-solution-inference-code\" target=\"_blank\">https://www.kaggle.com/code/hanx2smile/4th-place-solution-inference-code</a></li>\n<li>training code is made public: <a href=\"https://github.com/HanxSmile/lavis-kaggle\" target=\"_blank\">https://github.com/HanxSmile/lavis-kaggle</a></li>\n</ul>",
  "messages": [
    {
      "id": 2486602,
      "postDate": "2023-10-18T04:11:13.213Z",
      "content": "<p>Many thanks to the organizers for this interesting competition. Speech recognition is a very interesting direction and I did learn a lot from the discussion and public code of the many contestants in this competition. The past three months have been stressful but rewarding. I will try to make my solution clear in my broken English</p>\n<h2>Summary</h2>\n<p>My solution is relatively simple, using a wav2vec2 1b model as the pretrain model and training a Wav2Vec2ForCTC model. During the post-processing stage, a 6-gram language model is trained using KenLM, followed by further post-processing of normalization and dari on the output results.</p>\n<h2>Wav2Vec2ForCTC Training</h2>\n<p>Specifically, I use <a href=\"https://huggingface.co/facebook/wav2vec2-xls-r-1b\" target=\"_blank\">facebook/wav2vec2-xls-r-1b</a> as the pretrain model. Training this model requires three stages, with different random seeds and  consistent data augmentation and parameters in each stage:</p>\n<ul>\n<li>Optimizer: AdamW (weight_decay: 0.05; betas: (0.9, 0.999))</li>\n<li>Scheduler:  Modified linear_warmup_cosine scheduler (init_lr: 1e-5; min_lr: 5e-6; warmup_start_lr: 1e-6; warmup_steps: 1000; max_epoch: 120; iters_per_epoch: 1000)</li>\n</ul>\n<pre><code> :\n     ():\n        self.optimizer = optimizer\n\n        self.max_epoch = max_epoch\n        self.min_lr = min_lr\n\n        self.init_lr = init_lr\n        self.warmup_steps = warmup_steps\n        self.iters_per_epoch = iters_per_epoch\n        self.warmup_start_lr = warmup_start_lr  warmup_start_lr &gt;=   init_lr\n        self.max_iters = max_epoch * iters_per_epoch\n\n     ():\n        \n        total_steps = cur_epoch * self.iters_per_epoch + cur_step\n         total_steps &lt; self.warmup_steps:\n            warmup_lr_schedule(\n                step=cur_step,\n                optimizer=self.optimizer,\n                max_step=self.warmup_steps,\n                init_lr=self.warmup_start_lr,\n                max_lr=self.init_lr,\n            )\n         total_steps &lt;= self.max_iters // :\n            cosine_lr_schedule(\n                epoch=total_steps,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // ,\n                init_lr=self.init_lr,\n                min_lr=self.min_lr,\n            )\n         total_steps &lt;= self.max_iters // :\n            cosine_lr_schedule(\n                epoch=self.max_iters // ,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // ,\n                init_lr=self.init_lr,\n                min_lr=self.min_lr,\n            )\n        :  \n            cosine_lr_schedule(\n                epoch=total_steps - self.max_iters // ,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // ,\n                init_lr=self.min_lr,\n                min_lr=,\n            )\n\n\n ():\n    \n    lr = (init_lr - min_lr) *  * (\n             + math.cos(math.pi * epoch / max_epoch)\n    ) + min_lr\n     param_group  optimizer.param_groups:\n        param_group[] = lr\n</code></pre>\n<ul>\n<li>base DataAugmentation (denode as <code>base_aug</code>): </li>\n</ul>\n<pre><code> ():\n    trans = Compose(\n        [\n            TimeStretch(min_rate=, max_rate=, p=, leave_length_unchanged=),\n            Gain(min_gain_in_db=-, max_gain_in_db=, p=),\n            PitchShift(min_semitones=-, max_semitones=, p=),\n            OneOf(\n                [\n                    \n                    AddBackgroundNoise(sounds_path=musan_dir, min_snr_in_db=, max_snr_in_db=,\n                                       noise_transform=PolarityInversion(), p=),\n                    AddGaussianNoise(min_amplitude=, max_amplitude=, p=),\n                ]  musan_dir     [\n                    AddGaussianNoise(min_amplitude=, max_amplitude=, p=), ],\n                p=,\n            ),\n        ]\n    )\n     trans\n</code></pre>\n<ul>\n<li><p>composite DataAugmentation (denode as <code>comp_aug</code>)。I use three comp_augs：</p>\n<ol>\n<li>split an audio wave evenly into 3 segments and perform base_aug on each segment;</li>\n<li>randomly select two speeches from the dataset, perform base_aug on each of them separately, and then concatenate them together;</li>\n<li>combine the above two data augmentation methods.</li></ol></li>\n<li><p>dataset<br>\nFor Wav2Vec2ForCTC training, I didn't use external data, I filtered the competition dataset in the following steps:</p>\n<ol>\n<li>train a model based on <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">arijitx/wav2vec2-xls-r-300m-bengali</a> </li>\n<li>use the model above to inference the whole dataset, sort all sample scores from small to large and retain the top 70% of the data</li></ol></li>\n</ul>\n<h2>KenLM training</h2>\n<ul>\n<li><p>dataset:<br>\nI use IndicCorpv1 and IndicCorpv2 as corpus. After cleaning, the two corpus are combined without reduplicates.</p></li>\n<li><p>corpus cleaning:<br>\nI clean each sentence in the corpus using the below code:</p></li>\n</ul>\n<pre><code>chars_to_ignore = re.()\nlong_space_to_ignore = re.()\nbnorm = Normalizer()\n\n ():\n    \n    text = re.sub(chars_to_ignore, , text)\n    \n    text = re.sub(long_space_to_ignore, , text).strip()\n\n     text\n\n ():\n    sentence = normalize(sentence)\n    sentence = fix_text(sentence)\n    words = sentence.split()\n    :\n        all_words = [bnorm(word)[]  word  words]\n        all_words = [_  _  all_words  _]\n         (all_words) &lt; :\n             \n         .join(all_words).strip()\n     TypeError:\n         \n</code></pre>\n<ul>\n<li>6gram language model is trained.</li>\n</ul>\n<h2>Tried but not work</h2>\n<ol>\n<li>use external data, like openslr, shrutilipi</li>\n<li>use deepfilternet to denose the audio</li>\n<li>use larger model (wav2vec2-xls-r-2b)</li>\n<li>use whisper-small</li>\n<li>use more corpus to train an language model (BanglaLM)</li>\n<li>train a spelling error correction model</li>\n<li>train the fourth stage model</li>\n<li>…<br>\nThere's a lot more, off the top of my head</li>\n</ol>\n<h2>Update</h2>\n<ul>\n<li>I just found the 7gram-language kenlm model is chosen to be the final submission version, not the 6gram version, even though their scores are the same.</li>\n<li>inference code is made public: <a href=\"https://www.kaggle.com/code/hanx2smile/4th-place-solution-inference-code\" target=\"_blank\">https://www.kaggle.com/code/hanx2smile/4th-place-solution-inference-code</a></li>\n<li>training code is made public: <a href=\"https://github.com/HanxSmile/lavis-kaggle\" target=\"_blank\">https://github.com/HanxSmile/lavis-kaggle</a></li>\n</ul>",
      "rawMarkdown": "Many thanks to the organizers for this interesting competition. Speech recognition is a very interesting direction and I did learn a lot from the discussion and public code of the many contestants in this competition. The past three months have been stressful but rewarding. I will try to make my solution clear in my broken English\n\n## Summary\nMy solution is relatively simple, using a wav2vec2 1b model as the pretrain model and training a Wav2Vec2ForCTC model. During the post-processing stage, a 6-gram language model is trained using KenLM, followed by further post-processing of normalization and dari on the output results.\n\n## Wav2Vec2ForCTC Training\nSpecifically, I use [facebook/wav2vec2-xls-r-1b](https://huggingface.co/facebook/wav2vec2-xls-r-1b) as the pretrain model. Training this model requires three stages, with different random seeds and  consistent data augmentation and parameters in each stage:\n* Optimizer: AdamW (weight_decay: 0.05; betas: (0.9, 0.999))\n* Scheduler:  Modified linear_warmup_cosine scheduler (init_lr: 1e-5; min_lr: 5e-6; warmup_start_lr: 1e-6; warmup_steps: 1000; max_epoch: 120; iters_per_epoch: 1000)\n```python\nclass LinearWarmupCosine3LongTailLRScheduler:\n    def __init__(\n            self,\n            optimizer,\n            max_epoch,\n            min_lr,\n            init_lr,\n            iters_per_epoch,\n            warmup_steps=0,\n            warmup_start_lr=-1,\n            **kwargs\n    ):\n        self.optimizer = optimizer\n\n        self.max_epoch = max_epoch\n        self.min_lr = min_lr\n\n        self.init_lr = init_lr\n        self.warmup_steps = warmup_steps\n        self.iters_per_epoch = iters_per_epoch\n        self.warmup_start_lr = warmup_start_lr if warmup_start_lr >= 0 else init_lr\n        self.max_iters = max_epoch * iters_per_epoch\n\n    def step(self, cur_epoch, cur_step):\n        # assuming the warmup iters less than one epoch\n        total_steps = cur_epoch * self.iters_per_epoch + cur_step\n        if total_steps < self.warmup_steps:\n            warmup_lr_schedule(\n                step=cur_step,\n                optimizer=self.optimizer,\n                max_step=self.warmup_steps,\n                init_lr=self.warmup_start_lr,\n                max_lr=self.init_lr,\n            )\n        elif total_steps <= self.max_iters // 4:\n            cosine_lr_schedule(\n                epoch=total_steps,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // 4,\n                init_lr=self.init_lr,\n                min_lr=self.min_lr,\n            )\n        elif total_steps <= self.max_iters // 2:\n            cosine_lr_schedule(\n                epoch=self.max_iters // 4,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // 4,\n                init_lr=self.init_lr,\n                min_lr=self.min_lr,\n            )\n        else:  # total_steps > self.max_iters // 2\n            cosine_lr_schedule(\n                epoch=total_steps - self.max_iters // 2,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // 2,\n                init_lr=self.min_lr,\n                min_lr=0,\n            )\n\n\ndef cosine_lr_schedule(optimizer, epoch, max_epoch, init_lr, min_lr):\n    \"\"\"Decay the learning rate\"\"\"\n    lr = (init_lr - min_lr) * 0.5 * (\n            1.0 + math.cos(math.pi * epoch / max_epoch)\n    ) + min_lr\n    for param_group in optimizer.param_groups:\n        param_group[\"lr\"] = lr\n```\n\n* base DataAugmentation (denode as `base_aug`): \n```python\ndef get_transform(musan_dir):\n    trans = Compose(\n        [\n            TimeStretch(min_rate=0.9, max_rate=1.1, p=0.2, leave_length_unchanged=False),\n            Gain(min_gain_in_db=-6, max_gain_in_db=6, p=0.1),\n            PitchShift(min_semitones=-4, max_semitones=4, p=0.2),\n            OneOf(\n                [\n                    # AddBackgroundNoise(sounds_path=musan_dir, min_snr_in_db=1.0, max_snr_in_db=5.0,\n                    AddBackgroundNoise(sounds_path=musan_dir, min_snr_in_db=3.0, max_snr_in_db=30.0,\n                                       noise_transform=PolarityInversion(), p=1.0),\n                    AddGaussianNoise(min_amplitude=0.005, max_amplitude=0.015, p=1.0),\n                ] if musan_dir is not None else [\n                    AddGaussianNoise(min_amplitude=0.005, max_amplitude=0.015, p=1.0), ],\n                p=0.5,\n            ),\n        ]\n    )\n    return trans\n```\n\n* composite DataAugmentation (denode as `comp_aug`)。I use three comp_augs：\n    1. split an audio wave evenly into 3 segments and perform base_aug on each segment;\n    2. randomly select two speeches from the dataset, perform base_aug on each of them separately, and then concatenate them together;\n    3. combine the above two data augmentation methods.\n\n* dataset\nFor Wav2Vec2ForCTC training, I didn't use external data, I filtered the competition dataset in the following steps:\n    1. train a model based on [arijitx/wav2vec2-xls-r-300m-bengali](https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali) \n    2. use the model above to inference the whole dataset, sort all sample scores from small to large and retain the top 70% of the data\n\n\n## KenLM training\n* dataset:\nI use IndicCorpv1 and IndicCorpv2 as corpus. After cleaning, the two corpus are combined without reduplicates.\n\n* corpus cleaning:\nI clean each sentence in the corpus using the below code:\n```python\nchars_to_ignore = re.compile(r'[^\\u0980-\\u09FF\\s]')\nlong_space_to_ignore = re.compile(r'\\s+')\nbnorm = Normalizer()\n\ndef fix_text(text: str):\n    # remove punctuations\n    text = re.sub(chars_to_ignore, ' ', text)\n    # match multiple spaces and replace them with a single space\n    text = re.sub(long_space_to_ignore, ' ', text).strip()\n\n    return text\n\ndef norm_sentence(sentence):\n    sentence = normalize(sentence)\n    sentence = fix_text(sentence)\n    words = sentence.split()\n    try:\n        all_words = [bnorm(word)[\"normalized\"] for word in words]\n        all_words = [_ for _ in all_words if _]\n        if len(all_words) < 2:\n            return \"\"\n        return \" \".join(all_words).strip()\n    except TypeError:\n        return None\n```\n\n* 6gram language model is trained.\n\n## Tried but not work\n1. use external data, like openslr, shrutilipi\n2. use deepfilternet to denose the audio\n3. use larger model (wav2vec2-xls-r-2b)\n4. use whisper-small\n5. use more corpus to train an language model (BanglaLM)\n6. train a spelling error correction model\n7. train the fourth stage model\n8. ...\nThere's a lot more, off the top of my head\n\n## Update\n* I just found the 7gram-language kenlm model is chosen to be the final submission version, not the 6gram version, even though their scores are the same.\n* inference code is made public: https://www.kaggle.com/code/hanx2smile/4th-place-solution-inference-code\n* training code is made public: https://github.com/HanxSmile/lavis-kaggle",
      "votes": 15
    },
    {
      "id": 2486611,
      "postDate": "2023-10-18T04:18:23.367Z",
      "content": "<p>Thank you for sharing the solution!<br>\nI also tried deepfilternet, but it was not effective. Demucs worked well though it is not a model for denoising.</p>",
      "rawMarkdown": "Thank you for sharing the solution!\nI also tried deepfilternet, but it was not effective. Demucs worked well though it is not a model for denoising.",
      "replies": [
        {
          "id": 2486654,
          "postDate": "2023-10-18T05:04:46.580Z",
          "content": "<p>Thank you for your reply! It looks like that a souce separation model is more suitable for this competition. Your solution is very inspiring, especially the punctuation model part, Congratulations!</p>",
          "rawMarkdown": "Thank you for your reply! It looks like that a souce separation model is more suitable for this competition. Your solution is very inspiring, especially the punctuation model part, Congratulations!",
          "votes": 2,
          "replies": [
            {
              "id": 2489997,
              "postDate": "2023-10-20T10:57:05.113Z",
              "content": "<p>Thank you! Congratulations, too! Winning a gold medal is outstanding.</p>",
              "rawMarkdown": "Thank you! Congratulations, too! Winning a gold medal is outstanding."
            }
          ]
        }
      ]
    },
    {
      "id": 2487850,
      "postDate": "2023-10-18T19:50:06.717Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2486611,
      "author_name": "moto",
      "author_url": "",
      "post_date": "2023-10-18T04:18:23.367000",
      "content": "<p>Thank you for sharing the solution!<br>\nI also tried deepfilternet, but it was not effective. Demucs worked well though it is not a model for denoising.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2486654,
          "author_name": "hanx",
          "author_url": "",
          "post_date": "2023-10-18T05:04:46.580000",
          "content": "<p>Thank you for your reply! It looks like that a souce separation model is more suitable for this competition. Your solution is very inspiring, especially the punctuation model part, Congratulations!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2489997,
              "author_name": "moto",
              "author_url": "",
              "post_date": "2023-10-20T10:57:05.113000",
              "content": "<p>Thank you! Congratulations, too! Winning a gold medal is outstanding.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2487850,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:50:06.717000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486602": "Many thanks to the organizers for this interesting competition. Speech recognition is a very interesting direction and I did learn a lot from the discussion and public code of the many contestants in this competition. The past three months have been stressful but rewarding. I will try to make my solution clear in my broken English\n\n## Summary\nMy solution is relatively simple, using a wav2vec2 1b model as the pretrain model and training a Wav2Vec2ForCTC model. During the post-processing stage, a 6-gram language model is trained using KenLM, followed by further post-processing of normalization and dari on the output results.\n\n## Wav2Vec2ForCTC Training\nSpecifically, I use [facebook/wav2vec2-xls-r-1b](https://huggingface.co/facebook/wav2vec2-xls-r-1b) as the pretrain model. Training this model requires three stages, with different random seeds and  consistent data augmentation and parameters in each stage:\n* Optimizer: AdamW (weight_decay: 0.05; betas: (0.9, 0.999))\n* Scheduler:  Modified linear_warmup_cosine scheduler (init_lr: 1e-5; min_lr: 5e-6; warmup_start_lr: 1e-6; warmup_steps: 1000; max_epoch: 120; iters_per_epoch: 1000)\n```python\nclass LinearWarmupCosine3LongTailLRScheduler:\n    def __init__(\n            self,\n            optimizer,\n            max_epoch,\n            min_lr,\n            init_lr,\n            iters_per_epoch,\n            warmup_steps=0,\n            warmup_start_lr=-1,\n            **kwargs\n    ):\n        self.optimizer = optimizer\n\n        self.max_epoch = max_epoch\n        self.min_lr = min_lr\n\n        self.init_lr = init_lr\n        self.warmup_steps = warmup_steps\n        self.iters_per_epoch = iters_per_epoch\n        self.warmup_start_lr = warmup_start_lr if warmup_start_lr >= 0 else init_lr\n        self.max_iters = max_epoch * iters_per_epoch\n\n    def step(self, cur_epoch, cur_step):\n        # assuming the warmup iters less than one epoch\n        total_steps = cur_epoch * self.iters_per_epoch + cur_step\n        if total_steps < self.warmup_steps:\n            warmup_lr_schedule(\n                step=cur_step,\n                optimizer=self.optimizer,\n                max_step=self.warmup_steps,\n                init_lr=self.warmup_start_lr,\n                max_lr=self.init_lr,\n            )\n        elif total_steps <= self.max_iters // 4:\n            cosine_lr_schedule(\n                epoch=total_steps,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // 4,\n                init_lr=self.init_lr,\n                min_lr=self.min_lr,\n            )\n        elif total_steps <= self.max_iters // 2:\n            cosine_lr_schedule(\n                epoch=self.max_iters // 4,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // 4,\n                init_lr=self.init_lr,\n                min_lr=self.min_lr,\n            )\n        else:  # total_steps > self.max_iters // 2\n            cosine_lr_schedule(\n                epoch=total_steps - self.max_iters // 2,\n                optimizer=self.optimizer,\n                max_epoch=self.max_iters // 2,\n                init_lr=self.min_lr,\n                min_lr=0,\n            )\n\n\ndef cosine_lr_schedule(optimizer, epoch, max_epoch, init_lr, min_lr):\n    \"\"\"Decay the learning rate\"\"\"\n    lr = (init_lr - min_lr) * 0.5 * (\n            1.0 + math.cos(math.pi * epoch / max_epoch)\n    ) + min_lr\n    for param_group in optimizer.param_groups:\n        param_group[\"lr\"] = lr\n```\n\n* base DataAugmentation (denode as `base_aug`): \n```python\ndef get_transform(musan_dir):\n    trans = Compose(\n        [\n            TimeStretch(min_rate=0.9, max_rate=1.1, p=0.2, leave_length_unchanged=False),\n            Gain(min_gain_in_db=-6, max_gain_in_db=6, p=0.1),\n            PitchShift(min_semitones=-4, max_semitones=4, p=0.2),\n            OneOf(\n                [\n                    # AddBackgroundNoise(sounds_path=musan_dir, min_snr_in_db=1.0, max_snr_in_db=5.0,\n                    AddBackgroundNoise(sounds_path=musan_dir, min_snr_in_db=3.0, max_snr_in_db=30.0,\n                                       noise_transform=PolarityInversion(), p=1.0),\n                    AddGaussianNoise(min_amplitude=0.005, max_amplitude=0.015, p=1.0),\n                ] if musan_dir is not None else [\n                    AddGaussianNoise(min_amplitude=0.005, max_amplitude=0.015, p=1.0), ],\n                p=0.5,\n            ),\n        ]\n    )\n    return trans\n```\n\n* composite DataAugmentation (denode as `comp_aug`)。I use three comp_augs：\n    1. split an audio wave evenly into 3 segments and perform base_aug on each segment;\n    2. randomly select two speeches from the dataset, perform base_aug on each of them separately, and then concatenate them together;\n    3. combine the above two data augmentation methods.\n\n* dataset\nFor Wav2Vec2ForCTC training, I didn't use external data, I filtered the competition dataset in the following steps:\n    1. train a model based on [arijitx/wav2vec2-xls-r-300m-bengali](https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali) \n    2. use the model above to inference the whole dataset, sort all sample scores from small to large and retain the top 70% of the data\n\n\n## KenLM training\n* dataset:\nI use IndicCorpv1 and IndicCorpv2 as corpus. After cleaning, the two corpus are combined without reduplicates.\n\n* corpus cleaning:\nI clean each sentence in the corpus using the below code:\n```python\nchars_to_ignore = re.compile(r'[^\\u0980-\\u09FF\\s]')\nlong_space_to_ignore = re.compile(r'\\s+')\nbnorm = Normalizer()\n\ndef fix_text(text: str):\n    # remove punctuations\n    text = re.sub(chars_to_ignore, ' ', text)\n    # match multiple spaces and replace them with a single space\n    text = re.sub(long_space_to_ignore, ' ', text).strip()\n\n    return text\n\ndef norm_sentence(sentence):\n    sentence = normalize(sentence)\n    sentence = fix_text(sentence)\n    words = sentence.split()\n    try:\n        all_words = [bnorm(word)[\"normalized\"] for word in words]\n        all_words = [_ for _ in all_words if _]\n        if len(all_words) < 2:\n            return \"\"\n        return \" \".join(all_words).strip()\n    except TypeError:\n        return None\n```\n\n* 6gram language model is trained.\n\n## Tried but not work\n1. use external data, like openslr, shrutilipi\n2. use deepfilternet to denose the audio\n3. use larger model (wav2vec2-xls-r-2b)\n4. use whisper-small\n5. use more corpus to train an language model (BanglaLM)\n6. train a spelling error correction model\n7. train the fourth stage model\n8. ...\nThere's a lot more, off the top of my head\n\n## Update\n* I just found the 7gram-language kenlm model is chosen to be the final submission version, not the 6gram version, even though their scores are the same.\n* inference code is made public: https://www.kaggle.com/code/hanx2smile/4th-place-solution-inference-code\n* training code is made public: https://github.com/HanxSmile/lavis-kaggle",
    "2486611": "Thank you for sharing the solution!\nI also tried deepfilternet, but it was not effective. Demucs worked well though it is not a model for denoising.",
    "2487850": ""
  }
}