{
  "id": 447976,
  "title": "2nd place solution",
  "url": "/competitions/bengaliai-speech/discussion/447976",
  "author_name": "qdv206",
  "post_date": "2023-10-18T02:27:13.218000",
  "votes": 41,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Many thanks to Kaggle and Bengali.AI for hosting such an interesting competition. As I was about to start my PhD  research on speech processing and I knew nothing about the field, the competition came as the perfect opportunity for me to learn. In the end it was a highly rewarding experience.</p>\n<p>The solution consists of 3 components:</p>\n<ul>\n<li>ASR model</li>\n<li>Language model</li>\n<li>Punctuation model</li>\n</ul>\n<p><strong>1. ASR model</strong></p>\n<p>I used <code>ai4bharat/indicwav2vec_v1_bengali</code> as pretrained model.</p>\n<p><strong>Datasets:</strong></p>\n<ul>\n<li>Speech data: Competition data, Shrutilipi, MADASR, ULCA (for ULCA data most of the links are dead but some are still downloadable, only a few thousands samples though)</li>\n<li>Noise data: music data from MUSAN and noise data from DNS Challenge 2020</li>\n<li>All data are normalized and punctuation marks are removed except for dot (<code>.</code>) and hyphen (<code>-</code>)</li>\n</ul>\n<p><strong>Augmentation:</strong></p>\n<ul>\n<li>use augmentation from <code>audiomentations</code>, apply heavy augmentation for read speech (comp data, MADASR) and lighter augmentation for spontaneous speech (Shrutilipi, ULCA)<br>\nExample of read speech augmentation:</li>\n</ul>\n<pre><code>augments = Compose([\n    TimeStretch(=0.8, =2.0, =0.5, =),\n    RoomSimulator(=0.3),\n    OneOf([\n        AddBackgroundNoise(\n            sounds_path=[\n                ,\n            ],\n            =5.0,\n            =30.0,\n            =PolarityInversion(),\n            =1.0\n        ),\n        AddBackgroundNoise(\n            sounds_path=[\n                \n            ],\n            =5.0,\n            =30.0,\n            =PolarityInversion(),\n            =1.0\n        ),\n        AddGaussianNoise(=0.005, =0.015, =1.0),\n    ], =0.7),\n    Gain(=-6, =6, =0.2),\n    ])\n</code></pre>\n<p>For spontaneous speech augment probabilities are smaller and <code>TimeStretch</code> rate much less extreme.</p>\n<ul>\n<li>concat augment: randomly concatenate short samples together to make length distribution of training set closer to OOD test set.</li>\n<li>SpecAugment: mask_time_prob = 0.1, mask_feature_prob = 0.05.</li>\n</ul>\n<p><strong>Training:</strong></p>\n<ul>\n<li>First fit on all training data, then remove about 10% with highest WER after fitted from train set.</li>\n<li>Don't freeze feature encoder.</li>\n<li>Use cosine schedule with warmups and restarts: 1st cycle 5 epochs peak lr 4e-5, 2nd cycle 3 epochs peak lr 3e-5, third cycle 3 epochs peak lr 2e-5.</li>\n</ul>\n<p><strong>Inference:</strong><br>\nUse <code>AutomaticSpeechRecognitionPipeline</code> from <code>transformers</code> to apply inference with chunking and stride:</p>\n<pre><code> = pipe(w, chunk_length_s=, stride_length_s=(, ))[]\n</code></pre>\n<p><strong>2. Language model</strong></p>\n<p>6-gram kenlm model trained on multiple external Bengali corpus:</p>\n<ul>\n<li>IndicCorp V1+V2.</li>\n<li>Bharat Parallel Corpus Collection.</li>\n<li>Samanantar.</li>\n<li><a href=\"https://www.kaggle.com/datasets/truthr/free-bengali-poetry\" target=\"_blank\">Bengali poetry dataset</a>.</li>\n<li><a href=\"https://data.statmt.org/news-crawl/\" target=\"_blank\">WMT News Crawl</a>.</li>\n<li>Hate speech corpus from <a href=\"https://github.com/rezacsedu/Classification_Benchmarks_Benglai_NLP\" target=\"_blank\">https://github.com/rezacsedu/Classification_Benchmarks_Benglai_NLP</a>.</li>\n</ul>\n<p><strong>3. Punctuation model</strong></p>\n<p>Train token classification model to add the following punctuation set: <code>।,?!</code></p>\n<ul>\n<li>use <code>ai4bharat/IndicBERTv2-MLM-Sam-TLM</code> as backbone</li>\n<li>add LSTM head</li>\n<li>train for 6 epochs, cosine schedule, lr 3e-5 on competition data + subset of IndicCorp</li>\n<li>mask 15% of the tokens during training as augmentation</li>\n<li>ensemble 3 folds of model trained on 3 different subsets of IndicCorp</li>\n<li>beam search decoding for inference.</li>\n</ul>\n<p>Thank you very much for reading and please let me know if you have any questions.</p>\n<p>Update: </p>\n<ul>\n<li>training code: <a href=\"https://github.com/quangdao206/Kaggle_Bengali_Speech_Recognition_2nd_Place_Solution\" target=\"_blank\">https://github.com/quangdao206/Kaggle_Bengali_Speech_Recognition_2nd_Place_Solution</a></li>\n<li>inference notebook: <a href=\"https://www.kaggle.com/code/qdv206/2nd-place-bengali-speech-infer/\" target=\"_blank\">https://www.kaggle.com/code/qdv206/2nd-place-bengali-speech-infer/</a></li>\n</ul>",
  "messages": [
    {
      "id": 2486531,
      "postDate": "2023-10-18T02:27:13.220Z",
      "content": "<p>Many thanks to Kaggle and Bengali.AI for hosting such an interesting competition. As I was about to start my PhD  research on speech processing and I knew nothing about the field, the competition came as the perfect opportunity for me to learn. In the end it was a highly rewarding experience.</p>\n<p>The solution consists of 3 components:</p>\n<ul>\n<li>ASR model</li>\n<li>Language model</li>\n<li>Punctuation model</li>\n</ul>\n<p><strong>1. ASR model</strong></p>\n<p>I used <code>ai4bharat/indicwav2vec_v1_bengali</code> as pretrained model.</p>\n<p><strong>Datasets:</strong></p>\n<ul>\n<li>Speech data: Competition data, Shrutilipi, MADASR, ULCA (for ULCA data most of the links are dead but some are still downloadable, only a few thousands samples though)</li>\n<li>Noise data: music data from MUSAN and noise data from DNS Challenge 2020</li>\n<li>All data are normalized and punctuation marks are removed except for dot (<code>.</code>) and hyphen (<code>-</code>)</li>\n</ul>\n<p><strong>Augmentation:</strong></p>\n<ul>\n<li>use augmentation from <code>audiomentations</code>, apply heavy augmentation for read speech (comp data, MADASR) and lighter augmentation for spontaneous speech (Shrutilipi, ULCA)<br>\nExample of read speech augmentation:</li>\n</ul>\n<pre><code>augments = Compose([\n    TimeStretch(=0.8, =2.0, =0.5, =),\n    RoomSimulator(=0.3),\n    OneOf([\n        AddBackgroundNoise(\n            sounds_path=[\n                ,\n            ],\n            =5.0,\n            =30.0,\n            =PolarityInversion(),\n            =1.0\n        ),\n        AddBackgroundNoise(\n            sounds_path=[\n                \n            ],\n            =5.0,\n            =30.0,\n            =PolarityInversion(),\n            =1.0\n        ),\n        AddGaussianNoise(=0.005, =0.015, =1.0),\n    ], =0.7),\n    Gain(=-6, =6, =0.2),\n    ])\n</code></pre>\n<p>For spontaneous speech augment probabilities are smaller and <code>TimeStretch</code> rate much less extreme.</p>\n<ul>\n<li>concat augment: randomly concatenate short samples together to make length distribution of training set closer to OOD test set.</li>\n<li>SpecAugment: mask_time_prob = 0.1, mask_feature_prob = 0.05.</li>\n</ul>\n<p><strong>Training:</strong></p>\n<ul>\n<li>First fit on all training data, then remove about 10% with highest WER after fitted from train set.</li>\n<li>Don't freeze feature encoder.</li>\n<li>Use cosine schedule with warmups and restarts: 1st cycle 5 epochs peak lr 4e-5, 2nd cycle 3 epochs peak lr 3e-5, third cycle 3 epochs peak lr 2e-5.</li>\n</ul>\n<p><strong>Inference:</strong><br>\nUse <code>AutomaticSpeechRecognitionPipeline</code> from <code>transformers</code> to apply inference with chunking and stride:</p>\n<pre><code> = pipe(w, chunk_length_s=, stride_length_s=(, ))[]\n</code></pre>\n<p><strong>2. Language model</strong></p>\n<p>6-gram kenlm model trained on multiple external Bengali corpus:</p>\n<ul>\n<li>IndicCorp V1+V2.</li>\n<li>Bharat Parallel Corpus Collection.</li>\n<li>Samanantar.</li>\n<li><a href=\"https://www.kaggle.com/datasets/truthr/free-bengali-poetry\" target=\"_blank\">Bengali poetry dataset</a>.</li>\n<li><a href=\"https://data.statmt.org/news-crawl/\" target=\"_blank\">WMT News Crawl</a>.</li>\n<li>Hate speech corpus from <a href=\"https://github.com/rezacsedu/Classification_Benchmarks_Benglai_NLP\" target=\"_blank\">https://github.com/rezacsedu/Classification_Benchmarks_Benglai_NLP</a>.</li>\n</ul>\n<p><strong>3. Punctuation model</strong></p>\n<p>Train token classification model to add the following punctuation set: <code>।,?!</code></p>\n<ul>\n<li>use <code>ai4bharat/IndicBERTv2-MLM-Sam-TLM</code> as backbone</li>\n<li>add LSTM head</li>\n<li>train for 6 epochs, cosine schedule, lr 3e-5 on competition data + subset of IndicCorp</li>\n<li>mask 15% of the tokens during training as augmentation</li>\n<li>ensemble 3 folds of model trained on 3 different subsets of IndicCorp</li>\n<li>beam search decoding for inference.</li>\n</ul>\n<p>Thank you very much for reading and please let me know if you have any questions.</p>\n<p>Update: </p>\n<ul>\n<li>training code: <a href=\"https://github.com/quangdao206/Kaggle_Bengali_Speech_Recognition_2nd_Place_Solution\" target=\"_blank\">https://github.com/quangdao206/Kaggle_Bengali_Speech_Recognition_2nd_Place_Solution</a></li>\n<li>inference notebook: <a href=\"https://www.kaggle.com/code/qdv206/2nd-place-bengali-speech-infer/\" target=\"_blank\">https://www.kaggle.com/code/qdv206/2nd-place-bengali-speech-infer/</a></li>\n</ul>",
      "rawMarkdown": "Many thanks to Kaggle and Bengali.AI for hosting such an interesting competition. As I was about to start my PhD  research on speech processing and I knew nothing about the field, the competition came as the perfect opportunity for me to learn. In the end it was a highly rewarding experience.\n\nThe solution consists of 3 components:\n- ASR model\n- Language model\n- Punctuation model\n\n**1. ASR model**\n\nI used `ai4bharat/indicwav2vec_v1_bengali` as pretrained model.\n\n**Datasets:**\n- Speech data: Competition data, Shrutilipi, MADASR, ULCA (for ULCA data most of the links are dead but some are still downloadable, only a few thousands samples though)\n- Noise data: music data from MUSAN and noise data from DNS Challenge 2020\n- All data are normalized and punctuation marks are removed except for dot (`.`) and hyphen (`-`)\n\n**Augmentation:**\n- use augmentation from `audiomentations`, apply heavy augmentation for read speech (comp data, MADASR) and lighter augmentation for spontaneous speech (Shrutilipi, ULCA)\nExample of read speech augmentation:\n```\naugments = Compose([\n    TimeStretch(min_rate=0.8, max_rate=2.0, p=0.5, leave_length_unchanged=False),\n    RoomSimulator(p=0.3),\n    OneOf([\n        AddBackgroundNoise(\n            sounds_path=[\n                '/path_to_DNS_Challenge_noise',\n            ],\n            min_snr_in_db=5.0,\n            max_snr_in_db=30.0,\n            noise_transform=PolarityInversion(),\n            p=1.0\n        ),\n        AddBackgroundNoise(\n            sounds_path=[\n                '/path_to_MUSAN_music'\n            ],\n            min_snr_in_db=5.0,\n            max_snr_in_db=30.0,\n            noise_transform=PolarityInversion(),\n            p=1.0\n        ),\n        AddGaussianNoise(min_amplitude=0.005, max_amplitude=0.015, p=1.0),\n    ], p=0.7),\n    Gain(min_gain_in_db=-6, max_gain_in_db=6, p=0.2),\n    ])\n```\nFor spontaneous speech augment probabilities are smaller and `TimeStretch` rate much less extreme.\n- concat augment: randomly concatenate short samples together to make length distribution of training set closer to OOD test set.\n- SpecAugment: mask_time_prob = 0.1, mask_feature_prob = 0.05.\n\n**Training:**\n- First fit on all training data, then remove about 10% with highest WER after fitted from train set.\n- Don't freeze feature encoder.\n- Use cosine schedule with warmups and restarts: 1st cycle 5 epochs peak lr 4e-5, 2nd cycle 3 epochs peak lr 3e-5, third cycle 3 epochs peak lr 2e-5.\n\n**Inference:**\nUse `AutomaticSpeechRecognitionPipeline` from `transformers` to apply inference with chunking and stride:\n```\ntext = pipe(w, chunk_length_s=14, stride_length_s=(6, 3))[\"text\"]\n```\n\n**2. Language model**\n\n6-gram kenlm model trained on multiple external Bengali corpus:\n- IndicCorp V1+V2.\n- Bharat Parallel Corpus Collection.\n- Samanantar.\n- [Bengali poetry dataset](https://www.kaggle.com/datasets/truthr/free-bengali-poetry).\n- [WMT News Crawl](https://data.statmt.org/news-crawl/).\n- Hate speech corpus from https://github.com/rezacsedu/Classification_Benchmarks_Benglai_NLP.\n\n**3. Punctuation model**\n\nTrain token classification model to add the following punctuation set: `।,?!`\n\n- use `ai4bharat/IndicBERTv2-MLM-Sam-TLM` as backbone\n- add LSTM head\n- train for 6 epochs, cosine schedule, lr 3e-5 on competition data + subset of IndicCorp\n- mask 15% of the tokens during training as augmentation\n- ensemble 3 folds of model trained on 3 different subsets of IndicCorp\n- beam search decoding for inference.\n\nThank you very much for reading and please let me know if you have any questions.\n\nUpdate: \n- training code: https://github.com/quangdao206/Kaggle_Bengali_Speech_Recognition_2nd_Place_Solution\n- inference notebook: https://www.kaggle.com/code/qdv206/2nd-place-bengali-speech-infer/\n\n",
      "votes": 41
    },
    {
      "id": 2490732,
      "postDate": "2023-10-21T01:52:55.477Z",
      "content": "<p>Such neat work! Congratulations on becoming a GM <a href=\"https://www.kaggle.com/qdv206\" target=\"_blank\">@qdv206</a> </p>",
      "rawMarkdown": "Such neat work! Congratulations on becoming a GM @qdv206 ",
      "votes": 1
    },
    {
      "id": 2489780,
      "postDate": "2023-10-20T07:40:18.207Z",
      "content": "<p>Congratulations on your solo gold and becoming a grandmaster.<br>\nThank you for sharing a detailed solution. A small question, you said you retained \",\" and \".\" when training an ASR model, but can the model predict these characters correctly?</p>",
      "rawMarkdown": "Congratulations on your solo gold and becoming a grandmaster.\nThank you for sharing a detailed solution. A small question, you said you retained \",\" and \".\" when training an ASR model, but can the model predict these characters correctly?",
      "votes": 1,
      "replies": [
        {
          "id": 2490703,
          "postDate": "2023-10-21T00:50:29.807Z",
          "content": "<p>Thank you very much and congrats on your team's strong finish too!</p>\n<p>That's a good question. I gave dot and hyphen for ASR model to learn because they are the most problematic punctuations in the annotations.  Since their absence break a word into subwords, with WER metric the score detoriates signifincantly compared to punctuations only attached to one word. And looking at validation result, the model was actually able to learn some of the patterns. But you need to combine training with external data that have dot annotations (ex: MADASR), since competition data have a lot of hyphens but very few abbreviation dots.</p>",
          "rawMarkdown": "Thank you very much and congrats on your team's strong finish too!\n\nThat's a good question. I gave dot and hyphen for ASR model to learn because they are the most problematic punctuations in the annotations.  Since their absence break a word into subwords, with WER metric the score detoriates signifincantly compared to punctuations only attached to one word. And looking at validation result, the model was actually able to learn some of the patterns. But you need to combine training with external data that have dot annotations (ex: MADASR), since competition data have a lot of hyphens but very few abbreviation dots.",
          "votes": 1,
          "replies": [
            {
              "id": 2490725,
              "postDate": "2023-10-21T01:35:43.863Z",
              "content": "<p>I see. We made a punctuation model to predict them because they don't have apparent sounds, but it seems we should try. Thank you, I learned a lot!</p>",
              "rawMarkdown": "I see. We made a punctuation model to predict them because they don't have apparent sounds, but it seems we should try. Thank you, I learned a lot!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2488154,
      "postDate": "2023-10-19T04:30:14.157Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/qdv206\" target=\"_blank\">@qdv206</a> on your amazing achievement. I had some questions regarding building a n-gram Kenlm language model to be used with Beam Search.</p>\n<p>Let's say I have trained a wav2vec2 STT model using a custom Bengali vocabulary set of 70 Bengali unicodes and special tokens which is different from the vocabulary (vocab.json) of the pre-trained wav2vec2 STT model like ai4bharat/indicwav2vec_v1_bengali. How should I train a custom n-gram Kenlm model? As i have found that open source Kenlm models, for example: <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model</a> degrade the inference result when used with wav2vec2 STT models with custom vocabulary. Any suggestions regarding this problem?</p>\n<p>Apologies for any ignorant questions.</p>",
      "rawMarkdown": "Congratulations @qdv206 on your amazing achievement. I had some questions regarding building a n-gram Kenlm language model to be used with Beam Search.\n\nLet's say I have trained a wav2vec2 STT model using a custom Bengali vocabulary set of 70 Bengali unicodes and special tokens which is different from the vocabulary (vocab.json) of the pre-trained wav2vec2 STT model like ai4bharat/indicwav2vec_v1_bengali. How should I train a custom n-gram Kenlm model? As i have found that open source Kenlm models, for example: https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model degrade the inference result when used with wav2vec2 STT models with custom vocabulary. Any suggestions regarding this problem?\n\nApologies for any ignorant questions.",
      "votes": 1
    },
    {
      "id": 2487454,
      "postDate": "2023-10-18T15:45:27.270Z",
      "content": "<p>Great work and Congrats on becoming GM!! 🥳</p>\n<p>Just one small question, How much boost did you get with these augmentations settings, I never tried augs using ASR models before. </p>",
      "rawMarkdown": "Great work and Congrats on becoming GM!! 🥳\n\nJust one small question, How much boost did you get with these augmentations settings, I never tried augs using ASR models before. \n",
      "votes": 1,
      "replies": [
        {
          "id": 2487640,
          "postDate": "2023-10-18T17:33:42.213Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> , congrats on breaking into global top10, what an impressive run!</p>\n<p>Sorry I couldn't provide you with exact numbers as I fixed my augmentations about a month ago. I will instead try to talk about my reasonings behind these augs. When I spent some time listening to different samples of read speech (train data) and spontaneous speech (test data), I noticed that spontaneous speech generally:  </p>\n<ul>\n<li>is spoken with a much more diverse range of speed (can be much faster)</li>\n<li>contains various background noises</li>\n</ul>\n<p>So logically augmentations that simulate these situations should improve results. Then external data were added, some of which were spontaneous speech, so the same logic doesn't apply and naturally a separate scheme must be made for them.</p>\n<p>What I want to say is that most of the augmentations are there because logically it makes sense for them to be there and not the results of some grid search experiments. I try to take the same approach in CV comps too. Hopefully that can be interesting to you.</p>",
          "rawMarkdown": "Thanks @nischaydnk , congrats on breaking into global top10, what an impressive run!\n\nSorry I couldn't provide you with exact numbers as I fixed my augmentations about a month ago. I will instead try to talk about my reasonings behind these augs. When I spent some time listening to different samples of read speech (train data) and spontaneous speech (test data), I noticed that spontaneous speech generally:  \n- is spoken with a much more diverse range of speed (can be much faster)\n- contains various background noises\n\nSo logically augmentations that simulate these situations should improve results. Then external data were added, some of which were spontaneous speech, so the same logic doesn't apply and naturally a separate scheme must be made for them.\n\nWhat I want to say is that most of the augmentations are there because logically it makes sense for them to be there and not the results of some grid search experiments. I try to take the same approach in CV comps too. Hopefully that can be interesting to you.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2486856,
      "postDate": "2023-10-18T07:44:45.380Z",
      "content": "<p>Such neat work! Congratulations on becoming a GM!</p>",
      "rawMarkdown": "Such neat work! Congratulations on becoming a GM!",
      "votes": 1
    },
    {
      "id": 2486741,
      "postDate": "2023-10-18T06:32:10Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/qdv206\" target=\"_blank\">@qdv206</a> , Congrats on becoming a GM !, well deserved</p>",
      "rawMarkdown": "Great work @qdv206 , Congrats on becoming a GM !, well deserved",
      "votes": 1,
      "replies": [
        {
          "id": 2486783,
          "postDate": "2023-10-18T06:58:53.377Z",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> . You are well on your way too! Soon the old Icecube team will be full gold 😆</p>",
          "rawMarkdown": "Thanks a lot @ahmedelfazouan . You are well on your way too! Soon the old Icecube team will be full gold 😆",
          "votes": 2
        }
      ]
    },
    {
      "id": 2486584,
      "postDate": "2023-10-18T03:44:32.547Z",
      "content": "<p>Great job! Congratulation</p>",
      "rawMarkdown": "Great job! Congratulation",
      "votes": 1,
      "replies": [
        {
          "id": 2486779,
          "postDate": "2023-10-18T06:57:25.540Z",
          "content": "<p>Thanks a lot bro, huge congrats to you too!</p>",
          "rawMarkdown": "Thanks a lot bro, huge congrats to you too!"
        }
      ]
    },
    {
      "id": 2486553,
      "postDate": "2023-10-18T02:45:56.020Z",
      "content": "<p>Congratulations on another gold medal and becoming Grandmasters</p>",
      "rawMarkdown": "Congratulations on another gold medal and becoming Grandmasters",
      "votes": 1
    },
    {
      "id": 2754842,
      "postDate": "2024-04-16T08:24:38.753Z",
      "content": "<p>can you introduce your general idea and why you chose these three models? i am an undergraduate student and i would like to learn more about the</p>",
      "rawMarkdown": "can you introduce your general idea and why you chose these three models? i am an undergraduate student and i would like to learn more about the"
    },
    {
      "id": 2493203,
      "postDate": "2023-10-23T08:02:42.773Z",
      "content": "<p>Congratulations on top position in this competition. Thanks for sharing your solution details.</p>",
      "rawMarkdown": "Congratulations on top position in this competition. Thanks for sharing your solution details."
    },
    {
      "id": 2486568,
      "postDate": "2023-10-18T03:10:57.377Z",
      "content": "<p>Can you provide training hardware as well as training time for each model?</p>",
      "rawMarkdown": "Can you provide training hardware as well as training time for each model?",
      "replies": [
        {
          "id": 2486786,
          "postDate": "2023-10-18T07:01:22.823Z",
          "content": "<p>I used V100s to train my models. For ASR normally I only sampled 30% of the data for experimentations, but the final full train took ~5days. Training punctuation models are relatively quick though, only a few hours.</p>",
          "rawMarkdown": "I used V100s to train my models. For ASR normally I only sampled 30% of the data for experimentations, but the final full train took ~5days. Training punctuation models are relatively quick though, only a few hours.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2492713,
      "postDate": "2023-10-22T19:51:38.753Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2487847,
      "postDate": "2023-10-18T19:47:43.343Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2490732,
      "author_name": "Ranamalla Nithin Reddy",
      "author_url": "",
      "post_date": "2023-10-21T01:52:55.477000",
      "content": "<p>Such neat work! Congratulations on becoming a GM <a href=\"https://www.kaggle.com/qdv206\" target=\"_blank\">@qdv206</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2489780,
      "author_name": "luddite^",
      "author_url": "",
      "post_date": "2023-10-20T07:40:18.207000",
      "content": "<p>Congratulations on your solo gold and becoming a grandmaster.<br>\nThank you for sharing a detailed solution. A small question, you said you retained \",\" and \".\" when training an ASR model, but can the model predict these characters correctly?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2490703,
          "author_name": "qdv206",
          "author_url": "",
          "post_date": "2023-10-21T00:50:29.807000",
          "content": "<p>Thank you very much and congrats on your team's strong finish too!</p>\n<p>That's a good question. I gave dot and hyphen for ASR model to learn because they are the most problematic punctuations in the annotations.  Since their absence break a word into subwords, with WER metric the score detoriates signifincantly compared to punctuations only attached to one word. And looking at validation result, the model was actually able to learn some of the patterns. But you need to combine training with external data that have dot annotations (ex: MADASR), since competition data have a lot of hyphens but very few abbreviation dots.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2490725,
              "author_name": "luddite^",
              "author_url": "",
              "post_date": "2023-10-21T01:35:43.863000",
              "content": "<p>I see. We made a punctuation model to predict them because they don't have apparent sounds, but it seems we should try. Thank you, I learned a lot!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2488154,
      "author_name": "Fahim Shahriar Khan",
      "author_url": "",
      "post_date": "2023-10-19T04:30:14.157000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/qdv206\" target=\"_blank\">@qdv206</a> on your amazing achievement. I had some questions regarding building a n-gram Kenlm language model to be used with Beam Search.</p>\n<p>Let's say I have trained a wav2vec2 STT model using a custom Bengali vocabulary set of 70 Bengali unicodes and special tokens which is different from the vocabulary (vocab.json) of the pre-trained wav2vec2 STT model like ai4bharat/indicwav2vec_v1_bengali. How should I train a custom n-gram Kenlm model? As i have found that open source Kenlm models, for example: <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model</a> degrade the inference result when used with wav2vec2 STT models with custom vocabulary. Any suggestions regarding this problem?</p>\n<p>Apologies for any ignorant questions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2487454,
      "author_name": "Nischay Dhankhar",
      "author_url": "",
      "post_date": "2023-10-18T15:45:27.270000",
      "content": "<p>Great work and Congrats on becoming GM!! 🥳</p>\n<p>Just one small question, How much boost did you get with these augmentations settings, I never tried augs using ASR models before. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2487640,
          "author_name": "qdv206",
          "author_url": "",
          "post_date": "2023-10-18T17:33:42.213000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> , congrats on breaking into global top10, what an impressive run!</p>\n<p>Sorry I couldn't provide you with exact numbers as I fixed my augmentations about a month ago. I will instead try to talk about my reasonings behind these augs. When I spent some time listening to different samples of read speech (train data) and spontaneous speech (test data), I noticed that spontaneous speech generally:  </p>\n<ul>\n<li>is spoken with a much more diverse range of speed (can be much faster)</li>\n<li>contains various background noises</li>\n</ul>\n<p>So logically augmentations that simulate these situations should improve results. Then external data were added, some of which were spontaneous speech, so the same logic doesn't apply and naturally a separate scheme must be made for them.</p>\n<p>What I want to say is that most of the augmentations are there because logically it makes sense for them to be there and not the results of some grid search experiments. I try to take the same approach in CV comps too. Hopefully that can be interesting to you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2486856,
      "author_name": "Md Boktiar Mahbub Murad",
      "author_url": "",
      "post_date": "2023-10-18T07:44:45.380000",
      "content": "<p>Such neat work! Congratulations on becoming a GM!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2486741,
      "author_name": "Ahmed El Fazouani",
      "author_url": "",
      "post_date": "2023-10-18T06:32:10",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/qdv206\" target=\"_blank\">@qdv206</a> , Congrats on becoming a GM !, well deserved</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2486783,
          "author_name": "qdv206",
          "author_url": "",
          "post_date": "2023-10-18T06:58:53.377000",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> . You are well on your way too! Soon the old Icecube team will be full gold 😆</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2486584,
      "author_name": "Anh Pham",
      "author_url": "",
      "post_date": "2023-10-18T03:44:32.547000",
      "content": "<p>Great job! Congratulation</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2486779,
          "author_name": "qdv206",
          "author_url": "",
          "post_date": "2023-10-18T06:57:25.540000",
          "content": "<p>Thanks a lot bro, huge congrats to you too!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2486553,
      "author_name": "datnt114",
      "author_url": "",
      "post_date": "2023-10-18T02:45:56.020000",
      "content": "<p>Congratulations on another gold medal and becoming Grandmasters</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2754842,
      "author_name": "nanshui",
      "author_url": "",
      "post_date": "2024-04-16T08:24:38.753000",
      "content": "<p>can you introduce your general idea and why you chose these three models? i am an undergraduate student and i would like to learn more about the</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2493203,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-10-23T08:02:42.773000",
      "content": "<p>Congratulations on top position in this competition. Thanks for sharing your solution details.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2486568,
      "author_name": "datnt114",
      "author_url": "",
      "post_date": "2023-10-18T03:10:57.377000",
      "content": "<p>Can you provide training hardware as well as training time for each model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2486786,
          "author_name": "qdv206",
          "author_url": "",
          "post_date": "2023-10-18T07:01:22.823000",
          "content": "<p>I used V100s to train my models. For ASR normally I only sampled 30% of the data for experimentations, but the final full train took ~5days. Training punctuation models are relatively quick though, only a few hours.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2492713,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-22T19:51:38.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2487847,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:47:43.343000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486531": "Many thanks to Kaggle and Bengali.AI for hosting such an interesting competition. As I was about to start my PhD  research on speech processing and I knew nothing about the field, the competition came as the perfect opportunity for me to learn. In the end it was a highly rewarding experience.\n\nThe solution consists of 3 components:\n- ASR model\n- Language model\n- Punctuation model\n\n**1. ASR model**\n\nI used `ai4bharat/indicwav2vec_v1_bengali` as pretrained model.\n\n**Datasets:**\n- Speech data: Competition data, Shrutilipi, MADASR, ULCA (for ULCA data most of the links are dead but some are still downloadable, only a few thousands samples though)\n- Noise data: music data from MUSAN and noise data from DNS Challenge 2020\n- All data are normalized and punctuation marks are removed except for dot (`.`) and hyphen (`-`)\n\n**Augmentation:**\n- use augmentation from `audiomentations`, apply heavy augmentation for read speech (comp data, MADASR) and lighter augmentation for spontaneous speech (Shrutilipi, ULCA)\nExample of read speech augmentation:\n```\naugments = Compose([\n    TimeStretch(min_rate=0.8, max_rate=2.0, p=0.5, leave_length_unchanged=False),\n    RoomSimulator(p=0.3),\n    OneOf([\n        AddBackgroundNoise(\n            sounds_path=[\n                '/path_to_DNS_Challenge_noise',\n            ],\n            min_snr_in_db=5.0,\n            max_snr_in_db=30.0,\n            noise_transform=PolarityInversion(),\n            p=1.0\n        ),\n        AddBackgroundNoise(\n            sounds_path=[\n                '/path_to_MUSAN_music'\n            ],\n            min_snr_in_db=5.0,\n            max_snr_in_db=30.0,\n            noise_transform=PolarityInversion(),\n            p=1.0\n        ),\n        AddGaussianNoise(min_amplitude=0.005, max_amplitude=0.015, p=1.0),\n    ], p=0.7),\n    Gain(min_gain_in_db=-6, max_gain_in_db=6, p=0.2),\n    ])\n```\nFor spontaneous speech augment probabilities are smaller and `TimeStretch` rate much less extreme.\n- concat augment: randomly concatenate short samples together to make length distribution of training set closer to OOD test set.\n- SpecAugment: mask_time_prob = 0.1, mask_feature_prob = 0.05.\n\n**Training:**\n- First fit on all training data, then remove about 10% with highest WER after fitted from train set.\n- Don't freeze feature encoder.\n- Use cosine schedule with warmups and restarts: 1st cycle 5 epochs peak lr 4e-5, 2nd cycle 3 epochs peak lr 3e-5, third cycle 3 epochs peak lr 2e-5.\n\n**Inference:**\nUse `AutomaticSpeechRecognitionPipeline` from `transformers` to apply inference with chunking and stride:\n```\ntext = pipe(w, chunk_length_s=14, stride_length_s=(6, 3))[\"text\"]\n```\n\n**2. Language model**\n\n6-gram kenlm model trained on multiple external Bengali corpus:\n- IndicCorp V1+V2.\n- Bharat Parallel Corpus Collection.\n- Samanantar.\n- [Bengali poetry dataset](https://www.kaggle.com/datasets/truthr/free-bengali-poetry).\n- [WMT News Crawl](https://data.statmt.org/news-crawl/).\n- Hate speech corpus from https://github.com/rezacsedu/Classification_Benchmarks_Benglai_NLP.\n\n**3. Punctuation model**\n\nTrain token classification model to add the following punctuation set: `।,?!`\n\n- use `ai4bharat/IndicBERTv2-MLM-Sam-TLM` as backbone\n- add LSTM head\n- train for 6 epochs, cosine schedule, lr 3e-5 on competition data + subset of IndicCorp\n- mask 15% of the tokens during training as augmentation\n- ensemble 3 folds of model trained on 3 different subsets of IndicCorp\n- beam search decoding for inference.\n\nThank you very much for reading and please let me know if you have any questions.\n\nUpdate: \n- training code: https://github.com/quangdao206/Kaggle_Bengali_Speech_Recognition_2nd_Place_Solution\n- inference notebook: https://www.kaggle.com/code/qdv206/2nd-place-bengali-speech-infer/\n\n",
    "2490732": "Such neat work! Congratulations on becoming a GM @qdv206 ",
    "2489780": "Congratulations on your solo gold and becoming a grandmaster.\nThank you for sharing a detailed solution. A small question, you said you retained \",\" and \".\" when training an ASR model, but can the model predict these characters correctly?",
    "2488154": "Congratulations @qdv206 on your amazing achievement. I had some questions regarding building a n-gram Kenlm language model to be used with Beam Search.\n\nLet's say I have trained a wav2vec2 STT model using a custom Bengali vocabulary set of 70 Bengali unicodes and special tokens which is different from the vocabulary (vocab.json) of the pre-trained wav2vec2 STT model like ai4bharat/indicwav2vec_v1_bengali. How should I train a custom n-gram Kenlm model? As i have found that open source Kenlm models, for example: https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali/tree/main/language_model degrade the inference result when used with wav2vec2 STT models with custom vocabulary. Any suggestions regarding this problem?\n\nApologies for any ignorant questions.",
    "2487454": "Great work and Congrats on becoming GM!! 🥳\n\nJust one small question, How much boost did you get with these augmentations settings, I never tried augs using ASR models before. \n",
    "2486856": "Such neat work! Congratulations on becoming a GM!",
    "2486741": "Great work @qdv206 , Congrats on becoming a GM !, well deserved",
    "2486584": "Great job! Congratulation",
    "2486553": "Congratulations on another gold medal and becoming Grandmasters",
    "2754842": "can you introduce your general idea and why you chose these three models? i am an undergraduate student and i would like to learn more about the",
    "2493203": "Congratulations on top position in this competition. Thanks for sharing your solution details.",
    "2486568": "Can you provide training hardware as well as training time for each model?",
    "2492713": "",
    "2487847": ""
  }
}