{
  "id": 448030,
  "title": "31st place silver medal solution - My first competition medal",
  "url": "/competitions/bengaliai-speech/discussion/448030",
  "author_name": "Md Boktiar Mahbub Murad",
  "post_date": "2023-10-18T06:41:13.158000",
  "votes": 25,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First, we want to thank Bengali.ai and Kaggle for hosting such an excellent competition on Bengali ASR and releasing a large-scale dataset for this domain. This competition was challenging, to say the least. But I enjoyed the last three months working on it ngl.  It was a pretty good learning curve for me. I got the opportunity to contribute to the community and also engage with the community through this competition. I got 1 Gold, 2 silver and 8 bronze medals for my notebooks in this competition. So overall, a very busy and happy three months!</p>\n<h2>Summary :</h2>\n<p><strong>Acoustic Model</strong>: We fine-tuned the ai4bharat/indicwav2vec_v1_bengali model<br>\n<strong>Dataset</strong> : Competition Data + Openslr53+openslr37+ TTS dataset<br>\nIn the first phase, we fine-tuned the model with the whole competition data. which was making things worse, dragging the LB performance below the best public NB ( 0.445)<br>\nThen we filtered out the bad-quality audio from the competition data using the training metadata provided by the host ( We took the audios with MOS&gt;=2). <br>\nWe did a random split with both the train and validation data(since it was evident MaCro validation audio quality was better than train audio)  and trained with 95% of the audios. Then added the other datasets. This improved the LB score to 0.430<br>\nAugmentation : reverberation, Speed perturbation,Volume perturbation (0.125x ~ 2.0x)<br>\nAdding background noise from the example audios(Improved performance slightly)</p>\n<p><strong>Language Model</strong>: We built a 5-gram LM using competition sentences + <a href=\"https://github.com/csebuetnlp/banglanmt\" target=\"_blank\">banglanmt</a> + IndicCorp_V2</p>\n<h2>What did not work:</h2>\n<ol>\n<li><strong>Punctuation model</strong> : I tried to add the <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">xashru-punctuation-restoration</a> model but it always gave CUDA OOM. Now it's bugging me to see others implemented it and got a huge boost with it. I also tried to fine-tune T5 model for this task but didn't see much improvement. So I gave up the idea.</li>\n<li><strong>Speech enhancement</strong> : </li>\n</ol>\n<ul>\n<li><a href=\"https://github.com/Rikorose/DeepFilterNet\" target=\"_blank\">DeepFIlterNet</a></li>\n<li><a href=\"https://huggingface.co/speechbrain/sepformer-wham-enhancement\" target=\"_blank\">Speechbrain wham</a></li>\n</ul>\n<p>I spent a significant amount of time on them. First I tried to add a speech enhancer model to my inference pipeline. It didn't improve the performance. Then I tried to train with enhanced audio, but it didn't help either. The noise cancellation performance by these models was quite good, but it also affected the pitch and distorted the speech a little bit. </p>\n<p><strong>What could have been done to improve further/ Next steps</strong> : </p>\n<ol>\n<li>Adding the punctuation model.</li>\n<li>Train the acoustic model with more data (MADASR,Shrutilipi)</li>\n<li>Add more texts to the LM</li>\n</ol>",
  "messages": [
    {
      "id": 2486751,
      "postDate": "2023-10-18T06:41:13.160Z",
      "content": "<p>First, we want to thank Bengali.ai and Kaggle for hosting such an excellent competition on Bengali ASR and releasing a large-scale dataset for this domain. This competition was challenging, to say the least. But I enjoyed the last three months working on it ngl.  It was a pretty good learning curve for me. I got the opportunity to contribute to the community and also engage with the community through this competition. I got 1 Gold, 2 silver and 8 bronze medals for my notebooks in this competition. So overall, a very busy and happy three months!</p>\n<h2>Summary :</h2>\n<p><strong>Acoustic Model</strong>: We fine-tuned the ai4bharat/indicwav2vec_v1_bengali model<br>\n<strong>Dataset</strong> : Competition Data + Openslr53+openslr37+ TTS dataset<br>\nIn the first phase, we fine-tuned the model with the whole competition data. which was making things worse, dragging the LB performance below the best public NB ( 0.445)<br>\nThen we filtered out the bad-quality audio from the competition data using the training metadata provided by the host ( We took the audios with MOS&gt;=2). <br>\nWe did a random split with both the train and validation data(since it was evident MaCro validation audio quality was better than train audio)  and trained with 95% of the audios. Then added the other datasets. This improved the LB score to 0.430<br>\nAugmentation : reverberation, Speed perturbation,Volume perturbation (0.125x ~ 2.0x)<br>\nAdding background noise from the example audios(Improved performance slightly)</p>\n<p><strong>Language Model</strong>: We built a 5-gram LM using competition sentences + <a href=\"https://github.com/csebuetnlp/banglanmt\" target=\"_blank\">banglanmt</a> + IndicCorp_V2</p>\n<h2>What did not work:</h2>\n<ol>\n<li><strong>Punctuation model</strong> : I tried to add the <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">xashru-punctuation-restoration</a> model but it always gave CUDA OOM. Now it's bugging me to see others implemented it and got a huge boost with it. I also tried to fine-tune T5 model for this task but didn't see much improvement. So I gave up the idea.</li>\n<li><strong>Speech enhancement</strong> : </li>\n</ol>\n<ul>\n<li><a href=\"https://github.com/Rikorose/DeepFilterNet\" target=\"_blank\">DeepFIlterNet</a></li>\n<li><a href=\"https://huggingface.co/speechbrain/sepformer-wham-enhancement\" target=\"_blank\">Speechbrain wham</a></li>\n</ul>\n<p>I spent a significant amount of time on them. First I tried to add a speech enhancer model to my inference pipeline. It didn't improve the performance. Then I tried to train with enhanced audio, but it didn't help either. The noise cancellation performance by these models was quite good, but it also affected the pitch and distorted the speech a little bit. </p>\n<p><strong>What could have been done to improve further/ Next steps</strong> : </p>\n<ol>\n<li>Adding the punctuation model.</li>\n<li>Train the acoustic model with more data (MADASR,Shrutilipi)</li>\n<li>Add more texts to the LM</li>\n</ol>",
      "rawMarkdown": "First, we want to thank Bengali.ai and Kaggle for hosting such an excellent competition on Bengali ASR and releasing a large-scale dataset for this domain. This competition was challenging, to say the least. But I enjoyed the last three months working on it ngl.  It was a pretty good learning curve for me. I got the opportunity to contribute to the community and also engage with the community through this competition. I got 1 Gold, 2 silver and 8 bronze medals for my notebooks in this competition. So overall, a very busy and happy three months!\n\n## Summary : \n\n**Acoustic Model**: We fine-tuned the ai4bharat/indicwav2vec_v1_bengali model\n**Dataset** : Competition Data + Openslr53+openslr37+ TTS dataset\nIn the first phase, we fine-tuned the model with the whole competition data. which was making things worse, dragging the LB performance below the best public NB ( 0.445)\nThen we filtered out the bad-quality audio from the competition data using the training metadata provided by the host ( We took the audios with MOS>=2). \nWe did a random split with both the train and validation data(since it was evident MaCro validation audio quality was better than train audio)  and trained with 95% of the audios. Then added the other datasets. This improved the LB score to 0.430\nAugmentation : reverberation, Speed perturbation,Volume perturbation (0.125x ~ 2.0x)\nAdding background noise from the example audios(Improved performance slightly)\n\n**Language Model**: We built a 5-gram LM using competition sentences + [banglanmt](https://github.com/csebuetnlp/banglanmt) + IndicCorp_V2\n\n## What did not work: \n1. **Punctuation model** : I tried to add the [xashru-punctuation-restoration](https://github.com/xashru/punctuation-restoration) model but it always gave CUDA OOM. Now it's bugging me to see others implemented it and got a huge boost with it. I also tried to fine-tune T5 model for this task but didn't see much improvement. So I gave up the idea.\n2. **Speech enhancement** : \n- [DeepFIlterNet](https://github.com/Rikorose/DeepFilterNet)\n- [Speechbrain wham](https://huggingface.co/speechbrain/sepformer-wham-enhancement)\n\nI spent a significant amount of time on them. First I tried to add a speech enhancer model to my inference pipeline. It didn't improve the performance. Then I tried to train with enhanced audio, but it didn't help either. The noise cancellation performance by these models was quite good, but it also affected the pitch and distorted the speech a little bit. \n\n**What could have been done to improve further/ Next steps** : \n1. Adding the punctuation model.\n2. Train the acoustic model with more data (MADASR,Shrutilipi)\n3. Add more texts to the LM",
      "votes": 25
    },
    {
      "id": 2488035,
      "postDate": "2023-10-19T01:10:08.940Z",
      "content": "<p>Congratulations!  I have learned a lot from your discussions and notebooks,they take me many new ideas,thank you for your sharing and wish you better in next game!</p>",
      "rawMarkdown": "Congratulations!  I have learned a lot from your discussions and notebooks,they take me many new ideas,thank you for your sharing and wish you better in next game!",
      "votes": 2,
      "replies": [
        {
          "id": 2490196,
          "postDate": "2023-10-20T14:17:58.423Z",
          "content": "<p>Thanks a lot! Best of luck to you too!</p>",
          "rawMarkdown": "Thanks a lot! Best of luck to you too!"
        }
      ]
    },
    {
      "id": 2487268,
      "postDate": "2023-10-18T13:39:48.097Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> for your details EDA notebook. Congratulations for your achievement :)</p>",
      "rawMarkdown": "Thanks @mbmmurad for your details EDA notebook. Congratulations for your achievement :)",
      "votes": 2,
      "replies": [
        {
          "id": 2487394,
          "postDate": "2023-10-18T15:00:46.977Z",
          "content": "<p>Thanks! Congratulations to you too!</p>",
          "rawMarkdown": "Thanks! Congratulations to you too!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2487083,
      "postDate": "2023-10-18T11:00:35.770Z",
      "content": "<p>Thanks for sharing! Your notebooks and discussions were very enlightening to us. Thank you for your contribution!</p>",
      "rawMarkdown": "Thanks for sharing! Your notebooks and discussions were very enlightening to us. Thank you for your contribution!",
      "votes": 2,
      "replies": [
        {
          "id": 2487185,
          "postDate": "2023-10-18T12:11:45.273Z",
          "content": "<p>Thanks for your kind words! Glad I could help.</p>",
          "rawMarkdown": "Thanks for your kind words! Glad I could help.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2486942,
      "postDate": "2023-10-18T08:52:55.077Z",
      "content": "<p>Congratulations! What was surprising about this event was that there was no major shake-up in the leaderboard, unlike in other competitions here.</p>",
      "rawMarkdown": "Congratulations! What was surprising about this event was that there was no major shake-up in the leaderboard, unlike in other competitions here.",
      "votes": 2,
      "replies": [
        {
          "id": 2486968,
          "postDate": "2023-10-18T09:06:08.050Z",
          "content": "<p>True. I think the main reason behind it was the amount of data the top models were trained on. Almost all of them were trained on at least half a million audios, so the models are quite generalised and robust. </p>",
          "rawMarkdown": "True. I think the main reason behind it was the amount of data the top models were trained on. Almost all of them were trained on at least half a million audios, so the models are quite generalised and robust. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2486757,
      "postDate": "2023-10-18T06:47:02.030Z",
      "content": "<p>Congratulations!</p>\n<p>And your notebook was very helpful, thank you so much.🙏</p>",
      "rawMarkdown": "Congratulations!\n\nAnd your notebook was very helpful, thank you so much.🙏",
      "votes": 2,
      "replies": [
        {
          "id": 2486776,
          "postDate": "2023-10-18T06:55:51.620Z",
          "content": "<p>Thanks a lot! I'm glad my NB was helpful to you. Congratulations to you too!</p>",
          "rawMarkdown": "Thanks a lot! I'm glad my NB was helpful to you. Congratulations to you too!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2487860,
      "postDate": "2023-10-18T19:58:21.123Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2488035,
      "author_name": "LIUHAHA1888",
      "author_url": "",
      "post_date": "2023-10-19T01:10:08.940000",
      "content": "<p>Congratulations!  I have learned a lot from your discussions and notebooks,they take me many new ideas,thank you for your sharing and wish you better in next game!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2490196,
          "author_name": "Md Boktiar Mahbub Murad",
          "author_url": "",
          "post_date": "2023-10-20T14:17:58.423000",
          "content": "<p>Thanks a lot! Best of luck to you too!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2487268,
      "author_name": "AIFahim",
      "author_url": "",
      "post_date": "2023-10-18T13:39:48.097000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> for your details EDA notebook. Congratulations for your achievement :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2487394,
          "author_name": "Md Boktiar Mahbub Murad",
          "author_url": "",
          "post_date": "2023-10-18T15:00:46.977000",
          "content": "<p>Thanks! Congratulations to you too!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2487083,
      "author_name": "Mutian Hong",
      "author_url": "",
      "post_date": "2023-10-18T11:00:35.770000",
      "content": "<p>Thanks for sharing! Your notebooks and discussions were very enlightening to us. Thank you for your contribution!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2487185,
          "author_name": "Md Boktiar Mahbub Murad",
          "author_url": "",
          "post_date": "2023-10-18T12:11:45.273000",
          "content": "<p>Thanks for your kind words! Glad I could help.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2486942,
      "author_name": "Abdullah Al Asif",
      "author_url": "",
      "post_date": "2023-10-18T08:52:55.077000",
      "content": "<p>Congratulations! What was surprising about this event was that there was no major shake-up in the leaderboard, unlike in other competitions here.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2486968,
          "author_name": "Md Boktiar Mahbub Murad",
          "author_url": "",
          "post_date": "2023-10-18T09:06:08.050000",
          "content": "<p>True. I think the main reason behind it was the amount of data the top models were trained on. Almost all of them were trained on at least half a million audios, so the models are quite generalised and robust. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2486757,
      "author_name": "neilus",
      "author_url": "",
      "post_date": "2023-10-18T06:47:02.030000",
      "content": "<p>Congratulations!</p>\n<p>And your notebook was very helpful, thank you so much.🙏</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2486776,
          "author_name": "Md Boktiar Mahbub Murad",
          "author_url": "",
          "post_date": "2023-10-18T06:55:51.620000",
          "content": "<p>Thanks a lot! I'm glad my NB was helpful to you. Congratulations to you too!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2487860,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:58:21.123000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486751": "First, we want to thank Bengali.ai and Kaggle for hosting such an excellent competition on Bengali ASR and releasing a large-scale dataset for this domain. This competition was challenging, to say the least. But I enjoyed the last three months working on it ngl.  It was a pretty good learning curve for me. I got the opportunity to contribute to the community and also engage with the community through this competition. I got 1 Gold, 2 silver and 8 bronze medals for my notebooks in this competition. So overall, a very busy and happy three months!\n\n## Summary : \n\n**Acoustic Model**: We fine-tuned the ai4bharat/indicwav2vec_v1_bengali model\n**Dataset** : Competition Data + Openslr53+openslr37+ TTS dataset\nIn the first phase, we fine-tuned the model with the whole competition data. which was making things worse, dragging the LB performance below the best public NB ( 0.445)\nThen we filtered out the bad-quality audio from the competition data using the training metadata provided by the host ( We took the audios with MOS>=2). \nWe did a random split with both the train and validation data(since it was evident MaCro validation audio quality was better than train audio)  and trained with 95% of the audios. Then added the other datasets. This improved the LB score to 0.430\nAugmentation : reverberation, Speed perturbation,Volume perturbation (0.125x ~ 2.0x)\nAdding background noise from the example audios(Improved performance slightly)\n\n**Language Model**: We built a 5-gram LM using competition sentences + [banglanmt](https://github.com/csebuetnlp/banglanmt) + IndicCorp_V2\n\n## What did not work: \n1. **Punctuation model** : I tried to add the [xashru-punctuation-restoration](https://github.com/xashru/punctuation-restoration) model but it always gave CUDA OOM. Now it's bugging me to see others implemented it and got a huge boost with it. I also tried to fine-tune T5 model for this task but didn't see much improvement. So I gave up the idea.\n2. **Speech enhancement** : \n- [DeepFIlterNet](https://github.com/Rikorose/DeepFilterNet)\n- [Speechbrain wham](https://huggingface.co/speechbrain/sepformer-wham-enhancement)\n\nI spent a significant amount of time on them. First I tried to add a speech enhancer model to my inference pipeline. It didn't improve the performance. Then I tried to train with enhanced audio, but it didn't help either. The noise cancellation performance by these models was quite good, but it also affected the pitch and distorted the speech a little bit. \n\n**What could have been done to improve further/ Next steps** : \n1. Adding the punctuation model.\n2. Train the acoustic model with more data (MADASR,Shrutilipi)\n3. Add more texts to the LM",
    "2488035": "Congratulations!  I have learned a lot from your discussions and notebooks,they take me many new ideas,thank you for your sharing and wish you better in next game!",
    "2487268": "Thanks @mbmmurad for your details EDA notebook. Congratulations for your achievement :)",
    "2487083": "Thanks for sharing! Your notebooks and discussions were very enlightening to us. Thank you for your contribution!",
    "2486942": "Congratulations! What was surprising about this event was that there was no major shake-up in the leaderboard, unlike in other competitions here.",
    "2486757": "Congratulations!\n\nAnd your notebook was very helpful, thank you so much.🙏",
    "2487860": ""
  }
}