{
  "id": 448118,
  "title": "67th Place Solution - Bronze - LB 0.515",
  "url": "/competitions/bengaliai-speech/discussion/448118",
  "author_name": "Umong Sain",
  "post_date": "2023-10-18T14:31:38.156000",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, thank you to the organizers for arranging this exciting competition, and congratulations to the winners. I have learned a lot throughout this journey. The discussion section and public notebooks were quite helpful. I feel a bit regretful because I couldn't dedicate enough time due to my final exams.</p>\n<h2>My Approach</h2>\n<p>A few months back, I improved <a href=\"https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference\" target=\"_blank\">the public notebook</a> by <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> just by training my own <a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\" target=\"_blank\">n-gram language model</a>. It improved both my public LB (0.445 -&gt; 0.436) and private LB (0.534 -&gt; 0.515).</p>\n<h2>Preparing the Data</h2>\n<p>I normalized the training texts using <a href=\"https://github.com/mnansary/bnUnicodeNormalizer\" target=\"_blank\">bnunicodenormalizer</a> as well as the normalizer developed by <a href=\"https://github.com/csebuetnlp/normalizer\" target=\"_blank\">CSE BUET NLP</a>. Some data could not be normalized because of the confusion which contained the character \" ঃ\" (<a href=\"https://bn.wikipedia.org/wiki/%E0%A6%AC%E0%A6%BF%E0%A6%B8%E0%A6%B0%E0%A7%8D%E0%A6%97\" target=\"_blank\">বিসর্গ</a>) as a separate word. As it is mostly used in place of \":\" (colon), I replaced them with colons instead.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4772132%2Fc86799e65036ebc9228f9f709a98a309%2FScreenshot%20from%202023-10-20%2020-15-23.png?generation=1697811401235581&amp;alt=media\" alt=\"\"></p>\n<h2>Training the N-gram Model</h2>\n<p>Before training the n-gram model, I removed the punctuation from the normalized text data because Wav2Vec2 does not predict any punctuation. I trained the n-gram model using <a href=\"https://github.com/kensho-technologies/pyctcdecode\" target=\"_blank\">pyctcdecode</a>.</p>\n<h2>Solution Code</h2>\n<ul>\n<li>Data preparation: <a href=\"https://www.kaggle.com/code/umongsain/macro-normalization\" target=\"_blank\">https://www.kaggle.com/code/umongsain/macro-normalization</a></li>\n<li>Training code: <a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro/notebook\" target=\"_blank\">https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro/notebook</a></li>\n<li>Inference code: <a href=\"https://www.kaggle.com/code/umongsain/67th-place-solution-inference-notebook\" target=\"_blank\">https://www.kaggle.com/code/umongsain/67th-place-solution-inference-notebook</a></li>\n</ul>\n<h2>Things I Tried But Did Not Work</h2>\n<p>Later, I tried to fine-tune multiple XLS-R and MMS models from scratch. As I had limited resources, I focused on low-resource training, using only 30k samples from the Common Voice 13. Unfortunately, they could not beat my previous baseline. But what I found is that <a href=\"https://huggingface.co/facebook/mms-1b\" target=\"_blank\">MMS 1B</a> works quite well for low-resource languages. Sadly, it is available only for non-commercial uses. The checkpoints can be found here:</p>\n<ul>\n<li><a href=\"https://huggingface.co/Umong/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">https://huggingface.co/Umong/wav2vec2-xls-r-300m-bengali</a></li>\n<li><a href=\"https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali\" target=\"_blank\">https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali</a></li>\n</ul>\n<h2>Things That I Could Have Tried</h2>\n<p>I had plans to use QLoRa to fine-tune on a few out-of-domain (OOD) hand-labeled samples for domain adaptation after fine-tuning on the crowd-sourced data. I wish I had a bit more free time.</p>\n<p>I am glad to have my 1st Bronze competition medal. Again, thanks to the organizers.</p>",
  "messages": [
    {
      "id": 2487361,
      "postDate": "2023-10-18T14:31:38.157Z",
      "content": "<p>First of all, thank you to the organizers for arranging this exciting competition, and congratulations to the winners. I have learned a lot throughout this journey. The discussion section and public notebooks were quite helpful. I feel a bit regretful because I couldn't dedicate enough time due to my final exams.</p>\n<h2>My Approach</h2>\n<p>A few months back, I improved <a href=\"https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference\" target=\"_blank\">the public notebook</a> by <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> just by training my own <a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\" target=\"_blank\">n-gram language model</a>. It improved both my public LB (0.445 -&gt; 0.436) and private LB (0.534 -&gt; 0.515).</p>\n<h2>Preparing the Data</h2>\n<p>I normalized the training texts using <a href=\"https://github.com/mnansary/bnUnicodeNormalizer\" target=\"_blank\">bnunicodenormalizer</a> as well as the normalizer developed by <a href=\"https://github.com/csebuetnlp/normalizer\" target=\"_blank\">CSE BUET NLP</a>. Some data could not be normalized because of the confusion which contained the character \" ঃ\" (<a href=\"https://bn.wikipedia.org/wiki/%E0%A6%AC%E0%A6%BF%E0%A6%B8%E0%A6%B0%E0%A7%8D%E0%A6%97\" target=\"_blank\">বিসর্গ</a>) as a separate word. As it is mostly used in place of \":\" (colon), I replaced them with colons instead.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4772132%2Fc86799e65036ebc9228f9f709a98a309%2FScreenshot%20from%202023-10-20%2020-15-23.png?generation=1697811401235581&amp;alt=media\" alt=\"\"></p>\n<h2>Training the N-gram Model</h2>\n<p>Before training the n-gram model, I removed the punctuation from the normalized text data because Wav2Vec2 does not predict any punctuation. I trained the n-gram model using <a href=\"https://github.com/kensho-technologies/pyctcdecode\" target=\"_blank\">pyctcdecode</a>.</p>\n<h2>Solution Code</h2>\n<ul>\n<li>Data preparation: <a href=\"https://www.kaggle.com/code/umongsain/macro-normalization\" target=\"_blank\">https://www.kaggle.com/code/umongsain/macro-normalization</a></li>\n<li>Training code: <a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro/notebook\" target=\"_blank\">https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro/notebook</a></li>\n<li>Inference code: <a href=\"https://www.kaggle.com/code/umongsain/67th-place-solution-inference-notebook\" target=\"_blank\">https://www.kaggle.com/code/umongsain/67th-place-solution-inference-notebook</a></li>\n</ul>\n<h2>Things I Tried But Did Not Work</h2>\n<p>Later, I tried to fine-tune multiple XLS-R and MMS models from scratch. As I had limited resources, I focused on low-resource training, using only 30k samples from the Common Voice 13. Unfortunately, they could not beat my previous baseline. But what I found is that <a href=\"https://huggingface.co/facebook/mms-1b\" target=\"_blank\">MMS 1B</a> works quite well for low-resource languages. Sadly, it is available only for non-commercial uses. The checkpoints can be found here:</p>\n<ul>\n<li><a href=\"https://huggingface.co/Umong/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">https://huggingface.co/Umong/wav2vec2-xls-r-300m-bengali</a></li>\n<li><a href=\"https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali\" target=\"_blank\">https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali</a></li>\n</ul>\n<h2>Things That I Could Have Tried</h2>\n<p>I had plans to use QLoRa to fine-tune on a few out-of-domain (OOD) hand-labeled samples for domain adaptation after fine-tuning on the crowd-sourced data. I wish I had a bit more free time.</p>\n<p>I am glad to have my 1st Bronze competition medal. Again, thanks to the organizers.</p>",
      "rawMarkdown": "First of all, thank you to the organizers for arranging this exciting competition, and congratulations to the winners. I have learned a lot throughout this journey. The discussion section and public notebooks were quite helpful. I feel a bit regretful because I couldn't dedicate enough time due to my final exams.\n\n## My Approach\nA few months back, I improved [the public notebook](https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference) by @nischaydnk just by training my own [n-gram language model](https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro). It improved both my public LB (0.445 -> 0.436) and private LB (0.534 -> 0.515).\n\n## Preparing the Data\nI normalized the training texts using [bnunicodenormalizer](https://github.com/mnansary/bnUnicodeNormalizer) as well as the normalizer developed by [CSE BUET NLP](https://github.com/csebuetnlp/normalizer). Some data could not be normalized because of the confusion which contained the character \" ঃ\" ([বিসর্গ](https://bn.wikipedia.org/wiki/%E0%A6%AC%E0%A6%BF%E0%A6%B8%E0%A6%B0%E0%A7%8D%E0%A6%97)) as a separate word. As it is mostly used in place of \":\" (colon), I replaced them with colons instead.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4772132%2Fc86799e65036ebc9228f9f709a98a309%2FScreenshot%20from%202023-10-20%2020-15-23.png?generation=1697811401235581&alt=media)\n\n## Training the N-gram Model\nBefore training the n-gram model, I removed the punctuation from the normalized text data because Wav2Vec2 does not predict any punctuation. I trained the n-gram model using [pyctcdecode](https://github.com/kensho-technologies/pyctcdecode).\n\n## Solution Code\n- Data preparation: https://www.kaggle.com/code/umongsain/macro-normalization\n- Training code: https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro/notebook\n- Inference code: https://www.kaggle.com/code/umongsain/67th-place-solution-inference-notebook\n\n## Things I Tried But Did Not Work\nLater, I tried to fine-tune multiple XLS-R and MMS models from scratch. As I had limited resources, I focused on low-resource training, using only 30k samples from the Common Voice 13. Unfortunately, they could not beat my previous baseline. But what I found is that [MMS 1B](https://huggingface.co/facebook/mms-1b) works quite well for low-resource languages. Sadly, it is available only for non-commercial uses. The checkpoints can be found here:\n\n- https://huggingface.co/Umong/wav2vec2-xls-r-300m-bengali\n- https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali\n\n## Things That I Could Have Tried\nI had plans to use QLoRa to fine-tune on a few out-of-domain (OOD) hand-labeled samples for domain adaptation after fine-tuning on the crowd-sourced data. I wish I had a bit more free time.\n\nI am glad to have my 1st Bronze competition medal. Again, thanks to the organizers.",
      "votes": 9
    },
    {
      "id": 2487865,
      "postDate": "2023-10-18T20:00:03.470Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2487865,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T20:00:03.470000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2487361": "First of all, thank you to the organizers for arranging this exciting competition, and congratulations to the winners. I have learned a lot throughout this journey. The discussion section and public notebooks were quite helpful. I feel a bit regretful because I couldn't dedicate enough time due to my final exams.\n\n## My Approach\nA few months back, I improved [the public notebook](https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference) by @nischaydnk just by training my own [n-gram language model](https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro). It improved both my public LB (0.445 -> 0.436) and private LB (0.534 -> 0.515).\n\n## Preparing the Data\nI normalized the training texts using [bnunicodenormalizer](https://github.com/mnansary/bnUnicodeNormalizer) as well as the normalizer developed by [CSE BUET NLP](https://github.com/csebuetnlp/normalizer). Some data could not be normalized because of the confusion which contained the character \" ঃ\" ([বিসর্গ](https://bn.wikipedia.org/wiki/%E0%A6%AC%E0%A6%BF%E0%A6%B8%E0%A6%B0%E0%A7%8D%E0%A6%97)) as a separate word. As it is mostly used in place of \":\" (colon), I replaced them with colons instead.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4772132%2Fc86799e65036ebc9228f9f709a98a309%2FScreenshot%20from%202023-10-20%2020-15-23.png?generation=1697811401235581&alt=media)\n\n## Training the N-gram Model\nBefore training the n-gram model, I removed the punctuation from the normalized text data because Wav2Vec2 does not predict any punctuation. I trained the n-gram model using [pyctcdecode](https://github.com/kensho-technologies/pyctcdecode).\n\n## Solution Code\n- Data preparation: https://www.kaggle.com/code/umongsain/macro-normalization\n- Training code: https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro/notebook\n- Inference code: https://www.kaggle.com/code/umongsain/67th-place-solution-inference-notebook\n\n## Things I Tried But Did Not Work\nLater, I tried to fine-tune multiple XLS-R and MMS models from scratch. As I had limited resources, I focused on low-resource training, using only 30k samples from the Common Voice 13. Unfortunately, they could not beat my previous baseline. But what I found is that [MMS 1B](https://huggingface.co/facebook/mms-1b) works quite well for low-resource languages. Sadly, it is available only for non-commercial uses. The checkpoints can be found here:\n\n- https://huggingface.co/Umong/wav2vec2-xls-r-300m-bengali\n- https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali\n\n## Things That I Could Have Tried\nI had plans to use QLoRa to fine-tune on a few out-of-domain (OOD) hand-labeled samples for domain adaptation after fine-tuning on the crowd-sourced data. I wish I had a bit more free time.\n\nI am glad to have my 1st Bronze competition medal. Again, thanks to the organizers.",
    "2487865": ""
  }
}