{
  "id": 451471,
  "title": "57th Place Solution for the Bengali.AI Speech Recognition Competition (Top 8%)",
  "url": "/competitions/bengaliai-speech/writeups/overfitting-butter-chicken-57th-place-solution-for",
  "author_name": "",
  "post_date": "2023-10-29T03:55:02.190Z",
  "votes": 20,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thank you Bengali.AI &amp; Kaggle for organizing this competition. It was a very wonderful learning experience for my team <a href=\"https://www.kaggle.com/kamaruladha\" target=\"_blank\">@kamaruladha</a> <a href=\"https://www.kaggle.com/ariffnazhan\" target=\"_blank\">@ariffnazhan</a>. This is our first bronze medal and first-time dealing with audio data! </p>\n<p>Business context: <a href=\"http://www.kaggle.com/competitions/bengaliai-speech/overview\" target=\"_blank\">www.kaggle.com/competitions/bengaliai-speech/overview</a><br>\nData context: <a href=\"http://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">www.kaggle.com/competitions/bengaliai-speech/data</a></p>\n<h1>Overview of the Approach</h1>\n<h2>ASR Model</h2>\n<p>Our most performing model is from finetuning <code>ai4bharat/indicwav2vec_v1_bengali</code> trained for 2 epochs with training data containing both train and valid split. </p>\n<h2>Dataset</h2>\n<p>We took 80% of training data and 96% of validation split using random sample for our training set. Dataset is also preprocessed first using bnmnormalizer. We noticed a difference in performance for preprocessed data compared to unprocessed one.</p>\n<h2>Dataset Cleaning</h2>\n<p>We conducted data cleaning on the audio transcription by:</p>\n<ul>\n<li>Removing all kinds of punctuation. Manually added quotations into the list.</li>\n</ul>\n<pre><code>   \nbase = .punctuation \npunct = base + \npunct\n&gt;&gt;&gt; !</code></pre>\n<ul>\n<li>Strip any whitespaces at the front and at the back of string.</li>\n<li>Normalize the transcriptions using bnunicodenormalizer.</li>\n</ul>\n<pre><code>!pip install bnunicodenormalizer \n\n bnunicodenormalizer import Normalizer \nbnorm=Normalizer()\n\ndef norm(transcription):\n  text_list = []\n  texts = transcription.()\n\n     texts:\n     = bnorm()\n     ([]) &gt; :\n      text_list.append([][][])\n    :\n      text_list.append()\n\n  normalized_transcription = .join(text_list)\n</code></pre>\n<p>In order for the normalizer to work on our data, we had to split it into tokens (words). This is due to the normalizer not accepting a full sentence and can only process single words. We also noticed that the process could take some time. Hence, we implemented multi-threading to speed up the entire normalization process. We applied the data cleaning process on the base dataset from Bengali.AI and also Google Fleurs.</p>\n<ul>\n<li>Calculated the length for each of the audio files.</li>\n</ul>\n<p>Based on our analysis, we noticed the audio files had varying lengths. This resulted in unexpected OOM during the validation. Thus, we filtered the audio files and only took 18~20 seconds. </p>\n<ul>\n<li>Mixing Training and Validation Data</li>\n</ul>\n<p>The data given by Bengali.AI was already separated into 2 (training and validation). We noticed that the validation data was somewhat better as compared to the training data. Thus, we mixed some of the validation data into the training set. </p>\n<ul>\n<li>[Untested Process]</li>\n</ul>\n<p>In total we had almost 1 million audio files. From our analysis and preliminary training evaluation, we noticed that the data consists of very bad audio quality. This affected the performance of the model directly and prevented the model to improve. Initially we thought of calculating the amplitude of the audio files. But, determining the amplitude does not correlate to background noise. What we thought of doing is to filter the audio files that the model has trouble to transcribe. </p>\n<p>To put it simply, by training the model on the cleaned data, we would use the best checkpoint and transcribe the training and validation data. The data that produces the worst WER (word error rate) shall be excluded. This is to ensure that the model is only trained and validated on good quality data. </p>\n<h2>Language Model</h2>\n<p>We attached language model trained using kenlm library on top of our wav2vec model prediction to improve our score and use Word Error Rate (WER) and Character Error Rate (CER) as our evaluation metrics to benchmark the predictions.  We train the language model with the texts from competition datasets (train and validation) and google fleurs bengali.</p>\n<p>After rigorous evaluations, 3-gram showed the best performance among 4-gram and 5-gram models.</p>\n<h2>Details of Submission</h2>\n<h3>Model Saturation</h3>\n<p>In the early stages of our training phase, we dumped the entire dataset onto the model. Not only did we face OOM problems, but we slowly notice that the model became saturated. Due to not having an early stopping function and not stopping the training at certain points, the model was performing even worst as compared to those who only train with 10%-20% data. </p>\n<p>To prevent from the model being saturated, we had to feed the data into byte-sized pieces (pun intended). As it turns out, feeding a small partition of the data improved the model' performance quite drastically. From there, we tested the model using the test set. If the performance is increasing, we maintain the current data partition. If the performance is constant or decreasing, we would increase the data partition, i.e. add in more data. </p>\n<p>By using this strategy, we were able to prevent the model from being saturated. The only downside is that we had to keep track of the model's performance and constantly push out the checkpoints. </p>\n<p>All in all, implementing these strategies showed a significant improvement on the model's performance as compared to not cleaning the data at all. In our previous projects, we would only conduct extremely minor data cleaning process as to preserve the structure of the data. We wanted the model to learn the clean and also unclean data in the hopes of making it robust. Unfortunately, this was not the case. </p>\n<p>Not cleaning the data at all would degrade the performance and training the model for long periods of time will result in model saturation.</p>\n<h3>What didn't work for our team</h3>\n<p>We experimented with a few other pretrained models, but their result is not as good as our final model. We tried models from:</p>\n<ul>\n<li>facebook/mms-1b</li>\n<li>wav2vec2-large-xlsr-53</li>\n<li>whisper-small</li>\n<li>conformer-rnn-t</li>\n</ul>\n<p>We also noticed a degrading in performance when we train the model longer, our validation score no longer correlates with the public lb. </p>\n<p>We experimented with adding google fleur to our dataset for training, but it seems that the model trained on the dataset doesn't improve much of our submission score.</p>\n<p>We tried increasing the value in wav2vec2 parameters for mask_time_prob &amp; mask_feature_prob but we didn't notice any improvement in performance.</p>\n<p>We learned that it's very important that the dataset is properly cleaned and filtered in this competition. It could help boost the score if we manage to filter poor quality audio with high WER. We didn't manage to work more on preprocessing the dataset.</p>\n<h3>Sources</h3>\n<p>Pretrained model: <a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\" target=\"_blank\">https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali</a><br>\nSource Code: <a href=\"https://github.com/malaysia-ai/bengali.ai-stt-competition\" target=\"_blank\">https://github.com/malaysia-ai/bengali.ai-stt-competition</a><br>\nInference Notebook: <a href=\"https://www.kaggle.com/code/aisyahhrazak/inference-wav2vec2-kenlm\" target=\"_blank\">https://www.kaggle.com/code/aisyahhrazak/inference-wav2vec2-kenlm</a></p>",
  "messages": [
    {
      "id": "2503318",
      "postDate": "10/29/2023 00:31:02",
      "content": "<p>Thank you Bengali.AI &amp; Kaggle for organizing this competition. It was a very wonderful learning experience for my team <a href=\"https://www.kaggle.com/kamaruladha\" target=\"_blank\">@kamaruladha</a> <a href=\"https://www.kaggle.com/ariffnazhan\" target=\"_blank\">@ariffnazhan</a>. This is our first bronze medal and first-time dealing with audio data! </p>\n<p>Business context: <a href=\"http://www.kaggle.com/competitions/bengaliai-speech/overview\" target=\"_blank\">www.kaggle.com/competitions/bengaliai-speech/overview</a><br>\nData context: <a href=\"http://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">www.kaggle.com/competitions/bengaliai-speech/data</a></p>\n<h1>Overview of the Approach</h1>\n<h2>ASR Model</h2>\n<p>Our most performing model is from finetuning <code>ai4bharat/indicwav2vec_v1_bengali</code> trained for 2 epochs with training data containing both train and valid split. </p>\n<h2>Dataset</h2>\n<p>We took 80% of training data and 96% of validation split using random sample for our training set. Dataset is also preprocessed first using bnmnormalizer. We noticed a difference in performance for preprocessed data compared to unprocessed one.</p>\n<h2>Dataset Cleaning</h2>\n<p>We conducted data cleaning on the audio transcription by:</p>\n<ul>\n<li>Removing all kinds of punctuation. Manually added quotations into the list.</li>\n</ul>\n<pre><code>   \nbase = .punctuation \npunct = base + \npunct\n&gt;&gt;&gt; !</code></pre>\n<ul>\n<li>Strip any whitespaces at the front and at the back of string.</li>\n<li>Normalize the transcriptions using bnunicodenormalizer.</li>\n</ul>\n<pre><code>!pip install bnunicodenormalizer \n\n bnunicodenormalizer import Normalizer \nbnorm=Normalizer()\n\ndef norm(transcription):\n  text_list = []\n  texts = transcription.()\n\n     texts:\n     = bnorm()\n     ([]) &gt; :\n      text_list.append([][][])\n    :\n      text_list.append()\n\n  normalized_transcription = .join(text_list)\n</code></pre>\n<p>In order for the normalizer to work on our data, we had to split it into tokens (words). This is due to the normalizer not accepting a full sentence and can only process single words. We also noticed that the process could take some time. Hence, we implemented multi-threading to speed up the entire normalization process. We applied the data cleaning process on the base dataset from Bengali.AI and also Google Fleurs.</p>\n<ul>\n<li>Calculated the length for each of the audio files.</li>\n</ul>\n<p>Based on our analysis, we noticed the audio files had varying lengths. This resulted in unexpected OOM during the validation. Thus, we filtered the audio files and only took 18~20 seconds. </p>\n<ul>\n<li>Mixing Training and Validation Data</li>\n</ul>\n<p>The data given by Bengali.AI was already separated into 2 (training and validation). We noticed that the validation data was somewhat better as compared to the training data. Thus, we mixed some of the validation data into the training set. </p>\n<ul>\n<li>[Untested Process]</li>\n</ul>\n<p>In total we had almost 1 million audio files. From our analysis and preliminary training evaluation, we noticed that the data consists of very bad audio quality. This affected the performance of the model directly and prevented the model to improve. Initially we thought of calculating the amplitude of the audio files. But, determining the amplitude does not correlate to background noise. What we thought of doing is to filter the audio files that the model has trouble to transcribe. </p>\n<p>To put it simply, by training the model on the cleaned data, we would use the best checkpoint and transcribe the training and validation data. The data that produces the worst WER (word error rate) shall be excluded. This is to ensure that the model is only trained and validated on good quality data. </p>\n<h2>Language Model</h2>\n<p>We attached language model trained using kenlm library on top of our wav2vec model prediction to improve our score and use Word Error Rate (WER) and Character Error Rate (CER) as our evaluation metrics to benchmark the predictions.  We train the language model with the texts from competition datasets (train and validation) and google fleurs bengali.</p>\n<p>After rigorous evaluations, 3-gram showed the best performance among 4-gram and 5-gram models.</p>\n<h2>Details of Submission</h2>\n<h3>Model Saturation</h3>\n<p>In the early stages of our training phase, we dumped the entire dataset onto the model. Not only did we face OOM problems, but we slowly notice that the model became saturated. Due to not having an early stopping function and not stopping the training at certain points, the model was performing even worst as compared to those who only train with 10%-20% data. </p>\n<p>To prevent from the model being saturated, we had to feed the data into byte-sized pieces (pun intended). As it turns out, feeding a small partition of the data improved the model' performance quite drastically. From there, we tested the model using the test set. If the performance is increasing, we maintain the current data partition. If the performance is constant or decreasing, we would increase the data partition, i.e. add in more data. </p>\n<p>By using this strategy, we were able to prevent the model from being saturated. The only downside is that we had to keep track of the model's performance and constantly push out the checkpoints. </p>\n<p>All in all, implementing these strategies showed a significant improvement on the model's performance as compared to not cleaning the data at all. In our previous projects, we would only conduct extremely minor data cleaning process as to preserve the structure of the data. We wanted the model to learn the clean and also unclean data in the hopes of making it robust. Unfortunately, this was not the case. </p>\n<p>Not cleaning the data at all would degrade the performance and training the model for long periods of time will result in model saturation.</p>\n<h3>What didn't work for our team</h3>\n<p>We experimented with a few other pretrained models, but their result is not as good as our final model. We tried models from:</p>\n<ul>\n<li>facebook/mms-1b</li>\n<li>wav2vec2-large-xlsr-53</li>\n<li>whisper-small</li>\n<li>conformer-rnn-t</li>\n</ul>\n<p>We also noticed a degrading in performance when we train the model longer, our validation score no longer correlates with the public lb. </p>\n<p>We experimented with adding google fleur to our dataset for training, but it seems that the model trained on the dataset doesn't improve much of our submission score.</p>\n<p>We tried increasing the value in wav2vec2 parameters for mask_time_prob &amp; mask_feature_prob but we didn't notice any improvement in performance.</p>\n<p>We learned that it's very important that the dataset is properly cleaned and filtered in this competition. It could help boost the score if we manage to filter poor quality audio with high WER. We didn't manage to work more on preprocessing the dataset.</p>\n<h3>Sources</h3>\n<p>Pretrained model: <a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\" target=\"_blank\">https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali</a><br>\nSource Code: <a href=\"https://github.com/malaysia-ai/bengali.ai-stt-competition\" target=\"_blank\">https://github.com/malaysia-ai/bengali.ai-stt-competition</a><br>\nInference Notebook: <a href=\"https://www.kaggle.com/code/aisyahhrazak/inference-wav2vec2-kenlm\" target=\"_blank\">https://www.kaggle.com/code/aisyahhrazak/inference-wav2vec2-kenlm</a></p>",
      "rawMarkdown": "Thank you Bengali.AI & Kaggle for organizing this competition. It was a very wonderful learning experience for my team @kamaruladha @ariffnazhan. This is our first bronze medal and first-time dealing with audio data! \n\nBusiness context: www.kaggle.com/competitions/bengaliai-speech/overview\nData context: www.kaggle.com/competitions/bengaliai-speech/data\n\n#Overview of the Approach\n\n## ASR Model\nOur most performing model is from finetuning `ai4bharat/indicwav2vec_v1_bengali` trained for 2 epochs with training data containing both train and valid split. \n\n## Dataset\nWe took 80% of training data and 96% of validation split using random sample for our training set. Dataset is also preprocessed first using bnmnormalizer. We noticed a difference in performance for preprocessed data compared to unprocessed one.\n\n## Dataset Cleaning\nWe conducted data cleaning on the audio transcription by:\n\n- Removing all kinds of punctuation. Manually added quotations into the list.\n```\nimport string  \nbase = string.punctuation \npunct = base + '“”'\npunct\n>>> !\"#$%&\\'()*+,-./:;<=>?@[\\\\]^_`{|}~“”\n```\n\n- Strip any whitespaces at the front and at the back of string.\n- Normalize the transcriptions using bnunicodenormalizer.\n```\n!pip install bnunicodenormalizer \n\nfrom bnunicodenormalizer import Normalizer \nbnorm=Normalizer()\n\ndef norm(transcription):\n  text_list = []\n  texts = transcription.split()\n\n  for text in texts:\n    result = bnorm(text)\n    if len(result[\"ops\"]) > 0:\n      text_list.append(result[\"ops\"][0][\"after\"])\n    else:\n      text_list.append(text)\n\n  normalized_transcription = \", \".join(text_list)\n```\nIn order for the normalizer to work on our data, we had to split it into tokens (words). This is due to the normalizer not accepting a full sentence and can only process single words. We also noticed that the process could take some time. Hence, we implemented multi-threading to speed up the entire normalization process. We applied the data cleaning process on the base dataset from Bengali.AI and also Google Fleurs.\n\n- Calculated the length for each of the audio files.\n\nBased on our analysis, we noticed the audio files had varying lengths. This resulted in unexpected OOM during the validation. Thus, we filtered the audio files and only took 18~20 seconds. \n\n\n- Mixing Training and Validation Data\n\nThe data given by Bengali.AI was already separated into 2 (training and validation). We noticed that the validation data was somewhat better as compared to the training data. Thus, we mixed some of the validation data into the training set. \n\n\n- [Untested Process]\n\nIn total we had almost 1 million audio files. From our analysis and preliminary training evaluation, we noticed that the data consists of very bad audio quality. This affected the performance of the model directly and prevented the model to improve. Initially we thought of calculating the amplitude of the audio files. But, determining the amplitude does not correlate to background noise. What we thought of doing is to filter the audio files that the model has trouble to transcribe. \n\nTo put it simply, by training the model on the cleaned data, we would use the best checkpoint and transcribe the training and validation data. The data that produces the worst WER (word error rate) shall be excluded. This is to ensure that the model is only trained and validated on good quality data. \n\n\n## Language Model\nWe attached language model trained using kenlm library on top of our wav2vec model prediction to improve our score and use Word Error Rate (WER) and Character Error Rate (CER) as our evaluation metrics to benchmark the predictions.  We train the language model with the texts from competition datasets (train and validation) and google fleurs bengali.\n\nAfter rigorous evaluations, 3-gram showed the best performance among 4-gram and 5-gram models.\n\n## Details of Submission\n\n### Model Saturation\n\nIn the early stages of our training phase, we dumped the entire dataset onto the model. Not only did we face OOM problems, but we slowly notice that the model became saturated. Due to not having an early stopping function and not stopping the training at certain points, the model was performing even worst as compared to those who only train with 10%-20% data. \n\nTo prevent from the model being saturated, we had to feed the data into byte-sized pieces (pun intended). As it turns out, feeding a small partition of the data improved the model' performance quite drastically. From there, we tested the model using the test set. If the performance is increasing, we maintain the current data partition. If the performance is constant or decreasing, we would increase the data partition, i.e. add in more data. \n\nBy using this strategy, we were able to prevent the model from being saturated. The only downside is that we had to keep track of the model's performance and constantly push out the checkpoints. \n\n\nAll in all, implementing these strategies showed a significant improvement on the model's performance as compared to not cleaning the data at all. In our previous projects, we would only conduct extremely minor data cleaning process as to preserve the structure of the data. We wanted the model to learn the clean and also unclean data in the hopes of making it robust. Unfortunately, this was not the case. \n\nNot cleaning the data at all would degrade the performance and training the model for long periods of time will result in model saturation.\n\n### What didn't work for our team\nWe experimented with a few other pretrained models, but their result is not as good as our final model. We tried models from:\n\n- facebook/mms-1b\n- wav2vec2-large-xlsr-53\n- whisper-small\n- conformer-rnn-t\n\nWe also noticed a degrading in performance when we train the model longer, our validation score no longer correlates with the public lb. \n\nWe experimented with adding google fleur to our dataset for training, but it seems that the model trained on the dataset doesn't improve much of our submission score.\n\nWe tried increasing the value in wav2vec2 parameters for mask_time_prob & mask_feature_prob but we didn't notice any improvement in performance.\n\nWe learned that it's very important that the dataset is properly cleaned and filtered in this competition. It could help boost the score if we manage to filter poor quality audio with high WER. We didn't manage to work more on preprocessing the dataset.\n\n### Sources\n\nPretrained model: https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\nSource Code: https://github.com/malaysia-ai/bengali.ai-stt-competition\nInference Notebook: https://www.kaggle.com/code/aisyahhrazak/inference-wav2vec2-kenlm",
      "votes": null
    },
    {
      "id": "2503372",
      "postDate": "10/29/2023 03:34:03",
      "content": "<p>Very insightful!</p>",
      "rawMarkdown": "Very insightful!",
      "votes": null
    },
    {
      "id": "2503397",
      "postDate": "10/29/2023 04:03:33",
      "content": "<p>Greatwork!</p>",
      "rawMarkdown": "Greatwork!",
      "votes": null
    },
    {
      "id": "2503408",
      "postDate": "10/29/2023 04:18:36",
      "content": "<p>Well done!</p>",
      "rawMarkdown": "Well done!",
      "votes": null
    },
    {
      "id": "2508873",
      "postDate": "11/02/2023 03:58:25",
      "content": "<p>good read!</p>",
      "rawMarkdown": "good read!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2503372,
      "author_name": "kamaruladha",
      "author_url": "",
      "post_date": "10/29/2023 03:34:03",
      "content": "<p>Very insightful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2503397,
      "author_name": "qnqfbqfqo",
      "author_url": "",
      "post_date": "10/29/2023 04:03:33",
      "content": "<p>Greatwork!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2503408,
      "author_name": "asvinkumarmoghan",
      "author_url": "",
      "post_date": "10/29/2023 04:18:36",
      "content": "<p>Well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2508873,
      "author_name": "ariffnazhan",
      "author_url": "",
      "post_date": "11/02/2023 03:58:25",
      "content": "<p>good read!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2503318": "Thank you Bengali.AI & Kaggle for organizing this competition. It was a very wonderful learning experience for my team @kamaruladha @ariffnazhan. This is our first bronze medal and first-time dealing with audio data! \n\nBusiness context: www.kaggle.com/competitions/bengaliai-speech/overview\nData context: www.kaggle.com/competitions/bengaliai-speech/data\n\n#Overview of the Approach\n\n## ASR Model\nOur most performing model is from finetuning `ai4bharat/indicwav2vec_v1_bengali` trained for 2 epochs with training data containing both train and valid split. \n\n## Dataset\nWe took 80% of training data and 96% of validation split using random sample for our training set. Dataset is also preprocessed first using bnmnormalizer. We noticed a difference in performance for preprocessed data compared to unprocessed one.\n\n## Dataset Cleaning\nWe conducted data cleaning on the audio transcription by:\n\n- Removing all kinds of punctuation. Manually added quotations into the list.\n```\nimport string  \nbase = string.punctuation \npunct = base + '“”'\npunct\n>>> !\"#$%&\\'()*+,-./:;<=>?@[\\\\]^_`{|}~“”\n```\n\n- Strip any whitespaces at the front and at the back of string.\n- Normalize the transcriptions using bnunicodenormalizer.\n```\n!pip install bnunicodenormalizer \n\nfrom bnunicodenormalizer import Normalizer \nbnorm=Normalizer()\n\ndef norm(transcription):\n  text_list = []\n  texts = transcription.split()\n\n  for text in texts:\n    result = bnorm(text)\n    if len(result[\"ops\"]) > 0:\n      text_list.append(result[\"ops\"][0][\"after\"])\n    else:\n      text_list.append(text)\n\n  normalized_transcription = \", \".join(text_list)\n```\nIn order for the normalizer to work on our data, we had to split it into tokens (words). This is due to the normalizer not accepting a full sentence and can only process single words. We also noticed that the process could take some time. Hence, we implemented multi-threading to speed up the entire normalization process. We applied the data cleaning process on the base dataset from Bengali.AI and also Google Fleurs.\n\n- Calculated the length for each of the audio files.\n\nBased on our analysis, we noticed the audio files had varying lengths. This resulted in unexpected OOM during the validation. Thus, we filtered the audio files and only took 18~20 seconds. \n\n\n- Mixing Training and Validation Data\n\nThe data given by Bengali.AI was already separated into 2 (training and validation). We noticed that the validation data was somewhat better as compared to the training data. Thus, we mixed some of the validation data into the training set. \n\n\n- [Untested Process]\n\nIn total we had almost 1 million audio files. From our analysis and preliminary training evaluation, we noticed that the data consists of very bad audio quality. This affected the performance of the model directly and prevented the model to improve. Initially we thought of calculating the amplitude of the audio files. But, determining the amplitude does not correlate to background noise. What we thought of doing is to filter the audio files that the model has trouble to transcribe. \n\nTo put it simply, by training the model on the cleaned data, we would use the best checkpoint and transcribe the training and validation data. The data that produces the worst WER (word error rate) shall be excluded. This is to ensure that the model is only trained and validated on good quality data. \n\n\n## Language Model\nWe attached language model trained using kenlm library on top of our wav2vec model prediction to improve our score and use Word Error Rate (WER) and Character Error Rate (CER) as our evaluation metrics to benchmark the predictions.  We train the language model with the texts from competition datasets (train and validation) and google fleurs bengali.\n\nAfter rigorous evaluations, 3-gram showed the best performance among 4-gram and 5-gram models.\n\n## Details of Submission\n\n### Model Saturation\n\nIn the early stages of our training phase, we dumped the entire dataset onto the model. Not only did we face OOM problems, but we slowly notice that the model became saturated. Due to not having an early stopping function and not stopping the training at certain points, the model was performing even worst as compared to those who only train with 10%-20% data. \n\nTo prevent from the model being saturated, we had to feed the data into byte-sized pieces (pun intended). As it turns out, feeding a small partition of the data improved the model' performance quite drastically. From there, we tested the model using the test set. If the performance is increasing, we maintain the current data partition. If the performance is constant or decreasing, we would increase the data partition, i.e. add in more data. \n\nBy using this strategy, we were able to prevent the model from being saturated. The only downside is that we had to keep track of the model's performance and constantly push out the checkpoints. \n\n\nAll in all, implementing these strategies showed a significant improvement on the model's performance as compared to not cleaning the data at all. In our previous projects, we would only conduct extremely minor data cleaning process as to preserve the structure of the data. We wanted the model to learn the clean and also unclean data in the hopes of making it robust. Unfortunately, this was not the case. \n\nNot cleaning the data at all would degrade the performance and training the model for long periods of time will result in model saturation.\n\n### What didn't work for our team\nWe experimented with a few other pretrained models, but their result is not as good as our final model. We tried models from:\n\n- facebook/mms-1b\n- wav2vec2-large-xlsr-53\n- whisper-small\n- conformer-rnn-t\n\nWe also noticed a degrading in performance when we train the model longer, our validation score no longer correlates with the public lb. \n\nWe experimented with adding google fleur to our dataset for training, but it seems that the model trained on the dataset doesn't improve much of our submission score.\n\nWe tried increasing the value in wav2vec2 parameters for mask_time_prob & mask_feature_prob but we didn't notice any improvement in performance.\n\nWe learned that it's very important that the dataset is properly cleaned and filtered in this competition. It could help boost the score if we manage to filter poor quality audio with high WER. We didn't manage to work more on preprocessing the dataset.\n\n### Sources\n\nPretrained model: https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\nSource Code: https://github.com/malaysia-ai/bengali.ai-stt-competition\nInference Notebook: https://www.kaggle.com/code/aisyahhrazak/inference-wav2vec2-kenlm",
    "2503372": "Very insightful!",
    "2503397": "Greatwork!",
    "2503408": "Well done!",
    "2508873": "good read!"
  },
  "source": "meta"
}