{
  "id": 450572,
  "title": "66th Place Solution for the Bengali.AI Speech Recognition Competition",
  "url": "/competitions/bengaliai-speech/discussion/450572",
  "author_name": "Man of the year",
  "post_date": "2023-10-24T21:11:36.356000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Business context: <a href=\"http://www.kaggle.com/competitions/bengaliai-speech/overview\" target=\"_blank\">www.kaggle.com/competitions/bengaliai-speech/overview</a><br>\nData context: <a href=\"http://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">www.kaggle.com/competitions/bengaliai-speech/data</a></p>\n<h2>Overview of the Approach</h2>\n<p>Basically, we decided to use a very straightforward approach: finetuning only one model and using the language model afterward. We thought that the key to the success in this competition is the data, as the private dataset contains different OOD samples. In our final submission we have used the <code>ai4bharat/indicwav2vec_v1_bengali</code> model, that was finetuned for 3 epochs with the data. For the data we have used: SLR53 dataset, SLR37 dataset, Common Voice 13 data, 0.3 of the train split (random) competition data, 0.85 of the validation split competition data. We deleted Common Voice samples from the validation set, however, our CV score did not really correlated with the LB score. For the language model we just used the <code>arijitx/wav2vec2-xls-r-300m-bengali</code>.</p>\n<h2>Details of the submission</h2>\n<p>In conclusion, we have tried other pretrained checkpoints and other models, but they performed worse. We have also tried some augmentations for the volume and noise, but they did not work for us.</p>\n<p>It is important to clean the competition dataset, for example, by pseudolabeling it. Otherwise you get poor correlation between CV and LB as we did. It is also important to train model on several datasets, then model trains better. We suppose, that it is due to \"unlearning\". If we finetune the model only on a competition data, then the model starts forgetting previous knowledge. For the OOD data model should remember many variations of speech. </p>\n<p>We also did not do the punctuations part due to the lack of time, but it is essential, if you work with this type of problem.</p>\n<h2>Sources</h2>\n<p>Pretrained model: <a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\" target=\"_blank\">https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali</a><br>\nDataset 1: <a href=\"https://www.openslr.org/53/\" target=\"_blank\">https://www.openslr.org/53/</a><br>\nDataset 2: <a href=\"https://www.openslr.org/37/\" target=\"_blank\">https://www.openslr.org/37/</a><br>\nDataset 3: <a href=\"https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv\" target=\"_blank\">https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv</a></p>",
  "messages": [
    {
      "id": 2497747,
      "postDate": "2023-10-24T21:11:36.357Z",
      "content": "<p>Business context: <a href=\"http://www.kaggle.com/competitions/bengaliai-speech/overview\" target=\"_blank\">www.kaggle.com/competitions/bengaliai-speech/overview</a><br>\nData context: <a href=\"http://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">www.kaggle.com/competitions/bengaliai-speech/data</a></p>\n<h2>Overview of the Approach</h2>\n<p>Basically, we decided to use a very straightforward approach: finetuning only one model and using the language model afterward. We thought that the key to the success in this competition is the data, as the private dataset contains different OOD samples. In our final submission we have used the <code>ai4bharat/indicwav2vec_v1_bengali</code> model, that was finetuned for 3 epochs with the data. For the data we have used: SLR53 dataset, SLR37 dataset, Common Voice 13 data, 0.3 of the train split (random) competition data, 0.85 of the validation split competition data. We deleted Common Voice samples from the validation set, however, our CV score did not really correlated with the LB score. For the language model we just used the <code>arijitx/wav2vec2-xls-r-300m-bengali</code>.</p>\n<h2>Details of the submission</h2>\n<p>In conclusion, we have tried other pretrained checkpoints and other models, but they performed worse. We have also tried some augmentations for the volume and noise, but they did not work for us.</p>\n<p>It is important to clean the competition dataset, for example, by pseudolabeling it. Otherwise you get poor correlation between CV and LB as we did. It is also important to train model on several datasets, then model trains better. We suppose, that it is due to \"unlearning\". If we finetune the model only on a competition data, then the model starts forgetting previous knowledge. For the OOD data model should remember many variations of speech. </p>\n<p>We also did not do the punctuations part due to the lack of time, but it is essential, if you work with this type of problem.</p>\n<h2>Sources</h2>\n<p>Pretrained model: <a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\" target=\"_blank\">https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali</a><br>\nDataset 1: <a href=\"https://www.openslr.org/53/\" target=\"_blank\">https://www.openslr.org/53/</a><br>\nDataset 2: <a href=\"https://www.openslr.org/37/\" target=\"_blank\">https://www.openslr.org/37/</a><br>\nDataset 3: <a href=\"https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv\" target=\"_blank\">https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv</a></p>",
      "rawMarkdown": "Business context: www.kaggle.com/competitions/bengaliai-speech/overview\nData context: www.kaggle.com/competitions/bengaliai-speech/data\n\n## Overview of the Approach\n\nBasically, we decided to use a very straightforward approach: finetuning only one model and using the language model afterward. We thought that the key to the success in this competition is the data, as the private dataset contains different OOD samples. In our final submission we have used the `ai4bharat/indicwav2vec_v1_bengali` model, that was finetuned for 3 epochs with the data. For the data we have used: SLR53 dataset, SLR37 dataset, Common Voice 13 data, 0.3 of the train split (random) competition data, 0.85 of the validation split competition data. We deleted Common Voice samples from the validation set, however, our CV score did not really correlated with the LB score. For the language model we just used the `arijitx/wav2vec2-xls-r-300m-bengali`.\n\n## Details of the submission\n\nIn conclusion, we have tried other pretrained checkpoints and other models, but they performed worse. We have also tried some augmentations for the volume and noise, but they did not work for us.\n\nIt is important to clean the competition dataset, for example, by pseudolabeling it. Otherwise you get poor correlation between CV and LB as we did. It is also important to train model on several datasets, then model trains better. We suppose, that it is due to \"unlearning\". If we finetune the model only on a competition data, then the model starts forgetting previous knowledge. For the OOD data model should remember many variations of speech. \n\nWe also did not do the punctuations part due to the lack of time, but it is essential, if you work with this type of problem.\n\n## Sources\nPretrained model: https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\nDataset 1: https://www.openslr.org/53/\nDataset 2: https://www.openslr.org/37/\nDataset 3: https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2497747": "Business context: www.kaggle.com/competitions/bengaliai-speech/overview\nData context: www.kaggle.com/competitions/bengaliai-speech/data\n\n## Overview of the Approach\n\nBasically, we decided to use a very straightforward approach: finetuning only one model and using the language model afterward. We thought that the key to the success in this competition is the data, as the private dataset contains different OOD samples. In our final submission we have used the `ai4bharat/indicwav2vec_v1_bengali` model, that was finetuned for 3 epochs with the data. For the data we have used: SLR53 dataset, SLR37 dataset, Common Voice 13 data, 0.3 of the train split (random) competition data, 0.85 of the validation split competition data. We deleted Common Voice samples from the validation set, however, our CV score did not really correlated with the LB score. For the language model we just used the `arijitx/wav2vec2-xls-r-300m-bengali`.\n\n## Details of the submission\n\nIn conclusion, we have tried other pretrained checkpoints and other models, but they performed worse. We have also tried some augmentations for the volume and noise, but they did not work for us.\n\nIt is important to clean the competition dataset, for example, by pseudolabeling it. Otherwise you get poor correlation between CV and LB as we did. It is also important to train model on several datasets, then model trains better. We suppose, that it is due to \"unlearning\". If we finetune the model only on a competition data, then the model starts forgetting previous knowledge. For the OOD data model should remember many variations of speech. \n\nWe also did not do the punctuations part due to the lack of time, but it is essential, if you work with this type of problem.\n\n## Sources\nPretrained model: https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\nDataset 1: https://www.openslr.org/53/\nDataset 2: https://www.openslr.org/37/\nDataset 3: https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized?select=train.tsv"
  }
}