{
  "id": 450522,
  "title": "A low ranked solution!!!😅",
  "url": "/competitions/bengaliai-speech/discussion/450522",
  "author_name": "Tushar",
  "post_date": "2023-10-24T16:10:15.159000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Summary: I used the \"ai4bharat/indicwav2vec_v1_bengali\" acoustic model and worked with the competition dataset. Since I don't have any physical GPU, I relied on Kaggle's free GPU. Training the model with the entire competition dataset was impossible for me, so I decided to focus on 29k validation samples that has more accurate labels and randomly selected an equal number of training samples. I still couldn't train with these 58k samples, so I further reduced it to 30k training samples by random selection and 2k validation samples. Before training, I normalized the labels also.</p>\n<p>I experimented with various open-source acoustic models and found that \"ai4bharat/indicwav2vec_v1_bengali\" performed the best. I fine-tuned this model with my modified dataset in Kaggle. I saved checkpoints after 12 hours of training and then reload the checkpoint in another kernel and repeated this process three times (but you can't compare it with a continuous run, as the learning rate returns to its initial). Unfortunately, due to hardware limitations, I couldn't make further progress without a personal offline GPU workstation. That's why I gave up on the idea of using augmentations, incorporating a punctuation model, and so on.</p>\n<p>I noticed that many sentences ended with \"|\" , so I added \"|\" at the end of all labels. But then I realized, generally bengali imperative sentences start with words like {'কি', 'কী', 'কেন', 'কিভাবে','কবে','কখন'}, so I used the \"?\" sign instead of \"|\" , after all the sentences which started with the words that indicate imperatives in Bengali. Additionally, I incorporated a 5-gram language model into my approach. So, that's it.</p>",
  "messages": [
    {
      "id": 2497437,
      "postDate": "2023-10-24T16:10:15.160Z",
      "content": "<p>Summary: I used the \"ai4bharat/indicwav2vec_v1_bengali\" acoustic model and worked with the competition dataset. Since I don't have any physical GPU, I relied on Kaggle's free GPU. Training the model with the entire competition dataset was impossible for me, so I decided to focus on 29k validation samples that has more accurate labels and randomly selected an equal number of training samples. I still couldn't train with these 58k samples, so I further reduced it to 30k training samples by random selection and 2k validation samples. Before training, I normalized the labels also.</p>\n<p>I experimented with various open-source acoustic models and found that \"ai4bharat/indicwav2vec_v1_bengali\" performed the best. I fine-tuned this model with my modified dataset in Kaggle. I saved checkpoints after 12 hours of training and then reload the checkpoint in another kernel and repeated this process three times (but you can't compare it with a continuous run, as the learning rate returns to its initial). Unfortunately, due to hardware limitations, I couldn't make further progress without a personal offline GPU workstation. That's why I gave up on the idea of using augmentations, incorporating a punctuation model, and so on.</p>\n<p>I noticed that many sentences ended with \"|\" , so I added \"|\" at the end of all labels. But then I realized, generally bengali imperative sentences start with words like {'কি', 'কী', 'কেন', 'কিভাবে','কবে','কখন'}, so I used the \"?\" sign instead of \"|\" , after all the sentences which started with the words that indicate imperatives in Bengali. Additionally, I incorporated a 5-gram language model into my approach. So, that's it.</p>",
      "rawMarkdown": "Summary: I used the \"ai4bharat/indicwav2vec_v1_bengali\" acoustic model and worked with the competition dataset. Since I don't have any physical GPU, I relied on Kaggle's free GPU. Training the model with the entire competition dataset was impossible for me, so I decided to focus on 29k validation samples that has more accurate labels and randomly selected an equal number of training samples. I still couldn't train with these 58k samples, so I further reduced it to 30k training samples by random selection and 2k validation samples. Before training, I normalized the labels also.\n\nI experimented with various open-source acoustic models and found that \"ai4bharat/indicwav2vec_v1_bengali\" performed the best. I fine-tuned this model with my modified dataset in Kaggle. I saved checkpoints after 12 hours of training and then reload the checkpoint in another kernel and repeated this process three times (but you can't compare it with a continuous run, as the learning rate returns to its initial). Unfortunately, due to hardware limitations, I couldn't make further progress without a personal offline GPU workstation. That's why I gave up on the idea of using augmentations, incorporating a punctuation model, and so on.\n\nI noticed that many sentences ended with \"|\" , so I added \"|\" at the end of all labels. But then I realized, generally bengali imperative sentences start with words like {'কি', 'কী', 'কেন', 'কিভাবে','কবে','কখন'}, so I used the \"?\" sign instead of \"|\" , after all the sentences which started with the words that indicate imperatives in Bengali. Additionally, I incorporated a 5-gram language model into my approach. So, that's it.",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2497437": "Summary: I used the \"ai4bharat/indicwav2vec_v1_bengali\" acoustic model and worked with the competition dataset. Since I don't have any physical GPU, I relied on Kaggle's free GPU. Training the model with the entire competition dataset was impossible for me, so I decided to focus on 29k validation samples that has more accurate labels and randomly selected an equal number of training samples. I still couldn't train with these 58k samples, so I further reduced it to 30k training samples by random selection and 2k validation samples. Before training, I normalized the labels also.\n\nI experimented with various open-source acoustic models and found that \"ai4bharat/indicwav2vec_v1_bengali\" performed the best. I fine-tuned this model with my modified dataset in Kaggle. I saved checkpoints after 12 hours of training and then reload the checkpoint in another kernel and repeated this process three times (but you can't compare it with a continuous run, as the learning rate returns to its initial). Unfortunately, due to hardware limitations, I couldn't make further progress without a personal offline GPU workstation. That's why I gave up on the idea of using augmentations, incorporating a punctuation model, and so on.\n\nI noticed that many sentences ended with \"|\" , so I added \"|\" at the end of all labels. But then I realized, generally bengali imperative sentences start with words like {'কি', 'কী', 'কেন', 'কিভাবে','কবে','কখন'}, so I used the \"?\" sign instead of \"|\" , after all the sentences which started with the words that indicate imperatives in Bengali. Additionally, I incorporated a 5-gram language model into my approach. So, that's it."
  }
}