{
  "id": 448020,
  "title": "41st place solution",
  "url": "/competitions/bengaliai-speech/discussion/448020",
  "author_name": "Arunodhayan",
  "post_date": "2023-10-18T06:15:00.914000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong><em><em>41st place solution</em></em></strong><br>\nMany thanks to Kaggle and Bengali.AI for hosting such an interesting competition.</p>\n<p>The solution comprises of 2 components:</p>\n<p>ASR model<br>\nLanguage model<br>\n<strong>1. ASR model</strong></p>\n<p>I used wav2vec2-large-960h-lv60-self/   as prpretrained model.</p>\n<p><strong>Datasets:</strong></p>\n<p>Speech data: Competition data, Common voice, Fleur data, Bn data<br>\n<strong>Augmentation:</strong></p>\n<p>speedAugProb: float, optional<br>\n            Probability that speed augmentation will be applied.<br>\n            If &lt;= 0, speed augmentation is disabled.</p>\n<pre><code>    volAugProb: , optional\n        Probability that volume augmentation will be applied.\n         &lt;= , volume augmentation  disabled.\n\n    reverbAugProb: , optional\n        Probability that reverberation augmentation will be applied.\n         &lt;= , reverberation augmentation  disabled.\n\n    noiseAugProb: , optional\n        Probability that noise augmentation will be applied.\n         &lt;= , noise augmentation  disabled.\n\n    speedFactors: List[], optional\n        List  factors  which  speed up (&gt;)  slow down (&lt;)\n        audio . One factor  chosen randomly  provided. Otherwise,\n         speed factors are [, , ].\n\n    volScaleMinMax: Tuple[, ], optional\n        [Min, Max] range  volume scale factors. One factor \n        chose randomly  uniform probability  this range.\n         range  [, ].\n</code></pre>\n<p><strong>Training:</strong></p>\n<p>First Trained on Common voice, fleur and bn , then followed by 10% and 20 % data<br>\nwarmup for every 4000 steps 5e-5 &amp; 3e-5 as LR</p>\n<p><strong>Inference:</strong><br>\n                    corrections = {<br>\n                              \"ক ষ\": \"ক্ষ\",<br>\n                              \"ঞ চ\": \"ঞ্চ\",<br>\n                              \"ঞ ছ\": \"ঞ্ছ\",<br>\n                              \"ঞ জ\": \"ঞ্জ\",<br>\n                              \"জ ঞ\": \"জ্ঞ\",<br>\n                              \"ত র\": \"ত্র\",<br>\n                              \"শ র\": \"শ্র\",<br>\n                              \"স র\": \"স্র\",<br>\n                              \"স প\": \"স্প\",<br>\n                              \"স ফ\": \"স্ফ\",<br>\n                              \"স থ\": \"স্থ\",<br>\n                              \"স ম\": \"স্ম\",<br>\n                              \"হ ম\": \"হ্ম\",<br>\n                              \"হ ল\": \"হ্ল\",<br>\n                              \"হ র\": \"হ্র\",<br>\n                              }<br>\napplied these as sample to the inference post processing</p>\n<p><strong>Language model</strong></p>\n<p>6-gram kenlm model trained on the following data:</p>\n<ul>\n<li>Competition training data</li>\n<li>External training data etc</li>\n</ul>",
  "messages": [
    {
      "id": 2486724,
      "postDate": "2023-10-18T06:15:00.913Z",
      "content": "<p><strong><em><em>41st place solution</em></em></strong><br>\nMany thanks to Kaggle and Bengali.AI for hosting such an interesting competition.</p>\n<p>The solution comprises of 2 components:</p>\n<p>ASR model<br>\nLanguage model<br>\n<strong>1. ASR model</strong></p>\n<p>I used wav2vec2-large-960h-lv60-self/   as prpretrained model.</p>\n<p><strong>Datasets:</strong></p>\n<p>Speech data: Competition data, Common voice, Fleur data, Bn data<br>\n<strong>Augmentation:</strong></p>\n<p>speedAugProb: float, optional<br>\n            Probability that speed augmentation will be applied.<br>\n            If &lt;= 0, speed augmentation is disabled.</p>\n<pre><code>    volAugProb: , optional\n        Probability that volume augmentation will be applied.\n         &lt;= , volume augmentation  disabled.\n\n    reverbAugProb: , optional\n        Probability that reverberation augmentation will be applied.\n         &lt;= , reverberation augmentation  disabled.\n\n    noiseAugProb: , optional\n        Probability that noise augmentation will be applied.\n         &lt;= , noise augmentation  disabled.\n\n    speedFactors: List[], optional\n        List  factors  which  speed up (&gt;)  slow down (&lt;)\n        audio . One factor  chosen randomly  provided. Otherwise,\n         speed factors are [, , ].\n\n    volScaleMinMax: Tuple[, ], optional\n        [Min, Max] range  volume scale factors. One factor \n        chose randomly  uniform probability  this range.\n         range  [, ].\n</code></pre>\n<p><strong>Training:</strong></p>\n<p>First Trained on Common voice, fleur and bn , then followed by 10% and 20 % data<br>\nwarmup for every 4000 steps 5e-5 &amp; 3e-5 as LR</p>\n<p><strong>Inference:</strong><br>\n                    corrections = {<br>\n                              \"ক ষ\": \"ক্ষ\",<br>\n                              \"ঞ চ\": \"ঞ্চ\",<br>\n                              \"ঞ ছ\": \"ঞ্ছ\",<br>\n                              \"ঞ জ\": \"ঞ্জ\",<br>\n                              \"জ ঞ\": \"জ্ঞ\",<br>\n                              \"ত র\": \"ত্র\",<br>\n                              \"শ র\": \"শ্র\",<br>\n                              \"স র\": \"স্র\",<br>\n                              \"স প\": \"স্প\",<br>\n                              \"স ফ\": \"স্ফ\",<br>\n                              \"স থ\": \"স্থ\",<br>\n                              \"স ম\": \"স্ম\",<br>\n                              \"হ ম\": \"হ্ম\",<br>\n                              \"হ ল\": \"হ্ল\",<br>\n                              \"হ র\": \"হ্র\",<br>\n                              }<br>\napplied these as sample to the inference post processing</p>\n<p><strong>Language model</strong></p>\n<p>6-gram kenlm model trained on the following data:</p>\n<ul>\n<li>Competition training data</li>\n<li>External training data etc</li>\n</ul>",
      "rawMarkdown": "****41st place solution****\nMany thanks to Kaggle and Bengali.AI for hosting such an interesting competition.\n\nThe solution comprises of 2 components:\n\nASR model\nLanguage model\n**1. ASR model**\n\nI used wav2vec2-large-960h-lv60-self/   as prpretrained model.\n\n**Datasets:**\n\nSpeech data: Competition data, Common voice, Fleur data, Bn data\n**Augmentation:**\n\nspeedAugProb: float, optional\n            Probability that speed augmentation will be applied.\n            If <= 0, speed augmentation is disabled.\n\n        volAugProb: float, optional\n            Probability that volume augmentation will be applied.\n            If <= 0, volume augmentation is disabled.\n\n        reverbAugProb: float, optional\n            Probability that reverberation augmentation will be applied.\n            If <= 0, reverberation augmentation is disabled.\n\n        noiseAugProb: float, optional\n            Probability that noise augmentation will be applied.\n            If <= 0, noise augmentation is disabled.\n\n        speedFactors: List[float], optional\n            List of factors by which to speed up (>1) or slow down (<1)\n            audio by. One factor is chosen randomly if provided. Otherwise,\n            default speed factors are [0.9, 1.0, 1.0].\n\n        volScaleMinMax: Tuple[float, float], optional\n            [Min, Max] range for volume scale factors. One factor is\n            chose randomly with uniform probability from this range.\n            Default range is [0.125, 2.0].\n\n**Training:**\n\nFirst Trained on Common voice, fleur and bn , then followed by 10% and 20 % data\nwarmup for every 4000 steps 5e-5 & 3e-5 as LR\n\n**Inference:**\n                    corrections = {\n                              \"ক ষ\": \"ক্ষ\",\n                              \"ঞ চ\": \"ঞ্চ\",\n                              \"ঞ ছ\": \"ঞ্ছ\",\n                              \"ঞ জ\": \"ঞ্জ\",\n                              \"জ ঞ\": \"জ্ঞ\",\n                              \"ত র\": \"ত্র\",\n                              \"শ র\": \"শ্র\",\n                              \"স র\": \"স্র\",\n                              \"স প\": \"স্প\",\n                              \"স ফ\": \"স্ফ\",\n                              \"স থ\": \"স্থ\",\n                              \"স ম\": \"স্ম\",\n                              \"হ ম\": \"হ্ম\",\n                              \"হ ল\": \"হ্ল\",\n                              \"হ র\": \"হ্র\",\n                              }\napplied these as sample to the inference post processing\n\n**Language model**\n\n6-gram kenlm model trained on the following data:\n\n- Competition training data\n- External training data etc",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2486724": "****41st place solution****\nMany thanks to Kaggle and Bengali.AI for hosting such an interesting competition.\n\nThe solution comprises of 2 components:\n\nASR model\nLanguage model\n**1. ASR model**\n\nI used wav2vec2-large-960h-lv60-self/   as prpretrained model.\n\n**Datasets:**\n\nSpeech data: Competition data, Common voice, Fleur data, Bn data\n**Augmentation:**\n\nspeedAugProb: float, optional\n            Probability that speed augmentation will be applied.\n            If <= 0, speed augmentation is disabled.\n\n        volAugProb: float, optional\n            Probability that volume augmentation will be applied.\n            If <= 0, volume augmentation is disabled.\n\n        reverbAugProb: float, optional\n            Probability that reverberation augmentation will be applied.\n            If <= 0, reverberation augmentation is disabled.\n\n        noiseAugProb: float, optional\n            Probability that noise augmentation will be applied.\n            If <= 0, noise augmentation is disabled.\n\n        speedFactors: List[float], optional\n            List of factors by which to speed up (>1) or slow down (<1)\n            audio by. One factor is chosen randomly if provided. Otherwise,\n            default speed factors are [0.9, 1.0, 1.0].\n\n        volScaleMinMax: Tuple[float, float], optional\n            [Min, Max] range for volume scale factors. One factor is\n            chose randomly with uniform probability from this range.\n            Default range is [0.125, 2.0].\n\n**Training:**\n\nFirst Trained on Common voice, fleur and bn , then followed by 10% and 20 % data\nwarmup for every 4000 steps 5e-5 & 3e-5 as LR\n\n**Inference:**\n                    corrections = {\n                              \"ক ষ\": \"ক্ষ\",\n                              \"ঞ চ\": \"ঞ্চ\",\n                              \"ঞ ছ\": \"ঞ্ছ\",\n                              \"ঞ জ\": \"ঞ্জ\",\n                              \"জ ঞ\": \"জ্ঞ\",\n                              \"ত র\": \"ত্র\",\n                              \"শ র\": \"শ্র\",\n                              \"স র\": \"স্র\",\n                              \"স প\": \"স্প\",\n                              \"স ফ\": \"স্ফ\",\n                              \"স থ\": \"স্থ\",\n                              \"স ম\": \"স্ম\",\n                              \"হ ম\": \"হ্ম\",\n                              \"হ ল\": \"হ্ল\",\n                              \"হ র\": \"হ্র\",\n                              }\napplied these as sample to the inference post processing\n\n**Language model**\n\n6-gram kenlm model trained on the following data:\n\n- Competition training data\n- External training data etc"
  }
}