{
  "id": 448066,
  "title": "20th Rank Solution",
  "url": "/competitions/bengaliai-speech/discussion/448066",
  "author_name": "Balaji Selvaraj",
  "post_date": "2023-10-18T09:43:46.524000",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Our Final Solution is a single Wave2Vec model along with a N-Gram Model</p>\n<h1>Step 1:</h1>\n<p>Trained a model for 5 epochs on </p>\n<ol>\n<li>OpenSLR37</li>\n<li>OpenSLR53</li>\n<li>Fleurs</li>\n<li>Common voice</li>\n</ol>\n<p>Validation:<br>\nKaggle Validation data</p>\n<h1>Step 2:</h1>\n<p>Estimate the WER for the Kaggle training data with the model from Step 1.<br>\nRemove samples which have WER greater than 0.5<br>\nResulting cleaned version of Training data had 50-60% of overall Training data</p>\n<h1>Step 3:</h1>\n<p>Trained a model for 20 epochs on </p>\n<ol>\n<li>OpenSLR37</li>\n<li>OpenSLR53</li>\n<li>Fleurs</li>\n<li>Common voice</li>\n<li>Cleaned version of Training data</li>\n</ol>\n<h1>Step 4:</h1>\n<p>Build a N-Gram model with Common Voice dataset</p>\n<p>Our best Wave2Vec model                                       0.407 in LB<br>\nOur best Wave2Vec model + N-Gram Model          0.396 in Public LB</p>\n<p><strong>Things that worked:</strong></p>\n<ol>\n<li>Caching of the datasets reduced the overall training time. But it didnt allow us to add any augmentations on the fly</li>\n<li>Having more warmup steps ensured that the model doesn't throw NaN. We used warmup of 0.25</li>\n<li>MSD as Final layer instead of Linear layer</li>\n</ol>\n<p><strong>Things that didnt work:</strong></p>\n<ol>\n<li>We tried adding few more finetuning steps. But mostly the model started to overfit</li>\n</ol>\n<p>We will add few more points about our pipeline in next few hours</p>",
  "messages": [
    {
      "id": 2486996,
      "postDate": "2023-10-18T09:43:46.523Z",
      "content": "<p>Our Final Solution is a single Wave2Vec model along with a N-Gram Model</p>\n<h1>Step 1:</h1>\n<p>Trained a model for 5 epochs on </p>\n<ol>\n<li>OpenSLR37</li>\n<li>OpenSLR53</li>\n<li>Fleurs</li>\n<li>Common voice</li>\n</ol>\n<p>Validation:<br>\nKaggle Validation data</p>\n<h1>Step 2:</h1>\n<p>Estimate the WER for the Kaggle training data with the model from Step 1.<br>\nRemove samples which have WER greater than 0.5<br>\nResulting cleaned version of Training data had 50-60% of overall Training data</p>\n<h1>Step 3:</h1>\n<p>Trained a model for 20 epochs on </p>\n<ol>\n<li>OpenSLR37</li>\n<li>OpenSLR53</li>\n<li>Fleurs</li>\n<li>Common voice</li>\n<li>Cleaned version of Training data</li>\n</ol>\n<h1>Step 4:</h1>\n<p>Build a N-Gram model with Common Voice dataset</p>\n<p>Our best Wave2Vec model                                       0.407 in LB<br>\nOur best Wave2Vec model + N-Gram Model          0.396 in Public LB</p>\n<p><strong>Things that worked:</strong></p>\n<ol>\n<li>Caching of the datasets reduced the overall training time. But it didnt allow us to add any augmentations on the fly</li>\n<li>Having more warmup steps ensured that the model doesn't throw NaN. We used warmup of 0.25</li>\n<li>MSD as Final layer instead of Linear layer</li>\n</ol>\n<p><strong>Things that didnt work:</strong></p>\n<ol>\n<li>We tried adding few more finetuning steps. But mostly the model started to overfit</li>\n</ol>\n<p>We will add few more points about our pipeline in next few hours</p>",
      "rawMarkdown": "Our Final Solution is a single Wave2Vec model along with a N-Gram Model\n\n# Step 1:\nTrained a model for 5 epochs on \n1. OpenSLR37\n2. OpenSLR53\n3. Fleurs\n4. Common voice\n\nValidation:\nKaggle Validation data\n\n# Step 2:\nEstimate the WER for the Kaggle training data with the model from Step 1.\nRemove samples which have WER greater than 0.5\nResulting cleaned version of Training data had 50-60% of overall Training data\n\n# Step 3:\nTrained a model for 20 epochs on \n1. OpenSLR37\n2. OpenSLR53\n3. Fleurs\n4. Common voice\n5.  Cleaned version of Training data\n\n# Step 4:\nBuild a N-Gram model with Common Voice dataset\n\nOur best Wave2Vec model                                       0.407 in LB\nOur best Wave2Vec model + N-Gram Model          0.396 in Public LB\n\n**Things that worked:**\n1. Caching of the datasets reduced the overall training time. But it didnt allow us to add any augmentations on the fly\n2. Having more warmup steps ensured that the model doesn't throw NaN. We used warmup of 0.25\n3. MSD as Final layer instead of Linear layer\n\n**Things that didnt work:**\n1. We tried adding few more finetuning steps. But mostly the model started to overfit\n\nWe will add few more points about our pipeline in next few hours",
      "votes": 6
    },
    {
      "id": 2487856,
      "postDate": "2023-10-18T19:54:31.813Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2487856,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:54:31.813000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486996": "Our Final Solution is a single Wave2Vec model along with a N-Gram Model\n\n# Step 1:\nTrained a model for 5 epochs on \n1. OpenSLR37\n2. OpenSLR53\n3. Fleurs\n4. Common voice\n\nValidation:\nKaggle Validation data\n\n# Step 2:\nEstimate the WER for the Kaggle training data with the model from Step 1.\nRemove samples which have WER greater than 0.5\nResulting cleaned version of Training data had 50-60% of overall Training data\n\n# Step 3:\nTrained a model for 20 epochs on \n1. OpenSLR37\n2. OpenSLR53\n3. Fleurs\n4. Common voice\n5.  Cleaned version of Training data\n\n# Step 4:\nBuild a N-Gram model with Common Voice dataset\n\nOur best Wave2Vec model                                       0.407 in LB\nOur best Wave2Vec model + N-Gram Model          0.396 in Public LB\n\n**Things that worked:**\n1. Caching of the datasets reduced the overall training time. But it didnt allow us to add any augmentations on the fly\n2. Having more warmup steps ensured that the model doesn't throw NaN. We used warmup of 0.25\n3. MSD as Final layer instead of Linear layer\n\n**Things that didnt work:**\n1. We tried adding few more finetuning steps. But mostly the model started to overfit\n\nWe will add few more points about our pipeline in next few hours",
    "2487856": ""
  }
}