{
  "id": 448074,
  "title": "8th place solution",
  "url": "/competitions/bengaliai-speech/writeups/ktr-8th-place-solution",
  "author_name": "",
  "post_date": "2023-10-20T08:59:43.273Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thank you for organizing this competition.<br>\nI am new to speech recognition and have learned a lot. Thanks to the competition participants and hosts.</p>\n<h1>Solution Overview</h1>\n<ul>\n<li>Dataset cleaning</li>\n<li>Punctuation</li>\n<li>Simple Ensemble (Rescoring)</li>\n</ul>\n<h1>Dataset</h1>\n<p>First, I applied a data filter. Since I'm not familiar with the audio domain, I used a simple filter.<br>\nOne was the ratio between the transcribed text and the length of the audio data. I filtered the data based on a 20 characters per audio second. If it's higher than that, I considered the possibility that there are too many characters for the length of the audio data. It could be abnormally fast speech or corrupted audio data.<br>\nAnother filter I used was based on the prediction results from the yellow_king model. I filtered out data with scores above a certain threshold (CER &gt; 0.6).<br>\nBy training on clean data, I was able to start off with a good score.</p>\n<p>Because the competition data set was so large, training was repeated on a subset of it rather than using all of it. Once the model achieved some performance, the WER for the remaining data set was calculated and added to the data. This round was repeated several times.</p>\n<h2>external data</h2>\n<p>kathbath, openslr37, 53, ulca<br>\nThese external data were also filtered</p>\n<h1>Model</h1>\n<p><code>IndicWav2Vec</code><br>\nThis model was very powerful for the test set.<br>\nSince the previous competition report suggested that retraining a model once it had converged would not perform well, I used a pre-trained model instead of bengali's fine-tuned checkpoints (converted to HF model from faireseq checkpoints). Another reason was that I customized vocab (Mentioned in the section on punctuation.).</p>\n<h2>LM Model</h2>\n<p>5-gram Language Model (using all of Indiccorpus v1 and v2, BangaLM, and mC4). The text corpus totaled 45GB, with the 5gram.bin file being around 22GB in size.<br>\nUsing pyctcdecoder, alpha, beta and beam_width were adjusted with example audio CV. This initially provided a 0.01 boost to the LB score, but the benefit decreased as the score improved, eventually settling at around 0.003.</p>\n<h1>Punctuation</h1>\n<p>punctuation is very important.<br>\nInstead of creating a punctuation model, I took the simple approach of including punctuation in the vocab and LM corpus. Even when including punctuation, wav2vec was largely unable to predict it. However, it was found that including them in the vocab allowed them to be considered as candidates during LM decoding. Simply by including punctuation in both the model and the LM corpus, it was possible to improve the score from 0.44 to 0.428.</p>\n<h1>Ensemble (Rescoring)</h1>\n<p>Due to the structure of the output, it was not possible to simple ensemble at the logit level. Therefore, a strategy was adopted to calculate the LM Score for decoded sentences using the LM model, and the sentence with the highest score was selected.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197187%2F5e6ee0ffb1f1b28b93f1a744fe77c438%2Fbengali.png?generation=1697624102630269&amp;alt=media\" alt=\"\"></p>\n<p>Sentence used the output of three models with variations on training data and decoding parameters. Both public and private models improve by about 0.006 - 0.008.</p>\n<h1>Not working for me</h1>\n<ul>\n<li>xls-r-300m<ul>\n<li>Common Voice's CV is competitive with indicwav2vec, but performs very poorly for the test set. Not effective when added to an ensemble.</li></ul></li>\n<li>Speech Enhancement<ul>\n<li>It was found that the score was poor for example audio with reverb, so noise removal was performed using DeepFilterNet. There were cases where this boosted the LB score and cases where it didn't, so it was finally not used.</li></ul></li>\n<li>Adding Assamese<ul>\n<li>my understanding is that it shares characters with Bengali. Addition did not improve score.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "2487050",
      "postDate": "10/18/2023 10:33:06",
      "content": "<p>Thank you for organizing this competition.<br>\nI am new to speech recognition and have learned a lot. Thanks to the competition participants and hosts.</p>\n<h1>Solution Overview</h1>\n<ul>\n<li>Dataset cleaning</li>\n<li>Punctuation</li>\n<li>Simple Ensemble (Rescoring)</li>\n</ul>\n<h1>Dataset</h1>\n<p>First, I applied a data filter. Since I'm not familiar with the audio domain, I used a simple filter.<br>\nOne was the ratio between the transcribed text and the length of the audio data. I filtered the data based on a 20 characters per audio second. If it's higher than that, I considered the possibility that there are too many characters for the length of the audio data. It could be abnormally fast speech or corrupted audio data.<br>\nAnother filter I used was based on the prediction results from the yellow_king model. I filtered out data with scores above a certain threshold (CER &gt; 0.6).<br>\nBy training on clean data, I was able to start off with a good score.</p>\n<p>Because the competition data set was so large, training was repeated on a subset of it rather than using all of it. Once the model achieved some performance, the WER for the remaining data set was calculated and added to the data. This round was repeated several times.</p>\n<h2>external data</h2>\n<p>kathbath, openslr37, 53, ulca<br>\nThese external data were also filtered</p>\n<h1>Model</h1>\n<p><code>IndicWav2Vec</code><br>\nThis model was very powerful for the test set.<br>\nSince the previous competition report suggested that retraining a model once it had converged would not perform well, I used a pre-trained model instead of bengali's fine-tuned checkpoints (converted to HF model from faireseq checkpoints). Another reason was that I customized vocab (Mentioned in the section on punctuation.).</p>\n<h2>LM Model</h2>\n<p>5-gram Language Model (using all of Indiccorpus v1 and v2, BangaLM, and mC4). The text corpus totaled 45GB, with the 5gram.bin file being around 22GB in size.<br>\nUsing pyctcdecoder, alpha, beta and beam_width were adjusted with example audio CV. This initially provided a 0.01 boost to the LB score, but the benefit decreased as the score improved, eventually settling at around 0.003.</p>\n<h1>Punctuation</h1>\n<p>punctuation is very important.<br>\nInstead of creating a punctuation model, I took the simple approach of including punctuation in the vocab and LM corpus. Even when including punctuation, wav2vec was largely unable to predict it. However, it was found that including them in the vocab allowed them to be considered as candidates during LM decoding. Simply by including punctuation in both the model and the LM corpus, it was possible to improve the score from 0.44 to 0.428.</p>\n<h1>Ensemble (Rescoring)</h1>\n<p>Due to the structure of the output, it was not possible to simple ensemble at the logit level. Therefore, a strategy was adopted to calculate the LM Score for decoded sentences using the LM model, and the sentence with the highest score was selected.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197187%2F5e6ee0ffb1f1b28b93f1a744fe77c438%2Fbengali.png?generation=1697624102630269&amp;alt=media\" alt=\"\"></p>\n<p>Sentence used the output of three models with variations on training data and decoding parameters. Both public and private models improve by about 0.006 - 0.008.</p>\n<h1>Not working for me</h1>\n<ul>\n<li>xls-r-300m<ul>\n<li>Common Voice's CV is competitive with indicwav2vec, but performs very poorly for the test set. Not effective when added to an ensemble.</li></ul></li>\n<li>Speech Enhancement<ul>\n<li>It was found that the score was poor for example audio with reverb, so noise removal was performed using DeepFilterNet. There were cases where this boosted the LB score and cases where it didn't, so it was finally not used.</li></ul></li>\n<li>Adding Assamese<ul>\n<li>my understanding is that it shares characters with Bengali. Addition did not improve score.</li></ul></li>\n</ul>",
      "rawMarkdown": "Thank you for organizing this competition.\nI am new to speech recognition and have learned a lot. Thanks to the competition participants and hosts.\n\n# Solution Overview\n- Dataset cleaning\n- Punctuation\n- Simple Ensemble (Rescoring)\n\n# Dataset\nFirst, I applied a data filter. Since I'm not familiar with the audio domain, I used a simple filter.\nOne was the ratio between the transcribed text and the length of the audio data. I filtered the data based on a 20 characters per audio second. If it's higher than that, I considered the possibility that there are too many characters for the length of the audio data. It could be abnormally fast speech or corrupted audio data.\nAnother filter I used was based on the prediction results from the yellow_king model. I filtered out data with scores above a certain threshold (CER > 0.6).\nBy training on clean data, I was able to start off with a good score.\n\nBecause the competition data set was so large, training was repeated on a subset of it rather than using all of it. Once the model achieved some performance, the WER for the remaining data set was calculated and added to the data. This round was repeated several times.\n\n## external data\nkathbath, openslr37, 53, ulca\nThese external data were also filtered\n\n# Model\n`IndicWav2Vec`\nThis model was very powerful for the test set.\nSince the previous competition report suggested that retraining a model once it had converged would not perform well, I used a pre-trained model instead of bengali's fine-tuned checkpoints (converted to HF model from faireseq checkpoints). Another reason was that I customized vocab (Mentioned in the section on punctuation.).\n\n## LM Model\n5-gram Language Model (using all of Indiccorpus v1 and v2, BangaLM, and mC4). The text corpus totaled 45GB, with the 5gram.bin file being around 22GB in size.\nUsing pyctcdecoder, alpha, beta and beam_width were adjusted with example audio CV. This initially provided a 0.01 boost to the LB score, but the benefit decreased as the score improved, eventually settling at around 0.003.\n\n# Punctuation\npunctuation is very important.\nInstead of creating a punctuation model, I took the simple approach of including punctuation in the vocab and LM corpus. Even when including punctuation, wav2vec was largely unable to predict it. However, it was found that including them in the vocab allowed them to be considered as candidates during LM decoding. Simply by including punctuation in both the model and the LM corpus, it was possible to improve the score from 0.44 to 0.428.\n\n# Ensemble (Rescoring)\nDue to the structure of the output, it was not possible to simple ensemble at the logit level. Therefore, a strategy was adopted to calculate the LM Score for decoded sentences using the LM model, and the sentence with the highest score was selected.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197187%2F5e6ee0ffb1f1b28b93f1a744fe77c438%2Fbengali.png?generation=1697624102630269&alt=media)\n\nSentence used the output of three models with variations on training data and decoding parameters. Both public and private models improve by about 0.006 - 0.008.\n\n# Not working for me\n- xls-r-300m\n    - Common Voice's CV is competitive with indicwav2vec, but performs very poorly for the test set. Not effective when added to an ensemble.\n- Speech Enhancement\n - It was found that the score was poor for example audio with reverb, so noise removal was performed using DeepFilterNet. There were cases where this boosted the LB score and cases where it didn't, so it was finally not used.\n- Adding Assamese\n - my understanding is that it shares characters with Bengali. Addition did not improve score.",
      "votes": null
    },
    {
      "id": "2487921",
      "postDate": "10/18/2023 21:27:50",
      "content": "<p>Congratulations on your gold medal and thanks for sharing.</p>\n<p>A few questions:</p>\n<ol>\n<li>Could you please elaborate on how you predicted the LM score? Did you use a separate model, a second output of the encoder, …? What training loss did you use?</li>\n<li>What characters did you consider for punctuation? All (|?!.,\"'- etc.) or only some of them?</li>\n<li>Did you keep consecutive punctuation characters or reduced it to just one character?</li>\n<li>How did you add punctuation to the LM corpus? did you add it to all options or considered only what appeared in the available corpus?</li>\n<li>What hardware did you use for training?</li>\n</ol>",
      "rawMarkdown": "Congratulations on your gold medal and thanks for sharing.\n\nA few questions:\n1. Could you please elaborate on how you predicted the LM score? Did you use a separate model, a second output of the encoder, ...? What training loss did you use?\n2. What characters did you consider for punctuation? All (|?!.,\"'- etc.) or only some of them?\n3. Did you keep consecutive punctuation characters or reduced it to just one character?\n4. How did you add punctuation to the LM corpus? did you add it to all options or considered only what appeared in the available corpus?\n5. What hardware did you use for training?",
      "votes": null
    },
    {
      "id": "2488060",
      "postDate": "10/19/2023 01:57:35",
      "content": "<p>Thank you. Please feel free to ask additional questions if anything is unclear.</p>\n<p>1.No specific LMs have been created.<br>\nkenLM scores used in Wav2Vec2ProcessorWithLM were used.</p>\n<pre><code>  kenlm.LanguageModel()\n  model.score(pred_1)\n</code></pre>\n<p>My understanding is that the more natural the sentence, the better the score.</p>\n<p>2.The top four symbols in the competition data were included. <code>[',', '-', '?', '!']</code><br>\n3-4.The LM corpus only retains symbols and vocabulary contained in the tokenizer vocab. At this time, I did not single out consecutive symbols (but I think I should have done so). I replaced it with one in the predict sentence. <br>\n5.single A6000 or RTX 4090.</p>",
      "rawMarkdown": "Thank you. Please feel free to ask additional questions if anything is unclear.\n\n1.No specific LMs have been created.\nkenLM scores used in Wav2Vec2ProcessorWithLM were used.\n```\nmodel = kenlm.LanguageModel(\"5gram.bin\")\nscore = model.score(pred_1)\n```\nMy understanding is that the more natural the sentence, the better the score.\n\n2.The top four symbols in the competition data were included. `[',', '-', '?', '!']`\n3-4.The LM corpus only retains symbols and vocabulary contained in the tokenizer vocab. At this time, I did not single out consecutive symbols (but I think I should have done so). I replaced it with one in the predict sentence. \n5.single A6000 or RTX 4090.",
      "votes": null
    },
    {
      "id": "2490711",
      "postDate": "10/21/2023 01:16:31",
      "content": "<p>Thanks for the responses.</p>\n<p>If I understood well you response to 4., you included in the LM corpus only the words with the punctuations present in the input data. Let me illustrate with an English word (easier for me in light of my utter Bengali illiteracy):</p>\n<ul>\n<li>If the input data contains \"house\", \"house,\", \"house?\", you include the 3 options in the LM corpus;</li>\n<li>\"house-\" and \"house!\" would not be included.</li>\n</ul>\n<p>Did I get it right?</p>",
      "rawMarkdown": "Thanks for the responses.\n\nIf I understood well you response to 4., you included in the LM corpus only the words with the punctuations present in the input data. Let me illustrate with an English word (easier for me in light of my utter Bengali illiteracy):\n- If the input data contains \"house\", \"house,\", \"house?\", you include the 3 options in the LM corpus;\n- \"house-\" and \"house!\" would not be included.\n\nDid I get it right?",
      "votes": null
    },
    {
      "id": "2490978",
      "postDate": "10/21/2023 08:31:11",
      "content": "<p>Thanks for the additional questions.</p>\n<p>Assume the following case for the CTC tokenizer vocabulary.<br>\nHave: \",\" and \"?\"<br>\nDoes not have: \"-\" and \"!\"</p>\n<p>In this case, the possible words in the LM corpus are \"house\", \"house,\" and \"house?\"</p>\n<p>I wasn't sure if wav2vec could handle the symbols well, but I thought they should be recovered from the information in the acoustic model. (e.g., inflection and breath between words). However, in the top solution, the punctuation model with only textual information seems to work well enough. I was a bit surprised.</p>",
      "rawMarkdown": "Thanks for the additional questions.\n\nAssume the following case for the CTC tokenizer vocabulary.\nHave: \",\" and \"?\"\nDoes not have: \"-\" and \"!\"\n\nIn this case, the possible words in the LM corpus are \"house\", \"house,\" and \"house?\"\n\nI wasn't sure if wav2vec could handle the symbols well, but I thought they should be recovered from the information in the acoustic model. (e.g., inflection and breath between words). However, in the top solution, the punctuation model with only textual information seems to work well enough. I was a bit surprised.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2487921,
      "author_name": "vialactea",
      "author_url": "",
      "post_date": "10/18/2023 21:27:50",
      "content": "<p>Congratulations on your gold medal and thanks for sharing.</p>\n<p>A few questions:</p>\n<ol>\n<li>Could you please elaborate on how you predicted the LM score? Did you use a separate model, a second output of the encoder, …? What training loss did you use?</li>\n<li>What characters did you consider for punctuation? All (|?!.,\"'- etc.) or only some of them?</li>\n<li>Did you keep consecutive punctuation characters or reduced it to just one character?</li>\n<li>How did you add punctuation to the LM corpus? did you add it to all options or considered only what appeared in the available corpus?</li>\n<li>What hardware did you use for training?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2488060,
          "author_name": "kcotton21",
          "author_url": "",
          "post_date": "10/19/2023 01:57:35",
          "content": "<p>Thank you. Please feel free to ask additional questions if anything is unclear.</p>\n<p>1.No specific LMs have been created.<br>\nkenLM scores used in Wav2Vec2ProcessorWithLM were used.</p>\n<pre><code>  kenlm.LanguageModel()\n  model.score(pred_1)\n</code></pre>\n<p>My understanding is that the more natural the sentence, the better the score.</p>\n<p>2.The top four symbols in the competition data were included. <code>[',', '-', '?', '!']</code><br>\n3-4.The LM corpus only retains symbols and vocabulary contained in the tokenizer vocab. At this time, I did not single out consecutive symbols (but I think I should have done so). I replaced it with one in the predict sentence. <br>\n5.single A6000 or RTX 4090.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2490711,
              "author_name": "vialactea",
              "author_url": "",
              "post_date": "10/21/2023 01:16:31",
              "content": "<p>Thanks for the responses.</p>\n<p>If I understood well you response to 4., you included in the LM corpus only the words with the punctuations present in the input data. Let me illustrate with an English word (easier for me in light of my utter Bengali illiteracy):</p>\n<ul>\n<li>If the input data contains \"house\", \"house,\", \"house?\", you include the 3 options in the LM corpus;</li>\n<li>\"house-\" and \"house!\" would not be included.</li>\n</ul>\n<p>Did I get it right?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2490978,
                  "author_name": "kcotton21",
                  "author_url": "",
                  "post_date": "10/21/2023 08:31:11",
                  "content": "<p>Thanks for the additional questions.</p>\n<p>Assume the following case for the CTC tokenizer vocabulary.<br>\nHave: \",\" and \"?\"<br>\nDoes not have: \"-\" and \"!\"</p>\n<p>In this case, the possible words in the LM corpus are \"house\", \"house,\" and \"house?\"</p>\n<p>I wasn't sure if wav2vec could handle the symbols well, but I thought they should be recovered from the information in the acoustic model. (e.g., inflection and breath between words). However, in the top solution, the punctuation model with only textual information seems to work well enough. I was a bit surprised.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2487050": "Thank you for organizing this competition.\nI am new to speech recognition and have learned a lot. Thanks to the competition participants and hosts.\n\n# Solution Overview\n- Dataset cleaning\n- Punctuation\n- Simple Ensemble (Rescoring)\n\n# Dataset\nFirst, I applied a data filter. Since I'm not familiar with the audio domain, I used a simple filter.\nOne was the ratio between the transcribed text and the length of the audio data. I filtered the data based on a 20 characters per audio second. If it's higher than that, I considered the possibility that there are too many characters for the length of the audio data. It could be abnormally fast speech or corrupted audio data.\nAnother filter I used was based on the prediction results from the yellow_king model. I filtered out data with scores above a certain threshold (CER > 0.6).\nBy training on clean data, I was able to start off with a good score.\n\nBecause the competition data set was so large, training was repeated on a subset of it rather than using all of it. Once the model achieved some performance, the WER for the remaining data set was calculated and added to the data. This round was repeated several times.\n\n## external data\nkathbath, openslr37, 53, ulca\nThese external data were also filtered\n\n# Model\n`IndicWav2Vec`\nThis model was very powerful for the test set.\nSince the previous competition report suggested that retraining a model once it had converged would not perform well, I used a pre-trained model instead of bengali's fine-tuned checkpoints (converted to HF model from faireseq checkpoints). Another reason was that I customized vocab (Mentioned in the section on punctuation.).\n\n## LM Model\n5-gram Language Model (using all of Indiccorpus v1 and v2, BangaLM, and mC4). The text corpus totaled 45GB, with the 5gram.bin file being around 22GB in size.\nUsing pyctcdecoder, alpha, beta and beam_width were adjusted with example audio CV. This initially provided a 0.01 boost to the LB score, but the benefit decreased as the score improved, eventually settling at around 0.003.\n\n# Punctuation\npunctuation is very important.\nInstead of creating a punctuation model, I took the simple approach of including punctuation in the vocab and LM corpus. Even when including punctuation, wav2vec was largely unable to predict it. However, it was found that including them in the vocab allowed them to be considered as candidates during LM decoding. Simply by including punctuation in both the model and the LM corpus, it was possible to improve the score from 0.44 to 0.428.\n\n# Ensemble (Rescoring)\nDue to the structure of the output, it was not possible to simple ensemble at the logit level. Therefore, a strategy was adopted to calculate the LM Score for decoded sentences using the LM model, and the sentence with the highest score was selected.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197187%2F5e6ee0ffb1f1b28b93f1a744fe77c438%2Fbengali.png?generation=1697624102630269&alt=media)\n\nSentence used the output of three models with variations on training data and decoding parameters. Both public and private models improve by about 0.006 - 0.008.\n\n# Not working for me\n- xls-r-300m\n    - Common Voice's CV is competitive with indicwav2vec, but performs very poorly for the test set. Not effective when added to an ensemble.\n- Speech Enhancement\n - It was found that the score was poor for example audio with reverb, so noise removal was performed using DeepFilterNet. There were cases where this boosted the LB score and cases where it didn't, so it was finally not used.\n- Adding Assamese\n - my understanding is that it shares characters with Bengali. Addition did not improve score.",
    "2487921": "Congratulations on your gold medal and thanks for sharing.\n\nA few questions:\n1. Could you please elaborate on how you predicted the LM score? Did you use a separate model, a second output of the encoder, ...? What training loss did you use?\n2. What characters did you consider for punctuation? All (|?!.,\"'- etc.) or only some of them?\n3. Did you keep consecutive punctuation characters or reduced it to just one character?\n4. How did you add punctuation to the LM corpus? did you add it to all options or considered only what appeared in the available corpus?\n5. What hardware did you use for training?",
    "2488060": "Thank you. Please feel free to ask additional questions if anything is unclear.\n\n1.No specific LMs have been created.\nkenLM scores used in Wav2Vec2ProcessorWithLM were used.\n```\nmodel = kenlm.LanguageModel(\"5gram.bin\")\nscore = model.score(pred_1)\n```\nMy understanding is that the more natural the sentence, the better the score.\n\n2.The top four symbols in the competition data were included. `[',', '-', '?', '!']`\n3-4.The LM corpus only retains symbols and vocabulary contained in the tokenizer vocab. At this time, I did not single out consecutive symbols (but I think I should have done so). I replaced it with one in the predict sentence. \n5.single A6000 or RTX 4090.",
    "2490711": "Thanks for the responses.\n\nIf I understood well you response to 4., you included in the LM corpus only the words with the punctuations present in the input data. Let me illustrate with an English word (easier for me in light of my utter Bengali illiteracy):\n- If the input data contains \"house\", \"house,\", \"house?\", you include the 3 options in the LM corpus;\n- \"house-\" and \"house!\" would not be included.\n\nDid I get it right?",
    "2490978": "Thanks for the additional questions.\n\nAssume the following case for the CTC tokenizer vocabulary.\nHave: \",\" and \"?\"\nDoes not have: \"-\" and \"!\"\n\nIn this case, the possible words in the LM corpus are \"house\", \"house,\" and \"house?\"\n\nI wasn't sure if wav2vec could handle the symbols well, but I thought they should be recovered from the information in the acoustic model. (e.g., inflection and breath between words). However, in the top solution, the punctuation model with only textual information seems to work well enough. I was a bit surprised."
  },
  "source": "meta"
}