{
  "id": 447957,
  "title": "3rd place solution",
  "url": "/competitions/bengaliai-speech/discussion/447957",
  "author_name": "luddite^",
  "post_date": "2023-10-17T23:59:54.867000",
  "votes": 43,
  "comment_count": 20,
  "views": 0,
  "content": "<p>First of all, we want to thank the competition organizers who hosted this fantastic competition and gave us an excellent opportunity to learn how to improve the ASR model against low-resource languages like Bengali and other competitors who generously shared their knowledge.<br>\nThe following is a brief explanation of our solution. We will open-source more detailed codes later, but if you have questions, ask us freely.</p>\n<hr>\n<h1>Model Architecture</h1>\n<h3>・CTC</h3>\n<p>We fine-tuned \"ai4bharat/indicwav2vec_v1_bengali\" with competition data(CD).<br>\nWe observed low-quality samples in CD, and mindlessly fine-tuning a model with all the CD deteriorated its performance. So first, we fine-tuned a model only with split=”valid” CD (this improved the model’s performance) and predicted with it against split=”train” CD. After that, we included a high-quality split=’train’ CD (WER&lt;0.75) to split=’valid’ CD and fine-tuned  \"ai4bharat/indicwav2vec_v1_bengali\" from scratch. <br>\nThis improved the public baseline to LB=0.405.</p>\n<h3>・kenlm</h3>\n<p>Because there are many out-of-vocabulary in OOD data, we thought training strong LM with large external text data is important. So we downloaded text data and script of ASR datasets(CD, indicCorp v2, common voice, fleurs, openslr, openslr37, and oscar) and trained 5-gram LM.<br>\nThis LM improved LB score by about 0.01 compared with \"arijitx/wav2vec2-xls-r-300m-bengali\"</p>\n<h1>Data</h1>\n<h3>・audio data</h3>\n<p>We used only the CD. As mentioned in the  Model Architecture section, we did something like an adversarial validation using a CTC model trained with split=’valid’ and used about 70% of the all CD.</p>\n<h3>・text data</h3>\n<p>We used text data and script of ASR datasets(CD, indicCorp v2, common voice, fleurs, openslr, openslr37, and oscar). As a preprocessing, we normalized the text with bnUnicodeNormalizer and removed some characters('[\\,\\?.!-\\;:\\\"\\।\\—]').</p>\n<h1>Inference</h1>\n<h3>・sort data by audio length</h3>\n<p>Padding at the end of audios negatively affected CTC, so we sorted the data based on the audio length and dynamically padded each batch. This increased prediction speed and performance.</p>\n<h3>・Denoising with Demucs, a music source separation model.</h3>\n<p>We utilized Demucs to denoise audios. This improved the LB score by about 0.003.</p>\n<h3>・Judge if we use Demucs or not</h3>\n<p>Demucs sometimes make the audio worse, so we evaluated if the audio gets worse and we switched the audio used in a prediction. This improved the LB score by about 0.001.<br>\nThe procedure is as follows.</p>\n<ol>\n<li>we made two predictions: prediction with Demucs and prediction without Demucs. To speed up this prediction, we made the LM parameter beam_width=10.</li>\n<li>we compared the number of tokens in two predictions. If the number of tokens in the prediction with Demucs is shorter than the other, we predicted without Demucs. Otherwise, we predicted with Demucs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2Fc66948ac5c61830bd513e6af4564cfa4%2Fdenoising_solution.png?generation=1697586862740699&amp;alt=media\" alt=\"\"></li>\n</ol>\n<h1>Post Processing</h1>\n<h3>・punctuation model</h3>\n<p>We built models that predict punctuations that go into the spaces between words. These improved LB by more than 0.030.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2Fdc61bcc13bd81e9a0259a3df7b95ac6f%2Fesprit_solution.png?generation=1697586896853219&amp;alt=media\" alt=\"\"></p>\n<p>As a tip for training, the CV score was better improved by setting the loss weight of \"PAD\" to 0.0.<br>\nBackborn: xlm-roberta-large, xlm-roberta-base<br>\nTrainer: XLMRobertaForTokenClassification<br>\nDataset: train.csv (given), indicCorp v2<br>\nPunctuations: [ ,।?-]</p>\n<hr>\n<p>・<a href=\"https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/tree/main\" target=\"_blank\">CTC and LM train code</a><br>\n・<a href=\"https://github.com/espritmirai/bengali-punctuation-model\" target=\"_blank\">punctuation model code</a><br>\n・<a href=\"https://www.kaggle.com/code/takuji/3rd-place-solution?scriptVersionId=147177786\" target=\"_blank\">inference notebook</a></p>",
  "messages": [
    {
      "id": 2486446,
      "postDate": "2023-10-17T23:59:54.867Z",
      "content": "<p>First of all, we want to thank the competition organizers who hosted this fantastic competition and gave us an excellent opportunity to learn how to improve the ASR model against low-resource languages like Bengali and other competitors who generously shared their knowledge.<br>\nThe following is a brief explanation of our solution. We will open-source more detailed codes later, but if you have questions, ask us freely.</p>\n<hr>\n<h1>Model Architecture</h1>\n<h3>・CTC</h3>\n<p>We fine-tuned \"ai4bharat/indicwav2vec_v1_bengali\" with competition data(CD).<br>\nWe observed low-quality samples in CD, and mindlessly fine-tuning a model with all the CD deteriorated its performance. So first, we fine-tuned a model only with split=”valid” CD (this improved the model’s performance) and predicted with it against split=”train” CD. After that, we included a high-quality split=’train’ CD (WER&lt;0.75) to split=’valid’ CD and fine-tuned  \"ai4bharat/indicwav2vec_v1_bengali\" from scratch. <br>\nThis improved the public baseline to LB=0.405.</p>\n<h3>・kenlm</h3>\n<p>Because there are many out-of-vocabulary in OOD data, we thought training strong LM with large external text data is important. So we downloaded text data and script of ASR datasets(CD, indicCorp v2, common voice, fleurs, openslr, openslr37, and oscar) and trained 5-gram LM.<br>\nThis LM improved LB score by about 0.01 compared with \"arijitx/wav2vec2-xls-r-300m-bengali\"</p>\n<h1>Data</h1>\n<h3>・audio data</h3>\n<p>We used only the CD. As mentioned in the  Model Architecture section, we did something like an adversarial validation using a CTC model trained with split=’valid’ and used about 70% of the all CD.</p>\n<h3>・text data</h3>\n<p>We used text data and script of ASR datasets(CD, indicCorp v2, common voice, fleurs, openslr, openslr37, and oscar). As a preprocessing, we normalized the text with bnUnicodeNormalizer and removed some characters('[\\,\\?.!-\\;:\\\"\\।\\—]').</p>\n<h1>Inference</h1>\n<h3>・sort data by audio length</h3>\n<p>Padding at the end of audios negatively affected CTC, so we sorted the data based on the audio length and dynamically padded each batch. This increased prediction speed and performance.</p>\n<h3>・Denoising with Demucs, a music source separation model.</h3>\n<p>We utilized Demucs to denoise audios. This improved the LB score by about 0.003.</p>\n<h3>・Judge if we use Demucs or not</h3>\n<p>Demucs sometimes make the audio worse, so we evaluated if the audio gets worse and we switched the audio used in a prediction. This improved the LB score by about 0.001.<br>\nThe procedure is as follows.</p>\n<ol>\n<li>we made two predictions: prediction with Demucs and prediction without Demucs. To speed up this prediction, we made the LM parameter beam_width=10.</li>\n<li>we compared the number of tokens in two predictions. If the number of tokens in the prediction with Demucs is shorter than the other, we predicted without Demucs. Otherwise, we predicted with Demucs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2Fc66948ac5c61830bd513e6af4564cfa4%2Fdenoising_solution.png?generation=1697586862740699&amp;alt=media\" alt=\"\"></li>\n</ol>\n<h1>Post Processing</h1>\n<h3>・punctuation model</h3>\n<p>We built models that predict punctuations that go into the spaces between words. These improved LB by more than 0.030.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2Fdc61bcc13bd81e9a0259a3df7b95ac6f%2Fesprit_solution.png?generation=1697586896853219&amp;alt=media\" alt=\"\"></p>\n<p>As a tip for training, the CV score was better improved by setting the loss weight of \"PAD\" to 0.0.<br>\nBackborn: xlm-roberta-large, xlm-roberta-base<br>\nTrainer: XLMRobertaForTokenClassification<br>\nDataset: train.csv (given), indicCorp v2<br>\nPunctuations: [ ,।?-]</p>\n<hr>\n<p>・<a href=\"https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/tree/main\" target=\"_blank\">CTC and LM train code</a><br>\n・<a href=\"https://github.com/espritmirai/bengali-punctuation-model\" target=\"_blank\">punctuation model code</a><br>\n・<a href=\"https://www.kaggle.com/code/takuji/3rd-place-solution?scriptVersionId=147177786\" target=\"_blank\">inference notebook</a></p>",
      "rawMarkdown": "First of all, we want to thank the competition organizers who hosted this fantastic competition and gave us an excellent opportunity to learn how to improve the ASR model against low-resource languages like Bengali and other competitors who generously shared their knowledge.\nThe following is a brief explanation of our solution. We will open-source more detailed codes later, but if you have questions, ask us freely.\n\n----\n# Model Architecture\n### ・CTC\nWe fine-tuned \"ai4bharat/indicwav2vec_v1_bengali\" with competition data(CD).\nWe observed low-quality samples in CD, and mindlessly fine-tuning a model with all the CD deteriorated its performance. So first, we fine-tuned a model only with split=”valid” CD (this improved the model’s performance) and predicted with it against split=”train” CD. After that, we included a high-quality split=’train’ CD (WER<0.75) to split=’valid’ CD and fine-tuned  \"ai4bharat/indicwav2vec_v1_bengali\" from scratch. \nThis improved the public baseline to LB=0.405.\n\n### ・kenlm\nBecause there are many out-of-vocabulary in OOD data, we thought training strong LM with large external text data is important. So we downloaded text data and script of ASR datasets(CD, indicCorp v2, common voice, fleurs, openslr, openslr37, and oscar) and trained 5-gram LM.\nThis LM improved LB score by about 0.01 compared with \"arijitx/wav2vec2-xls-r-300m-bengali\"\n\n# Data\n### ・audio data\nWe used only the CD. As mentioned in the  Model Architecture section, we did something like an adversarial validation using a CTC model trained with split=’valid’ and used about 70% of the all CD.\n\n\n### ・text data\nWe used text data and script of ASR datasets(CD, indicCorp v2, common voice, fleurs, openslr, openslr37, and oscar). As a preprocessing, we normalized the text with bnUnicodeNormalizer and removed some characters('[\\,\\?\\.\\!\\-\\;\\:\\\"\\।\\—]').\n\n# Inference\n### ・sort data by audio length\nPadding at the end of audios negatively affected CTC, so we sorted the data based on the audio length and dynamically padded each batch. This increased prediction speed and performance.\n\n### ・Denoising with Demucs, a music source separation model.\nWe utilized Demucs to denoise audios. This improved the LB score by about 0.003.\n\n### ・Judge if we use Demucs or not\nDemucs sometimes make the audio worse, so we evaluated if the audio gets worse and we switched the audio used in a prediction. This improved the LB score by about 0.001.\nThe procedure is as follows.\n1. we made two predictions: prediction with Demucs and prediction without Demucs. To speed up this prediction, we made the LM parameter beam_width=10.\n2. we compared the number of tokens in two predictions. If the number of tokens in the prediction with Demucs is shorter than the other, we predicted without Demucs. Otherwise, we predicted with Demucs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2Fc66948ac5c61830bd513e6af4564cfa4%2Fdenoising_solution.png?generation=1697586862740699&alt=media)\n\n# Post Processing\n### ・punctuation model\nWe built models that predict punctuations that go into the spaces between words. These improved LB by more than 0.030.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2Fdc61bcc13bd81e9a0259a3df7b95ac6f%2Fesprit_solution.png?generation=1697586896853219&alt=media)\n\nAs a tip for training, the CV score was better improved by setting the loss weight of \"PAD\" to 0.0.\nBackborn: xlm-roberta-large, xlm-roberta-base\nTrainer: XLMRobertaForTokenClassification\nDataset: train.csv (given), indicCorp v2\nPunctuations: [ ,।?-]\n\n\n----\n・[CTC and LM train code](https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/tree/main)\n・[punctuation model code](https://github.com/espritmirai/bengali-punctuation-model)\n・[inference notebook](https://www.kaggle.com/code/takuji/3rd-place-solution?scriptVersionId=147177786)",
      "votes": 43
    },
    {
      "id": 2486840,
      "postDate": "2023-10-18T07:24:02.460Z",
      "content": "<p>Congratulations on the 3rd place! <br>\nDid xlm-roberta-large fit into GPU memory? I was always getting CUDA OOM</p>",
      "rawMarkdown": "Congratulations on the 3rd place! \nDid xlm-roberta-large fit into GPU memory? I was always getting CUDA OOM",
      "votes": 1,
      "replies": [
        {
          "id": 2487118,
          "postDate": "2023-10-18T11:29:07.950Z",
          "content": "<p>I was able to train with <code>batch_size=32</code> using a single RTX3090 or 4090.<br>\nAnd, your public notebook was very helpful. Thank you very much.</p>",
          "rawMarkdown": "I was able to train with `batch_size=32` using a single RTX3090 or 4090.\nAnd, your public notebook was very helpful. Thank you very much.",
          "votes": 1,
          "replies": [
            {
              "id": 2487187,
              "postDate": "2023-10-18T12:12:48.190Z",
              "content": "<p>Glad I could help. Thanks for your kind words!</p>",
              "rawMarkdown": "Glad I could help. Thanks for your kind words!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2486567,
      "postDate": "2023-10-18T03:08:13.700Z",
      "content": "<p>Congratulations on the 3rd place and thanks for sharing. </p>\n<p>Yesterday, I was contemplating two optimizations, one being a punctuation model. Considering the percentage of words with punctuation it seemed a good bet to improve a few points. I first thought about predicting punctuation in free spaces with a classification model, but after compiling occurrences, I noticed that there are many combinations of two or more consecutive characters, e.g., \"| and |||. After a quick look online (unfortunately I'm totally ignorant on Bengali), I felt that a few might be mistakes, but most reflected correct punctuation. </p>\n<p>Hence, to predict punctuation in blank spaces I'd need to consider multiple possibilities, even after discarding the ones with few occurrences. I felt that a seq2seq model would be better suited to learn how to do that than a classification model. I ended up pursuing the other idea (it had a higher potential, though unfortunately I ran out of time), but I'm still curious about what would have been the way to go. I'd appreciate your thoughts on it.</p>",
      "rawMarkdown": "Congratulations on the 3rd place and thanks for sharing. \n\nYesterday, I was contemplating two optimizations, one being a punctuation model. Considering the percentage of words with punctuation it seemed a good bet to improve a few points. I first thought about predicting punctuation in free spaces with a classification model, but after compiling occurrences, I noticed that there are many combinations of two or more consecutive characters, e.g., \"| and |||. After a quick look online (unfortunately I'm totally ignorant on Bengali), I felt that a few might be mistakes, but most reflected correct punctuation. \n\nHence, to predict punctuation in blank spaces I'd need to consider multiple possibilities, even after discarding the ones with few occurrences. I felt that a seq2seq model would be better suited to learn how to do that than a classification model. I ended up pursuing the other idea (it had a higher potential, though unfortunately I ran out of time), but I'm still curious about what would have been the way to go. I'd appreciate your thoughts on it.",
      "votes": 1,
      "replies": [
        {
          "id": 2486571,
          "postDate": "2023-10-18T03:23:31.457Z",
          "content": "<p>In our model, consecutive <code>[।!?]</code> in the train data were replaced by a single character <code>\"।\"</code> or <code>\"?\"</code>. We also removed sentences containing consecutive <code>[,-]</code> from the train data. However, we have not tried anything else, so we cannot say for sure if this was the best idea.</p>\n<p>For punctuation frequency, we decided to predict <code>\"-\"</code> as well, since there were many <code>\"-\"</code> in given train.csv.</p>\n<p>We also tried to create a seq2seq model (backborn: T5), but this did not work at all. If there is a success solution I would like to see it too.</p>",
          "rawMarkdown": "In our model, consecutive `[।!?]` in the train data were replaced by a single character `\"।\"` or `\"?\"`. We also removed sentences containing consecutive `[,-]` from the train data. However, we have not tried anything else, so we cannot say for sure if this was the best idea.\n\nFor punctuation frequency, we decided to predict `\"-\"` as well, since there were many `\"-\"` in given train.csv.\n\nWe also tried to create a seq2seq model (backborn: T5), but this did not work at all. If there is a success solution I would like to see it too.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2486454,
      "postDate": "2023-10-18T00:10:55.273Z",
      "content": "<p>Thank you for sharing! Liked the postprocessing steps.</p>\n<p>For how many epochs did you train your final model and with how much data in total?</p>",
      "rawMarkdown": "Thank you for sharing! Liked the postprocessing steps.\n\nFor how many epochs did you train your final model and with how much data in total?",
      "votes": 1,
      "replies": [
        {
          "id": 2486494,
          "postDate": "2023-10-18T01:16:17.297Z",
          "content": "<p>Which model (CTC, LM, punctuation)?</p>",
          "rawMarkdown": "Which model (CTC, LM, punctuation)?",
          "votes": 1
        },
        {
          "id": 2486539,
          "postDate": "2023-10-18T02:36:15.130Z",
          "content": "<p>In the training of the punctuation model, 17M Bengali sentences were trained for 1~2 epochs.</p>",
          "rawMarkdown": "In the training of the punctuation model, 17M Bengali sentences were trained for 1~2 epochs.",
          "votes": 1,
          "replies": [
            {
              "id": 2487949,
              "postDate": "2023-10-18T22:36:40.603Z",
              "content": "<p>I was curious about the CTC model, thanks for sharing the details for the punctuation model as well though.</p>",
              "rawMarkdown": "I was curious about the CTC model, thanks for sharing the details for the punctuation model as well though.",
              "votes": 1
            },
            {
              "id": 2488041,
              "postDate": "2023-10-19T01:22:35.647Z",
              "content": "<p>We trained the CTC model for 10 epochs with 671231 data(split='valid': 28855, cleaned split='train': 642376).</p>",
              "rawMarkdown": "We trained the CTC model for 10 epochs with 671231 data(split='valid': 28855, cleaned split='train': 642376).",
              "votes": 1
            },
            {
              "id": 2488400,
              "postDate": "2023-10-19T08:39:49.550Z",
              "content": "<p>Oh, nice! Thanks for sharing. Was wondering if it was manageable or not and apparently yes! Well done again! </p>",
              "rawMarkdown": "Oh, nice! Thanks for sharing. Was wondering if it was manageable or not and apparently yes! Well done again! ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2486505,
      "postDate": "2023-10-18T01:27:31.777Z",
      "content": "<p>Congratulations for your 3rd place. \"Predicting punctuations that go into the spaces between words\" sounds amazing for a newbie like me.</p>",
      "rawMarkdown": "Congratulations for your 3rd place. \"Predicting punctuations that go into the spaces between words\" sounds amazing for a newbie like me.",
      "votes": 2
    },
    {
      "id": 2491953,
      "postDate": "2023-10-22T06:55:28.637Z",
      "content": "<p>Congratulations on the 3rd place 🎉<br>\nand thanks for sharing your knowledge!</p>\n<p>Two questions about \"sort data by audio length\".</p>\n<ol>\n<li><p>It means shuffle=False? <br>\nref. <a href=\"https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L239\" target=\"_blank\">https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L239</a></p></li>\n<li><p>Sorted by ascending order?<br>\nref. <a href=\"https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L199\" target=\"_blank\">https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L199</a></p></li>\n</ol>",
      "rawMarkdown": "Congratulations on the 3rd place :tada:\nand thanks for sharing your knowledge!\n\nTwo questions about \"sort data by audio length\".\n\n1. It means shuffle=False? \nref. https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L239\n\n2. Sorted by ascending order?\nref. https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L199\n",
      "replies": [
        {
          "id": 2492310,
          "postDate": "2023-10-22T11:31:25.110Z",
          "content": "<p>During inference, we sorted data based on audio_length in ascending order and set the dataloader's variable shuffle=False.</p>",
          "rawMarkdown": "During inference, we sorted data based on audio_length in ascending order and set the dataloader's variable shuffle=False.",
          "votes": 1,
          "replies": [
            {
              "id": 2492458,
              "postDate": "2023-10-22T14:33:49.100Z",
              "content": "<p>thank you for answering!</p>",
              "rawMarkdown": "thank you for answering!"
            }
          ]
        }
      ]
    },
    {
      "id": 2486459,
      "postDate": "2023-10-18T00:12:49.757Z",
      "content": "<p>Congratulations! I have a question:</p>\n<blockquote>\n  <p>So first, we fine-tuned a model only with split=”valid” CD (this improved the model’s performance) and predicted with it against split=”train” CD. After that, we included a high-quality split=’train’ CD (WER&lt;0.75) to split=’valid’ CD</p>\n</blockquote>\n<p>In the second step did you use the original train split data or the new pseudo labeled data to train the model?</p>",
      "rawMarkdown": "Congratulations! I have a question:\n>So first, we fine-tuned a model only with split=”valid” CD (this improved the model’s performance) and predicted with it against split=”train” CD. After that, we included a high-quality split=’train’ CD (WER<0.75) to split=’valid’ CD\n\nIn the second step did you use the original train split data or the new pseudo labeled data to train the model?",
      "replies": [
        {
          "id": 2486472,
          "postDate": "2023-10-18T00:41:50.950Z",
          "content": "<p>Thank you!<br>\nWe used the original train split data at 2nd step.</p>",
          "rawMarkdown": "Thank you!\nWe used the original train split data at 2nd step.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2486452,
      "postDate": "2023-10-18T00:08:26.947Z",
      "content": "<p>Awesome solution! I wonder how to let the LM be uploaded in kaggle input without exceeding the 20G(?) capacity limit. ARPA may seem too large when I train a LM with too much corpus.</p>",
      "rawMarkdown": "Awesome solution! I wonder how to let the LM be uploaded in kaggle input without exceeding the 20G(?) capacity limit. ARPA may seem too large when I train a LM with too much corpus.",
      "replies": [
        {
          "id": 2486477,
          "postDate": "2023-10-18T00:50:27.227Z",
          "content": "<p>Thank you.<br>\nI don't understand what 20GB limit is. The limit of data uploaded to kaggle 107GB. By deleting or making public your datasets, you can upload over 20GB files.</p>",
          "rawMarkdown": "Thank you.\nI don't understand what 20GB limit is. The limit of data uploaded to kaggle 107GB. By deleting or making public your datasets, you can upload over 20GB files."
        }
      ]
    },
    {
      "id": 2487849,
      "postDate": "2023-10-18T19:48:46.843Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2486840,
      "author_name": "Md Boktiar Mahbub Murad",
      "author_url": "",
      "post_date": "2023-10-18T07:24:02.460000",
      "content": "<p>Congratulations on the 3rd place! <br>\nDid xlm-roberta-large fit into GPU memory? I was always getting CUDA OOM</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2487118,
          "author_name": "esprit",
          "author_url": "",
          "post_date": "2023-10-18T11:29:07.950000",
          "content": "<p>I was able to train with <code>batch_size=32</code> using a single RTX3090 or 4090.<br>\nAnd, your public notebook was very helpful. Thank you very much.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2487187,
              "author_name": "Md Boktiar Mahbub Murad",
              "author_url": "",
              "post_date": "2023-10-18T12:12:48.190000",
              "content": "<p>Glad I could help. Thanks for your kind words!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2486567,
      "author_name": "vialactea",
      "author_url": "",
      "post_date": "2023-10-18T03:08:13.700000",
      "content": "<p>Congratulations on the 3rd place and thanks for sharing. </p>\n<p>Yesterday, I was contemplating two optimizations, one being a punctuation model. Considering the percentage of words with punctuation it seemed a good bet to improve a few points. I first thought about predicting punctuation in free spaces with a classification model, but after compiling occurrences, I noticed that there are many combinations of two or more consecutive characters, e.g., \"| and |||. After a quick look online (unfortunately I'm totally ignorant on Bengali), I felt that a few might be mistakes, but most reflected correct punctuation. </p>\n<p>Hence, to predict punctuation in blank spaces I'd need to consider multiple possibilities, even after discarding the ones with few occurrences. I felt that a seq2seq model would be better suited to learn how to do that than a classification model. I ended up pursuing the other idea (it had a higher potential, though unfortunately I ran out of time), but I'm still curious about what would have been the way to go. I'd appreciate your thoughts on it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2486571,
          "author_name": "esprit",
          "author_url": "",
          "post_date": "2023-10-18T03:23:31.457000",
          "content": "<p>In our model, consecutive <code>[।!?]</code> in the train data were replaced by a single character <code>\"।\"</code> or <code>\"?\"</code>. We also removed sentences containing consecutive <code>[,-]</code> from the train data. However, we have not tried anything else, so we cannot say for sure if this was the best idea.</p>\n<p>For punctuation frequency, we decided to predict <code>\"-\"</code> as well, since there were many <code>\"-\"</code> in given train.csv.</p>\n<p>We also tried to create a seq2seq model (backborn: T5), but this did not work at all. If there is a success solution I would like to see it too.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2486454,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2023-10-18T00:10:55.273000",
      "content": "<p>Thank you for sharing! Liked the postprocessing steps.</p>\n<p>For how many epochs did you train your final model and with how much data in total?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2486494,
          "author_name": "luddite^",
          "author_url": "",
          "post_date": "2023-10-18T01:16:17.297000",
          "content": "<p>Which model (CTC, LM, punctuation)?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2486539,
          "author_name": "esprit",
          "author_url": "",
          "post_date": "2023-10-18T02:36:15.130000",
          "content": "<p>In the training of the punctuation model, 17M Bengali sentences were trained for 1~2 epochs.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2487949,
              "author_name": "Sinan Calisir",
              "author_url": "",
              "post_date": "2023-10-18T22:36:40.603000",
              "content": "<p>I was curious about the CTC model, thanks for sharing the details for the punctuation model as well though.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2488041,
              "author_name": "luddite^",
              "author_url": "",
              "post_date": "2023-10-19T01:22:35.647000",
              "content": "<p>We trained the CTC model for 10 epochs with 671231 data(split='valid': 28855, cleaned split='train': 642376).</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2488400,
              "author_name": "Sinan Calisir",
              "author_url": "",
              "post_date": "2023-10-19T08:39:49.550000",
              "content": "<p>Oh, nice! Thanks for sharing. Was wondering if it was manageable or not and apparently yes! Well done again! </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2486505,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2023-10-18T01:27:31.777000",
      "content": "<p>Congratulations for your 3rd place. \"Predicting punctuations that go into the spaces between words\" sounds amazing for a newbie like me.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2491953,
      "author_name": "sadahry",
      "author_url": "",
      "post_date": "2023-10-22T06:55:28.637000",
      "content": "<p>Congratulations on the 3rd place 🎉<br>\nand thanks for sharing your knowledge!</p>\n<p>Two questions about \"sort data by audio length\".</p>\n<ol>\n<li><p>It means shuffle=False? <br>\nref. <a href=\"https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L239\" target=\"_blank\">https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L239</a></p></li>\n<li><p>Sorted by ascending order?<br>\nref. <a href=\"https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L199\" target=\"_blank\">https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L199</a></p></li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 2492310,
          "author_name": "luddite^",
          "author_url": "",
          "post_date": "2023-10-22T11:31:25.110000",
          "content": "<p>During inference, we sorted data based on audio_length in ascending order and set the dataloader's variable shuffle=False.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2492458,
              "author_name": "sadahry",
              "author_url": "",
              "post_date": "2023-10-22T14:33:49.100000",
              "content": "<p>thank you for answering!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2486459,
      "author_name": "Man of the year",
      "author_url": "",
      "post_date": "2023-10-18T00:12:49.757000",
      "content": "<p>Congratulations! I have a question:</p>\n<blockquote>\n  <p>So first, we fine-tuned a model only with split=”valid” CD (this improved the model’s performance) and predicted with it against split=”train” CD. After that, we included a high-quality split=’train’ CD (WER&lt;0.75) to split=’valid’ CD</p>\n</blockquote>\n<p>In the second step did you use the original train split data or the new pseudo labeled data to train the model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2486472,
          "author_name": "luddite^",
          "author_url": "",
          "post_date": "2023-10-18T00:41:50.950000",
          "content": "<p>Thank you!<br>\nWe used the original train split data at 2nd step.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2486452,
      "author_name": "iiiiitsu",
      "author_url": "",
      "post_date": "2023-10-18T00:08:26.947000",
      "content": "<p>Awesome solution! I wonder how to let the LM be uploaded in kaggle input without exceeding the 20G(?) capacity limit. ARPA may seem too large when I train a LM with too much corpus.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2486477,
          "author_name": "luddite^",
          "author_url": "",
          "post_date": "2023-10-18T00:50:27.227000",
          "content": "<p>Thank you.<br>\nI don't understand what 20GB limit is. The limit of data uploaded to kaggle 107GB. By deleting or making public your datasets, you can upload over 20GB files.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2487849,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:48:46.843000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486446": "First of all, we want to thank the competition organizers who hosted this fantastic competition and gave us an excellent opportunity to learn how to improve the ASR model against low-resource languages like Bengali and other competitors who generously shared their knowledge.\nThe following is a brief explanation of our solution. We will open-source more detailed codes later, but if you have questions, ask us freely.\n\n----\n# Model Architecture\n### ・CTC\nWe fine-tuned \"ai4bharat/indicwav2vec_v1_bengali\" with competition data(CD).\nWe observed low-quality samples in CD, and mindlessly fine-tuning a model with all the CD deteriorated its performance. So first, we fine-tuned a model only with split=”valid” CD (this improved the model’s performance) and predicted with it against split=”train” CD. After that, we included a high-quality split=’train’ CD (WER<0.75) to split=’valid’ CD and fine-tuned  \"ai4bharat/indicwav2vec_v1_bengali\" from scratch. \nThis improved the public baseline to LB=0.405.\n\n### ・kenlm\nBecause there are many out-of-vocabulary in OOD data, we thought training strong LM with large external text data is important. So we downloaded text data and script of ASR datasets(CD, indicCorp v2, common voice, fleurs, openslr, openslr37, and oscar) and trained 5-gram LM.\nThis LM improved LB score by about 0.01 compared with \"arijitx/wav2vec2-xls-r-300m-bengali\"\n\n# Data\n### ・audio data\nWe used only the CD. As mentioned in the  Model Architecture section, we did something like an adversarial validation using a CTC model trained with split=’valid’ and used about 70% of the all CD.\n\n\n### ・text data\nWe used text data and script of ASR datasets(CD, indicCorp v2, common voice, fleurs, openslr, openslr37, and oscar). As a preprocessing, we normalized the text with bnUnicodeNormalizer and removed some characters('[\\,\\?\\.\\!\\-\\;\\:\\\"\\।\\—]').\n\n# Inference\n### ・sort data by audio length\nPadding at the end of audios negatively affected CTC, so we sorted the data based on the audio length and dynamically padded each batch. This increased prediction speed and performance.\n\n### ・Denoising with Demucs, a music source separation model.\nWe utilized Demucs to denoise audios. This improved the LB score by about 0.003.\n\n### ・Judge if we use Demucs or not\nDemucs sometimes make the audio worse, so we evaluated if the audio gets worse and we switched the audio used in a prediction. This improved the LB score by about 0.001.\nThe procedure is as follows.\n1. we made two predictions: prediction with Demucs and prediction without Demucs. To speed up this prediction, we made the LM parameter beam_width=10.\n2. we compared the number of tokens in two predictions. If the number of tokens in the prediction with Demucs is shorter than the other, we predicted without Demucs. Otherwise, we predicted with Demucs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2Fc66948ac5c61830bd513e6af4564cfa4%2Fdenoising_solution.png?generation=1697586862740699&alt=media)\n\n# Post Processing\n### ・punctuation model\nWe built models that predict punctuations that go into the spaces between words. These improved LB by more than 0.030.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2Fdc61bcc13bd81e9a0259a3df7b95ac6f%2Fesprit_solution.png?generation=1697586896853219&alt=media)\n\nAs a tip for training, the CV score was better improved by setting the loss weight of \"PAD\" to 0.0.\nBackborn: xlm-roberta-large, xlm-roberta-base\nTrainer: XLMRobertaForTokenClassification\nDataset: train.csv (given), indicCorp v2\nPunctuations: [ ,।?-]\n\n\n----\n・[CTC and LM train code](https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/tree/main)\n・[punctuation model code](https://github.com/espritmirai/bengali-punctuation-model)\n・[inference notebook](https://www.kaggle.com/code/takuji/3rd-place-solution?scriptVersionId=147177786)",
    "2486840": "Congratulations on the 3rd place! \nDid xlm-roberta-large fit into GPU memory? I was always getting CUDA OOM",
    "2486567": "Congratulations on the 3rd place and thanks for sharing. \n\nYesterday, I was contemplating two optimizations, one being a punctuation model. Considering the percentage of words with punctuation it seemed a good bet to improve a few points. I first thought about predicting punctuation in free spaces with a classification model, but after compiling occurrences, I noticed that there are many combinations of two or more consecutive characters, e.g., \"| and |||. After a quick look online (unfortunately I'm totally ignorant on Bengali), I felt that a few might be mistakes, but most reflected correct punctuation. \n\nHence, to predict punctuation in blank spaces I'd need to consider multiple possibilities, even after discarding the ones with few occurrences. I felt that a seq2seq model would be better suited to learn how to do that than a classification model. I ended up pursuing the other idea (it had a higher potential, though unfortunately I ran out of time), but I'm still curious about what would have been the way to go. I'd appreciate your thoughts on it.",
    "2486454": "Thank you for sharing! Liked the postprocessing steps.\n\nFor how many epochs did you train your final model and with how much data in total?",
    "2486505": "Congratulations for your 3rd place. \"Predicting punctuations that go into the spaces between words\" sounds amazing for a newbie like me.",
    "2491953": "Congratulations on the 3rd place :tada:\nand thanks for sharing your knowledge!\n\nTwo questions about \"sort data by audio length\".\n\n1. It means shuffle=False? \nref. https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L239\n\n2. Sorted by ascending order?\nref. https://github.com/sagawatatsuya/BengaliAI_Speech_Recognition_3rd_solution/blob/bedefebe763409cbc6b2a0461a325d8f90a5e166/train_CTC/stage1.py#L199\n",
    "2486459": "Congratulations! I have a question:\n>So first, we fine-tuned a model only with split=”valid” CD (this improved the model’s performance) and predicted with it against split=”train” CD. After that, we included a high-quality split=’train’ CD (WER<0.75) to split=’valid’ CD\n\nIn the second step did you use the original train split data or the new pseudo labeled data to train the model?",
    "2486452": "Awesome solution! I wonder how to let the LM be uploaded in kaggle input without exceeding the 20G(?) capacity limit. ARPA may seem too large when I train a LM with too much corpus.",
    "2487849": ""
  }
}