{
  "id": 447961,
  "title": "1st place solution",
  "url": "/competitions/bengaliai-speech/discussion/447961",
  "author_name": "tugstugi",
  "post_date": "2023-10-18T00:15:42.524000",
  "votes": 107,
  "comment_count": 38,
  "views": 0,
  "content": "<p>STT Model:</p>\n<ul>\n<li>OpenAI whisper-medium</li>\n<li>Huggingface trainer</li>\n<li>Trained on 8x 48GB RTX A6000</li>\n<li>bs=8 and lr=1e-5</li>\n<li>Train steps 50k</li>\n<li>Spectrogram dithering</li>\n<li>Spectrogram time and frequency masking</li>\n<li>Resampling 16khz-&gt;8khz-&gt;16khz as augmentation</li>\n<li>Inference with max_length=260, num_beams=4 and chunk_length_s=20.1s</li>\n<li>Libsonic based speed/pitch augmentation</li>\n<li>Datasets: OpenSLR 37, OpenSLR 53, MadASR, Shrutilipi, Macro, Kathbath, GoogleTTS generated audios and pseudo labeled YouTube videos</li>\n</ul>\n<p>Punctuation Model:</p>\n<ul>\n<li>AutoModelForTokenClassification google/muril-base-cased</li>\n<li>Huggingface trainer</li>\n<li>Labels: period, comma and question mark</li>\n<li>bs=64, lr=2e-4 and max_seq_length=512</li>\n<li>Ensemble of 4 models (using 6, 8, 11 and 12 layers of google/muril-base-cased)</li>\n<li>Normalized IndicCorp v2 Bangla dataset</li>\n</ul>\n<p>In my daily job, I do speech speech recognition for low resource central Asian languages. From my experience, OpenAI Whisper works really well for OOD audios and can even transribe song lyrics. The downside is it is very sensitive to the annotation noise. So fixing the annotation noise, is the most crucial part of this competition.</p>\n<p>Because the competition dataset was not validated, the initial model was trained on OpenSLR datasets. We normalized the texts and filtered out texts containing Bengali digits. All punctuation was also removed. Additionally, we sampled 420k texts from the IndicCorp and synthesized audios using GoogleTTS, which were then used as training datasets.</p>\n<p>Following the training of an initial Whisper-medium model on OpenSLR and GoogleTTS, we conducted inference on MadASR, Shrutilipi, Macro, and Kathbath. We included audios with a WER of less than 15% in the next training phase. After three rounds of training, the model achieved an 8% WER on the Macro validation dataset and a public leaderboard score of approximately 0.380.</p>\n<p>Since most of the training set audios were short, we merged some short audios to create around 70k longer audios. Subsequently, we achieved a public leaderboard score of approximately 0.370.</p>\n<p>Whisper with the original tokenizer was slow on Bengali audios. Therefore, we trained a Whisper tokenizer with a 12k vocabulary on Bengali texts. With this tokenizer, we were able to perform inference with a num_beam value of up to 8 and a chunk_length_s of 20.1 seconds in less than 7 hours.</p>\n<p>In the next step, we applied pseudo labeling to some YouTube videos, which enabled us to achieve a public leaderboard score of approximately 0.360. When we combined the predictions of four punctuation models, our public leaderboard score improved to around 0.325.</p>\n<p>By adding more pseudo-labeled YouTube videos, our public leaderboard score further improved to 0.312 (private LB 0.372)</p>\n<p>model weight and inference notebook: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/447970\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/447970</a><br>\ncleaned/long/pseudo data: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110</a></p>",
  "messages": [
    {
      "id": 2486461,
      "postDate": "2023-10-18T00:15:42.523Z",
      "content": "<p>STT Model:</p>\n<ul>\n<li>OpenAI whisper-medium</li>\n<li>Huggingface trainer</li>\n<li>Trained on 8x 48GB RTX A6000</li>\n<li>bs=8 and lr=1e-5</li>\n<li>Train steps 50k</li>\n<li>Spectrogram dithering</li>\n<li>Spectrogram time and frequency masking</li>\n<li>Resampling 16khz-&gt;8khz-&gt;16khz as augmentation</li>\n<li>Inference with max_length=260, num_beams=4 and chunk_length_s=20.1s</li>\n<li>Libsonic based speed/pitch augmentation</li>\n<li>Datasets: OpenSLR 37, OpenSLR 53, MadASR, Shrutilipi, Macro, Kathbath, GoogleTTS generated audios and pseudo labeled YouTube videos</li>\n</ul>\n<p>Punctuation Model:</p>\n<ul>\n<li>AutoModelForTokenClassification google/muril-base-cased</li>\n<li>Huggingface trainer</li>\n<li>Labels: period, comma and question mark</li>\n<li>bs=64, lr=2e-4 and max_seq_length=512</li>\n<li>Ensemble of 4 models (using 6, 8, 11 and 12 layers of google/muril-base-cased)</li>\n<li>Normalized IndicCorp v2 Bangla dataset</li>\n</ul>\n<p>In my daily job, I do speech speech recognition for low resource central Asian languages. From my experience, OpenAI Whisper works really well for OOD audios and can even transribe song lyrics. The downside is it is very sensitive to the annotation noise. So fixing the annotation noise, is the most crucial part of this competition.</p>\n<p>Because the competition dataset was not validated, the initial model was trained on OpenSLR datasets. We normalized the texts and filtered out texts containing Bengali digits. All punctuation was also removed. Additionally, we sampled 420k texts from the IndicCorp and synthesized audios using GoogleTTS, which were then used as training datasets.</p>\n<p>Following the training of an initial Whisper-medium model on OpenSLR and GoogleTTS, we conducted inference on MadASR, Shrutilipi, Macro, and Kathbath. We included audios with a WER of less than 15% in the next training phase. After three rounds of training, the model achieved an 8% WER on the Macro validation dataset and a public leaderboard score of approximately 0.380.</p>\n<p>Since most of the training set audios were short, we merged some short audios to create around 70k longer audios. Subsequently, we achieved a public leaderboard score of approximately 0.370.</p>\n<p>Whisper with the original tokenizer was slow on Bengali audios. Therefore, we trained a Whisper tokenizer with a 12k vocabulary on Bengali texts. With this tokenizer, we were able to perform inference with a num_beam value of up to 8 and a chunk_length_s of 20.1 seconds in less than 7 hours.</p>\n<p>In the next step, we applied pseudo labeling to some YouTube videos, which enabled us to achieve a public leaderboard score of approximately 0.360. When we combined the predictions of four punctuation models, our public leaderboard score improved to around 0.325.</p>\n<p>By adding more pseudo-labeled YouTube videos, our public leaderboard score further improved to 0.312 (private LB 0.372)</p>\n<p>model weight and inference notebook: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/447970\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/447970</a><br>\ncleaned/long/pseudo data: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110</a></p>",
      "rawMarkdown": "STT Model:\n- OpenAI whisper-medium\n- Huggingface trainer\n- Trained on 8x 48GB RTX A6000\n- bs=8 and lr=1e-5\n- Train steps 50k\n- Spectrogram dithering\n- Spectrogram time and frequency masking\n- Resampling 16khz->8khz->16khz as augmentation\n- Inference with max_length=260, num_beams=4 and chunk_length_s=20.1s\n- Libsonic based speed/pitch augmentation\n- Datasets: OpenSLR 37, OpenSLR 53, MadASR, Shrutilipi, Macro, Kathbath, GoogleTTS generated audios and pseudo labeled YouTube videos\n\n\nPunctuation Model:\n- AutoModelForTokenClassification google/muril-base-cased\n- Huggingface trainer\n- Labels: period, comma and question mark\n- bs=64, lr=2e-4 and max_seq_length=512\n- Ensemble of 4 models (using 6, 8, 11 and 12 layers of google/muril-base-cased)\n- Normalized IndicCorp v2 Bangla dataset\n\nIn my daily job, I do speech speech recognition for low resource central Asian languages. From my experience, OpenAI Whisper works really well for OOD audios and can even transribe song lyrics. The downside is it is very sensitive to the annotation noise. So fixing the annotation noise, is the most crucial part of this competition.\n\nBecause the competition dataset was not validated, the initial model was trained on OpenSLR datasets. We normalized the texts and filtered out texts containing Bengali digits. All punctuation was also removed. Additionally, we sampled 420k texts from the IndicCorp and synthesized audios using GoogleTTS, which were then used as training datasets.\n\nFollowing the training of an initial Whisper-medium model on OpenSLR and GoogleTTS, we conducted inference on MadASR, Shrutilipi, Macro, and Kathbath. We included audios with a WER of less than 15% in the next training phase. After three rounds of training, the model achieved an 8% WER on the Macro validation dataset and a public leaderboard score of approximately 0.380.\n\nSince most of the training set audios were short, we merged some short audios to create around 70k longer audios. Subsequently, we achieved a public leaderboard score of approximately 0.370.\n\nWhisper with the original tokenizer was slow on Bengali audios. Therefore, we trained a Whisper tokenizer with a 12k vocabulary on Bengali texts. With this tokenizer, we were able to perform inference with a num_beam value of up to 8 and a chunk_length_s of 20.1 seconds in less than 7 hours.\n\nIn the next step, we applied pseudo labeling to some YouTube videos, which enabled us to achieve a public leaderboard score of approximately 0.360. When we combined the predictions of four punctuation models, our public leaderboard score improved to around 0.325.\n\nBy adding more pseudo-labeled YouTube videos, our public leaderboard score further improved to 0.312 (private LB 0.372)\n\nmodel weight and inference notebook: https://www.kaggle.com/competitions/bengaliai-speech/discussion/447970\ncleaned/long/pseudo data: https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110",
      "votes": 107
    },
    {
      "id": 2486466,
      "postDate": "2023-10-18T00:20:57.933Z",
      "content": "<p>So the whisper do performs better. But why most of us didn't successed with finetuning whisper? Do you have any insights into it? </p>",
      "rawMarkdown": "So the whisper do performs better. But why most of us didn't successed with finetuning whisper? Do you have any insights into it? ",
      "votes": 5,
      "replies": [
        {
          "id": 2486470,
          "postDate": "2023-10-18T00:39:57.673Z",
          "content": "<p>I have updated the solution summary, the important part is you have to remove wrong annotations from the dataset.</p>",
          "rawMarkdown": "I have updated the solution summary, the important part is you have to remove wrong annotations from the dataset.",
          "votes": 9
        }
      ]
    },
    {
      "id": 2493597,
      "postDate": "2023-10-23T13:49:51.150Z",
      "content": "<p>it's amazing congrats for achievement</p>",
      "rawMarkdown": "it's amazing congrats for achievement\n",
      "votes": 1
    },
    {
      "id": 2491249,
      "postDate": "2023-10-21T13:05:39.633Z",
      "content": "<p>congratulations on your 1st place and thanks for sharing.</p>",
      "rawMarkdown": "congratulations on your 1st place and thanks for sharing.",
      "votes": 1
    },
    {
      "id": 2491230,
      "postDate": "2023-10-21T12:39:24.543Z",
      "content": "<p>Congratz on 1st place!!! that is crazy and thank you for sharing </p>",
      "rawMarkdown": "Congratz on 1st place!!! that is crazy and thank you for sharing \n",
      "votes": 1
    },
    {
      "id": 2490782,
      "postDate": "2023-10-21T04:21:16.427Z",
      "content": "<p>Good job to you my sir!</p>",
      "rawMarkdown": "Good job to you my sir!",
      "votes": 1
    },
    {
      "id": 2489470,
      "postDate": "2023-10-20T03:00:52.900Z",
      "content": "<p>Thanks for sharing your approach <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> </p>",
      "rawMarkdown": "Thanks for sharing your approach @tugstugi ",
      "votes": 1
    },
    {
      "id": 2488738,
      "postDate": "2023-10-19T13:08:02.387Z",
      "content": "<p>Congratulations on 1st place and thank you for sharing a great solution write-up.<br>\nCould you tell me how you downloaded YouTube videos which specifically contains Bengali speaks? Is there any criteria to which audio used?</p>",
      "rawMarkdown": "Congratulations on 1st place and thank you for sharing a great solution write-up.\nCould you tell me how you downloaded YouTube videos which specifically contains Bengali speaks? Is there any criteria to which audio used?",
      "votes": 1,
      "replies": [
        {
          "id": 2488746,
          "postDate": "2023-10-19T13:13:06.830Z",
          "content": "<p>Browse YouTube and identify Bangle channels like news, cartoon, poetry etc. Then download all videos from the channels and do pseudo labeling. Nothing fancy. I have released the data here: <a href=\"https://www.kaggle.com/datasets/tugstugi/bengali-asr-data\" target=\"_blank\">https://www.kaggle.com/datasets/tugstugi/bengali-asr-data</a> The audio file name starts with the channel name + YouTube id.</p>",
          "rawMarkdown": "Browse YouTube and identify Bangle channels like news, cartoon, poetry etc. Then download all videos from the channels and do pseudo labeling. Nothing fancy. I have released the data here: https://www.kaggle.com/datasets/tugstugi/bengali-asr-data The audio file name starts with the channel name + YouTube id.",
          "votes": 3,
          "replies": [
            {
              "id": 2489323,
              "postDate": "2023-10-19T22:21:53.947Z",
              "content": "<p>I understood. Thank you!</p>",
              "rawMarkdown": "I understood. Thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2487792,
      "postDate": "2023-10-18T18:54:06.917Z",
      "content": "<p>impressive score. <br>\n<strong>Congratulations</strong></p>",
      "rawMarkdown": "impressive score. \n**Congratulations**",
      "votes": 1
    },
    {
      "id": 2487668,
      "postDate": "2023-10-18T17:49:15.863Z",
      "content": "<p>It is very good and nice speech </p>",
      "rawMarkdown": "It is very good and nice speech ",
      "votes": 1
    },
    {
      "id": 2487089,
      "postDate": "2023-10-18T11:05:40.580Z",
      "content": "<p>Very impressive score. Congratulations on your win.<br>\nI have a question about the pseudo labels.<br>\nHow did you handle the alignment between audio and transcription? For example, a simple approach of audio separation at 30 seconds and transcription?</p>",
      "rawMarkdown": "Very impressive score. Congratulations on your win.\nI have a question about the pseudo labels.\nHow did you handle the alignment between audio and transcription? For example, a simple approach of audio separation at 30 seconds and transcription?",
      "votes": 1,
      "replies": [
        {
          "id": 2487099,
          "postDate": "2023-10-18T11:13:35.593Z",
          "content": "<p>Long youtube audios are splitted using with VAD. Reject too short or too long audio segments and let's say use only segments between 5 and 22 seconds.</p>",
          "rawMarkdown": "Long youtube audios are splitted using with VAD. Reject too short or too long audio segments and let's say use only segments between 5 and 22 seconds.",
          "votes": 2,
          "replies": [
            {
              "id": 2489385,
              "postDate": "2023-10-20T00:48:53.893Z",
              "content": "<p>Thank you. An additional question, did you use any libraries for the VAD?</p>",
              "rawMarkdown": "Thank you. An additional question, did you use any libraries for the VAD?",
              "votes": 3
            },
            {
              "id": 2492009,
              "postDate": "2023-10-22T07:49:51.883Z",
              "content": "<p>you can use any VAD library, but i have used webrtcvad.</p>",
              "rawMarkdown": "you can use any VAD library, but i have used webrtcvad.",
              "votes": 3
            },
            {
              "id": 2493148,
              "postDate": "2023-10-23T07:14:40.910Z",
              "content": "<p>Thanks for the clarification!</p>",
              "rawMarkdown": "Thanks for the clarification!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2486703,
      "postDate": "2023-10-18T05:49:15.990Z",
      "content": "<p>Unique approach in this competition, outstanding score in private and public lb. Impressive! Well done!</p>",
      "rawMarkdown": "Unique approach in this competition, outstanding score in private and public lb. Impressive! Well done!",
      "votes": 1
    },
    {
      "id": 2486503,
      "postDate": "2023-10-18T01:25:10.607Z",
      "content": "<p>A huge congratulations Tugstugi for your 1st place!</p>",
      "rawMarkdown": "A huge congratulations Tugstugi for your 1st place!",
      "votes": 1
    },
    {
      "id": 2490749,
      "postDate": "2023-10-21T02:44:30.287Z",
      "content": "<p>Congratulations on your 1st place. Thanks for sharing this great solution. I will certainlyCongratulations on your 1st place. Thanks for sharing this great solution. I will certain get a 1sst like that .</p>",
      "rawMarkdown": "Congratulations on your 1st place. Thanks for sharing this great solution. I will certainlyCongratulations on your 1st place. Thanks for sharing this great solution. I will certain get a 1sst like that .",
      "votes": 2
    },
    {
      "id": 2489757,
      "postDate": "2023-10-20T07:23:32.973Z",
      "content": "<p>amazing, congratulations! </p>",
      "rawMarkdown": "amazing, congratulations! ",
      "votes": 2
    },
    {
      "id": 2488163,
      "postDate": "2023-10-19T04:39:16.587Z",
      "content": "<p>That's so cool!</p>",
      "rawMarkdown": "That's so cool!",
      "votes": 2
    },
    {
      "id": 2486467,
      "postDate": "2023-10-18T00:21:02.333Z",
      "content": "<p>Learned a lot from Whisper performance with long train time, outperforming Wav2Vec2😃</p>",
      "rawMarkdown": "Learned a lot from Whisper performance with long train time, outperforming Wav2Vec2😃",
      "votes": 2
    },
    {
      "id": 2606989,
      "postDate": "2024-01-18T01:58:19.450Z",
      "content": "<p>Hi Tugstugi,</p>\n<p>Congratulations on your phenomenal victory.</p>\n<p>Would it be possible to replicate your training scripts on a 24GB RTX4090 by reducing batch sizes? How many GPU hours did it take to train your model 8x48GB RTX A6000 setup? What was the aggregated size of the datasets for the STT models? Thank you. </p>",
      "rawMarkdown": "Hi Tugstugi,\n\nCongratulations on your phenomenal victory.\n\nWould it be possible to replicate your training scripts on a 24GB RTX4090 by reducing batch sizes? How many GPU hours did it take to train your model 8x48GB RTX A6000 setup? What was the aggregated size of the datasets for the STT models? Thank you. "
    },
    {
      "id": 2590693,
      "postDate": "2024-01-07T11:39:22.077Z",
      "content": "<p>oh good good.</p>",
      "rawMarkdown": "oh good good."
    },
    {
      "id": 2588964,
      "postDate": "2024-01-05T22:45:59.767Z",
      "content": "<p>hey pal congratulations 🎉</p>\n<p>sorry if my question is off-topic; it is just that your use of external GPU power contradicts my understanding of code competition, has this competition changed rules after your submission? or does code competition allow the use of external cloud resources during training?</p>",
      "rawMarkdown": "hey pal congratulations 🎉\n\nsorry if my question is off-topic; it is just that your use of external GPU power contradicts my understanding of code competition, has this competition changed rules after your submission? or does code competition allow the use of external cloud resources during training?\n\n",
      "replies": [
        {
          "id": 2614873,
          "postDate": "2024-01-22T20:44:34.750Z",
          "content": "<p>in a code competition, only the inference time is limited. So you can upload the trained weights (as a dataset) beforehand, then make a blank notebook that imports those weights and performs the inference within the timeframe.</p>",
          "rawMarkdown": "in a code competition, only the inference time is limited. So you can upload the trained weights (as a dataset) beforehand, then make a blank notebook that imports those weights and performs the inference within the timeframe."
        }
      ]
    },
    {
      "id": 2499682,
      "postDate": "2023-10-26T07:12:55.287Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> inspirational news.</p>",
      "rawMarkdown": "Congratulations @tugstugi inspirational news."
    },
    {
      "id": 2499672,
      "postDate": "2023-10-26T07:09:02.807Z",
      "content": "<p>can you please share the training notebook as well?</p>",
      "rawMarkdown": "can you please share the training notebook as well?"
    },
    {
      "id": 2496590,
      "postDate": "2023-10-24T06:39:26.500Z",
      "content": "<p>amazing, congratulations!!!.  On your 1st place. </p>",
      "rawMarkdown": "amazing, congratulations!!!.  On your 1st place. \n"
    },
    {
      "id": 2486923,
      "postDate": "2023-10-18T08:41:20.410Z",
      "content": "<p>Just curious how much was the cost of the training? Was a cloud hosted GPU used? Fantastic write-up by the way!</p>",
      "rawMarkdown": "Just curious how much was the cost of the training? Was a cloud hosted GPU used? Fantastic write-up by the way!"
    },
    {
      "id": 2486480,
      "postDate": "2023-10-18T00:57:37.300Z",
      "content": "<p>Did you preprocess the dataset beforehand or did you read the audio files before each batch? </p>",
      "rawMarkdown": "Did you preprocess the dataset beforehand or did you read the audio files before each batch? ",
      "replies": [
        {
          "id": 2486482,
          "postDate": "2023-10-18T01:00:48.290Z",
          "content": "<p>We don't do any preprocessing except converting to 16khz. During the training, audios are read from the disk and applied ramdom speed/pitch etc before computing features.</p>",
          "rawMarkdown": "We don't do any preprocessing except converting to 16khz. During the training, audios are read from the disk and applied ramdom speed/pitch etc before computing features.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2486479,
      "postDate": "2023-10-18T00:56:02.317Z",
      "content": "<p>Great solution! I also noticed how inefficient the tokenizer was for whisper. How did you initialize the word embeddings after changing the tokenizer?  </p>",
      "rawMarkdown": "Great solution! I also noticed how inefficient the tokenizer was for whisper. How did you initialize the word embeddings after changing the tokenizer?  ",
      "replies": [
        {
          "id": 2486518,
          "postDate": "2023-10-18T01:38:54.157Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/nbroad\" target=\"_blank\">@nbroad</a>, The original Whisper tokenizer employs character-level tokens for low-resource languages, which can be time-consuming.  So we trained a BPE tokenizer with 12,000 tokens specifically for Bengali text. We then replaced some tokens in the Whisper tokenizer with these. We carefully replaced tokens after the 10,000th position of the original Whisper's tokens.</p>",
          "rawMarkdown": "Thanks, @nbroad, The original Whisper tokenizer employs character-level tokens for low-resource languages, which can be time-consuming.  So we trained a BPE tokenizer with 12,000 tokens specifically for Bengali text. We then replaced some tokens in the Whisper tokenizer with these. We carefully replaced tokens after the 10,000th position of the original Whisper's tokens.",
          "votes": 8,
          "replies": [
            {
              "id": 2491715,
              "postDate": "2023-10-21T23:53:09.053Z",
              "content": "<p>Wow that's a great approach and idea. For the text, do you use BPE for just text from speech examples or did you even include written text like stuff from news articles or wiki dumps? My first thought is that written data would would be different enough from speech text to cause issue, but maybe those minor difference are drowned out from the overall tokenization benefit?</p>",
              "rawMarkdown": "Wow that's a great approach and idea. For the text, do you use BPE for just text from speech examples or did you even include written text like stuff from news articles or wiki dumps? My first thought is that written data would would be different enough from speech text to cause issue, but maybe those minor difference are drowned out from the overall tokenization benefit?"
            }
          ]
        }
      ]
    },
    {
      "id": 2861401,
      "postDate": "2024-06-08T07:00:49.323Z",
      "content": "<p>Can you please tell me how you handled the noise annotation , I am facing the same issue in whisper-medium , its keep getting hallucinated whenever it hears annotation noise, can you tell me how to reduce the hallucination and handle the annotation noise , I've tried the repetition penalty , changed the beam size , changed the -log probability  , compression ratio , nothing seems to me working good</p>",
      "rawMarkdown": "Can you please tell me how you handled the noise annotation , I am facing the same issue in whisper-medium , its keep getting hallucinated whenever it hears annotation noise, can you tell me how to reduce the hallucination and handle the annotation noise , I've tried the repetition penalty , changed the beam size , changed the -log probability  , compression ratio , nothing seems to me working good",
      "isDeleted": true
    },
    {
      "id": 2487845,
      "postDate": "2023-10-18T19:46:39.940Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2486466,
      "author_name": "Mutian Hong",
      "author_url": "",
      "post_date": "2023-10-18T00:20:57.933000",
      "content": "<p>So the whisper do performs better. But why most of us didn't successed with finetuning whisper? Do you have any insights into it? </p>",
      "votes": 5,
      "replies": [
        {
          "id": 2486470,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2023-10-18T00:39:57.673000",
          "content": "<p>I have updated the solution summary, the important part is you have to remove wrong annotations from the dataset.</p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 2493597,
      "author_name": "Sameer",
      "author_url": "",
      "post_date": "2023-10-23T13:49:51.150000",
      "content": "<p>it's amazing congrats for achievement</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2491249,
      "author_name": "Ali Aghayari",
      "author_url": "",
      "post_date": "2023-10-21T13:05:39.633000",
      "content": "<p>congratulations on your 1st place and thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2491230,
      "author_name": "Max Marriott-Clarke",
      "author_url": "",
      "post_date": "2023-10-21T12:39:24.543000",
      "content": "<p>Congratz on 1st place!!! that is crazy and thank you for sharing </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2490782,
      "author_name": "Brian The Messiah 410",
      "author_url": "",
      "post_date": "2023-10-21T04:21:16.427000",
      "content": "<p>Good job to you my sir!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2489470,
      "author_name": "Indra Sonowal",
      "author_url": "",
      "post_date": "2023-10-20T03:00:52.900000",
      "content": "<p>Thanks for sharing your approach <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2488738,
      "author_name": "luddite^",
      "author_url": "",
      "post_date": "2023-10-19T13:08:02.387000",
      "content": "<p>Congratulations on 1st place and thank you for sharing a great solution write-up.<br>\nCould you tell me how you downloaded YouTube videos which specifically contains Bengali speaks? Is there any criteria to which audio used?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2488746,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2023-10-19T13:13:06.830000",
          "content": "<p>Browse YouTube and identify Bangle channels like news, cartoon, poetry etc. Then download all videos from the channels and do pseudo labeling. Nothing fancy. I have released the data here: <a href=\"https://www.kaggle.com/datasets/tugstugi/bengali-asr-data\" target=\"_blank\">https://www.kaggle.com/datasets/tugstugi/bengali-asr-data</a> The audio file name starts with the channel name + YouTube id.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2489323,
              "author_name": "luddite^",
              "author_url": "",
              "post_date": "2023-10-19T22:21:53.947000",
              "content": "<p>I understood. Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2487792,
      "author_name": "Abdallah Hesham",
      "author_url": "",
      "post_date": "2023-10-18T18:54:06.917000",
      "content": "<p>impressive score. <br>\n<strong>Congratulations</strong></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2487668,
      "author_name": "Konde Dinesh Kumar ",
      "author_url": "",
      "post_date": "2023-10-18T17:49:15.863000",
      "content": "<p>It is very good and nice speech </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2487089,
      "author_name": "ktr",
      "author_url": "",
      "post_date": "2023-10-18T11:05:40.580000",
      "content": "<p>Very impressive score. Congratulations on your win.<br>\nI have a question about the pseudo labels.<br>\nHow did you handle the alignment between audio and transcription? For example, a simple approach of audio separation at 30 seconds and transcription?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2487099,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2023-10-18T11:13:35.593000",
          "content": "<p>Long youtube audios are splitted using with VAD. Reject too short or too long audio segments and let's say use only segments between 5 and 22 seconds.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2489385,
              "author_name": "ktr",
              "author_url": "",
              "post_date": "2023-10-20T00:48:53.893000",
              "content": "<p>Thank you. An additional question, did you use any libraries for the VAD?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2492009,
              "author_name": "tugstugi",
              "author_url": "",
              "post_date": "2023-10-22T07:49:51.883000",
              "content": "<p>you can use any VAD library, but i have used webrtcvad.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2493148,
              "author_name": "ktr",
              "author_url": "",
              "post_date": "2023-10-23T07:14:40.910000",
              "content": "<p>Thanks for the clarification!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2486703,
      "author_name": "Benedikt Droste",
      "author_url": "",
      "post_date": "2023-10-18T05:49:15.990000",
      "content": "<p>Unique approach in this competition, outstanding score in private and public lb. Impressive! Well done!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2486503,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2023-10-18T01:25:10.607000",
      "content": "<p>A huge congratulations Tugstugi for your 1st place!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2490749,
      "author_name": "Tran Le Minh Nhat",
      "author_url": "",
      "post_date": "2023-10-21T02:44:30.287000",
      "content": "<p>Congratulations on your 1st place. Thanks for sharing this great solution. I will certainlyCongratulations on your 1st place. Thanks for sharing this great solution. I will certain get a 1sst like that .</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2489757,
      "author_name": "YuanXu",
      "author_url": "",
      "post_date": "2023-10-20T07:23:32.973000",
      "content": "<p>amazing, congratulations! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2488163,
      "author_name": "Vishrut Grover",
      "author_url": "",
      "post_date": "2023-10-19T04:39:16.587000",
      "content": "<p>That's so cool!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2486467,
      "author_name": "iiiiitsu",
      "author_url": "",
      "post_date": "2023-10-18T00:21:02.333000",
      "content": "<p>Learned a lot from Whisper performance with long train time, outperforming Wav2Vec2😃</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2606989,
      "author_name": "tahsin1221",
      "author_url": "",
      "post_date": "2024-01-18T01:58:19.450000",
      "content": "<p>Hi Tugstugi,</p>\n<p>Congratulations on your phenomenal victory.</p>\n<p>Would it be possible to replicate your training scripts on a 24GB RTX4090 by reducing batch sizes? How many GPU hours did it take to train your model 8x48GB RTX A6000 setup? What was the aggregated size of the datasets for the STT models? Thank you. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2590693,
      "author_name": "BIraj Ar",
      "author_url": "",
      "post_date": "2024-01-07T11:39:22.077000",
      "content": "<p>oh good good.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2588964,
      "author_name": "Mansour Hassan",
      "author_url": "",
      "post_date": "2024-01-05T22:45:59.767000",
      "content": "<p>hey pal congratulations 🎉</p>\n<p>sorry if my question is off-topic; it is just that your use of external GPU power contradicts my understanding of code competition, has this competition changed rules after your submission? or does code competition allow the use of external cloud resources during training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2614873,
          "author_name": "Curtis Chong",
          "author_url": "",
          "post_date": "2024-01-22T20:44:34.750000",
          "content": "<p>in a code competition, only the inference time is limited. So you can upload the trained weights (as a dataset) beforehand, then make a blank notebook that imports those weights and performs the inference within the timeframe.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2499682,
      "author_name": "Tariq Mahmood",
      "author_url": "",
      "post_date": "2023-10-26T07:12:55.287000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> inspirational news.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2499672,
      "author_name": "Shubham Agnihotri",
      "author_url": "",
      "post_date": "2023-10-26T07:09:02.807000",
      "content": "<p>can you please share the training notebook as well?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2496590,
      "author_name": "zisa123",
      "author_url": "",
      "post_date": "2023-10-24T06:39:26.500000",
      "content": "<p>amazing, congratulations!!!.  On your 1st place. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2486923,
      "author_name": "Wayne Lau",
      "author_url": "",
      "post_date": "2023-10-18T08:41:20.410000",
      "content": "<p>Just curious how much was the cost of the training? Was a cloud hosted GPU used? Fantastic write-up by the way!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2486480,
      "author_name": "Nicholas Broad",
      "author_url": "",
      "post_date": "2023-10-18T00:57:37.300000",
      "content": "<p>Did you preprocess the dataset beforehand or did you read the audio files before each batch? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2486482,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2023-10-18T01:00:48.290000",
          "content": "<p>We don't do any preprocessing except converting to 16khz. During the training, audios are read from the disk and applied ramdom speed/pitch etc before computing features.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2486479,
      "author_name": "Nicholas Broad",
      "author_url": "",
      "post_date": "2023-10-18T00:56:02.317000",
      "content": "<p>Great solution! I also noticed how inefficient the tokenizer was for whisper. How did you initialize the word embeddings after changing the tokenizer?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2486518,
          "author_name": "lhagiimn",
          "author_url": "",
          "post_date": "2023-10-18T01:38:54.157000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/nbroad\" target=\"_blank\">@nbroad</a>, The original Whisper tokenizer employs character-level tokens for low-resource languages, which can be time-consuming.  So we trained a BPE tokenizer with 12,000 tokens specifically for Bengali text. We then replaced some tokens in the Whisper tokenizer with these. We carefully replaced tokens after the 10,000th position of the original Whisper's tokens.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2491715,
              "author_name": "Matt S.",
              "author_url": "",
              "post_date": "2023-10-21T23:53:09.053000",
              "content": "<p>Wow that's a great approach and idea. For the text, do you use BPE for just text from speech examples or did you even include written text like stuff from news articles or wiki dumps? My first thought is that written data would would be different enough from speech text to cause issue, but maybe those minor difference are drowned out from the overall tokenization benefit?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2861401,
      "author_name": "Navaneetha Krishnan",
      "author_url": "",
      "post_date": "2024-06-08T07:00:49.323000",
      "content": "<p>Can you please tell me how you handled the noise annotation , I am facing the same issue in whisper-medium , its keep getting hallucinated whenever it hears annotation noise, can you tell me how to reduce the hallucination and handle the annotation noise , I've tried the repetition penalty , changed the beam size , changed the -log probability  , compression ratio , nothing seems to me working good</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2487845,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:46:39.940000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486461": "STT Model:\n- OpenAI whisper-medium\n- Huggingface trainer\n- Trained on 8x 48GB RTX A6000\n- bs=8 and lr=1e-5\n- Train steps 50k\n- Spectrogram dithering\n- Spectrogram time and frequency masking\n- Resampling 16khz->8khz->16khz as augmentation\n- Inference with max_length=260, num_beams=4 and chunk_length_s=20.1s\n- Libsonic based speed/pitch augmentation\n- Datasets: OpenSLR 37, OpenSLR 53, MadASR, Shrutilipi, Macro, Kathbath, GoogleTTS generated audios and pseudo labeled YouTube videos\n\n\nPunctuation Model:\n- AutoModelForTokenClassification google/muril-base-cased\n- Huggingface trainer\n- Labels: period, comma and question mark\n- bs=64, lr=2e-4 and max_seq_length=512\n- Ensemble of 4 models (using 6, 8, 11 and 12 layers of google/muril-base-cased)\n- Normalized IndicCorp v2 Bangla dataset\n\nIn my daily job, I do speech speech recognition for low resource central Asian languages. From my experience, OpenAI Whisper works really well for OOD audios and can even transribe song lyrics. The downside is it is very sensitive to the annotation noise. So fixing the annotation noise, is the most crucial part of this competition.\n\nBecause the competition dataset was not validated, the initial model was trained on OpenSLR datasets. We normalized the texts and filtered out texts containing Bengali digits. All punctuation was also removed. Additionally, we sampled 420k texts from the IndicCorp and synthesized audios using GoogleTTS, which were then used as training datasets.\n\nFollowing the training of an initial Whisper-medium model on OpenSLR and GoogleTTS, we conducted inference on MadASR, Shrutilipi, Macro, and Kathbath. We included audios with a WER of less than 15% in the next training phase. After three rounds of training, the model achieved an 8% WER on the Macro validation dataset and a public leaderboard score of approximately 0.380.\n\nSince most of the training set audios were short, we merged some short audios to create around 70k longer audios. Subsequently, we achieved a public leaderboard score of approximately 0.370.\n\nWhisper with the original tokenizer was slow on Bengali audios. Therefore, we trained a Whisper tokenizer with a 12k vocabulary on Bengali texts. With this tokenizer, we were able to perform inference with a num_beam value of up to 8 and a chunk_length_s of 20.1 seconds in less than 7 hours.\n\nIn the next step, we applied pseudo labeling to some YouTube videos, which enabled us to achieve a public leaderboard score of approximately 0.360. When we combined the predictions of four punctuation models, our public leaderboard score improved to around 0.325.\n\nBy adding more pseudo-labeled YouTube videos, our public leaderboard score further improved to 0.312 (private LB 0.372)\n\nmodel weight and inference notebook: https://www.kaggle.com/competitions/bengaliai-speech/discussion/447970\ncleaned/long/pseudo data: https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110",
    "2486466": "So the whisper do performs better. But why most of us didn't successed with finetuning whisper? Do you have any insights into it? ",
    "2493597": "it's amazing congrats for achievement\n",
    "2491249": "congratulations on your 1st place and thanks for sharing.",
    "2491230": "Congratz on 1st place!!! that is crazy and thank you for sharing \n",
    "2490782": "Good job to you my sir!",
    "2489470": "Thanks for sharing your approach @tugstugi ",
    "2488738": "Congratulations on 1st place and thank you for sharing a great solution write-up.\nCould you tell me how you downloaded YouTube videos which specifically contains Bengali speaks? Is there any criteria to which audio used?",
    "2487792": "impressive score. \n**Congratulations**",
    "2487668": "It is very good and nice speech ",
    "2487089": "Very impressive score. Congratulations on your win.\nI have a question about the pseudo labels.\nHow did you handle the alignment between audio and transcription? For example, a simple approach of audio separation at 30 seconds and transcription?",
    "2486703": "Unique approach in this competition, outstanding score in private and public lb. Impressive! Well done!",
    "2486503": "A huge congratulations Tugstugi for your 1st place!",
    "2490749": "Congratulations on your 1st place. Thanks for sharing this great solution. I will certainlyCongratulations on your 1st place. Thanks for sharing this great solution. I will certain get a 1sst like that .",
    "2489757": "amazing, congratulations! ",
    "2488163": "That's so cool!",
    "2486467": "Learned a lot from Whisper performance with long train time, outperforming Wav2Vec2😃",
    "2606989": "Hi Tugstugi,\n\nCongratulations on your phenomenal victory.\n\nWould it be possible to replicate your training scripts on a 24GB RTX4090 by reducing batch sizes? How many GPU hours did it take to train your model 8x48GB RTX A6000 setup? What was the aggregated size of the datasets for the STT models? Thank you. ",
    "2590693": "oh good good.",
    "2588964": "hey pal congratulations 🎉\n\nsorry if my question is off-topic; it is just that your use of external GPU power contradicts my understanding of code competition, has this competition changed rules after your submission? or does code competition allow the use of external cloud resources during training?\n\n",
    "2499682": "Congratulations @tugstugi inspirational news.",
    "2499672": "can you please share the training notebook as well?",
    "2496590": "amazing, congratulations!!!.  On your 1st place. \n",
    "2486923": "Just curious how much was the cost of the training? Was a cloud hosted GPU used? Fantastic write-up by the way!",
    "2486480": "Did you preprocess the dataset beforehand or did you read the audio files before each batch? ",
    "2486479": "Great solution! I also noticed how inefficient the tokenizer was for whisper. How did you initialize the word embeddings after changing the tokenizer?  ",
    "2861401": "Can you please tell me how you handled the noise annotation , I am facing the same issue in whisper-medium , its keep getting hallucinated whenever it hears annotation noise, can you tell me how to reduce the hallucination and handle the annotation noise , I've tried the repetition penalty , changed the beam size , changed the -log probability  , compression ratio , nothing seems to me working good",
    "2487845": ""
  }
}