{
  "id": 425496,
  "title": "[lb 0.481] My experimental results",
  "url": "/competitions/bengaliai-speech/discussion/425496",
  "author_name": "",
  "post_date": "2023-07-19T04:52:14.642257300Z",
  "votes": 52,
  "comment_count": 37,
  "views": 0,
  "content": "<h1>more details later,  … here is the lb0.481 recipe:</h1>\n<h1>[a]model:</h1>\n<p><a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\" target=\"_blank\">https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali</a> (CTC model only)</p>\n<h1>[b]decoder:</h1>\n<p>add a LM yourself, you can train one or use e.g. <a href=\"https://huggingface.co/shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\" target=\"_blank\">https://huggingface.co/shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice</a></p>\n<h1>[c]post-process:</h1>\n<p>normalise + dari</p>\n<pre><code>some lbscore:\n[]+[b]+[c]: 0.481\n[]+[c]: 0.520\n[]+[b]: 0.550\n[]: \n[]+[b]+[c/only normalise]: 0.490\n</code></pre>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd1010fbf802e033d78dfb6ed2043dbea%2FSelection_999(2747).png?generation=1689742315359191&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer\" target=\"_blank\">https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer</a><br>\ninitial results …. to be updated as experiments progress</p>",
  "messages": [
    {
      "id": "2350183",
      "postDate": "07/19/2023 04:52:14",
      "content": "<h1>more details later,  … here is the lb0.481 recipe:</h1>\n<h1>[a]model:</h1>\n<p><a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\" target=\"_blank\">https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali</a> (CTC model only)</p>\n<h1>[b]decoder:</h1>\n<p>add a LM yourself, you can train one or use e.g. <a href=\"https://huggingface.co/shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\" target=\"_blank\">https://huggingface.co/shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice</a></p>\n<h1>[c]post-process:</h1>\n<p>normalise + dari</p>\n<pre><code>some lbscore:\n[]+[b]+[c]: 0.481\n[]+[c]: 0.520\n[]+[b]: 0.550\n[]: \n[]+[b]+[c/only normalise]: 0.490\n</code></pre>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd1010fbf802e033d78dfb6ed2043dbea%2FSelection_999(2747).png?generation=1689742315359191&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer\" target=\"_blank\">https://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer</a><br>\ninitial results …. to be updated as experiments progress</p>",
      "rawMarkdown": "#more details later,  ... here is the lb0.481 recipe:\n#[a]model: \nhttps://huggingface.co/ai4bharat/indicwav2vec_v1_bengali (CTC model only)\n\n#[b]decoder: \nadd a LM yourself, you can train one or use e.g. https://huggingface.co/shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\n\n#[c]post-process:\nnormalise + dari\n\n```\nsome lbscore:\n[a]+[b]+[c]: 0.481\n[a]+[c]: 0.520\n[a]+[b]: 0.550\n[a]: 0.585\n[a]+[b]+[c/only normalise]: 0.490\n```\n\n----\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd1010fbf802e033d78dfb6ed2043dbea%2FSelection_999(2747).png?generation=1689742315359191&alt=media)\n\nhttps://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer\ninitial results .... to be updated as experiments progress",
      "votes": null
    },
    {
      "id": "2350184",
      "postDate": "07/19/2023 04:54:18",
      "content": "<p>some important todo list:</p>\n<ul>\n<li>speed up whsiper inference (e.g. JAX or openAI C api)</li>\n</ul>",
      "rawMarkdown": "some important todo list:\n- speed up whsiper inference (e.g. JAX or openAI C api)",
      "votes": null
    },
    {
      "id": "2350187",
      "postDate": "07/19/2023 05:00:08",
      "content": "<p>the trick to winning is really the OOD (out of distribution) private test set</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa28e51cd1564da2a042d70893ec20d51%2FSelection_999(2748).png?generation=1689742664897606&amp;alt=media\" alt=\"\"></p>\n<p>Note:</p>\n<ol>\n<li>from <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/data</a><br>\nThe full test set contains about 20 hours of speech in almost 8000 MP3 audio files, public LB is 46% of the test data</li>\n<li>from dataset paper</li>\n</ol>\n<ul>\n<li>OOD = 2681</li>\n<li>Macro Test = 4872</li>\n<li>all = 7553</li>\n</ul>",
      "rawMarkdown": "the trick to winning is really the OOD (out of distribution) private test set\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa28e51cd1564da2a042d70893ec20d51%2FSelection_999(2748).png?generation=1689742664897606&alt=media)\n\nNote:\n1. from https://www.kaggle.com/competitions/bengaliai-speech/data\nThe full test set contains about 20 hours of speech in almost 8000 MP3 audio files, public LB is 46% of the test data\n2. from dataset paper\n- OOD = 2681\n- Macro Test = 4872\n- all = 7553",
      "votes": null
    },
    {
      "id": "2350194",
      "postDate": "07/19/2023 05:08:29",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F261acc12207cdb46587c27f675c55cdb%2FSelection_999(2750).png?generation=1689743292178787&amp;alt=media\" alt=\"\"></p>\n<p>you should read the paper very carefully</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F261acc12207cdb46587c27f675c55cdb%2FSelection_999(2750).png?generation=1689743292178787&alt=media)\n\nyou should read the paper very carefully",
      "votes": null
    },
    {
      "id": "2350229",
      "postDate": "07/19/2023 05:38:34",
      "content": "<p>papers related to out-of-distribution and asr:</p>\n<ul>\n<li>to be updated</li>\n</ul>",
      "rawMarkdown": "papers related to out-of-distribution and asr:\n- to be updated",
      "votes": null
    },
    {
      "id": "2351573",
      "postDate": "07/20/2023 08:40:56",
      "content": "<p>using lanuage model is another trick<br>\n<a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali</a></p>\n<p><a href=\"https://huggingface.co/blog/wav2vec2-with-ngram\" target=\"_blank\">https://huggingface.co/blog/wav2vec2-with-ngram</a></p>\n<pre><code> language model :\n\n: .\n: .\n  gram language model trained  M sentences randomly chosen from AI4Bharat IndicCorp dataset :\n\n: .\n: .\n</code></pre>",
      "rawMarkdown": "using lanuage model is another trick\nhttps://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\n\nhttps://huggingface.co/blog/wav2vec2-with-ngram\n```\nWithout language model :\n\nWER: 0.21726385291857586\nCER: 0.04725010353701041\nWith 5 gram language model trained on 30M sentences randomly chosen from AI4Bharat IndicCorp dataset :\n\nWER: 0.15322879016421437\nCER: 0.03413696666806267\n\n```",
      "votes": null
    },
    {
      "id": "2353514",
      "postDate": "07/21/2023 17:50:58",
      "content": "<p>i think i find some OOD dataset:</p>\n<p><a href=\"https://github.com/Open-Speech-EkStep/ULCA-asr-dataset-corpus\" target=\"_blank\">https://github.com/Open-Speech-EkStep/ULCA-asr-dataset-corpus</a></p>\n<p>e.g. youtube data<br>\nEntertainment    Mirchi_Bangla_1 Labelled    Mirchi_Bangla_1 54 (hrs)<br>\nEntertainment    Mirchi_Bangla_2 Labelled    Mirchi_Bangla_2 <br>\nEntertainment    Mirchi_Bangla_3 Labelled    Mirchi_Bangla_3 <br>\nGeneral    CTVN_AKD_PLUS_1 Labelled    CTVN_AKD_PLUS_1 21.7 (hrs)<br>\nGeneral    CTVN_AKD_PLUS_2 Labelled    CTVN_AKD_PLUS_2 </p>\n<p>however, the data are not available ….<br>\nbut the pretrain model are here<br>\n<a href=\"https://github.com/Open-Speech-EkStep/vakyansh-models\" target=\"_blank\">https://github.com/Open-Speech-EkStep/vakyansh-models</a><br>\nVakyansh-Conformer-SSL</p>\n<p>\"This model was pre-trained using Nemo toolkit with 34,000 hours unlabeled audio in 39 Indian languages. This includes 15,000 hours of news recordings available on the internet, 10,000 hours of YouTube audios and other audio data.\"</p>\n<p>model config yaml:<br>\ntrain_ds:<br>\n  manifest_filepath: /mnt/lustre/megh/indic-ssl/manifest/crisil_noa_nptel_yt_indic_ssl_train.json</p>\n<p>i am guessing<br>\nnoa : news on air<br>\nyt: youtube</p>\n<p>paper: <br>\nVakyansh: ASR Toolkit for Low Resource Indic languages<br>\n<a href=\"https://arxiv.org/pdf/2203.16512.pdf\" target=\"_blank\">https://arxiv.org/pdf/2203.16512.pdf</a></p>",
      "rawMarkdown": "i think i find some OOD dataset:\n\nhttps://github.com/Open-Speech-EkStep/ULCA-asr-dataset-corpus\n\ne.g. youtube data\nEntertainment\tMirchi_Bangla_1\tLabelled\tMirchi_Bangla_1\t54 (hrs)\nEntertainment\tMirchi_Bangla_2\tLabelled\tMirchi_Bangla_2\t\nEntertainment\tMirchi_Bangla_3\tLabelled\tMirchi_Bangla_3\t\nGeneral\tCTVN_AKD_PLUS_1\tLabelled\tCTVN_AKD_PLUS_1\t21.7 (hrs)\nGeneral\tCTVN_AKD_PLUS_2\tLabelled\tCTVN_AKD_PLUS_2\t\n\n\nhowever, the data are not available ....\nbut the pretrain model are here\nhttps://github.com/Open-Speech-EkStep/vakyansh-models\nVakyansh-Conformer-SSL\n\n\"This model was pre-trained using Nemo toolkit with 34,000 hours unlabeled audio in 39 Indian languages. This includes 15,000 hours of news recordings available on the internet, 10,000 hours of YouTube audios and other audio data.\"\n\nmodel config yaml:\ntrain_ds:\n  manifest_filepath: /mnt/lustre/megh/indic-ssl/manifest/crisil_noa_nptel_yt_indic_ssl_train.json\n\ni am guessing\nnoa : news on air\nyt: youtube\n\n\npaper: \nVakyansh: ASR Toolkit for Low Resource Indic languages\nhttps://arxiv.org/pdf/2203.16512.pdf",
      "votes": null
    },
    {
      "id": "2353528",
      "postDate": "07/21/2023 18:20:49",
      "content": "<p>sampling rate as TTA?</p>\n<pre><code>sampling_rate=\n\n :   one\n    d = valid_df\n    mp3_file = f\n\n     = (mp3_file) \n    with (mp3_file, ) as f:\n        bpayload = f()\n     = ( bpayload , sampling_rate)\n     = (a) \n    (, p)\n    (, d.sentence)\n</code></pre>\n<pre><code> Hz\npredict ও বলেছে আপনার টিকাপ\ntruth   ও বলেছে আপনার ঠিকানা!\n\n Hz\npredict ওবলেছে আপনার টিকা\ntruth   ও বলেছে আপনার ঠিকানা!\n\n Hz\npredict এগুন শ অ\ntruth   ও বলেছে আপনার ঠিকানা!\n\n Hz\npredict অবিচয় আনাটি\ntruth   ও বলেছে আপনার ঠিকানা!\n\n Hz\npredict ও বলেছে আপনার টিকা\ntruth   ও বলেছে আপনার ঠিকানা!\n</code></pre>",
      "rawMarkdown": "sampling rate as TTA?\n\n```\nsampling_rate=16000\n\nif 1:  #test one\n\td = valid_df.iloc[0]\n\tmp3_file = f'{mp3_dir}/{d[\"id\"]}.mp3'\n\n\t#p = pipe(mp3_file)['text'] \n\twith open(mp3_file, 'rb') as f:\n\t\tbpayload = f.read()\n\ta = ffmpeg_read( bpayload , sampling_rate)\n\tp = pipe(a)['text'] \n\tprint('predict', p)\n\tprint('truth  ', d.sentence)\n\n```\n\n```\n18000 Hz\npredict ও বলেছে আপনার টিকাপ\ntruth   ও বলেছে আপনার ঠিকানা!\n\n16000 Hz\npredict ওবলেছে আপনার টিকা\ntruth   ও বলেছে আপনার ঠিকানা!\n\n32000 Hz\npredict এগুন শ অ\ntruth   ও বলেছে আপনার ঠিকানা!\n\n14000 Hz\npredict অবিচয় আনাটি\ntruth   ও বলেছে আপনার ঠিকানা!\n\n20000 Hz\npredict ও বলেছে আপনার টিকা\ntruth   ও বলেছে আপনার ঠিকানা!\n```",
      "votes": null
    },
    {
      "id": "2353969",
      "postDate": "07/22/2023 06:43:19",
      "content": "<p>this youtube channel has many tutorials:</p>\n<p><a href=\"https://www.youtube.com/@ai4bharat/videos\" target=\"_blank\">https://www.youtube.com/@ai4bharat/videos</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e96ec049ef64542d26eac80c7249bc4%2FSelection_999(2766).png?generation=1690008184783299&amp;alt=media\" alt=\"\"></p>\n<p>e.g<br>\n<a href=\"https://www.youtube.com/watch?v=iSyipKKgleo\" target=\"_blank\">https://www.youtube.com/watch?v=iSyipKKgleo</a></p>",
      "rawMarkdown": "this youtube channel has many tutorials:\n\nhttps://www.youtube.com/@ai4bharat/videos\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e96ec049ef64542d26eac80c7249bc4%2FSelection_999(2766).png?generation=1690008184783299&alt=media)\n\ne.g\nhttps://www.youtube.com/watch?v=iSyipKKgleo",
      "votes": null
    },
    {
      "id": "2355813",
      "postDate": "07/23/2023 16:15:41",
      "content": "<p>visualise distance edit</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa32c0720b98a08aedccc73603163a0bf%2FSelection_999(2767).png?generation=1690128936873031&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://github.com/ukiuki-satoshi/visedit\" target=\"_blank\">https://github.com/ukiuki-satoshi/visedit</a></p>",
      "rawMarkdown": "visualise distance edit\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa32c0720b98a08aedccc73603163a0bf%2FSelection_999(2767).png?generation=1690128936873031&alt=media)\n\nhttps://github.com/ukiuki-satoshi/visedit",
      "votes": null
    },
    {
      "id": "2356341",
      "postDate": "07/24/2023 05:36:49",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd719f707bd940500540adeb99191780d%2FSelection_999(2811).png?generation=1690403372321495&amp;alt=media\" alt=\"\"></p>\n<p><br>\n<br>\n</p>\n<p><br>\n</p>\n<p></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd719f707bd940500540adeb99191780d%2FSelection_999(2811).png?generation=1690403372321495&alt=media)\n\n\n~~should we be using kaggle dataset \"as it is\" at all?~~\n~~i note that the baseline conformer-CTC has good results on train.csv valid split (clean label). But it only has LB0.68.~~\n~~On the other hands, public huggingface/ai4bharat/indicwav2vec_v1_bengali only trained on their datase ~~~~(radio boradcast and youtube) has less performance gap:  valid split  0.39, LB0.52.~~\n\n~~This makes me suspect kaggle data is not good on hidden OOD public dataset.~~\n~~In my experiments abovem, the better results you get for train.csv valid split (or train split), you get worse ~~~~results for publuc LB.~~\n\n~~Kagglers may want to confirm this observation~~",
      "votes": null
    },
    {
      "id": "2356475",
      "postDate": "07/24/2023 08:14:01",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F11ad141e94c7a32b3148618428ab7b11%2FSelection_999(2797).png?generation=1690186435573149&amp;alt=media\" alt=\"\"></p>\n<p>OOD results</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F11ad141e94c7a32b3148618428ab7b11%2FSelection_999(2797).png?generation=1690186435573149&alt=media)\n\nOOD results",
      "votes": null
    },
    {
      "id": "2356515",
      "postDate": "07/24/2023 08:56:36",
      "content": "<p>\" We first show an interesting finding that while Whisper is very robust against real-world background sounds (e.g., music), its audio representation is actually not noise-invariant, but is instead highly correlated to non-speech sounds, indicating that Whisper recognizes speech conditioned on the noise type\"</p>\n<p><a href=\"https://arxiv.org/pdf/2307.03183.pdf\" target=\"_blank\">https://arxiv.org/pdf/2307.03183.pdf</a><br>\nWhisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers</p>",
      "rawMarkdown": "\" We first show an interesting finding that while Whisper is very robust against real-world background sounds (e.g., music), its audio representation is actually not noise-invariant, but is instead highly correlated to non-speech sounds, indicating that Whisper recognizes speech conditioned on the noise type\"\n\nhttps://arxiv.org/pdf/2307.03183.pdf\nWhisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers",
      "votes": null
    },
    {
      "id": "2356526",
      "postDate": "07/24/2023 09:01:11",
      "content": "<p>End-to-end Music-mixed Speech Recognition<br>\n<a href=\"https://arxiv.org/pdf/2008.12048.pdf\" target=\"_blank\">https://arxiv.org/pdf/2008.12048.pdf</a></p>",
      "rawMarkdown": "End-to-end Music-mixed Speech Recognition\nhttps://arxiv.org/pdf/2008.12048.pdf",
      "votes": null
    },
    {
      "id": "2356542",
      "postDate": "07/24/2023 09:08:24",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3fbba9c75392b0f594274f017d7bbeec%2FSelection_999(2798).png?generation=1690189696857157&amp;alt=media\" alt=\"\"></p>\n<p>Wav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3fbba9c75392b0f594274f017d7bbeec%2FSelection_999(2798).png?generation=1690189696857157&alt=media)\n\nWav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition",
      "votes": null
    },
    {
      "id": "2356736",
      "postDate": "07/24/2023 10:52:08",
      "content": "<p><a href=\"https://sites.google.com/iitdh.ac.in/vssasr2021/resources\" target=\"_blank\">https://sites.google.com/iitdh.ac.in/vssasr2021/resources</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F55dbc9f676130c29ea69dd771132459e%2FSelection_999(2799).png?generation=1690195926421740&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "https://sites.google.com/iitdh.ac.in/vssasr2021/resources\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F55dbc9f676130c29ea69dd771132459e%2FSelection_999(2799).png?generation=1690195926421740&alt=media)",
      "votes": null
    },
    {
      "id": "2358541",
      "postDate": "07/25/2023 15:23:32",
      "content": "<p>CTC bert?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F40e8bbfbd5da3f9ddd2af5e56f66ed62%2FSelection_999(2803).png?generation=1690298599425663&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F231e9aa6db041b8f6fe453217fbedd6f%2FSelection_999(2804).png?generation=1690298610687049&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "CTC bert?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F40e8bbfbd5da3f9ddd2af5e56f66ed62%2FSelection_999(2803).png?generation=1690298599425663&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F231e9aa6db041b8f6fe453217fbedd6f%2FSelection_999(2804).png?generation=1690298610687049&alt=media)",
      "votes": null
    },
    {
      "id": "2359089",
      "postDate": "07/26/2023 01:41:46",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  How did you get the ground truth for the test_example  to calculate the wer ? It's not available in the Data section </p>\n<p>EDIT: sorry, I've found the source <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932</a></p>",
      "rawMarkdown": "hengck23  How did you get the ground truth for the test_example  to calculate the wer ? It's not available in the Data section \n\nEDIT: sorry, I've found the source https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932",
      "votes": null
    },
    {
      "id": "2359216",
      "postDate": "07/26/2023 04:31:33",
      "content": "<p>massive external data:<br>\n<a href=\"https://github.com/AI4Bharat/vistaar#download-training-datasets-and-benchmarks\" target=\"_blank\">https://github.com/AI4Bharat/vistaar#download-training-datasets-and-benchmarks</a></p>",
      "rawMarkdown": "massive external data:\nhttps://github.com/AI4Bharat/vistaar#download-training-datasets-and-benchmarks",
      "votes": null
    },
    {
      "id": "2360593",
      "postDate": "07/26/2023 20:31:06",
      "content": "<p>i made a mistake. i forget about normalisation and others in post-processing</p>",
      "rawMarkdown": "i made a mistake. i forget about normalisation and others in post-processing",
      "votes": null
    },
    {
      "id": "2360720",
      "postDate": "07/27/2023 00:33:30",
      "content": "<p>paper on effects  of in and out domain :<br>\nROBUST WAV2VEC 2.0: ANALYZING DOMAIN SHIFT IN SELF-SUPERVISED PRE-TRAINING<br>\n<a href=\"https://arxiv.org/pdf/2104.01027.pdf\" target=\"_blank\">https://arxiv.org/pdf/2104.01027.pdf</a></p>\n<p>\"In this paper, we explore more general setups where the domain of the unlabeled data for pre-training data differs from the domain of the labeled data for fine-tuning, which in turn may differ from the test data domain\"</p>",
      "rawMarkdown": "paper on effects  of in and out domain :\nROBUST WAV2VEC 2.0: ANALYZING DOMAIN SHIFT IN SELF-SUPERVISED PRE-TRAINING\nhttps://arxiv.org/pdf/2104.01027.pdf\n\n\"In this paper, we explore more general setups where the domain of the unlabeled data for pre-training data differs from the domain of the labeled data for fine-tuning, which in turn may differ from the test data domain\"",
      "votes": null
    },
    {
      "id": "2360878",
      "postDate": "07/27/2023 04:36:08",
      "content": "<p>yet another external data<br>\n<a href=\"https://globalrecordings.net\" target=\"_blank\">https://globalrecordings.net</a></p>",
      "rawMarkdown": "yet another external data\nhttps://globalrecordings.net",
      "votes": null
    },
    {
      "id": "2365899",
      "postDate": "07/30/2023 14:44:37",
      "content": "<p>i wonder if the each sentence in the test is unique?<br>\nif not, this is a leak?<br>\n(e.g. one may collect speech data by asking different people <strong>to speak the same sentence</strong>)</p>",
      "rawMarkdown": "i wonder if the each sentence in the test is unique?\nif not, this is a leak?\n(e.g. one may collect speech data by asking different people **to speak the same sentence**)",
      "votes": null
    },
    {
      "id": "2367644",
      "postDate": "07/31/2023 17:20:34",
      "content": "<p>Thanks for posting all the insights. May I ask how long does it take you to run whsiper for leaderboard submssion?</p>",
      "rawMarkdown": "Thanks for posting all the insights. May I ask how long does it take you to run whsiper for leaderboard submssion?",
      "votes": null
    },
    {
      "id": "2368194",
      "postDate": "08/01/2023 04:04:52",
      "content": "<p>whisper takes in 30sec audio chunk. so i did a probe on audio clip duration of hidden test data:</p>\n<pre><code>valid_df = pd()\nmp3_dir = f\n\nduration=\n t,d  valid_df():\n    audio_file = f\n\n    sampling_rate = _000\n    with (audio_file, ) as f:\n        bpayload = f()\n     = (bpayload, sampling_rate)\n    dur = (a)/sampling_rate\n    (dur)\n    duration(dur)\n\nmax_duration=np(duration)\n\n\n</code></pre>\n<p>submission failed for : max_duration&lt;=30 sec<br>\nsubmission passed for : max_duration&lt;=32 sec</p>",
      "rawMarkdown": "whisper takes in 30sec audio chunk. so i did a probe on audio clip duration of hidden test data:\n```\nvalid_df = pd.read_csv('/kaggle/input/bengaliai-speech/sample_submission.csv')\nmp3_dir = f'/kaggle/input/bengaliai-speech/test_mp3s'\n\nduration=[]\nfor t,d in valid_df.iterrows():\n    audio_file = f'{mp3_dir}/{d[\"id\"]}.mp3'\n\n    sampling_rate = 16_000\n    with open(audio_file, 'rb') as f:\n        bpayload = f.read()\n    a = ffmpeg_read(bpayload, sampling_rate)\n    dur = len(a)/sampling_rate\n    print(dur)\n    duration.append(dur)\n\nmax_duration=np.max(duration)\nprint('max_duration',max_duration)\nassert(max_duration<=DURATION)\n\n```\n\nsubmission failed for : max_duration<=30 sec\nsubmission passed for : max_duration<=32 sec",
      "votes": null
    },
    {
      "id": "2368354",
      "postDate": "08/01/2023 06:08:38",
      "content": "<p>Thank you for the info … so much research on the topic great …all the best</p>",
      "rawMarkdown": "Thank you for the info ... so much research on the topic great ...all the best",
      "votes": null
    },
    {
      "id": "2369358",
      "postDate": "08/01/2023 17:34:50",
      "content": "<p>another dataset<br>\nDataset - RESPIN<br>\n<a href=\"https://slt2022.org/projects/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri.pdf\" target=\"_blank\">https://slt2022.org/projects/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri.pdf</a></p>\n<p><a href=\"https://sites.google.com/view/respinasrchallenge2023/home\" target=\"_blank\">https://sites.google.com/view/respinasrchallenge2023/home</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6ff8e18a7c97388e287e41fd5bcf58bb%2FSelection_999(2829).png?generation=1690911460374350&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "another dataset\nDataset - RESPIN\nhttps://slt2022.org/projects/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri.pdf\n\nhttps://sites.google.com/view/respinasrchallenge2023/home\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6ff8e18a7c97388e287e41fd5bcf58bb%2FSelection_999(2829).png?generation=1690911460374350&alt=media)",
      "votes": null
    },
    {
      "id": "2369394",
      "postDate": "08/01/2023 17:57:23",
      "content": "<p>code switch dataset:<br>\n<a href=\"https://github.com/navana-tech/baseline_recipe_is21s_indic_asr_challenge\" target=\"_blank\">https://github.com/navana-tech/baseline_recipe_is21s_indic_asr_challenge</a></p>",
      "rawMarkdown": "code switch dataset:\nhttps://github.com/navana-tech/baseline_recipe_is21s_indic_asr_challenge",
      "votes": null
    },
    {
      "id": "2370469",
      "postDate": "08/02/2023 12:43:43",
      "content": "<p>Hi, is indicwav2vec_v1_bengali simply transformed from the fairseq model on indicwav2vec github repo?</p>",
      "rawMarkdown": "Hi, is indicwav2vec_v1_bengali simply transformed from the fairseq model on indicwav2vec github repo?",
      "votes": null
    },
    {
      "id": "2370496",
      "postDate": "08/02/2023 13:09:32",
      "content": "<p>i have submitted the version from hugging face.</p>\n<p>i may choose to use the github version later</p>",
      "rawMarkdown": "i have submitted the version from hugging face.\n\ni may choose to use the github version later",
      "votes": null
    },
    {
      "id": "2370590",
      "postDate": "08/02/2023 14:04:55",
      "content": "<p><br>\nThey are the same model</p>",
      "rawMarkdown": "~~I try the fairseq model a little, it's fast but the wer is high. Maybe they are not the same model, or I did something wrong. ~~\nThey are the same model",
      "votes": null
    },
    {
      "id": "2372470",
      "postDate": "08/03/2023 17:32:09",
      "content": "<p>In my experiment, directly binding IndicWav2Vec2 to Yellowking's decoder has led to some problems, that is the vocabularie sizes have not matched The former is 87 and the latter is 112. How can I solve it? Thanks!</p>",
      "rawMarkdown": "In my experiment, directly binding IndicWav2Vec2 to Yellowking's decoder has led to some problems, that is the vocabularie sizes have not matched The former is 87 and the latter is 112. How can I solve it? Thanks!",
      "votes": null
    },
    {
      "id": "2372494",
      "postDate": "08/03/2023 17:58:12",
      "content": "<p>Is this OK by just setting <code>ignore_mismatched_sizes=True</code></p>",
      "rawMarkdown": "Is this OK by just setting `ignore_mismatched_sizes=True`",
      "votes": null
    },
    {
      "id": "2379732",
      "postDate": "08/08/2023 09:07:32",
      "content": "<p>There may be some problems, because the vocab.json of <code>indicwav2vec_v1_bengali</code> and other models (bengali, with LM) are not the same. I tried rewriting the alphabet.json and the model infers some nonsense.</p>",
      "rawMarkdown": "There may be some problems, because the vocab.json of `indicwav2vec_v1_bengali` and other models (bengali, with LM) are not the same. I tried rewriting the alphabet.json and the model infers some nonsense.",
      "votes": null
    },
    {
      "id": "2381841",
      "postDate": "08/09/2023 12:20:28",
      "content": "<p>this is my code</p>\n<pre><code> transformers  pipeline\n transformers  Wav2Vec2CTCTokenizer, Wav2Vec2FeatureExtractor, Wav2Vec2Processor,Wav2Vec2ProcessorWithLM\n transformers  Wav2Vec2ForCTC\n\n\npretrain_model = \\\n    \nvocab_dir = \\\n    \narpa_file= \\\n     \n   \n\n\n\nsampling_rate=\ntokenizer = Wav2Vec2CTCTokenizer.from_pretrained(\n    vocab_dir, \n    unk_token=,\n    pad_token=,\n    word_delimiter_token=, \n    bos_token=,\n    eos_token=,\n)\n\n\n\nfeature_extractor = Wav2Vec2FeatureExtractor(\n    feature_size=,\n    sampling_rate=sampling_rate,\n    padding_value=,\n    padding_side=,\n    do_normalize=,\n    return_attention_mask=,\n)\n\n :\n     pyctcdecode  BeamSearchDecoderCTC\n     pyctcdecode  build_ctcdecoder\n\n    vocab_dict = tokenizer.get_vocab()\n\n    vocab_dict = {k: v  k, v  (vocab_dict.items(), key= item: item[])}\n    decoder = build_ctcdecoder(\n        labels=(vocab_dict.keys()),\n        kenlm_model_path=arpa_file,\n    )\n\nprocessor = Wav2Vec2ProcessorWithLM(\n    feature_extractor=feature_extractor,\n    tokenizer=tokenizer,\n    decoder=decoder,\n)\n\n\n (nn.Module):\n     ():\n        ().__init__()\n        self.output_type = []\n        self.model = Wav2Vec2ForCTC.from_pretrained(\n            pretrain_model,\n            ignore_mismatched_sizes=,\n            attention_dropout= ,\n            hidden_dropout=  ,\n            feat_proj_dropout=  ,\n            mask_time_prob= ,\n            layerdrop= ,\n            ctc_loss_reduction=,\n            pad_token_id=processor.tokenizer.pad_token_id,\n            vocab_size=(processor.tokenizer),\n        )\n        \n     ():\n        B,L = batch[].shape\n\n        \n        out = self.model (\n            batch[],\n            batch[],\n            labels = batch.get(,),\n        )\n\n\n        output = {}\n           self.output_type:\n            output[] = out.logits\n            \n            \n            output[] = postprocess_lm_to_text(out, batch[])\n            \n\n           self.output_type:\n            output[] = out.loss\n\n         output\n\n\n ():\n    B = (out.logits)\n    \n    logit = out.logits\n    token = logit.argmax(dim=-)\n    text = []\n     b  (B):\n        t = token[b][:length[b]//]\n        t = tokenizer.decode(t, skip_special_tokens=)\n        text.append(t)\n     text\n\n ():\n    B = (out.logits)\n    \n    logit = out.logits.data.cpu().numpy()\n    text = []\n     b  (B):\n        l = logit[b]\n        beam = decoder.decode_beams(l, beam_width=)\n        t = beam[][]\n        text.append(t)\n     text\n</code></pre>",
      "rawMarkdown": "this is my code\n\n```\nfrom transformers import pipeline\nfrom transformers import Wav2Vec2CTCTokenizer, Wav2Vec2FeatureExtractor, Wav2Vec2Processor,Wav2Vec2ProcessorWithLM\nfrom transformers import Wav2Vec2ForCTC\n\n\npretrain_model = \\\n    '/kaggle/input/ai4bharat-indicwav2vec-v1-bengali'\nvocab_dir = \\\n    '/kaggle/input/ai4bharat-indicwav2vec-v1-bengali'\narpa_file= \\\n    '/kaggle/input/my-weight-bengali-asr-01/my-lm-5gram.arpa' #or binary file\n   #'/kaggle/input/my-weight-bengali-asr-01/language_model/commonvoice-bn.5.arpa'\n\n##########################################################################\n#model    \nsampling_rate=16_000\ntokenizer = Wav2Vec2CTCTokenizer.from_pretrained(\n    vocab_dir, #'my_tokenizer',\n    unk_token='<unk>',\n    pad_token='<pad>',\n    word_delimiter_token='|', ##<todo>???\n    bos_token='<s>',\n    eos_token='</s>',\n)\n'''\ntokenizer.convert_tokens_to_ids(['|','<s>', '</s>', '<unk>', '<pad>'])\nOut[1]: [62, 1, 2, 3, 0]\n'''\n\n# just for padding, etc : audio to pad_audio, mask\nfeature_extractor = Wav2Vec2FeatureExtractor(\n    feature_size=1,\n    sampling_rate=sampling_rate,\n    padding_value=0.0,\n    padding_side='right',\n    do_normalize=True,\n    return_attention_mask=True,\n)\n\nif 1:\n\tfrom pyctcdecode import BeamSearchDecoderCTC\n\tfrom pyctcdecode import build_ctcdecoder\n\n\tvocab_dict = tokenizer.get_vocab()\n\n\tvocab_dict = {k: v for k, v in sorted(vocab_dict.items(), key=lambda item: item[1])}\n\tdecoder = build_ctcdecoder(\n\t    labels=list(vocab_dict.keys()),\n\t    kenlm_model_path=arpa_file,\n\t)\n\nprocessor = Wav2Vec2ProcessorWithLM(\n    feature_extractor=feature_extractor,\n    tokenizer=tokenizer,\n\tdecoder=decoder,\n)\n\n\nclass Net(nn.Module):\n\tdef __init__(self, ):\n\t\tsuper().__init__()\n\t\tself.output_type = ['inference']\n\t\tself.model = Wav2Vec2ForCTC.from_pretrained(\n\t\t    pretrain_model,\n\t\t    ignore_mismatched_sizes=False,\n\t\t    attention_dropout= 0,#0.1,\n\t\t    hidden_dropout=  0,#0.1,\n\t\t    feat_proj_dropout=  0,#0.1,\n\t\t    mask_time_prob= 0,#0.05,\n\t\t    layerdrop= 0,#0.1,\n\t\t    ctc_loss_reduction='mean',\n\t\t    pad_token_id=processor.tokenizer.pad_token_id,\n\t\t    vocab_size=len(processor.tokenizer),\n\t\t)\n\t\t#self.model.config.ctc_zero_infinity = True\n\tdef forward(self, batch):\n\t\tB,L = batch['input_values'].shape\n\n\t\t#class Wav2Vec2ForCTC(Wav2Vec2PreTrainedModel):\n\t\tout = self.model (\n\t\t\tbatch['input_values'],\n\t\t\tbatch['attention_mask'],\n\t\t\tlabels = batch.get('labels',None),\n\t\t)#CausalLMOutput\n\n\n\t\toutput = {}\n\t\tif 'inference' in self.output_type:\n\t\t\toutput['logit'] = out.logits\n\t\t\t#output['hidden_state'] = out.hidden_states\n\t\t\t#'attention' : out.attentions,\n\t\t\toutput['text'] = postprocess_lm_to_text(out, batch['length'])\n\t\t\t#output['text'] = postprocess_ctc_to_text(out, batch['length'])\n\n\t\tif 'loss' in self.output_type:\n\t\t\toutput['ctc_loss'] = out.loss\n\n\t\treturn output\n\n\ndef postprocess_ctc_to_text(out, length):\n\tB = len(out.logits)\n\t# AutomaticSpeechRecognitionPipeline(ChunkPipeline)\n\tlogit = out.logits\n\ttoken = logit.argmax(dim=-1)\n\ttext = []\n\tfor b in range(B):\n\t\tt = token[b][:length[b]//320]\n\t\tt = tokenizer.decode(t, skip_special_tokens=False)\n\t\ttext.append(t)\n\treturn text\n \ndef postprocess_lm_to_text(out, length):\n\tB = len(out.logits)\n\t# AutomaticSpeechRecognitionPipeline(ChunkPipeline)\n\tlogit = out.logits.data.cpu().numpy()\n\ttext = []\n\tfor b in range(B):\n\t\tl = logit[b]#[:length[b]//320]\n\t\tbeam = decoder.decode_beams(l, beam_width=512)\n\t\tt = beam[0][0]\n\t\ttext.append(t)\n\treturn text\n```",
      "votes": null
    },
    {
      "id": "2381844",
      "postDate": "08/09/2023 12:21:31",
      "content": "<p>LM model and wave2vec models are independent.<br>\ni have been mixing different wav2vec and LM from different sources</p>\n<p>i also tried  Flashlight decoder, and it would also work</p>",
      "rawMarkdown": "LM model and wave2vec models are independent.\ni have been mixing different wav2vec and LM from different sources\n\ni also tried  Flashlight decoder, and it would also work",
      "votes": null
    },
    {
      "id": "2381973",
      "postDate": "08/09/2023 13:44:30",
      "content": "<p>Thanks a lot for your detailed help! I'm really appreciating that</p>",
      "rawMarkdown": "Thanks a lot for your detailed help! I'm really appreciating that",
      "votes": null
    },
    {
      "id": "2438832",
      "postDate": "09/14/2023 14:10:06",
      "content": "<p>Thank you </p>",
      "rawMarkdown": "Thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2350184,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/19/2023 04:54:18",
      "content": "<p>some important todo list:</p>\n<ul>\n<li>speed up whsiper inference (e.g. JAX or openAI C api)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2367644,
          "author_name": "zhenlanwang",
          "author_url": "",
          "post_date": "07/31/2023 17:20:34",
          "content": "<p>Thanks for posting all the insights. May I ask how long does it take you to run whsiper for leaderboard submssion?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2350187,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/19/2023 05:00:08",
      "content": "<p>the trick to winning is really the OOD (out of distribution) private test set</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa28e51cd1564da2a042d70893ec20d51%2FSelection_999(2748).png?generation=1689742664897606&amp;alt=media\" alt=\"\"></p>\n<p>Note:</p>\n<ol>\n<li>from <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/data</a><br>\nThe full test set contains about 20 hours of speech in almost 8000 MP3 audio files, public LB is 46% of the test data</li>\n<li>from dataset paper</li>\n</ol>\n<ul>\n<li>OOD = 2681</li>\n<li>Macro Test = 4872</li>\n<li>all = 7553</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2350194,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/19/2023 05:08:29",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F261acc12207cdb46587c27f675c55cdb%2FSelection_999(2750).png?generation=1689743292178787&amp;alt=media\" alt=\"\"></p>\n<p>you should read the paper very carefully</p>",
          "votes": null,
          "replies": [
            {
              "id": 2351573,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "07/20/2023 08:40:56",
              "content": "<p>using lanuage model is another trick<br>\n<a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali</a></p>\n<p><a href=\"https://huggingface.co/blog/wav2vec2-with-ngram\" target=\"_blank\">https://huggingface.co/blog/wav2vec2-with-ngram</a></p>\n<pre><code> language model :\n\n: .\n: .\n  gram language model trained  M sentences randomly chosen from AI4Bharat IndicCorp dataset :\n\n: .\n: .\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2350229,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/19/2023 05:38:34",
      "content": "<p>papers related to out-of-distribution and asr:</p>\n<ul>\n<li>to be updated</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2353514,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/21/2023 17:50:58",
      "content": "<p>i think i find some OOD dataset:</p>\n<p><a href=\"https://github.com/Open-Speech-EkStep/ULCA-asr-dataset-corpus\" target=\"_blank\">https://github.com/Open-Speech-EkStep/ULCA-asr-dataset-corpus</a></p>\n<p>e.g. youtube data<br>\nEntertainment    Mirchi_Bangla_1 Labelled    Mirchi_Bangla_1 54 (hrs)<br>\nEntertainment    Mirchi_Bangla_2 Labelled    Mirchi_Bangla_2 <br>\nEntertainment    Mirchi_Bangla_3 Labelled    Mirchi_Bangla_3 <br>\nGeneral    CTVN_AKD_PLUS_1 Labelled    CTVN_AKD_PLUS_1 21.7 (hrs)<br>\nGeneral    CTVN_AKD_PLUS_2 Labelled    CTVN_AKD_PLUS_2 </p>\n<p>however, the data are not available ….<br>\nbut the pretrain model are here<br>\n<a href=\"https://github.com/Open-Speech-EkStep/vakyansh-models\" target=\"_blank\">https://github.com/Open-Speech-EkStep/vakyansh-models</a><br>\nVakyansh-Conformer-SSL</p>\n<p>\"This model was pre-trained using Nemo toolkit with 34,000 hours unlabeled audio in 39 Indian languages. This includes 15,000 hours of news recordings available on the internet, 10,000 hours of YouTube audios and other audio data.\"</p>\n<p>model config yaml:<br>\ntrain_ds:<br>\n  manifest_filepath: /mnt/lustre/megh/indic-ssl/manifest/crisil_noa_nptel_yt_indic_ssl_train.json</p>\n<p>i am guessing<br>\nnoa : news on air<br>\nyt: youtube</p>\n<p>paper: <br>\nVakyansh: ASR Toolkit for Low Resource Indic languages<br>\n<a href=\"https://arxiv.org/pdf/2203.16512.pdf\" target=\"_blank\">https://arxiv.org/pdf/2203.16512.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2353528,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/21/2023 18:20:49",
      "content": "<p>sampling rate as TTA?</p>\n<pre><code>sampling_rate=\n\n :   one\n    d = valid_df\n    mp3_file = f\n\n     = (mp3_file) \n    with (mp3_file, ) as f:\n        bpayload = f()\n     = ( bpayload , sampling_rate)\n     = (a) \n    (, p)\n    (, d.sentence)\n</code></pre>\n<pre><code> Hz\npredict ও বলেছে আপনার টিকাপ\ntruth   ও বলেছে আপনার ঠিকানা!\n\n Hz\npredict ওবলেছে আপনার টিকা\ntruth   ও বলেছে আপনার ঠিকানা!\n\n Hz\npredict এগুন শ অ\ntruth   ও বলেছে আপনার ঠিকানা!\n\n Hz\npredict অবিচয় আনাটি\ntruth   ও বলেছে আপনার ঠিকানা!\n\n Hz\npredict ও বলেছে আপনার টিকা\ntruth   ও বলেছে আপনার ঠিকানা!\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2353969,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/22/2023 06:43:19",
      "content": "<p>this youtube channel has many tutorials:</p>\n<p><a href=\"https://www.youtube.com/@ai4bharat/videos\" target=\"_blank\">https://www.youtube.com/@ai4bharat/videos</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e96ec049ef64542d26eac80c7249bc4%2FSelection_999(2766).png?generation=1690008184783299&amp;alt=media\" alt=\"\"></p>\n<p>e.g<br>\n<a href=\"https://www.youtube.com/watch?v=iSyipKKgleo\" target=\"_blank\">https://www.youtube.com/watch?v=iSyipKKgleo</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2355813,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/23/2023 16:15:41",
      "content": "<p>visualise distance edit</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa32c0720b98a08aedccc73603163a0bf%2FSelection_999(2767).png?generation=1690128936873031&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://github.com/ukiuki-satoshi/visedit\" target=\"_blank\">https://github.com/ukiuki-satoshi/visedit</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2356341,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2023 05:36:49",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd719f707bd940500540adeb99191780d%2FSelection_999(2811).png?generation=1690403372321495&amp;alt=media\" alt=\"\"></p>\n<p><br>\n<br>\n</p>\n<p><br>\n</p>\n<p></p>",
      "votes": null,
      "replies": [
        {
          "id": 2360593,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/26/2023 20:31:06",
          "content": "<p>i made a mistake. i forget about normalisation and others in post-processing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2356475,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2023 08:14:01",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F11ad141e94c7a32b3148618428ab7b11%2FSelection_999(2797).png?generation=1690186435573149&amp;alt=media\" alt=\"\"></p>\n<p>OOD results</p>",
      "votes": null,
      "replies": [
        {
          "id": 2359089,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "07/26/2023 01:41:46",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  How did you get the ground truth for the test_example  to calculate the wer ? It's not available in the Data section </p>\n<p>EDIT: sorry, I've found the source <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2356515,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2023 08:56:36",
      "content": "<p>\" We first show an interesting finding that while Whisper is very robust against real-world background sounds (e.g., music), its audio representation is actually not noise-invariant, but is instead highly correlated to non-speech sounds, indicating that Whisper recognizes speech conditioned on the noise type\"</p>\n<p><a href=\"https://arxiv.org/pdf/2307.03183.pdf\" target=\"_blank\">https://arxiv.org/pdf/2307.03183.pdf</a><br>\nWhisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers</p>",
      "votes": null,
      "replies": [
        {
          "id": 2356526,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/24/2023 09:01:11",
          "content": "<p>End-to-end Music-mixed Speech Recognition<br>\n<a href=\"https://arxiv.org/pdf/2008.12048.pdf\" target=\"_blank\">https://arxiv.org/pdf/2008.12048.pdf</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2356542,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "07/24/2023 09:08:24",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3fbba9c75392b0f594274f017d7bbeec%2FSelection_999(2798).png?generation=1690189696857157&amp;alt=media\" alt=\"\"></p>\n<p>Wav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2358541,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "07/25/2023 15:23:32",
                  "content": "<p>CTC bert?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F40e8bbfbd5da3f9ddd2af5e56f66ed62%2FSelection_999(2803).png?generation=1690298599425663&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F231e9aa6db041b8f6fe453217fbedd6f%2FSelection_999(2804).png?generation=1690298610687049&amp;alt=media\" alt=\"\"></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2356736,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2023 10:52:08",
      "content": "<p><a href=\"https://sites.google.com/iitdh.ac.in/vssasr2021/resources\" target=\"_blank\">https://sites.google.com/iitdh.ac.in/vssasr2021/resources</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F55dbc9f676130c29ea69dd771132459e%2FSelection_999(2799).png?generation=1690195926421740&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2359216,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/26/2023 04:31:33",
      "content": "<p>massive external data:<br>\n<a href=\"https://github.com/AI4Bharat/vistaar#download-training-datasets-and-benchmarks\" target=\"_blank\">https://github.com/AI4Bharat/vistaar#download-training-datasets-and-benchmarks</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2360878,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/27/2023 04:36:08",
          "content": "<p>yet another external data<br>\n<a href=\"https://globalrecordings.net\" target=\"_blank\">https://globalrecordings.net</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2360720,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/27/2023 00:33:30",
      "content": "<p>paper on effects  of in and out domain :<br>\nROBUST WAV2VEC 2.0: ANALYZING DOMAIN SHIFT IN SELF-SUPERVISED PRE-TRAINING<br>\n<a href=\"https://arxiv.org/pdf/2104.01027.pdf\" target=\"_blank\">https://arxiv.org/pdf/2104.01027.pdf</a></p>\n<p>\"In this paper, we explore more general setups where the domain of the unlabeled data for pre-training data differs from the domain of the labeled data for fine-tuning, which in turn may differ from the test data domain\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2365899,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/30/2023 14:44:37",
      "content": "<p>i wonder if the each sentence in the test is unique?<br>\nif not, this is a leak?<br>\n(e.g. one may collect speech data by asking different people <strong>to speak the same sentence</strong>)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2368194,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/01/2023 04:04:52",
      "content": "<p>whisper takes in 30sec audio chunk. so i did a probe on audio clip duration of hidden test data:</p>\n<pre><code>valid_df = pd()\nmp3_dir = f\n\nduration=\n t,d  valid_df():\n    audio_file = f\n\n    sampling_rate = _000\n    with (audio_file, ) as f:\n        bpayload = f()\n     = (bpayload, sampling_rate)\n    dur = (a)/sampling_rate\n    (dur)\n    duration(dur)\n\nmax_duration=np(duration)\n\n\n</code></pre>\n<p>submission failed for : max_duration&lt;=30 sec<br>\nsubmission passed for : max_duration&lt;=32 sec</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2368354,
      "author_name": "aniltk",
      "author_url": "",
      "post_date": "08/01/2023 06:08:38",
      "content": "<p>Thank you for the info … so much research on the topic great …all the best</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2369358,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/01/2023 17:34:50",
      "content": "<p>another dataset<br>\nDataset - RESPIN<br>\n<a href=\"https://slt2022.org/projects/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri.pdf\" target=\"_blank\">https://slt2022.org/projects/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri.pdf</a></p>\n<p><a href=\"https://sites.google.com/view/respinasrchallenge2023/home\" target=\"_blank\">https://sites.google.com/view/respinasrchallenge2023/home</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6ff8e18a7c97388e287e41fd5bcf58bb%2FSelection_999(2829).png?generation=1690911460374350&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2369394,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/01/2023 17:57:23",
          "content": "<p>code switch dataset:<br>\n<a href=\"https://github.com/navana-tech/baseline_recipe_is21s_indic_asr_challenge\" target=\"_blank\">https://github.com/navana-tech/baseline_recipe_is21s_indic_asr_challenge</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2370469,
      "author_name": "aphysict",
      "author_url": "",
      "post_date": "08/02/2023 12:43:43",
      "content": "<p>Hi, is indicwav2vec_v1_bengali simply transformed from the fairseq model on indicwav2vec github repo?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2370496,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/02/2023 13:09:32",
          "content": "<p>i have submitted the version from hugging face.</p>\n<p>i may choose to use the github version later</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2370590,
          "author_name": "aphysict",
          "author_url": "",
          "post_date": "08/02/2023 14:04:55",
          "content": "<p><br>\nThey are the same model</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2372470,
      "author_name": "nisshokuitsuki",
      "author_url": "",
      "post_date": "08/03/2023 17:32:09",
      "content": "<p>In my experiment, directly binding IndicWav2Vec2 to Yellowking's decoder has led to some problems, that is the vocabularie sizes have not matched The former is 87 and the latter is 112. How can I solve it? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2372494,
          "author_name": "nisshokuitsuki",
          "author_url": "",
          "post_date": "08/03/2023 17:58:12",
          "content": "<p>Is this OK by just setting <code>ignore_mismatched_sizes=True</code></p>",
          "votes": null,
          "replies": [
            {
              "id": 2379732,
              "author_name": "nisshokuitsuki",
              "author_url": "",
              "post_date": "08/08/2023 09:07:32",
              "content": "<p>There may be some problems, because the vocab.json of <code>indicwav2vec_v1_bengali</code> and other models (bengali, with LM) are not the same. I tried rewriting the alphabet.json and the model infers some nonsense.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2381841,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "08/09/2023 12:20:28",
                  "content": "<p>this is my code</p>\n<pre><code> transformers  pipeline\n transformers  Wav2Vec2CTCTokenizer, Wav2Vec2FeatureExtractor, Wav2Vec2Processor,Wav2Vec2ProcessorWithLM\n transformers  Wav2Vec2ForCTC\n\n\npretrain_model = \\\n    \nvocab_dir = \\\n    \narpa_file= \\\n     \n   \n\n\n\nsampling_rate=\ntokenizer = Wav2Vec2CTCTokenizer.from_pretrained(\n    vocab_dir, \n    unk_token=,\n    pad_token=,\n    word_delimiter_token=, \n    bos_token=,\n    eos_token=,\n)\n\n\n\nfeature_extractor = Wav2Vec2FeatureExtractor(\n    feature_size=,\n    sampling_rate=sampling_rate,\n    padding_value=,\n    padding_side=,\n    do_normalize=,\n    return_attention_mask=,\n)\n\n :\n     pyctcdecode  BeamSearchDecoderCTC\n     pyctcdecode  build_ctcdecoder\n\n    vocab_dict = tokenizer.get_vocab()\n\n    vocab_dict = {k: v  k, v  (vocab_dict.items(), key= item: item[])}\n    decoder = build_ctcdecoder(\n        labels=(vocab_dict.keys()),\n        kenlm_model_path=arpa_file,\n    )\n\nprocessor = Wav2Vec2ProcessorWithLM(\n    feature_extractor=feature_extractor,\n    tokenizer=tokenizer,\n    decoder=decoder,\n)\n\n\n (nn.Module):\n     ():\n        ().__init__()\n        self.output_type = []\n        self.model = Wav2Vec2ForCTC.from_pretrained(\n            pretrain_model,\n            ignore_mismatched_sizes=,\n            attention_dropout= ,\n            hidden_dropout=  ,\n            feat_proj_dropout=  ,\n            mask_time_prob= ,\n            layerdrop= ,\n            ctc_loss_reduction=,\n            pad_token_id=processor.tokenizer.pad_token_id,\n            vocab_size=(processor.tokenizer),\n        )\n        \n     ():\n        B,L = batch[].shape\n\n        \n        out = self.model (\n            batch[],\n            batch[],\n            labels = batch.get(,),\n        )\n\n\n        output = {}\n           self.output_type:\n            output[] = out.logits\n            \n            \n            output[] = postprocess_lm_to_text(out, batch[])\n            \n\n           self.output_type:\n            output[] = out.loss\n\n         output\n\n\n ():\n    B = (out.logits)\n    \n    logit = out.logits\n    token = logit.argmax(dim=-)\n    text = []\n     b  (B):\n        t = token[b][:length[b]//]\n        t = tokenizer.decode(t, skip_special_tokens=)\n        text.append(t)\n     text\n\n ():\n    B = (out.logits)\n    \n    logit = out.logits.data.cpu().numpy()\n    text = []\n     b  (B):\n        l = logit[b]\n        beam = decoder.decode_beams(l, beam_width=)\n        t = beam[][]\n        text.append(t)\n     text\n</code></pre>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2381844,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "08/09/2023 12:21:31",
                  "content": "<p>LM model and wave2vec models are independent.<br>\ni have been mixing different wav2vec and LM from different sources</p>\n<p>i also tried  Flashlight decoder, and it would also work</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2381973,
                      "author_name": "nisshokuitsuki",
                      "author_url": "",
                      "post_date": "08/09/2023 13:44:30",
                      "content": "<p>Thanks a lot for your detailed help! I'm really appreciating that</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2438832,
      "author_name": "barnobarno",
      "author_url": "",
      "post_date": "09/14/2023 14:10:06",
      "content": "<p>Thank you </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2350183": "#more details later,  ... here is the lb0.481 recipe:\n#[a]model: \nhttps://huggingface.co/ai4bharat/indicwav2vec_v1_bengali (CTC model only)\n\n#[b]decoder: \nadd a LM yourself, you can train one or use e.g. https://huggingface.co/shahruk10/wav2vec2-xls-r-300m-bengali-commonvoice\n\n#[c]post-process:\nnormalise + dari\n\n```\nsome lbscore:\n[a]+[b]+[c]: 0.481\n[a]+[c]: 0.520\n[a]+[b]: 0.550\n[a]: 0.585\n[a]+[b]+[c/only normalise]: 0.490\n```\n\n----\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd1010fbf802e033d78dfb6ed2043dbea%2FSelection_999(2747).png?generation=1689742315359191&alt=media)\n\nhttps://www.kaggle.com/code/hengck23/local-wer-0-2600-nemo-baseline-conformer\ninitial results .... to be updated as experiments progress",
    "2350184": "some important todo list:\n- speed up whsiper inference (e.g. JAX or openAI C api)",
    "2350187": "the trick to winning is really the OOD (out of distribution) private test set\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa28e51cd1564da2a042d70893ec20d51%2FSelection_999(2748).png?generation=1689742664897606&alt=media)\n\nNote:\n1. from https://www.kaggle.com/competitions/bengaliai-speech/data\nThe full test set contains about 20 hours of speech in almost 8000 MP3 audio files, public LB is 46% of the test data\n2. from dataset paper\n- OOD = 2681\n- Macro Test = 4872\n- all = 7553",
    "2350194": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F261acc12207cdb46587c27f675c55cdb%2FSelection_999(2750).png?generation=1689743292178787&alt=media)\n\nyou should read the paper very carefully",
    "2350229": "papers related to out-of-distribution and asr:\n- to be updated",
    "2351573": "using lanuage model is another trick\nhttps://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\n\nhttps://huggingface.co/blog/wav2vec2-with-ngram\n```\nWithout language model :\n\nWER: 0.21726385291857586\nCER: 0.04725010353701041\nWith 5 gram language model trained on 30M sentences randomly chosen from AI4Bharat IndicCorp dataset :\n\nWER: 0.15322879016421437\nCER: 0.03413696666806267\n\n```",
    "2353514": "i think i find some OOD dataset:\n\nhttps://github.com/Open-Speech-EkStep/ULCA-asr-dataset-corpus\n\ne.g. youtube data\nEntertainment\tMirchi_Bangla_1\tLabelled\tMirchi_Bangla_1\t54 (hrs)\nEntertainment\tMirchi_Bangla_2\tLabelled\tMirchi_Bangla_2\t\nEntertainment\tMirchi_Bangla_3\tLabelled\tMirchi_Bangla_3\t\nGeneral\tCTVN_AKD_PLUS_1\tLabelled\tCTVN_AKD_PLUS_1\t21.7 (hrs)\nGeneral\tCTVN_AKD_PLUS_2\tLabelled\tCTVN_AKD_PLUS_2\t\n\n\nhowever, the data are not available ....\nbut the pretrain model are here\nhttps://github.com/Open-Speech-EkStep/vakyansh-models\nVakyansh-Conformer-SSL\n\n\"This model was pre-trained using Nemo toolkit with 34,000 hours unlabeled audio in 39 Indian languages. This includes 15,000 hours of news recordings available on the internet, 10,000 hours of YouTube audios and other audio data.\"\n\nmodel config yaml:\ntrain_ds:\n  manifest_filepath: /mnt/lustre/megh/indic-ssl/manifest/crisil_noa_nptel_yt_indic_ssl_train.json\n\ni am guessing\nnoa : news on air\nyt: youtube\n\n\npaper: \nVakyansh: ASR Toolkit for Low Resource Indic languages\nhttps://arxiv.org/pdf/2203.16512.pdf",
    "2353528": "sampling rate as TTA?\n\n```\nsampling_rate=16000\n\nif 1:  #test one\n\td = valid_df.iloc[0]\n\tmp3_file = f'{mp3_dir}/{d[\"id\"]}.mp3'\n\n\t#p = pipe(mp3_file)['text'] \n\twith open(mp3_file, 'rb') as f:\n\t\tbpayload = f.read()\n\ta = ffmpeg_read( bpayload , sampling_rate)\n\tp = pipe(a)['text'] \n\tprint('predict', p)\n\tprint('truth  ', d.sentence)\n\n```\n\n```\n18000 Hz\npredict ও বলেছে আপনার টিকাপ\ntruth   ও বলেছে আপনার ঠিকানা!\n\n16000 Hz\npredict ওবলেছে আপনার টিকা\ntruth   ও বলেছে আপনার ঠিকানা!\n\n32000 Hz\npredict এগুন শ অ\ntruth   ও বলেছে আপনার ঠিকানা!\n\n14000 Hz\npredict অবিচয় আনাটি\ntruth   ও বলেছে আপনার ঠিকানা!\n\n20000 Hz\npredict ও বলেছে আপনার টিকা\ntruth   ও বলেছে আপনার ঠিকানা!\n```",
    "2353969": "this youtube channel has many tutorials:\n\nhttps://www.youtube.com/@ai4bharat/videos\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e96ec049ef64542d26eac80c7249bc4%2FSelection_999(2766).png?generation=1690008184783299&alt=media)\n\ne.g\nhttps://www.youtube.com/watch?v=iSyipKKgleo",
    "2355813": "visualise distance edit\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa32c0720b98a08aedccc73603163a0bf%2FSelection_999(2767).png?generation=1690128936873031&alt=media)\n\nhttps://github.com/ukiuki-satoshi/visedit",
    "2356341": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd719f707bd940500540adeb99191780d%2FSelection_999(2811).png?generation=1690403372321495&alt=media)\n\n\n~~should we be using kaggle dataset \"as it is\" at all?~~\n~~i note that the baseline conformer-CTC has good results on train.csv valid split (clean label). But it only has LB0.68.~~\n~~On the other hands, public huggingface/ai4bharat/indicwav2vec_v1_bengali only trained on their datase ~~~~(radio boradcast and youtube) has less performance gap:  valid split  0.39, LB0.52.~~\n\n~~This makes me suspect kaggle data is not good on hidden OOD public dataset.~~\n~~In my experiments abovem, the better results you get for train.csv valid split (or train split), you get worse ~~~~results for publuc LB.~~\n\n~~Kagglers may want to confirm this observation~~",
    "2356475": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F11ad141e94c7a32b3148618428ab7b11%2FSelection_999(2797).png?generation=1690186435573149&alt=media)\n\nOOD results",
    "2356515": "\" We first show an interesting finding that while Whisper is very robust against real-world background sounds (e.g., music), its audio representation is actually not noise-invariant, but is instead highly correlated to non-speech sounds, indicating that Whisper recognizes speech conditioned on the noise type\"\n\nhttps://arxiv.org/pdf/2307.03183.pdf\nWhisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers",
    "2356526": "End-to-end Music-mixed Speech Recognition\nhttps://arxiv.org/pdf/2008.12048.pdf",
    "2356542": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3fbba9c75392b0f594274f017d7bbeec%2FSelection_999(2798).png?generation=1690189696857157&alt=media)\n\nWav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition",
    "2356736": "https://sites.google.com/iitdh.ac.in/vssasr2021/resources\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F55dbc9f676130c29ea69dd771132459e%2FSelection_999(2799).png?generation=1690195926421740&alt=media)",
    "2358541": "CTC bert?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F40e8bbfbd5da3f9ddd2af5e56f66ed62%2FSelection_999(2803).png?generation=1690298599425663&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F231e9aa6db041b8f6fe453217fbedd6f%2FSelection_999(2804).png?generation=1690298610687049&alt=media)",
    "2359089": "hengck23  How did you get the ground truth for the test_example  to calculate the wer ? It's not available in the Data section \n\nEDIT: sorry, I've found the source https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932",
    "2359216": "massive external data:\nhttps://github.com/AI4Bharat/vistaar#download-training-datasets-and-benchmarks",
    "2360593": "i made a mistake. i forget about normalisation and others in post-processing",
    "2360720": "paper on effects  of in and out domain :\nROBUST WAV2VEC 2.0: ANALYZING DOMAIN SHIFT IN SELF-SUPERVISED PRE-TRAINING\nhttps://arxiv.org/pdf/2104.01027.pdf\n\n\"In this paper, we explore more general setups where the domain of the unlabeled data for pre-training data differs from the domain of the labeled data for fine-tuning, which in turn may differ from the test data domain\"",
    "2360878": "yet another external data\nhttps://globalrecordings.net",
    "2365899": "i wonder if the each sentence in the test is unique?\nif not, this is a leak?\n(e.g. one may collect speech data by asking different people **to speak the same sentence**)",
    "2367644": "Thanks for posting all the insights. May I ask how long does it take you to run whsiper for leaderboard submssion?",
    "2368194": "whisper takes in 30sec audio chunk. so i did a probe on audio clip duration of hidden test data:\n```\nvalid_df = pd.read_csv('/kaggle/input/bengaliai-speech/sample_submission.csv')\nmp3_dir = f'/kaggle/input/bengaliai-speech/test_mp3s'\n\nduration=[]\nfor t,d in valid_df.iterrows():\n    audio_file = f'{mp3_dir}/{d[\"id\"]}.mp3'\n\n    sampling_rate = 16_000\n    with open(audio_file, 'rb') as f:\n        bpayload = f.read()\n    a = ffmpeg_read(bpayload, sampling_rate)\n    dur = len(a)/sampling_rate\n    print(dur)\n    duration.append(dur)\n\nmax_duration=np.max(duration)\nprint('max_duration',max_duration)\nassert(max_duration<=DURATION)\n\n```\n\nsubmission failed for : max_duration<=30 sec\nsubmission passed for : max_duration<=32 sec",
    "2368354": "Thank you for the info ... so much research on the topic great ...all the best",
    "2369358": "another dataset\nDataset - RESPIN\nhttps://slt2022.org/projects/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri/13%20-%20Dialectical%20speech%20recognition%20for%20two%20Indian%20languages%20-%20Bengali%20and%20Bhojpuri.pdf\n\nhttps://sites.google.com/view/respinasrchallenge2023/home\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6ff8e18a7c97388e287e41fd5bcf58bb%2FSelection_999(2829).png?generation=1690911460374350&alt=media)",
    "2369394": "code switch dataset:\nhttps://github.com/navana-tech/baseline_recipe_is21s_indic_asr_challenge",
    "2370469": "Hi, is indicwav2vec_v1_bengali simply transformed from the fairseq model on indicwav2vec github repo?",
    "2370496": "i have submitted the version from hugging face.\n\ni may choose to use the github version later",
    "2370590": "~~I try the fairseq model a little, it's fast but the wer is high. Maybe they are not the same model, or I did something wrong. ~~\nThey are the same model",
    "2372470": "In my experiment, directly binding IndicWav2Vec2 to Yellowking's decoder has led to some problems, that is the vocabularie sizes have not matched The former is 87 and the latter is 112. How can I solve it? Thanks!",
    "2372494": "Is this OK by just setting `ignore_mismatched_sizes=True`",
    "2379732": "There may be some problems, because the vocab.json of `indicwav2vec_v1_bengali` and other models (bengali, with LM) are not the same. I tried rewriting the alphabet.json and the model infers some nonsense.",
    "2381841": "this is my code\n\n```\nfrom transformers import pipeline\nfrom transformers import Wav2Vec2CTCTokenizer, Wav2Vec2FeatureExtractor, Wav2Vec2Processor,Wav2Vec2ProcessorWithLM\nfrom transformers import Wav2Vec2ForCTC\n\n\npretrain_model = \\\n    '/kaggle/input/ai4bharat-indicwav2vec-v1-bengali'\nvocab_dir = \\\n    '/kaggle/input/ai4bharat-indicwav2vec-v1-bengali'\narpa_file= \\\n    '/kaggle/input/my-weight-bengali-asr-01/my-lm-5gram.arpa' #or binary file\n   #'/kaggle/input/my-weight-bengali-asr-01/language_model/commonvoice-bn.5.arpa'\n\n##########################################################################\n#model    \nsampling_rate=16_000\ntokenizer = Wav2Vec2CTCTokenizer.from_pretrained(\n    vocab_dir, #'my_tokenizer',\n    unk_token='<unk>',\n    pad_token='<pad>',\n    word_delimiter_token='|', ##<todo>???\n    bos_token='<s>',\n    eos_token='</s>',\n)\n'''\ntokenizer.convert_tokens_to_ids(['|','<s>', '</s>', '<unk>', '<pad>'])\nOut[1]: [62, 1, 2, 3, 0]\n'''\n\n# just for padding, etc : audio to pad_audio, mask\nfeature_extractor = Wav2Vec2FeatureExtractor(\n    feature_size=1,\n    sampling_rate=sampling_rate,\n    padding_value=0.0,\n    padding_side='right',\n    do_normalize=True,\n    return_attention_mask=True,\n)\n\nif 1:\n\tfrom pyctcdecode import BeamSearchDecoderCTC\n\tfrom pyctcdecode import build_ctcdecoder\n\n\tvocab_dict = tokenizer.get_vocab()\n\n\tvocab_dict = {k: v for k, v in sorted(vocab_dict.items(), key=lambda item: item[1])}\n\tdecoder = build_ctcdecoder(\n\t    labels=list(vocab_dict.keys()),\n\t    kenlm_model_path=arpa_file,\n\t)\n\nprocessor = Wav2Vec2ProcessorWithLM(\n    feature_extractor=feature_extractor,\n    tokenizer=tokenizer,\n\tdecoder=decoder,\n)\n\n\nclass Net(nn.Module):\n\tdef __init__(self, ):\n\t\tsuper().__init__()\n\t\tself.output_type = ['inference']\n\t\tself.model = Wav2Vec2ForCTC.from_pretrained(\n\t\t    pretrain_model,\n\t\t    ignore_mismatched_sizes=False,\n\t\t    attention_dropout= 0,#0.1,\n\t\t    hidden_dropout=  0,#0.1,\n\t\t    feat_proj_dropout=  0,#0.1,\n\t\t    mask_time_prob= 0,#0.05,\n\t\t    layerdrop= 0,#0.1,\n\t\t    ctc_loss_reduction='mean',\n\t\t    pad_token_id=processor.tokenizer.pad_token_id,\n\t\t    vocab_size=len(processor.tokenizer),\n\t\t)\n\t\t#self.model.config.ctc_zero_infinity = True\n\tdef forward(self, batch):\n\t\tB,L = batch['input_values'].shape\n\n\t\t#class Wav2Vec2ForCTC(Wav2Vec2PreTrainedModel):\n\t\tout = self.model (\n\t\t\tbatch['input_values'],\n\t\t\tbatch['attention_mask'],\n\t\t\tlabels = batch.get('labels',None),\n\t\t)#CausalLMOutput\n\n\n\t\toutput = {}\n\t\tif 'inference' in self.output_type:\n\t\t\toutput['logit'] = out.logits\n\t\t\t#output['hidden_state'] = out.hidden_states\n\t\t\t#'attention' : out.attentions,\n\t\t\toutput['text'] = postprocess_lm_to_text(out, batch['length'])\n\t\t\t#output['text'] = postprocess_ctc_to_text(out, batch['length'])\n\n\t\tif 'loss' in self.output_type:\n\t\t\toutput['ctc_loss'] = out.loss\n\n\t\treturn output\n\n\ndef postprocess_ctc_to_text(out, length):\n\tB = len(out.logits)\n\t# AutomaticSpeechRecognitionPipeline(ChunkPipeline)\n\tlogit = out.logits\n\ttoken = logit.argmax(dim=-1)\n\ttext = []\n\tfor b in range(B):\n\t\tt = token[b][:length[b]//320]\n\t\tt = tokenizer.decode(t, skip_special_tokens=False)\n\t\ttext.append(t)\n\treturn text\n \ndef postprocess_lm_to_text(out, length):\n\tB = len(out.logits)\n\t# AutomaticSpeechRecognitionPipeline(ChunkPipeline)\n\tlogit = out.logits.data.cpu().numpy()\n\ttext = []\n\tfor b in range(B):\n\t\tl = logit[b]#[:length[b]//320]\n\t\tbeam = decoder.decode_beams(l, beam_width=512)\n\t\tt = beam[0][0]\n\t\ttext.append(t)\n\treturn text\n```",
    "2381844": "LM model and wave2vec models are independent.\ni have been mixing different wav2vec and LM from different sources\n\ni also tried  Flashlight decoder, and it would also work",
    "2381973": "Thanks a lot for your detailed help! I'm really appreciating that",
    "2438832": "Thank you"
  },
  "source": "meta"
}