{
  "id": 428235,
  "title": "whisper will perform pretty poorly in Bengali?",
  "url": "/competitions/bengaliai-speech/discussion/428235",
  "author_name": "",
  "post_date": "2023-07-31T16:58:12.458202Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>When I read through the Whisper paper ie Robust Speech Recognition via Large-Scale Weak Supervision by OpenAI team. I noticed the performance of bengali is really bad based on bench-marking results in two datasets - Common Voice9 and Fleurs.</p>\n<ul>\n<li>COMMON VOICE 9</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1310697%2F353eddf6b198743d96e68da92839dfd4%2FScreenshot%202023-07-31%20at%2022-23-37%20whisper.pdf.png?generation=1690822464550409&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>FLEURS</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1310697%2Fc25dc14105013abbc40adcd6598a5687%2FScreenshot%202023-07-31%20at%2022-23-57%20whisper.pdf.png?generation=1690822545094901&amp;alt=media\" alt=\"\"></p>\n<p>I suspect because of this directly using whisper model weights might be a bad idea. Yet fine-tuning whisper might be promising as in low-resource languages like Malayalam we have been able to fine tune and reach an WER of below 0.1 in Common Voice 11 dataset.</p>",
  "messages": [
    {
      "id": "2367609",
      "postDate": "07/31/2023 16:58:12",
      "content": "<p>When I read through the Whisper paper ie Robust Speech Recognition via Large-Scale Weak Supervision by OpenAI team. I noticed the performance of bengali is really bad based on bench-marking results in two datasets - Common Voice9 and Fleurs.</p>\n<ul>\n<li>COMMON VOICE 9</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1310697%2F353eddf6b198743d96e68da92839dfd4%2FScreenshot%202023-07-31%20at%2022-23-37%20whisper.pdf.png?generation=1690822464550409&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>FLEURS</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1310697%2Fc25dc14105013abbc40adcd6598a5687%2FScreenshot%202023-07-31%20at%2022-23-57%20whisper.pdf.png?generation=1690822545094901&amp;alt=media\" alt=\"\"></p>\n<p>I suspect because of this directly using whisper model weights might be a bad idea. Yet fine-tuning whisper might be promising as in low-resource languages like Malayalam we have been able to fine tune and reach an WER of below 0.1 in Common Voice 11 dataset.</p>",
      "rawMarkdown": "When I read through the Whisper paper ie Robust Speech Recognition via Large-Scale Weak Supervision by OpenAI team. I noticed the performance of bengali is really bad based on bench-marking results in two datasets - Common Voice9 and Fleurs.\n\n- COMMON VOICE 9\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1310697%2F353eddf6b198743d96e68da92839dfd4%2FScreenshot%202023-07-31%20at%2022-23-37%20whisper.pdf.png?generation=1690822464550409&alt=media)\n\n- FLEURS\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1310697%2Fc25dc14105013abbc40adcd6598a5687%2FScreenshot%202023-07-31%20at%2022-23-57%20whisper.pdf.png?generation=1690822545094901&alt=media)\n\n\nI suspect because of this directly using whisper model weights might be a bad idea. Yet fine-tuning whisper might be promising as in low-resource languages like Malayalam we have been able to fine tune and reach an WER of below 0.1 in Common Voice 11 dataset.",
      "votes": null
    },
    {
      "id": "2367622",
      "postDate": "07/31/2023 17:03:25",
      "content": "<p>this is very interesting to see, thanks</p>",
      "rawMarkdown": "this is very interesting to see, thanks",
      "votes": null
    },
    {
      "id": "2367746",
      "postDate": "07/31/2023 19:07:19",
      "content": "<p>it depends on your training data, check this instead:<br>\n<a href=\"https://github.com/AI4Bharat/vistaar\" target=\"_blank\">https://github.com/AI4Bharat/vistaar</a></p>\n<p>the downside of whisper i think is its speed and( unexpected behavior which results in erroneous prediction for some few simple cases)</p>",
      "rawMarkdown": "it depends on your training data, check this instead:\nhttps://github.com/AI4Bharat/vistaar\n\nthe downside of whisper i think is its speed and( unexpected behavior which results in erroneous prediction for some few simple cases)",
      "votes": null
    },
    {
      "id": "2369253",
      "postDate": "08/01/2023 16:08:20",
      "content": "<p>Thanks for sharing about vistaar. TIL moment.</p>\n<blockquote>\n  <p>the downside of whisper i think is its speed and( unexpected behavior which results in erroneous prediction for some few simple cases)</p>\n</blockquote>\n<p>I agree the inference of whisper is pretty slow. Yet <a href=\"https://github.com/guillaumekln/faster-whisper\" target=\"_blank\">faster-whisper</a> project kind of makes the inference lot more faster now.</p>",
      "rawMarkdown": "Thanks for sharing about vistaar. TIL moment.\n\n> the downside of whisper i think is its speed and( unexpected behavior which results in erroneous prediction for some few simple cases)\n\nI agree the inference of whisper is pretty slow. Yet [faster-whisper](https://github.com/guillaumekln/faster-whisper) project kind of makes the inference lot more faster now.",
      "votes": null
    },
    {
      "id": "2369333",
      "postDate": "08/01/2023 17:05:55",
      "content": "<p>after my experiments, here are the good news:</p>\n<p>huggingface whisper-medium:</p>\n<ul>\n<li>on old pascal titan x gpu (12 gb) : batch=8 takes about 2 hr for 8000 mp3s files.</li>\n<li>since whisper pad all audio input to 30sec, i think we can use whisper in submission.</li>\n<li>i also probe that hidden test mp3s are mostly less than 30sec, maybe with a few within 30 to32 sec</li>\n</ul>\n<p>i think whisper-medium may performance better than wav2vec-large?</p>\n<hr>\n<p>i tried faster-whisper, but there is some loss of accuracy (maybe there is a bug in my code). it is faster but cannot do batching? there is another wispherX = faster-whisper  + batch<br>\n<a href=\"https://github.com/m-bain/whisperX\" target=\"_blank\">https://github.com/m-bain/whisperX</a></p>",
      "rawMarkdown": "after my experiments, here are the good news:\n\nhuggingface whisper-medium:\n- on old pascal titan x gpu (12 gb) : batch=8 takes about 2 hr for 8000 mp3s files.\n- since whisper pad all audio input to 30sec, i think we can use whisper in submission.\n- i also probe that hidden test mp3s are mostly less than 30sec, maybe with a few within 30 to32 sec\n\ni think whisper-medium may performance better than wav2vec-large?\n\n---\n\ni tried faster-whisper, but there is some loss of accuracy (maybe there is a bug in my code). it is faster but cannot do batching? there is another wispherX = faster-whisper  + batch\nhttps://github.com/m-bain/whisperX",
      "votes": null
    },
    {
      "id": "2369346",
      "postDate": "08/01/2023 17:17:19",
      "content": "<p>do note that most of SOTA asr benchmark you found in papers and github may be different from our kaggle competition.<br>\nFor example, the punctuation are removed from ground truth and text normalization/capitalization are applied.</p>\n<p>actually the trick to win this competition may be these post processing, e.g.</p>\n<ol>\n<li>according to my statistics, test mp3 has average length of e.g. 5sec, with  e.g. 10 words.</li>\n<li>there is on average 1.5 punctuation per clip.</li>\n<li>the bari (full stop is attached to the last word) and so if you don't detect it. it would be a substitution error in WER.</li>\n<li>even if you detect all words perfectly without punctuation marks, you have a WER of 0.15.</li>\n<li>just by building good punctuation model (either within your model or as good post-processing), you probably gain +0.05 to 0.10 in LB score</li>\n</ol>",
      "rawMarkdown": "do note that most of SOTA asr benchmark you found in papers and github may be different from our kaggle competition.\nFor example, the punctuation are removed from ground truth and text normalization/capitalization are applied.\n\nactually the trick to win this competition may be these post processing, e.g.\n1. according to my statistics, test mp3 has average length of e.g. 5sec, with  e.g. 10 words.\n2. there is on average 1.5 punctuation per clip.\n3. the bari (full stop is attached to the last word) and so if you don't detect it. it would be a substitution error in WER.\n4. even if you detect all words perfectly without punctuation marks, you have a WER of 0.15.\n5. just by building good punctuation model (either within your model or as good post-processing), you probably gain +0.05 to 0.10 in LB score",
      "votes": null
    },
    {
      "id": "2369370",
      "postDate": "08/01/2023 17:44:51",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb60f4f52a147ceb769724e5d0c067965%2FSelection_999(2830).png?generation=1690911869019500&amp;alt=media\" alt=\"\"></p>\n<p>another whisper results</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb60f4f52a147ceb769724e5d0c067965%2FSelection_999(2830).png?generation=1690911869019500&alt=media)\n\nanother whisper results",
      "votes": null
    },
    {
      "id": "2369372",
      "postDate": "08/01/2023 17:46:48",
      "content": "<p>Under review as a conference paper at ICLR 2023<br>\nESC: A BENCHMARK FOR MULTI-DOMAIN END-TO-END SPEECH RECOGNITION<br>\n<a href=\"https://openreview.net/pdf?id=9OL2fIfDLK\" target=\"_blank\">https://openreview.net/pdf?id=9OL2fIfDLK</a></p>\n<p>yet another whisper results</p>",
      "rawMarkdown": "Under review as a conference paper at ICLR 2023\nESC: A BENCHMARK FOR MULTI-DOMAIN END-TO-END SPEECH RECOGNITION\nhttps://openreview.net/pdf?id=9OL2fIfDLK\n\nyet another whisper results",
      "votes": null
    },
    {
      "id": "2369657",
      "postDate": "08/01/2023 23:13:09",
      "content": "<p>I tried submitting whisper-medium using the same code that worked for whisper-small and it failed with a scoring error.  Were you able to get it submit fully?</p>",
      "rawMarkdown": "I tried submitting whisper-medium using the same code that worked for whisper-small and it failed with a scoring error.  Were you able to get it submit fully?",
      "votes": null
    },
    {
      "id": "2369763",
      "postDate": "08/02/2023 02:36:33",
      "content": "<p>scoring error probably due to this<br>\n<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425942\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/425942</a></p>",
      "rawMarkdown": "scoring error probably due to this\nhttps://www.kaggle.com/competitions/bengaliai-speech/discussion/425942",
      "votes": null
    },
    {
      "id": "2369855",
      "postDate": "08/02/2023 04:13:37",
      "content": "<p>Interesting! I thought punctuation was not taken into account when calculating WER. It is a pity that this is not specified anywhere in the description.</p>",
      "rawMarkdown": "Interesting! I thought punctuation was not taken into account when calculating WER. It is a pity that this is not specified anywhere in the description.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2367622,
      "author_name": "anivana",
      "author_url": "",
      "post_date": "07/31/2023 17:03:25",
      "content": "<p>this is very interesting to see, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2367746,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2023 19:07:19",
      "content": "<p>it depends on your training data, check this instead:<br>\n<a href=\"https://github.com/AI4Bharat/vistaar\" target=\"_blank\">https://github.com/AI4Bharat/vistaar</a></p>\n<p>the downside of whisper i think is its speed and( unexpected behavior which results in erroneous prediction for some few simple cases)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2369253,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "08/01/2023 16:08:20",
          "content": "<p>Thanks for sharing about vistaar. TIL moment.</p>\n<blockquote>\n  <p>the downside of whisper i think is its speed and( unexpected behavior which results in erroneous prediction for some few simple cases)</p>\n</blockquote>\n<p>I agree the inference of whisper is pretty slow. Yet <a href=\"https://github.com/guillaumekln/faster-whisper\" target=\"_blank\">faster-whisper</a> project kind of makes the inference lot more faster now.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2369333,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "08/01/2023 17:05:55",
              "content": "<p>after my experiments, here are the good news:</p>\n<p>huggingface whisper-medium:</p>\n<ul>\n<li>on old pascal titan x gpu (12 gb) : batch=8 takes about 2 hr for 8000 mp3s files.</li>\n<li>since whisper pad all audio input to 30sec, i think we can use whisper in submission.</li>\n<li>i also probe that hidden test mp3s are mostly less than 30sec, maybe with a few within 30 to32 sec</li>\n</ul>\n<p>i think whisper-medium may performance better than wav2vec-large?</p>\n<hr>\n<p>i tried faster-whisper, but there is some loss of accuracy (maybe there is a bug in my code). it is faster but cannot do batching? there is another wispherX = faster-whisper  + batch<br>\n<a href=\"https://github.com/m-bain/whisperX\" target=\"_blank\">https://github.com/m-bain/whisperX</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2369370,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "08/01/2023 17:44:51",
                  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb60f4f52a147ceb769724e5d0c067965%2FSelection_999(2830).png?generation=1690911869019500&amp;alt=media\" alt=\"\"></p>\n<p>another whisper results</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2369372,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "08/01/2023 17:46:48",
                      "content": "<p>Under review as a conference paper at ICLR 2023<br>\nESC: A BENCHMARK FOR MULTI-DOMAIN END-TO-END SPEECH RECOGNITION<br>\n<a href=\"https://openreview.net/pdf?id=9OL2fIfDLK\" target=\"_blank\">https://openreview.net/pdf?id=9OL2fIfDLK</a></p>\n<p>yet another whisper results</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2369657,
                          "author_name": "nbroad",
                          "author_url": "",
                          "post_date": "08/01/2023 23:13:09",
                          "content": "<p>I tried submitting whisper-medium using the same code that worked for whisper-small and it failed with a scoring error.  Were you able to get it submit fully?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2369763,
                              "author_name": "hengck23",
                              "author_url": "",
                              "post_date": "08/02/2023 02:36:33",
                              "content": "<p>scoring error probably due to this<br>\n<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425942\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/425942</a></p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2369346,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/01/2023 17:17:19",
      "content": "<p>do note that most of SOTA asr benchmark you found in papers and github may be different from our kaggle competition.<br>\nFor example, the punctuation are removed from ground truth and text normalization/capitalization are applied.</p>\n<p>actually the trick to win this competition may be these post processing, e.g.</p>\n<ol>\n<li>according to my statistics, test mp3 has average length of e.g. 5sec, with  e.g. 10 words.</li>\n<li>there is on average 1.5 punctuation per clip.</li>\n<li>the bari (full stop is attached to the last word) and so if you don't detect it. it would be a substitution error in WER.</li>\n<li>even if you detect all words perfectly without punctuation marks, you have a WER of 0.15.</li>\n<li>just by building good punctuation model (either within your model or as good post-processing), you probably gain +0.05 to 0.10 in LB score</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2369855,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "08/02/2023 04:13:37",
          "content": "<p>Interesting! I thought punctuation was not taken into account when calculating WER. It is a pity that this is not specified anywhere in the description.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2367609": "When I read through the Whisper paper ie Robust Speech Recognition via Large-Scale Weak Supervision by OpenAI team. I noticed the performance of bengali is really bad based on bench-marking results in two datasets - Common Voice9 and Fleurs.\n\n- COMMON VOICE 9\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1310697%2F353eddf6b198743d96e68da92839dfd4%2FScreenshot%202023-07-31%20at%2022-23-37%20whisper.pdf.png?generation=1690822464550409&alt=media)\n\n- FLEURS\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1310697%2Fc25dc14105013abbc40adcd6598a5687%2FScreenshot%202023-07-31%20at%2022-23-57%20whisper.pdf.png?generation=1690822545094901&alt=media)\n\n\nI suspect because of this directly using whisper model weights might be a bad idea. Yet fine-tuning whisper might be promising as in low-resource languages like Malayalam we have been able to fine tune and reach an WER of below 0.1 in Common Voice 11 dataset.",
    "2367622": "this is very interesting to see, thanks",
    "2367746": "it depends on your training data, check this instead:\nhttps://github.com/AI4Bharat/vistaar\n\nthe downside of whisper i think is its speed and( unexpected behavior which results in erroneous prediction for some few simple cases)",
    "2369253": "Thanks for sharing about vistaar. TIL moment.\n\n> the downside of whisper i think is its speed and( unexpected behavior which results in erroneous prediction for some few simple cases)\n\nI agree the inference of whisper is pretty slow. Yet [faster-whisper](https://github.com/guillaumekln/faster-whisper) project kind of makes the inference lot more faster now.",
    "2369333": "after my experiments, here are the good news:\n\nhuggingface whisper-medium:\n- on old pascal titan x gpu (12 gb) : batch=8 takes about 2 hr for 8000 mp3s files.\n- since whisper pad all audio input to 30sec, i think we can use whisper in submission.\n- i also probe that hidden test mp3s are mostly less than 30sec, maybe with a few within 30 to32 sec\n\ni think whisper-medium may performance better than wav2vec-large?\n\n---\n\ni tried faster-whisper, but there is some loss of accuracy (maybe there is a bug in my code). it is faster but cannot do batching? there is another wispherX = faster-whisper  + batch\nhttps://github.com/m-bain/whisperX",
    "2369346": "do note that most of SOTA asr benchmark you found in papers and github may be different from our kaggle competition.\nFor example, the punctuation are removed from ground truth and text normalization/capitalization are applied.\n\nactually the trick to win this competition may be these post processing, e.g.\n1. according to my statistics, test mp3 has average length of e.g. 5sec, with  e.g. 10 words.\n2. there is on average 1.5 punctuation per clip.\n3. the bari (full stop is attached to the last word) and so if you don't detect it. it would be a substitution error in WER.\n4. even if you detect all words perfectly without punctuation marks, you have a WER of 0.15.\n5. just by building good punctuation model (either within your model or as good post-processing), you probably gain +0.05 to 0.10 in LB score",
    "2369370": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb60f4f52a147ceb769724e5d0c067965%2FSelection_999(2830).png?generation=1690911869019500&alt=media)\n\nanother whisper results",
    "2369372": "Under review as a conference paper at ICLR 2023\nESC: A BENCHMARK FOR MULTI-DOMAIN END-TO-END SPEECH RECOGNITION\nhttps://openreview.net/pdf?id=9OL2fIfDLK\n\nyet another whisper results",
    "2369657": "I tried submitting whisper-medium using the same code that worked for whisper-small and it failed with a scoring error.  Were you able to get it submit fully?",
    "2369763": "scoring error probably due to this\nhttps://www.kaggle.com/competitions/bengaliai-speech/discussion/425942",
    "2369855": "Interesting! I thought punctuation was not taken into account when calculating WER. It is a pity that this is not specified anywhere in the description."
  },
  "source": "meta"
}