{
  "id": 436218,
  "title": "Trying to understand the metric",
  "url": "/competitions/bengaliai-speech/discussion/436218",
  "author_name": "",
  "post_date": "2023-09-01T12:18:25.636108100Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In speech recognition, it seems WER is usually given as a percent? For example, for <a href=\"https://huggingface.co/docs/transformers/model_doc/wav2vec2\" target=\"_blank\">wav2vec2</a> they report a WER of 1.8/3.3 on clean/other test sets of Librispeech. So here is my question: does ~0.4 WER on LB equal to 4 WER (close to SOTA) or 40 WER (very, very far behind…)<br>\nOr does this competition use a different definition of WER?<br>\nIn short, I want to understand how the current LB scores are compared to reported SOTA, such as wav2vec2/whisper, etc. Thanks!</p>",
  "messages": [
    {
      "id": "2418608",
      "postDate": "09/01/2023 12:18:25",
      "content": "<p>In speech recognition, it seems WER is usually given as a percent? For example, for <a href=\"https://huggingface.co/docs/transformers/model_doc/wav2vec2\" target=\"_blank\">wav2vec2</a> they report a WER of 1.8/3.3 on clean/other test sets of Librispeech. So here is my question: does ~0.4 WER on LB equal to 4 WER (close to SOTA) or 40 WER (very, very far behind…)<br>\nOr does this competition use a different definition of WER?<br>\nIn short, I want to understand how the current LB scores are compared to reported SOTA, such as wav2vec2/whisper, etc. Thanks!</p>",
      "rawMarkdown": "In speech recognition, it seems WER is usually given as a percent? For example, for [wav2vec2](https://huggingface.co/docs/transformers/model_doc/wav2vec2) they report a WER of 1.8/3.3 on clean/other test sets of Librispeech. So here is my question: does ~0.4 WER on LB equal to 4 WER (close to SOTA) or 40 WER (very, very far behind...)\nOr does this competition use a different definition of WER?\nIn short, I want to understand how the current LB scores are compared to reported SOTA, such as wav2vec2/whisper, etc. Thanks!",
      "votes": null
    },
    {
      "id": "2418640",
      "postDate": "09/01/2023 12:51:07",
      "content": "<p>Great question!</p>\n<p><strong>For individual examples,</strong></p>\n<p>$$<br>\nWER(y_{true}, y_{pred}) = \\frac{N_{insertions} + N_{substitutions} + N_{deletions}}{\\text{Length of $y_{true}$}}<br>\n$$</p>\n<p>Suppose, $y_{true}$ is \"Do Re Mi Fa\" and $y_{pred}$ is \"Do\", WER is 3/4 = 0.75.<br>\nBut if $y_{true}$ is \"Do\" and $y_{pred}$ is \"Do Re Mi Fa\", WER is 3/1 = 3.</p>\n<p>The reported WER in the papers are in percentages. So, 0.4 in LB means 40. Pretty far behind actually!</p>\n<p>References:</p>\n<ol>\n<li><a href=\"https://huggingface.co/openai/whisper-large-v2#evaluation\" target=\"_blank\">https://huggingface.co/openai/whisper-large-v2#evaluation</a></li>\n<li><a href=\"https://paperswithcode.com/sota/automatic-speech-recognition-on-librispeech-7\" target=\"_blank\">https://paperswithcode.com/sota/automatic-speech-recognition-on-librispeech-7</a></li>\n</ol>",
      "rawMarkdown": "Great question!\n\n**For individual examples,**\n\n$$\nWER(y_{true}, y_{pred}) = \\frac{N_{insertions} + N_{substitutions} + N_{deletions}}{\\text{Length of $y_{true}$}}\n$$\n\nSuppose, $y_{true}$ is \"Do Re Mi Fa\" and $y_{pred}$ is \"Do\", WER is 3/4 = 0.75.\nBut if $y_{true}$ is \"Do\" and $y_{pred}$ is \"Do Re Mi Fa\", WER is 3/1 = 3.\n\nThe reported WER in the papers are in percentages. So, 0.4 in LB means 40. Pretty far behind actually!\n\nReferences:\n1. https://huggingface.co/openai/whisper-large-v2#evaluation\n2. https://paperswithcode.com/sota/automatic-speech-recognition-on-librispeech-7",
      "votes": null
    },
    {
      "id": "2418738",
      "postDate": "09/01/2023 13:42:25",
      "content": "<p>The LB scores are in percentage. So 0.4 in LB would mean 40% WER. According to SOTA it'd be 40</p>",
      "rawMarkdown": "The LB scores are in percentage. So 0.4 in LB would mean 40% WER. According to SOTA it'd be 40",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2418640,
      "author_name": "umongsain",
      "author_url": "",
      "post_date": "09/01/2023 12:51:07",
      "content": "<p>Great question!</p>\n<p><strong>For individual examples,</strong></p>\n<p>$$<br>\nWER(y_{true}, y_{pred}) = \\frac{N_{insertions} + N_{substitutions} + N_{deletions}}{\\text{Length of $y_{true}$}}<br>\n$$</p>\n<p>Suppose, $y_{true}$ is \"Do Re Mi Fa\" and $y_{pred}$ is \"Do\", WER is 3/4 = 0.75.<br>\nBut if $y_{true}$ is \"Do\" and $y_{pred}$ is \"Do Re Mi Fa\", WER is 3/1 = 3.</p>\n<p>The reported WER in the papers are in percentages. So, 0.4 in LB means 40. Pretty far behind actually!</p>\n<p>References:</p>\n<ol>\n<li><a href=\"https://huggingface.co/openai/whisper-large-v2#evaluation\" target=\"_blank\">https://huggingface.co/openai/whisper-large-v2#evaluation</a></li>\n<li><a href=\"https://paperswithcode.com/sota/automatic-speech-recognition-on-librispeech-7\" target=\"_blank\">https://paperswithcode.com/sota/automatic-speech-recognition-on-librispeech-7</a></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2418738,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "09/01/2023 13:42:25",
      "content": "<p>The LB scores are in percentage. So 0.4 in LB would mean 40% WER. According to SOTA it'd be 40</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2418608": "In speech recognition, it seems WER is usually given as a percent? For example, for [wav2vec2](https://huggingface.co/docs/transformers/model_doc/wav2vec2) they report a WER of 1.8/3.3 on clean/other test sets of Librispeech. So here is my question: does ~0.4 WER on LB equal to 4 WER (close to SOTA) or 40 WER (very, very far behind...)\nOr does this competition use a different definition of WER?\nIn short, I want to understand how the current LB scores are compared to reported SOTA, such as wav2vec2/whisper, etc. Thanks!",
    "2418640": "Great question!\n\n**For individual examples,**\n\n$$\nWER(y_{true}, y_{pred}) = \\frac{N_{insertions} + N_{substitutions} + N_{deletions}}{\\text{Length of $y_{true}$}}\n$$\n\nSuppose, $y_{true}$ is \"Do Re Mi Fa\" and $y_{pred}$ is \"Do\", WER is 3/4 = 0.75.\nBut if $y_{true}$ is \"Do\" and $y_{pred}$ is \"Do Re Mi Fa\", WER is 3/1 = 3.\n\nThe reported WER in the papers are in percentages. So, 0.4 in LB means 40. Pretty far behind actually!\n\nReferences:\n1. https://huggingface.co/openai/whisper-large-v2#evaluation\n2. https://paperswithcode.com/sota/automatic-speech-recognition-on-librispeech-7",
    "2418738": "The LB scores are in percentage. So 0.4 in LB would mean 40% WER. According to SOTA it'd be 40"
  },
  "source": "meta"
}