{
  "id": 425248,
  "title": "WER (Substitutions + Insertions + Deletions ). Suffering to talk Bengali (=adding data is taking too loooong)",
  "url": "/competitions/bengaliai-speech/discussion/425248",
  "author_name": "",
  "post_date": "2023-07-17T22:45:52.189437600Z",
  "votes": 28,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>WER (Word Error Rate)</h1>\n<p>\"Automatic speech recognition (ASR) technology uses machines and software to identify and process spoken language. It can also be used to authenticate a person’s identity by their voice.\" </p>\n<p>\"In the process of recognizing speech and translating it into text form, some words may be left out or mistranslated. If you have used ASR in some capacity, you probably have encountered the phrase “word error rate” (WER). \"</p>\n<p>\"To get the WER, start by adding up the substitutions, insertions, and deletions that occur in a sequence of recognized words. Divide that number by the total number of words originally spoken. The result is the WER.\"</p>\n<p>\"To put it in a simple formula, Word Error Rate = (Substitutions + Insertions + Deletions) / Number of Words Spoken\"</p>\n<p>Let’s look at each one:</p>\n<ul>\n<li>\"A substitution occurs when a word gets replaced (for example, “noose” is transcribed as “moose”)\"</li>\n<li>\"An insertion is when a word is added that wasn’t said (for example, “SAT” becomes “essay tea”)\"</li>\n<li>\"A deletion happens when a word is left out of the transcript completely (for example, “turn it around” becomes “turn around”)\"</li>\n</ul>\n<p>\"Let’s say that a person speaks 29 total words in an original transcription file. Among those words spoken, the transcription included 11 substitutions, insertions, and deletions.\" </p>\n<h1>Where did the word error rate calculation come from? Levenshtein distance.</h1>\n<p>\"The WER calculation is based on a measurement called the “Levenshtein distance.” The Levenshtein distance is a measurement of the differences between two “strings.” In this case, the strings are sequences of letters that make up the words in a transcription.\"</p>\n<h1>Why does word error rate matter?</h1>\n<p>\"WER is an important, common metric used to measure the performance of the speech recognition APIs used to power interactive voice-based technology, like Siri or the Amazon Echo.\"</p>\n<p>\"Lower WER often indicates that the ASR software is more accurate in recognizing speech. A higher WER, then, often indicates lower ASR accuracy.\"</p>\n<p><a href=\"https://www.rev.com/blog/resources/what-is-wer-what-does-word-error-rate-mean\" target=\"_blank\">https://www.rev.com/blog/resources/what-is-wer-what-does-word-error-rate-mean</a></p>\n<h1>Suffering to add data</h1>\n<p>I've just tried to intall the libraries. Adding the Bengali.Ai data is taking a lot of time (more than 20 min.)<br>\nIs there a short cut to perform that?</p>",
  "messages": [
    {
      "id": "2348720",
      "postDate": "07/17/2023 22:45:52",
      "content": "<h1>WER (Word Error Rate)</h1>\n<p>\"Automatic speech recognition (ASR) technology uses machines and software to identify and process spoken language. It can also be used to authenticate a person’s identity by their voice.\" </p>\n<p>\"In the process of recognizing speech and translating it into text form, some words may be left out or mistranslated. If you have used ASR in some capacity, you probably have encountered the phrase “word error rate” (WER). \"</p>\n<p>\"To get the WER, start by adding up the substitutions, insertions, and deletions that occur in a sequence of recognized words. Divide that number by the total number of words originally spoken. The result is the WER.\"</p>\n<p>\"To put it in a simple formula, Word Error Rate = (Substitutions + Insertions + Deletions) / Number of Words Spoken\"</p>\n<p>Let’s look at each one:</p>\n<ul>\n<li>\"A substitution occurs when a word gets replaced (for example, “noose” is transcribed as “moose”)\"</li>\n<li>\"An insertion is when a word is added that wasn’t said (for example, “SAT” becomes “essay tea”)\"</li>\n<li>\"A deletion happens when a word is left out of the transcript completely (for example, “turn it around” becomes “turn around”)\"</li>\n</ul>\n<p>\"Let’s say that a person speaks 29 total words in an original transcription file. Among those words spoken, the transcription included 11 substitutions, insertions, and deletions.\" </p>\n<h1>Where did the word error rate calculation come from? Levenshtein distance.</h1>\n<p>\"The WER calculation is based on a measurement called the “Levenshtein distance.” The Levenshtein distance is a measurement of the differences between two “strings.” In this case, the strings are sequences of letters that make up the words in a transcription.\"</p>\n<h1>Why does word error rate matter?</h1>\n<p>\"WER is an important, common metric used to measure the performance of the speech recognition APIs used to power interactive voice-based technology, like Siri or the Amazon Echo.\"</p>\n<p>\"Lower WER often indicates that the ASR software is more accurate in recognizing speech. A higher WER, then, often indicates lower ASR accuracy.\"</p>\n<p><a href=\"https://www.rev.com/blog/resources/what-is-wer-what-does-word-error-rate-mean\" target=\"_blank\">https://www.rev.com/blog/resources/what-is-wer-what-does-word-error-rate-mean</a></p>\n<h1>Suffering to add data</h1>\n<p>I've just tried to intall the libraries. Adding the Bengali.Ai data is taking a lot of time (more than 20 min.)<br>\nIs there a short cut to perform that?</p>",
      "rawMarkdown": "#WER (Word Error Rate)\n\n\"Automatic speech recognition (ASR) technology uses machines and software to identify and process spoken language. It can also be used to authenticate a person’s identity by their voice.\" \n\n\"In the process of recognizing speech and translating it into text form, some words may be left out or mistranslated. If you have used ASR in some capacity, you probably have encountered the phrase “word error rate” (WER). \"\n\n\"To get the WER, start by adding up the substitutions, insertions, and deletions that occur in a sequence of recognized words. Divide that number by the total number of words originally spoken. The result is the WER.\"\n\n\"To put it in a simple formula, Word Error Rate = (Substitutions + Insertions + Deletions) / Number of Words Spoken\"\n\nLet’s look at each one:\n\n- \"A substitution occurs when a word gets replaced (for example, “noose” is transcribed as “moose”)\"\n- \"An insertion is when a word is added that wasn’t said (for example, “SAT” becomes “essay tea”)\"\n- \"A deletion happens when a word is left out of the transcript completely (for example, “turn it around” becomes “turn around”)\"\n \n\n\"Let’s say that a person speaks 29 total words in an original transcription file. Among those words spoken, the transcription included 11 substitutions, insertions, and deletions.\" \n\n#Where did the word error rate calculation come from? Levenshtein distance.\n \n\"The WER calculation is based on a measurement called the “Levenshtein distance.” The Levenshtein distance is a measurement of the differences between two “strings.” In this case, the strings are sequences of letters that make up the words in a transcription.\"\n\n#Why does word error rate matter?\n\n\"WER is an important, common metric used to measure the performance of the speech recognition APIs used to power interactive voice-based technology, like Siri or the Amazon Echo.\"\n\n\"Lower WER often indicates that the ASR software is more accurate in recognizing speech. A higher WER, then, often indicates lower ASR accuracy.\"\n\nhttps://www.rev.com/blog/resources/what-is-wer-what-does-word-error-rate-mean\n\n#Suffering to add data\n\nI've just tried to intall the libraries. Adding the Bengali.Ai data is taking a lot of time (more than 20 min.)\nIs there a short cut to perform that?",
      "votes": null
    },
    {
      "id": "2348722",
      "postDate": "07/17/2023 22:47:58",
      "content": "<p>Yep, it's taking wayyyyyyyy too long to add the data</p>",
      "rawMarkdown": "Yep, it's taking wayyyyyyyy too long to add the data",
      "votes": null
    },
    {
      "id": "2348723",
      "postDate": "07/17/2023 22:50:01",
      "content": "<p>Finally loaded :)</p>",
      "rawMarkdown": "Finally loaded :)",
      "votes": null
    },
    {
      "id": "2348727",
      "postDate": "07/17/2023 22:55:55",
      "content": "<p>More than 30 min. On the other old Notebook is stiiiiiiiiLLLLL running or walking : ) </p>",
      "rawMarkdown": "More than 30 min. On the other old Notebook is stiiiiiiiiLLLLL running or walking : )",
      "votes": null
    },
    {
      "id": "2348728",
      "postDate": "07/17/2023 22:58:23",
      "content": "<p>Crawling !!!</p>",
      "rawMarkdown": "Crawling !!!",
      "votes": null
    },
    {
      "id": "2349289",
      "postDate": "07/18/2023 09:37:25",
      "content": "<p>Thank for sharing this topics,it is new for me </p>",
      "rawMarkdown": "Thank for sharing this topics,it is new for me",
      "votes": null
    },
    {
      "id": "2403229",
      "postDate": "08/22/2023 15:10:46",
      "content": "<p>Thank for sharing. </p>",
      "rawMarkdown": "Thank for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2348722,
      "author_name": "abhranta",
      "author_url": "",
      "post_date": "07/17/2023 22:47:58",
      "content": "<p>Yep, it's taking wayyyyyyyy too long to add the data</p>",
      "votes": null,
      "replies": [
        {
          "id": 2348727,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "07/17/2023 22:55:55",
          "content": "<p>More than 30 min. On the other old Notebook is stiiiiiiiiLLLLL running or walking : ) </p>",
          "votes": null,
          "replies": [
            {
              "id": 2348728,
              "author_name": "abhranta",
              "author_url": "",
              "post_date": "07/17/2023 22:58:23",
              "content": "<p>Crawling !!!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2348723,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "07/17/2023 22:50:01",
      "content": "<p>Finally loaded :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2349289,
      "author_name": "mohimaakter",
      "author_url": "",
      "post_date": "07/18/2023 09:37:25",
      "content": "<p>Thank for sharing this topics,it is new for me </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2403229,
      "author_name": "berserker408",
      "author_url": "",
      "post_date": "08/22/2023 15:10:46",
      "content": "<p>Thank for sharing. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2348720": "#WER (Word Error Rate)\n\n\"Automatic speech recognition (ASR) technology uses machines and software to identify and process spoken language. It can also be used to authenticate a person’s identity by their voice.\" \n\n\"In the process of recognizing speech and translating it into text form, some words may be left out or mistranslated. If you have used ASR in some capacity, you probably have encountered the phrase “word error rate” (WER). \"\n\n\"To get the WER, start by adding up the substitutions, insertions, and deletions that occur in a sequence of recognized words. Divide that number by the total number of words originally spoken. The result is the WER.\"\n\n\"To put it in a simple formula, Word Error Rate = (Substitutions + Insertions + Deletions) / Number of Words Spoken\"\n\nLet’s look at each one:\n\n- \"A substitution occurs when a word gets replaced (for example, “noose” is transcribed as “moose”)\"\n- \"An insertion is when a word is added that wasn’t said (for example, “SAT” becomes “essay tea”)\"\n- \"A deletion happens when a word is left out of the transcript completely (for example, “turn it around” becomes “turn around”)\"\n \n\n\"Let’s say that a person speaks 29 total words in an original transcription file. Among those words spoken, the transcription included 11 substitutions, insertions, and deletions.\" \n\n#Where did the word error rate calculation come from? Levenshtein distance.\n \n\"The WER calculation is based on a measurement called the “Levenshtein distance.” The Levenshtein distance is a measurement of the differences between two “strings.” In this case, the strings are sequences of letters that make up the words in a transcription.\"\n\n#Why does word error rate matter?\n\n\"WER is an important, common metric used to measure the performance of the speech recognition APIs used to power interactive voice-based technology, like Siri or the Amazon Echo.\"\n\n\"Lower WER often indicates that the ASR software is more accurate in recognizing speech. A higher WER, then, often indicates lower ASR accuracy.\"\n\nhttps://www.rev.com/blog/resources/what-is-wer-what-does-word-error-rate-mean\n\n#Suffering to add data\n\nI've just tried to intall the libraries. Adding the Bengali.Ai data is taking a lot of time (more than 20 min.)\nIs there a short cut to perform that?",
    "2348722": "Yep, it's taking wayyyyyyyy too long to add the data",
    "2348723": "Finally loaded :)",
    "2348727": "More than 30 min. On the other old Notebook is stiiiiiiiiLLLLL running or walking : )",
    "2348728": "Crawling !!!",
    "2349289": "Thank for sharing this topics,it is new for me",
    "2403229": "Thank for sharing."
  },
  "source": "meta"
}