{
  "id": 425244,
  "title": "Papers for Bengali Speech Recognition",
  "url": "/competitions/bengaliai-speech/discussion/425244",
  "author_name": "",
  "post_date": "2023-07-17T22:21:09.698779Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Competition Paper:</strong></p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2305.09688\" target=\"_blank\">OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking</a></li>\n</ul>\n<blockquote>\n  <p>We present OOD-Speech, the first out-of-distribution (OOD) benchmarking dataset for Bengali automatic speech recognition (ASR). Being one of the most spoken languages globally, Bengali portrays large diversity in dialects and prosodic features, which demands ASR frameworks to be robust towards distribution shifts. For example, islamic religious sermons in Bengali are delivered with a tonality that is significantly different from regular speech. Our training dataset is collected via massively online crowdsourcing campaigns which resulted in 1177.94 hours collected and curated from 22,645 native Bengali speakers from South Asia. Our test dataset comprises 23.03 hours of speech collected and manually annotated from 17 different sources, e.g., Bengali TV drama, Audiobook, Talk show, Online class, and Islamic sermons to name a few. OOD-Speech is jointly the largest publicly available speech dataset, as well as the first out-of-distribution ASR benchmarking dataset for Bengali.</p>\n</blockquote>\n<p><a href=\"https://huggingface.co/bengaliAI/BanglaConformer\" target=\"_blank\"><strong>BanglaConformer Model</strong></a><br>\n <a href=\"https://huggingface.co/datasets/bengaliAI/CommonVoiceBangla\" target=\"_blank\"><strong>External Data</strong></a></p>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/2209.08119.pdf\" target=\"_blank\">An Automatic Speech Recognition System for\nBengali Language based on Wav2Vec2 and Transfer\nLearning</a></li>\n</ul>\n<blockquote>\n  <p>An independent, automated method of decoding and transcribing oral speech is known as automatic speech recognition (ASR). A typical ASR system extracts feature from audio recordings or streams and run one or more algorithms to map the features to corresponding texts. Numerous of research has been done in the field of speech signal processing in recent years. When given adequate resources, both conventional ASR and emerging end-to-end (E2E) speech recognition have produced promising results. However, for low-resource languages like Bengali, the current state of ASR lags behind, although the low resource state does not reflect upon the fact that this language is spoken by over 500 million people all over the world. Despite its popularity, there aren’t many diverse open-source datasets available, which makes it difficult to conduct research on Bengali speech recognition systems. This paper is a part of the competition named ‘BUET CSE Fest DL Sprint’. The purpose of this paper is to improve the speech recognition performance of the Bengali language by adopting speech recognition technology on the E2E structure based on the transfer learning framework. The proposed method effectively models the Bengali language and achieves 3.819 score in ‘Levenshtein Mean Distance’ on the test dataset of 7747 samples, when only 1000 samples of train dataset were used to train</p>\n</blockquote>\n<p>*<a href=\"https://ieeexplore.ieee.org/document/8821629\" target=\"_blank\">A Speech Recognition System for Bengali Language using Recurrent Neural Network</a></p>\n<blockquote>\n  <p>Speech recognition is the most interactive technology between a human and a machine. Over the past 70 years, tremendous work has been accomplished in this fundamental area of speech communication. However, implementation of the Bengali language is unsubstantial in the field of Human-Computer Interaction. This research paper tried to implement convolution neural network technique for creating a speech recognition system in the Bengali language. We also implemented recurrent neural network to find the Bengali character probabilities which were then improved further by using CTC loss function and language model. This paper implemented Bengali language consisting of diacritic characters and such languages are very much difficult to train in a model.[</p>\n</blockquote>",
  "messages": [
    {
      "id": "2348713",
      "postDate": "07/17/2023 22:21:09",
      "content": "<p><strong>Competition Paper:</strong></p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2305.09688\" target=\"_blank\">OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking</a></li>\n</ul>\n<blockquote>\n  <p>We present OOD-Speech, the first out-of-distribution (OOD) benchmarking dataset for Bengali automatic speech recognition (ASR). Being one of the most spoken languages globally, Bengali portrays large diversity in dialects and prosodic features, which demands ASR frameworks to be robust towards distribution shifts. For example, islamic religious sermons in Bengali are delivered with a tonality that is significantly different from regular speech. Our training dataset is collected via massively online crowdsourcing campaigns which resulted in 1177.94 hours collected and curated from 22,645 native Bengali speakers from South Asia. Our test dataset comprises 23.03 hours of speech collected and manually annotated from 17 different sources, e.g., Bengali TV drama, Audiobook, Talk show, Online class, and Islamic sermons to name a few. OOD-Speech is jointly the largest publicly available speech dataset, as well as the first out-of-distribution ASR benchmarking dataset for Bengali.</p>\n</blockquote>\n<p><a href=\"https://huggingface.co/bengaliAI/BanglaConformer\" target=\"_blank\"><strong>BanglaConformer Model</strong></a><br>\n <a href=\"https://huggingface.co/datasets/bengaliAI/CommonVoiceBangla\" target=\"_blank\"><strong>External Data</strong></a></p>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/2209.08119.pdf\" target=\"_blank\">An Automatic Speech Recognition System for\nBengali Language based on Wav2Vec2 and Transfer\nLearning</a></li>\n</ul>\n<blockquote>\n  <p>An independent, automated method of decoding and transcribing oral speech is known as automatic speech recognition (ASR). A typical ASR system extracts feature from audio recordings or streams and run one or more algorithms to map the features to corresponding texts. Numerous of research has been done in the field of speech signal processing in recent years. When given adequate resources, both conventional ASR and emerging end-to-end (E2E) speech recognition have produced promising results. However, for low-resource languages like Bengali, the current state of ASR lags behind, although the low resource state does not reflect upon the fact that this language is spoken by over 500 million people all over the world. Despite its popularity, there aren’t many diverse open-source datasets available, which makes it difficult to conduct research on Bengali speech recognition systems. This paper is a part of the competition named ‘BUET CSE Fest DL Sprint’. The purpose of this paper is to improve the speech recognition performance of the Bengali language by adopting speech recognition technology on the E2E structure based on the transfer learning framework. The proposed method effectively models the Bengali language and achieves 3.819 score in ‘Levenshtein Mean Distance’ on the test dataset of 7747 samples, when only 1000 samples of train dataset were used to train</p>\n</blockquote>\n<p>*<a href=\"https://ieeexplore.ieee.org/document/8821629\" target=\"_blank\">A Speech Recognition System for Bengali Language using Recurrent Neural Network</a></p>\n<blockquote>\n  <p>Speech recognition is the most interactive technology between a human and a machine. Over the past 70 years, tremendous work has been accomplished in this fundamental area of speech communication. However, implementation of the Bengali language is unsubstantial in the field of Human-Computer Interaction. This research paper tried to implement convolution neural network technique for creating a speech recognition system in the Bengali language. We also implemented recurrent neural network to find the Bengali character probabilities which were then improved further by using CTC loss function and language model. This paper implemented Bengali language consisting of diacritic characters and such languages are very much difficult to train in a model.[</p>\n</blockquote>",
      "rawMarkdown": "**Competition Paper:**\n* [OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking](https://arxiv.org/abs/2305.09688)\n> We present OOD-Speech, the first out-of-distribution (OOD) benchmarking dataset for Bengali automatic speech recognition (ASR). Being one of the most spoken languages globally, Bengali portrays large diversity in dialects and prosodic features, which demands ASR frameworks to be robust towards distribution shifts. For example, islamic religious sermons in Bengali are delivered with a tonality that is significantly different from regular speech. Our training dataset is collected via massively online crowdsourcing campaigns which resulted in 1177.94 hours collected and curated from 22,645 native Bengali speakers from South Asia. Our test dataset comprises 23.03 hours of speech collected and manually annotated from 17 different sources, e.g., Bengali TV drama, Audiobook, Talk show, Online class, and Islamic sermons to name a few. OOD-Speech is jointly the largest publicly available speech dataset, as well as the first out-of-distribution ASR benchmarking dataset for Bengali.\n\n [**BanglaConformer Model**](https://huggingface.co/bengaliAI/BanglaConformer)\n [**External Data**](https://huggingface.co/datasets/bengaliAI/CommonVoiceBangla)\n\n* [An Automatic Speech Recognition System for\nBengali Language based on Wav2Vec2 and Transfer\nLearning](https://arxiv.org/pdf/2209.08119.pdf)\n> An independent, automated method of decoding and transcribing oral speech is known as automatic speech recognition (ASR). A typical ASR system extracts feature from audio recordings or streams and run one or more algorithms to map the features to corresponding texts. Numerous of research has been done in the field of speech signal processing in recent years. When given adequate resources, both conventional ASR and emerging end-to-end (E2E) speech recognition have produced promising results. However, for low-resource languages like Bengali, the current state of ASR lags behind, although the low resource state does not reflect upon the fact that this language is spoken by over 500 million people all over the world. Despite its popularity, there aren’t many diverse open-source datasets available, which makes it difficult to conduct research on Bengali speech recognition systems. This paper is a part of the competition named ‘BUET CSE Fest DL Sprint’. The purpose of this paper is to improve the speech recognition performance of the Bengali language by adopting speech recognition technology on the E2E structure based on the transfer learning framework. The proposed method effectively models the Bengali language and achieves 3.819 score in ‘Levenshtein Mean Distance’ on the test dataset of 7747 samples, when only 1000 samples of train dataset were used to train\n\n*[A Speech Recognition System for Bengali Language using Recurrent Neural Network](https://ieeexplore.ieee.org/document/8821629)\n> Speech recognition is the most interactive technology between a human and a machine. Over the past 70 years, tremendous work has been accomplished in this fundamental area of speech communication. However, implementation of the Bengali language is unsubstantial in the field of Human-Computer Interaction. This research paper tried to implement convolution neural network technique for creating a speech recognition system in the Bengali language. We also implemented recurrent neural network to find the Bengali character probabilities which were then improved further by using CTC loss function and language model. This paper implemented Bengali language consisting of diacritic characters and such languages are very much difficult to train in a model.[",
      "votes": null
    },
    {
      "id": "2349238",
      "postDate": "07/18/2023 08:50:53",
      "content": "<p>check the dataset paper and the baseline methods (e.g. whisper, confromer, wave2vec)</p>",
      "rawMarkdown": "check the dataset paper and the baseline methods (e.g. whisper, confromer, wave2vec)",
      "votes": null
    },
    {
      "id": "2359188",
      "postDate": "07/26/2023 03:51:00",
      "content": "<p>Thank you for this compilation</p>",
      "rawMarkdown": "Thank you for this compilation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2349238,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/18/2023 08:50:53",
      "content": "<p>check the dataset paper and the baseline methods (e.g. whisper, confromer, wave2vec)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2359188,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "07/26/2023 03:51:00",
      "content": "<p>Thank you for this compilation</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2348713": "**Competition Paper:**\n* [OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking](https://arxiv.org/abs/2305.09688)\n> We present OOD-Speech, the first out-of-distribution (OOD) benchmarking dataset for Bengali automatic speech recognition (ASR). Being one of the most spoken languages globally, Bengali portrays large diversity in dialects and prosodic features, which demands ASR frameworks to be robust towards distribution shifts. For example, islamic religious sermons in Bengali are delivered with a tonality that is significantly different from regular speech. Our training dataset is collected via massively online crowdsourcing campaigns which resulted in 1177.94 hours collected and curated from 22,645 native Bengali speakers from South Asia. Our test dataset comprises 23.03 hours of speech collected and manually annotated from 17 different sources, e.g., Bengali TV drama, Audiobook, Talk show, Online class, and Islamic sermons to name a few. OOD-Speech is jointly the largest publicly available speech dataset, as well as the first out-of-distribution ASR benchmarking dataset for Bengali.\n\n [**BanglaConformer Model**](https://huggingface.co/bengaliAI/BanglaConformer)\n [**External Data**](https://huggingface.co/datasets/bengaliAI/CommonVoiceBangla)\n\n* [An Automatic Speech Recognition System for\nBengali Language based on Wav2Vec2 and Transfer\nLearning](https://arxiv.org/pdf/2209.08119.pdf)\n> An independent, automated method of decoding and transcribing oral speech is known as automatic speech recognition (ASR). A typical ASR system extracts feature from audio recordings or streams and run one or more algorithms to map the features to corresponding texts. Numerous of research has been done in the field of speech signal processing in recent years. When given adequate resources, both conventional ASR and emerging end-to-end (E2E) speech recognition have produced promising results. However, for low-resource languages like Bengali, the current state of ASR lags behind, although the low resource state does not reflect upon the fact that this language is spoken by over 500 million people all over the world. Despite its popularity, there aren’t many diverse open-source datasets available, which makes it difficult to conduct research on Bengali speech recognition systems. This paper is a part of the competition named ‘BUET CSE Fest DL Sprint’. The purpose of this paper is to improve the speech recognition performance of the Bengali language by adopting speech recognition technology on the E2E structure based on the transfer learning framework. The proposed method effectively models the Bengali language and achieves 3.819 score in ‘Levenshtein Mean Distance’ on the test dataset of 7747 samples, when only 1000 samples of train dataset were used to train\n\n*[A Speech Recognition System for Bengali Language using Recurrent Neural Network](https://ieeexplore.ieee.org/document/8821629)\n> Speech recognition is the most interactive technology between a human and a machine. Over the past 70 years, tremendous work has been accomplished in this fundamental area of speech communication. However, implementation of the Bengali language is unsubstantial in the field of Human-Computer Interaction. This research paper tried to implement convolution neural network technique for creating a speech recognition system in the Bengali language. We also implemented recurrent neural network to find the Bengali character probabilities which were then improved further by using CTC loss function and language model. This paper implemented Bengali language consisting of diacritic characters and such languages are very much difficult to train in a model.[",
    "2349238": "check the dataset paper and the baseline methods (e.g. whisper, confromer, wave2vec)",
    "2359188": "Thank you for this compilation"
  },
  "source": "meta"
}