{
  "id": 452663,
  "title": "WHISPER - Reproducing its training",
  "url": "/competitions/bengaliai-speech/discussion/452663",
  "author_name": "yukiya",
  "post_date": "2023-11-03T03:49:08.422000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Something I find interesting : <a href=\"https://arxiv.org/pdf/2309.13876.pdf\" target=\"_blank\">https://arxiv.org/pdf/2309.13876.pdf</a></p>\n<p>\"REPRODUCING WHISPER-STYLE TRAINING USING AN OPEN-SOURCE TOOLKIT AND<br>\nPUBLICLY AVAILABLE DATA\"</p>\n<p><em>Pre-training speech models on large volumes of data has achieved\nremarkable success. OpenAI Whisper is a multilingual multitask\nmodel trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and translation benchmarks\neven in a zero-shot setup. However, the full pipeline for developing such models (from data collection to training) is not publicly\naccessible, which makes it difficult for researchers to further improve its performance and address training-related issues such as efficiency, robustness, fairness, and bias. This work presents an Open\nWhisper-style Speech Model (OWSM), which reproduces Whisperstyle training using an open-source toolkit and publicly available\ndata. OWSM even supports more translation directions and can\nbe more efficient to train. We will publicly release all scripts used\nfor data preparation, training, inference, and scoring as well as pretrained models and training logs to promote open science.</em></p>",
  "messages": [
    {
      "id": 2510495,
      "postDate": "2023-11-03T03:49:08.423Z",
      "content": "<p>Something I find interesting : <a href=\"https://arxiv.org/pdf/2309.13876.pdf\" target=\"_blank\">https://arxiv.org/pdf/2309.13876.pdf</a></p>\n<p>\"REPRODUCING WHISPER-STYLE TRAINING USING AN OPEN-SOURCE TOOLKIT AND<br>\nPUBLICLY AVAILABLE DATA\"</p>\n<p><em>Pre-training speech models on large volumes of data has achieved\nremarkable success. OpenAI Whisper is a multilingual multitask\nmodel trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and translation benchmarks\neven in a zero-shot setup. However, the full pipeline for developing such models (from data collection to training) is not publicly\naccessible, which makes it difficult for researchers to further improve its performance and address training-related issues such as efficiency, robustness, fairness, and bias. This work presents an Open\nWhisper-style Speech Model (OWSM), which reproduces Whisperstyle training using an open-source toolkit and publicly available\ndata. OWSM even supports more translation directions and can\nbe more efficient to train. We will publicly release all scripts used\nfor data preparation, training, inference, and scoring as well as pretrained models and training logs to promote open science.</em></p>",
      "rawMarkdown": "Something I find interesting : https://arxiv.org/pdf/2309.13876.pdf\n\n\"REPRODUCING WHISPER-STYLE TRAINING USING AN OPEN-SOURCE TOOLKIT AND\nPUBLICLY AVAILABLE DATA\"\n\n*Pre-training speech models on large volumes of data has achieved\nremarkable success. OpenAI Whisper is a multilingual multitask\nmodel trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and translation benchmarks\neven in a zero-shot setup. However, the full pipeline for developing such models (from data collection to training) is not publicly\naccessible, which makes it difficult for researchers to further improve its performance and address training-related issues such as efficiency, robustness, fairness, and bias. This work presents an Open\nWhisper-style Speech Model (OWSM), which reproduces Whisperstyle training using an open-source toolkit and publicly available\ndata. OWSM even supports more translation directions and can\nbe more efficient to train. We will publicly release all scripts used\nfor data preparation, training, inference, and scoring as well as pretrained models and training logs to promote open science.*",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2510495": "Something I find interesting : https://arxiv.org/pdf/2309.13876.pdf\n\n\"REPRODUCING WHISPER-STYLE TRAINING USING AN OPEN-SOURCE TOOLKIT AND\nPUBLICLY AVAILABLE DATA\"\n\n*Pre-training speech models on large volumes of data has achieved\nremarkable success. OpenAI Whisper is a multilingual multitask\nmodel trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and translation benchmarks\neven in a zero-shot setup. However, the full pipeline for developing such models (from data collection to training) is not publicly\naccessible, which makes it difficult for researchers to further improve its performance and address training-related issues such as efficiency, robustness, fairness, and bias. This work presents an Open\nWhisper-style Speech Model (OWSM), which reproduces Whisperstyle training using an open-source toolkit and publicly available\ndata. OWSM even supports more translation directions and can\nbe more efficient to train. We will publicly release all scripts used\nfor data preparation, training, inference, and scoring as well as pretrained models and training logs to promote open science.*"
  }
}