{
  "id": 433835,
  "title": "MetaAI SeamlessM4T",
  "url": "/competitions/bengaliai-speech/discussion/433835",
  "author_name": "tugstugi",
  "post_date": "2023-08-23T03:55:49.744000",
  "votes": 11,
  "comment_count": 13,
  "views": 0,
  "content": "<p>A new MetaAI MNT/ASR/TTS model:<br>\nHuggingface: <a href=\"https://huggingface.co/facebook/seamless-m4t-large\" target=\"_blank\">https://huggingface.co/facebook/seamless-m4t-large</a><br>\nMetaAI blog: <a href=\"https://ai.meta.com/research/publications/seamless-m4t/\" target=\"_blank\">https://ai.meta.com/research/publications/seamless-m4t/</a></p>\n<p><strong>SeamlessM4T—Massively Multilingual &amp; Multimodal Machine Translation</strong></p>\n<p><strong>Abstract</strong></p>\n<p>What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200 languages, unified speech-to-speech translation models have yet to achieve similar strides. More specifically, conventional speech-to-speech translation systems rely on cascaded systems composed of multiple subsystems performing translation progressively, putting scalable and high-performing unified speech translation systems out of reach. To address these gaps, we introduce SeamlessM4T—Massively Multilingual &amp; Multimodal Machine Translation—a single model that supports speech-to-speech translation, speech-to-text translation, text-to-speech translation, text-to-text translation, and automatic speech recognition for up to 100 languages. To build this, we used 1 million hours of open speech audio data to learn self-supervised speech representations with w2v-BERT 2.0. Subsequently, we created a multimodal corpus of automatically aligned speech translations, dubbed SeamlessAlign. Filtered and combined with human labeled and pseudo-labeled data (totaling 406,000 hours), we developed the first multilingual system capable of translating from and into English for both speech and text. On Fleurs, SeamlessM4T sets a new standard for translations into multiple target languages, achieving an improvement of 20% BLEU over the previous state-of-the-art in direct speech-to-text translation. Compared to strong cascaded models, SeamlessM4T improves the quality of into-English translation by 1.3 BLEU points in speech-to-text and by 2.6 ASR-BLEU points in speech-to-speech. On CVSS and compared to a 2-stage cascaded model for speech-to-speech translation, SeamlessM4T-Large’s performance is stronger by 58%. Preliminary human evaluations of speech-to-text translation outputs evinced similarly impressive results; for translations from English, XSTS scores for 24 evaluated languages are consistently above 4 (out of 5). For into English directions, we see significant improvement over WhisperLarge-v2’s baseline for 7 out of 24 languages. To further evaluate our system, we developed Blaser 2.0, which enables evaluation across speech and text with similar accuracy compared to its predecessor when it comes to quality estimation. Tested for robustness, our system performs better against background noises and speaker variations in speech-to-text tasks (average improvements of 38% and 49%, respectively) compared to the current state-of-the-art model. Critically, we evaluated SeamlessM4T on gender bias and added toxicity to assess translation safety. Compared to the state-of-the-art, we report up to 63% of reduction in added toxicity in our translation outputs. Finally, all contributions in this work—including models, inference code, finetuning recipes backed by our improved modeling toolkit Fairseq2, and metadata to recreate the unfiltered 470,000 hours of SeamlessAlign — are open-sourced and accessible at <a href=\"https://github.com/facebookresearch/seamless_communication\" target=\"_blank\">https://github.com/facebookresearch/seamless_communication</a>.</p>",
  "messages": [
    {
      "id": 2404094,
      "postDate": "2023-08-23T03:55:49.743Z",
      "content": "<p>A new MetaAI MNT/ASR/TTS model:<br>\nHuggingface: <a href=\"https://huggingface.co/facebook/seamless-m4t-large\" target=\"_blank\">https://huggingface.co/facebook/seamless-m4t-large</a><br>\nMetaAI blog: <a href=\"https://ai.meta.com/research/publications/seamless-m4t/\" target=\"_blank\">https://ai.meta.com/research/publications/seamless-m4t/</a></p>\n<p><strong>SeamlessM4T—Massively Multilingual &amp; Multimodal Machine Translation</strong></p>\n<p><strong>Abstract</strong></p>\n<p>What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200 languages, unified speech-to-speech translation models have yet to achieve similar strides. More specifically, conventional speech-to-speech translation systems rely on cascaded systems composed of multiple subsystems performing translation progressively, putting scalable and high-performing unified speech translation systems out of reach. To address these gaps, we introduce SeamlessM4T—Massively Multilingual &amp; Multimodal Machine Translation—a single model that supports speech-to-speech translation, speech-to-text translation, text-to-speech translation, text-to-text translation, and automatic speech recognition for up to 100 languages. To build this, we used 1 million hours of open speech audio data to learn self-supervised speech representations with w2v-BERT 2.0. Subsequently, we created a multimodal corpus of automatically aligned speech translations, dubbed SeamlessAlign. Filtered and combined with human labeled and pseudo-labeled data (totaling 406,000 hours), we developed the first multilingual system capable of translating from and into English for both speech and text. On Fleurs, SeamlessM4T sets a new standard for translations into multiple target languages, achieving an improvement of 20% BLEU over the previous state-of-the-art in direct speech-to-text translation. Compared to strong cascaded models, SeamlessM4T improves the quality of into-English translation by 1.3 BLEU points in speech-to-text and by 2.6 ASR-BLEU points in speech-to-speech. On CVSS and compared to a 2-stage cascaded model for speech-to-speech translation, SeamlessM4T-Large’s performance is stronger by 58%. Preliminary human evaluations of speech-to-text translation outputs evinced similarly impressive results; for translations from English, XSTS scores for 24 evaluated languages are consistently above 4 (out of 5). For into English directions, we see significant improvement over WhisperLarge-v2’s baseline for 7 out of 24 languages. To further evaluate our system, we developed Blaser 2.0, which enables evaluation across speech and text with similar accuracy compared to its predecessor when it comes to quality estimation. Tested for robustness, our system performs better against background noises and speaker variations in speech-to-text tasks (average improvements of 38% and 49%, respectively) compared to the current state-of-the-art model. Critically, we evaluated SeamlessM4T on gender bias and added toxicity to assess translation safety. Compared to the state-of-the-art, we report up to 63% of reduction in added toxicity in our translation outputs. Finally, all contributions in this work—including models, inference code, finetuning recipes backed by our improved modeling toolkit Fairseq2, and metadata to recreate the unfiltered 470,000 hours of SeamlessAlign — are open-sourced and accessible at <a href=\"https://github.com/facebookresearch/seamless_communication\" target=\"_blank\">https://github.com/facebookresearch/seamless_communication</a>.</p>",
      "rawMarkdown": "A new MetaAI MNT/ASR/TTS model:\nHuggingface: https://huggingface.co/facebook/seamless-m4t-large\nMetaAI blog: https://ai.meta.com/research/publications/seamless-m4t/\n\n\n\n**SeamlessM4T—Massively Multilingual & Multimodal Machine Translation**\n\n**Abstract**\n\nWhat does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200 languages, unified speech-to-speech translation models have yet to achieve similar strides. More specifically, conventional speech-to-speech translation systems rely on cascaded systems composed of multiple subsystems performing translation progressively, putting scalable and high-performing unified speech translation systems out of reach. To address these gaps, we introduce SeamlessM4T—Massively Multilingual & Multimodal Machine Translation—a single model that supports speech-to-speech translation, speech-to-text translation, text-to-speech translation, text-to-text translation, and automatic speech recognition for up to 100 languages. To build this, we used 1 million hours of open speech audio data to learn self-supervised speech representations with w2v-BERT 2.0. Subsequently, we created a multimodal corpus of automatically aligned speech translations, dubbed SeamlessAlign. Filtered and combined with human labeled and pseudo-labeled data (totaling 406,000 hours), we developed the first multilingual system capable of translating from and into English for both speech and text. On Fleurs, SeamlessM4T sets a new standard for translations into multiple target languages, achieving an improvement of 20% BLEU over the previous state-of-the-art in direct speech-to-text translation. Compared to strong cascaded models, SeamlessM4T improves the quality of into-English translation by 1.3 BLEU points in speech-to-text and by 2.6 ASR-BLEU points in speech-to-speech. On CVSS and compared to a 2-stage cascaded model for speech-to-speech translation, SeamlessM4T-Large’s performance is stronger by 58%. Preliminary human evaluations of speech-to-text translation outputs evinced similarly impressive results; for translations from English, XSTS scores for 24 evaluated languages are consistently above 4 (out of 5). For into English directions, we see significant improvement over WhisperLarge-v2’s baseline for 7 out of 24 languages. To further evaluate our system, we developed Blaser 2.0, which enables evaluation across speech and text with similar accuracy compared to its predecessor when it comes to quality estimation. Tested for robustness, our system performs better against background noises and speaker variations in speech-to-text tasks (average improvements of 38% and 49%, respectively) compared to the current state-of-the-art model. Critically, we evaluated SeamlessM4T on gender bias and added toxicity to assess translation safety. Compared to the state-of-the-art, we report up to 63% of reduction in added toxicity in our translation outputs. Finally, all contributions in this work—including models, inference code, finetuning recipes backed by our improved modeling toolkit Fairseq2, and metadata to recreate the unfiltered 470,000 hours of SeamlessAlign — are open-sourced and accessible at https://github.com/facebookresearch/seamless_communication.",
      "votes": 11
    },
    {
      "id": 2412869,
      "postDate": "2023-08-28T14:49:58.200Z",
      "content": "<p>Anybody attempting distillation with this model ? </p>",
      "rawMarkdown": "Anybody attempting distillation with this model ? ",
      "votes": 1,
      "replies": [
        {
          "id": 2412945,
          "postDate": "2023-08-28T15:42:39.590Z",
          "content": "<p>Distillation would be a great way to reduce parameter a and make the model smaller indeed. You can use pruning or quantization to reduce the model!</p>",
          "rawMarkdown": "Distillation would be a great way to reduce parameter a and make the model smaller indeed. You can use pruning or quantization to reduce the model!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2405795,
      "postDate": "2023-08-24T05:44:59.737Z",
      "content": "<p>Isn't this a translation model? How is that usable here?</p>",
      "rawMarkdown": "Isn't this a translation model? How is that usable here?\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 2406461,
          "postDate": "2023-08-24T12:59:31.643Z",
          "content": "<p>Look at the paper and github! The model has S2TT, T2TT and <strong>ASR</strong> capabilities. The paper includes Bengali as one of the languages, so it's interesting to see how it would perform with the extra data mined and its massive amount of parameters (2.3B for large, 1.2B for medium).</p>",
          "rawMarkdown": "Look at the paper and github! The model has S2TT, T2TT and **ASR** capabilities. The paper includes Bengali as one of the languages, so it's interesting to see how it would perform with the extra data mined and its massive amount of parameters (2.3B for large, 1.2B for medium)."
        }
      ]
    },
    {
      "id": 2404430,
      "postDate": "2023-08-23T08:50:35.920Z",
      "content": "<p>1.4 Billion parameters…. that's huge, compared to wav2vec2(~100M), I don't think it will be easy to use SeamlessM4T in this competition. Yet, really curious how well they actually are on all these tasks. </p>",
      "rawMarkdown": "1.4 Billion parameters.... that's huge, compared to wav2vec2(~100M), I don't think it will be easy to use SeamlessM4T in this competition. Yet, really curious how well they actually are on all these tasks. ",
      "votes": 2,
      "replies": [
        {
          "id": 2412946,
          "postDate": "2023-08-28T15:43:34.713Z",
          "content": "<p>I found when submit wav2vec2 model it takes lot' of time to infer!</p>",
          "rawMarkdown": "I found when submit wav2vec2 model it takes lot' of time to infer!"
        }
      ]
    },
    {
      "id": 2421863,
      "postDate": "2023-09-03T14:43:01.683Z",
      "content": "<p>Hi! I would like to try this model for inference. <br>\nand during the inference, the internet is off, so I wanted to load the model from local like this</p>\n<pre><code>translator = Translator(\n    ,\n    ,\n    torch.device()\n)\n</code></pre>\n<p>however, I got an error:<br>\nValueError: <code>name</code> must be a valid filename, but is '/kaggle/input/meta-ai-seamlessm4t-large/seamless-m4t-large' instead.</p>\n<p>Do anyone know how to fix this?</p>",
      "rawMarkdown": "Hi! I would like to try this model for inference. \nand during the inference, the internet is off, so I wanted to load the model from local like this\n```\ntranslator = Translator(\n    \"/kaggle/input/meta-ai-seamlessm4t-large/seamless-m4t-large\",\n    \"vocoder_36langs\",\n    torch.device(\"cuda:0\")\n)\n```\nhowever, I got an error:\nValueError: `name` must be a valid filename, but is '/kaggle/input/meta-ai-seamlessm4t-large/seamless-m4t-large' instead.\n\nDo anyone know how to fix this?",
      "replies": [
        {
          "id": 2421972,
          "postDate": "2023-09-03T16:14:45.363Z",
          "content": "<p>I've already tested the medium model, but it gave poor results, I'll publish the notebook later.</p>",
          "rawMarkdown": "I've already tested the medium model, but it gave poor results, I'll publish the notebook later.",
          "votes": 1,
          "replies": [
            {
              "id": 2422802,
              "postDate": "2023-09-04T08:56:57.030Z",
              "content": "<p>Thanks! Looking forward to your notebook.</p>",
              "rawMarkdown": "Thanks! Looking forward to your notebook."
            }
          ]
        },
        {
          "id": 2422899,
          "postDate": "2023-09-04T10:28:07.447Z",
          "content": "<p>I release a notebook <a href=\"https://www.kaggle.com/code/baohaoliao/bengali-m4t-submission\" target=\"_blank\">here</a>.</p>\n<p>Basically, you can download the model first and save it in a dataset. Then in the submission notebook, you just need to copy the model from the dataset to the default downloaded path.</p>",
          "rawMarkdown": "I release a notebook [here](https://www.kaggle.com/code/baohaoliao/bengali-m4t-submission).\n\nBasically, you can download the model first and save it in a dataset. Then in the submission notebook, you just need to copy the model from the dataset to the default downloaded path.",
          "votes": 1,
          "replies": [
            {
              "id": 2423969,
              "postDate": "2023-09-05T01:00:56.043Z",
              "content": "<p>Thanks a lot! This helps!</p>",
              "rawMarkdown": "Thanks a lot! This helps!"
            }
          ]
        }
      ]
    },
    {
      "id": 2417729,
      "postDate": "2023-08-31T19:40:33.943Z",
      "content": "<p>Works very well ~~</p>",
      "rawMarkdown": "Works very well ~~"
    },
    {
      "id": 2408442,
      "postDate": "2023-08-25T16:21:59.157Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2412869,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2023-08-28T14:49:58.200000",
      "content": "<p>Anybody attempting distillation with this model ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2412945,
          "author_name": "AIFahim",
          "author_url": "",
          "post_date": "2023-08-28T15:42:39.590000",
          "content": "<p>Distillation would be a great way to reduce parameter a and make the model smaller indeed. You can use pruning or quantization to reduce the model!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2405795,
      "author_name": "Sawradip Saha",
      "author_url": "",
      "post_date": "2023-08-24T05:44:59.737000",
      "content": "<p>Isn't this a translation model? How is that usable here?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2406461,
          "author_name": "SeanInAction",
          "author_url": "",
          "post_date": "2023-08-24T12:59:31.643000",
          "content": "<p>Look at the paper and github! The model has S2TT, T2TT and <strong>ASR</strong> capabilities. The paper includes Bengali as one of the languages, so it's interesting to see how it would perform with the extra data mined and its massive amount of parameters (2.3B for large, 1.2B for medium).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2404430,
      "author_name": "Nischay Dhankhar",
      "author_url": "",
      "post_date": "2023-08-23T08:50:35.920000",
      "content": "<p>1.4 Billion parameters…. that's huge, compared to wav2vec2(~100M), I don't think it will be easy to use SeamlessM4T in this competition. Yet, really curious how well they actually are on all these tasks. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2412946,
          "author_name": "AIFahim",
          "author_url": "",
          "post_date": "2023-08-28T15:43:34.713000",
          "content": "<p>I found when submit wav2vec2 model it takes lot' of time to infer!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2421863,
      "author_name": "Mutian Hong",
      "author_url": "",
      "post_date": "2023-09-03T14:43:01.683000",
      "content": "<p>Hi! I would like to try this model for inference. <br>\nand during the inference, the internet is off, so I wanted to load the model from local like this</p>\n<pre><code>translator = Translator(\n    ,\n    ,\n    torch.device()\n)\n</code></pre>\n<p>however, I got an error:<br>\nValueError: <code>name</code> must be a valid filename, but is '/kaggle/input/meta-ai-seamlessm4t-large/seamless-m4t-large' instead.</p>\n<p>Do anyone know how to fix this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2421972,
          "author_name": "HuBERT",
          "author_url": "",
          "post_date": "2023-09-03T16:14:45.363000",
          "content": "<p>I've already tested the medium model, but it gave poor results, I'll publish the notebook later.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2422802,
              "author_name": "Mutian Hong",
              "author_url": "",
              "post_date": "2023-09-04T08:56:57.030000",
              "content": "<p>Thanks! Looking forward to your notebook.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2422899,
          "author_name": "bliao",
          "author_url": "",
          "post_date": "2023-09-04T10:28:07.447000",
          "content": "<p>I release a notebook <a href=\"https://www.kaggle.com/code/baohaoliao/bengali-m4t-submission\" target=\"_blank\">here</a>.</p>\n<p>Basically, you can download the model first and save it in a dataset. Then in the submission notebook, you just need to copy the model from the dataset to the default downloaded path.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2423969,
              "author_name": "Mutian Hong",
              "author_url": "",
              "post_date": "2023-09-05T01:00:56.043000",
              "content": "<p>Thanks a lot! This helps!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2417729,
      "author_name": "ThomCat",
      "author_url": "",
      "post_date": "2023-08-31T19:40:33.943000",
      "content": "<p>Works very well ~~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2408442,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-25T16:21:59.157000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2404094": "A new MetaAI MNT/ASR/TTS model:\nHuggingface: https://huggingface.co/facebook/seamless-m4t-large\nMetaAI blog: https://ai.meta.com/research/publications/seamless-m4t/\n\n\n\n**SeamlessM4T—Massively Multilingual & Multimodal Machine Translation**\n\n**Abstract**\n\nWhat does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200 languages, unified speech-to-speech translation models have yet to achieve similar strides. More specifically, conventional speech-to-speech translation systems rely on cascaded systems composed of multiple subsystems performing translation progressively, putting scalable and high-performing unified speech translation systems out of reach. To address these gaps, we introduce SeamlessM4T—Massively Multilingual & Multimodal Machine Translation—a single model that supports speech-to-speech translation, speech-to-text translation, text-to-speech translation, text-to-text translation, and automatic speech recognition for up to 100 languages. To build this, we used 1 million hours of open speech audio data to learn self-supervised speech representations with w2v-BERT 2.0. Subsequently, we created a multimodal corpus of automatically aligned speech translations, dubbed SeamlessAlign. Filtered and combined with human labeled and pseudo-labeled data (totaling 406,000 hours), we developed the first multilingual system capable of translating from and into English for both speech and text. On Fleurs, SeamlessM4T sets a new standard for translations into multiple target languages, achieving an improvement of 20% BLEU over the previous state-of-the-art in direct speech-to-text translation. Compared to strong cascaded models, SeamlessM4T improves the quality of into-English translation by 1.3 BLEU points in speech-to-text and by 2.6 ASR-BLEU points in speech-to-speech. On CVSS and compared to a 2-stage cascaded model for speech-to-speech translation, SeamlessM4T-Large’s performance is stronger by 58%. Preliminary human evaluations of speech-to-text translation outputs evinced similarly impressive results; for translations from English, XSTS scores for 24 evaluated languages are consistently above 4 (out of 5). For into English directions, we see significant improvement over WhisperLarge-v2’s baseline for 7 out of 24 languages. To further evaluate our system, we developed Blaser 2.0, which enables evaluation across speech and text with similar accuracy compared to its predecessor when it comes to quality estimation. Tested for robustness, our system performs better against background noises and speaker variations in speech-to-text tasks (average improvements of 38% and 49%, respectively) compared to the current state-of-the-art model. Critically, we evaluated SeamlessM4T on gender bias and added toxicity to assess translation safety. Compared to the state-of-the-art, we report up to 63% of reduction in added toxicity in our translation outputs. Finally, all contributions in this work—including models, inference code, finetuning recipes backed by our improved modeling toolkit Fairseq2, and metadata to recreate the unfiltered 470,000 hours of SeamlessAlign — are open-sourced and accessible at https://github.com/facebookresearch/seamless_communication.",
    "2412869": "Anybody attempting distillation with this model ? ",
    "2405795": "Isn't this a translation model? How is that usable here?\n\n",
    "2404430": "1.4 Billion parameters.... that's huge, compared to wav2vec2(~100M), I don't think it will be easy to use SeamlessM4T in this competition. Yet, really curious how well they actually are on all these tasks. ",
    "2421863": "Hi! I would like to try this model for inference. \nand during the inference, the internet is off, so I wanted to load the model from local like this\n```\ntranslator = Translator(\n    \"/kaggle/input/meta-ai-seamlessm4t-large/seamless-m4t-large\",\n    \"vocoder_36langs\",\n    torch.device(\"cuda:0\")\n)\n```\nhowever, I got an error:\nValueError: `name` must be a valid filename, but is '/kaggle/input/meta-ai-seamlessm4t-large/seamless-m4t-large' instead.\n\nDo anyone know how to fix this?",
    "2417729": "Works very well ~~",
    "2408442": ""
  }
}