{
  "id": 437033,
  "title": "Any idea about ensemble the ASR models like wav2vec2.0, whisper?",
  "url": "/competitions/bengaliai-speech/discussion/437033",
  "author_name": "",
  "post_date": "2023-09-05T06:51:57.028863Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Curious to know about ASR model ensembling. Did anyone try or approach it??<br>\nWhisper is Auto Regressive and Wav2vec2.0 is NAR!! <br>\nThanks in advance :)</p>",
  "messages": [
    {
      "id": "2424257",
      "postDate": "09/05/2023 06:51:57",
      "content": "<p>Curious to know about ASR model ensembling. Did anyone try or approach it??<br>\nWhisper is Auto Regressive and Wav2vec2.0 is NAR!! <br>\nThanks in advance :)</p>",
      "rawMarkdown": "Curious to know about ASR model ensembling. Did anyone try or approach it??\nWhisper is Auto Regressive and Wav2vec2.0 is NAR!! \nThanks in advance :)",
      "votes": null
    },
    {
      "id": "2425272",
      "postDate": "09/05/2023 18:42:17",
      "content": "<p>Very good question. I don't know about whisper and wav2vec together, but you could ensemble multiple wav2vec models since they should have the same output shape.</p>\n<p>I'm sure there are some research papers out there on this topic. This paper on <a href=\"https://arxiv.org/pdf/2306.02561.pdf\" target=\"_blank\">ensembling LLM's </a> seems cool and thinks of a way to combine multiple generative models. Maybe you could get some inspiration and make your own research paper based on this competition :)</p>",
      "rawMarkdown": "Very good question. I don't know about whisper and wav2vec together, but you could ensemble multiple wav2vec models since they should have the same output shape.\n\nI'm sure there are some research papers out there on this topic. This paper on [ensembling LLM's ](https://arxiv.org/pdf/2306.02561.pdf) seems cool and thinks of a way to combine multiple generative models. Maybe you could get some inspiration and make your own research paper based on this competition :)",
      "votes": null
    },
    {
      "id": "2425773",
      "postDate": "09/06/2023 07:03:18",
      "content": "<p>Oh thanks <a href=\"https://www.kaggle.com/msthil\" target=\"_blank\">@msthil</a> :) Something need to build candidate predictor for multiple ASR outputs!  </p>",
      "rawMarkdown": "Oh thanks @msthil :) Something need to build candidate predictor for multiple ASR outputs!",
      "votes": null
    },
    {
      "id": "2445121",
      "postDate": "09/18/2023 16:19:55",
      "content": "<p>I have the same question. Maybe multiple non-autoregressive models with same sequence length of logits, which can be fused as ensembling, but I have no ideas in detail.</p>",
      "rawMarkdown": "I have the same question. Maybe multiple non-autoregressive models with same sequence length of logits, which can be fused as ensembling, but I have no ideas in detail.",
      "votes": null
    },
    {
      "id": "2445134",
      "postDate": "09/18/2023 16:23:56",
      "content": "<p>My idea is pretty similar to <a href=\"https://www.kaggle.com/msthil\" target=\"_blank\">@msthil</a> 's, but I wonder whether there are any instances of fusing output logits. Neural candidate predictor may be too complicated and its inference may consume time, I think we can just take the weighted mean of logits?</p>",
      "rawMarkdown": "My idea is pretty similar to @msthil 's, but I wonder whether there are any instances of fusing output logits. Neural candidate predictor may be too complicated and its inference may consume time, I think we can just take the weighted mean of logits?",
      "votes": null
    },
    {
      "id": "2445151",
      "postDate": "09/18/2023 16:27:59",
      "content": "<p>Though all Bengali as language, the <code>vocab.json</code>s of the models published by users vary a lot, so it's a bit hard to ensemble.</p>",
      "rawMarkdown": "Though all Bengali as language, the `vocab.json`s of the models published by users vary a lot, so it's a bit hard to ensemble.",
      "votes": null
    },
    {
      "id": "2445846",
      "postDate": "09/19/2023 05:46:41",
      "content": "<p>This is good idea I think! &gt;weighted mean of logits</p>",
      "rawMarkdown": "This is good idea I think! >weighted mean of logits",
      "votes": null
    },
    {
      "id": "2445928",
      "postDate": "09/19/2023 06:36:29",
      "content": "<p>Update, we can just fuse the hidden states before the decoder, which always have its dim=1024 or so, so vocab is not a concern. However, what we should concern is sequence length…</p>",
      "rawMarkdown": "Update, we can just fuse the hidden states before the decoder, which always have its dim=1024 or so, so vocab is not a concern. However, what we should concern is sequence length...",
      "votes": null
    },
    {
      "id": "2445930",
      "postDate": "09/19/2023 06:38:09",
      "content": "<p>idk how to fuse when the actual seq length from prediction is not the same, even non-autoregresive models have these problems. Is wav2vec2 non-autoregressive? I searched it on google, but i have still no idea.</p>",
      "rawMarkdown": "idk how to fuse when the actual seq length from prediction is not the same, even non-autoregresive models have these problems. Is wav2vec2 non-autoregressive? I searched it on google, but i have still no idea.",
      "votes": null
    },
    {
      "id": "2445948",
      "postDate": "09/19/2023 06:51:10",
      "content": "<p>oh i forgot to check the caption behind the title. w2v2 is non-autoregressive</p>",
      "rawMarkdown": "oh i forgot to check the caption behind the title. w2v2 is non-autoregressive",
      "votes": null
    },
    {
      "id": "2445951",
      "postDate": "09/19/2023 06:51:48",
      "content": "<p>though the seq length predicted may be not the same, it's still ok to just add seqs?</p>",
      "rawMarkdown": "though the seq length predicted may be not the same, it's still ok to just add seqs?",
      "votes": null
    },
    {
      "id": "2446011",
      "postDate": "09/19/2023 07:36:48",
      "content": "<p>It's very much possible for different non-auto regressive models sequence lengths would be different!! </p>",
      "rawMarkdown": "It's very much possible for different non-auto regressive models sequence lengths would be different!!",
      "votes": null
    },
    {
      "id": "2446182",
      "postDate": "09/19/2023 09:44:53",
      "content": "<p>yea, so i wonder how to fuse this</p>",
      "rawMarkdown": "yea, so i wonder how to fuse this",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2425272,
      "author_name": "msthil",
      "author_url": "",
      "post_date": "09/05/2023 18:42:17",
      "content": "<p>Very good question. I don't know about whisper and wav2vec together, but you could ensemble multiple wav2vec models since they should have the same output shape.</p>\n<p>I'm sure there are some research papers out there on this topic. This paper on <a href=\"https://arxiv.org/pdf/2306.02561.pdf\" target=\"_blank\">ensembling LLM's </a> seems cool and thinks of a way to combine multiple generative models. Maybe you could get some inspiration and make your own research paper based on this competition :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2425773,
          "author_name": "aifahim",
          "author_url": "",
          "post_date": "09/06/2023 07:03:18",
          "content": "<p>Oh thanks <a href=\"https://www.kaggle.com/msthil\" target=\"_blank\">@msthil</a> :) Something need to build candidate predictor for multiple ASR outputs!  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2445151,
          "author_name": "nisshokuitsuki",
          "author_url": "",
          "post_date": "09/18/2023 16:27:59",
          "content": "<p>Though all Bengali as language, the <code>vocab.json</code>s of the models published by users vary a lot, so it's a bit hard to ensemble.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2445928,
              "author_name": "nisshokuitsuki",
              "author_url": "",
              "post_date": "09/19/2023 06:36:29",
              "content": "<p>Update, we can just fuse the hidden states before the decoder, which always have its dim=1024 or so, so vocab is not a concern. However, what we should concern is sequence length…</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2445121,
      "author_name": "nisshokuitsuki",
      "author_url": "",
      "post_date": "09/18/2023 16:19:55",
      "content": "<p>I have the same question. Maybe multiple non-autoregressive models with same sequence length of logits, which can be fused as ensembling, but I have no ideas in detail.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2445134,
          "author_name": "nisshokuitsuki",
          "author_url": "",
          "post_date": "09/18/2023 16:23:56",
          "content": "<p>My idea is pretty similar to <a href=\"https://www.kaggle.com/msthil\" target=\"_blank\">@msthil</a> 's, but I wonder whether there are any instances of fusing output logits. Neural candidate predictor may be too complicated and its inference may consume time, I think we can just take the weighted mean of logits?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2445846,
              "author_name": "aifahim",
              "author_url": "",
              "post_date": "09/19/2023 05:46:41",
              "content": "<p>This is good idea I think! &gt;weighted mean of logits</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2445930,
                  "author_name": "nisshokuitsuki",
                  "author_url": "",
                  "post_date": "09/19/2023 06:38:09",
                  "content": "<p>idk how to fuse when the actual seq length from prediction is not the same, even non-autoregresive models have these problems. Is wav2vec2 non-autoregressive? I searched it on google, but i have still no idea.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2445948,
                      "author_name": "nisshokuitsuki",
                      "author_url": "",
                      "post_date": "09/19/2023 06:51:10",
                      "content": "<p>oh i forgot to check the caption behind the title. w2v2 is non-autoregressive</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2445951,
                          "author_name": "nisshokuitsuki",
                          "author_url": "",
                          "post_date": "09/19/2023 06:51:48",
                          "content": "<p>though the seq length predicted may be not the same, it's still ok to just add seqs?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2446011,
                              "author_name": "aifahim",
                              "author_url": "",
                              "post_date": "09/19/2023 07:36:48",
                              "content": "<p>It's very much possible for different non-auto regressive models sequence lengths would be different!! </p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2446182,
                                  "author_name": "nisshokuitsuki",
                                  "author_url": "",
                                  "post_date": "09/19/2023 09:44:53",
                                  "content": "<p>yea, so i wonder how to fuse this</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2424257": "Curious to know about ASR model ensembling. Did anyone try or approach it??\nWhisper is Auto Regressive and Wav2vec2.0 is NAR!! \nThanks in advance :)",
    "2425272": "Very good question. I don't know about whisper and wav2vec together, but you could ensemble multiple wav2vec models since they should have the same output shape.\n\nI'm sure there are some research papers out there on this topic. This paper on [ensembling LLM's ](https://arxiv.org/pdf/2306.02561.pdf) seems cool and thinks of a way to combine multiple generative models. Maybe you could get some inspiration and make your own research paper based on this competition :)",
    "2425773": "Oh thanks @msthil :) Something need to build candidate predictor for multiple ASR outputs!",
    "2445121": "I have the same question. Maybe multiple non-autoregressive models with same sequence length of logits, which can be fused as ensembling, but I have no ideas in detail.",
    "2445134": "My idea is pretty similar to @msthil 's, but I wonder whether there are any instances of fusing output logits. Neural candidate predictor may be too complicated and its inference may consume time, I think we can just take the weighted mean of logits?",
    "2445151": "Though all Bengali as language, the `vocab.json`s of the models published by users vary a lot, so it's a bit hard to ensemble.",
    "2445846": "This is good idea I think! >weighted mean of logits",
    "2445928": "Update, we can just fuse the hidden states before the decoder, which always have its dim=1024 or so, so vocab is not a concern. However, what we should concern is sequence length...",
    "2445930": "idk how to fuse when the actual seq length from prediction is not the same, even non-autoregresive models have these problems. Is wav2vec2 non-autoregressive? I searched it on google, but i have still no idea.",
    "2445948": "oh i forgot to check the caption behind the title. w2v2 is non-autoregressive",
    "2445951": "though the seq length predicted may be not the same, it's still ok to just add seqs?",
    "2446011": "It's very much possible for different non-auto regressive models sequence lengths would be different!!",
    "2446182": "yea, so i wonder how to fuse this"
  },
  "source": "meta"
}