{
  "id": 444510,
  "title": "Has anyone had success with Whisper or Nemo ?",
  "url": "/competitions/bengaliai-speech/discussion/444510",
  "author_name": "Dmitriy Gerasimov",
  "post_date": "2023-10-02T11:40:49.557000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'm testing the following:</p>\n<ol>\n<li>Whisper, training with peft. Medium doesn't pass the time limit for inference, Small model gives around LB 0.6, even not 0.54 (as public notebook for Whisper). </li>\n<li>Nemo, CTC Conformer BPE, Large.  At this moment the  training still diverges, making loss more (predictions become empty at some moment). I'm continue with issue search.</li>\n</ol>\n<p>I see that in official paper (Table 2: Benchmarking pre-trained networks on the<br>\nOOD-Speech test subsets): </p>\n<ol>\n<li>Whisper-small gives 0.25 WER (!) for Macro DS</li>\n<li>Conformer-CTC gives 0.55 WER for Macro DS</li>\n</ol>\n<p>Is there any success with Whisper/Nemo?</p>",
  "messages": [
    {
      "id": 2464620,
      "postDate": "2023-10-02T11:40:49.557Z",
      "content": "<p>I'm testing the following:</p>\n<ol>\n<li>Whisper, training with peft. Medium doesn't pass the time limit for inference, Small model gives around LB 0.6, even not 0.54 (as public notebook for Whisper). </li>\n<li>Nemo, CTC Conformer BPE, Large.  At this moment the  training still diverges, making loss more (predictions become empty at some moment). I'm continue with issue search.</li>\n</ol>\n<p>I see that in official paper (Table 2: Benchmarking pre-trained networks on the<br>\nOOD-Speech test subsets): </p>\n<ol>\n<li>Whisper-small gives 0.25 WER (!) for Macro DS</li>\n<li>Conformer-CTC gives 0.55 WER for Macro DS</li>\n</ol>\n<p>Is there any success with Whisper/Nemo?</p>",
      "rawMarkdown": "I'm testing the following:\n1. Whisper, training with peft. Medium doesn't pass the time limit for inference, Small model gives around LB 0.6, even not 0.54 (as public notebook for Whisper). \n2. Nemo, CTC Conformer BPE, Large.  At this moment the  training still diverges, making loss more (predictions become empty at some moment). I'm continue with issue search.\n\nI see that in official paper (Table 2: Benchmarking pre-trained networks on the\nOOD-Speech test subsets): \n1. Whisper-small gives 0.25 WER (!) for Macro DS\n2. Conformer-CTC gives 0.55 WER for Macro DS\n\nIs there any success with Whisper/Nemo?",
      "votes": 4
    },
    {
      "id": 2465453,
      "postDate": "2023-10-03T01:28:00.733Z",
      "content": "<p>I guess nobody could reproduce the whisper results from the OOD paper ? </p>",
      "rawMarkdown": "I guess nobody could reproduce the whisper results from the OOD paper ? ",
      "votes": 1
    },
    {
      "id": 2469331,
      "postDate": "2023-10-06T09:25:07.717Z",
      "content": "<p>Hi Dimitriy,</p>\n<p>Just to give some context, the Macro test set is the in-distribution subset from the test set. It is part of the public test set in the competition (not all of it since there are OOD recordings in public as well). Therefore, the WER you see on the public test set wont match exactly with the Macro Test set results from the paper!</p>",
      "rawMarkdown": "Hi Dimitriy,\n\nJust to give some context, the Macro test set is the in-distribution subset from the test set. It is part of the public test set in the competition (not all of it since there are OOD recordings in public as well). Therefore, the WER you see on the public test set wont match exactly with the Macro Test set results from the paper!"
    },
    {
      "id": 2469113,
      "postDate": "2023-10-06T07:02:50.480Z",
      "content": "<p>I tryed to tune whisper-large-v2 int-8 via QLORA on 80% of train data in 2 epochs and check score on examples.<br>\nI got 0.76 with greedy decoding and 0.66 with beam_width=4, but the inference took too long in case of beam search. <br>\nMain problem with whisper in my opinion - long inference with large beam_width param.</p>",
      "rawMarkdown": "I tryed to tune whisper-large-v2 int-8 via QLORA on 80% of train data in 2 epochs and check score on examples.\nI got 0.76 with greedy decoding and 0.66 with beam_width=4, but the inference took too long in case of beam search. \nMain problem with whisper in my opinion - long inference with large beam_width param."
    },
    {
      "id": 2464648,
      "postDate": "2023-10-02T12:07:47.203Z",
      "content": "<p>Isn't <a href=\"https://www.kaggle.com/code/nbroad/whisper-inference\" target=\"_blank\">https://www.kaggle.com/code/nbroad/whisper-inference</a> a whisper medium model? If yes, the notebook says <code>2.5 hours for whisper medium</code>.</p>",
      "rawMarkdown": "Isn't https://www.kaggle.com/code/nbroad/whisper-inference a whisper medium model? If yes, the notebook says `2.5 hours for whisper medium`.",
      "replies": [
        {
          "id": 2464849,
          "postDate": "2023-10-02T14:57:57.470Z",
          "content": "<p>Looks like you right, thank you!  Not possible to inference with Medium and 8 bit.  But Medium and full/half inference working much faster.</p>\n<p>see <a href=\"https://github.com/huggingface/peft/discussions/477\" target=\"_blank\">https://github.com/huggingface/peft/discussions/477</a> </p>\n<p>\"Takeaway: PEFT is great for stable, low-resource training in 8-bit. We can then leverage the fine-tuned checkpoints for fast inference in full or half precision and negate possible hallucinations\"</p>",
          "rawMarkdown": "Looks like you right, thank you!  Not possible to inference with Medium and 8 bit.  But Medium and full/half inference working much faster.\n\nsee https://github.com/huggingface/peft/discussions/477 \n\n\"Takeaway: PEFT is great for stable, low-resource training in 8-bit. We can then leverage the fine-tuned checkpoints for fast inference in full or half precision and negate possible hallucinations\""
        }
      ]
    },
    {
      "id": 2464666,
      "postDate": "2023-10-02T12:20:17.500Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2465453,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2023-10-03T01:28:00.733000",
      "content": "<p>I guess nobody could reproduce the whisper results from the OOD paper ? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2469331,
      "author_name": "Ahmed Imtiaz Humayun",
      "author_url": "",
      "post_date": "2023-10-06T09:25:07.717000",
      "content": "<p>Hi Dimitriy,</p>\n<p>Just to give some context, the Macro test set is the in-distribution subset from the test set. It is part of the public test set in the competition (not all of it since there are OOD recordings in public as well). Therefore, the WER you see on the public test set wont match exactly with the Macro Test set results from the paper!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2469113,
      "author_name": "Ivan Ilyushchenko",
      "author_url": "",
      "post_date": "2023-10-06T07:02:50.480000",
      "content": "<p>I tryed to tune whisper-large-v2 int-8 via QLORA on 80% of train data in 2 epochs and check score on examples.<br>\nI got 0.76 with greedy decoding and 0.66 with beam_width=4, but the inference took too long in case of beam search. <br>\nMain problem with whisper in my opinion - long inference with large beam_width param.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2464648,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "2023-10-02T12:07:47.203000",
      "content": "<p>Isn't <a href=\"https://www.kaggle.com/code/nbroad/whisper-inference\" target=\"_blank\">https://www.kaggle.com/code/nbroad/whisper-inference</a> a whisper medium model? If yes, the notebook says <code>2.5 hours for whisper medium</code>.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2464849,
          "author_name": "Dmitriy Gerasimov",
          "author_url": "",
          "post_date": "2023-10-02T14:57:57.470000",
          "content": "<p>Looks like you right, thank you!  Not possible to inference with Medium and 8 bit.  But Medium and full/half inference working much faster.</p>\n<p>see <a href=\"https://github.com/huggingface/peft/discussions/477\" target=\"_blank\">https://github.com/huggingface/peft/discussions/477</a> </p>\n<p>\"Takeaway: PEFT is great for stable, low-resource training in 8-bit. We can then leverage the fine-tuned checkpoints for fast inference in full or half precision and negate possible hallucinations\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2464666,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-02T12:20:17.500000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2464620": "I'm testing the following:\n1. Whisper, training with peft. Medium doesn't pass the time limit for inference, Small model gives around LB 0.6, even not 0.54 (as public notebook for Whisper). \n2. Nemo, CTC Conformer BPE, Large.  At this moment the  training still diverges, making loss more (predictions become empty at some moment). I'm continue with issue search.\n\nI see that in official paper (Table 2: Benchmarking pre-trained networks on the\nOOD-Speech test subsets): \n1. Whisper-small gives 0.25 WER (!) for Macro DS\n2. Conformer-CTC gives 0.55 WER for Macro DS\n\nIs there any success with Whisper/Nemo?",
    "2465453": "I guess nobody could reproduce the whisper results from the OOD paper ? ",
    "2469331": "Hi Dimitriy,\n\nJust to give some context, the Macro test set is the in-distribution subset from the test set. It is part of the public test set in the competition (not all of it since there are OOD recordings in public as well). Therefore, the WER you see on the public test set wont match exactly with the Macro Test set results from the paper!",
    "2469113": "I tryed to tune whisper-large-v2 int-8 via QLORA on 80% of train data in 2 epochs and check score on examples.\nI got 0.76 with greedy decoding and 0.66 with beam_width=4, but the inference took too long in case of beam search. \nMain problem with whisper in my opinion - long inference with large beam_width param.",
    "2464648": "Isn't https://www.kaggle.com/code/nbroad/whisper-inference a whisper medium model? If yes, the notebook says `2.5 hours for whisper medium`.",
    "2464666": ""
  }
}