{
  "id": 431928,
  "title": "Bad performance of whisper small.",
  "url": "/competitions/bengaliai-speech/discussion/431928",
  "author_name": "",
  "post_date": "2023-08-15T11:35:53.379220200Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Due to the limited resources I have, I decided to stick with Whisper Small for initial trials. </p>\n<p>I tried the following</p>\n<p>OpenSLR, Common Voice 14, and this dataset (Bengali.AI)</p>\n<p>Overall, I finetuned 3 models - </p>\n<p>OpenSLR only<br>\nOpenSLR + Bengali.AI<br>\nCommonVoice + Bengali.AI (shuffled)</p>\n<p>Each of the variants gave me a WER of 1. While the indic whisper medium gives 0.529 without any work.</p>\n<p>My training logic</p>\n<p>data-collator from hugging face example<br>\nTokenizer limit to model's limit<br>\nOutputs and labels normalized with bi-unicode normalizer</p>\n<p>Am I doing something wrong? or is the model bad?</p>",
  "messages": [
    {
      "id": "2391931",
      "postDate": "08/15/2023 11:35:53",
      "content": "<p>Due to the limited resources I have, I decided to stick with Whisper Small for initial trials. </p>\n<p>I tried the following</p>\n<p>OpenSLR, Common Voice 14, and this dataset (Bengali.AI)</p>\n<p>Overall, I finetuned 3 models - </p>\n<p>OpenSLR only<br>\nOpenSLR + Bengali.AI<br>\nCommonVoice + Bengali.AI (shuffled)</p>\n<p>Each of the variants gave me a WER of 1. While the indic whisper medium gives 0.529 without any work.</p>\n<p>My training logic</p>\n<p>data-collator from hugging face example<br>\nTokenizer limit to model's limit<br>\nOutputs and labels normalized with bi-unicode normalizer</p>\n<p>Am I doing something wrong? or is the model bad?</p>",
      "rawMarkdown": "Due to the limited resources I have, I decided to stick with Whisper Small for initial trials. \n\nI tried the following\n\nOpenSLR, Common Voice 14, and this dataset (Bengali.AI)\n\nOverall, I finetuned 3 models - \n\nOpenSLR only\nOpenSLR + Bengali.AI\nCommonVoice + Bengali.AI (shuffled)\n\nEach of the variants gave me a WER of 1. While the indic whisper medium gives 0.529 without any work.\n\n\nMy training logic\n\ndata-collator from hugging face example\nTokenizer limit to model's limit\nOutputs and labels normalized with bi-unicode normalizer\n\nAm I doing something wrong? or is the model bad?",
      "votes": null
    },
    {
      "id": "2391965",
      "postDate": "08/15/2023 11:50:51",
      "content": "<p>How much data are you utilising for your experiments? the whisper starting kit <a href=\"https://www.kaggle.com/code/nbroad/whisper-training-starter-kit\" target=\"_blank\">https://www.kaggle.com/code/nbroad/whisper-training-starter-kit</a> here by <a href=\"https://www.kaggle.com/nbroad\" target=\"_blank\">@nbroad</a> already has .69~ WER with 10k samples. So might be related your setup.</p>",
      "rawMarkdown": "How much data are you utilising for your experiments? the whisper starting kit https://www.kaggle.com/code/nbroad/whisper-training-starter-kit here by @nbroad already has .69~ WER with 10k samples. So might be related your setup.",
      "votes": null
    },
    {
      "id": "2391984",
      "postDate": "08/15/2023 12:03:34",
      "content": "<p>I think then there is something seriously wrong with my setup.</p>\n<p>I used the entire dataset 😶</p>",
      "rawMarkdown": "I think then there is something seriously wrong with my setup.\n\nI used the entire dataset 😶",
      "votes": null
    },
    {
      "id": "2392508",
      "postDate": "08/15/2023 17:50:35",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/SuperSecureHuman\" target=\"_blank\">@SuperSecureHuman</a>,</p>\n<p>The common voice dataset is what our training set is based on as well, we basically crowdsourced the data using the common voice platform and then cleaned it up for the training set here. So you can ignore the CV + B.AI (shuffled) setting. Check out the dataset paper linked under the data tab for more info.</p>\n<p>If you want to start with a smaller subset of the dataset to train, you can start with the validation split of the training set (marked as 'valid' under <code>split</code> column of <code>train.csv</code>). This portion has been manually validated and cleaned while the rest of the dataset has mostly been cleaned algorithmically.</p>",
      "rawMarkdown": "Hi @SuperSecureHuman,\n\nThe common voice dataset is what our training set is based on as well, we basically crowdsourced the data using the common voice platform and then cleaned it up for the training set here. So you can ignore the CV + B.AI (shuffled) setting. Check out the dataset paper linked under the data tab for more info.\n\nIf you want to start with a smaller subset of the dataset to train, you can start with the validation split of the training set (marked as 'valid' under `split` column of `train.csv`). This portion has been manually validated and cleaned while the rest of the dataset has mostly been cleaned algorithmically.",
      "votes": null
    },
    {
      "id": "2393095",
      "postDate": "08/16/2023 06:12:08",
      "content": "<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/SuperSecureHuman\" target=\"_blank\">@SuperSecureHuman</a>,</p>\n  <p>The common voice dataset is what our training set is based on as well, we basically crowdsourced the data using the common voice platform and then cleaned it up for the training set here. So you can ignore the CV + B.AI (shuffled) setting. Check out the dataset paper linked under the data tab for more info.</p>\n  <p>If you want to start with a smaller subset of the dataset to train, you can start with the validation split of the training set (marked as 'valid' under <code>split</code> column of <code>train.csv</code>). This portion has been manually validated and cleaned while the rest of the dataset has mostly been cleaned algorithmically.</p>\n</blockquote>\n<p>Thanks!</p>\n<p>This gives a lot of insight into the dataset. I will check the paper.</p>",
      "rawMarkdown": "> Hi @SuperSecureHuman,\n> \n> The common voice dataset is what our training set is based on as well, we basically crowdsourced the data using the common voice platform and then cleaned it up for the training set here. So you can ignore the CV + B.AI (shuffled) setting. Check out the dataset paper linked under the data tab for more info.\n> \n> If you want to start with a smaller subset of the dataset to train, you can start with the validation split of the training set (marked as 'valid' under `split` column of `train.csv`). This portion has been manually validated and cleaned while the rest of the dataset has mostly been cleaned algorithmically.\n\n\nThanks!\n\nThis gives a lot of insight into the dataset. I will check the paper.",
      "votes": null
    },
    {
      "id": "2395774",
      "postDate": "08/17/2023 20:03:51",
      "content": "<p><a href=\"https://www.kaggle.com/supersecurehuman\" target=\"_blank\">@supersecurehuman</a> Have you used LM while using Indic whisper medium?</p>",
      "rawMarkdown": "supersecurehuman Have you used LM while using Indic whisper medium?",
      "votes": null
    },
    {
      "id": "2396365",
      "postDate": "08/18/2023 08:08:29",
      "content": "<p>I did not use LM.</p>",
      "rawMarkdown": "I did not use LM.",
      "votes": null
    },
    {
      "id": "2397603",
      "postDate": "08/19/2023 06:48:08",
      "content": "<p>Hello! I'm interested in utilizing the LM with Whisper for my project. However, I encountered an issue where the tokenizer encoded the text differently. I have been unable to find a solution to integrate the LM(Yellow King's version) with Whisper. Could you please guide me on how to effectively utilize the Language Model in conjunction with the Whisper?</p>",
      "rawMarkdown": "Hello! I'm interested in utilizing the LM with Whisper for my project. However, I encountered an issue where the tokenizer encoded the text differently. I have been unable to find a solution to integrate the LM(Yellow King's version) with Whisper. Could you please guide me on how to effectively utilize the Language Model in conjunction with the Whisper?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2391965,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "08/15/2023 11:50:51",
      "content": "<p>How much data are you utilising for your experiments? the whisper starting kit <a href=\"https://www.kaggle.com/code/nbroad/whisper-training-starter-kit\" target=\"_blank\">https://www.kaggle.com/code/nbroad/whisper-training-starter-kit</a> here by <a href=\"https://www.kaggle.com/nbroad\" target=\"_blank\">@nbroad</a> already has .69~ WER with 10k samples. So might be related your setup.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2391984,
      "author_name": "supersecurehuman",
      "author_url": "",
      "post_date": "08/15/2023 12:03:34",
      "content": "<p>I think then there is something seriously wrong with my setup.</p>\n<p>I used the entire dataset 😶</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2392508,
      "author_name": "imtiazprio",
      "author_url": "",
      "post_date": "08/15/2023 17:50:35",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/SuperSecureHuman\" target=\"_blank\">@SuperSecureHuman</a>,</p>\n<p>The common voice dataset is what our training set is based on as well, we basically crowdsourced the data using the common voice platform and then cleaned it up for the training set here. So you can ignore the CV + B.AI (shuffled) setting. Check out the dataset paper linked under the data tab for more info.</p>\n<p>If you want to start with a smaller subset of the dataset to train, you can start with the validation split of the training set (marked as 'valid' under <code>split</code> column of <code>train.csv</code>). This portion has been manually validated and cleaned while the rest of the dataset has mostly been cleaned algorithmically.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2393095,
          "author_name": "supersecurehuman",
          "author_url": "",
          "post_date": "08/16/2023 06:12:08",
          "content": "<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/SuperSecureHuman\" target=\"_blank\">@SuperSecureHuman</a>,</p>\n  <p>The common voice dataset is what our training set is based on as well, we basically crowdsourced the data using the common voice platform and then cleaned it up for the training set here. So you can ignore the CV + B.AI (shuffled) setting. Check out the dataset paper linked under the data tab for more info.</p>\n  <p>If you want to start with a smaller subset of the dataset to train, you can start with the validation split of the training set (marked as 'valid' under <code>split</code> column of <code>train.csv</code>). This portion has been manually validated and cleaned while the rest of the dataset has mostly been cleaned algorithmically.</p>\n</blockquote>\n<p>Thanks!</p>\n<p>This gives a lot of insight into the dataset. I will check the paper.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2395774,
      "author_name": "rajgothi",
      "author_url": "",
      "post_date": "08/17/2023 20:03:51",
      "content": "<p><a href=\"https://www.kaggle.com/supersecurehuman\" target=\"_blank\">@supersecurehuman</a> Have you used LM while using Indic whisper medium?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2396365,
          "author_name": "supersecurehuman",
          "author_url": "",
          "post_date": "08/18/2023 08:08:29",
          "content": "<p>I did not use LM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2397603,
          "author_name": "hongori",
          "author_url": "",
          "post_date": "08/19/2023 06:48:08",
          "content": "<p>Hello! I'm interested in utilizing the LM with Whisper for my project. However, I encountered an issue where the tokenizer encoded the text differently. I have been unable to find a solution to integrate the LM(Yellow King's version) with Whisper. Could you please guide me on how to effectively utilize the Language Model in conjunction with the Whisper?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2391931": "Due to the limited resources I have, I decided to stick with Whisper Small for initial trials. \n\nI tried the following\n\nOpenSLR, Common Voice 14, and this dataset (Bengali.AI)\n\nOverall, I finetuned 3 models - \n\nOpenSLR only\nOpenSLR + Bengali.AI\nCommonVoice + Bengali.AI (shuffled)\n\nEach of the variants gave me a WER of 1. While the indic whisper medium gives 0.529 without any work.\n\n\nMy training logic\n\ndata-collator from hugging face example\nTokenizer limit to model's limit\nOutputs and labels normalized with bi-unicode normalizer\n\nAm I doing something wrong? or is the model bad?",
    "2391965": "How much data are you utilising for your experiments? the whisper starting kit https://www.kaggle.com/code/nbroad/whisper-training-starter-kit here by @nbroad already has .69~ WER with 10k samples. So might be related your setup.",
    "2391984": "I think then there is something seriously wrong with my setup.\n\nI used the entire dataset 😶",
    "2392508": "Hi @SuperSecureHuman,\n\nThe common voice dataset is what our training set is based on as well, we basically crowdsourced the data using the common voice platform and then cleaned it up for the training set here. So you can ignore the CV + B.AI (shuffled) setting. Check out the dataset paper linked under the data tab for more info.\n\nIf you want to start with a smaller subset of the dataset to train, you can start with the validation split of the training set (marked as 'valid' under `split` column of `train.csv`). This portion has been manually validated and cleaned while the rest of the dataset has mostly been cleaned algorithmically.",
    "2393095": "> Hi @SuperSecureHuman,\n> \n> The common voice dataset is what our training set is based on as well, we basically crowdsourced the data using the common voice platform and then cleaned it up for the training set here. So you can ignore the CV + B.AI (shuffled) setting. Check out the dataset paper linked under the data tab for more info.\n> \n> If you want to start with a smaller subset of the dataset to train, you can start with the validation split of the training set (marked as 'valid' under `split` column of `train.csv`). This portion has been manually validated and cleaned while the rest of the dataset has mostly been cleaned algorithmically.\n\n\nThanks!\n\nThis gives a lot of insight into the dataset. I will check the paper.",
    "2395774": "supersecurehuman Have you used LM while using Indic whisper medium?",
    "2396365": "I did not use LM.",
    "2397603": "Hello! I'm interested in utilizing the LM with Whisper for my project. However, I encountered an issue where the tokenizer encoded the text differently. I have been unable to find a solution to integrate the LM(Yellow King's version) with Whisper. Could you please guide me on how to effectively utilize the Language Model in conjunction with the Whisper?"
  },
  "source": "meta"
}