{
  "id": 435300,
  "title": "Datasets and Model Checkpoint for resource efficient training",
  "url": "/competitions/bengaliai-speech/discussion/435300",
  "author_name": "Umong Sain",
  "post_date": "2023-08-28T19:22:16.477000",
  "votes": 36,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Training on 960K+ data is massive. Training for just a single epoch might take over 100 hours, and the worst part is that a lot of the data is mislabelled or of bad audio quality, which could do more harm than good. Instead of using the entire dataset, using a subset of good data would yield better results. <a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">This notebook</a>, by <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> and <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a>, explores various methods to filter out good-quality data. Since Common Voice is a subset of the MaCro dataset, I have chosen data samples from the Common Voice dataset, specifically those with more upvotes than downvotes. I have already <strong>removed punctuations</strong>, <strong>normalized the texts</strong>, and fixed problematic entries. For those with limited resources, I recommend using only the <code>train</code> split for training, which contains 20k samples. If more samples are needed, the <code>other</code> split (47k+ samples) can be combined with the <code>train</code> split.</p>\n<p>I have already trained a model using only the <code>train</code> split. When coupled with a language model, it outperforms the model created by YellowKing in the public LB.</p>\n<ol>\n<li>Dataset link: <a href=\"https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized\" target=\"_blank\">https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized</a></li>\n<li>Pretrained Model: <a href=\"https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali\" target=\"_blank\">https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali</a></li>\n</ol>\n<p>Feel free to provide feedback. Happy Kaggling!</p>",
  "messages": [
    {
      "id": 2413302,
      "postDate": "2023-08-28T19:22:16.477Z",
      "content": "<p>Training on 960K+ data is massive. Training for just a single epoch might take over 100 hours, and the worst part is that a lot of the data is mislabelled or of bad audio quality, which could do more harm than good. Instead of using the entire dataset, using a subset of good data would yield better results. <a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">This notebook</a>, by <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> and <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a>, explores various methods to filter out good-quality data. Since Common Voice is a subset of the MaCro dataset, I have chosen data samples from the Common Voice dataset, specifically those with more upvotes than downvotes. I have already <strong>removed punctuations</strong>, <strong>normalized the texts</strong>, and fixed problematic entries. For those with limited resources, I recommend using only the <code>train</code> split for training, which contains 20k samples. If more samples are needed, the <code>other</code> split (47k+ samples) can be combined with the <code>train</code> split.</p>\n<p>I have already trained a model using only the <code>train</code> split. When coupled with a language model, it outperforms the model created by YellowKing in the public LB.</p>\n<ol>\n<li>Dataset link: <a href=\"https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized\" target=\"_blank\">https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized</a></li>\n<li>Pretrained Model: <a href=\"https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali\" target=\"_blank\">https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali</a></li>\n</ol>\n<p>Feel free to provide feedback. Happy Kaggling!</p>",
      "rawMarkdown": "Training on 960K+ data is massive. Training for just a single epoch might take over 100 hours, and the worst part is that a lot of the data is mislabelled or of bad audio quality, which could do more harm than good. Instead of using the entire dataset, using a subset of good data would yield better results. [This notebook](https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda), by @imtiazprio and @reasat, explores various methods to filter out good-quality data. Since Common Voice is a subset of the MaCro dataset, I have chosen data samples from the Common Voice dataset, specifically those with more upvotes than downvotes. I have already **removed punctuations**, **normalized the texts**, and fixed problematic entries. For those with limited resources, I recommend using only the `train` split for training, which contains 20k samples. If more samples are needed, the `other` split (47k+ samples) can be combined with the `train` split.\n\nI have already trained a model using only the `train` split. When coupled with a language model, it outperforms the model created by YellowKing in the public LB.\n\n1. Dataset link: https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized\n2. Pretrained Model: https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali\n\nFeel free to provide feedback. Happy Kaggling!",
      "votes": 36
    },
    {
      "id": 2419777,
      "postDate": "2023-09-02T08:09:29.733Z",
      "content": "<p>Thank you for sharing.</p>\n<p>I just made submission by your model and got 0.451 public score.  <br>\nIt seems that the model will get higher score than <code>ai4bharat/indicwav2vec_v1_bengali</code> by fine-tuning.</p>",
      "rawMarkdown": "Thank you for sharing.\n\nI just made submission by your model and got 0.451 public score.  \nIt seems that the model will get higher score than `ai4bharat/indicwav2vec_v1_bengali` by fine-tuning.",
      "votes": 3,
      "replies": [
        {
          "id": 2419808,
          "postDate": "2023-09-02T08:21:49.717Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, for the feedback. Curious to see how people will use methods like LoRA or adapters for parameter efficient fine-tuning.</p>",
          "rawMarkdown": "Thanks @ttahara, for the feedback. Curious to see how people will use methods like LoRA or adapters for parameter efficient fine-tuning.",
          "votes": 1
        },
        {
          "id": 2420766,
          "postDate": "2023-09-02T21:27:24.547Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, I am also trying to use <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> model to see the results. However, I am facing an error while I am building the decoder with 5gram. It seems the tokenizer vocab has <code>ben</code> (nested), <code>&lt;s&gt;</code>,  and <code>&lt;/s&gt;</code> keys. How did you build the decoder with this nested tokenizer vocab list? Thanks in advance.</p>",
          "rawMarkdown": "Hi @ttahara, I am also trying to use @umongsain model to see the results. However, I am facing an error while I am building the decoder with 5gram. It seems the tokenizer vocab has `ben` (nested), `<s>`,  and `</s>` keys. How did you build the decoder with this nested tokenizer vocab list? Thanks in advance.",
          "replies": [
            {
              "id": 2420790,
              "postDate": "2023-09-02T22:38:50.033Z",
              "content": "<p>As <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> said in <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300#2416881\" target=\"_blank\">this comment</a>, I use <code>Wav2Vec2Processor</code> and <code>pyctcdecode</code> separately.</p>\n<p>I'll share inference notebook later.</p>",
              "rawMarkdown": "As @umongsain said in [this comment](https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300#2416881), I use `Wav2Vec2Processor` and `pyctcdecode` separately.\n\nI'll share inference notebook later.",
              "votes": 1
            },
            {
              "id": 2420869,
              "postDate": "2023-09-03T02:55:20.330Z",
              "content": "<p><a href=\"https://www.kaggle.com/pritamsinha23\" target=\"_blank\">@pritamsinha23</a> <br>\nI've just pulished inference notebook. Check it :)<br>\n<a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-umong-sain-s-wav2vec2-0-w-lm-baseline\" target=\"_blank\">https://www.kaggle.com/code/ttahara/bengali-sr-umong-sain-s-wav2vec2-0-w-lm-baseline</a></p>",
              "rawMarkdown": "@pritamsinha23 \nI've just pulished inference notebook. Check it :)\nhttps://www.kaggle.com/code/ttahara/bengali-sr-umong-sain-s-wav2vec2-0-w-lm-baseline",
              "votes": 1
            },
            {
              "id": 2421873,
              "postDate": "2023-09-03T14:49:26.923Z",
              "content": "<p>Thank you for showing us all these <br>\nThe model is 3 times bigger than your previous baseline, but the vocabulary is lesser (from 87 to 64).  </p>",
              "rawMarkdown": "Thank you for showing us all these \nThe model is 3 times bigger than your previous baseline, but the vocabulary is lesser (from 87 to 64).  ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2413311,
      "postDate": "2023-08-28T19:27:15.833Z",
      "content": "<p>Text corpus is available for those who are using a language model: <a href=\"https://www.kaggle.com/datasets/umongsain/indiccorp-v2-bengali\" target=\"_blank\">https://www.kaggle.com/datasets/umongsain/indiccorp-v2-bengali</a><br>\nEnsure that you <strong>normalize the data</strong> before training the language model.</p>",
      "rawMarkdown": "Text corpus is available for those who are using a language model: https://www.kaggle.com/datasets/umongsain/indiccorp-v2-bengali\nEnsure that you **normalize the data** before training the language model.",
      "votes": 4,
      "replies": [
        {
          "id": 2420272,
          "postDate": "2023-09-02T14:29:26.427Z",
          "content": "<p>Is there already pre-trained model for this corpus or do we need to train it ourselves ?  Care to share your experience using this corpus ? </p>",
          "rawMarkdown": "Is there already pre-trained model for this corpus or do we need to train it ourselves ?  Care to share your experience using this corpus ? "
        }
      ]
    },
    {
      "id": 2413994,
      "postDate": "2023-08-29T09:40:25.643Z",
      "content": "<p>Hello Umong. </p>\n<p>Did you employ the vocabulary and normalisation strategy which you described in the following link</p>\n<p><a href=\"https://www.kaggle.com/code/umongsain/macro-normalization\" target=\"_blank\">https://www.kaggle.com/code/umongsain/macro-normalization</a></p>",
      "rawMarkdown": "Hello Umong. \n\nDid you employ the vocabulary and normalisation strategy which you described in the following link\n\nhttps://www.kaggle.com/code/umongsain/macro-normalization",
      "votes": 1,
      "replies": [
        {
          "id": 2414037,
          "postDate": "2023-08-29T10:40:32.293Z",
          "content": "<p>Yes, I did.</p>",
          "rawMarkdown": "Yes, I did.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2416128,
      "postDate": "2023-08-30T19:55:03.853Z",
      "content": "<blockquote>\n  <p>a lot of the data is mislabelled or of bad audio quality</p>\n</blockquote>\n<p>Just curious, how did you know that?</p>",
      "rawMarkdown": "> a lot of the data is mislabelled or of bad audio quality\n\nJust curious, how did you know that?",
      "votes": 2,
      "replies": [
        {
          "id": 2416710,
          "postDate": "2023-08-31T07:11:57.740Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">This notebook</a> showed few examples where WER was high for YellowKing's model just because the audio were mislabelled. I tried to filter out audio using different thresholds for MOS and YKG's WER, but failed. Everytime I found some examples with mislabelled data or data where the speakers were mumbling a lot.</p>",
          "rawMarkdown": "[This notebook](https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda) showed few examples where WER was high for YellowKing's model just because the audio were mislabelled. I tried to filter out audio using different thresholds for MOS and YKG's WER, but failed. Everytime I found some examples with mislabelled data or data where the speakers were mumbling a lot.",
          "votes": 2,
          "replies": [
            {
              "id": 2416776,
              "postDate": "2023-08-31T07:49:48.020Z",
              "content": "<p>This is interesting, could you please elaborate how you were thresholding using YKG WER+MOS?<br>\nAlso, we have provided clientID which could help find out which contributors were mumbling too much or were very noisy. Remember that a little bit of label noise would be allowable and even help regularize the network.</p>",
              "rawMarkdown": "This is interesting, could you please elaborate how you were thresholding using YKG WER+MOS?\nAlso, we have provided clientID which could help find out which contributors were mumbling too much or were very noisy. Remember that a little bit of label noise would be allowable and even help regularize the network.",
              "votes": 3
            }
          ]
        },
        {
          "id": 2416786,
          "postDate": "2023-08-31T07:59:14.043Z",
          "content": "<p>It seems only 60-70% of the train data is usable.</p>",
          "rawMarkdown": "It seems only 60-70% of the train data is usable.",
          "votes": 4,
          "replies": [
            {
              "id": 2424327,
              "postDate": "2023-09-05T07:24:02.590Z",
              "content": "<p>how could you identify this 60-70% of the data that is usable? did you rely on MOS &amp; YKG's WER correlation?</p>",
              "rawMarkdown": "how could you identify this 60-70% of the data that is usable? did you rely on MOS & YKG's WER correlation?",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2415110,
      "postDate": "2023-08-30T05:51:44.180Z",
      "content": "<p>Good job, does this dataset have overlap with the competition dataset? The valid split of the competition dataset already has 29588 samples, so we get 20k more clean samples? </p>",
      "rawMarkdown": "Good job, does this dataset have overlap with the competition dataset? The valid split of the competition dataset already has 29588 samples, so we get 20k more clean samples? ",
      "votes": 2,
      "replies": [
        {
          "id": 2415160,
          "postDate": "2023-08-30T06:44:19.330Z",
          "content": "<p>Not sure though. The organizer gave the mappings to common voice id for the train split. So the validation split might have data from common voice.</p>",
          "rawMarkdown": "Not sure though. The organizer gave the mappings to common voice id for the train split. So the validation split might have data from common voice.",
          "votes": 2
        },
        {
          "id": 2416762,
          "postDate": "2023-08-31T07:41:34.217Z",
          "content": "<p><a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> <a href=\"https://www.kaggle.com/aphysict\" target=\"_blank\">@aphysict</a> </p>\n<p>Will be making the validation split quality+mappings available in a few days. The validation split (in the competition dataset) has all the upvoted samples from CV12 + more of which were evaluated by our in house teams.</p>",
          "rawMarkdown": "@umongsain @aphysict \n\nWill be making the validation split quality+mappings available in a few days. The validation split (in the competition dataset) has all the upvoted samples from CV12 + more of which were evaluated by our in house teams.",
          "votes": 5
        }
      ]
    },
    {
      "id": 2443608,
      "postDate": "2023-09-17T20:14:08.337Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> </p>\n<p>Thank you for sharing this dataset. How did you prepare the other part of the dataset? Still the same approach as train - more upvotes than the downvotes?</p>",
      "rawMarkdown": "Hi @umongsain \n\nThank you for sharing this dataset. How did you prepare the other part of the dataset? Still the same approach as train - more upvotes than the downvotes?"
    },
    {
      "id": 2419534,
      "postDate": "2023-09-02T04:54:46.713Z",
      "content": "<p>Hi, I want to know how did you train/finetuning the model. Can you share the training notebook if possible? Thanks</p>",
      "rawMarkdown": "Hi, I want to know how did you train/finetuning the model. Can you share the training notebook if possible? Thanks"
    },
    {
      "id": 2417044,
      "postDate": "2023-08-31T11:16:47.113Z",
      "content": "<p>I am following similar setup as Tawara<br>\n<a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\" target=\"_blank\">https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline</a></p>\n<p>When I kept the vocab as such and I finetuned my model, I was able to get 0.439 in 2 epochs</p>\n<p>When I changed the vocab to the one show in the following link, my performance dropped to 0.503<br>\n<a href=\"https://www.kaggle.com/code/umongsain/macro-normalization\" target=\"_blank\">https://www.kaggle.com/code/umongsain/macro-normalization</a></p>\n<p>Is there some thing that I need to do ?</p>",
      "rawMarkdown": "I am following similar setup as Tawara\nhttps://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\n\nWhen I kept the vocab as such and I finetuned my model, I was able to get 0.439 in 2 epochs\n\nWhen I changed the vocab to the one show in the following link, my performance dropped to 0.503\nhttps://www.kaggle.com/code/umongsain/macro-normalization\n\nIs there some thing that I need to do ?",
      "replies": [
        {
          "id": 2417696,
          "postDate": "2023-08-31T19:10:38.847Z",
          "content": "<p>Changing vocab is not enough. You have tune after changing the vocab</p>",
          "rawMarkdown": "Changing vocab is not enough. You have tune after changing the vocab",
          "votes": 1,
          "replies": [
            {
              "id": 2418352,
              "postDate": "2023-09-01T08:53:02.603Z",
              "content": "<p>I finetuned for 2 epochs after changing vocab.. <br>\nThe eval loss and wer reduced a lot but the LB score became 0.503</p>",
              "rawMarkdown": "I finetuned for 2 epochs after changing vocab.. \nThe eval loss and wer reduced a lot but the LB score became 0.503",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2416856,
      "postDate": "2023-08-31T08:55:59.287Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4350309%2F820297538c297ca97453dc05b450a466%2F2023-08-31_14-54.png?generation=1693472073318037&amp;alt=media\" alt=\"\"><br>\nThanks for you contribution <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a>. How did you include LM part? Did you use Wav2Vec2ProcessorWithLM?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4350309%2F820297538c297ca97453dc05b450a466%2F2023-08-31_14-54.png?generation=1693472073318037&alt=media)\nThanks for you contribution @umongsain. How did you include LM part? Did you use Wav2Vec2ProcessorWithLM?",
      "replies": [
        {
          "id": 2416881,
          "postDate": "2023-08-31T09:13:29.713Z",
          "content": "<p>I used <code>pyctcdecode</code> to decode the logits. <code>Wav2Vec2ProcessorWithLM</code> is not compatible with new MMS Adapters yet.</p>",
          "rawMarkdown": "I used `pyctcdecode` to decode the logits. `Wav2Vec2ProcessorWithLM` is not compatible with new MMS Adapters yet.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2413797,
      "postDate": "2023-08-29T07:06:31.447Z",
      "content": "<p>How did you manage to submit such a large model? Did you do it with the Kenlm model?</p>",
      "rawMarkdown": "How did you manage to submit such a large model? Did you do it with the Kenlm model?",
      "replies": [
        {
          "id": 2414036,
          "postDate": "2023-08-29T10:40:11.990Z",
          "content": "<p>With a KenLM 5-gram model, the inference time is almost 4 hours.</p>",
          "rawMarkdown": "With a KenLM 5-gram model, the inference time is almost 4 hours.",
          "votes": 2,
          "replies": [
            {
              "id": 2415524,
              "postDate": "2023-08-30T12:19:53.173Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2419777,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2023-09-02T08:09:29.733000",
      "content": "<p>Thank you for sharing.</p>\n<p>I just made submission by your model and got 0.451 public score.  <br>\nIt seems that the model will get higher score than <code>ai4bharat/indicwav2vec_v1_bengali</code> by fine-tuning.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2419808,
          "author_name": "Umong Sain",
          "author_url": "",
          "post_date": "2023-09-02T08:21:49.717000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, for the feedback. Curious to see how people will use methods like LoRA or adapters for parameter efficient fine-tuning.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2420766,
          "author_name": "Pritam Sinha",
          "author_url": "",
          "post_date": "2023-09-02T21:27:24.547000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, I am also trying to use <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> model to see the results. However, I am facing an error while I am building the decoder with 5gram. It seems the tokenizer vocab has <code>ben</code> (nested), <code>&lt;s&gt;</code>,  and <code>&lt;/s&gt;</code> keys. How did you build the decoder with this nested tokenizer vocab list? Thanks in advance.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2420790,
              "author_name": "Tawara",
              "author_url": "",
              "post_date": "2023-09-02T22:38:50.033000",
              "content": "<p>As <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> said in <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300#2416881\" target=\"_blank\">this comment</a>, I use <code>Wav2Vec2Processor</code> and <code>pyctcdecode</code> separately.</p>\n<p>I'll share inference notebook later.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2420869,
              "author_name": "Tawara",
              "author_url": "",
              "post_date": "2023-09-03T02:55:20.330000",
              "content": "<p><a href=\"https://www.kaggle.com/pritamsinha23\" target=\"_blank\">@pritamsinha23</a> <br>\nI've just pulished inference notebook. Check it :)<br>\n<a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-umong-sain-s-wav2vec2-0-w-lm-baseline\" target=\"_blank\">https://www.kaggle.com/code/ttahara/bengali-sr-umong-sain-s-wav2vec2-0-w-lm-baseline</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2421873,
              "author_name": "yukiya",
              "author_url": "",
              "post_date": "2023-09-03T14:49:26.923000",
              "content": "<p>Thank you for showing us all these <br>\nThe model is 3 times bigger than your previous baseline, but the vocabulary is lesser (from 87 to 64).  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2413311,
      "author_name": "Umong Sain",
      "author_url": "",
      "post_date": "2023-08-28T19:27:15.833000",
      "content": "<p>Text corpus is available for those who are using a language model: <a href=\"https://www.kaggle.com/datasets/umongsain/indiccorp-v2-bengali\" target=\"_blank\">https://www.kaggle.com/datasets/umongsain/indiccorp-v2-bengali</a><br>\nEnsure that you <strong>normalize the data</strong> before training the language model.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2420272,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2023-09-02T14:29:26.427000",
          "content": "<p>Is there already pre-trained model for this corpus or do we need to train it ourselves ?  Care to share your experience using this corpus ? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2413994,
      "author_name": "Balaji Selvaraj",
      "author_url": "",
      "post_date": "2023-08-29T09:40:25.643000",
      "content": "<p>Hello Umong. </p>\n<p>Did you employ the vocabulary and normalisation strategy which you described in the following link</p>\n<p><a href=\"https://www.kaggle.com/code/umongsain/macro-normalization\" target=\"_blank\">https://www.kaggle.com/code/umongsain/macro-normalization</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2414037,
          "author_name": "Umong Sain",
          "author_url": "",
          "post_date": "2023-08-29T10:40:32.293000",
          "content": "<p>Yes, I did.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2416128,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2023-08-30T19:55:03.853000",
      "content": "<blockquote>\n  <p>a lot of the data is mislabelled or of bad audio quality</p>\n</blockquote>\n<p>Just curious, how did you know that?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2416710,
          "author_name": "Umong Sain",
          "author_url": "",
          "post_date": "2023-08-31T07:11:57.740000",
          "content": "<p><a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">This notebook</a> showed few examples where WER was high for YellowKing's model just because the audio were mislabelled. I tried to filter out audio using different thresholds for MOS and YKG's WER, but failed. Everytime I found some examples with mislabelled data or data where the speakers were mumbling a lot.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2416776,
              "author_name": "Ahmed Imtiaz Humayun",
              "author_url": "",
              "post_date": "2023-08-31T07:49:48.020000",
              "content": "<p>This is interesting, could you please elaborate how you were thresholding using YKG WER+MOS?<br>\nAlso, we have provided clientID which could help find out which contributors were mumbling too much or were very noisy. Remember that a little bit of label noise would be allowable and even help regularize the network.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2416786,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2023-08-31T07:59:14.043000",
          "content": "<p>It seems only 60-70% of the train data is usable.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2424327,
              "author_name": "AIFahim",
              "author_url": "",
              "post_date": "2023-09-05T07:24:02.590000",
              "content": "<p>how could you identify this 60-70% of the data that is usable? did you rely on MOS &amp; YKG's WER correlation?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2415110,
      "author_name": "Aphysict",
      "author_url": "",
      "post_date": "2023-08-30T05:51:44.180000",
      "content": "<p>Good job, does this dataset have overlap with the competition dataset? The valid split of the competition dataset already has 29588 samples, so we get 20k more clean samples? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2415160,
          "author_name": "Umong Sain",
          "author_url": "",
          "post_date": "2023-08-30T06:44:19.330000",
          "content": "<p>Not sure though. The organizer gave the mappings to common voice id for the train split. So the validation split might have data from common voice.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2416762,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-08-31T07:41:34.217000",
          "content": "<p><a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> <a href=\"https://www.kaggle.com/aphysict\" target=\"_blank\">@aphysict</a> </p>\n<p>Will be making the validation split quality+mappings available in a few days. The validation split (in the competition dataset) has all the upvoted samples from CV12 + more of which were evaluated by our in house teams.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2443608,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2023-09-17T20:14:08.337000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> </p>\n<p>Thank you for sharing this dataset. How did you prepare the other part of the dataset? Still the same approach as train - more upvotes than the downvotes?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2419534,
      "author_name": "Pritam Sinha",
      "author_url": "",
      "post_date": "2023-09-02T04:54:46.713000",
      "content": "<p>Hi, I want to know how did you train/finetuning the model. Can you share the training notebook if possible? Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2417044,
      "author_name": "Balaji Selvaraj",
      "author_url": "",
      "post_date": "2023-08-31T11:16:47.113000",
      "content": "<p>I am following similar setup as Tawara<br>\n<a href=\"https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\" target=\"_blank\">https://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline</a></p>\n<p>When I kept the vocab as such and I finetuned my model, I was able to get 0.439 in 2 epochs</p>\n<p>When I changed the vocab to the one show in the following link, my performance dropped to 0.503<br>\n<a href=\"https://www.kaggle.com/code/umongsain/macro-normalization\" target=\"_blank\">https://www.kaggle.com/code/umongsain/macro-normalization</a></p>\n<p>Is there some thing that I need to do ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2417696,
          "author_name": "Umong Sain",
          "author_url": "",
          "post_date": "2023-08-31T19:10:38.847000",
          "content": "<p>Changing vocab is not enough. You have tune after changing the vocab</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2418352,
              "author_name": "Balaji Selvaraj",
              "author_url": "",
              "post_date": "2023-09-01T08:53:02.603000",
              "content": "<p>I finetuned for 2 epochs after changing vocab.. <br>\nThe eval loss and wer reduced a lot but the LB score became 0.503</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2416856,
      "author_name": "AIFahim",
      "author_url": "",
      "post_date": "2023-08-31T08:55:59.287000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4350309%2F820297538c297ca97453dc05b450a466%2F2023-08-31_14-54.png?generation=1693472073318037&amp;alt=media\" alt=\"\"><br>\nThanks for you contribution <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a>. How did you include LM part? Did you use Wav2Vec2ProcessorWithLM?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2416881,
          "author_name": "Umong Sain",
          "author_url": "",
          "post_date": "2023-08-31T09:13:29.713000",
          "content": "<p>I used <code>pyctcdecode</code> to decode the logits. <code>Wav2Vec2ProcessorWithLM</code> is not compatible with new MMS Adapters yet.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2413797,
      "author_name": "HuBERT",
      "author_url": "",
      "post_date": "2023-08-29T07:06:31.447000",
      "content": "<p>How did you manage to submit such a large model? Did you do it with the Kenlm model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2414036,
          "author_name": "Umong Sain",
          "author_url": "",
          "post_date": "2023-08-29T10:40:11.990000",
          "content": "<p>With a KenLM 5-gram model, the inference time is almost 4 hours.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2415524,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-30T12:19:53.173000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2413302": "Training on 960K+ data is massive. Training for just a single epoch might take over 100 hours, and the worst part is that a lot of the data is mislabelled or of bad audio quality, which could do more harm than good. Instead of using the entire dataset, using a subset of good data would yield better results. [This notebook](https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda), by @imtiazprio and @reasat, explores various methods to filter out good-quality data. Since Common Voice is a subset of the MaCro dataset, I have chosen data samples from the Common Voice dataset, specifically those with more upvotes than downvotes. I have already **removed punctuations**, **normalized the texts**, and fixed problematic entries. For those with limited resources, I recommend using only the `train` split for training, which contains 20k samples. If more samples are needed, the `other` split (47k+ samples) can be combined with the `train` split.\n\nI have already trained a model using only the `train` split. When coupled with a language model, it outperforms the model created by YellowKing in the public LB.\n\n1. Dataset link: https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized\n2. Pretrained Model: https://huggingface.co/Umong/wav2vec2-large-mms-1b-bengali\n\nFeel free to provide feedback. Happy Kaggling!",
    "2419777": "Thank you for sharing.\n\nI just made submission by your model and got 0.451 public score.  \nIt seems that the model will get higher score than `ai4bharat/indicwav2vec_v1_bengali` by fine-tuning.",
    "2413311": "Text corpus is available for those who are using a language model: https://www.kaggle.com/datasets/umongsain/indiccorp-v2-bengali\nEnsure that you **normalize the data** before training the language model.",
    "2413994": "Hello Umong. \n\nDid you employ the vocabulary and normalisation strategy which you described in the following link\n\nhttps://www.kaggle.com/code/umongsain/macro-normalization",
    "2416128": "> a lot of the data is mislabelled or of bad audio quality\n\nJust curious, how did you know that?",
    "2415110": "Good job, does this dataset have overlap with the competition dataset? The valid split of the competition dataset already has 29588 samples, so we get 20k more clean samples? ",
    "2443608": "Hi @umongsain \n\nThank you for sharing this dataset. How did you prepare the other part of the dataset? Still the same approach as train - more upvotes than the downvotes?",
    "2419534": "Hi, I want to know how did you train/finetuning the model. Can you share the training notebook if possible? Thanks",
    "2417044": "I am following similar setup as Tawara\nhttps://www.kaggle.com/code/ttahara/bengali-sr-public-wav2vec2-0-w-lm-baseline\n\nWhen I kept the vocab as such and I finetuned my model, I was able to get 0.439 in 2 epochs\n\nWhen I changed the vocab to the one show in the following link, my performance dropped to 0.503\nhttps://www.kaggle.com/code/umongsain/macro-normalization\n\nIs there some thing that I need to do ?",
    "2416856": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4350309%2F820297538c297ca97453dc05b450a466%2F2023-08-31_14-54.png?generation=1693472073318037&alt=media)\nThanks for you contribution @umongsain. How did you include LM part? Did you use Wav2Vec2ProcessorWithLM?",
    "2413797": "How did you manage to submit such a large model? Did you do it with the Kenlm model?"
  }
}