{
  "id": 433722,
  "title": "[LB 0.445] Finetuning is the key 🙆‍♂️🙆‍♂️",
  "url": "/competitions/bengaliai-speech/discussion/433722",
  "author_name": "Nischay Dhankhar",
  "post_date": "2023-08-22T17:19:15.030000",
  "votes": 47,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Just started off with the competition and sharing the current progress publicly. I think many teams are still using pretrained models from HuggingFace, but at the end most of the winning solutions are going to be the one which are finetuned on competition data. </p>\n<p>Anyways, current solution is based on <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> great <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/432791\" target=\"_blank\">baseline</a>. I basically finetuned the <em>Wav2vec2CTC</em> model &amp; kept the pretrained LM model same as in the baseline. </p>\n<p>Here is the trained model checkpoint: <a href=\"https://www.kaggle.com/datasets/nischaydnk/bengali-wav2vec2-finetuned\" target=\"_blank\">dataset</a><br>\nInference notebook: <a href=\"https://www.kaggle.com/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference\" target=\"_blank\">link</a></p>\n<p>I will release the training code in upcoming days, yet to clean it. </p>\n<p>Here are current settings I used in training:</p>\n<ul>\n<li>Only 10% of training data was used.</li>\n<li>Freeze encoder layers</li>\n<li>loss function: CTC</li>\n<li>5 epochs, Batch size 4, lr: 5e-5</li>\n</ul>\n<p>I applied random kfold on train+valid samples, hence cv seems quite high. </p>\n<p>CV: 0.465 ---&gt; 0.433<br>\nLB: 0.471  ---&gt; 0.445</p>\n<p>Will keep updated of more experiments in the thread, happy kaggling!! </p>",
  "messages": [
    {
      "id": 2403426,
      "postDate": "2023-08-22T17:19:15.030Z",
      "content": "<p>Just started off with the competition and sharing the current progress publicly. I think many teams are still using pretrained models from HuggingFace, but at the end most of the winning solutions are going to be the one which are finetuned on competition data. </p>\n<p>Anyways, current solution is based on <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> great <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/432791\" target=\"_blank\">baseline</a>. I basically finetuned the <em>Wav2vec2CTC</em> model &amp; kept the pretrained LM model same as in the baseline. </p>\n<p>Here is the trained model checkpoint: <a href=\"https://www.kaggle.com/datasets/nischaydnk/bengali-wav2vec2-finetuned\" target=\"_blank\">dataset</a><br>\nInference notebook: <a href=\"https://www.kaggle.com/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference\" target=\"_blank\">link</a></p>\n<p>I will release the training code in upcoming days, yet to clean it. </p>\n<p>Here are current settings I used in training:</p>\n<ul>\n<li>Only 10% of training data was used.</li>\n<li>Freeze encoder layers</li>\n<li>loss function: CTC</li>\n<li>5 epochs, Batch size 4, lr: 5e-5</li>\n</ul>\n<p>I applied random kfold on train+valid samples, hence cv seems quite high. </p>\n<p>CV: 0.465 ---&gt; 0.433<br>\nLB: 0.471  ---&gt; 0.445</p>\n<p>Will keep updated of more experiments in the thread, happy kaggling!! </p>",
      "rawMarkdown": "Just started off with the competition and sharing the current progress publicly. I think many teams are still using pretrained models from HuggingFace, but at the end most of the winning solutions are going to be the one which are finetuned on competition data. \n\nAnyways, current solution is based on @ttahara great [baseline](https://www.kaggle.com/competitions/bengaliai-speech/discussion/432791). I basically finetuned the *Wav2vec2CTC* model & kept the pretrained LM model same as in the baseline. \n\nHere is the trained model checkpoint: [dataset](https://www.kaggle.com/datasets/nischaydnk/bengali-wav2vec2-finetuned)\nInference notebook: [link](https://www.kaggle.com/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference)\n\nI will release the training code in upcoming days, yet to clean it. \n\nHere are current settings I used in training:\n- Only 10% of training data was used.\n- Freeze encoder layers\n-  loss function: CTC\n- 5 epochs, Batch size 4, lr: 5e-5\n\nI applied random kfold on train+valid samples, hence cv seems quite high. \n\nCV: 0.465 ---> 0.433\nLB: 0.471  ---> 0.445\n\nWill keep updated of more experiments in the thread, happy kaggling!! ",
      "votes": 47
    },
    {
      "id": 2449745,
      "postDate": "2023-09-21T11:59:17.127Z",
      "content": "<p>Hello! How did you implement KFold in Wav2Vec2 training? I wonder if I can do it with a continuous wandb or huggingface record</p>",
      "rawMarkdown": "Hello! How did you implement KFold in Wav2Vec2 training? I wonder if I can do it with a continuous wandb or huggingface record",
      "votes": 1
    },
    {
      "id": 2431075,
      "postDate": "2023-09-09T18:51:50.360Z",
      "content": "<p>Onyx shadows cast,<br>\nDarkened whispers from the past,<br>\nSorrows that will last.</p>",
      "rawMarkdown": "Onyx shadows cast,\nDarkened whispers from the past,\nSorrows that will last.",
      "votes": 1
    },
    {
      "id": 2413731,
      "postDate": "2023-08-29T06:08:52.633Z",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a><br>\nPosting my questione here too as I got no reply when commenting under code.<br>\n<a href=\"https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference/comments\" target=\"_blank\">https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference/comments</a></p>\n<p>Question: When I run code cells in this notebook it fails and when I submit it, it get 0.4455 score. <br>\nThe link includes the screen shot when it fails.</p>\n<p>Can others try and report back if possible.</p>",
      "rawMarkdown": "@nischaydnk\nPosting my questione here too as I got no reply when commenting under code.\nhttps://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference/comments\n\nQuestion: When I run code cells in this notebook it fails and when I submit it, it get 0.4455 score. \nThe link includes the screen shot when it fails.\n\nCan others try and report back if possible.\n",
      "votes": 1,
      "replies": [
        {
          "id": 2421957,
          "postDate": "2023-09-03T16:00:20.093Z",
          "content": "<p>I got the same issue like you, not sure how to solve it though.  I tried installing the requested numpy, use latest environment, enabling internet - all didn't work. </p>",
          "rawMarkdown": "I got the same issue like you, not sure how to solve it though.  I tried installing the requested numpy, use latest environment, enabling internet - all didn't work. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2411352,
      "postDate": "2023-08-27T14:54:59.633Z",
      "content": "<p>Thank you for sharing. But I think there is something missing with the model checkpoints. I got \"Some weights of Wav2Vec2ForCTC were not initialized from the model checkpoint at /kaggle/input/bengali-wav2vec2-finetuned and are newly initialized: ['wav2vec2.masked_spec_embed']\" just running the inference notebook. Can you take a look? Thank you!</p>",
      "rawMarkdown": "Thank you for sharing. But I think there is something missing with the model checkpoints. I got \"Some weights of Wav2Vec2ForCTC were not initialized from the model checkpoint at /kaggle/input/bengali-wav2vec2-finetuned and are newly initialized: ['wav2vec2.masked_spec_embed']\" just running the inference notebook. Can you take a look? Thank you!",
      "votes": 1
    },
    {
      "id": 2404636,
      "postDate": "2023-08-23T11:57:18.107Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> Great job! </p>",
      "rawMarkdown": "Thanks for sharing @nischaydnk Great job! ",
      "votes": 1
    },
    {
      "id": 2403672,
      "postDate": "2023-08-22T19:35:38.050Z",
      "content": "<p>Thanks for sharing your knowkedge. I'm glad that my baseline was helpful to you :)</p>",
      "rawMarkdown": "Thanks for sharing your knowkedge. I'm glad that my baseline was helpful to you :)",
      "votes": 1,
      "replies": [
        {
          "id": 2404017,
          "postDate": "2023-08-23T02:36:41.757Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, have been following your work since long time 🫡</p>",
          "rawMarkdown": "Thank you @ttahara, have been following your work since long time 🫡"
        }
      ]
    },
    {
      "id": 2403649,
      "postDate": "2023-08-22T19:09:45.233Z",
      "content": "<p>Very nice my friend. Fine tuning and pre-training all the way. </p>\n<p>Any plans to do fine-tuning with adapters? Plan on reading this Hugging Face blog post on it, but seems like adapters aren't as easy for audio as text or CV: <a href=\"url\" target=\"_blank\">https://huggingface.co/blog/mms_adapters</a></p>",
      "rawMarkdown": "Very nice my friend. Fine tuning and pre-training all the way. \n\nAny plans to do fine-tuning with adapters? Plan on reading this Hugging Face blog post on it, but seems like adapters aren't as easy for audio as text or CV: [https://huggingface.co/blog/mms_adapters](url)",
      "votes": 1,
      "replies": [
        {
          "id": 2404025,
          "postDate": "2023-08-23T02:40:31.983Z",
          "content": "<p>Thanks for sharing. Definitely make sense to use the adapters, I will try. Currently, I froze the entire encoder layers and finetuned it.</p>",
          "rawMarkdown": "Thanks for sharing. Definitely make sense to use the adapters, I will try. Currently, I froze the entire encoder layers and finetuned it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2403447,
      "postDate": "2023-08-22T17:32:19.723Z",
      "content": "<p>Great result! Why did you use only 10% of training data?</p>",
      "rawMarkdown": "Great result! Why did you use only 10% of training data?",
      "votes": 1,
      "replies": [
        {
          "id": 2403451,
          "postDate": "2023-08-22T17:34:15.173Z",
          "content": "<p>ASR training is quite computationally expensive, took me around 8 hours to train on just 10% data. Using more data should improve the score further. </p>",
          "rawMarkdown": "ASR training is quite computationally expensive, took me around 8 hours to train on just 10% data. Using more data should improve the score further. ",
          "votes": 3,
          "replies": [
            {
              "id": 2403542,
              "postDate": "2023-08-22T17:58:59.290Z",
              "content": "<p>Ah, understandable 😂… Never worked with ASR yet, so I thought maybe it is some special strategy for not overfitting on this kind of data. Thank you for your post. Very interested in this competition, will learn more about everything soon. </p>",
              "rawMarkdown": "Ah, understandable 😂... Never worked with ASR yet, so I thought maybe it is some special strategy for not overfitting on this kind of data. Thank you for your post. Very interested in this competition, will learn more about everything soon. ",
              "votes": 1
            },
            {
              "id": 2405618,
              "postDate": "2023-08-24T03:00:23.217Z",
              "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>  I hope you don't mind I continue my question here. You mentioned using A6000 machine which has 48 gig on board memory. Is this a necessity to train ?  <br>\nIn short, do you think it's possible to train the model with a smaller GPU (24 gig memory)?</p>",
              "rawMarkdown": "@nischaydnk  I hope you don't mind I continue my question here. You mentioned using A6000 machine which has 48 gig on board memory. Is this a necessity to train ?  \nIn short, do you think it's possible to train the model with a smaller GPU (24 gig memory)?",
              "votes": 1
            },
            {
              "id": 2405656,
              "postDate": "2023-08-24T03:44:43.177Z",
              "content": "<p>Not a necessity at all. model should be able to fit with batch size 4 (with gradient checkpointing) using 24 gigs. You can try it out. </p>",
              "rawMarkdown": "Not a necessity at all. model should be able to fit with batch size 4 (with gradient checkpointing) using 24 gigs. You can try it out. ",
              "votes": 2
            },
            {
              "id": 2411439,
              "postDate": "2023-08-27T16:04:40.327Z",
              "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>  You're right it can run, but it can't complete in 8 hour. In fact it can't even complete one epoch <br>\nAny advice ? I'm using the training code from here: <a href=\"https://www.kaggle.com/code/mbmmurad/lb-0-49-wav2vec2-baseline-train-and-infer\" target=\"_blank\">https://www.kaggle.com/code/mbmmurad/lb-0-49-wav2vec2-baseline-train-and-infer</a></p>",
              "rawMarkdown": "@nischaydnk  You're right it can run, but it can't complete in 8 hour. In fact it can't even complete one epoch \nAny advice ? I'm using the training code from here: https://www.kaggle.com/code/mbmmurad/lb-0-49-wav2vec2-baseline-train-and-infer"
            }
          ]
        }
      ]
    },
    {
      "id": 2412360,
      "postDate": "2023-08-28T08:33:59.733Z",
      "content": "<p>How long does it take to train 10% in 5 epochs? In my environment, it will take 100h or so.</p>",
      "rawMarkdown": "How long does it take to train 10% in 5 epochs? In my environment, it will take 100h or so.",
      "votes": 2,
      "replies": [
        {
          "id": 2412995,
          "postDate": "2023-08-28T16:09:11.490Z",
          "content": "<p>I think they froze the entire encoder layers and fine-tuned.</p>",
          "rawMarkdown": "I think they froze the entire encoder layers and fine-tuned.",
          "replies": [
            {
              "id": 2413012,
              "postDate": "2023-08-28T16:21:48.373Z",
              "content": "<p>I froze those too with<br>\n<code>model.freeze_feature_encoder()</code></p>",
              "rawMarkdown": "I froze those too with\n`model.freeze_feature_encoder()`"
            },
            {
              "id": 2413177,
              "postDate": "2023-08-28T17:41:30.740Z",
              "content": "<p><code>model.freeze_feature_encoder()</code> only freezes the latent feature encoder (1D CNN). The output of the feature encoder is then passed into the transformer encoder. To freeze all the encoder layers, use <code>model.freeze_base_model()</code>.</p>",
              "rawMarkdown": "`model.freeze_feature_encoder()` only freezes the latent feature encoder (1D CNN). The output of the feature encoder is then passed into the transformer encoder. To freeze all the encoder layers, use `model.freeze_base_model()`.",
              "votes": 3
            },
            {
              "id": 2413189,
              "postDate": "2023-08-28T17:50:54.633Z",
              "content": "<p><code>model.freeze_base_model</code> freezes almost all layers except last lm_head.<br>\nBut yes, it improves training speed(and sacrifices accuracy). I will see the result, thanks.</p>",
              "rawMarkdown": "`model.freeze_base_model` freezes almost all layers except last lm_head.\nBut yes, it improves training speed(and sacrifices accuracy). I will see the result, thanks.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2413196,
          "postDate": "2023-08-28T17:56:54.467Z",
          "content": "<p>If you want to speed up training, I explored using LORA and the speed up was pretty good. I don't think hugging face's PEFT package has an easy plug in for wav2vec2 models, so I manually added LORA to linear layers like this:</p>\n<pre><code> peft  LoraConfig, TaskType, get_peft_model\n\npeft_config = LoraConfig(task_type=TaskType.FEATURE_EXTRACTION, \n                         inference_mode=, \n                         r=, \n                         lora_alpha=, \n                         lora_dropout=,\n                         target_modules=[\n                             name  name, module  model.named_modules()  (module, torch.nn.Linear)\n                        ])\n\nlora_model = get_peft_model(model, peft_config)\n</code></pre>",
          "rawMarkdown": "If you want to speed up training, I explored using LORA and the speed up was pretty good. I don't think hugging face's PEFT package has an easy plug in for wav2vec2 models, so I manually added LORA to linear layers like this:\n\n```python\nfrom peft import LoraConfig, TaskType, get_peft_model\n\npeft_config = LoraConfig(task_type=TaskType.FEATURE_EXTRACTION, \n                         inference_mode=False, \n                         r=8, \n                         lora_alpha=32, \n                         lora_dropout=0.1,\n                         target_modules=[\n                             name for name, module in model.named_modules() if isinstance(module, torch.nn.Linear)\n                        ])\n\nlora_model = get_peft_model(model, peft_config)\n```",
          "votes": 3
        },
        {
          "id": 2430044,
          "postDate": "2023-09-09T03:23:17.887Z",
          "content": "<p>How to you train the model? Any hint. I am kind of stuck…</p>",
          "rawMarkdown": "How to you train the model? Any hint. I am kind of stuck..."
        }
      ]
    },
    {
      "id": 2403954,
      "postDate": "2023-08-23T01:09:44.020Z",
      "content": "<p>Thanks for your sharing! Such a great progress.<br>\nCan't wait to learn from your training code!</p>",
      "rawMarkdown": "Thanks for your sharing! Such a great progress.\nCan't wait to learn from your training code!",
      "votes": 2
    },
    {
      "id": 2403917,
      "postDate": "2023-08-23T00:03:34.573Z",
      "content": "<p>Great job! Thanks for sharing!</p>",
      "rawMarkdown": "Great job! Thanks for sharing!",
      "votes": 2,
      "replies": [
        {
          "id": 2404019,
          "postDate": "2023-08-23T02:37:51.173Z",
          "content": "<p>Thank you 🫡 curious, are you planning to join the competition? 🙃</p>",
          "rawMarkdown": "Thank you 🫡 curious, are you planning to join the competition? 🙃",
          "replies": [
            {
              "id": 2404559,
              "postDate": "2023-08-23T10:09:11.740Z",
              "content": "<p>both finger spell and this speech competition are using the same methods, CTC, RNNT, LM,. etc…<br>\nI will be back after the finger spell.</p>",
              "rawMarkdown": "both finger spell and this speech competition are using the same methods, CTC, RNNT, LM,. etc...\nI will be back after the finger spell.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2404246,
      "postDate": "2023-08-23T06:12:03.377Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>\n<p>thanks for your explanation and details about your training. I really appreciate that you are sharing your insights and details. However, I would like you to remove the inference notebook and the weights. You are sharing weights here that were in the gold medal zone yesterday. A lot of people have been working hard on this competition for weeks and one day later there is a notebook that is being forked and submitted en masse which scores in the  ~top 10. I realize it's still early in the competition, but it's been going on for a few weeks. Therefore, I think it would be better if you give tips to get results yourself instead of sharing the weights and the code 1-to-1.</p>\n<p>Best wishes,<br>\nBenedikt</p>",
      "rawMarkdown": "Hey @nischaydnk \n\nthanks for your explanation and details about your training. I really appreciate that you are sharing your insights and details. However, I would like you to remove the inference notebook and the weights. You are sharing weights here that were in the gold medal zone yesterday. A lot of people have been working hard on this competition for weeks and one day later there is a notebook that is being forked and submitted en masse which scores in the  ~top 10. I realize it's still early in the competition, but it's been going on for a few weeks. Therefore, I think it would be better if you give tips to get results yourself instead of sharing the weights and the code 1-to-1.\n\nBest wishes,\nBenedikt",
      "votes": 1,
      "replies": [
        {
          "id": 2404353,
          "postDate": "2023-08-23T07:39:16.410Z",
          "content": "<p>Hey, I completely understand what you are trying to say and apologies if it levelled up the leaderboard for now. But it's going to be temporary, the thing is it's still too early in the competition. 2 months from now, I don't think this even score as good as 0.440 will be in medal zone. My only purpose was to show impact of fine-tuning in the competition, the model is just trained on little data with no parameter tuning. Noting your point, I will not share the training script complete end to end or any better scoring notebook. </p>\n<p>Those who completely rely on fork and high scoring public won't be able to improve score further anyways. So, I will not be worried much about the weights. </p>",
          "rawMarkdown": "Hey, I completely understand what you are trying to say and apologies if it levelled up the leaderboard for now. But it's going to be temporary, the thing is it's still too early in the competition. 2 months from now, I don't think this even score as good as 0.440 will be in medal zone. My only purpose was to show impact of fine-tuning in the competition, the model is just trained on little data with no parameter tuning. Noting your point, I will not share the training script complete end to end or any better scoring notebook. \n\nThose who completely rely on fork and high scoring public won't be able to improve score further anyways. So, I will not be worried much about the weights. \n\n",
          "votes": 12,
          "replies": [
            {
              "id": 2404357,
              "postDate": "2023-08-23T07:44:22.107Z",
              "content": "<p>Thank you for your kind reply. I would like to emphasize again that I think it is very good that you share your insights here. I'm just a fan of giving more guidance on how to help yourself instead of presenting the solutions ;). Anyway, I wish you only the best in competition.</p>\n<p>Best wishes,<br>\nBenedikt</p>",
              "rawMarkdown": "Thank you for your kind reply. I would like to emphasize again that I think it is very good that you share your insights here. I'm just a fan of giving more guidance on how to help yourself instead of presenting the solutions ;). Anyway, I wish you only the best in competition.\n\nBest wishes,\nBenedikt",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2462485,
      "postDate": "2023-09-30T11:03:10.973Z",
      "content": "<p>How much does unfreezing decrease the error rate?</p>",
      "rawMarkdown": "How much does unfreezing decrease the error rate?"
    },
    {
      "id": 2460367,
      "postDate": "2023-09-28T18:15:40.280Z",
      "content": "<p>thanks for the notebook<br>\nI cloned the LM_PATH (!git clone <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali</a><br>\n)but I can't find the 5gram.bin (I can only find pytorch_model.bin and when I use it gives  error:<br>\nCannot read model '/home/asr_r02/.cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin' (lm/read_arpa.cc:65 in void lm::ReadARPACounts(util::FilePiece&amp;, std::vector&amp;) threw FormatLoadException. first non-empty line was \"PK\u0003\u0004) \")</p>\n<p>thanks </p>",
      "rawMarkdown": "thanks for the notebook\nI cloned the LM_PATH (!git clone https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\n)but I can't find the 5gram.bin (I can only find pytorch_model.bin and when I use it gives  error:\nCannot read model '/home/asr_r02/.cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin' (lm/read_arpa.cc:65 in void lm::ReadARPACounts(util::FilePiece&, std::vector<long unsigned int>&) threw FormatLoadException. first non-empty line was \"PK\u0003\u0004) \")\n\nthanks "
    },
    {
      "id": 2450115,
      "postDate": "2023-09-21T16:14:42.533Z",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> Great Notebook! Did you finetune the \"facebook/wav2vec2-large-960h-lv60-self\" model or the \"ai4bharat/indicwav2vec_v1_bengali\" model?</p>",
      "rawMarkdown": "@nischaydnk Great Notebook! Did you finetune the \"facebook/wav2vec2-large-960h-lv60-self\" model or the \"ai4bharat/indicwav2vec_v1_bengali\" model?"
    },
    {
      "id": 2422822,
      "postDate": "2023-09-04T09:18:51.013Z",
      "content": "<p>Hi,</p>\n<p>If you don't mind, which model did you fine-tune? the \"facebook/wav2vec2-large-960h-lv60-self\" or the \"ai4bharat/indicwav2vec_v1_bengali\" model? I think the latter started from the former already but I am not sure. Thanks!</p>",
      "rawMarkdown": "Hi,\n\nIf you don't mind, which model did you fine-tune? the \"facebook/wav2vec2-large-960h-lv60-self\" or the \"ai4bharat/indicwav2vec_v1_bengali\" model? I think the latter started from the former already but I am not sure. Thanks!"
    },
    {
      "id": 2422183,
      "postDate": "2023-09-03T18:50:45.107Z",
      "content": "<p>lora finetuning?</p>",
      "rawMarkdown": "lora finetuning?"
    },
    {
      "id": 2421182,
      "postDate": "2023-09-03T08:10:11.397Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>. It will be quite beneficial.</p>",
      "rawMarkdown": "Thanks for sharing @nischaydnk. It will be quite beneficial."
    },
    {
      "id": 2408912,
      "postDate": "2023-08-25T22:45:15.607Z",
      "content": "<p>I know you're probably still working on the training notebook, but curious if you used the datasets library from Hugging Face or created your own torch dataset. I was having issues with datasets and resorted to using a torch dataset to process the raw waveform and do feature extraction.</p>",
      "rawMarkdown": "I know you're probably still working on the training notebook, but curious if you used the datasets library from Hugging Face or created your own torch dataset. I was having issues with datasets and resorted to using a torch dataset to process the raw waveform and do feature extraction."
    },
    {
      "id": 2408877,
      "postDate": "2023-08-25T21:40:02.203Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> Great work!</p>",
      "rawMarkdown": "Thanks for sharing @nischaydnk Great work!"
    },
    {
      "id": 2406964,
      "postDate": "2023-08-24T18:11:03.313Z",
      "content": "<p>Thanks for the information. I would like to see how to fine-tune the Wav2Vec2CTC model. Please do upload the training code and your approach whenever possible. Is there any resource showing how to fine-tune the Wav2Vec2CTC model as you did?</p>",
      "rawMarkdown": "Thanks for the information. I would like to see how to fine-tune the Wav2Vec2CTC model. Please do upload the training code and your approach whenever possible. Is there any resource showing how to fine-tune the Wav2Vec2CTC model as you did?"
    },
    {
      "id": 2406893,
      "postDate": "2023-08-24T17:39:03.497Z",
      "content": "<p>sir you are a inspiration!</p>",
      "rawMarkdown": "sir you are a inspiration!"
    },
    {
      "id": 2468969,
      "postDate": "2023-10-06T03:24:21.797Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2424081,
      "postDate": "2023-09-05T03:26:59.750Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2449745,
      "author_name": "iiiiitsu",
      "author_url": "",
      "post_date": "2023-09-21T11:59:17.127000",
      "content": "<p>Hello! How did you implement KFold in Wav2Vec2 training? I wonder if I can do it with a continuous wandb or huggingface record</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2431075,
      "author_name": "Pizzaboi",
      "author_url": "",
      "post_date": "2023-09-09T18:51:50.360000",
      "content": "<p>Onyx shadows cast,<br>\nDarkened whispers from the past,<br>\nSorrows that will last.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2413731,
      "author_name": "Blue",
      "author_url": "",
      "post_date": "2023-08-29T06:08:52.633000",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a><br>\nPosting my questione here too as I got no reply when commenting under code.<br>\n<a href=\"https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference/comments\" target=\"_blank\">https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference/comments</a></p>\n<p>Question: When I run code cells in this notebook it fails and when I submit it, it get 0.4455 score. <br>\nThe link includes the screen shot when it fails.</p>\n<p>Can others try and report back if possible.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2421957,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2023-09-03T16:00:20.093000",
          "content": "<p>I got the same issue like you, not sure how to solve it though.  I tried installing the requested numpy, use latest environment, enabling internet - all didn't work. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2411352,
      "author_name": "Zhenlan",
      "author_url": "",
      "post_date": "2023-08-27T14:54:59.633000",
      "content": "<p>Thank you for sharing. But I think there is something missing with the model checkpoints. I got \"Some weights of Wav2Vec2ForCTC were not initialized from the model checkpoint at /kaggle/input/bengali-wav2vec2-finetuned and are newly initialized: ['wav2vec2.masked_spec_embed']\" just running the inference notebook. Can you take a look? Thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2404636,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2023-08-23T11:57:18.107000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> Great job! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2403672,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2023-08-22T19:35:38.050000",
      "content": "<p>Thanks for sharing your knowkedge. I'm glad that my baseline was helpful to you :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2404017,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-08-23T02:36:41.757000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, have been following your work since long time 🫡</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2403649,
      "author_name": "Matt S.",
      "author_url": "",
      "post_date": "2023-08-22T19:09:45.233000",
      "content": "<p>Very nice my friend. Fine tuning and pre-training all the way. </p>\n<p>Any plans to do fine-tuning with adapters? Plan on reading this Hugging Face blog post on it, but seems like adapters aren't as easy for audio as text or CV: <a href=\"url\" target=\"_blank\">https://huggingface.co/blog/mms_adapters</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2404025,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-08-23T02:40:31.983000",
          "content": "<p>Thanks for sharing. Definitely make sense to use the adapters, I will try. Currently, I froze the entire encoder layers and finetuned it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2403447,
      "author_name": "Man of the year",
      "author_url": "",
      "post_date": "2023-08-22T17:32:19.723000",
      "content": "<p>Great result! Why did you use only 10% of training data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2403451,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-08-22T17:34:15.173000",
          "content": "<p>ASR training is quite computationally expensive, took me around 8 hours to train on just 10% data. Using more data should improve the score further. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2403542,
              "author_name": "Man of the year",
              "author_url": "",
              "post_date": "2023-08-22T17:58:59.290000",
              "content": "<p>Ah, understandable 😂… Never worked with ASR yet, so I thought maybe it is some special strategy for not overfitting on this kind of data. Thank you for your post. Very interested in this competition, will learn more about everything soon. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2405618,
              "author_name": "yukiya",
              "author_url": "",
              "post_date": "2023-08-24T03:00:23.217000",
              "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>  I hope you don't mind I continue my question here. You mentioned using A6000 machine which has 48 gig on board memory. Is this a necessity to train ?  <br>\nIn short, do you think it's possible to train the model with a smaller GPU (24 gig memory)?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2405656,
              "author_name": "Nischay Dhankhar",
              "author_url": "",
              "post_date": "2023-08-24T03:44:43.177000",
              "content": "<p>Not a necessity at all. model should be able to fit with batch size 4 (with gradient checkpointing) using 24 gigs. You can try it out. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2411439,
              "author_name": "yukiya",
              "author_url": "",
              "post_date": "2023-08-27T16:04:40.327000",
              "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>  You're right it can run, but it can't complete in 8 hour. In fact it can't even complete one epoch <br>\nAny advice ? I'm using the training code from here: <a href=\"https://www.kaggle.com/code/mbmmurad/lb-0-49-wav2vec2-baseline-train-and-infer\" target=\"_blank\">https://www.kaggle.com/code/mbmmurad/lb-0-49-wav2vec2-baseline-train-and-infer</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2412360,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2023-08-28T08:33:59.733000",
      "content": "<p>How long does it take to train 10% in 5 epochs? In my environment, it will take 100h or so.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2412995,
          "author_name": "Umong Sain",
          "author_url": "",
          "post_date": "2023-08-28T16:09:11.490000",
          "content": "<p>I think they froze the entire encoder layers and fine-tuned.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2413012,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2023-08-28T16:21:48.373000",
              "content": "<p>I froze those too with<br>\n<code>model.freeze_feature_encoder()</code></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2413177,
              "author_name": "Umong Sain",
              "author_url": "",
              "post_date": "2023-08-28T17:41:30.740000",
              "content": "<p><code>model.freeze_feature_encoder()</code> only freezes the latent feature encoder (1D CNN). The output of the feature encoder is then passed into the transformer encoder. To freeze all the encoder layers, use <code>model.freeze_base_model()</code>.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2413189,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2023-08-28T17:50:54.633000",
              "content": "<p><code>model.freeze_base_model</code> freezes almost all layers except last lm_head.<br>\nBut yes, it improves training speed(and sacrifices accuracy). I will see the result, thanks.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2413196,
          "author_name": "Matt S.",
          "author_url": "",
          "post_date": "2023-08-28T17:56:54.467000",
          "content": "<p>If you want to speed up training, I explored using LORA and the speed up was pretty good. I don't think hugging face's PEFT package has an easy plug in for wav2vec2 models, so I manually added LORA to linear layers like this:</p>\n<pre><code> peft  LoraConfig, TaskType, get_peft_model\n\npeft_config = LoraConfig(task_type=TaskType.FEATURE_EXTRACTION, \n                         inference_mode=, \n                         r=, \n                         lora_alpha=, \n                         lora_dropout=,\n                         target_modules=[\n                             name  name, module  model.named_modules()  (module, torch.nn.Linear)\n                        ])\n\nlora_model = get_peft_model(model, peft_config)\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2430044,
          "author_name": "Roy Wei",
          "author_url": "",
          "post_date": "2023-09-09T03:23:17.887000",
          "content": "<p>How to you train the model? Any hint. I am kind of stuck…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2403954,
      "author_name": "Rakuwa",
      "author_url": "",
      "post_date": "2023-08-23T01:09:44.020000",
      "content": "<p>Thanks for your sharing! Such a great progress.<br>\nCan't wait to learn from your training code!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2403917,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-08-23T00:03:34.573000",
      "content": "<p>Great job! Thanks for sharing!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2404019,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-08-23T02:37:51.173000",
          "content": "<p>Thank you 🫡 curious, are you planning to join the competition? 🙃</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2404559,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-23T10:09:11.740000",
              "content": "<p>both finger spell and this speech competition are using the same methods, CTC, RNNT, LM,. etc…<br>\nI will be back after the finger spell.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2404246,
      "author_name": "Benedikt Droste",
      "author_url": "",
      "post_date": "2023-08-23T06:12:03.377000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>\n<p>thanks for your explanation and details about your training. I really appreciate that you are sharing your insights and details. However, I would like you to remove the inference notebook and the weights. You are sharing weights here that were in the gold medal zone yesterday. A lot of people have been working hard on this competition for weeks and one day later there is a notebook that is being forked and submitted en masse which scores in the  ~top 10. I realize it's still early in the competition, but it's been going on for a few weeks. Therefore, I think it would be better if you give tips to get results yourself instead of sharing the weights and the code 1-to-1.</p>\n<p>Best wishes,<br>\nBenedikt</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2404353,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-08-23T07:39:16.410000",
          "content": "<p>Hey, I completely understand what you are trying to say and apologies if it levelled up the leaderboard for now. But it's going to be temporary, the thing is it's still too early in the competition. 2 months from now, I don't think this even score as good as 0.440 will be in medal zone. My only purpose was to show impact of fine-tuning in the competition, the model is just trained on little data with no parameter tuning. Noting your point, I will not share the training script complete end to end or any better scoring notebook. </p>\n<p>Those who completely rely on fork and high scoring public won't be able to improve score further anyways. So, I will not be worried much about the weights. </p>",
          "votes": 12,
          "replies": [
            {
              "id": 2404357,
              "author_name": "Benedikt Droste",
              "author_url": "",
              "post_date": "2023-08-23T07:44:22.107000",
              "content": "<p>Thank you for your kind reply. I would like to emphasize again that I think it is very good that you share your insights here. I'm just a fan of giving more guidance on how to help yourself instead of presenting the solutions ;). Anyway, I wish you only the best in competition.</p>\n<p>Best wishes,<br>\nBenedikt</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2462485,
      "author_name": "stas444",
      "author_url": "",
      "post_date": "2023-09-30T11:03:10.973000",
      "content": "<p>How much does unfreezing decrease the error rate?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2460367,
      "author_name": "Antar",
      "author_url": "",
      "post_date": "2023-09-28T18:15:40.280000",
      "content": "<p>thanks for the notebook<br>\nI cloned the LM_PATH (!git clone <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali</a><br>\n)but I can't find the 5gram.bin (I can only find pytorch_model.bin and when I use it gives  error:<br>\nCannot read model '/home/asr_r02/.cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin' (lm/read_arpa.cc:65 in void lm::ReadARPACounts(util::FilePiece&amp;, std::vector&amp;) threw FormatLoadException. first non-empty line was \"PK\u0003\u0004) \")</p>\n<p>thanks </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2450115,
      "author_name": "Md Boktiar Mahbub Murad",
      "author_url": "",
      "post_date": "2023-09-21T16:14:42.533000",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> Great Notebook! Did you finetune the \"facebook/wav2vec2-large-960h-lv60-self\" model or the \"ai4bharat/indicwav2vec_v1_bengali\" model?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2422822,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2023-09-04T09:18:51.013000",
      "content": "<p>Hi,</p>\n<p>If you don't mind, which model did you fine-tune? the \"facebook/wav2vec2-large-960h-lv60-self\" or the \"ai4bharat/indicwav2vec_v1_bengali\" model? I think the latter started from the former already but I am not sure. Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2422183,
      "author_name": "DestructiveForce",
      "author_url": "",
      "post_date": "2023-09-03T18:50:45.107000",
      "content": "<p>lora finetuning?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2421182,
      "author_name": "Sparsh Jain",
      "author_url": "",
      "post_date": "2023-09-03T08:10:11.397000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>. It will be quite beneficial.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2408912,
      "author_name": "Matt S.",
      "author_url": "",
      "post_date": "2023-08-25T22:45:15.607000",
      "content": "<p>I know you're probably still working on the training notebook, but curious if you used the datasets library from Hugging Face or created your own torch dataset. I was having issues with datasets and resorted to using a torch dataset to process the raw waveform and do feature extraction.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2408877,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-08-25T21:40:02.203000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> Great work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2406964,
      "author_name": "Pritam Sinha",
      "author_url": "",
      "post_date": "2023-08-24T18:11:03.313000",
      "content": "<p>Thanks for the information. I would like to see how to fine-tune the Wav2Vec2CTC model. Please do upload the training code and your approach whenever possible. Is there any resource showing how to fine-tune the Wav2Vec2CTC model as you did?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2406893,
      "author_name": "Mystic Shadow",
      "author_url": "",
      "post_date": "2023-08-24T17:39:03.497000",
      "content": "<p>sir you are a inspiration!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2468969,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-06T03:24:21.797000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2424081,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-05T03:26:59.750000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2403426": "Just started off with the competition and sharing the current progress publicly. I think many teams are still using pretrained models from HuggingFace, but at the end most of the winning solutions are going to be the one which are finetuned on competition data. \n\nAnyways, current solution is based on @ttahara great [baseline](https://www.kaggle.com/competitions/bengaliai-speech/discussion/432791). I basically finetuned the *Wav2vec2CTC* model & kept the pretrained LM model same as in the baseline. \n\nHere is the trained model checkpoint: [dataset](https://www.kaggle.com/datasets/nischaydnk/bengali-wav2vec2-finetuned)\nInference notebook: [link](https://www.kaggle.com/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference)\n\nI will release the training code in upcoming days, yet to clean it. \n\nHere are current settings I used in training:\n- Only 10% of training data was used.\n- Freeze encoder layers\n-  loss function: CTC\n- 5 epochs, Batch size 4, lr: 5e-5\n\nI applied random kfold on train+valid samples, hence cv seems quite high. \n\nCV: 0.465 ---> 0.433\nLB: 0.471  ---> 0.445\n\nWill keep updated of more experiments in the thread, happy kaggling!! ",
    "2449745": "Hello! How did you implement KFold in Wav2Vec2 training? I wonder if I can do it with a continuous wandb or huggingface record",
    "2431075": "Onyx shadows cast,\nDarkened whispers from the past,\nSorrows that will last.",
    "2413731": "@nischaydnk\nPosting my questione here too as I got no reply when commenting under code.\nhttps://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference/comments\n\nQuestion: When I run code cells in this notebook it fails and when I submit it, it get 0.4455 score. \nThe link includes the screen shot when it fails.\n\nCan others try and report back if possible.\n",
    "2411352": "Thank you for sharing. But I think there is something missing with the model checkpoints. I got \"Some weights of Wav2Vec2ForCTC were not initialized from the model checkpoint at /kaggle/input/bengali-wav2vec2-finetuned and are newly initialized: ['wav2vec2.masked_spec_embed']\" just running the inference notebook. Can you take a look? Thank you!",
    "2404636": "Thanks for sharing @nischaydnk Great job! ",
    "2403672": "Thanks for sharing your knowkedge. I'm glad that my baseline was helpful to you :)",
    "2403649": "Very nice my friend. Fine tuning and pre-training all the way. \n\nAny plans to do fine-tuning with adapters? Plan on reading this Hugging Face blog post on it, but seems like adapters aren't as easy for audio as text or CV: [https://huggingface.co/blog/mms_adapters](url)",
    "2403447": "Great result! Why did you use only 10% of training data?",
    "2412360": "How long does it take to train 10% in 5 epochs? In my environment, it will take 100h or so.",
    "2403954": "Thanks for your sharing! Such a great progress.\nCan't wait to learn from your training code!",
    "2403917": "Great job! Thanks for sharing!",
    "2404246": "Hey @nischaydnk \n\nthanks for your explanation and details about your training. I really appreciate that you are sharing your insights and details. However, I would like you to remove the inference notebook and the weights. You are sharing weights here that were in the gold medal zone yesterday. A lot of people have been working hard on this competition for weeks and one day later there is a notebook that is being forked and submitted en masse which scores in the  ~top 10. I realize it's still early in the competition, but it's been going on for a few weeks. Therefore, I think it would be better if you give tips to get results yourself instead of sharing the weights and the code 1-to-1.\n\nBest wishes,\nBenedikt",
    "2462485": "How much does unfreezing decrease the error rate?",
    "2460367": "thanks for the notebook\nI cloned the LM_PATH (!git clone https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\n)but I can't find the 5gram.bin (I can only find pytorch_model.bin and when I use it gives  error:\nCannot read model '/home/asr_r02/.cache/huggingface/hub/models--arijitx--wav2vec2-xls-r-300m-bengali/snapshots/45ed7c704f276acb9ed6ef234b66e67a1a2cb864/pytorch_model.bin' (lm/read_arpa.cc:65 in void lm::ReadARPACounts(util::FilePiece&, std::vector<long unsigned int>&) threw FormatLoadException. first non-empty line was \"PK\u0003\u0004) \")\n\nthanks ",
    "2450115": "@nischaydnk Great Notebook! Did you finetune the \"facebook/wav2vec2-large-960h-lv60-self\" model or the \"ai4bharat/indicwav2vec_v1_bengali\" model?",
    "2422822": "Hi,\n\nIf you don't mind, which model did you fine-tune? the \"facebook/wav2vec2-large-960h-lv60-self\" or the \"ai4bharat/indicwav2vec_v1_bengali\" model? I think the latter started from the former already but I am not sure. Thanks!",
    "2422183": "lora finetuning?",
    "2421182": "Thanks for sharing @nischaydnk. It will be quite beneficial.",
    "2408912": "I know you're probably still working on the training notebook, but curious if you used the datasets library from Hugging Face or created your own torch dataset. I was having issues with datasets and resorted to using a torch dataset to process the raw waveform and do feature extraction.",
    "2408877": "Thanks for sharing @nischaydnk Great work!",
    "2406964": "Thanks for the information. I would like to see how to fine-tune the Wav2Vec2CTC model. Please do upload the training code and your approach whenever possible. Is there any resource showing how to fine-tune the Wav2Vec2CTC model as you did?",
    "2406893": "sir you are a inspiration!",
    "2468969": "",
    "2424081": ""
  }
}