{
  "id": 448579,
  "title": "17th place : Curriculum learning ?? ",
  "url": "/competitions/bengaliai-speech/discussion/448579",
  "author_name": "",
  "post_date": "2023-10-20T09:35:05.517181500Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Frankly I'm not sure if it works ^^   I need to do ablation test if I can achieve the same results without ordering the training data. I'm packing it in 6 steps based on the WER and Mospred ,plus stage 1 includes CV13 data shared <a href=\"https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized\" target=\"_blank\">here</a> </p>\n<p>Training set :  </p>\n<ul>\n<li>Almost half of competition data plus CV13  (total 497574 samples)</li>\n</ul>\n<p>Augmentation:</p>\n<ul>\n<li>None except the default specaugment enabled  in the config. </li>\n</ul>\n<p>Model starts from : <br>\n<a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\" target=\"_blank\">https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali</a></p>\n<p>Loss curve : <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1440116%2Fcd7897061cf30485770715685785fb5c%2FlosscurveLB.jpg?generation=1697793073506232&amp;alt=media\" alt=\"\"></p>\n<p>The 0.405 model is my final submission, but with modified Language Model (kenLM 5 gram model consisting : ai4bharat text, hate speech, Quran text, poetry,  Shrutilipi, Samanantar, indicTTS, and BanglaLM text), and it improves the LB to 0.385</p>\n<p>I tried deepfilternet to denoise (trained to work on 16khz, no resampling) , but it consistently degrade the LB by 0.002 to 0.003. My offline test shows that it can only improve the score of \"certain\" class in OOD, but I didn't have time to build some sort of selection method. It's enlightening to see the 3rd solution manage to do it with demucs  (Personally, I'm a big fan of deepfilternet, demucs is too big ! ) </p>\n<p>Finally, something I miss : punctuations<br>\nI've plugged in the punctuation model from <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/447965\" target=\"_blank\">this solution</a> and it improve LB to 0.359 !   The private score is still not good enough for gold though, so at least that give me some peace at heart. Will try to do better next time 😁</p>\n<p>Thank for holding audio based competition, and special thanks to <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> and @ <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> without whose notebook, I would not have known how to start </p>\n<p>Inference code <a href=\"https://www.kaggle.com/code/nyleve/giantlm-dat6-step2-at-111k-add-punctuation?scriptVersionId=146943071\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "2489911",
      "postDate": "10/20/2023 09:35:05",
      "content": "<p>Frankly I'm not sure if it works ^^   I need to do ablation test if I can achieve the same results without ordering the training data. I'm packing it in 6 steps based on the WER and Mospred ,plus stage 1 includes CV13 data shared <a href=\"https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized\" target=\"_blank\">here</a> </p>\n<p>Training set :  </p>\n<ul>\n<li>Almost half of competition data plus CV13  (total 497574 samples)</li>\n</ul>\n<p>Augmentation:</p>\n<ul>\n<li>None except the default specaugment enabled  in the config. </li>\n</ul>\n<p>Model starts from : <br>\n<a href=\"https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\" target=\"_blank\">https://huggingface.co/ai4bharat/indicwav2vec_v1_bengali</a></p>\n<p>Loss curve : <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1440116%2Fcd7897061cf30485770715685785fb5c%2FlosscurveLB.jpg?generation=1697793073506232&amp;alt=media\" alt=\"\"></p>\n<p>The 0.405 model is my final submission, but with modified Language Model (kenLM 5 gram model consisting : ai4bharat text, hate speech, Quran text, poetry,  Shrutilipi, Samanantar, indicTTS, and BanglaLM text), and it improves the LB to 0.385</p>\n<p>I tried deepfilternet to denoise (trained to work on 16khz, no resampling) , but it consistently degrade the LB by 0.002 to 0.003. My offline test shows that it can only improve the score of \"certain\" class in OOD, but I didn't have time to build some sort of selection method. It's enlightening to see the 3rd solution manage to do it with demucs  (Personally, I'm a big fan of deepfilternet, demucs is too big ! ) </p>\n<p>Finally, something I miss : punctuations<br>\nI've plugged in the punctuation model from <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/447965\" target=\"_blank\">this solution</a> and it improve LB to 0.359 !   The private score is still not good enough for gold though, so at least that give me some peace at heart. Will try to do better next time 😁</p>\n<p>Thank for holding audio based competition, and special thanks to <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> and @ <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> without whose notebook, I would not have known how to start </p>\n<p>Inference code <a href=\"https://www.kaggle.com/code/nyleve/giantlm-dat6-step2-at-111k-add-punctuation?scriptVersionId=146943071\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Frankly I'm not sure if it works ^^   I need to do ablation test if I can achieve the same results without ordering the training data. I'm packing it in 6 steps based on the WER and Mospred ,plus stage 1 includes CV13 data shared [here](https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized) \n\nTraining set :  \n- Almost half of competition data plus CV13  (total 497574 samples)\n\nAugmentation:\n- None except the default specaugment enabled  in the config. \n\nModel starts from : \nhttps://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\n\nLoss curve : \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1440116%2Fcd7897061cf30485770715685785fb5c%2FlosscurveLB.jpg?generation=1697793073506232&alt=media)\n\nThe 0.405 model is my final submission, but with modified Language Model (kenLM 5 gram model consisting : ai4bharat text, hate speech, Quran text, poetry,  Shrutilipi, Samanantar, indicTTS, and BanglaLM text), and it improves the LB to 0.385\n\nI tried deepfilternet to denoise (trained to work on 16khz, no resampling) , but it consistently degrade the LB by 0.002 to 0.003. My offline test shows that it can only improve the score of \"certain\" class in OOD, but I didn't have time to build some sort of selection method. It's enlightening to see the 3rd solution manage to do it with demucs  (Personally, I'm a big fan of deepfilternet, demucs is too big ! ) \n\nFinally, something I miss : punctuations\nI've plugged in the punctuation model from [this solution](https://www.kaggle.com/competitions/bengaliai-speech/discussion/447965) and it improve LB to 0.359 !   The private score is still not good enough for gold though, so at least that give me some peace at heart. Will try to do better next time 😁\n\nThank for holding audio based competition, and special thanks to @mbmmurad and @ @ttahara without whose notebook, I would not have known how to start \n\nInference code [here](https://www.kaggle.com/code/nyleve/giantlm-dat6-step2-at-111k-add-punctuation?scriptVersionId=146943071)",
      "votes": null
    },
    {
      "id": "2491347",
      "postDate": "10/21/2023 14:37:50",
      "content": "<p>Congratulations on your placement. How does the WER metric compare to loss?</p>",
      "rawMarkdown": "Congratulations on your placement. How does the WER metric compare to loss?",
      "votes": null
    },
    {
      "id": "2492189",
      "postDate": "10/22/2023 10:16:31",
      "content": "<p>Curriculum Learning? I found you continue training for some steps! Confused</p>",
      "rawMarkdown": "Curriculum Learning? I found you continue training for some steps! Confused",
      "votes": null
    },
    {
      "id": "2492904",
      "postDate": "10/23/2023 02:42:00",
      "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> Therefore the question mark I put in the title :)  It's not exactly curriculum learning, just that the way I presented the training data follows some curriculum (decided from mos and wer) <br>\nJust one cycle did not give me anything, I let it continue few more cycles.<br>\nThe two dips in the loss are about 34K apart, and my training data is 500K with batch size 16. My guess is that it was one exact cycle (going through from easy to hard) </p>",
      "rawMarkdown": "aifahim Therefore the question mark I put in the title :)  It's not exactly curriculum learning, just that the way I presented the training data follows some curriculum (decided from mos and wer) \nJust one cycle did not give me anything, I let it continue few more cycles.\nThe two dips in the loss are about 34K apart, and my training data is 500K with batch size 16. My guess is that it was one exact cycle (going through from easy to hard)",
      "votes": null
    },
    {
      "id": "2492912",
      "postDate": "10/23/2023 02:57:58",
      "content": "<p>Thanks, possible that I just got lucky with the checkpoint ^^<br>\nThe WER is not as fancy, I post it below just for reference (I'm probing LB most of the time) <br>\nWhat I find interesting is the two dips in the loss, which I'm guessing is the time the cycle restarted (going through from easy samples to hard samples) <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1440116%2Fa9d96670559352f7314b2fe00008f88e%2Fwercurve.jpg?generation=1698029840132999&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks, possible that I just got lucky with the checkpoint ^^\nThe WER is not as fancy, I post it below just for reference (I'm probing LB most of the time) \nWhat I find interesting is the two dips in the loss, which I'm guessing is the time the cycle restarted (going through from easy samples to hard samples) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1440116%2Fa9d96670559352f7314b2fe00008f88e%2Fwercurve.jpg?generation=1698029840132999&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2491347,
      "author_name": "hubert101",
      "author_url": "",
      "post_date": "10/21/2023 14:37:50",
      "content": "<p>Congratulations on your placement. How does the WER metric compare to loss?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2492912,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "10/23/2023 02:57:58",
          "content": "<p>Thanks, possible that I just got lucky with the checkpoint ^^<br>\nThe WER is not as fancy, I post it below just for reference (I'm probing LB most of the time) <br>\nWhat I find interesting is the two dips in the loss, which I'm guessing is the time the cycle restarted (going through from easy samples to hard samples) <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1440116%2Fa9d96670559352f7314b2fe00008f88e%2Fwercurve.jpg?generation=1698029840132999&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2492189,
      "author_name": "aifahim",
      "author_url": "",
      "post_date": "10/22/2023 10:16:31",
      "content": "<p>Curriculum Learning? I found you continue training for some steps! Confused</p>",
      "votes": null,
      "replies": [
        {
          "id": 2492904,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "10/23/2023 02:42:00",
          "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> Therefore the question mark I put in the title :)  It's not exactly curriculum learning, just that the way I presented the training data follows some curriculum (decided from mos and wer) <br>\nJust one cycle did not give me anything, I let it continue few more cycles.<br>\nThe two dips in the loss are about 34K apart, and my training data is 500K with batch size 16. My guess is that it was one exact cycle (going through from easy to hard) </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2489911": "Frankly I'm not sure if it works ^^   I need to do ablation test if I can achieve the same results without ordering the training data. I'm packing it in 6 steps based on the WER and Mospred ,plus stage 1 includes CV13 data shared [here](https://www.kaggle.com/datasets/umongsain/common-voice-13-bengali-normalized) \n\nTraining set :  \n- Almost half of competition data plus CV13  (total 497574 samples)\n\nAugmentation:\n- None except the default specaugment enabled  in the config. \n\nModel starts from : \nhttps://huggingface.co/ai4bharat/indicwav2vec_v1_bengali\n\nLoss curve : \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1440116%2Fcd7897061cf30485770715685785fb5c%2FlosscurveLB.jpg?generation=1697793073506232&alt=media)\n\nThe 0.405 model is my final submission, but with modified Language Model (kenLM 5 gram model consisting : ai4bharat text, hate speech, Quran text, poetry,  Shrutilipi, Samanantar, indicTTS, and BanglaLM text), and it improves the LB to 0.385\n\nI tried deepfilternet to denoise (trained to work on 16khz, no resampling) , but it consistently degrade the LB by 0.002 to 0.003. My offline test shows that it can only improve the score of \"certain\" class in OOD, but I didn't have time to build some sort of selection method. It's enlightening to see the 3rd solution manage to do it with demucs  (Personally, I'm a big fan of deepfilternet, demucs is too big ! ) \n\nFinally, something I miss : punctuations\nI've plugged in the punctuation model from [this solution](https://www.kaggle.com/competitions/bengaliai-speech/discussion/447965) and it improve LB to 0.359 !   The private score is still not good enough for gold though, so at least that give me some peace at heart. Will try to do better next time 😁\n\nThank for holding audio based competition, and special thanks to @mbmmurad and @ @ttahara without whose notebook, I would not have known how to start \n\nInference code [here](https://www.kaggle.com/code/nyleve/giantlm-dat6-step2-at-111k-add-punctuation?scriptVersionId=146943071)",
    "2491347": "Congratulations on your placement. How does the WER metric compare to loss?",
    "2492189": "Curriculum Learning? I found you continue training for some steps! Confused",
    "2492904": "aifahim Therefore the question mark I put in the title :)  It's not exactly curriculum learning, just that the way I presented the training data follows some curriculum (decided from mos and wer) \nJust one cycle did not give me anything, I let it continue few more cycles.\nThe two dips in the loss are about 34K apart, and my training data is 500K with batch size 16. My guess is that it was one exact cycle (going through from easy to hard)",
    "2492912": "Thanks, possible that I just got lucky with the checkpoint ^^\nThe WER is not as fancy, I post it below just for reference (I'm probing LB most of the time) \nWhat I find interesting is the two dips in the loss, which I'm guessing is the time the cycle restarted (going through from easy samples to hard samples) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1440116%2Fa9d96670559352f7314b2fe00008f88e%2Fwercurve.jpg?generation=1698029840132999&alt=media)"
  },
  "source": "meta"
}