{
  "id": 447986,
  "title": "11th place solution",
  "url": "/competitions/bengaliai-speech/discussion/447986",
  "author_name": "Roy Wei",
  "post_date": "2023-10-18T03:24:58.780000",
  "votes": 13,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi, all the fellow Kagglers! </p>\n<p>I'd like to first give a massive gratitude to Kaggle, and Bengali AI on behalf of my team. Besides, I want to thank lots of fellow competitors for providing inspirational insights. Thanks to authors of reposLastly, thanks to my teammates, particularly <a href=\"https://www.kaggle.com/rkxuan\" target=\"_blank\">@rkxuan</a> for spotting the repo that forms the basis of the punctuation model. </p>\n<hr>\n<h2>Model Architecture:</h2>\n<ul>\n<li><strong>ASR model:</strong> Wav2vec2 CTC model</li>\n<li><strong>Ngram:</strong> Kenlm</li>\n<li><strong>Punctuation Model:</strong> xlm-roberta-large</li>\n</ul>\n<h2>What Worked:</h2>\n<ul>\n<li>Fine-tuning on <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> filtered Common Voice dataset. Public LB: <code>0.439 -&gt; 0.428</code></li>\n<li>Fine-tuning on Openslr 37 resulted in: Public LB <code>0.428 -&gt; 0.414</code></li>\n<li>Incorporating Oscar corpus into ngram</li>\n<li>Data augmentation with <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\">audiomentations</a>.</li>\n<li>xlm-roberta-large configuration from this <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">repo</a>.</li>\n<li>Optuna search for decoding hyperparameters, as demonstrated in this <a href=\"https://www.kaggle.com/code/royalacecat/lb-0-442-the-best-decoding-parameters\" target=\"_blank\">notebook</a>.</li>\n</ul>\n<h2>Challenges:</h2>\n<h3>Datasets:</h3>\n<ul>\n<li><strong>Openslr 53:</strong> Voluminous and might've led to overfitting during training. It being crowdsourced could be a factor, especially when compared to the more refined Openslr 37.</li>\n<li><strong>Fleurs:</strong> Plagued with inconsistent quality. A consistent sharp noise mars the dataset.</li>\n<li><strong>Competition ds:</strong> Exhibits quality diversity.</li>\n</ul>\n<h3>Punctuation Restoration:</h3>\n<ul>\n<li>Difficulties in restoring five punctuations: <strong>!</strong>, <strong>,</strong>, <strong>?</strong>, <strong>।</strong>, and <strong>-</strong>.<ul>\n<li>Notably, hyphens appear to be tricky. Maybe isolating it for training could help.</li></ul></li>\n</ul>\n<h3>Audio Augmentation:</h3>\n<ul>\n<li>Might've overdone with BGMs. Modulating pitch could potentially be more effective.</li>\n</ul>\n<h3>Speech Enhancement:</h3>\n<ul>\n<li>Both FAIR Denoiser and Nvidia CleanUnet fell short of expectations. It's perplexing, but perhaps they inadvertently degraded human voice quality.</li>\n</ul>\n<h2>Pro-Tips:</h2>\n<ul>\n<li>Prefer ARPA over BIN. Kenlm seems to lose unigram post-conversion, but this trick can amplify the score by <code>0.007</code>. Mind the 13GB RAM constraint on Kaggle. I eventually settled with a trimmed 4gram, approximately 12GB.</li>\n</ul>\n<h2>Observations:</h2>\n<ul>\n<li>Local CV, grounded in annotated OOD examples, aligns well with the public LB. This might explain the negligible shake-up in the end.</li>\n<li>Oscar primarily features formal <strong>articles</strong>. So, the ngram, despite its magnitude, might overlook the niche vocabularies in OOD test sets, especially with the high OOV observed.</li>\n</ul>\n<hr>\n<p>Feel free to share your thoughts and experiences!</p>",
  "messages": [
    {
      "id": 2486574,
      "postDate": "2023-10-18T03:24:58.780Z",
      "content": "<p>Hi, all the fellow Kagglers! </p>\n<p>I'd like to first give a massive gratitude to Kaggle, and Bengali AI on behalf of my team. Besides, I want to thank lots of fellow competitors for providing inspirational insights. Thanks to authors of reposLastly, thanks to my teammates, particularly <a href=\"https://www.kaggle.com/rkxuan\" target=\"_blank\">@rkxuan</a> for spotting the repo that forms the basis of the punctuation model. </p>\n<hr>\n<h2>Model Architecture:</h2>\n<ul>\n<li><strong>ASR model:</strong> Wav2vec2 CTC model</li>\n<li><strong>Ngram:</strong> Kenlm</li>\n<li><strong>Punctuation Model:</strong> xlm-roberta-large</li>\n</ul>\n<h2>What Worked:</h2>\n<ul>\n<li>Fine-tuning on <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> filtered Common Voice dataset. Public LB: <code>0.439 -&gt; 0.428</code></li>\n<li>Fine-tuning on Openslr 37 resulted in: Public LB <code>0.428 -&gt; 0.414</code></li>\n<li>Incorporating Oscar corpus into ngram</li>\n<li>Data augmentation with <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\">audiomentations</a>.</li>\n<li>xlm-roberta-large configuration from this <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">repo</a>.</li>\n<li>Optuna search for decoding hyperparameters, as demonstrated in this <a href=\"https://www.kaggle.com/code/royalacecat/lb-0-442-the-best-decoding-parameters\" target=\"_blank\">notebook</a>.</li>\n</ul>\n<h2>Challenges:</h2>\n<h3>Datasets:</h3>\n<ul>\n<li><strong>Openslr 53:</strong> Voluminous and might've led to overfitting during training. It being crowdsourced could be a factor, especially when compared to the more refined Openslr 37.</li>\n<li><strong>Fleurs:</strong> Plagued with inconsistent quality. A consistent sharp noise mars the dataset.</li>\n<li><strong>Competition ds:</strong> Exhibits quality diversity.</li>\n</ul>\n<h3>Punctuation Restoration:</h3>\n<ul>\n<li>Difficulties in restoring five punctuations: <strong>!</strong>, <strong>,</strong>, <strong>?</strong>, <strong>।</strong>, and <strong>-</strong>.<ul>\n<li>Notably, hyphens appear to be tricky. Maybe isolating it for training could help.</li></ul></li>\n</ul>\n<h3>Audio Augmentation:</h3>\n<ul>\n<li>Might've overdone with BGMs. Modulating pitch could potentially be more effective.</li>\n</ul>\n<h3>Speech Enhancement:</h3>\n<ul>\n<li>Both FAIR Denoiser and Nvidia CleanUnet fell short of expectations. It's perplexing, but perhaps they inadvertently degraded human voice quality.</li>\n</ul>\n<h2>Pro-Tips:</h2>\n<ul>\n<li>Prefer ARPA over BIN. Kenlm seems to lose unigram post-conversion, but this trick can amplify the score by <code>0.007</code>. Mind the 13GB RAM constraint on Kaggle. I eventually settled with a trimmed 4gram, approximately 12GB.</li>\n</ul>\n<h2>Observations:</h2>\n<ul>\n<li>Local CV, grounded in annotated OOD examples, aligns well with the public LB. This might explain the negligible shake-up in the end.</li>\n<li>Oscar primarily features formal <strong>articles</strong>. So, the ngram, despite its magnitude, might overlook the niche vocabularies in OOD test sets, especially with the high OOV observed.</li>\n</ul>\n<hr>\n<p>Feel free to share your thoughts and experiences!</p>",
      "rawMarkdown": "Hi, all the fellow Kagglers! \n\nI'd like to first give a massive gratitude to Kaggle, and Bengali AI on behalf of my team. Besides, I want to thank lots of fellow competitors for providing inspirational insights. Thanks to authors of reposLastly, thanks to my teammates, particularly @rkxuan for spotting the repo that forms the basis of the punctuation model. \n\n---\n\n## Model Architecture:\n\n- **ASR model:** Wav2vec2 CTC model\n- **Ngram:** Kenlm\n- **Punctuation Model:** xlm-roberta-large\n\n## What Worked:\n\n- Fine-tuning on @umongsain filtered Common Voice dataset. Public LB: `0.439 -> 0.428`\n- Fine-tuning on Openslr 37 resulted in: Public LB `0.428 -> 0.414`\n- Incorporating Oscar corpus into ngram\n- Data augmentation with [audiomentations](https://github.com/asteroid-team/torch-audiomentations).\n- xlm-roberta-large configuration from this [repo](https://github.com/xashru/punctuation-restoration).\n- Optuna search for decoding hyperparameters, as demonstrated in this [notebook](https://www.kaggle.com/code/royalacecat/lb-0-442-the-best-decoding-parameters).\n\n## Challenges:\n\n### Datasets:\n\n- **Openslr 53:** Voluminous and might've led to overfitting during training. It being crowdsourced could be a factor, especially when compared to the more refined Openslr 37.\n- **Fleurs:** Plagued with inconsistent quality. A consistent sharp noise mars the dataset.\n- **Competition ds:** Exhibits quality diversity.\n\n### Punctuation Restoration:\n\n- Difficulties in restoring five punctuations: **!**, **,**, **?**, **।**, and **-**.\n  - Notably, hyphens appear to be tricky. Maybe isolating it for training could help.\n\n### Audio Augmentation:\n\n- Might've overdone with BGMs. Modulating pitch could potentially be more effective.\n\n### Speech Enhancement:\n\n- Both FAIR Denoiser and Nvidia CleanUnet fell short of expectations. It's perplexing, but perhaps they inadvertently degraded human voice quality.\n\n## Pro-Tips:\n\n- Prefer ARPA over BIN. Kenlm seems to lose unigram post-conversion, but this trick can amplify the score by `0.007`. Mind the 13GB RAM constraint on Kaggle. I eventually settled with a trimmed 4gram, approximately 12GB.\n\n## Observations:\n\n- Local CV, grounded in annotated OOD examples, aligns well with the public LB. This might explain the negligible shake-up in the end.\n- Oscar primarily features formal **articles**. So, the ngram, despite its magnitude, might overlook the niche vocabularies in OOD test sets, especially with the high OOV observed.\n\n---\n\nFeel free to share your thoughts and experiences!",
      "votes": 12
    },
    {
      "id": 2486580,
      "postDate": "2023-10-18T03:30:34.623Z",
      "content": "<p>I have also realized that my notebook only runs for approximately 6 hours, which means I may get some easy money by simply increasing <strong>beam width</strong> as long as it fits into the time limit! That may be missing element that makes me lost the gold medal 🙃</p>\n<p>Stay tune to my reflection on the competition! I will post it soon!</p>",
      "rawMarkdown": "I have also realized that my notebook only runs for approximately 6 hours, which means I may get some easy money by simply increasing **beam width** as long as it fits into the time limit! That may be missing element that makes me lost the gold medal 🙃\n\nStay tune to my reflection on the competition! I will post it soon!",
      "votes": 3,
      "replies": [
        {
          "id": 2486595,
          "postDate": "2023-10-18T04:03:36.433Z",
          "content": "<p>Yeah! I didnt invest so much in this competition, so I tried to find easy ways to improve WER. In my observations increasing <code>beam_width</code> to 2048 leads to 8h for Wav2vec2 model and +0.003 boost compared to default</p>",
          "rawMarkdown": "Yeah! I didnt invest so much in this competition, so I tried to find easy ways to improve WER. In my observations increasing `beam_width` to 2048 leads to 8h for Wav2vec2 model and +0.003 boost compared to default",
          "votes": 2,
          "replies": [
            {
              "id": 2486624,
              "postDate": "2023-10-18T04:31:50.870Z",
              "content": "<p>if I had remembered to do that, I may be placed 11th and get the gold medal🥲. Lots of sadness as I realize that. This is truly a simple method to boost up the rank at the cost of time… The beam_width in my final submission is 1024…</p>",
              "rawMarkdown": "if I had remembered to do that, I may be placed 11th and get the gold medal🥲. Lots of sadness as I realize that. This is truly a simple method to boost up the rank at the cost of time... The beam_width in my final submission is 1024...",
              "votes": 2
            }
          ]
        },
        {
          "id": 2490185,
          "postDate": "2023-10-20T14:03:01.183Z",
          "content": "<p>You got your desired gold! :3 Congratulations! </p>",
          "rawMarkdown": "You got your desired gold! :3 Congratulations! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2486975,
      "postDate": "2023-10-18T09:11:47.617Z",
      "content": "<p>Nice work ! I got excited when you mentioned you got improvement with public repo, but I never thought it was punctuations 😭 My guess was Language model, and I went all out with searching for LM (my bin file is 40gig is size 🤣 )<br>\nAnyway, congratulations for the teamwork!</p>",
      "rawMarkdown": "Nice work ! I got excited when you mentioned you got improvement with public repo, but I never thought it was punctuations 😭 My guess was Language model, and I went all out with searching for LM (my bin file is 40gig is size 🤣 )\nAnyway, congratulations for the teamwork!",
      "votes": 1,
      "replies": [
        {
          "id": 2487078,
          "postDate": "2023-10-18T10:54:43.983Z",
          "content": "<p>Hahah, can't be too explicit about the tricks I use🤣. I think I ngram is more about data rather than the architecture, 'cause ngram is a statistical model. But, 40 giga! wow you are really stretching it to the limit!</p>\n<p>Also, rly thank you a lot for almost posting a comment in all of my posts🫡. We may be able to see each other soon!</p>",
          "rawMarkdown": "Hahah, can't be too explicit about the tricks I use🤣. I think I ngram is more about data rather than the architecture, 'cause ngram is a statistical model. But, 40 giga! wow you are really stretching it to the limit!\n\nAlso, rly thank you a lot for almost posting a comment in all of my posts🫡. We may be able to see each other soon!",
          "votes": 1,
          "replies": [
            {
              "id": 2489397,
              "postDate": "2023-10-20T01:30:00.623Z",
              "content": "<p>And you got gold on your first competition, congratulations 🥳<br>\nLooking forward to see you in other competition :) </p>",
              "rawMarkdown": "And you got gold on your first competition, congratulations 🥳\nLooking forward to see you in other competition :) "
            }
          ]
        }
      ]
    },
    {
      "id": 2490128,
      "postDate": "2023-10-20T12:59:14.273Z",
      "content": "<blockquote>\n  <p>Local CV, grounded in annotated OOD examples, aligns well with the public LB. This might explain the negligible shake-up in the end.</p>\n</blockquote>\n<p>Could you clarify, how did you make your CV?</p>",
      "rawMarkdown": ">Local CV, grounded in annotated OOD examples, aligns well with the public LB. This might explain the negligible shake-up in the end.\n\nCould you clarify, how did you make your CV?",
      "replies": [
        {
          "id": 2490136,
          "postDate": "2023-10-20T13:16:09.860Z",
          "content": "<p>Of course. Also, sorry for misusing the term 'CV', it just a OOD validation set I hold out 🤔</p>\n<p>It is based on this <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\" target=\"_blank\">post</a>. The author annotates the OOD examples, and after correcting some mistakes, I calculate my model OOD wer rate against these ground truths. The improvements in public LB is positively correlated with OOD wer rate. </p>\n<p>I interpret that as a sign that public LB will be correlated with private LB. Thus, there isn't much shake up/down if the margin is relatively large.</p>\n<p>It's just a hypothesis :)</p>",
          "rawMarkdown": "Of course. Also, sorry for misusing the term 'CV', it just a OOD validation set I hold out 🤔\n\nIt is based on this [post](https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932). The author annotates the OOD examples, and after correcting some mistakes, I calculate my model OOD wer rate against these ground truths. The improvements in public LB is positively correlated with OOD wer rate. \n\nI interpret that as a sign that public LB will be correlated with private LB. Thus, there isn't much shake up/down if the margin is relatively large.\n\nIt's just a hypothesis :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2487853,
      "postDate": "2023-10-18T19:52:20.383Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2486580,
      "author_name": "Roy Wei",
      "author_url": "",
      "post_date": "2023-10-18T03:30:34.623000",
      "content": "<p>I have also realized that my notebook only runs for approximately 6 hours, which means I may get some easy money by simply increasing <strong>beam width</strong> as long as it fits into the time limit! That may be missing element that makes me lost the gold medal 🙃</p>\n<p>Stay tune to my reflection on the competition! I will post it soon!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2486595,
          "author_name": "Yurnero",
          "author_url": "",
          "post_date": "2023-10-18T04:03:36.433000",
          "content": "<p>Yeah! I didnt invest so much in this competition, so I tried to find easy ways to improve WER. In my observations increasing <code>beam_width</code> to 2048 leads to 8h for Wav2vec2 model and +0.003 boost compared to default</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2486624,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2023-10-18T04:31:50.870000",
              "content": "<p>if I had remembered to do that, I may be placed 11th and get the gold medal🥲. Lots of sadness as I realize that. This is truly a simple method to boost up the rank at the cost of time… The beam_width in my final submission is 1024…</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2490185,
          "author_name": "Md Boktiar Mahbub Murad",
          "author_url": "",
          "post_date": "2023-10-20T14:03:01.183000",
          "content": "<p>You got your desired gold! :3 Congratulations! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2486975,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2023-10-18T09:11:47.617000",
      "content": "<p>Nice work ! I got excited when you mentioned you got improvement with public repo, but I never thought it was punctuations 😭 My guess was Language model, and I went all out with searching for LM (my bin file is 40gig is size 🤣 )<br>\nAnyway, congratulations for the teamwork!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2487078,
          "author_name": "Roy Wei",
          "author_url": "",
          "post_date": "2023-10-18T10:54:43.983000",
          "content": "<p>Hahah, can't be too explicit about the tricks I use🤣. I think I ngram is more about data rather than the architecture, 'cause ngram is a statistical model. But, 40 giga! wow you are really stretching it to the limit!</p>\n<p>Also, rly thank you a lot for almost posting a comment in all of my posts🫡. We may be able to see each other soon!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2489397,
              "author_name": "yukiya",
              "author_url": "",
              "post_date": "2023-10-20T01:30:00.623000",
              "content": "<p>And you got gold on your first competition, congratulations 🥳<br>\nLooking forward to see you in other competition :) </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2490128,
      "author_name": "Man of the year",
      "author_url": "",
      "post_date": "2023-10-20T12:59:14.273000",
      "content": "<blockquote>\n  <p>Local CV, grounded in annotated OOD examples, aligns well with the public LB. This might explain the negligible shake-up in the end.</p>\n</blockquote>\n<p>Could you clarify, how did you make your CV?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2490136,
          "author_name": "Roy Wei",
          "author_url": "",
          "post_date": "2023-10-20T13:16:09.860000",
          "content": "<p>Of course. Also, sorry for misusing the term 'CV', it just a OOD validation set I hold out 🤔</p>\n<p>It is based on this <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\" target=\"_blank\">post</a>. The author annotates the OOD examples, and after correcting some mistakes, I calculate my model OOD wer rate against these ground truths. The improvements in public LB is positively correlated with OOD wer rate. </p>\n<p>I interpret that as a sign that public LB will be correlated with private LB. Thus, there isn't much shake up/down if the margin is relatively large.</p>\n<p>It's just a hypothesis :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2487853,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:52:20.383000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486574": "Hi, all the fellow Kagglers! \n\nI'd like to first give a massive gratitude to Kaggle, and Bengali AI on behalf of my team. Besides, I want to thank lots of fellow competitors for providing inspirational insights. Thanks to authors of reposLastly, thanks to my teammates, particularly @rkxuan for spotting the repo that forms the basis of the punctuation model. \n\n---\n\n## Model Architecture:\n\n- **ASR model:** Wav2vec2 CTC model\n- **Ngram:** Kenlm\n- **Punctuation Model:** xlm-roberta-large\n\n## What Worked:\n\n- Fine-tuning on @umongsain filtered Common Voice dataset. Public LB: `0.439 -> 0.428`\n- Fine-tuning on Openslr 37 resulted in: Public LB `0.428 -> 0.414`\n- Incorporating Oscar corpus into ngram\n- Data augmentation with [audiomentations](https://github.com/asteroid-team/torch-audiomentations).\n- xlm-roberta-large configuration from this [repo](https://github.com/xashru/punctuation-restoration).\n- Optuna search for decoding hyperparameters, as demonstrated in this [notebook](https://www.kaggle.com/code/royalacecat/lb-0-442-the-best-decoding-parameters).\n\n## Challenges:\n\n### Datasets:\n\n- **Openslr 53:** Voluminous and might've led to overfitting during training. It being crowdsourced could be a factor, especially when compared to the more refined Openslr 37.\n- **Fleurs:** Plagued with inconsistent quality. A consistent sharp noise mars the dataset.\n- **Competition ds:** Exhibits quality diversity.\n\n### Punctuation Restoration:\n\n- Difficulties in restoring five punctuations: **!**, **,**, **?**, **।**, and **-**.\n  - Notably, hyphens appear to be tricky. Maybe isolating it for training could help.\n\n### Audio Augmentation:\n\n- Might've overdone with BGMs. Modulating pitch could potentially be more effective.\n\n### Speech Enhancement:\n\n- Both FAIR Denoiser and Nvidia CleanUnet fell short of expectations. It's perplexing, but perhaps they inadvertently degraded human voice quality.\n\n## Pro-Tips:\n\n- Prefer ARPA over BIN. Kenlm seems to lose unigram post-conversion, but this trick can amplify the score by `0.007`. Mind the 13GB RAM constraint on Kaggle. I eventually settled with a trimmed 4gram, approximately 12GB.\n\n## Observations:\n\n- Local CV, grounded in annotated OOD examples, aligns well with the public LB. This might explain the negligible shake-up in the end.\n- Oscar primarily features formal **articles**. So, the ngram, despite its magnitude, might overlook the niche vocabularies in OOD test sets, especially with the high OOV observed.\n\n---\n\nFeel free to share your thoughts and experiences!",
    "2486580": "I have also realized that my notebook only runs for approximately 6 hours, which means I may get some easy money by simply increasing **beam width** as long as it fits into the time limit! That may be missing element that makes me lost the gold medal 🙃\n\nStay tune to my reflection on the competition! I will post it soon!",
    "2486975": "Nice work ! I got excited when you mentioned you got improvement with public repo, but I never thought it was punctuations 😭 My guess was Language model, and I went all out with searching for LM (my bin file is 40gig is size 🤣 )\nAnyway, congratulations for the teamwork!",
    "2490128": ">Local CV, grounded in annotated OOD examples, aligns well with the public LB. This might explain the negligible shake-up in the end.\n\nCould you clarify, how did you make your CV?",
    "2487853": ""
  }
}