{
  "id": 450531,
  "title": "40th Place Solution without External Dataset!",
  "url": "/competitions/bengaliai-speech/writeups/bengalx-40th-place-solution-without-external-datas",
  "author_name": "",
  "post_date": "2023-10-24T17:07:20.073Z",
  "votes": 12,
  "comment_count": 3,
  "views": 0,
  "content": "<p>We achieved 40th Place (Silver Medal). Congratulations to all of Team Members from <strong>BengalX</strong>: <a href=\"https://www.kaggle.com/iftekharamin\" target=\"_blank\">@iftekharamin</a> , <a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a> , <a href=\"https://www.kaggle.com/fahimshahriarkhan\" target=\"_blank\">@fahimshahriarkhan</a> 🎉</p>\n<p><em>I would like to acknowledge my Team Lead <a href=\"https://www.kaggle.com/iftekharamin\" target=\"_blank\">@iftekharamin</a> Bhaiya for giving opportunity to do this competition and giving proper guideline for achieving silver medal.</em></p>\n<p><strong>Dataset:</strong> We subset the dataset Based on <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> published train metadata. We used condition to filter clean dataset from train meta features (yellowking_preds &amp; google_preds) wer = 90% similar. After this we further filter dataset which mos_pred &gt; 2. And we found around 100k+ datapoints.</p>\n<p><strong>Data Cleaning:</strong> We filter audios which duration is less then 1 sec. as outlier and to not mislead model performance.</p>\n<p><strong>Augmentations:</strong> We used audio augmentations i.e, Noise , Background sound mixing, Speed up-down, SpecAug, Changing different Sampling Rates. </p>\n<p><strong>STT Modeling:</strong> We used Indic wav2vec2 pretrained model. And we finetune with the Subset augmented dataset. </p>\n<p><strong>Post-processing- LM Decode:</strong> We used arijit indic pretrained KenLM.</p>\n<p><strong>Post-processing-Punctuation:</strong> We used xashru/punctuation-restoration repo with xlm-roberta-base model and fine tune this competition dataset as punctuation restoration. We only consider 4 punctuation classes : {'O': 0, 'COMMA': 1, 'PERIOD': 2, 'QUESTION': 3} </p>\n<p><strong>Post-processing-Erro Correction:</strong> We used this repo solution as further error correction <a href=\"https://github.com/Tawkat/Bengali-Spell-Checker-and-Auto-Correction-Suggestion-for-MS-Word\" target=\"_blank\">https://github.com/Tawkat/Bengali-Spell-Checker-and-Auto-Correction-Suggestion-for-MS-Word</a></p>",
  "messages": [
    {
      "id": "2497488",
      "postDate": "10/24/2023 16:43:34",
      "content": "<p>We achieved 40th Place (Silver Medal). Congratulations to all of Team Members from <strong>BengalX</strong>: <a href=\"https://www.kaggle.com/iftekharamin\" target=\"_blank\">@iftekharamin</a> , <a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a> , <a href=\"https://www.kaggle.com/fahimshahriarkhan\" target=\"_blank\">@fahimshahriarkhan</a> 🎉</p>\n<p><em>I would like to acknowledge my Team Lead <a href=\"https://www.kaggle.com/iftekharamin\" target=\"_blank\">@iftekharamin</a> Bhaiya for giving opportunity to do this competition and giving proper guideline for achieving silver medal.</em></p>\n<p><strong>Dataset:</strong> We subset the dataset Based on <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> published train metadata. We used condition to filter clean dataset from train meta features (yellowking_preds &amp; google_preds) wer = 90% similar. After this we further filter dataset which mos_pred &gt; 2. And we found around 100k+ datapoints.</p>\n<p><strong>Data Cleaning:</strong> We filter audios which duration is less then 1 sec. as outlier and to not mislead model performance.</p>\n<p><strong>Augmentations:</strong> We used audio augmentations i.e, Noise , Background sound mixing, Speed up-down, SpecAug, Changing different Sampling Rates. </p>\n<p><strong>STT Modeling:</strong> We used Indic wav2vec2 pretrained model. And we finetune with the Subset augmented dataset. </p>\n<p><strong>Post-processing- LM Decode:</strong> We used arijit indic pretrained KenLM.</p>\n<p><strong>Post-processing-Punctuation:</strong> We used xashru/punctuation-restoration repo with xlm-roberta-base model and fine tune this competition dataset as punctuation restoration. We only consider 4 punctuation classes : {'O': 0, 'COMMA': 1, 'PERIOD': 2, 'QUESTION': 3} </p>\n<p><strong>Post-processing-Erro Correction:</strong> We used this repo solution as further error correction <a href=\"https://github.com/Tawkat/Bengali-Spell-Checker-and-Auto-Correction-Suggestion-for-MS-Word\" target=\"_blank\">https://github.com/Tawkat/Bengali-Spell-Checker-and-Auto-Correction-Suggestion-for-MS-Word</a></p>",
      "rawMarkdown": "We achieved 40th Place (Silver Medal). Congratulations to all of Team Members from **BengalX**: @iftekharamin , @mdfahimreshm , @fahimshahriarkhan 🎉\n\n*I would like to acknowledge my Team Lead @iftekharamin Bhaiya for giving opportunity to do this competition and giving proper guideline for achieving silver medal.*\n\n**Dataset:** We subset the dataset Based on @imtiazprio published train metadata. We used condition to filter clean dataset from train meta features (yellowking_preds & google_preds) wer = 90% similar. After this we further filter dataset which mos_pred > 2. And we found around 100k+ datapoints.\n\n**Data Cleaning:** We filter audios which duration is less then 1 sec. as outlier and to not mislead model performance.\n\n**Augmentations:** We used audio augmentations i.e, Noise , Background sound mixing, Speed up-down, SpecAug, Changing different Sampling Rates. \n\n**STT Modeling:** We used Indic wav2vec2 pretrained model. And we finetune with the Subset augmented dataset. \n\n**Post-processing- LM Decode:** We used arijit indic pretrained KenLM.\n\n**Post-processing-Punctuation:** We used xashru/punctuation-restoration repo with xlm-roberta-base model and fine tune this competition dataset as punctuation restoration. We only consider 4 punctuation classes : {'O': 0, 'COMMA': 1, 'PERIOD': 2, 'QUESTION': 3} \n\n**Post-processing-Erro Correction:** We used this repo solution as further error correction https://github.com/Tawkat/Bengali-Spell-Checker-and-Auto-Correction-Suggestion-for-MS-Word",
      "votes": null
    },
    {
      "id": "2497717",
      "postDate": "10/24/2023 20:09:16",
      "content": "<p>Great job! Congrats on the Silver Medal. </p>",
      "rawMarkdown": "Great job! Congrats on the Silver Medal.",
      "votes": null
    },
    {
      "id": "2497842",
      "postDate": "10/25/2023 00:26:24",
      "content": "<p>Thank you for sharing your solution.<br>\nWe couldn't make a spell checker work. How did it affectEd your LBand PB score?</p>",
      "rawMarkdown": "Thank you for sharing your solution.\nWe couldn't make a spell checker work. How did it affectEd your LBand PB score?",
      "votes": null
    },
    {
      "id": "2497895",
      "postDate": "10/25/2023 01:50:28",
      "content": "<p>It’s improved 0.001 in LB &amp; PB. Thank you for your comment and congratulations achieving 3rd place 🙌</p>",
      "rawMarkdown": "It’s improved 0.001 in LB & PB. Thank you for your comment and congratulations achieving 3rd place 🙌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2497717,
      "author_name": "scientiapotentia",
      "author_url": "",
      "post_date": "10/24/2023 20:09:16",
      "content": "<p>Great job! Congrats on the Silver Medal. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2497842,
      "author_name": "ludditep",
      "author_url": "",
      "post_date": "10/25/2023 00:26:24",
      "content": "<p>Thank you for sharing your solution.<br>\nWe couldn't make a spell checker work. How did it affectEd your LBand PB score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2497895,
          "author_name": "aifahim",
          "author_url": "",
          "post_date": "10/25/2023 01:50:28",
          "content": "<p>It’s improved 0.001 in LB &amp; PB. Thank you for your comment and congratulations achieving 3rd place 🙌</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2497488": "We achieved 40th Place (Silver Medal). Congratulations to all of Team Members from **BengalX**: @iftekharamin , @mdfahimreshm , @fahimshahriarkhan 🎉\n\n*I would like to acknowledge my Team Lead @iftekharamin Bhaiya for giving opportunity to do this competition and giving proper guideline for achieving silver medal.*\n\n**Dataset:** We subset the dataset Based on @imtiazprio published train metadata. We used condition to filter clean dataset from train meta features (yellowking_preds & google_preds) wer = 90% similar. After this we further filter dataset which mos_pred > 2. And we found around 100k+ datapoints.\n\n**Data Cleaning:** We filter audios which duration is less then 1 sec. as outlier and to not mislead model performance.\n\n**Augmentations:** We used audio augmentations i.e, Noise , Background sound mixing, Speed up-down, SpecAug, Changing different Sampling Rates. \n\n**STT Modeling:** We used Indic wav2vec2 pretrained model. And we finetune with the Subset augmented dataset. \n\n**Post-processing- LM Decode:** We used arijit indic pretrained KenLM.\n\n**Post-processing-Punctuation:** We used xashru/punctuation-restoration repo with xlm-roberta-base model and fine tune this competition dataset as punctuation restoration. We only consider 4 punctuation classes : {'O': 0, 'COMMA': 1, 'PERIOD': 2, 'QUESTION': 3} \n\n**Post-processing-Erro Correction:** We used this repo solution as further error correction https://github.com/Tawkat/Bengali-Spell-Checker-and-Auto-Correction-Suggestion-for-MS-Word",
    "2497717": "Great job! Congrats on the Silver Medal.",
    "2497842": "Thank you for sharing your solution.\nWe couldn't make a spell checker work. How did it affectEd your LBand PB score?",
    "2497895": "It’s improved 0.001 in LB & PB. Thank you for your comment and congratulations achieving 3rd place 🙌"
  },
  "source": "meta"
}