{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":73047,"databundleVersionId":8149390,"sourceType":"competition"},{"sourceId":8150853,"sourceType":"datasetVersion","datasetId":4820621},{"sourceId":8159145,"sourceType":"datasetVersion","datasetId":4826969},{"sourceId":8174826,"sourceType":"datasetVersion","datasetId":4838779}],"dockerImageVersionId":30699,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Team Black</center></div>","metadata":{}},{"cell_type":"markdown","source":"### **Team Members**\n* Mohammad Sadat Hossain\n* Asif Azad\n* Ashrafur Rahman Khan","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Problem & Dataset </center></div>","metadata":{}},{"cell_type":"markdown","source":"## **Problem Overview**\nThe prime goal of this competition is to come up with a solution system to transcribe Bengali speech with various regional dialects following the orthography set by linguists. The speech corpus for this competition has been developed by Bengali.AI from one of their ongoing large-scale research projects. The corpus consists of spontaneous speech on everyday topics from 373 individuals with various regional dialects from ten different geographical locations such as Rangpur, Kishoreganj, Narail, Chittagong, Narsingdi, Tangail, Barishal, Habiganj, Sylhet & Sandwip. The cumulative length of this entire speech corpus is over 79+ hours. Your efforts could improve the Bengali speech recognition technology using this unique speech recognition dataset for the very first time in history which is dealing with the regional speech domain. In addition, your submission will be among the first open-source speech recognition methods for Bengali.","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-24T16:44:21.897301Z","iopub.execute_input":"2024-04-24T16:44:21.897691Z","iopub.status.idle":"2024-04-24T16:44:22.919639Z","shell.execute_reply.started":"2024-04-24T16:44:21.897659Z","shell.execute_reply":"2024-04-24T16:44:22.918742Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_audio_dir = '/kaggle/input/ben10/ben10/16_kHz_train_audio/'\ntest_audio_dir = '/kaggle/input/ben10/ben10/16_kHz_valid_audio/'\ntrain_df = pd.read_csv('/kaggle/input/ben10/ben10/train.csv')\nprint(f'\\nThere are {train_df.shape[0]} train samples.\\n')\nprint(f'\\nThe columns are:\\n')\n\ncolumn_names = train_df.columns\nfor column_name in column_names:\n    print(f'{column_name}')\ntrain_df.sample(5)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-24T16:44:22.921902Z","iopub.execute_input":"2024-04-24T16:44:22.922402Z","iopub.status.idle":"2024-04-24T16:44:23.185436Z","shell.execute_reply.started":"2024-04-24T16:44:22.922372Z","shell.execute_reply":"2024-04-24T16:44:23.184485Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"split2path = {\n    \"train\": train_audio_dir,\n    \"test\": test_audio_dir,\n}\n\ndef extract_split(filename):\n    filename_ = filename.split(\"_\")\n    split = filename_[0]\n    return split\n\ndef extract_district(filename):\n    filename_ = filename.split(\" \")[0]\n    district = filename_.split(\"_\")[1]\n    return district\n\ndef beautify_dataset(data):\n    splits = []\n    districts = []\n    newpaths = []\n    transcripts = []\n    \n    for i in range(len(data)):\n        filename, transcript = data.iloc[i]\n        split = extract_split(filename)\n        district = extract_district(filename)\n        dir_path = split2path[split]\n        composed_path = f\"{dir_path}{filename}\"\n        \n        if os.path.exists(composed_path) == False:\n            print(f\"{composed_path} does not exist.\")\n            continue\n        \n        # replace any newline characters\n        transcript = transcript.replace(\"\\n\", \" \")        \n        punctuations = ['।', '?', ',', ';', '!',':', '<>']\n        for p in punctuations:\n            transcript = transcript.replace(p, \" \")\n        transcript = \" \".join(transcript.split())\n        # transcript = normalize(transcript)\n        \n        splits.append(split)\n        districts.append(district)\n        newpaths.append(composed_path)\n        transcripts.append(transcript)\n    \n    data['file_path'] = newpaths\n    data['district'] = districts\n    data['split'] = splits\n    data['transcripts'] = transcripts\n    \n#     data.drop(columns=['file_name'], inplace=True)\n    \n    return data\n\ntrain_df = beautify_dataset(train_df)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-24T16:44:23.186870Z","iopub.execute_input":"2024-04-24T16:44:23.187216Z","iopub.status.idle":"2024-04-24T16:44:39.780403Z","shell.execute_reply.started":"2024-04-24T16:44:23.187189Z","shell.execute_reply":"2024-04-24T16:44:39.779419Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_districts = sorted(train_df['district'].unique())\nprint(f'\\nThere are dialects from {len(unique_districts)} districts.\\n')\nfor district in unique_districts:\n    print(district)\nprint('\\n\\n')\n\ndistrict_counts = train_df['district'].value_counts().reset_index()\ndistrict_counts.columns = ['district', 'frequency']\nsns.set(style=\"whitegrid\")\n\n# Create the bar plot\nplt.figure(figsize=(12, 8))  # Adjust the figure size\nbarplot = sns.barplot(x='district', y='frequency', data=district_counts, palette='viridis')\n\n# Improve the aesthetics and readability\nplt.xticks(rotation=45, ha='right', fontsize=12)  # Rotate x-axis labels for better readability\nplt.yticks(fontsize=12)\nplt.xlabel('District', fontsize=14)  # Label the x-axis\nplt.ylabel('Number of Samples', fontsize=14)  # Label the y-axis\nplt.title('Samples from Each District', fontsize=16)  # Add a title to the plot\n\n# Optional: Add the frequency values on top of the bars\nfor p in barplot.patches:\n    barplot.annotate(format(int(p.get_height())), \n                     (p.get_x() + p.get_width() / 2., p.get_height()), \n                     ha = 'center', va = 'center', \n                     xytext = (0, 9), \n                     textcoords = 'offset points')\n\nplt.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-24T16:44:39.782624Z","iopub.execute_input":"2024-04-24T16:44:39.782950Z","iopub.status.idle":"2024-04-24T16:44:40.376239Z","shell.execute_reply.started":"2024-04-24T16:44:39.782923Z","shell.execute_reply":"2024-04-24T16:44:40.375121Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Our Final Approach</center></div>","metadata":{}},{"cell_type":"markdown","source":"### **🥲 After the competition, we realized just replacing csebuetnlp normalizer with bnunicode normalizer in our exact same approach improve the score below 0.7 🥲**\n\n**Complete Pipeline:**\n\n> Audio Classification → District-Specific ASR Processing → Punctuation Restoration\n\n**1. Audio Classification:**\n\n* *Purpose:* Classify speech samples by their respective districts.\n\n* *Model:* Used a fine-tuned [whisper-medium](https://huggingface.co/openai/whisper-medium) model adapted for recognizing ten distinct Bangla dialects.\n\n**2. District-Specific ASR Models:**\n\n* *Purpose:* Convert dialect-specific speech into text.\n\n* *Details:* Fine-tuned separate ASR models for each district, fine-tuned on regional speech data without punctuation. All of them were multi-lingual [whisper-medium](https://huggingface.co/openai/whisper-medium) model fine-tuned on Bengali Dataset.\n\n**3. Punctuation Restoration:**\n\n* *Purpose:* Enhance the accuracy of ASR output by adding punctuation.\n\n* *Model:* Employed and fine-tuned the [MuRIL: Multilingual Representations for Indian Languages](https://scholar.google.com/scholar_lookup?arxiv_id=2103.10730) model for accurate punctuation insertion in Bangla text.\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Dataset Pre-processing</center></div>","metadata":{}},{"cell_type":"markdown","source":"### Basic Steps:\n\n- **Normalization:** All transcriptions were normalized. Normalizer from [this](https://github.com/csebuetnlp/normalizer) was used.\n\n- **Removing Punctuation:** We removed the punctuations *(\"।\", \",\", \"?\")* before training ASR model which was later added by our punctuation model.\n\n- **Removal of Empty Transcripts:** All samples with empty transcriptions were removed.","metadata":{}},{"cell_type":"code","source":"before = train_df.shape[0]\ntrain_df.drop(train_df[train_df['transcripts'] == ''].index, inplace=True)\nafter = train_df.shape[0]\nprint(f'\\nThere were {before - after} such samples.\\n\\n')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T16:44:40.377547Z","iopub.execute_input":"2024-04-24T16:44:40.377843Z","iopub.status.idle":"2024-04-24T16:44:40.398025Z","shell.execute_reply.started":"2024-04-24T16:44:40.377818Z","shell.execute_reply":"2024-04-24T16:44:40.396954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **Handling of Incomprehensible Audio:** Samples with transcriptions solely consisting of \"<>\"—indicating entirely unclear audio—were removed. However, samples containing \"<>\" within the transcriptions, signaling only partial audio clarity issues, were retained due to their significant quantity.","metadata":{}},{"cell_type":"code","source":"before = train_df.shape[0]\ntrain_df.drop(train_df[train_df['transcripts'] == \"<>\"].index, inplace=True)\nafter = train_df.shape[0]\nprint(f'\\nThere were {before - after} such samples.\\n\\n')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T16:44:40.399624Z","iopub.execute_input":"2024-04-24T16:44:40.399976Z","iopub.status.idle":"2024-04-24T16:44:40.413710Z","shell.execute_reply.started":"2024-04-24T16:44:40.399947Z","shell.execute_reply":"2024-04-24T16:44:40.412569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **Exclusion of Placeholder Transcripts:** Samples that contained \"..\" as their entire transcription, likely placeholders for inaudible or missing content, were also removed.","metadata":{}},{"cell_type":"code","source":"before = train_df.shape[0]\ntrain_df.drop(train_df[train_df['transcripts'] == \"..\"].index, inplace=True)\nafter = train_df.shape[0]\nprint(f'\\nThere were {before - after} such samples.\\n\\n')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T16:44:40.415174Z","iopub.execute_input":"2024-04-24T16:44:40.415587Z","iopub.status.idle":"2024-04-24T16:44:40.432036Z","shell.execute_reply.started":"2024-04-24T16:44:40.415557Z","shell.execute_reply":"2024-04-24T16:44:40.430840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Data Augmentations**\n\n- **Time Stretching:**\n    - Rate: Ranges from 0.8 to 1.1.\n    - Probability: Applied with a 30% chance.\n\n\n- **Pitch Shifting:**\n    - Semitone Variation: The pitch is shifted down by 1.0 semitone.\n    - Probability: Applied with a 20% chance.\n\n\n- **Aliasing:**\n    - Sample Rate Reduction: Audio is downsampled to a fixed rate of 8000 Hz.\n    - Probability: Applied with a 30% chance.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Dialect Identification</center></div>","metadata":{}},{"cell_type":"markdown","source":"* We explored two approaches:\n    1. Training a audio classification model to identify the region of dialect of the speech. \n    2.  Extracting the region of dialect from the filename.\n\n\n* We explored two popular models for audio classification:\n    1. [Wav2Vec2](https://huggingface.co/docs/transformers/en/model_doc/wav2vec2) | **Accuray after training: 0.7286**\n    2. [Fine-tuned Whisper medium](https://www.kaggle.com/datasets/tugstugi/bengali-ai-asr-submission) | **Accuracy after training: 0.9962**","metadata":{}},{"cell_type":"code","source":"id2label = {\n    0: 'sylhet',\n    1: 'kishoreganj',\n    2: 'narail',\n    3: 'chittagong',\n    4: 'narsingdi',\n    5: 'tangail',\n    6: 'rangpur',\n    7: 'sandwip',\n    8: 'habiganj',\n    9: 'barishal'\n    \n}\n\nfrom pydub import AudioSegment\n\nexample = train_df.iloc[0]\n\nprint('\\n')\ndisplay(AudioSegment.from_file(example['file_path']))\n\nprint(f\"\\nDistrict: {example['district']}\\n\")","metadata":{"execution":{"iopub.status.busy":"2024-04-24T17:41:47.179514Z","iopub.execute_input":"2024-04-24T17:41:47.179913Z","iopub.status.idle":"2024-04-24T17:41:47.383245Z","shell.execute_reply.started":"2024-04-24T17:41:47.179885Z","shell.execute_reply":"2024-04-24T17:41:47.382244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from transformers import pipeline\n\nclassifier = pipeline(\"audio-classification\", model='/kaggle/input/whisper-md-bn-dialect-identification/regdia-bn-audio-classification')\nclassifier_preds = classifier(example['file_path'])\nclassifier_preds_df = pd.DataFrame(classifier_preds)\nprint('\\n\\n')\nprint(classifier_preds_df)\nprint('\\n')\nclassified_district = classifier_preds[0][\"label\"]\nprint(f'Classified District: {classified_district}')\nprint('\\n\\n')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T17:41:58.994473Z","iopub.execute_input":"2024-04-24T17:41:58.994867Z","iopub.status.idle":"2024-04-24T17:42:06.801175Z","shell.execute_reply.started":"2024-04-24T17:41:58.994836Z","shell.execute_reply":"2024-04-24T17:42:06.799800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Automatic Speech Recognition | ASR</center></div>","metadata":{}},{"cell_type":"markdown","source":"* *Purpose:* Convert dialect-specific speech into text.\n\n* *Details:* Fine-tuned separate ASR models for each district, fine-tuned on regional speech data without punctuation. All of them were multi-lingual [whisper-medium](https://huggingface.co/openai/whisper-medium) model fine-tuned on Bengali Dataset.\n\n* In [Bengali.AI Speech Recognition](https://www.kaggle.com/competitions/bengaliai-speech/overview) competition, team *chimege* became champion fine-tuning whisper-medium on competition and external bangla datasets. We fine-tuned their [fine-tuned model](https://www.kaggle.com/datasets/tugstugi/bengali-ai-asr-submission) for our regional dialects of each regions/districts.\n\n### Training Specifications:\n> Hardware spec: 1 * NVIDIA RTX A6000 GPU\n\n### Training Arguments:\n* Batch size: 8\n* Cosine learning rate scheduler with learning rate 0.00001\n* Max sequence length: 448\n* MAX_STEPS = {\n    \"sylhet\": 1201,\n    \"kishoreganj\": 701,\n    \"narail\": 601,\n    \"chittagong\": 601,\n    \"narsingdi\": 421,\n    \"tangail\": 421,\n    \"rangpur\": 421,\n    \"sandwip\": 421,\n    \"habiganj\": 361,\n    \"barishal\": 301,\n}","metadata":{}},{"cell_type":"code","source":"import torch\nimport logging\nimport transformers\nfrom transformers import WhisperForConditionalGeneration, WhisperFeatureExtractor, WhisperTokenizer\n\ntransformers.logging.set_verbosity_error()  # Show only errors, not warnings\n\nasr_model = f'/kaggle/input/whisper-md-bn-dialect-v2/whisper-md-bn-{classified_district}'\n\nmodel = WhisperForConditionalGeneration.from_pretrained(asr_model, torch_dtype=torch.float16)\nfeature_extractor = WhisperFeatureExtractor.from_pretrained(\"openai/whisper-medium\", torch_dtype=torch.float16)\ntokenizer = WhisperTokenizer.from_pretrained(asr_model, torch_dtype=torch.float16)\n\ntranscriber = pipeline(\n    task=\"automatic-speech-recognition\",\n    model=model,\n    feature_extractor=feature_extractor,\n    tokenizer=tokenizer,\n    chunk_length_s=30,\n    device=0,\n    batch_size=1,\n    torch_dtype=torch.float16\n)\n\ntranscriber.model.config.forced_decoder_ids = (\n    transcriber.tokenizer.get_decoder_prompt_ids(language=\"bn\", task=\"transcribe\")\n)\n\nasr_pred = transcriber(example['file_path'], generate_kwargs={\"max_length\": 448, \"num_beams\": 4})\n\nprint(f\"\\n\\nInitial Prediction Text:\\n\\n{asr_pred['text']}\\n\\n\")","metadata":{"execution":{"iopub.status.busy":"2024-04-24T17:36:01.258145Z","iopub.execute_input":"2024-04-24T17:36:01.258564Z","iopub.status.idle":"2024-04-24T17:36:09.298506Z","shell.execute_reply.started":"2024-04-24T17:36:01.258532Z","shell.execute_reply":"2024-04-24T17:36:09.297425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ASR Model Post-processing","metadata":{}},{"cell_type":"code","source":"def fix_repetition(text, max_count):\n    uniq_word_counter = {}\n    words = text.split()\n    for word in text.split():\n        if word not in uniq_word_counter:\n            uniq_word_counter[word] = 1\n        else:\n            uniq_word_counter[word] += 1\n\n    for word, count in uniq_word_counter.items():\n        if count > max_count:\n            words = [w for w in words if w != word]\n    text = \" \".join(words)\n    return text\n\npred_text = fix_repetition(asr_pred['text'].strip(), max_count=8)\nprint(f\"\\n\\nPost-processed Prediction Text:\\n\\n{pred_text}\\n\\n\")","metadata":{"execution":{"iopub.status.busy":"2024-04-24T16:45:56.432686Z","iopub.execute_input":"2024-04-24T16:45:56.433032Z","iopub.status.idle":"2024-04-24T16:45:56.442018Z","shell.execute_reply.started":"2024-04-24T16:45:56.433002Z","shell.execute_reply":"2024-04-24T16:45:56.440963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Punctuating Predicted Sentences</center></div>","metadata":{}},{"cell_type":"markdown","source":"* *Purpose:* Enhance the accuracy of ASR output by adding punctuation.\n\n* *Model:* Employed and fine-tuned the [MuRIL: Multilingual Representations for Indian Languages](https://scholar.google.com/scholar_lookup?arxiv_id=2103.10730) model for accurate punctuation insertion in Bangla text.\n\n* *Problem:* We modelled this problem as a token classification problem. \n\n* There are 4 labels.\n    - No puncutation: \"\"\n    - Full-stop: \"।\"\n    - Comma: \",\"\n    - Question-mark: \"?\"\n    \n### Training Specifications:\n> Hardware spec: 1 * NVIDIA RTX A6000 GPU\n\n### Training Arguments:\n* Batch size: 16\n* Learning rate 0.00002\n* Number of epochs: 20","metadata":{}},{"cell_type":"code","source":"from transformers import AutoTokenizer, AutoModelForTokenClassification\n\npunct_model_path = f\"/kaggle/input/muril-base-cased-bn-dialect-punctuation/muril-base-cased-bn-{classified_district}\"\ntokenizer = AutoTokenizer.from_pretrained(punct_model_path)\npunct_model = AutoModelForTokenClassification.from_pretrained(punct_model_path).cuda()\n\ndef punctuate(text, model, tokenizer):\n    input_ids = tokenizer(text).input_ids\n    with torch.no_grad():\n        logits = torch.nn.functional.softmax(\n            model(input_ids=torch.LongTensor([input_ids]).cuda()).logits[0, 1:-1], dim=1\n        ).cpu()\n        label_ids = torch.argmax(logits, dim=-1)\n\n        tokens = tokenizer(text, add_special_tokens=False).input_ids\n        punct_text = \"\"\n        for index, token in enumerate(tokens):\n            token_str = tokenizer.decode(token)\n            if \"##\" not in token_str:\n                punct_text += \" \" + token_str\n            else:\n                punct_text += token_str[2:]\n            punct_text += [\"\", \"।\", \",\", \"?\"][label_ids[index].item()]\n\n    punct_text = punct_text.strip()\n    return punct_text\n\ndef process(text, punct_model, tokenizer):\n    text = punctuate(text, punct_model, tokenizer)\n    words = text.split()\n    for i in range(len(words)):\n        if words[i] == '[UNK]':\n            words[i] = '<>'\n    text = \" \".join(words)\n    text = text.strip()\n    text = text.replace(' - ', '-')\n    if len(text) == 0:\n        text = '<>'\n    elif text[-1] not in [\"।\", \",\", \"?\"]:\n        text = text + \"।\"\n#     text = normalize(text)\n    return text\n\npunctuated_pred_text = process(pred_text, punct_model, tokenizer)\nprint(f'\\n\\nPunctuated Prediction Text:\\n\\n{punctuated_pred_text}\\n\\n')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T16:54:42.995524Z","iopub.execute_input":"2024-04-24T16:54:42.995999Z","iopub.status.idle":"2024-04-24T16:54:44.057473Z","shell.execute_reply.started":"2024-04-24T16:54:42.995962Z","shell.execute_reply":"2024-04-24T16:54:44.056410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Model Performances</center></div>","metadata":{}},{"cell_type":"markdown","source":"#### **One Model for All Regional Dialects -> WER: 0.76035**\n\n#### **Model Experts for Each Regional Dialects -> WER: 0.75720**","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; color:Black\"> Other Strategies Explored</center></div>","metadata":{}},{"cell_type":"markdown","source":"- **Observation:** Whisper model is particularly sensitive to noisy data. And the competiton dataset have a lot of bad sample ( multiple voices, unclear speech etc. ). So we tried to filter out these kind of bad samples.\n- **Strategy:** \n    - We inferred the whole training set by our trained model, calculated word-error-rate (WER) with the ground truth. \n    - We filtered out the train set removing the sample having WER greater than 0.9\n    - We trained our ASR models on filtered dataset. \n    - Though the evaluation WER improved on all models but the leaderboard score didn't improve.","metadata":{}}]}