{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-08T13:30:19.936223Z","iopub.execute_input":"2022-07-08T13:30:19.936611Z","iopub.status.idle":"2022-07-08T13:30:30.555819Z","shell.execute_reply.started":"2022-07-08T13:30:19.936581Z","shell.execute_reply":"2022-07-08T13:30:30.554729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# [WEBINAR LINK](https://www.facebook.com/bengaliAI/videos/2911076665863989)","metadata":{}},{"cell_type":"markdown","source":"# Webinar:Vocabulary selection\n* normalized and non normalized text\n* legacy symbols\n* you don't get what you see\n* \"\\u200d\"","metadata":{}},{"cell_type":"code","source":"!pip install bnunicodenormalizer","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:32:18.821097Z","iopub.execute_input":"2022-07-08T19:32:18.821544Z","iopub.status.idle":"2022-07-08T19:32:34.496713Z","shell.execute_reply.started":"2022-07-08T19:32:18.821447Z","shell.execute_reply":"2022-07-08T19:32:34.495237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#-------------------------------\n# imports\n#-------------------------------\nimport os\nimport pandas as pd \nimport numpy as np \nfrom tqdm.auto import tqdm\nfrom pandarallel import pandarallel\nfrom bnunicodenormalizer import Normalizer \npandarallel.initialize(progress_bar=True,nb_workers=8)\ntqdm.pandas()\nbnorm=Normalizer()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:33:10.510769Z","iopub.execute_input":"2022-07-08T19:33:10.511276Z","iopub.status.idle":"2022-07-08T19:33:10.746541Z","shell.execute_reply.started":"2022-07-08T19:33:10.511237Z","shell.execute_reply":"2022-07-08T19:33:10.745283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* read the data and find unique unicodes","metadata":{}},{"cell_type":"code","source":"train_df=pd.read_csv(\"../input/dlsprint/train.csv\")\ntrain_df=train_df[[\"sentence\"]]\nval_df=pd.read_csv(\"../input/dlsprint/validation.csv\")\nval_df=val_df[[\"sentence\"]]\n\nsens=train_df[\"sentence\"].tolist()+val_df[\"sentence\"].tolist()\nprint(\"number of total sentences:\",len(sens))\nvocab=[]\nfor sen in tqdm(sens):\n    sen=sen.replace('”','')\n    for c in sen:\n        if c not in vocab:\n            vocab.append(c)\nnon_norm_vocab=sorted(vocab)\nprint(\"non normalized vocab(unicodes):\")\nfor idx,c in enumerate(non_norm_vocab):\n    if idx%20==0:print()\n    print(f\"'{c}'\",end=\",\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:35:00.341383Z","iopub.execute_input":"2022-07-08T19:35:00.341797Z","iopub.status.idle":"2022-07-08T19:35:08.819768Z","shell.execute_reply.started":"2022-07-08T19:35:00.341763Z","shell.execute_reply":"2022-07-08T19:35:08.818291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**unicodes that look the same but are not same**","metadata":{}},{"cell_type":"code","source":"print(f\"index of '–':\",non_norm_vocab.index('–'))\nprint(f\"index of '–':\",non_norm_vocab.index('—'))\nprint(\"are they same?:\",'–'=='—')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:38:25.052398Z","iopub.execute_input":"2022-07-08T19:38:25.052857Z","iopub.status.idle":"2022-07-08T19:38:25.060593Z","shell.execute_reply.started":"2022-07-08T19:38:25.052822Z","shell.execute_reply":"2022-07-08T19:38:25.059633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"index of '।':\",non_norm_vocab.index('।'))\nprint(f\"index of '৷':\",non_norm_vocab.index('৷'))\nprint(\"are they same?:\",'।'=='৷')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:40:52.606472Z","iopub.execute_input":"2022-07-08T19:40:52.606959Z","iopub.status.idle":"2022-07-08T19:40:52.614861Z","shell.execute_reply.started":"2022-07-08T19:40:52.606921Z","shell.execute_reply":"2022-07-08T19:40:52.613515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"index of ',':\",non_norm_vocab.index(','))\nprint(f\"index of '‚':\",non_norm_vocab.index('‚'))\nprint(\"are they same?:\",','=='‚')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* '৵' is not ৯ but a legacy symbol used in bangla\n\n**Unicode legacy blocks**\n* source: https://github.com/mnansary/bnUnicodeNormalizer/blob/main/bnunicodenormalizer/normalizer.py#L25\n```\n'৺':Isshar \n'৻':Ganda\n'ঀ':Anji (not '৭')\n'ঌ':li\n'ৡ':dirgho li\n'ঽ':Avagraha\n'ৠ':Vocalic Rr (not 'ঋ')\n'৲':rupi\n'৴':currency numerator 1\n'৵':currency numerator 2\n'৶':currency numerator 3\n'৷':currency numerator 4\n'৸':currency numerator one less than the denominator\n'৹':Currency Denominator Sixteen\n```\n\n* 'ৰ' is an assamese unicode \n\n**BUT THE MOST DANGEROUS THING IN THIS TEXT IS** Nukta **'়'**","metadata":{}},{"cell_type":"code","source":"print(f\"index of nukta:\",non_norm_vocab.index('়'))","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:47:10.798905Z","iopub.execute_input":"2022-07-08T19:47:10.799965Z","iopub.status.idle":"2022-07-08T19:47:10.805230Z","shell.execute_reply.started":"2022-07-08T19:47:10.799911Z","shell.execute_reply":"2022-07-08T19:47:10.804313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* What this means is text is getting broken\n* lets consider an example: 'কেন্দ্রীয়'=='কেন্দ্রীয়'\n    * they both look the same but are they?\n","metadata":{}},{"cell_type":"code","source":"'কেন্দ্রীয়'=='কেন্দ্রীয়'","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:48:31.014596Z","iopub.execute_input":"2022-07-08T19:48:31.015010Z","iopub.status.idle":"2022-07-08T19:48:31.023451Z","shell.execute_reply.started":"2022-07-08T19:48:31.014969Z","shell.execute_reply":"2022-07-08T19:48:31.022405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"But why? **because the first one contains nukta** \n","metadata":{}},{"cell_type":"code","source":"print(\"first one:\",[f for f in 'কেন্দ্রীয়'])\nprint(\"second one:\",[f for f in 'কেন্দ্রীয়'])\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:49:48.194228Z","iopub.execute_input":"2022-07-08T19:49:48.194646Z","iopub.status.idle":"2022-07-08T19:49:48.201750Z","shell.execute_reply.started":"2022-07-08T19:49:48.194611Z","shell.execute_reply":"2022-07-08T19:49:48.200766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**So now lets solve these problems by normalizing**\n\nvisit this github for normalization process details: https://github.com/mnansary/bnUnicodeNormalizer","metadata":{}},{"cell_type":"code","source":"def normalize(sen):\n    _words = [bnorm(word)['normalized']  for word in sen.split()]\n    return \" \".join([word for word in _words if word is not None]) \n\nval_df[\"sentence\"]=val_df[\"sentence\"].parallel_apply(lambda x:normalize(x))\ntrain_df[\"sentence\"]=train_df[\"sentence\"].parallel_apply(lambda x:normalize(x))\n\nsens=train_df[\"sentence\"].tolist()+val_df[\"sentence\"].tolist()\nprint(\"number of total sentences:\",len(sens))\nvocab=[]\nfor sen in tqdm(sens):\n    sen=sen.replace('”','')\n    for c in sen:\n        if c not in vocab:\n            vocab.append(c)\nnorm_vocab=sorted(vocab)\nprint(\"vocab(unicodes):\")\nfor idx,c in enumerate(norm_vocab):\n    if idx%20==0:print()\n    print(f\"'{c}'\",end=\",\")        ","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:53:20.374182Z","iopub.execute_input":"2022-07-08T19:53:20.374556Z","iopub.status.idle":"2022-07-08T20:01:35.572720Z","shell.execute_reply.started":"2022-07-08T19:53:20.374525Z","shell.execute_reply":"2022-07-08T20:01:35.571523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Removed symbols","metadata":{}},{"cell_type":"code","source":"for c in non_norm_vocab:\n    if c not in norm_vocab:\n        print(f\"'{c}'\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T20:02:27.436580Z","iopub.execute_input":"2022-07-08T20:02:27.436999Z","iopub.status.idle":"2022-07-08T20:02:27.443344Z","shell.execute_reply.started":"2022-07-08T20:02:27.436958Z","shell.execute_reply":"2022-07-08T20:02:27.442210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will also remove '—' from normalized-vocab since '-' has the same functional value","metadata":{}},{"cell_type":"markdown","source":"* we also have to consider \"\\u200d\" to cover words like-","metadata":{}},{"cell_type":"code","source":"words=[\"র‍্যাব\", \"র‍্যাকেট\", \"র‍্যাশানাল\"]\nfor word in words:\n    print(word)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T20:07:11.016650Z","iopub.execute_input":"2022-07-08T20:07:11.017054Z","iopub.status.idle":"2022-07-08T20:07:11.022732Z","shell.execute_reply.started":"2022-07-08T20:07:11.017020Z","shell.execute_reply":"2022-07-08T20:07:11.021925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"basically any word that has","metadata":{}},{"cell_type":"code","source":"print('র'+'\\u200d'+'্'+'য')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T20:07:32.785382Z","iopub.execute_input":"2022-07-08T20:07:32.785749Z","iopub.status.idle":"2022-07-08T20:07:32.791638Z","shell.execute_reply.started":"2022-07-08T20:07:32.785720Z","shell.execute_reply":"2022-07-08T20:07:32.790461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"in the text. Simply adding 'র'+'্'+'য' results in","metadata":{}},{"cell_type":"code","source":"print('র'+'্'+'য')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T20:08:36.281560Z","iopub.execute_input":"2022-07-08T20:08:36.281942Z","iopub.status.idle":"2022-07-08T20:08:36.288027Z","shell.execute_reply.started":"2022-07-08T20:08:36.281910Z","shell.execute_reply":"2022-07-08T20:08:36.286816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"final vocab we can go with (or even remove some puntuations and numbers if we want)","metadata":{}},{"cell_type":"code","source":"\nvocab=[ '\\u200d',\n        ' ','!',\"'\",',','-','.',':',';','=','?','।',\n        'ঁ','ং','ঃ',\n        'অ','আ','ই','ঈ','উ','ঊ','ঋ','এ','ঐ','ও','ঔ',\n        'ক','খ','গ','ঘ','ঙ',\n        'চ','ছ','জ','ঝ','ঞ',\n        'ট','ঠ','ড','ঢ','ণ',\n        'ত','থ','দ','ধ','ন',\n        'প','ফ','ব','ভ','ম',\n        'য','র','ল',\n        'শ','ষ','স','হ',\n        'া','ি','ী','ু','ূ','ৃ','ে','ৈ','ো','ৌ','্',\n        'ৎ','ড়','ঢ়','য়',\n        '০','১','২','৩','৪','৫','৬','৭','৮','৯']","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"to explore non norm text try the following","metadata":{}},{"cell_type":"code","source":"import unicodedata\ntext=\"বায়ান্ন\"\nnew_text = unicodedata.normalize(\"NFKC\", text)\nprint(\"text:\",text)\nprint(\"new text:\",new_text)\nprint(\"are they same?:\",text==new_text)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T20:11:44.101792Z","iopub.execute_input":"2022-07-08T20:11:44.102224Z","iopub.status.idle":"2022-07-08T20:11:44.110544Z","shell.execute_reply.started":"2022-07-08T20:11:44.102190Z","shell.execute_reply":"2022-07-08T20:11:44.108847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"this will cause problem in **WER** and **CER** calculation","metadata":{}},{"cell_type":"markdown","source":"# Webinar:Understanding audio files\n### * mp3 loading takes a lot of time. Converting to wav is the way to go. \n\n[**THIS DISCUSSION THREAD AND FIRST COMMENT HOLDS THE DATASETS AND NOTEBOOKS USED TO CONVERT mp3 to wav**](https://www.kaggle.com/competitions/dlsprint/discussion/334477)\n\n### * avoiding problematic audio files and filtering votes\n\nlist of problematic files : [from this notebook](https://www.kaggle.com/code/nazmuddhohaansary/tfrecords-for-tpu-training?scriptVersionId=100316555)\n\n```python\nerrors=[\"common_voice_bn_31727562\",\n        'common_voice_bn_30998934',\n        'common_voice_bn_31595526',\n        'common_voice_bn_31534853',\n        'common_voice_bn_31518061',\n        'common_voice_bn_31518373',\n        'common_voice_bn_31613621',\n        'common_voice_bn_31555333',\n        'common_voice_bn_31772113',\n        'common_voice_bn_31605391',\n        'common_voice_bn_31631175',\n        'common_voice_bn_31563901',\n        'common_voice_bn_31691690',\n        'common_voice_bn_31692010',\n        'common_voice_bn_31683653',\n        'common_voice_bn_31692182',\n        'common_voice_bn_31519976',\n        'common_voice_bn_31675793',\n        'common_voice_bn_31019914',\n        'common_voice_bn_31660287',\n        'common_voice_bn_31660384',\n        'common_voice_bn_31557261',\n        'common_voice_bn_31633101',\n        'common_voice_bn_31599243',\n        'common_voice_bn_31521515',\n        'common_voice_bn_31777802',\n        'common_voice_bn_31777848',\n        'common_voice_bn_31669646',\n        'common_voice_bn_31566083',\n        'common_voice_bn_31530331',\n        'common_voice_bn_31727697',\n        'common_voice_bn_31513270',\n        'common_voice_bn_31686295',\n        'common_voice_bn_31753693',\n        'common_voice_bn_31686334',\n        'common_voice_bn_31765546',\n        'common_voice_bn_31765548',\n        'common_voice_bn_31662742',\n        'common_voice_bn_31704856',\n        'common_voice_bn_31635344',\n        'common_voice_bn_31618327',\n        'common_voice_bn_31743074',\n        'common_voice_bn_31678862',\n        'common_voice_bn_31626674',\n        'common_voice_bn_31626677',\n        'common_voice_bn_31523889',\n        'common_voice_bn_31610804',\n        'common_voice_bn_31769538',\n        'common_voice_bn_31533273',\n        'common_voice_bn_31445621',\n        'common_voice_bn_31620650']\n```\n### * exploring spectograms as a way to view audio as images and apply cnn models\n[**This Notebook was explored for log mel spectogram and its design descisions while the webiner was going on**](https://www.kaggle.com/code/nazmuddhohaansary/logmelspctogram-basic-cnn-attention-modeling?scriptVersionId=100250051)\n* see the Log Mel Spectrogram section for details\n* [see this notebook for other ways of visualizing the data](https://www.kaggle.com/code/umongsain/spectrograms-quick-tour)\n* **THE MOST IMPORTANT DESIGN DESCISIONS ARE DISCUSSED IN THE WEBINER** \n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Webinar:Modeling\n* [Wave2vec2](https://www.kaggle.com/code/nazmuddhohaansary/wave2vec2-starter-for-dl-sprint-commonvoice)\n    * use huggingface hub to upload checkpoints and restore training\n    * why wave2vec2? [example leader board on asr](https://paperswithcode.com/task/speech-recognition)\n    \n* Non end-2-end deep learning approached: hybrid-kaldi,hmm-gmm\n* deep learning models: NVIDIA - Nemo covers almost all state of the art models: https://github.com/NVIDIA/NeMo\n* [custom positional attention modeling](https://www.kaggle.com/code/nazmuddhohaansary/logmelspctogram-basic-cnn-attention-modeling?scriptVersionId=100250051)\n    * this reaches Char acc > 0.1 unlike CTC models that needs lots of training\n    * this type of model can be of atleast 100 variants and given the option to extend upto 700 with\n        * seq2seq decoder\n        * different feature/mixed feature \n        * different attention strategy \n    * **BUT TRAINING WITH GPU IS REALLY TIME CONSUMING**: SOLUTION--> TFRecords and TPUS\n* [tfrecord creation for TPU](https://www.kaggle.com/code/nazmuddhohaansary/tfrecords-for-tpu-training?scriptVersionId=100276326)\n* [faster training with TPU](https://www.kaggle.com/code/nazmuddhohaansary/tpu-training-basic-cnn-attention-dl-sprint)\n* [using colab to save kaggle TPU time since kaggle TPU time is limited](https://drive.google.com/file/d/1dBKOTTbVxnBiwvITKfXMih6A6k-VrcBt/view?usp=sharing)\n* [tf-tpu-jasper](https://www.kaggle.com/code/nazmuddhohaansary/tfasr-jasper)\n* [tf-tpu-deepspeechv2](https://www.kaggle.com/code/nazmuddhohaansary/tfasr-deepspeechv2)\n**NOTE: Training these models from scratch takes a lot of time**  \n    \n\n\n# Acknowledgement for Technical Support:[APSIS Solutions Limited](https://apsissolutions.com/)\n* to test data\n* to process data locally\n* to create datasets that exceeds kaggle output limit\n\n![](https://apsissolutions.com/wp-content/uploads/MicrosoftTeams-image-21-1.png)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}