{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Training a custom tokenizer\n\nTokenizer breaks the raw text into smaller parts. It helps in developing models for the NLP tasks. Here, we trained a custom tokenizer for the competition.\n\nDocumentation: https://huggingface.co/docs/tokenizers/index","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\n\nfrom itertools import chain\nfrom tokenizers import BertWordPieceTokenizer\nfrom tokenizers.processors import TemplateProcessing","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-01T17:20:53.672068Z","iopub.execute_input":"2022-07-01T17:20:53.672861Z","iopub.status.idle":"2022-07-01T17:20:53.799215Z","shell.execute_reply.started":"2022-07-01T17:20:53.672762Z","shell.execute_reply":"2022-07-01T17:20:53.798115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv('../input/dlsprint/train.csv')\nvalid_df = pd.read_csv('../input/dlsprint/validation.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-01T17:21:01.847404Z","iopub.execute_input":"2022-07-01T17:21:01.847791Z","iopub.status.idle":"2022-07-01T17:21:03.978481Z","shell.execute_reply.started":"2022-07-01T17:21:01.847761Z","shell.execute_reply":"2022-07-01T17:21:03.977558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We put the texts into files for training our custom tokenizer.\n\nTEXT_PATH = 'texts'\n\ntry:\n    os.mkdir(TEXT_PATH)\nexcept FileExistsError:\n    pass\n\npaths = []\n\nfor i, text in enumerate(chain(train_df['sentence'], valid_df['sentence'])):\n    path = f'{TEXT_PATH}/text{i}.txt'\n    with open(path, 'w') as f:\n        f.write(text)\n    paths.append(path)\n\nprint(f'Successfully written {i+1} text files!')","metadata":{"execution":{"iopub.status.busy":"2022-07-01T17:21:21.887669Z","iopub.execute_input":"2022-07-01T17:21:21.888404Z","iopub.status.idle":"2022-07-01T17:21:32.585889Z","shell.execute_reply.started":"2022-07-01T17:21:21.888370Z","shell.execute_reply":"2022-07-01T17:21:32.583645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tokenizer = BertWordPieceTokenizer(\n    clean_text=True,\n    handle_chinese_chars=False,\n    strip_accents=False,\n    lowercase=False\n)\n\ntokenizer.train(\n    files=paths, vocab_size=100_000, min_frequency=1,\n    limit_alphabet=1000, wordpieces_prefix='##',\n    special_tokens=[\n        '[PAD]', '[UNK]', '[CLS]', '[SEP]', '[MASK]'\n])","metadata":{"execution":{"iopub.status.busy":"2022-07-01T17:21:39.517874Z","iopub.execute_input":"2022-07-01T17:21:39.518306Z","iopub.status.idle":"2022-07-01T17:21:46.659985Z","shell.execute_reply.started":"2022-07-01T17:21:39.518259Z","shell.execute_reply":"2022-07-01T17:21:46.658795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We save the model after training for reusability. You can download the './tokenizer-vocab.txt' file if you want\n\n[vocab_path] = tokenizer.save_model('./', 'tokenizer')\nprint(f'The tokenizer is saved in the file \"{vocab_path}\"')","metadata":{"execution":{"iopub.status.busy":"2022-07-01T17:23:02.627509Z","iopub.execute_input":"2022-07-01T17:23:02.628461Z","iopub.status.idle":"2022-07-01T17:23:02.667589Z","shell.execute_reply.started":"2022-07-01T17:23:02.628413Z","shell.execute_reply":"2022-07-01T17:23:02.666509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tokenizer = BertWordPieceTokenizer('./tokenizer-vocab.txt')","metadata":{"execution":{"iopub.status.busy":"2022-07-01T17:23:11.368589Z","iopub.execute_input":"2022-07-01T17:23:11.369319Z","iopub.status.idle":"2022-07-01T17:23:11.434720Z","shell.execute_reply.started":"2022-07-01T17:23:11.369266Z","shell.execute_reply":"2022-07-01T17:23:11.433811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tokenizer.post_processor = TemplateProcessing(\n    single=\"[CLS] $A [SEP]\",\n    special_tokens=[\n        (\"[CLS]\", 2),\n        (\"[SEP]\", 3),\n    ],\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-01T17:23:47.828034Z","iopub.execute_input":"2022-07-01T17:23:47.828431Z","iopub.status.idle":"2022-07-01T17:23:47.833981Z","shell.execute_reply.started":"2022-07-01T17:23:47.828400Z","shell.execute_reply":"2022-07-01T17:23:47.833010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<code>[CLS]</code> token is used as the start token. <code>[SEP]</code> token is used as the end token","metadata":{}},{"cell_type":"code","source":"encodings = tokenizer.encode('আমি বাংলার গান গাই')\nencodings.tokens","metadata":{"execution":{"iopub.status.busy":"2022-07-01T17:30:14.935693Z","iopub.execute_input":"2022-07-01T17:30:14.936695Z","iopub.status.idle":"2022-07-01T17:30:14.945000Z","shell.execute_reply.started":"2022-07-01T17:30:14.936650Z","shell.execute_reply":"2022-07-01T17:30:14.943961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluation Metric\n\nTo evaluate models, we need an evaluation metric. For this competition (phase1), Levenshtein Mean Distance for evaluation. Luckily, the implementation of the metric can be found in <code>nltk</code>! A great explanation of the algorithm can be found [in this video](https://www.youtube.com/watch?v=XYi2-LPrwm4).","metadata":{}},{"cell_type":"code","source":"from nltk.metrics.distance import edit_distance\n\nstring1 = \"আসিফ\"\nstring2 = \"আসিব\"\n  \n# the Levenshtein distance between string1 and string2\nprint(edit_distance(string1, string2))","metadata":{"execution":{"iopub.status.busy":"2022-07-01T17:35:31.939138Z","iopub.execute_input":"2022-07-01T17:35:31.940099Z","iopub.status.idle":"2022-07-01T17:35:31.947194Z","shell.execute_reply.started":"2022-07-01T17:35:31.940047Z","shell.execute_reply":"2022-07-01T17:35:31.945935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Additional Resources for beginners\n\nThough we used word piece tokenizer in this notebook, you can try out other subword tokenizers if you want. Such as:\n\n1. [Byte Pair Encoding Tokenizer](https://www.youtube.com/watch?v=-0IjF-7OB3s&t=432s)\n2. [Sentence Piece Tokenizer](https://www.youtube.com/watch?v=U51ranzJBpY)\n3. [Unigram Tokenizer](https://www.youtube.com/watch?v=TGZfZVuF9Yc)\n\nIf you are looking for a deep learning model for the task, you might wanna check out the transformer models. Here are are some resources if you want to learn more about transformers:\n\n1. [The Illustrated Transformer](http://jalammar.github.io/illustrated-transformer/)\n2. [Attention Is All You Need](https://papers.nips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf)\n\nI hope you learn something new out of this competition. Best of luck!😊","metadata":{}}]}