{"cells":[{"metadata":{},"cell_type":"markdown","source":"# MarianMT","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"This notebook is mostly for educational purpose. Transformers released a new function called [MarianMT](https://huggingface.co/transformers/model_doc/marian.html) which seems to be very powerful to translate text data.\n\nIn this notebook I used a small subset of the validation data because performing translations using this method takes a lot of time and probably doesn't give an edge compared to dataset transleted using tradtionnal methods.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install -U transformers","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load data","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nfrom tqdm.notebook import tqdm","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_csv('../input/jigsaw-multilingual-toxic-comment-classification/validation.csv')\ndf = df.sample(100, random_state=12)\ndf.head(3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Tokenize & translate","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df['lang'].unique()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"from transformers import MarianMTModel, MarianTokenizer","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df['content_english'] = ''","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for i, lang in tqdm(enumerate(['es', 'it', 'tr'])):\n    if lang in ['es', 'it']:\n        model_name = 'Helsinki-NLP/opus-mt-ROMANCE-en'\n        df_lang = df.loc[df['lang']==lang, 'comment_text'].apply(lambda x: '>>{}<< '.format(lang) + x)\n    else:\n        model_name = 'Helsinki-NLP/opus-mt-{}-en'.format(lang)\n        df_lang = df.loc[df['lang']==lang, 'comment_text']\n    \n    tokenizer = MarianTokenizer.from_pretrained(model_name)\n    model = MarianMTModel.from_pretrained(model_name, output_loading_info=False)\n        \n    batch = tokenizer.prepare_translation_batch(df_lang.values,\n                                               max_length=192,\n                                               pad_to_max_length=True)\n    translated = model.generate(**batch)\n\n    df.loc[df['lang']==lang, 'content_english'] = [tokenizer.decode(t, skip_special_tokens=True) for t in translated]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.to_csv(\"df_translated.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"MarianMT offers a wider range of things you can do, pleaste take a look at [the official documentation](https://huggingface.co/transformers/model_doc/marian.html).","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}