{"cells":[{"metadata":{"_cell_guid":"b8ed3608-a001-450f-a4ff-df084cfb066c","_uuid":"ed0967a09a1b39bb7bf5d7c551f7d93b860598cb"},"cell_type":"markdown","source":"This notebook is currently a test site for augmenting data by looking at previous ideas and possibly trying new ones. Here are three ideas, one for adding translations, one for creating synthetic data, and one for interjecting noise/variations.\n\n\n### Translations\n\nThis idea originally came from the first Toxic challenge in which translation was used as an encoder-decoder. Pavel Ostyakov's [A simple technique for extending dataset](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038) explains how translating an english comment to another language, and then back to english, can improve model accuracy. With this being a multilingual competition, there have been similar ideas implemented but maybe not in this exact way.\n\nIn this notebook I'm using Google Translator via the googletrans package."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"collapsed":true,"trusted":false},"cell_type":"code","source":"!pip install git+https://github.com/ssut/py-googletrans.git","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"trusted":false},"cell_type":"code","source":"import os\nimport pandas as pd\nfrom googletrans import Translator","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"trusted":false},"cell_type":"code","source":"text = pd.read_csv('../input/jigsaw-multilingual-toxic-comment-classification/' \\\n                       'jigsaw-toxic-comment-train.csv', \n                        nrows=10_000)\n\n\n# TODO: Get the proportion of languages in the test set and set a randomized language per comment with np.random.choice()\n\ntranslator = Translator()\nfor i,t in enumerate(text.comment_text[19:22]):\n    try:\n        encoded = translator.translate(t, dest='fr').text\n        decoded = translator.translate(encoded, dest='en').text\n        print(f\"\\nSet {i}\\n\"\n              f\"Original: {t}\\n\\n\"\n              f\"Recoded: {decoded}\\n\")\n    except: pass","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Synthetic comments\n\nHere I'm using [Markovify](https://github.com/jsvine/markovify) to generate additional toxic commnents. This package uses Markov chains to string together new sequences of words based on previous sequences."},{"metadata":{"_cell_guid":"2788bcf1-935a-4725-8ac3-ceade79617e8","_uuid":"2962fea7e7e0c971957396ce555f54cbfc95f59c","collapsed":true,"trusted":false},"cell_type":"code","source":"import markovify as mk","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0bbd53b2-8f73-4d28-881c-a9955f4e9101","_uuid":"79365925736a09d9995f7e6c0c004a6a2786a1b5","collapsed":true,"trusted":false},"cell_type":"code","source":"doc = text.loc[text.toxic == 1, 'comment_text'].tolist()\ntext_model = mk.Text(doc)\nfor i in range(10):\n    print(text_model.make_sentence())\n","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"46fc3439-f61a-43d0-83ff-f6ac4b9d6790","_uuid":"6e0d70e140d6f93f8cc232f08addc628d3fc02cb"},"cell_type":"markdown","source":"Only a few lines of code and you too can sound like an angry 5th grader! Sometimes this technique produces a bit of nonsense but the toxic keywords are in there. It may be a way to add toxic comments and get a more balanced dataset."},{"metadata":{},"cell_type":"markdown","source":"### Various variations\n\nThe [nlpaug](https://github.com/makcedward/nlpaug) package contains a variety of augmentations to supplement text data and introduce noise that may help your model generalize. Here is a summary of a few functions:\n\n<img src=\"https://github.com/makcedward/nlpaug/blob/master/res/textual_example.png?raw=true\" width=\"600\">\n\nThe package has many more augmentations -at the character, word, and sentence levels. \n\n - Character Augmenter\n    - OCR\n    - Keyboard\n    - Random\n - Word Augmenter\n    - Spelling\n    - Word Embeddings\n    - TF-IDF\n    - Contextual Word Embeddings\n    - Synonym\n    - Antonym\n    - Random Word\n    - Split\n - Sentence Augmenter\n    - Contextual Word Embeddings for Sentence"},{"metadata":{"_cell_guid":"ef010c82-dc1f-49a4-a06b-274cc70e5de2","_uuid":"a7e7e4a092bc9da488ec3387d628c8edb592c815","collapsed":true},"cell_type":"markdown","source":"There's definiteiy more to explore here. Maybe these ideas can improve your model. Good luck!"},{"metadata":{"collapsed":true,"trusted":false},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.6.6"},"varInspector":{"cols":{"lenName":16,"lenType":16,"lenVar":40},"kernels_config":{"python":{"delete_cmd_postfix":"","delete_cmd_prefix":"del ","library":"var_list.py","varRefreshCmd":"print(var_dic_list())"},"r":{"delete_cmd_postfix":") ","delete_cmd_prefix":"rm(","library":"var_list.r","varRefreshCmd":"cat(var_dic_list()) "}},"position":{"height":"652px","left":"1458px","right":"20px","top":"172px","width":"350px"},"types_to_exclude":["module","function","builtin_function_or_method","instance","_Feature"],"window_display":true}},"nbformat":4,"nbformat_minor":4}