{"cells":[{"metadata":{"_uuid":"b5bbc8a12ad4a9b8f5d27269c27301d2af09cc7d","_cell_guid":"2b094da0-433f-4bbd-a660-91db4e33503e"},"cell_type":"markdown","source":"# NLP Improvements: Peculiarities of Russian Inflection Diminish TF-IDF Quality\n### Maximizing TF-IDF scores in Russian NLP applications\n\n\nAnother competitor [brilliantly shared](https://www.kaggle.com/iggisv9t/handling-russian-language-inflectional-structure) that the Russian language has peculiar inflectional structure, so the same word can be spelled different ways within different contexts.  This can fool our TF-IDF which is built for languages without these structural pecularities.\n\n**For example:**\n* Dog -> Собак**а**\n* No dog -> нет собак**и**\n* Give a dog a bone -> Дай собак**е** кость.\n\nAs you can see, the last letter of the word \"dog\" changed within different contexts.  It can get even more complicated.  We need to account for this in our TF-IDF.\n\nWith package `pymorph2` and TF-IDF, it will be simple to *normalize* the text.  First, we must define a function called `normalize`.  `normalize` depends on `retoken` and `morph`, which we need to import `re` and `pymorphy2` for; `pymorphy2` requires a special installation from the kernel settings."},{"metadata":{"collapsed":true,"_cell_guid":"ff7ac2e3-0bc0-43ea-ab54-9057d6a80e04","_uuid":"edb525aadac17439fd39a330e94215dad05d4a9e","trusted":false},"cell_type":"code","source":"import os\nimport pymorphy2\nimport re\nimport pandas as pd\n\ntrain = pd.read_csv('../input/train.csv')\ntest = pd.read_csv('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"_cell_guid":"7018aaf0-5d1f-44ca-93cb-80ebc9e0d97b","_uuid":"cddc436c4f908c64aab6f6a6e28405acf01184ea","trusted":false},"cell_type":"code","source":"morph = pymorphy2.MorphAnalyzer()\nretoken = re.compile(r'[\\'\\w\\-]+')\ndef normalize(text):\n    text = retoken.findall(text.lower()) # make all text lowercase\n    text = [morph.parse(x)[0].normal_form for x in text] # morphological analysis\n    return ' '.join(text)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"018d71dc2847da1cb2dee931de665a0a71df5559","_cell_guid":"17e9fa1c-c5dc-4ede-b9bf-717293d87f8b"},"cell_type":"markdown","source":"Here we normalize all of the text... it takes a while:"},{"metadata":{"collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"code","source":"%%time\ntrain['title'] = train['title'].astype(str)\ntrain['description'] = train['description'].astype(str)\ntest['title'] = test['title'].astype(str)\ntest['description'] = test['description'].astype(str)","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"_cell_guid":"39678a01-1bf5-46fe-809e-bd9652f776b4","_uuid":"8409f3b71e0a4b9048cc62c6cbfd08ba06cb8614","trusted":false},"cell_type":"code","source":"%%time\ntrain['title'] = train['title'].apply(normalize)\ntrain['description'] = train['description'].apply(normalize)\ntrain.to_csv(\"updnlp-train.csv\"); del train\ntest['title'] = test['title'].apply(normalize)\ntest['description'] = test['description'].apply(normalize)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"57c1994e48e75a3849a4cd1e76eb05d2a6388ec4","_cell_guid":"d8f5a64f-94d8-4e57-876e-0ab6d20d511a"},"cell_type":"markdown","source":"You can see the results of this function (it does work!):"},{"metadata":{"collapsed":true,"_cell_guid":"c2a4588e-0717-41e0-af19-8fa14e245d73","_uuid":"ee5d90d97d607dd0d5113d1f5d70af746de81fee","trusted":false},"cell_type":"code","source":"print(normalize('собака'))\nprint(normalize('нет собаки'))\nprint(normalize('Дай собаке кость.'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"15c7799eb73c832bfab99cb1d670cb001aebd1a8","_cell_guid":"66a33660-80cb-4e3a-a985-f83fa88f8ce4"},"cell_type":"markdown","source":"Now, we will write the normalized text data to a `csv` file, which you can use to train your TF-IDF.  If you're running your model on the Kaggle platform, I suggest importing this kernel's output into your own kernel to perform TF-IDF.  Otherwise, you can download the output.  **This will decrease document size and improve the accuracy of your TF-IDF!**"},{"metadata":{"collapsed":true,"_cell_guid":"0b8259f1-37ef-4414-97a3-63cd50c5e5ba","_uuid":"ebe4d9b3f4e288540d8c79bb572af7fa74dfc93d","trusted":false},"cell_type":"code","source":"test.to_csv(\"updnlp-test.csv\")","execution_count":null,"outputs":[]}],"metadata":{"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}