{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":10737,"databundleVersionId":290346,"sourceType":"competition"}],"dockerImageVersionId":13801,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this kernel I want to illustrate how I do come up with meaningful preprocessing when building deep learning NLP models. \n\nI start with two golden rules:\n\n1.  **Don't use standard preprocessing steps like stemming or stopword removal when you have pre-trained embeddings** \n\nSome of you might used standard preprocessing steps when doing word count based feature extraction (e.g. TFIDF) such as removing stopwords, stemming etc. \nThe reason is simple: You loose valuable information, which would help your NN to figure things out.  \n\n2. **Get your vocabulary as close to the embeddings as possible**\n\nI will focus in this notebook, how to achieve that. For an example I take the GoogleNews pretrained embeddings, there is no deeper reason for this choice.","metadata":{"_uuid":"5fb07eae7b445e3bf358222065133a144b7c4ade"}},{"cell_type":"markdown","source":"","metadata":{"_uuid":"2b5f2c75786e9b8194938a420438bc1166e55ec3"}},{"cell_type":"markdown","source":"We start with a neat little trick that enables us to see a progressbar when applying functions to a pandas Dataframe","metadata":{"_uuid":"ec55ed96a9cab0ac75b22f9710c0393b17cfefd2"}},{"cell_type":"code","source":"import pandas as pd\nfrom tqdm import tqdm\ntqdm.pandas()\n","metadata":{"trusted":true,"_uuid":"b378958a9606ac48fe0dc54e24bed4cd503e0ac7"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Lets load our data","metadata":{"_uuid":"6e432bb170a329c79f58f78527bbbe4e857b2c41"}},{"cell_type":"code","source":"train = pd.read_csv(\"../input/train.csv\")\ntest = pd.read_csv(\"../input/test.csv\")\nprint(\"Train shape : \",train.shape)\nprint(\"Test shape : \",test.shape)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I will use the following function to track our training vocabulary, which goes through all our text and counts the occurance of the contained words. ","metadata":{"_uuid":"59fd43e3727bd265df08d2d8254230c655c51f9d"}},{"cell_type":"code","source":"def build_vocab(sentences, verbose =  True):\n    \"\"\"\n    :param sentences: list of list of words\n    :return: dictionary of words and their count\n    \"\"\"\n    vocab = {}\n    for sentence in tqdm(sentences, disable = (not verbose)):\n        for word in sentence:\n            try:\n                vocab[word] += 1\n            except KeyError:\n                vocab[word] = 1\n    return vocab","metadata":{"trusted":true,"_uuid":"3e050f2fa9668c765466d766cfd79720cf6cc819"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"So lets populate the vocabulary and display the first 5 elements and their count. Note that now we can use progess_apply to see progress bar","metadata":{"_uuid":"a453d10106aea7f6bf3f9bd952f2e02dd7a7472c"}},{"cell_type":"code","source":"sentences = train[\"question_text\"].progress_apply(lambda x: x.split()).values\nvocab = build_vocab(sentences)\nprint({k: vocab[k] for k in list(vocab)[:5]})","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Next we import the embeddings we want to use in our model later. For illustration I use GoogleNews here.","metadata":{"_uuid":"c5c1c2c96aecd40e969f13913e850a36d8e3a13d"}},{"cell_type":"code","source":"from gensim.models import KeyedVectors\n\nnews_path = '../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin'\nembeddings_index = KeyedVectors.load_word2vec_format(news_path, binary=True)","metadata":{"trusted":true,"_uuid":"807e734e0ce617480c824f8bf26f0672f38397d5"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Next I define a function that checks the intersection between our vocabulary and the embeddings. It will output a list of out of vocabulary (oov) words that we can use to improve our preprocessing","metadata":{"_uuid":"b5f9048e53a5cb91bdaa518d33fda756f1038562"}},{"cell_type":"code","source":"import operator \n\ndef check_coverage(vocab,embeddings_index):\n    a = {}\n    oov = {}\n    k = 0\n    i = 0\n    for word in tqdm(vocab):\n        try:\n            a[word] = embeddings_index[word]\n            k += vocab[word]\n        except:\n\n            oov[word] = vocab[word]\n            i += vocab[word]\n            pass\n\n    print('Found embeddings for {:.2%} of vocab'.format(len(a) / len(vocab)))\n    print('Found embeddings for  {:.2%} of all text'.format(k / (k + i)))\n    sorted_x = sorted(oov.items(), key=operator.itemgetter(1))[::-1]\n\n    return sorted_x","metadata":{"trusted":true,"_uuid":"f35a7213fc9a7e80a7c210d11b3a8094d3a8e07e"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"oov = check_coverage(vocab,embeddings_index)","metadata":{"trusted":true,"_uuid":"42c2b82740ac82f5678a46b7c82b5525616c1304"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Ouch only 24% of our vocabulary will have embeddings, making 21% of our data more or less useless. So lets have a look and start improving. For this we can easily have a look at the top oov words.","metadata":{"_uuid":"9ff8e507e571e671ed4249a95fb0eb339098446e"}},{"cell_type":"code","source":"oov[:10]","metadata":{"trusted":true,"_uuid":"de0a78175815413d204d9b7f70b5facba97a7b58"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"On first place there is \"to\". Why? Simply because \"to\" was removed when the GoogleNews Embeddings were trained. We will fix this later, for now we take care about the splitting of punctuation as this also seems to be a Problem. But what do we do with the punctuation then - Do we want to delete or consider as a token? I would say: It depends. If the token has an embedding, keep it, if it doesn't we don't need it anymore. So lets check:","metadata":{"_uuid":"dba213d3a9fefce81a3ab41e2ba95846f87da997"}},{"cell_type":"code","source":"'?' in embeddings_index","metadata":{"trusted":true,"_uuid":"69129cad1f08fda5912051d273db012f9adb081a"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"'&' in embeddings_index","metadata":{"trusted":true,"_uuid":"63deed2445710aef37e504a4d252b5ef02598f21"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Interesting. While \"&\" is in the Google News Embeddings, \"?\" is not. So we basically define a function that splits off \"&\" and removes other punctuation.","metadata":{"_uuid":"361cbf32713d89b11d62c444cc87968460c454b8"}},{"cell_type":"code","source":"def clean_text(x):\n\n    x = str(x)\n    for punct in \"/-'\":\n        x = x.replace(punct, ' ')\n    for punct in '&':\n        x = x.replace(punct, f' {punct} ')\n    for punct in '?!.,\"#$%\\'()*+-/:;<=>@[\\\\]^_`{|}~' + '“”’':\n        x = x.replace(punct, '')\n    return x","metadata":{"trusted":true,"_uuid":"afd570d1160b826a84c4c7af76950fd7a30e4471"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train[\"question_text\"] = train[\"question_text\"].progress_apply(lambda x: clean_text(x))\nsentences = train[\"question_text\"].apply(lambda x: x.split())\nvocab = build_vocab(sentences)","metadata":{"trusted":true,"_uuid":"b47c945e6f02c2a9f72aedd72da4681695cb20dd"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"oov = check_coverage(vocab,embeddings_index)\n","metadata":{"trusted":true,"_uuid":"9f90189a592e93f563d84d68a02ff48cc08315c0"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true,"_uuid":"75365c43419e36635e719b2dd281a96b37a4c94e"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Nice! We were able to increase our embeddings ratio from 24% to 57% by just handling punctiation. Ok lets check on thos oov words.","metadata":{"_uuid":"a06bdc24e539ea6d3fd17afd637c978fcaf36e7f"}},{"cell_type":"code","source":"oov[:10]","metadata":{"trusted":true,"_uuid":"5dff2b2ed016358b734b49e4d23d22cce746f43a"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Hmm seems like numbers also are a problem. Lets check the top 10 embeddings to get a clue.","metadata":{"_uuid":"55f7d6d8960ec03f0cf73e7e97d317065bf940dd"}},{"cell_type":"code","source":"for i in range(10):\n    print(embeddings_index.index2entity[i])","metadata":{"trusted":true,"_uuid":"385ba7b4747ca28e1fea5d25c6a6edbeb9532cc6"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"hmm why is \"##\" in there? Simply because as a reprocessing all numbers bigger tha 9 have been replaced by hashs. I.e. 15 becomes ## while 123 becomes ### or 15.80€ becomes ##.##€. So lets mimic this preprocessing step to further improve our embeddings coverage","metadata":{"_uuid":"c2340ec62a0f679ffffe94bfafc323ecf517f4c9"}},{"cell_type":"code","source":"import re\n\ndef clean_numbers(x):\n\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x","metadata":{"trusted":true,"_uuid":"bbc48ed45237a1096e29fac87c52d258b38fa1ee"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train[\"question_text\"] = train[\"question_text\"].progress_apply(lambda x: clean_numbers(x))\nsentences = train[\"question_text\"].progress_apply(lambda x: x.split())\nvocab = build_vocab(sentences)","metadata":{"trusted":true,"_uuid":"48da9d68b5d8911247a17a927649d33877570e0e"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"oov = check_coverage(vocab,embeddings_index)","metadata":{"trusted":true,"_uuid":"fb6a821cfdb49be15023fe4579c063394eaedbb9"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Nice! Another 3% increase. Now as much as with handling the puntuation, but every bit helps. Lets check the oov words","metadata":{"_uuid":"38618cca73324e598a8610c2ac8648e6aa7eb42b"}},{"cell_type":"code","source":"oov[:20]","metadata":{"trusted":true,"_uuid":"14638ae1defa12014ed7b6be63f3a3bbebb9ed3e"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Ok now we  take care of common misspellings when using american/ british vocab and replacing a few \"modern\" words with \"social media\" for this task I use a multi regex script I found some time ago on stack overflow. Additionally we will simply remove the words \"a\",\"to\",\"and\" and \"of\" since those have obviously been downsampled when training the GoogleNews Embeddings. \n","metadata":{"_uuid":"17ca5651a141e21572e4a5abb23a7b368b5337ba"}},{"cell_type":"code","source":"def _get_mispell(mispell_dict):\n    mispell_re = re.compile('(%s)' % '|'.join(mispell_dict.keys()))\n    return mispell_dict, mispell_re\n\n\nmispell_dict = {'colour':'color',\n                'centre':'center',\n                'didnt':'did not',\n                'doesnt':'does not',\n                'isnt':'is not',\n                'shouldnt':'should not',\n                'favourite':'favorite',\n                'travelling':'traveling',\n                'counselling':'counseling',\n                'theatre':'theater',\n                'cancelled':'canceled',\n                'labour':'labor',\n                'organisation':'organization',\n                'wwii':'world war 2',\n                'citicise':'criticize',\n                'instagram': 'social medium',\n                'whatsapp': 'social medium',\n                'snapchat': 'social medium'\n\n                }\nmispellings, mispellings_re = _get_mispell(mispell_dict)\n\ndef replace_typical_misspell(text):\n    def replace(match):\n        return mispellings[match.group(0)]\n\n    return mispellings_re.sub(replace, text)","metadata":{"trusted":true,"_uuid":"f1c3084cd7503c9148b681de897e0d38d7a90741"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train[\"question_text\"] = train[\"question_text\"].progress_apply(lambda x: replace_typical_misspell(x))\nsentences = train[\"question_text\"].progress_apply(lambda x: x.split())\nto_remove = ['a','to','of','and']\nsentences = [[word for word in sentence if not word in to_remove] for sentence in tqdm(sentences)]\nvocab = build_vocab(sentences)","metadata":{"trusted":true,"_uuid":"c887a0a6f7498724a2c377e9e29fac683ea3a59b"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"oov = check_coverage(vocab,embeddings_index)","metadata":{"trusted":true,"_uuid":"63aba43671245f1ebc87540a8bd49a174791c378"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We see that although we improved on the amount of embeddings found for all our text from 89% to 99%. Lets check the oov words again ","metadata":{"_uuid":"a5543990c3d10f9c77d58840a66bbcad11f233b5"}},{"cell_type":"code","source":"oov[:20]","metadata":{"trusted":true,"_uuid":"21f6da9413c717da789f1068387ac83e3bfd85ae"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Looks good. No obvious oov words there we could quickly fix.\nThank you for reading and happy kaggling","metadata":{"_uuid":"fe32acac0c1280d58ef66ac1611f63bbb98f0aec"}},{"cell_type":"code","source":"","metadata":{"trusted":true,"_uuid":"73df67fc23dd4e8d999d6c0cbb3f70cb952747fb"},"outputs":[],"execution_count":null}]}