{"cells":[{"metadata":{"_uuid":"01b5e1e20f0d00b2644cec3a98359fe101dd368d"},"cell_type":"markdown","source":"# Introduction\nThe Quora Insincere Question Classification competition allows us to use the four embeddings:  glove.840B.300d (GloVe), paragram_300_sl999 (paragram),  wiki-news-300d-1M (wiki) and GoogleNews-vectors-negative300 (GoogleNews). In a kernel titled: [\"How to: Preprocessing when Using Embeddings\"](https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings), the author raises the issue of tokenization and its effect on how much of the training vocabulary is covered by words in an embedding. The author uses Google news embeddings to illustrate this point. ** In this kernel I expand on this point by exploring the effect of tokenization assumptions on the other three embeddings: GloVe, Paragram, and Wiki News.** \n\nOur base tokenization method simply defines words as sequences of letters, sequences of letters with an apostrophe somewhere in the sequence, or a puctuation mark. To reduce the amount of preprocessing, nothing was removed during the initial tokenization. More preprocessing is added gradually to improve coverage of the training vocabulary."},{"metadata":{"_uuid":"40d303d1d140eb269fee05f9ddb7e29d41e00d73"},"cell_type":"markdown","source":"# Setting Up"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\nimport re\n\n# Use tqdm to show progress of an pandas function we use\ntqdm.pandas()\n\nfrom gensim.models import KeyedVectors as kv\nfrom gensim.scripts.glove2word2vec import glove2word2vec\n\nembedding_path_dict= {'googlenews':{\n                            'path':'../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin',\n                            'format':'word2vec',\n                            'binary': True\n                      },\n                      'glove':{\n                            'path':'../input/embeddings/glove.840B.300d/glove.840B.300d.txt',\n                            'format': 'glove',\n                            'binary': ''\n                      },\n                      'glove_word2vec':{\n                            'path':'../input/glove.840B.300d.txt.word2vec',\n                            'format': 'word2vec',\n                            'binary': False\n                      },\n                      'wiki':{\n                            'path': '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec',\n                            'format': 'word2vec',\n                            'binary': False\n                      },\n                      'paragram':{\n                            'path': '../input/embeddings/paragram_300_sl999/paragram_300_sl999.txt',\n                            'format': '',\n                            'binary': False\n                      }\n                    }\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"de1615767f5c4de575679313d2b7983f032c345a"},"cell_type":"markdown","source":"## Get Training and Test Data"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train=pd.read_csv(\"../input/train.csv\")\ntest= pd.read_csv(\"../input/test.csv\")\nprint(\"Train shape:\", train.shape)\nprint(\"Test shape:\", test.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e31b085ffe907fdf8d072f319baf43b8baf272ca"},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9ee00b7e5db1b212400206b5d73a91930e0bfa84"},"cell_type":"code","source":"train = train.loc[train.question_text.str.len()>100]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"70a090cedfeb97764bfc14dfbf63c03ab40f8a81"},"cell_type":"code","source":"len(train.loc[train['target']==0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b8bcad6f51d6ef07a8c66b43b30222e8a4d5943b"},"cell_type":"code","source":"num_pos= len(train.loc[train['target']==1])\nprint(num_pos)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3962dd3f63820949c27ceb5580df6aa5059f1fa9","scrolled":true},"cell_type":"code","source":"len(train['target'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"782b237b993dc687620e200526fac4328b200e23"},"cell_type":"markdown","source":"## Functions\nHere I define functions that will be used repeatedly in this notebook. More functions will be added as we learn about the embeddings"},{"metadata":{"_uuid":"c22655d390c33bdb1923493609549920a383e649"},"cell_type":"markdown","source":"### Functions: Embedding-Related Functions"},{"metadata":{"trusted":true,"_uuid":"b526d8f9de87a2195727d43df0658d540ba416fb","_kg_hide-input":false},"cell_type":"code","source":"# Get word embeddings\ndef get_embeddings(embedding_path_dict, emb_name):\n    \"\"\"\n    :params embedding_path_dict: a dictionary containing the path, binary flag, and format of the desired embedding,\n            emb_name: the name of the embedding to retrieve\n    :return embedding index: a dictionary containing the embeddings\"\"\"\n    \n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    \n    embeddings_index = {}\n    if (emb_name == 'googlenews'):\n        emb_path = embedding_path_dict[emb_name]['path']\n        bin_flag = embedding_path_dict[emb_name]['binary']\n        embeddings_index = kv.load_word2vec_format(emb_path, binary=bin_flag).vectors\n    elif (emb_name in ['glove', 'wiki']):\n        embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(embedding_path_dict[emb_name]['path']) if len(o)>100)    \n    elif (emb_name == 'paragram'):\n        embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(embedding_path_dict[emb_name]['path'], encoding=\"utf8\", errors='ignore'))\n    return embeddings_index\n\n#Convert GLoVe format into word2vec format\ndef glove_to_word2vec(embedding_path_dict, emb_name='glove', output_emb='glove_word2vec'):\n    \"\"\"\n    Convert the GLOVE embedding format to a word2vec format\n    :params embedding_path_dict: a dictionary containing the path, binary flag, and format of the desired embedding,\n            glove_path: the name of the GLOVE embedding\n            output_file_path: the name of the converted embedding in embedding_path_dict. \n    :return output from the glove2word2vec script\n    \"\"\"\n    glove_input_file = embedding_path_dict[emb_name]['path']\n    word2vec_output_file = embedding_path_dict[output_emb]['path']                \n    return glove2word2vec(glove_input_file, word2vec_output_file)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dc88daa8454e41983bfa890253a7c4ce60e4dc9c"},"cell_type":"code","source":"# Get stats of a given embeddings index\ndef get_emb_stats(embeddings_index):\n\n    # Put all embeddings in a numpy matrix\n    all_embs= np.stack(embeddings_index.values())\n\n    # Get embedding stats\n    emb_mean = all_embs.mean()\n    emb_std = all_embs.std()\n    \n    num_embs = all_embs.shape[0]\n    \n    emb_size = all_embs.shape[1]\n    \n    return emb_mean,emb_std, num_embs, emb_size ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"098bd88715aacc048326d6ed766eb816556d5634"},"cell_type":"markdown","source":"### Functions:  Tokenize Training Sentences"},{"metadata":{"trusted":true,"_uuid":"c279ced2e2ad7e249cb0bc6218fbe6bdb6c101f0"},"cell_type":"code","source":"# Converts sentences into lists of tokens\n# We use this function to allow more control over what constitutes a word\n# It also allows us to explore ways to cover more the pre-defined word embeddings.\n\ndef tokenize(sentences, restrict_to_len=-1):\n    \"\"\"\n    :params sentence_list: list of strings\n    :returns tok_sentences: list of list of tokens\n    \"\"\"\n    \n    if restrict_to_len>0:\n        tok_sentences = [re.findall(r\"[\\w]+[']*[\\w]+|[\\w]+|[.,!?;]\", x ) \\\n                         for x in sentences if len(x)>restrict_to_len]\n    else:\n       tok_sentences = [re.findall(r\"[\\w]+[']*[\\w]+|[\\w]+|[.,!?;]\", x ) \\\n                         for x in sentences] \n    return tok_sentences\n\n#Build the vocabulary given a list of sentence words\ndef get_vocab(sentences, verbose= True):\n    \"\"\"\n    :param sentences: a list of list of words\n    :return: a dictionary of words and their frequency \n    \"\"\"\n    vocab={}\n    for sentence in tqdm(sentences, disable = (not verbose)):\n        for word in sentence:\n            try:\n                vocab[word] +=1\n            except KeyError:\n                vocab[word] = 1\n    return vocab\n\ndef repl(m):\n    return '#' * len(m.group())\n\n#Convert numerals to a # sign\ndef convert_num_to_pound(sentences):\n    return sentences.progress_apply(lambda x: re.sub(\"[1-9][\\d]+\", repl, x)).values\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e590901198dcf3c349f393e02db32d5904fdd893"},"cell_type":"markdown","source":"### Functions: Compare Training and Embedding Vocabulary"},{"metadata":{"trusted":true,"_uuid":"37a513678e45b857beeb11fe4191aa896fda04bb"},"cell_type":"code","source":"\n#find words in common between a given embedding and our vocabulary\ndef compare_vocab_and_embeddings(vocab, embeddings_index):\n    \"\"\"\n    :params vocab: our corpus vocabulary (a dictionary of word frquencies)\n            embeddings_index: a genim object containing loaded embeddings.\n    :returns in_common: words in common,\n             in_common_freq: total frequency in the corpus vocabulary of \n                             all words in common\n             oov: out of vocabulary words\n             oov_frequency: total frequency in vocab of oov words\n    \"\"\"\n    in_common={}\n    oov=[]\n    in_common=[]\n    in_common_freq = 0\n    oov_freq = 0\n    \n    # Compose the vocabulary given the sentence tokens\n    vocab = get_vocab(sentences)\n\n    for word in tqdm(vocab):\n        if word in embeddings_index:\n            in_common.append(word)\n            in_common_freq += vocab[word]\n        else: \n            oov.append(word)\n            oov_freq += vocab[word]\n    \n    print('Found embeddings for {:.2%} of vocab'.format(len(in_common) / len(vocab)))\n    print('Found embeddings for  {:.2%} of all text'.format(in_common_freq / (in_common_freq + oov_freq)))\n\n    return sorted(in_common)[::-1], sorted(oov)[::-1], in_common_freq, oov_freq, vocab\n\n# print the list of out-of-vocabulary words sorted by their frequency in teh training text\ndef show_oov_words(oov, vocab,  num_to_show=15):\n    # Sort oov words by their frequency in the text\n    sorted_oov= sorted(oov, key =lambda x: vocab[x], reverse=True )\n\n    # Show oov words and their frequencies\n    if (len(sorted_oov)>0):\n        print(\"oov words:\")\n        for word in sorted_oov[:num_to_show]:\n            print(\"%s\\t%s\"%(word, vocab[word]))\n    else:\n        print(\"No words were out of vocabulary.\")\n        \n    return len(sorted_oov);\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c3f1201b139bcaedaa9c64bec8ebd0fbb1d6ddfc"},"cell_type":"markdown","source":"# Exploring Embeddings\n\nWe are now ready to explore each embedding and tokenization techniques that maximize its coverage of the training vocabulary"},{"metadata":{"_uuid":"8303bbe4fb2a6052f6390e2bd800641ee5a5cedb"},"cell_type":"markdown","source":"## GloVe"},{"metadata":{"_uuid":"a694d32ea01cf6659d22c921940c5f0d13779487"},"cell_type":"markdown","source":"### Choose Embedding"},{"metadata":{"trusted":true,"_uuid":"e136ce606f51aa2a6b044bde58295445660c870a"},"cell_type":"code","source":"embedding_name = 'glove'\nembeddings_index= get_embeddings(embedding_path_dict, embedding_name)\nimport gc; gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ff8d3f9d989ce99521d66aa38eb8657b39555cb1","scrolled":true},"cell_type":"code","source":"# Get embedding stats\nemb_mean,emb_std, num_embs, emb_size = get_emb_stats(embeddings_index)\nprint(\"mean: %5.5f\\nstd: %5.5f\\nnumber of embeddings: %d\\nembedding vector size:%d\" \\\n      %(emb_mean,emb_std, num_embs, emb_size))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4765ec86d61a0c20849f21fa0729c5cbb0bff242"},"cell_type":"markdown","source":"### Tokenize Training Text"},{"metadata":{"trusted":true,"_uuid":"606a001ed9cadf2606065c469dd92720e506b6ff"},"cell_type":"code","source":"question_text = train[\"question_text\"]\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(question_text)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c46b98952890c4a3ba1a4f46877c148ff06c5374","scrolled":true},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fafa2c60aba0b1ea52a226dd26a87886bf14dfd1"},"cell_type":"markdown","source":"good coverage but the top missing words all have contractions. We deal with those next..."},{"metadata":{"trusted":true,"_uuid":"22c2890a1267ab4819181b4a4680caf0db2f30f5"},"cell_type":"code","source":"contr_dict={\"I\\'m\": \"I am\",\n            \"won\\'t\": \"will not\",\n            \"\\'s\" : \"\", \n            \"\\'ll\":\"will\",\n            \"\\'ve\":\"have\",\n            \"n\\'t\":\"not\",\n            \"\\'re\": \"are\",\n            \"\\'d\": \"would\",\n            \"y'all\": \"all of you\"}\n\ndef replace_contractions(sentences, contr_dict=contr_dict):\n    res_sentences=[]\n    for sent in sentences:\n        for contr in contr_dict:\n            sent = sent.replace(contr, \" \"+contr_dict[contr])\n        res_sentences.append(sent)\n    return res_sentences","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"770d0f5da9adf608eb1efe2387ab6b772f17c491"},"cell_type":"code","source":"# start by replacing contractions\nsentences = replace_contractions(question_text)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3b79b0aed327add1a838ace97b2c13a1b7282a35"},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e2c069582719871b9eaa778d8b051e27da39cecd"},"cell_type":"markdown","source":"Much better... Quorans is the top word missed now...I wonder if Quora or quora is in the embeddings_index vocabulary.."},{"metadata":{"trusted":true,"_uuid":"51f29cb0f25df16c22d5d3dddff49e46f9e8a0bb"},"cell_type":"code","source":"print(\"Is 'Quora' in the wiki embeddings index?\",'Quora' in embeddings_index)\nprint(\"Is 'quora' in the wiki embeddings index?\",'quora' in embeddings_index)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cc5ccff7d801a6f57ee404d3319f6ddace0b11c4"},"cell_type":"markdown","source":"We can replace Quorans with Quora contributors..."},{"metadata":{"trusted":true,"_uuid":"1fdc2a46aa6bb876a394cb912e013ffd6d1d1e33"},"cell_type":"code","source":"w_quoran_contr_dict={\"I\\'m\": \"I am\",\n                    \"won\\'t\": \"will not\",\n                    \"\\'s\" : \"\", \n                    \"\\'ll\":\"will\",\n                    \"\\'ve\":\"have\",\n                    \"n\\'t\":\"not\",\n                    \"\\'re\": \"are\",\n                    \"\\'d\": \"would\",\n                    \"y'all\": \"all of you\",\n                    \"Quoran\": \"Quora contributor\",\n                    \"quoran\": \"quora contributor\"\n                    }","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c00b115014fefe03ccc525fcabc607ee19090e95"},"cell_type":"code","source":"# replace contractions using a contr dict containing replacement for Quoran\nsentences = replace_contractions(question_text, contr_dict = w_quoran_contr_dict)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3b79b0aed327add1a838ace97b2c13a1b7282a35","scrolled":true},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"734dc0a78cbe6c951fceede879d80b42c72a0129"},"cell_type":"markdown","source":"No change. Probably because it is only a single word whose frequency is small compared to the size of the vocabulary.\n\nwhat about heights? There are some tokens that mention height such as 5'2 and 6'4.. can we convert those to a longer, more compatible, format? \n\n(Note: Different height values  show up frequently further down the list. To see it in your own notebook use show_oov_words(oov, vocab, num_to_show=100). I am only showing a small list of oov words here to make the notebook more readable) "},{"metadata":{"_uuid":"eb7637f7ba9f28723b1c97992ac5aba3060dc115"},"cell_type":"markdown","source":"First, does the embeddings index contain digits? (Google News replaces numbers > 9 with # signs)"},{"metadata":{"trusted":true,"_uuid":"693f4a56dd24225abe71b00a1453b3e1cd65b7a4"},"cell_type":"code","source":"print(\"0 in embedding index?\", ('0' in embeddings_index))\nprint(\"Other digits?\", ('1' in embeddings_index) and ('2' in embeddings_index))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c79eca14cf09e462e7c01a6e2b968b1aa0bbf9e9"},"cell_type":"markdown","source":"Good...So let us replace all heights of the form \\d\\'\\d such as 5'4 with the \"5 foot 4\".."},{"metadata":{"trusted":true,"_uuid":"dae241d4902065657f9452a4076f47d5e70c4465"},"cell_type":"code","source":"import re\n\ndef convert_height(sentences):\n    res_sentences = []\n    for sent in sentences:\n        res_sent = re.sub( \"(\\d+)\\'(\\d+)\", \"\\1 foot \\2\", sent)\n        res_sentences.append(res_sent)\n    return res_sentences","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dc00dec8667b3fe39ebfde4a421d3d44d579e244"},"cell_type":"code","source":"# start by converting heights such as 5'4 to longer format 5 foot 4\nsentences = convert_height(question_text)\n\n# replace contractions\nsentences = replace_contractions(sentences)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3b79b0aed327add1a838ace97b2c13a1b7282a35","scrolled":true},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2541b0b512cfce25e9da75b68c20e0487b712788"},"cell_type":"markdown","source":"Very slight improvement. But overall, we managed to improve our coverage from 83.54% to 86.43%. For The GloVe embedding the main issue affecting our results was contractions. Replacing height with a longer format helped very slightly."},{"metadata":{"_uuid":"e05fbc7651737d5de5b0829258a8f244ff99c752"},"cell_type":"markdown","source":"## Paragram"},{"metadata":{"_uuid":"44575fad64e28e585392e6f3f77cc85010fe1d70"},"cell_type":"markdown","source":"### Choose Embedding"},{"metadata":{"trusted":true,"_uuid":"e136ce606f51aa2a6b044bde58295445660c870a"},"cell_type":"code","source":"embedding_name = 'paragram'\nembeddings_index= get_embeddings(embedding_path_dict, embedding_name)\nimport gc; gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ff8d3f9d989ce99521d66aa38eb8657b39555cb1","scrolled":true},"cell_type":"code","source":"# Get embedding stats\nemb_mean,emb_std, num_embs, emb_size = get_emb_stats(embeddings_index)\nprint(\"mean: %5.5f\\nstd: %5.5f\\nnumber of embeddings: %d\\nembedding vector size:%d\" \\\n      %(emb_mean,emb_std, num_embs, emb_size))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4765ec86d61a0c20849f21fa0729c5cbb0bff242"},"cell_type":"markdown","source":"### Tokenize Training Text"},{"metadata":{"trusted":true,"_uuid":"606a001ed9cadf2606065c469dd92720e506b6ff"},"cell_type":"code","source":"question_text = train[\"question_text\"]\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(question_text)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c46b98952890c4a3ba1a4f46877c148ff06c5374"},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b1564d75c34a746c3b5c7101d601f2755957115a"},"cell_type":"markdown","source":"Interesting! Very few vocabulary words are in common with the paragrams vocabulary. The top missing ones all have capital letters so let us convert to lower case"},{"metadata":{"trusted":true,"_uuid":"2dfb0843e920baaa72b02b8f4522a57301a1f2d1"},"cell_type":"code","source":"def convert_to_lower(sentences):\n    res_sentences = []\n    for sent in sentences:\n        lower_sent = sent.lower()\n        res_sentences.append(lower_sent)\n    return res_sentences","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"770d0f5da9adf608eb1efe2387ab6b772f17c491"},"cell_type":"code","source":"# convert capitals to lowercase\nsentences = convert_to_lower(question_text)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3b79b0aed327add1a838ace97b2c13a1b7282a35"},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3033bf25d5123dd6658a620f7c78a654e6c8fa9f"},"cell_type":"markdown","source":"Much better!  Lets deal with the contractions now..."},{"metadata":{"trusted":true,"_uuid":"3c64f4b8c5abd06f08b6737f75e20b955839eecc"},"cell_type":"code","source":"# start by converting capitals to lowercase\nsentences = convert_to_lower(question_text)\n\n# replace contractions\nsentences = replace_contractions(sentences)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0874926eb33ec8cf087eeb098401d45d068400a3"},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"04ba759a7e12cb6b550989a568b55b9d3082c269"},"cell_type":"markdown","source":"There are also frequent mentions of height..I wonder how that will affect the results..."},{"metadata":{"trusted":true,"_uuid":"dc00dec8667b3fe39ebfde4a421d3d44d579e244"},"cell_type":"code","source":"# start by replacing heights such as 5'4 to a longer format (5 foot 4)\nsentences = convert_height(question_text)\n\n# convert capitals to lowercase\nsentences = convert_to_lower(sentences)\n\n# replace contractions\nsentences = replace_contractions(sentences)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3b79b0aed327add1a838ace97b2c13a1b7282a35","scrolled":true},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2f3e9c73135a3b50ec2459b72d4994376d23b610"},"cell_type":"markdown","source":"Slightly better. \n\nWe managed to improve the vocabulary coverage from a mere 50.03% to 85.97% by replacing capitals with lower case letters and expanding contractions. Replacing heights with a longer form improved results very slightly."},{"metadata":{"_uuid":"c0f61001a6eeae6f0ba8058fb4d465ff3f0b8e03"},"cell_type":"markdown","source":"## Wiki"},{"metadata":{"_uuid":"a6bc05357d230c26640d9247734ff80700a3c97e"},"cell_type":"markdown","source":"### Choose Embedding"},{"metadata":{"trusted":true,"_uuid":"e136ce606f51aa2a6b044bde58295445660c870a"},"cell_type":"code","source":"embedding_name = 'wiki'\nembeddings_index= get_embeddings(embedding_path_dict, embedding_name)\nimport gc; gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ff8d3f9d989ce99521d66aa38eb8657b39555cb1","scrolled":true},"cell_type":"code","source":"# Get embedding stats\nemb_mean,emb_std, num_embs, emb_size = get_emb_stats(embeddings_index)\nprint(\"mean: %5.5f\\nstd: %5.5f\\nnumber of embeddings: %d\\nembedding vector size:%d\" \\\n      %(emb_mean,emb_std, num_embs, emb_size))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4765ec86d61a0c20849f21fa0729c5cbb0bff242"},"cell_type":"markdown","source":"### Tokenize Training Text"},{"metadata":{"trusted":true,"_uuid":"606a001ed9cadf2606065c469dd92720e506b6ff"},"cell_type":"code","source":"question_text = train[\"question_text\"]\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(question_text)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c46b98952890c4a3ba1a4f46877c148ff06c5374"},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c2e3018089b63be6c4dd8a97411e9dbdd6ba502b"},"cell_type":"markdown","source":"Contractions seem to be the main issue here.."},{"metadata":{"trusted":true,"_uuid":"54e472351009f2817e1022e0ddc8c8a6c6efab62"},"cell_type":"code","source":"# start by replacing contractions\nsentences = replace_contractions(question_text)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3b79b0aed327add1a838ace97b2c13a1b7282a35","scrolled":true},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"006d0fb446a77e181e4d5059a05477d0a8929f71"},"cell_type":"markdown","source":"Better! Quorans seems to be a repeatedly missed word in this embedding as well... "},{"metadata":{"trusted":true,"_uuid":"51f29cb0f25df16c22d5d3dddff49e46f9e8a0bb"},"cell_type":"code","source":"print(\"Is 'Quora' in the wiki embeddings index?\",'Quora' in embeddings_index)\nprint(\"Is 'quora' in the wiki embeddings index?\",'quora' in embeddings_index)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cc5ccff7d801a6f57ee404d3319f6ddace0b11c4"},"cell_type":"markdown","source":"We can replace Quorans with Quora contributors..."},{"metadata":{"trusted":true,"_uuid":"c00b115014fefe03ccc525fcabc607ee19090e95"},"cell_type":"code","source":"# start by replacing contractions using the contractions dict containing replacements for Quoran\nsentences = replace_contractions(question_text, contr_dict = w_quoran_contr_dict)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3b79b0aed327add1a838ace97b2c13a1b7282a35","scrolled":true},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c3eac3d64f77ff1ca6088709ff9c88426b2c4a5b"},"cell_type":"markdown","source":"Fixed that but with no effect on overall result since the qord frequency is small compared to the total vocabulary size. \n\nLets now look at the effect of switching to heights.. First, does this embedding contain numbers as tokens?"},{"metadata":{"trusted":true,"_uuid":"7e6072146f96efaa83af2ae7797cad11270666fa"},"cell_type":"code","source":"print(\"0 in embedding index?\", ('0' in embeddings_index))\nprint(\"Other digits?\", ('1' in embeddings_index) and ('2' in embeddings_index))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7f6596d8e3c58492a963cb6d388499cb05f1bf89"},"cell_type":"markdown","source":"Lets convert heights to a longer format..."},{"metadata":{"trusted":true,"_uuid":"dc00dec8667b3fe39ebfde4a421d3d44d579e244"},"cell_type":"code","source":"# start by converting height to longer form\nsentences = convert_height(question_text)\n\n# replace contractions\nsentences = replace_contractions(sentences,  contr_dict = w_quoran_contr_dict)\n\n# Get a list of token for each question text\n# restrict_to_len is approximately the mean sentence length+ 0.5std\nsentences = tokenize(sentences)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3b79b0aed327add1a838ace97b2c13a1b7282a35","scrolled":true},"cell_type":"code","source":"# Does our tokenization method produce a good match with \n# the words in the selected embedding type?\n\n# Get words in common and out of vocabulary words\nin_common, oov, in_common_freq, oov_freq, vocab = compare_vocab_and_embeddings(sentences, embeddings_index)\n\n# Print a sorted list of the oov words\nshow_oov_words(oov, vocab)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"62c76d888dd94f4d33c19ef610ae931879a7ba6a"},"cell_type":"markdown","source":"Slightly better!\n\nOverall coverage for the Wiki embeddings improved from 79.36% to 82.09%. The main issue in this embedding was dealing with contractions. As with the other two embeddings, replacing height with a longer format had a minor effect on the overall result."},{"metadata":{"_uuid":"88bec87c2285cfd92d9397f8fe59228613351d84"},"cell_type":"markdown","source":"## Conclusion"},{"metadata":{"_uuid":"63a4ae0e2afdfdb34e9221f472e3f733f381fdf2"},"cell_type":"markdown","source":"In this notebook We looked at three different embeddings: glove.840B.300d (GloVe), paragram_300_sl999 (paragram), and wiki-news-300d-1M (wiki) and how to best tokenize our Quora training text so that we maximize the percentage of words represented by the embeddings index.  In general, capetalization, contractions, and, to a lesser extent, height measurements, had the most impact on how much of the training vocabulary was covered by an embedding.\n\nOur base tokenization method simply defined words as sequences of letters, sequences of letters with an apostrophe somewhere in the sequence, or a puctuation mark. To reduce the amount of preprocessing nothing was removed in this method. More preprocessing was done based on our observations of which words were missed by the embedding. Coverage of an embedding was improved by about 3 percentage point for the GloVe and Wiki embeddings. The most significant improvement was for the paragram embedding which improved  36 percentage points from 50.03% to 85.97%.\n\nAwaerness of the intricacies of each embedding should help improve the accuracy of our Quora Insincere Questions Classification networks. "},{"metadata":{"_uuid":"30707cdde37c4d70ad34b0da0a1ed51ba7ac8e43"},"cell_type":"markdown","source":"### Acknowledgments"},{"metadata":{"_uuid":"b6cca48b95a145d1b7cddfad70705afd7668186c"},"cell_type":"markdown","source":"* [https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings](http://https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings)\n* [https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings](http://https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings)"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}