{"cells":[{"metadata":{"_uuid":"85034373c4deeef19f1b62eebffb4c2ff5f47395"},"cell_type":"markdown","source":"In this kernel, i will help you speed up your preprocessing at:\n\n1. replace text in `question_text`:\n\n    replace the text in the dict, e.g.: x = x.replace(\"?\", \" ? \")\n\n2. load embeddings"},{"metadata":{"_uuid":"89353d08cf96d5ec36275179bf441087d6f2dac1"},"cell_type":"markdown","source":"There is a quick report about the kernel. If you want to see the details, follow the codes :)"},{"metadata":{"_uuid":"8d192be1f75cea876ef814acfd71bac60d6ad65c"},"cell_type":"markdown","source":"| No | Category                 | Type   | DictLength   | Run time (total)  | Run time(s/w) |\n|-----|-----------------------|--------|----------------|---------------------| ----------------- |\n|1      |     replace text         |  slow  |          130      |                  43.4 s  |     0.3338          |\n|2     |     replace text         |   fast   |          130      |                    8.8 s  |     0.0677          |\n|3     |     replace text         |   slow  |          65       |                  23.9 s  |     0.3677          |\n|4     |     replace text         |   fast   |          65       |                  6.05 s  |     0.0931          |\n|5     |     load embedding |   slow  |          ---       |                   51.6 s  |             ---         |\n|6     |     load embedding |   fast   |          ---       |                    19.1 s  |             ---         |\n"},{"metadata":{"_uuid":"acde5a2e25bcf6bb1ba732d77791240dba3b0aab"},"cell_type":"markdown","source":"First, let's import the packages and load the datas."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c7540604ced9e7f08c50ada5f021b0e28e2abfd3"},"cell_type":"code","source":"!ls ../input/","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cb49294cfa3de44551454eb0b57aeeaf9ca75a25"},"cell_type":"code","source":"train_org = pd.read_csv(\"../input/train.csv\")\nprint(\"train shape:\", train_org.shape)\ntrain_org.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"markdown","source":"## 1. replace text in question_text"},{"metadata":{"_uuid":"c01f583744b7a2ce09749cb835c0c1a80f75b6c6"},"cell_type":"markdown","source":"In this section, i will use `clean_text_fast` to speed up. The original replace function is `clean_text_slow`.\n\nThe two functions are defined as blow:"},{"metadata":{"trusted":true,"_uuid":"65d30a1fbd15cf0bc5494d152b627ecb36c00a7b"},"cell_type":"code","source":"def clean_text_slow(x, maxlen=None):\n    puncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n    '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', \n    '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', \n    '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', \n    '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]\n    x = x.lower()\n    for punct in puncts[:maxlen]:\n        x = x.replace(punct, f' {punct} ')\n    return x\n\ndef clean_text_fast(x, maxlen=None):\n    puncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n    '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', \n    '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', \n    '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', \n    '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]\n    x = x.lower()\n    for punct in puncts[:maxlen]:\n        if punct in x:  # add this line\n            x = x.replace(punct, f' {punct} ')\n    return x","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e25ccafccae8c98ba3ae2743acfc81d647865e8d"},"cell_type":"markdown","source":"the `puncts` contains 130 words or characters\n\nthe `fast` function only add one line: `if punct in x:`\n\nLet's look at the run time first."},{"metadata":{"trusted":true,"_uuid":"d1f5997c5ad66aa9ae5bbc7e03883c34a04c9c41"},"cell_type":"code","source":"%%time\n_ = train_org.question_text.apply(lambda x: clean_text_slow(x, maxlen=None))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c3e461f078a06acba8aeae2a91e83c97bfe5fe60"},"cell_type":"code","source":"%%time\n_ = train_org.question_text.apply(lambda x: clean_text_fast(x, maxlen=None))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c83d2bf1d3522a3bd571513e2c251a22466d6da4"},"cell_type":"markdown","source":"This is because the `in` operation is more fast than `create a new str object` in python.\n\nIn the `slow` function, we create a new str in every iteration.\n\nNext, let's use a half `puncts` length to calculate the run time."},{"metadata":{"trusted":true,"_uuid":"944ff967fb1a527a8a824932f0ccf7832d6b30f1"},"cell_type":"code","source":"%%time\n_ = train_org.question_text.apply(lambda x: clean_text_slow(x, maxlen=65))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6b462b8e0e90ab3f179d3c71b103ad9c1d93aa64"},"cell_type":"code","source":"%%time\n_ = train_org.question_text.apply(lambda x: clean_text_fast(x, maxlen=65))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7ac6b57b397fa27c3e4711f4dd5b9f4cbdf12d94"},"cell_type":"markdown","source":"As we can see. The `slow` function runs **double** while the `puncts` become **double**\n\nThe `fast` only use extra **1/3** seconds while the `puncts` become  **double**."},{"metadata":{"_uuid":"8dcd963bd26a9dd59fd097ecc0e579e9f3f3ae3b"},"cell_type":"markdown","source":"As our `puncts` grows longer and longer... I think i dont need to say anymore."},{"metadata":{"_uuid":"db4144accf9bb203b5d902a64a880eb2ad3c2908"},"cell_type":"markdown","source":"> **do not create a `new string object` if you can use `in operation`** in python."},{"metadata":{"trusted":true,"_uuid":"513aa25dc132eb60db22b0ff27fb44c9efb143be"},"cell_type":"markdown","source":"## 2. load embeddings."},{"metadata":{"trusted":true,"_uuid":"20a826513165d5233c4d3a1f023ed5adacc27548"},"cell_type":"markdown","source":"use hard code to speed up."},{"metadata":{"trusted":true,"_uuid":"ebaa1bfefbf2fecb64bf6c2e8e44568138889d7e"},"cell_type":"code","source":"def load_glove_slow(word_index, max_words=200000, embed_size=300):\n    EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE) if o.split(\" \")[0] in word_index)\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (max_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_words: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n            \n    return embedding_matrix \n\ndef load_glove_fast(word_index, max_words=200000, embed_size=300):\n    EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n    emb_mean, emb_std = -0.005838499, 0.48782197\n\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (max_words, embed_size))\n    with open(EMBEDDING_FILE, 'r', encoding=\"utf8\") as f:\n        for line in f:\n            word, vec = line.split(' ', 1)\n            if word not in word_index:\n                continue\n            i = word_index[word]\n            if i >= max_words:\n                continue\n            embedding_vector = np.asarray(vec.split(' '), dtype='float32')[:300]\n            if len(embedding_vector) == 300:\n                embedding_matrix[i] = embedding_vector\n    return embedding_matrix","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6bc0b9d558933f7c00897f3450790b83398a333a"},"cell_type":"markdown","source":"In the `load_glove_slow`, we calculate the `emb_mean` and  `emb_std` on every loading.\n\nIn the `load_glove_fast`, we write the `emb_mean` and  `emb_std` in the code, and avoid create `dict`"},{"metadata":{"_uuid":"c3a8db129c0bf62324082c736dfd5d41f33339f2"},"cell_type":"markdown","source":"Let create the word_index:"},{"metadata":{"trusted":true,"_uuid":"23e5d6d60984f31e2280e3d2ebbf5a6bb70e8c16"},"cell_type":"code","source":"from keras.preprocessing.text import Tokenizer","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5c3c4f672056bb823f3577a59c7d953ead957a1c"},"cell_type":"code","source":"tokenizer = Tokenizer()\ntokenizer.fit_on_texts(train_org.question_text.values)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2162924e005b7f04d60f03ad4197afba0afa9a78"},"cell_type":"markdown","source":"run the functions:"},{"metadata":{"trusted":true,"_uuid":"f465aa4b74817f2298e4ab17c453926e100286ce"},"cell_type":"code","source":"%%time\n_ = load_glove_slow(tokenizer.word_index, len(tokenizer.word_index) + 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"70aaa87c84968a71f497ff09178fb17d18a1821c"},"cell_type":"code","source":"%%time\n_ = load_glove_fast(tokenizer.word_index, len(tokenizer.word_index) + 1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"49e5e7dc62421c7c14fffcffccd7664fcb7b4f18"},"cell_type":"markdown","source":"the `fast` function use less **32.5** seconds the `slow`\n\n**And in my code, the slow code runs totally 5mins while the `fast` runs only 40~50 seconds.**"},{"metadata":{"_uuid":"e25169e29969b1966b2b31b99c4152e968afca3d"},"cell_type":"markdown","source":"> Try use hard code if it is possible**?**"},{"metadata":{"trusted":true,"_uuid":"4bc9025203fd638171ad800d64c1c4e1500c3144"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}