{"cells":[{"metadata":{"_uuid":"c53a05d41aa883c65db82bc91f956340a66e06b8"},"cell_type":"markdown","source":"# Using Keras OOV tokens\n\nIn this quick kernel I'm going to demonstrate how you can use an OOV token with Keras' tokenizer. If you are using an RNN hopefully this will give you a slight edge in training"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"from keras.preprocessing.text import Tokenizer","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1cc910e6b4fa9e087048fb6092c556f66421ec6d"},"cell_type":"markdown","source":"So let's read in some data and then do what basically every public kernel is doing in the Quora Insincere Questions Classification competition is doing"},{"metadata":{"trusted":true,"_uuid":"de0c61b2b4af390643d056a6728374c3353d9e4c"},"cell_type":"code","source":"df = pd.read_csv('../input/train.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f50400b2c7746142436068af79383944168ff527"},"cell_type":"markdown","source":"Just a note on features - in this competition you see people using `max_features = 95000` a lot. If you're not aware already, `max_features` corresponds to the number of unique words you're interested in. 95000 is far less than the total number of unique words in the training set. This is over 200,000 words depending on how you do your cleaning.\n\nSo to train the tokenizer you would do this"},{"metadata":{"trusted":true,"_uuid":"7c9bc81eca1850580be5335bfb20d1f3739c0e88"},"cell_type":"code","source":"max_features = 95000\nmaxlen = 60\n\n# I'm just going to limit cleaning to lowering the string and putting spaces around stuff for now, you could do far more I guess\ndef clean_str(x):\n    x = str(x)\n    x = x.lower()\n    \n    specials = [',', '?']\n    for s in specials:\n        x = x.replace(s, f' {s} ')\n        \n    return x\n\n\ndf['question_text'] = df['question_text'].apply(clean_str)\n\ntokenizer = Tokenizer(num_words=None)\ntokenizer.fit_on_texts(list(df['question_text'].values))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0ed77f1b32ad3efc0cee45397c18dc4c5cf194f7"},"cell_type":"markdown","source":"OK, so what does that actually do? Let's tokenize a simple sentence with a misspelt word and a question mark"},{"metadata":{"trusted":true,"_uuid":"8cfc1b81af376732821b1a7f171bd009cfd9b25a"},"cell_type":"code","source":"some_string = \"burger king doesn't sell hamberders, does maccy ds?\"\nsome_string = clean_str(some_string)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"be8c10797f9b38cc8172f43f4c4acd26526215ff"},"cell_type":"code","source":"our_sent = tokenizer.texts_to_sequences([some_string])\nour_sent","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d7d007cf80ae69b8c6fd6d13c30ef212dc42b9ba"},"cell_type":"markdown","source":"Let's now use the tokenizer to return this vector back to a string and see what we get"},{"metadata":{"trusted":true,"_uuid":"d3d258d5ff8b2c0fd11324444d6f3723322a3228"},"cell_type":"code","source":"tokenizer.sequences_to_texts(our_sent)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c73f23a33e75ccd62373a94a2750b8b46ba9d266"},"cell_type":"markdown","source":"There are two really important things here - firstly our misspelt and rare words are just gone. That's really bad, we're trying to judge if a sentence is sincere and part of Quora's critera is that the sentence is gramatically correct - we've just broken that. There is also information in the fact that the word was uncommon enough to not be in the tokenizer.\n\nAnother issue is the tokenizer has stripped `,` and `?`. We might not care so much about `,`s but part of the critera for a sincere question is it is in fact a question, a `?` undoubtably helps us here.\n\n\n## Second attempt - use an OOV token\n\nKeras lets us define an Out Of Vocab token - this will replace any unknown words with a token of our choosing. This is better than just throwing away unknown words since it tells our model there was information here.\n\nLet's do that"},{"metadata":{"trusted":true,"_uuid":"ef5e8d8850b2d7042ffe4b36a1e2fd448da8d96b"},"cell_type":"code","source":"tokenizer_2 = Tokenizer(num_words=max_features, oov_token='OOV')\ntokenizer_2.fit_on_texts(list(df['question_text'].values))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"558e3c9695f5027aee925c5ae430706c6d2351d2"},"cell_type":"code","source":"our_sent_2 = tokenizer_2.texts_to_sequences([some_string])\nour_sent_2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"682f081fe619000d8b13197111aa91b53d460150"},"cell_type":"code","source":"tokenizer_2.sequences_to_texts(our_sent_2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f1b02fd9a2467bb3bdd019e1fc58cff5be1acd98"},"cell_type":"markdown","source":"## Third attempt - use question marks\n\nFinally, let's fix the `?` issue. The `?` is being filtered out by the tokenizer, we can solve this by specifying the filters ourselves"},{"metadata":{"trusted":true,"_uuid":"b0df945b5f80c47890fd676ea94d444e46da1077"},"cell_type":"code","source":"tokenizer_3 = Tokenizer(num_words=max_features, oov_token='OOV', filters='!\"#$%&()*+,-./:;<=>@[\\]^_`{|}~ ')\ntokenizer_3.fit_on_texts(list(df['question_text'].values))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"34789915b8e3d94469810d5ae4f96b59705498b0"},"cell_type":"code","source":"our_sent_3 = tokenizer_3.texts_to_sequences([some_string])\nour_sent_3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"042eefd1e27ee6f5a652e87e067f08b7438e7b68"},"cell_type":"code","source":"tokenizer_3.sequences_to_texts(our_sent_3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5330db66e3513452e27d3701d1b395d5230a611d"},"cell_type":"markdown","source":"This looks much more like something we'd like to train against"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}