{"cells":[{"metadata":{"_uuid":"d3d6451a0f6029825c07e3efadcd77169b31db9e"},"cell_type":"markdown","source":"## A Detailed Guide to understand the Word Embeddings and Embedding Layer in Keras."},{"metadata":{"_uuid":"7acfc8938e6bc84b3fc3f458cb74e7aeb296e12d"},"cell_type":"markdown","source":"## [Don't forget to upvote ;) ]"},{"metadata":{"_uuid":"3017fc7c972715b9236e2e0d77b74e1b36dbe8e9"},"cell_type":"markdown","source":"In this kernel I have explained the keras embedding layer. To do so I have created a sample corpus of just 3 documents and that should be sufficient to explain the working of the keras embedding layer.\n"},{"metadata":{"_uuid":"ed3010328fb2965b8c7665d54bb5690bbf3a9714"},"cell_type":"markdown","source":"Embeddings are useful in a variety of machine learning applications. Because of the fact I have attached many data sources to the kernel where I fell that embeddings and Keras embedding layer may prove to be useful."},{"metadata":{"_uuid":"7a5c0786e978d93aa5de2a4315cb817d66f3da38"},"cell_type":"markdown","source":"Before diving in let us skim through some of the applilcations of the embeddings : \n\n**1 ) The first application that strikes me is in the Collaborative Filtering based Recommender Systems where we have to create the user embeddings and the movie embeddings by decomposing the utility matrix which contains the user-item ratings.**\n\nTo see a complete tutorial on CF based recommender systems using embeddings in Keras you can follow **[this](https://www.kaggle.com/rajmehra03/cf-based-recsys-by-low-rank-matrix-factorization)** kernel of mine.\n\n\n**2 ) The second use is in the Natural Language Processing and its related applications whre we have to create the word embeddings for all the words present in the documents of our corpus.**\n\nThis is the terminology that I shall use in this kernel.\n\n\n**Thus the embedding layer in Keras can be used when we want to create the embeddings to embed higher dimensional data into lower dimensional vector space.**"},{"metadata":{"trusted":true,"_uuid":"c314b8c3775da2307ae38404ce755c298cdf27c3"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d48312c5146ad38bab97d505dd35a147b946de4a"},"cell_type":"markdown","source":"#### IMPORTING MODULES"},{"metadata":{"trusted":true,"_uuid":"cedb06de49a94e693e8de60941b7c57d6997f63f"},"cell_type":"code","source":"# Ignore  the warnings\nimport warnings\nwarnings.filterwarnings('always')\nwarnings.filterwarnings('ignore')\n\n# data visualisation and manipulation\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom matplotlib import style\nimport seaborn as sns\n#configure\n# sets matplotlib to inline and displays graphs below the corressponding cell.\n%matplotlib inline  \nstyle.use('fivethirtyeight')\nsns.set(style='whitegrid',color_codes=True)\n\n#nltk\nimport nltk\n\n#stop-words\nfrom nltk.corpus import stopwords\nstop_words=set(nltk.corpus.stopwords.words('english'))\n\n# tokenizing\nfrom nltk import word_tokenize,sent_tokenize\n\n#keras\nimport keras\nfrom keras.preprocessing.text import one_hot,Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.models import Sequential\nfrom keras.layers import Dense , Flatten ,Embedding,Input\nfrom keras.models import Model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ae1d19871544eb656e1bd44e3471e8a4d1ad9dfe"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"604b9fdec82053d0170e98af3e76594001ee9e43"},"cell_type":"markdown","source":"#### CREATING SAMPLE CORPUS OF DOCUMENTS ie TEXTS"},{"metadata":{"trusted":true,"_uuid":"b789f1783a653c9a6b74541fd27348f5680c6569"},"cell_type":"code","source":"sample_text_1=\"bitty bought a bit of butter\"\nsample_text_2=\"but the bit of butter was a bit bitter\"\nsample_text_3=\"so she bought some better butter to make the bitter butter better\"\n\ncorp=[sample_text_1,sample_text_2,sample_text_3]\nno_docs=len(corp)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2fc3f964bec6dab85e0b9e5874d065f6f5637f1d"},"cell_type":"markdown","source":"#### INTEGER ENCODING ALL THE DOCUMENTS"},{"metadata":{"_uuid":"46bd3ac6a6dfb9e73708302c20cd7091092b596e"},"cell_type":"markdown","source":"After this all the unique words will be reprsented by an integer. For this we are using **one_hot** function from the Keras. Note that the **vocab_size**  is specified large enough so as to ensure **unique integer encoding**  for each and every word.\n\n**Note one important thing that the integer encoding for the word remains same in different docs. eg 'butter' is  denoted by 31 in each and every document.**"},{"metadata":{"trusted":true,"_uuid":"04453f82e7bd7cb51c8a2c7d17637218060e398a"},"cell_type":"code","source":"vocab_size=50 \nencod_corp=[]\nfor i,doc in enumerate(corp):\n    encod_corp.append(one_hot(doc,50))\n    print(\"The encoding for document\",i+1,\" is : \",one_hot(doc,50))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e58da092bf54cc13ea8eedb3b930b5857dc94b40"},"cell_type":"markdown","source":"#### PADDING THE DOCS (to make very doc of same length)"},{"metadata":{"_uuid":"7dcb24072c22e932b98dcb15f17372430cd715f6"},"cell_type":"markdown","source":"**The Keras Embedding layer requires all individual documents to be of same length.**  Hence we wil pad the shorter documents with 0 for now. Therefore now in Keras Embedding layer the **'input_length'**  will be equal to the length  (ie no of words) of the document with maximum length or maximum number of words.\n\nTo pad the shorter documents I am using **pad_sequences** functon from the Keras library."},{"metadata":{"trusted":true,"_uuid":"0a0328b18b81c70f0db2dfbb7200fcde3ed665a5"},"cell_type":"code","source":"# length of maximum document. will be nedded whenever create embeddings for the words\nmaxlen=-1\nfor doc in corp:\n    tokens=nltk.word_tokenize(doc)\n    if(maxlen<len(tokens)):\n        maxlen=len(tokens)\nprint(\"The maximum number of words in any document is : \",maxlen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"22c873edc5421b21eec6419948d0e82d465c6373"},"cell_type":"code","source":"# now to create embeddings all of our docs need to be of same length. hence we can pad the docs with zeros.\npad_corp=pad_sequences(encod_corp,maxlen=maxlen,padding='post',value=0.0)\nprint(\"No of padded documents: \",len(pad_corp))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f3cfbb0c557cd741bde8678dbab207590ea57960"},"cell_type":"code","source":"for i,doc in enumerate(pad_corp):\n     print(\"The padded encoding for document\",i+1,\" is : \",doc)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"95e0fa1b14f37562d504ef5b093e7d3812a26868"},"cell_type":"markdown","source":"#### ACTUALLY CREATING THE EMBEDDINGS using KERAS EMBEDDING LAYER"},{"metadata":{"_uuid":"bb781af84c8e51a23241415749cb2bd1834f9676"},"cell_type":"markdown","source":"Now all the documents are of same length (after padding). And so now we are ready to create and use the embeddings.\n\n**I will embed the words into vectors of 8 dimensions.**"},{"metadata":{"trusted":true,"_uuid":"bdbc0bd1f475e8f6026b24e99fe8b353d226978c"},"cell_type":"code","source":"# specifying the input shape\ninput=Input(shape=(no_docs,maxlen),dtype='float64')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6a2d829400ece99acebacbba69845fd999142f54"},"cell_type":"code","source":"'''\nshape of input. \neach document has 12 element or words which is the value of our maxlen variable.\n\n'''\nword_input=Input(shape=(maxlen,),dtype='float64')  \n\n# creating the embedding\nword_embedding=Embedding(input_dim=vocab_size,output_dim=8,input_length=maxlen)(word_input)\n\nword_vec=Flatten()(word_embedding) # flatten\nembed_model =Model([word_input],word_vec) # combining all into a Keras model","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c71ce6d1c9719ed627989d985168f65129f6d9c6"},"cell_type":"markdown","source":"**PARAMETERS OF THE EMBEDDING LAYER --- **\n\n**'input_dim' = the vocab size that we will choose**. \nIn other words it is the number of unique words in the vocab.\n\n**'output_dim'  = the number of dimensions we wish to embed into**. Each word will be represented by a vector of this much dimensions.\n\n**'input_length' = lenght of the maximum document**. which is stored in maxlen variable in our case."},{"metadata":{"trusted":true,"_uuid":"68bc772020f0fbcefe2a67e9d0422678522c6868"},"cell_type":"code","source":"embed_model.compile(optimizer=keras.optimizers.Adam(lr=1e-3),loss='binary_crossentropy',metrics=['acc']) \n# compiling the model. parameters can be tuned as always.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cb509558221ab7e1df170340ac035af9f029f0f5"},"cell_type":"code","source":"print(type(word_embedding))\nprint(word_embedding)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0215812fb65979237ade681faa50af16ed1f0ac3"},"cell_type":"code","source":"print(embed_model.summary()) # summary of the model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"59db72f4b3c25b426087731f16317e24bf219e76"},"cell_type":"code","source":"embeddings=embed_model.predict(pad_corp) # finally getting the embeddings.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c7d5ce3b7afd01b2afc9ba6bd9beb5269cdb5d68"},"cell_type":"code","source":"print(\"Shape of embeddings : \",embeddings.shape)\nprint(embeddings)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1ce0944cbfab1c714c74ccc06f8eb1a846411ca8"},"cell_type":"code","source":"embeddings=embeddings.reshape(-1,maxlen,8)\nprint(\"Shape of embeddings : \",embeddings.shape) \nprint(embeddings)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7ebe56daafe993cd6b6845231928ec237831707c"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d7cd46a4aaa6d7d1f43a81202add33c5b7f3b528"},"cell_type":"markdown","source":"The resulting shape is (3,12,8).\n\n**3---> no of documents**\n\n**12---> each document is made of 12 words which was our maximum length of any document.**\n\n**& 8---> each word is 8 dimensional.**\n\n "},{"metadata":{"trusted":true,"_uuid":"4312e187fb3604f9ba14f4ad5a46e2a44e7b49a9"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d8862c2ce0926f3fc06b813faa35d06ce72e87b6"},"cell_type":"markdown","source":"#### GETTING ENCODING FOR A PARTICULAR WORD IN A SPECIFIC DOCUMENT"},{"metadata":{"trusted":true,"_uuid":"6ac216640a7db94d9794a88256851d4e1e0c8858"},"cell_type":"code","source":"for i,doc in enumerate(embeddings):\n    for j,word in enumerate(doc):\n        print(\"The encoding for \",j+1,\"th word\",\"in\",i+1,\"th document is : \\n\\n\",word)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"19c665f54917a6aca7f65ea2ff7dbc47b96e90db"},"cell_type":"markdown","source":"#### Now this makes it easier to visualize that we have 3(size of corp) documents with each consisting of 12(maxlen) words and each word mapped to a 8-dimensional vector."},{"metadata":{"trusted":true,"_uuid":"f916e869a05cdf8f1e7e6030f7f7c108f64ddb23"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fdd3d2fb2fc58e889f78dc63ff31a5d55052c686"},"cell_type":"markdown","source":"#### HOW TO WORK WITH A REAL PIECE OF TEXT"},{"metadata":{"_uuid":"5d9d15978ec2de3cf27fd52d08ddfd1b756b1a83"},"cell_type":"markdown","source":"Just like above we can now use any other document. We can sent_tokenize the doc into sentences.\n\nEach sentence has a list of words which we will integer encode using the 'one_hot' function as below. \n\nNow each sentence will be having different number of words. So we will need to pad the sequences to the sentence with maximum words.\n\n**At this point we are ready to feed the input to Keras Embedding layer as shown above.**\n\n**'input_dim' = the vocab size that we will choose**\n\n**'output_dim'  = the number of dimensions we wish to embed into**\n\n**'input_length' = lenght of the maximum document**"},{"metadata":{"trusted":true,"_uuid":"ed5997bd883965d864d75c86ae3bd050dd336efc"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8143b2d5460305133b30d1e8e0e70ee70839747f"},"cell_type":"markdown","source":"## THE END !!!"},{"metadata":{"trusted":true,"_uuid":"49efc1cd691c5964f2efa1b46a757b84233dd176"},"cell_type":"markdown","source":"**If you want to see the application of Keras embedding layer on a real task eg text classification then please check out my [this](https://github.com/mrc03/IMDB-Movie-Review-Sentiment-Analysis) repo on Github in which I have used the embeddings to perform sentiment analysis on IMdb movie review dataset.**"},{"metadata":{"trusted":true,"_uuid":"93078907fa0009405dbbcb046ca7cb65a8aa2af2"},"cell_type":"markdown","source":"## [ Please Do upvote the kernel;) ]"},{"metadata":{"trusted":true,"_uuid":"82790b65e6ccb380e3470ac3cb05b9793681d811"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}