{"cells":[{"metadata":{"_uuid":"0adf4f008233f87f62901b09885e68d45cacb6ce"},"cell_type":"markdown","source":"# **Word embedding with Python**\n**word2vec, doc2vec, GloVe implementation with Python**\n\n---\n![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcS7lrYJSLlPvn3Hoeo24Y2NAze3ZLMsRdxibZR1MsMCiHkwHXAS)\n![](https://s3-ap-south-1.amazonaws.com/av-blog-media/wp-content/uploads/2017/06/06062705/Word-Vectors.png)\n\n---\n\n### **Table of Contents**\n\n---\n* [**1.What are Word Embeddings?**](#1.What-are-Word-Embeddings?)\n* [**2.Different types of Word Embedding**](#2.-Different-types-of-Word-Embedding)\n    * [**2.1.Frequency based Embedding**](#2.1.Frequency-based-Embedding)  \n        * [**2.1.1.Count Vectors**](#2.1.1.Count-Vectors)  \n        * [**2.1.2.TF-IDF**](#2.1.2.TF-IDF)  \n        * [**2.1.3.Co-Occurrence Matrix**](#2.1.3.Co-Occurrence-Matrix)  \n    * [**2.2.Prediction based Embedding**](#2.2.Prediction-based-Embedding)  \n        * [**2.2.1.CBOW**](#2.2.1.CBOW)  \n        * [**2.2.2.Skip-Gram**](#2.2.2.Skip-Gram)  \n* [**3.Using pre-trained Word Vectors**](#3.Using-pre-trained-Word-Vectors)\n* [**4.Training your own Word Vectors**](#5.Training-your-own-Word-Vectors)\n\n---\n\n# ***1.What are Word Embeddings?***\n\n---\n# **Defination**\n\n> ## **Word embeddings are a type of word representation that allows words with similar meaning to have a similar representation.** ***...By Jason Brownlee.***  \n---\n### **Example 1**\n![](https://i.stack.imgur.com/oJEie.png)\n\n### **Example 2**\n![](https://cdn-images-1.medium.com/max/1600/1*YEJf9BQQh0ma1ECs6x_7yQ.png)\n\n* A very basic definition of a word embedding is a real number, vector representation of a word. Typically, these days, words with similar meaning will have vector representations that are close together in the embedding space (though this hasn’t always been the case).\n\n* ***Word embedding is a dense representation of words in the form of numeric vectors. It can be learned using a variety of language models. The word embedding representation is able to reveal many hidden relationships between words. For example, vector(“cat”) - vector(“kitten”) is similar to vector(“dog”) - vector(“puppy”). This post introduces several models for learning word embedding and how their loss functions are designed for the purpose.***\n\n"},{"metadata":{"_uuid":"847128166ca8b7a314c25c71b3420ae7aead4cf7"},"cell_type":"markdown","source":"---\n\n# ***2.Different types of Word Embedding***\n\n---\n\n### The different types of word embeddings can be broadly classified into two categories\n\n1. **Frequency based Embedding**\n1. **Prediction based Embedding**\n\n![](https://multithreaded.stitchfix.com/assets/posts/2017-10-18-stop-using-word2vec/fig_006.png)\n\n---\n## ***2.1.Frequency based Embedding***\n\n### **2.1.1.Count Vectors**\n\n---\n\n* Extract the corpus C {d1, d2 ... dD} of the document D and the N unique tokens (words) from the corpus C. N unque form our dictionary and the size of the count vector matrix M by DX N. D (i) is the number of times each row of the matrix contains M tokens in the document.\n\n#### Let us understand this with a simple example.\n\n* **D1: He is a lazy boy. She is also lazy.**\n* **D2: Neeraj is a lazy person.**\n\nThe dictionary created can be a word with a **unique tag in the corpus**: ***['He', 'She', 'lazy', 'boy', 'Neeraj', 'person']***  \n* Here, **D = 2, N = 6**, The count matrix M of size 2 X 6 will be represented as –\n\n||He|She|lazy|boy|Neeraj|person|\n|--|--|--|--|--|--|--|\n|D1|1|1|2|1|0|0|\n|D2|0|0|1|0|1|1|\n\n\n### **Practical Example**\n\n### **1. Count Vectorization**\n"},{"metadata":{"trusted":true,"_uuid":"4c0526b2315e40a8df8fb0898040662b64c82230"},"cell_type":"code","source":"from sklearn.feature_extraction.text import CountVectorizer\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.metrics.pairwise import cosine_similarity\nimport pandas as pd","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3c3944ff129083d6d4fd8a6a60aab923f36c419d"},"cell_type":"code","source":"text = ['The quick brown fox jumped over the lazy dog']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1b889013c041b2108d2bbc7298afe7612dcfc62d"},"cell_type":"code","source":"vectorizer = CountVectorizer()\nvectorizer.fit(text)\nprint(vectorizer.vocabulary_)\n#encode the document\nvector = vectorizer.transform(text)\nprint(vector.shape)\nprint(type(vector))\nprint(vector.toarray())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6a4aa57e48231eb67526875705243ffef9cbdae4"},"cell_type":"code","source":"vector = vectorizer.transform(text)\nprint(vector.shape)\nprint(type(vector))\nprint(vector.toarray())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5e5c2bd216a04ed8c927c945fc416a9f25a44cba"},"cell_type":"code","source":"vector.toarray()\ndf = pd.DataFrame(vector.todense())\ndf.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"675d4547991c985548fa82507e9c698c10845908"},"cell_type":"markdown","source":"## **2.1.2.TF-IDF**\n\n---\n\n\n![](https://lh4.googleusercontent.com/zSHLtG-IVPQG_raN67XTAuzIlKpJs8dkZtFP3VhN7W8Ur4keIzgRt8_w1eyqQ8lyX1flOqyf4xhrOoXUzLRHfzgCQhurjouJFyQaPMahHb8Ar5TH5L96T8QTQGKF7C90wvYhjvPOshzbyK4zSA)\n\n### **Formula**\n![](https://cdn-images-1.medium.com/max/1600/1*jNnpbGPxkjehlvTCXq9B8g.png)\n1.  **TF Score (Term Frequency)** :\nConsiders documents as bag of words, agnostic to order of words. A document with 10 occurrences of the term is more relevant than a document with term frequency 1. But it is not 10 times more relevant, relevance is not proportional to frequency\n\n2.  **IDF Score (Inverse Document Frequency)**\nWe also want to use the frequency of the term in the collection for weighting and ranking. Rare terms are more informative than frequent terms. We want low positive weights for frequent terms and high weights for rare terms.\n\n### **Mathematical Example**\n\n![image.png](https://s3-ap-south-1.amazonaws.com/av-blog-media/wp-content/uploads/2017/06/04171138/Tf-IDF.png)\n\n \n### ***Term Frequency***\n\n---\nTF(***Term Frequency***) = (**Number of times term t appears in a document)/(Number of terms in the document)**\n\n* **TF(This,Document1)** = 1/8\n* **TF(This, Document2)**=1/5\n\nIt **denotes the contribution of the word to the document i.e words relevant to the document should be frequent.** eg: **A document about Messi should contain the word ‘Messi’ in large number.**\n\n### ***Inverse Document Frequency***\n\n---\nIDF(***Inverse Document Frequency***) = **log(N/n)**, where, **N is the number of documents and n is the number of documents a term t has appeared in.**\n\n* where **N is the number of documents** and **n is the number of documents a term t has appeared in.**\n\n* **IDF(This) = log(2/2) = 0.**\n\n* So, how do we explain the reasoning behind IDF? Ideally, if a word has appeared in all the document, then probably that word is not relevant to a particular document. But if it has appeared in a subset of documents then probably the word is of some relevance to the documents it is present in.\n\nLet us compute IDF for the word ‘Messi’.\n\n* **IDF(Messi)** = log(2/1) = 0.301.\n\nNow, let us compare the TF-IDF for a common word ‘This’ and a word ‘Messi’ which seems to be of relevance to Document 1.\n\n* TF-IDF(This,Document1) = (1/8) * (0) = 0\n* TF-IDF(This, Document2) = (1/5) * (0) = 0\n* TF-IDF(Messi, Document1) = (4/8)*0.301 = 0.15"},{"metadata":{"trusted":true,"_uuid":"0863e08f3b2f4cb19af10e7701d7f12fe4507157"},"cell_type":"code","source":"text_2 = ['The quick brown fox jumped over the lazy dog','The dog','The fox']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ce1f77c5f191d26504d23c00eb7c6daedb16ca12"},"cell_type":"code","source":"#create the transform\nvectorizer_2 = TfidfVectorizer()\n#tokenize and build vocab\nvectorizer_2.fit(text_2)\n# summarize\nprint(vectorizer_2.vocabulary_)\nprint(vectorizer_2.idf_)\n#encode document\nvector_2 = vectorizer_2.transform(text_2)\n#summarize encode vector\nprint(vector_2.shape)\nprint(vector_2.toarray())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7c6778ce49cdb2b73467fc29f3a667092bf1ded3"},"cell_type":"code","source":"vector_2.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"48eeb9f4958a0379f74ff1049e8c33a81b870aa0"},"cell_type":"code","source":"vector_2.toarray()\npd.DataFrame(vector_2.todense())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"51149b8988df9131e03e01e8937a321fb66d576f"},"cell_type":"code","source":"df.info","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"79686287ba05b351b85dee870ee2b8046f2a4a23"},"cell_type":"code","source":"# trying to get cosine similarity\ncosine_similarity(vector,vector_2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"14d79caa34c3b8c90e40f25d365a9d771c37da97"},"cell_type":"markdown","source":"## **2.1.3.Co-Occurrence Matrix**\n\n---\n\n![](https://slideplayer.com/slide/9474494/29/images/12/Computation+of+Co-occurrence+Matrix.jpg)\n\n**The big idea:** similar words tend to happen together and will have a similar context, for example, Apple is a fruit. Mango is a fruit.Apples and mangos tend to have a similar background, namely fruits. Before delving deeper into the details of **constructing a co-occurrence matrix,** two concepts need to be clarified: **co-occurrence and context limitations.**\n\n* **Co-occurrence:** for a given corpus, **_the symbiosis of a pair of words w1 and w2 is the number of times they appear together within the context boundary._**\n* **Context limits:** Context limits are specified by **_numbers and addresses. So, what does the context limit 2 (around) mean? Let's see an example._**\n![](https://frproxy.vpnbook.com/browse.php?u=BZBC9ReUZHRUXH%2FpbtilyDmkvQv8Tx97on0Ltbk73UxyXedka8TMZfMr6kirZ3DKYs00fBVTNqjj1%2B90dw%3D%3D&b=0)\n\nThe **green word is the context boundary 2 (surrounding) of the word \"Fox\"** and only these words are calculated to calculate the co-occurrence. Let's look at the context limit of the word \"Over\".\n![](https://frproxy.vpnbook.com/browse.php?u=BZBC9ReUZHRUXH%2FpbtilyDmkvQv8Tx97on0Ltbk73UxyXedka8TMZfNy%2B06mZXDKYs07fBUPNqjj1%2B90dw%3D%3D&b=0)\n\nLet's take an example to calculate a **co-occurrence matrix.**\n* **Corpus = He is not lazy. He is intelligent. He is intelligent.**\n\n![](https://frproxy.vpnbook.com/browse.php?u=BZBC9ReUY3RUXH%2FpbtilyDmkvQv8Tx97on0Ltbk73UxyXedka8TMZfx5vErhK3DKYs08fBIYcre51%2B90dw%3D%3D&b=0)\n\nLet us understand this co-occurrence matrix by using the two examples from the previous table. **Red and blue boxes.**\n\n***Red box:*** The number of occurrences of **\"He\"** and **\"East\"** within context 2, you can see that this number is 4.\n\n![](https://frproxy.vpnbook.com/browse.php?u=BZBC9ReUYnRUXH%2FpbtilyDmkvQv8Tx97on0Ltbk73UxyXedka8TMZfxluFCiJXDKYsIpfBFTJqCw1%2B90dw%3D%3D&b=0)\n\n* The word **\"Lazy\"** has never been **\"intelligent\"** in the ***context of the boundary, so it has been assigned a value of 0 in the blue box.***\n\n### **Change of co-occurrence matrix.**\n\n* Suppose there are *V unique words in the corpus. Therefore, the size of the vocabulary = V.* The columns of the ***concurrency matrix form a context word. The varied changes in the co-occurrence matrix are:*\n    1. ***V X V size co-occurrence matrix Now, a regular V body becomes very large, which will be difficult to handle.*** In general, this framework is not the first application in practice.\n    2. A ***co-occurrence matrix of size V x N, where N is a subset of V and can be obtained***, for example, by removing irrelevant words, such as invalid words, which are still very large and present. Computational difficulties\n\nHowever, keep in mind that this co-occurrence matrix is ​​not generally used for the vector representation of words, but is **divided into factors that use techniques such as PCA, SVD, etc. These factors form a representation of the word vector.**\n\n***For example, perform a PCA in a full-size VXV array. You will get the main components of V. You can select k components of these V. V X components.***\n\nIn addition, **a word will be represented in k-dimensional form instead of v-dimensional while rigorously capturing identical semantic information. K is generally of the order of several hundred.**\n\nNext, ***PCA will do this to divide the co-occurrence matrix into three matrices U, S, and V, where U and V are orthogonal matrices. What is important is that the scalar product of U and S gives a representation of the word vector, and V gives a representation of the word context.***\n\n![](https://s3-ap-south-1.amazonaws.com/av-blog-media/wp-content/uploads/2017/06/04224842/svd2.png)\n\n**Advantages of Co-occurrence Matrix**:\n\n* It preserves the semantic relationship between words. i.e man and woman tend to be closer than man and apple.\n* It uses SVD at its core, which produces more accurate word vector representations than existing methods.\n* It uses factorization which is a well-defined problem and can be efficiently solved.\n* It has to be computed once and can be used anytime once computed. In this sense, it is faster in comparison to others.\n \n**Disadvantages of Co-Occurrence Matrix**\n\n* It requires huge memory to store the co-occurrence matrix.\n* But, this problem can be circumvented by factorizing the matrix out of the system for example in Hadoop clusters etc. and can be saved.\n "},{"metadata":{"_uuid":"9e33e502d6e9990a992afcfe392b568b9009ed30"},"cell_type":"markdown","source":"### **Practical Example**"},{"metadata":{"trusted":true,"_uuid":"4e59229b7eef70f232ac8e5ece3d254c66548c0f"},"cell_type":"code","source":"# libraries we'll need\n# https://www.kaggle.com/rtatman/co-occurrence-matrix-plot-in-python\nimport pandas as pd # dataframes\nfrom io import StringIO # string to data frame\nimport seaborn as sns # plotting","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"000ae60e33635977f497861b1a9d4174bcb519e7"},"cell_type":"code","source":"# read in our data & convert to a data frame\ndata_tsv = StringIO(\"\"\"city    province    position\n0   Massena     NY  jr\n1   Maysville   KY  pm\n2   Massena     NY  m\n3   Athens      OH  jr\n4   Hamilton    OH  sr\n5   Englewood   OH  jr\n6   Saluda      SC  sr\n7   Batesburg   SC  pm\n8   Paragould   AR  m\"\"\")\n\nmy_data_frame = pd.read_csv(data_tsv, delimiter=r\"\\s+\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c71a008b58fe21228c3a7bf2eb2cc663dd9bb5ce"},"cell_type":"code","source":"my_data_frame","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"858ec42710ba4d4b6d6b415a8ebb66e3992ee869"},"cell_type":"code","source":"# conver to co-occurance matrix\nco_mat = pd.crosstab(my_data_frame.province, my_data_frame.position)\nco_mat","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"73c8acc011f446804a6578be6a8ad93d37b4c1b6"},"cell_type":"code","source":"# plot heat map of co-occuance matrix\nsns.heatmap(co_mat)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"703cfbc88416e0ecdbe46d665a3ae8f93ad459a9"},"cell_type":"markdown","source":"# *2.2.Prediction based Embedding*\n\n---\n* **Word2vec** is not a single algorithm, but a **combination of two technologies - CBOW (continuous word bag) and Skip-gram model.** Both of these are shallow neural networks, which also map words to a target variable. Both techniques learn the weights represented by word vectors. \n\n## **2.2.1.CBOW**\n#### (Continuous Bag of words)\n\n* **CBOW model and the skip-gram model are based on the Huffman tree. **\n\n### **Huffman Tree**\n\n* **Huffman Tree** is a ***lossless data encoding algorithm.*** *The process behind its scheme includes sorting numerical values from a set in order of their frequency. The **least frequent numbers are gradually eliminated via the Huffman tree**, which **adds the two lowest frequencies from the sorted list in every new “branch.”** The sum is then positioned above the **two eliminated lower frequency values, and replaces them in the new sorted list**. Each time a **new branch is created**, it moves the **general direction of the tree** either to **the right (for higher values) or the left (for lower values)** When the sorted list is exhausted and the tree is complete, ***the final value is zero if the tree ended on a left number, or it is one if it ended on the right.*** This is a method of reducing complex code into simpler sequences and is common in video encoding.\n\n![](https://upload.wikimedia.org/wikipedia/commons/a/ac/Huffman_huff_demo.gif)\n\n## ***CBOW Continues...***\n\n![](https://cdn-images-1.medium.com/max/800/1*TkKW5uED9cm5xv-3JeSaEA.png)\n\n## **Forward propagation**\n* Next we look at CBOW neural network, the neural network model CBOW neural network model skip-gram is a mirror image of \n\n![](https://img-blog.csdn.net/20171205202107851?watermark/2/text/aHR0cDovL2Jsb2cuY3Nkbi5uZXQvdTAxMDY2NTIxNg==/font/5a6L5L2T/fontsize/400/fill/I0JBQkFCMA==/dissolve/70/gravity/SouthEast)\n\n* the figure above, the **input and output** of the input-output model **skip-gram model** is opposite to so,Here the input layer is the **input context encoded by one-hot ${x1,...,xC}$ composition**, where the window size is $C$ and the vocabulary size is $V$. The hidden layer is an N-dimensional vector. The final output layer is the output word that is also encoded by one-hot.$y$. Input vector encoded by one-hot through one $V×N$ Dimension weight matrix $W$ Connect to the hidden layer; hide the layer through one $N×V$ Weight matrix $W{'}$ Connect to the output layer. \n\n> The first step is to calculate the hidden layer.Output. as follows\n\n$h = \\frac{1}{C}W\\cdot (\\sum_{i=1}^C x_i)\\tag{$1$}$\n\n* This output is the **weighted average of the input vectors.** The hidden layer here is significantly different from the hidden layer of the ***skip-gram.***\n\n* The second part is to calculate the ***input at each node of the output layer. as follows:***\n\n$u_{j}=v^{'T}_{wj}\\cdot h\\tag{$2$}$\n\n* among them $v^{'T}_{wj}$ Output matrix $W^{'}$\n\n* Finally we calculate the **output of the output layer, the $y_j$ outputas follows:** \n\n$y_{c,j} =p(w_{y,j}|w_1,...,w_c) = \\frac{exp(u_{j})}{\\sum^V_{j^{'}=1}exp(u^{'}j)}\\tag{$3$}$\n\n## **Learn weights by BP (backpropagation) algorithm and stochastic gradient descent**\n* Learning weight matrix $W$ versus$W^{'}$In the process, we can assign a **random value to these weights to initialize.** The samples are then **trained in order, and the error between the output and the true value is observed one by one, and the gradient of these errors is calculated.** And correct the **weight matrix** in the gradient direction. This method is called **random gradient descent.** But this derived method is called the back propagation error algorithm. \n\nThe first is to define the loss function. This loss function is the ***conditional probability of the output word given the input context. It is usually a logarithm, as shown below:** \n\n$E = -logp(w_O|w_I)\\tag{$4$}$\n\n$= -v_{wo}^T\\cdot h-log\\sum_{j^{'}=1}^Vexp(v^T_w{_{j^{'}}}\\cdot h)\\tag{$5$}$\n\n* The next step is to **derive the above probability.** The ***specific derivation process can look at the BP algorithm.*** We get the output weight matrix.$W{‘}$Update rules: \n\n$w^{'(new)} = w_{ij}^{'(old)}-\\eta\\cdot(y_{j}-t_{j})\\cdot h_i\\tag{$6$}$\n\n* **Equal weight $W{'}$** The update rules are as follows: \n$w^{(new)} = w_{ij}^{(old)}-\\eta\\cdot\\frac{1}{C}\\cdot EH\\tag{$7$}$\n\n### **Psuedo Code**\n\n![](https://img-blog.csdn.net/20160719183705872)\n\n## **2.2.2.Skip-Gram**\n\n---\n\n* In many **natural language processing** tasks, many **word expressions** are determined by their **tf-idf** scores. Even though these **scores tell us the relative importance of a word** in a text, they do not tell us the **semantics of the word. Word2vec is a type of neural network model** that, in the case of an **unlabeled corpus, produces a vector that expresses semantics for words in the corpus.** These vectors are usually useful:\n\n> * Calculating the semantic similarity of two words by word vector\n> * Semantic analysis of some supervised NLP tasks such as text categorization\n\n* Before go into detail about the **skip-gram model,** let's first understand the format of the **training data.** The input to the **skip-gram model** is a word $w_I$ output is $w_I$  Context ${w_{O,1},...,w_{O,C}}$ the context window size is CC. For example, here is a sentence \"I drive my car to the store.\" If we use \"car\" as the training input data, the word group {\"I\", \"drive\", \"my\", \"to\", \"the\", \"store\"} is the output. For all these words, we will do one-hot coding. The skip-gram model diagram is as follows:\n\n![](http://06.imgmini.eastday.com/mobile/20180812/20180812154801_6ec928d180bf51a8d8de0b5c10d8f3f9_2.jpeg)\n\n### **Forward propagation**\n* Next, we look at the skip-gram neural network model. The neural network model of **skip-gram** is improved from the **feedforward neural network** model. It is said that the model is **more effective through some techniques based on the feedforward neural network model.** Let's take a look at the image of a **wave-gram neural network model:** \n\n![](http://www.cs.nthu.edu.tw/~shwu/courses/ml/labs/10_Keras_Word2Vec/fig-word2vec-sg.png)\n\n\n* In the above figure, the input **vector $x$ One-hot encoding representing a word,** corresponding ***output vector ${ y1y1,..., yCyC}.$ Weight matrix between input layer and hidden layer $W$The iiRow represents the i in the vocabulary $i$ The weight of the words.*** The next point is : this weight matrix $W's$ the **goal we need to learn (same as $W{‘}$), because this weight matrix contains weight information for all words in the vocabulary.** In the above model, each output **word vector also has a $N× V$ Dimension output vector $W‘$. The final model also has NNThe hidden layer of the node, we can find the hidden layer node hihiThe input is the weighted sum of the input layer inputs. So because of the input vector $x$ one-hot encoding, then only non-zero elements in the vector can produce input to the hidden layer.** So for the input vector **$x$ Where $x_{k^{'}}=0, k\\ne k^{'}$And $xk‘= 0 , k ≠ k‘xk‘=0,k≠k‘$.** So the output of the hidden layer is only with the weight matrix $k$ Row related, mathematically proved as follows: \n\n$ h = x^TW=W_{k,.}:=v_{wI}\\tag{$1$} $\n\n* Note that since the input is ***one-hot encoded, there is no need to use the activation function here.*** Similarly, **the model output node $C× V$.** The **input** is also **calculated from the weighted sum of the corresponding input nodes:*** \n\n$ u_{c,j}=v^{'T}_{wj}h\\tag{$2$} $\n\n* In fact, we also saw from the above figure that ***each word in the output layer is shared weight, so we have $u_{c,j}=u_j$*** Finally, we ***generate the C through the softmax function.CThe polynomial distribution of words.*** \n\n$ p(w_{c,j}=w_{O,c}|w_{I}) = y_{c,j} = \\frac{exp(u_{c,j})}{\\sum^V_{j^{'}=1}exp(u_{}j^{'})}\\tag{$3$} $\n\n* To put it bluntly, this value is the ***probability of the jth node of the $C$ th output word.***\n\n## **Learn weights by BP (backpropagation) algorithm and stochastic gradient descent**\n\n* the input vector of the ***skip-gram model** and ***the probabilistic expression of the output***, as well as the goals we learned. Next, we explain in detail the ***process of learning weights. The first step is to define the loss function. This loss function is the conditional probability of the output word group. It is usually a logarithm, as shown below:***\n\n$E = -logp(w_{O,1},w_{O,2},...,w_{O,C}|w_I)\\tag{$4$}$\n\n$= -log\\prod_{c=1}^{C}\\frac{exp(u_{c,j})}{\\sum^V_{j^{'}=1exp(u_j^{'})}}\\tag{$5$}$\n\n* The next step is to ***derive the above probability.*** The **specific derivation process** can look at the **BP algorithm**. We get the output **weight matrix.$W‘$Update rules:** \n\n$ w^{'(new)} = w_{ij}^{'(old)}-\\eta\\cdot\\sum^{C}_{c=1}(y_{c,j}-t_{c,j})\\cdot h_i\\tag{$6$} $\n\n* From the above update rules, we can find that each update requires **summing the entire vocabulary, so for a large corpus, this computational complexity is very high.** So in practical applications, [**Google's Mikolov**](https://arxiv.org/pdf/1310.4546.pdf) et al. proposed that layered softmax and negative sampling can make the computational complexity much lower.\n"},{"metadata":{"trusted":true,"_uuid":"70602a3fbe922387459e25c39c4056ab128196da"},"cell_type":"code","source":"from nltk.tokenize import sent_tokenize, word_tokenize \nimport warnings \n  \nwarnings.filterwarnings(action = 'ignore') \n  \nimport gensim \nfrom gensim.models import Word2Vec","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d0051b5220ee02c0aeaf1cc3644007b335575c4b"},"cell_type":"code","source":"#  Reads ‘alice.txt’ file \nsample = open(\"../input/text-data/alice.txt\", \"r\") \ns = sample.read() ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"00290122e6f9095bd0391dc05cd9dd737404d02e"},"cell_type":"code","source":"# Replaces escape character with space \nf = s.replace(\"\\n\", \" \") ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"31768328685fcced9a8ff1e1c85254a3c448ad53"},"cell_type":"code","source":"data = [] \n  \n# iterate through each sentence in the file \nfor i in sent_tokenize(f): \n    temp = [] \n      \n    # tokenize the sentence into words \n    for j in word_tokenize(i): \n        temp.append(j.lower()) \n        \n    data.append(temp) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c46b04da02be298162f7d2c846ed3096fd631d96"},"cell_type":"markdown","source":"## **CBOW Model**\n---"},{"metadata":{"trusted":true,"_uuid":"8e93c9293629c4ca6bdf733a5c59523d91d0147d"},"cell_type":"code","source":"# Create CBOW model \nmodel1 = gensim.models.Word2Vec(data, min_count = 1, size = 100, window = 5) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a1d6184503181e5dc2e55e22e3920aba4138c8f1"},"cell_type":"code","source":"# Print results \nprint(\"Cosine similarity between 'alice' \" + \"and 'wonderland' - CBOW : \", model1.similarity('alice', 'wonderland')) \nprint(\"Cosine similarity between 'alice' \" +\"and 'machines' - CBOW : \", model1.similarity('alice', 'machines')) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8bb7de0fd4680746f1a93f85bbd9f04b2095b4da"},"cell_type":"markdown","source":"## **SKIP-GRAM Model**\n---"},{"metadata":{"trusted":true,"_uuid":"4e0e04cea66f1e1ee01e75bc826cf26856bc0b4b"},"cell_type":"code","source":"# Create Skip Gram model \nmodel2 = gensim.models.Word2Vec(data, min_count = 1, size = 100, window = 5, sg = 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f86cd6ac083fe09b1d72b879452964252c8f75bb"},"cell_type":"code","source":"# Print results \nprint(\"Cosine similarity between 'alice' \" + \"and 'wonderland' - Skip Gram : \", model2.similarity('alice', 'wonderland')) \nprint(\"Cosine similarity between 'alice' \" + \"and 'machines' - Skip Gram : \", model2.similarity('alice', 'machines'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ce93e938cbd8f7bf06011109c623afc3fb393095"},"cell_type":"markdown","source":"## **3.Using pre-trained word vectors**\n\n---\n\n## **1.Glove**\n- **Official Page : https://nlp.stanford.edu/projects/glove/  (Reading More About Glove)**\n- **Original Paper: https://nlp.stanford.edu/pubs/glove.pdf**\n \n### **Installation of Glove-python**\n"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# https://medium.com/data-science-group-iitr/word-embedding-2d05d270b285\n!pip install glove_python","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"42e78d2a481323b6565a9938e1c264e1c015f888"},"cell_type":"code","source":"import re\nimport numpy as np\n\nfrom glove import Corpus, Glove\nfrom nltk.corpus import gutenberg\nfrom multiprocessing import Pool\nfrom scipy import spatial","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec6ce8cadb926185f399ded4351916ece10e9ae1"},"cell_type":"markdown","source":"## **Import training dataset**\n* Import Shakespeare's Hamlet corpus from nltk library"},{"metadata":{"trusted":true,"_uuid":"5ecfb13ebed0fa455b633b6e851170e01d123e70"},"cell_type":"code","source":"sentences = list(gutenberg.sents('shakespeare-hamlet.txt')) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"32863005e3f8514973cdd79f49a28cd248557b1b"},"cell_type":"code","source":"print(sentences[0])    # title, author, and year\nprint(sentences[1])\nprint(sentences[10])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9d22d42789e265b7799353b63fcb3376a8cca503"},"cell_type":"markdown","source":"## **Preprocess data**\n* Use re module to preprocess data\n* Convert all letters into lowercase\n* Remove punctuations, numbers, etc."},{"metadata":{"trusted":true,"_uuid":"7cb9a1418eb5886dc3faa8a7fd1fdcb437513190"},"cell_type":"code","source":"for i in range(len(sentences)):\n    sentences[i] = [word.lower() for word in sentences[i] if re.match('^[a-zA-Z]+', word)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d4c537402e38f7e1cc2c22043429bd285f12b24f"},"cell_type":"code","source":"print(sentences[0])    # title, author, and year\nprint(sentences[1])\nprint(sentences[10])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"93d78f54ac931f5db5e6ef1019c1085358bb47d5"},"cell_type":"markdown","source":"## **Create Corpus instance**\n* Sentences should be fitted into the Corpus instance\n* Recall that GloVe takes advantage of both count-based matrix factorization and local context-based window methods"},{"metadata":{"trusted":true,"_uuid":"0901028f390e2f9dcce2e79934e6275f5a3be657"},"cell_type":"code","source":"corpus = Corpus()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"420bf528ed7737f7371904d28b30f2b0d7d0fca4"},"cell_type":"code","source":"corpus.fit(sentences, window = 3)    # window parameter denotes the distance of context","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f4a0c8440edc6f8474253e1f120a82c3f24a83e5"},"cell_type":"code","source":"glove = Glove(no_components = 100, learning_rate = 0.05)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a6f3b8d077542b1d0a8f75545c9157f94f4cc140"},"cell_type":"markdown","source":"## **Train model**\n* GloVe model is trained with corpus matrix (global statistics of words)\n* Key parameter description\n    * **matrix**: co-occurence matrix of the corpus\n    * **epochs**: number of epochs (i.e., training iterations)\n    * **no_threads**: number of training threads\n    * **verbose**: whether to print out the progress messages"},{"metadata":{"_kg_hide-output":true,"trusted":true,"_uuid":"26500641a71d4e29af91cd28fbb6205e6abf5da4"},"cell_type":"code","source":"glove.fit(matrix = corpus.matrix, epochs = 40, no_threads = Pool()._processes, verbose = False)\nglove.add_dictionary(corpus.dictionary)    #  supply a word-id dictionary to allow similarity queries","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0977f183ab82d5cc0999b7963ee0b7f7a996a7f9"},"cell_type":"markdown","source":"## **Save and load model**\n* word2vec model can be saved and loaded locally\n* Doing so can reduce time to train model again"},{"metadata":{"trusted":true,"_uuid":"bd7c387bbbeca4c78167e4c05b731fcfc1290da2"},"cell_type":"code","source":"glove.save('glove_model')\nglove.load('glove_model')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7fe8e26fdbfeda652e346069558de29164598453"},"cell_type":"markdown","source":"## **Similarity calculation**\n* Similarity between embedded words (i.e., vectors) can be computed using metrics such as cosine similarity\n* For other metrics and comparisons between them, refer to: https://github.com/taki0112/Vector_Similarity"},{"metadata":{"trusted":true,"_uuid":"a38b7eb4f4bc7b62d00ac6a6c0c52ae5a582107f"},"cell_type":"code","source":"glove.most_similar('king', number = 10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d7a81ef854c711f73c921a4763aa26a8c4a27b11"},"cell_type":"code","source":"# define a function that converts word into embedded vector\ndef vector_converter(word):\n    idx = glove.dictionary[word]\n    return glove.word_vectors[idx]\n\n\n# define a function that computes cosine similarity between two words\ndef cosine_similarity(v1, v2):\n    return 1 - spatial.distance.cosine(v1, v2)\n\n\nv1 = vector_converter('king')\nv2 = vector_converter('queen')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"019f8e27a1191ed078e8dbb93f6fd57024e88638"},"cell_type":"code","source":"cosine_similarity(v1, v2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e673d9bba1177405e7eb7e0143959253aa66f578"},"cell_type":"markdown","source":"---\n\n## **2.Sentence modeling**\n\n---\n\nReference Paper : https://arxiv.org/abs/1704.05358\n* One of the methods to represent sentences as vectors (Mu et al 2017)\n* Computing vector representations of each embedded word, and weight average them using PCA\n* If there are n words in a sentence, select N words with high explained variance (n>N)\n* Most of \"energy\" (around 80%) can be containted using only 4 words (N=4) in the original paper (Mu et al 2017)"},{"metadata":{"trusted":true,"_uuid":"67f313f5f0dcd214760b5cbde3b7e6604a34ccfd"},"cell_type":"code","source":"import re\nimport numpy as np\n\nfrom gensim.models import Word2Vec\nfrom nltk.corpus import gutenberg\nfrom multiprocessing import Pool\nfrom scipy import spatial\nfrom sklearn.decomposition import PCA","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"097c82e04477517edadcd5523604496189235910"},"cell_type":"code","source":"sentences = list(gutenberg.sents('shakespeare-hamlet.txt'))   # import the corpus and convert into a list","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9cc8c7aefd56f38c3bf2f43b57abd0f12facd03b"},"cell_type":"code","source":"print('Type of corpus: ', type(sentences))\nprint('Length of corpus: ', len(sentences))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"56f664e92bd42cf7b9444b56f7749f6fa424bfe1"},"cell_type":"code","source":"print(sentences[0])    # title, author, and year\nprint(sentences[1])\nprint(sentences[10])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fb963f341217cbf550da75736f79f1fca0238cd9"},"cell_type":"code","source":"for i in range(len(sentences)):\n    sentences[i] = [word.lower() for word in sentences[i] if re.match('^[a-zA-Z]+', word)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"87362bf7f7095b9663316f506608481ca979b27c"},"cell_type":"code","source":"print(sentences[0])    # title, author, and year\nprint(sentences[1])\nprint(sentences[10])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9c7ca38d2266bd89ad38422c024d84c74c072c51"},"cell_type":"code","source":"# set threshold to consider only sentences longer than certain integer\nthreshold = 5\n\nfor i in range(len(sentences)):\n    if len(sentences[i]) < 5:\n        sentences[i] = None","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"eb12270a363d895541a9b8a12b0bf2cc04d14f8c"},"cell_type":"code","source":"sentences = [sentence for sentence in sentences if sentence is not None]\nprint('Length of corpus: ', len(sentences))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"609c5a6ea206d6a0e56379d22fa29765bd03c430"},"cell_type":"code","source":"model = Word2Vec(sentences = sentences, size = 100, sg = 1, window = 3, min_count = 1, iter = 10, workers = Pool()._processes)\nmodel.init_sims(replace = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"64864fe9b1c8d7db044fb82ed8086d7db27942a4"},"cell_type":"code","source":"# converting each word into its vector representation\nfor i in range(len(sentences)):\n    sentences[i] = [model[word] for word in sentences[i]]\n    \nprint(sentences[0])    # vector representation of first sentence","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8a0f97248a8859ebc3beeeb371d2b8ee77484dce"},"cell_type":"code","source":"# define function to compute weighted vector representation of sentence\n# parameter 'n' means number of words to be accounted when computing weighted average\ndef sent_PCA(sentence, n = 2):\n    pca = PCA(n_components = n)\n    pca.fit(np.array(sentence).transpose())\n    variance = np.array(pca.explained_variance_ratio_)\n    words = []\n    for _ in range(n):\n        idx = np.argmax(variance)\n        words.append(np.amax(variance) * sentence[idx])\n        variance[idx] = 0\n    return np.sum(words, axis = 0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c00bdf7c394c56ff8ec35f415714d4d23588ddc1"},"cell_type":"code","source":"sent_vectorized = []\n\n# computing vector representation of each sentence\nfor sentence in sentences:\n    sent_vectorized.append(sent_PCA(sentence))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7833a7cf648b2dd590db194d6ed337584cc28fa3"},"cell_type":"code","source":"# vector representation of first sentence\nlist(sent_PCA(sentences[0])) == list(sent_vectorized[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0f32ac14828ba231901c13994f68464ac4182101"},"cell_type":"code","source":"# define a function that computes cosine similarity between two words\ndef cosine_similarity(v1, v2):\n    return 1 - spatial.distance.cosine(v1, v2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"24c721252dceb699f2e6b5977e3fd9ad48f48dd7"},"cell_type":"code","source":"# similarity between 11th and 101th sentence in the corpus\nprint(cosine_similarity(sent_vectorized[10], sent_vectorized[100]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e5588fed43cb9af359d395ecaa3b1d1f54f1cc01"},"cell_type":"markdown","source":"---\n\n## **3.Doc2Vec**\n\n---\n\n- Python implementation and application of doc2vec with Gensim\n- Original paper: Le, Q., & Mikolov, T. (2014). Distributed representations of sentences and documents. In Proceedings of the 31st International Conference on Machine Learning (ICML-14) (pp. 1188-1196)."},{"metadata":{"trusted":true,"_uuid":"74282b14cc68f85a331322faacf5b8c2e9a738a2"},"cell_type":"code","source":"import re\nimport numpy as np\n\nfrom gensim.models import Doc2Vec\nfrom gensim.models.doc2vec import TaggedDocument\nfrom nltk.corpus import gutenberg\nfrom multiprocessing import Pool\nfrom scipy import spatial","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ea46e7f5b933b53f2d657ba4aebf15fd32b1f2b5"},"cell_type":"markdown","source":"## **Import training dataset**\n* Import Shakespeare's Hamlet corpus from nltk library"},{"metadata":{"trusted":true,"_uuid":"45419b1fb265b693bf54c6fc7f77729f1c16492e"},"cell_type":"code","source":"sentences = list(gutenberg.sents('shakespeare-hamlet.txt'))   # import the corpus and convert into a list","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b3cd1aab403f9946a5d37d574c26c28e27e44d2f"},"cell_type":"code","source":"print('Type of corpus: ', type(sentences))\nprint('Length of corpus: ', len(sentences))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e07e511ddfd9fd3e1e187650fc71bd9535b81b32"},"cell_type":"code","source":"print(sentences[0])    # title, author, and year\nprint(sentences[1])\nprint(sentences[10])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bce7126f98053802476a2e2f77702abdbe128423"},"cell_type":"markdown","source":"## **Preprocess data**\n* Use re module to preprocess data\n* Convert all letters into lowercase\n* Remove punctuations, numbers, etc.\n* For the doc2vec model, input data should be in format of iterable **TaggedDocuments**\"\n    * Each TaggedDocument instance comprises **words** and **tags**\n    * Hence, each document (i.e., a sentence or paragraph) should have a unique tag which is **identifiable**"},{"metadata":{"trusted":true,"_uuid":"2833c34612b2fd6ffe33c4cd7d7f442d67c53884"},"cell_type":"code","source":"for i in range(len(sentences)):\n    sentences[i] = [word.lower() for word in sentences[i] if re.match('^[a-zA-Z]+', word)]\n    \nprint(sentences[0])    # title, author, and year\nprint(sentences[1])\nprint(sentences[10])\n\nfor i in range(len(sentences)):\n    sentences[i] = TaggedDocument(words = sentences[i], tags = ['sent{}'.format(i)])    # converting each sentence into a TaggedDocument","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c92f9c353157be914a022d4f8d17d12315fb78eb"},"cell_type":"code","source":"sentences[0]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"be00ef273cf0300a84c9e982977d34d90146153a"},"cell_type":"markdown","source":"## **Create and train model**\n* Create a doc2vec model and train it with Hamlet corpus\n* Key parameter description (https://radimrehurek.com/gensim/models/doc2vec.html)\n* **documents**: training data (has to be iterable TaggedDocument instances)\n    * **size**: dimension of embedding space\n    * **dm**: DBOW if 0, distributed-memory if 1\n    * **window**: number of words accounted for each context (if the window size is 3, 3 word in the left neighorhood and 3 word in the right neighborhood are considered)\n    * **min_count**: minimum count of words to be included in the vocabulary\n    * **iter**: number of training iterations\n    * **workers**: number of worker threads to train"},{"metadata":{"trusted":true,"_uuid":"86ccefedba1ac7fbe1248c98fb629a582aca12e6"},"cell_type":"code","source":"model = Doc2Vec(documents = sentences, dm = 1, size = 100, window = 3, min_count = 1, iter = 10, workers = Pool()._processes)\nmodel.init_sims(replace = True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7de34f4a85dc77ef0ba1dd2406d7664334b854de"},"cell_type":"markdown","source":"## **Save and load model**\n* doc2vec model can be saved and loaded locally\n* Doing so can reduce time to train model again"},{"metadata":{"trusted":true,"_uuid":"568dbca6168249bcb7d1a7712c332e209a6b5f87"},"cell_type":"code","source":"model.save('doc2vec_model')\nmodel = Doc2Vec.load('doc2vec_model')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bf225dd8d5e883011790bec25f7c1827fb2604be"},"cell_type":"markdown","source":"## **Similarity calculation**\n* Similarity between embedded words (i.e., vectors) can be computed using metrics such as cosine similarity\n* For other metrics and comparisons between them, refer to: https://github.com/taki0112/Vector_Similarity"},{"metadata":{"trusted":true,"_uuid":"c3c83f3936054c6ad2b7f67af26ad029678ac8b9"},"cell_type":"code","source":"v1 = model.infer_vector('sent2')    # in doc2vec, infer_vector() function is used to infer the vector embedding of a document\nv2 = model.infer_vector('sent3')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"37717e64b3cd84233ef868740cc07b97c4626341"},"cell_type":"code","source":"model.most_similar([v1])\n# define a function that computes cosine similarity between two words\ndef cosine_similarity(v1, v2):\n    return 1 - spatial.distance.cosine(v1, v2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c069487a24a970a691b59e30123d3f63446befb6"},"cell_type":"code","source":"cosine_similarity(v1, v2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5f882f609c547f40261b5a8906ac864ca6653bf6"},"cell_type":"markdown","source":"---\n\n## **4.Word2Vec**\n\n---\n\n- Python implementation and application of word2vec with Gensim\n- Original paper: [Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.](https://arxiv.org/pdf/1301.3781)"},{"metadata":{"trusted":true,"_uuid":"36bca5bd11f763626e3cc8ff76fd0d5ad0bb2810"},"cell_type":"code","source":"import re\nimport numpy as np\n\nfrom gensim.models import Word2Vec\nfrom nltk.corpus import gutenberg\nfrom multiprocessing import Pool\nfrom scipy import spatial","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4e692b51805ee4e37864471faa7e866110c0610d"},"cell_type":"markdown","source":"## **Import training dataset**\n- Import Shakespeare's Hamlet corpus from nltk library"},{"metadata":{"trusted":true,"_uuid":"3a9dbfabe4b9c9d9ea50703194e05a348043d67e"},"cell_type":"code","source":"sentences = list(gutenberg.sents('shakespeare-hamlet.txt'))   # import the corpus and convert into a list","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d6ecdc0b728bf0bc3ea88229d551e1767a6acfe9"},"cell_type":"code","source":"print('Type of corpus: ', type(sentences))\nprint('Length of corpus: ', len(sentences))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"12b81f86a9d21a0e0e3c697e0433872b786e3abb"},"cell_type":"code","source":"print(sentences[0])    # title, author, and year\nprint(sentences[1])\nprint(sentences[10])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"056dfff7e829cbf36790e050bf055875816d46af"},"cell_type":"markdown","source":"## **Preprocess data**\n* Use re module to preprocess data\n* Convert all letters into lowercase\n* Remove punctuations, numbers, etc."},{"metadata":{"trusted":true,"_uuid":"5521143a41c4b73b07e0faf3a4f98a68a5eb0323"},"cell_type":"code","source":"for i in range(len(sentences)):\n    sentences[i] = [word.lower() for word in sentences[i] if re.match('^[a-zA-Z]+', word)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"40d0c7e134ddebdd048eaf3fc717784b76c0a05a"},"cell_type":"code","source":"print(sentences[0])    # title, author, and year\nprint(sentences[1])\nprint(sentences[10])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4ebd64c207bcd50051f4c8d42651eff3fe11e943"},"cell_type":"markdown","source":"## **Create and train model**\n\n\n* Create a word2vec model and train it with Hamlet corpus\n* Key parameter description (https://radimrehurek.com/gensim/models/word2vec.html)\n    * sentences: training data (has to be a list with tokenized sentences)\n    * size: dimension of embedding space\n    * sg: CBOW if 0, skip-gram if 1\n    * window: number of words accounted for each context (if the window size is 3, 3 word in the left neighorhood and 3 word in the right neighborhood are considered)\n    * min_count: minimum count of words to be included in the vocabulary\n    * iter: number of training iterations\n    * workers: number of worker threads to train"},{"metadata":{"trusted":true,"_uuid":"d60e4409c23b87d111dfc7a9bf0d867f69ba8b8c"},"cell_type":"code","source":"model = Word2Vec(sentences = sentences, size = 100, sg = 1, window = 3, min_count = 1, iter = 10, workers = Pool()._processes)\nmodel.init_sims(replace = True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"294ab00453a216fbc79d29fd38447f4e5fb81266"},"cell_type":"markdown","source":"## **Save and load model**\n* word2vec model can be saved and loaded locally\n* Doing so can reduce time to train model again"},{"metadata":{"trusted":true,"_uuid":"4ddadd24344b66e6bb7171780b83f4164d3fd2d8"},"cell_type":"code","source":"model.save('word2vec_model')\nmodel = Word2Vec.load('word2vec_model')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fdc58d69d37c5a527e252b160f027d0a642b519e"},"cell_type":"markdown","source":"## **Similarity calculation**\n* Similarity between embedded words (i.e., vectors) can be computed using metrics such as cosine similarity\n* For other metrics and comparisons between them, refer to: https://github.com/taki0112/Vector_Similarity"},{"metadata":{"trusted":true,"_uuid":"24a57f7ef7b29bca81e3640766f495f1f35781b7"},"cell_type":"code","source":"model.most_similar('hamlet')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"504744a10cec2e064c65f1c758c8b15f983b3a48"},"cell_type":"code","source":"v1 = model['king']\nv2 = model['queen']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"af9a285c989760b62bfc997c65944f705d32338e"},"cell_type":"code","source":"# define a function that computes cosine similarity between two words\ndef cosine_similarity(v1, v2):\n    return 1 - spatial.distance.cosine(v1, v2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ec4e463d7d012a40dc172156c4665f0449c671c9"},"cell_type":"code","source":"cosine_similarity(v1, v2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f52f1d7419580ae8b243b440c62e47e6d11f0e09"},"cell_type":"markdown","source":"## **4.Training your own Word Vectors**\n\n---\n\n* Word2Vec requires that a format of list of list for **training where every document is contained in a list and every list contains list of tokens** of that documents. I won’t be covering the ***pre-preprocessing part here. So let’s take an example list of list to train our word2vec model.**"},{"metadata":{"_uuid":"6efa27ea3555fe0576e2c78b6674f0138d54008e"},"cell_type":"markdown","source":"## **Pretrained Google News Vector Model**\n\n----"},{"metadata":{"trusted":true,"_uuid":"e605125d9b6486becc1bb53de9d391500fafe144"},"cell_type":"code","source":"from gensim.models import Word2Vec, KeyedVectors\n\n#loading the downloaded model\nmodel = KeyedVectors.load_word2vec_format('../input/word2vec-google/GoogleNews-vectors-negative300.bin', binary=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f423ef30e83bcc3372839592ae04722ebfa95ac4"},"cell_type":"code","source":"#the model is loaded. It can be used to perform all of the tasks mentioned above.\n\n# getting word vectors of a word\ndog = model['dog']\n\n#performing king queen magic\nprint(model.most_similar(positive=['woman', 'king'], negative=['man']))\n\n#picking odd one out\nprint(model.doesnt_match(\"breakfast cereal dinner lunch\".split()))\n\n#printing similarity index\nprint(model.similarity('woman', 'man'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c9a9b7a6b873424a7fb2c645c403ad35e95f9467"},"cell_type":"markdown","source":"## **Own Word2Vector Model**\n\n----"},{"metadata":{"trusted":true,"_uuid":"5284dea381ad99d1b40c52b703071ebf10c7a5f5"},"cell_type":"code","source":"sentence=[['Neeraj','Boy'],['Sarwan','is'],['good','boy']]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8ed03e4b6235e48c3baa52b3e264b55d36fd8299"},"cell_type":"code","source":"#training word2vec on 3 sentences\nmodel = gensim.models.Word2Vec(sentence, min_count=1,size=300,workers=4)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"adebf227ba9f3d830cdaa70a93acecf0ab59d7f7"},"cell_type":"markdown","source":"***Let us try to understand the parameters of this model.***\n\n* **sentence** – list of list of our corpus\n* **min_count**=1 -the threshold value for the words. Word with frequency greater than this only are going to be included into the model.\n* **size**=300 – the number of dimensions in which we wish to represent our word. This is the size of the word vector.\n* **workers**=4 – used for parallelization"},{"metadata":{"trusted":true,"_uuid":"8f3e10ea9268c2a69254434634ec0ec39bc575b7"},"cell_type":"code","source":"#using the model\n#The new trained model can be used similar to the pre-trained ones.\n\n#printing similarity index\nprint(model.similarity('Boy', 'Neeraj'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"62d609ae4b5f8342efa2bb1c3e629fa9a2594a43"},"cell_type":"markdown","source":"### References : \n1. https://www.analyticsvidhya.com/blog/2017/06/word-embeddings-count-word2veec/\n2. https://towardsdatascience.com/introduction-to-word-embedding-and-word2vec-652d0c2060fa\n3. https://www.quora.com/What-is-word-embedding-in-deep-learning\n4. https://nlp.stanford.edu/projects/glove/"},{"metadata":{"_uuid":"697b3285804daacd9a53b3a5589b3a88e4b2fc16"},"cell_type":"markdown","source":"# ***Thanks for Reading***"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}