{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h3>Word2Vec - Word Embeddings Technique:</h3>\n\n<p>\n    Word2vec is a deep learning based technique to create word embeddings for each data/ text present in the corpus while keeping their contextual meanining entact. Means, words that are usually used in similar context, will have the same or nearly same representation of embeddings. Embeddings refer here the numeric representation of textual data in high dimensioanl space which is most important for any text data to analyze and build predictive models.\n    \nThere are also some other techniques such as BoW, TF-IDF which are the basic ones and have some drawbacks which are overcome by the wrod2vec. Those shortcomings are: \n<ul>\n    <li>Length of Vector: Earlier approaches usually have very high dimensioanl space such as 500, 1000 etc, depends on the unique words in the corpous, which is computationally very expensive whereas Word2Vec has fixed legnth of numeric representation and can be controlled by the user.</li>\n    <li>Sparse Vectors: Other mentionaed appraoches are frequency based approaches that's why they have highly sparse vectors. Unlikely, word2vec can have sparse vetcor negligibly.</li>\n    <li>Semantic Meaning: Most important, word2vec captures the semantic meaning of words. Thus, words used in similar context will have same vector as numeric representation whereas other approaches lost the semantic meanings of words.</li>\n</ul>\n\n\nThere are two main variants of Word2Vec: Continuous Bag of Words (CBOW) and Skip-gram. Here's a brief overview of how Word2Vec works:\n<ol>\n    <li>Continuous Bag of Words (CBOW): In CBOW, the model predicts a target word based on its surrounding context words. The input to the model is a context window of words, and the output is the target word. CBOW is efficient and tends to work well for smaller datasets.</li>\n    <li>Skip-gram: Skip-gram, on the other hand, predicts the context words (surrounding words) given a target word. It's more computationally intensive but often performs better when you have a larger corpus.</li>\n</ol>\n\nIn essence, word2vec provides a way to represent words that preserves their semantic relationships, making them a powerful tool for understanding and working with language in a more meaningful way than traditional methods like BoW and TF-IDF.\n\n</p>","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-11-09T15:20:27.310472Z","iopub.execute_input":"2023-11-09T15:20:27.310875Z","iopub.status.idle":"2023-11-09T15:20:27.816563Z","shell.execute_reply.started":"2023-11-09T15:20:27.310842Z","shell.execute_reply":"2023-11-09T15:20:27.815078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gensim\nfrom gensim.utils import simple_preprocess","metadata":{"execution":{"iopub.status.busy":"2023-11-09T15:20:34.292641Z","iopub.execute_input":"2023-11-09T15:20:34.293205Z","iopub.status.idle":"2023-11-09T15:20:46.639512Z","shell.execute_reply.started":"2023-11-09T15:20:34.293169Z","shell.execute_reply":"2023-11-09T15:20:46.638303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import nltk\nfrom nltk import sent_tokenize\nfrom nltk.corpus import stopwords","metadata":{"execution":{"iopub.status.busy":"2023-11-09T15:20:46.642164Z","iopub.execute_input":"2023-11-09T15:20:46.643050Z","iopub.status.idle":"2023-11-09T15:20:47.496766Z","shell.execute_reply.started":"2023-11-09T15:20:46.643000Z","shell.execute_reply":"2023-11-09T15:20:47.495610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nltk.download('stopwords')","metadata":{"execution":{"iopub.status.busy":"2023-11-09T15:20:54.226160Z","iopub.execute_input":"2023-11-09T15:20:54.226566Z","iopub.status.idle":"2023-11-09T15:20:54.325794Z","shell.execute_reply.started":"2023-11-09T15:20:54.226536Z","shell.execute_reply":"2023-11-09T15:20:54.324476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"file_path = '''/kaggle/input/game-of-thrones-books/'''\nos.listdir(file_path)","metadata":{"execution":{"iopub.status.busy":"2023-11-09T15:20:58.433856Z","iopub.execute_input":"2023-11-09T15:20:58.434333Z","iopub.status.idle":"2023-11-09T15:20:58.450826Z","shell.execute_reply.started":"2023-11-09T15:20:58.434298Z","shell.execute_reply":"2023-11-09T15:20:58.449918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"books = []\nfor book in os.listdir(file_path):\n    book_name = os.path.join(file_path, book)\n    with open(book_name, encoding='utf-8', errors='ignore') as file:\n        sentences = sent_tokenize(file.read())\n        for sent in sentences:\n            preprocessed_sent = simple_preprocess(sent)\n            books.append([each for each in preprocessed_sent if each.lower() not in stopwords.words('english')])","metadata":{"execution":{"iopub.status.busy":"2023-11-09T15:21:54.487834Z","iopub.execute_input":"2023-11-09T15:21:54.488278Z","iopub.status.idle":"2023-11-09T15:26:11.344053Z","shell.execute_reply.started":"2023-11-09T15:21:54.488225Z","shell.execute_reply":"2023-11-09T15:26:11.342861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(books)","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:11:03.018071Z","iopub.execute_input":"2023-11-09T16:11:03.018466Z","iopub.status.idle":"2023-11-09T16:11:03.026252Z","shell.execute_reply.started":"2023-11-09T16:11:03.018437Z","shell.execute_reply":"2023-11-09T16:11:03.025118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"books[:5]","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:11:16.693617Z","iopub.execute_input":"2023-11-09T16:11:16.693995Z","iopub.status.idle":"2023-11-09T16:11:16.701532Z","shell.execute_reply.started":"2023-11-09T16:11:16.693966Z","shell.execute_reply":"2023-11-09T16:11:16.700402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.shape(books)","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:11:16.894295Z","iopub.execute_input":"2023-11-09T16:11:16.894677Z","iopub.status.idle":"2023-11-09T16:11:16.958890Z","shell.execute_reply.started":"2023-11-09T16:11:16.894647Z","shell.execute_reply":"2023-11-09T16:11:16.957952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = gensim.models.Word2Vec(window=5, min_count=5)","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:11:47.999169Z","iopub.execute_input":"2023-11-09T16:11:47.999529Z","iopub.status.idle":"2023-11-09T16:11:48.006038Z","shell.execute_reply.started":"2023-11-09T16:11:47.999502Z","shell.execute_reply":"2023-11-09T16:11:48.004940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel.build_vocab(books)","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:11:49.078007Z","iopub.execute_input":"2023-11-09T16:11:49.078890Z","iopub.status.idle":"2023-11-09T16:11:49.765151Z","shell.execute_reply.started":"2023-11-09T16:11:49.078858Z","shell.execute_reply":"2023-11-09T16:11:49.764025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.corpus_count","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:11:50.950657Z","iopub.execute_input":"2023-11-09T16:11:50.951043Z","iopub.status.idle":"2023-11-09T16:11:50.958299Z","shell.execute_reply.started":"2023-11-09T16:11:50.951013Z","shell.execute_reply":"2023-11-09T16:11:50.957233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel.train(books, total_examples=model.corpus_count, epochs=model.epochs)","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:12:11.183351Z","iopub.execute_input":"2023-11-09T16:12:11.183951Z","iopub.status.idle":"2023-11-09T16:12:16.488776Z","shell.execute_reply.started":"2023-11-09T16:12:11.183920Z","shell.execute_reply":"2023-11-09T16:12:16.487317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.shape(model.wv.get_normed_vectors())","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:12:16.490386Z","iopub.execute_input":"2023-11-09T16:12:16.490638Z","iopub.status.idle":"2023-11-09T16:12:16.503232Z","shell.execute_reply.started":"2023-11-09T16:12:16.490616Z","shell.execute_reply":"2023-11-09T16:12:16.502056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.wv.get_normed_vectors()[:5]","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:12:16.504525Z","iopub.execute_input":"2023-11-09T16:12:16.504884Z","iopub.status.idle":"2023-11-09T16:12:16.520891Z","shell.execute_reply.started":"2023-11-09T16:12:16.504842Z","shell.execute_reply":"2023-11-09T16:12:16.519927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" y = model.wv.index_to_key","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:12:16.818210Z","iopub.execute_input":"2023-11-09T16:12:16.818561Z","iopub.status.idle":"2023-11-09T16:12:16.823389Z","shell.execute_reply.started":"2023-11-09T16:12:16.818534Z","shell.execute_reply":"2023-11-09T16:12:16.822255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y[:5]","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:12:30.350724Z","iopub.execute_input":"2023-11-09T16:12:30.351094Z","iopub.status.idle":"2023-11-09T16:12:30.357478Z","shell.execute_reply.started":"2023-11-09T16:12:30.351066Z","shell.execute_reply":"2023-11-09T16:12:30.356506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.shape(y)","metadata":{"execution":{"iopub.status.busy":"2023-11-09T16:12:35.798275Z","iopub.execute_input":"2023-11-09T16:12:35.799317Z","iopub.status.idle":"2023-11-09T16:12:35.811893Z","shell.execute_reply.started":"2023-11-09T16:12:35.799266Z","shell.execute_reply":"2023-11-09T16:12:35.810972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}