{"cells":[{"metadata":{"_uuid":"158bfb3140844a3e6153b53185ada33785965e0d"},"cell_type":"markdown","source":"![](https://www.themavencircle.com/wp-content/uploads/2016/03/Neuro-linguistic-Programming-650x435.jpg)"},{"metadata":{"_uuid":"c50cb3e8f98287e935fa9bd3a071c615a73e33ed"},"cell_type":"markdown","source":"<a id=\"0\"></a> <br>\n## Kernel Headlines\n1. [Introduction](#1)\n    1.  [Concept of Text Mining](#2)\n    2.  [Text Mining Importance and Difficulties](#3)\n    3.  [Text Mining In Buisiness](#4)\n    4.  [Dealing With Unstructured Data](#5)\n    5.  [Dealing With Dirty Space](#6)\n    6.  [Why Representation is Complex ?!](#7)\n    7. [Wath Does BagOfWords Means?!](#8)\n    8. [Ignoring Grammer and NLP-Based Features](#9)\n    9. [Feature Extraction Using BOW](#10)\n    10. [Term Frequency (TF)](#11)\n    11. [Inverse Term Frequency (IDF)](#12)\n    12. [Simple Example (TF-IDF)](#13)\n    13. [Preprocessing In Traditional Manners](#14)\n2. [Embedding Methods](#15)\n     1. [What Does Embedding Means?!](#16)\n     2. [What Does N-Gram Means?!](#17)\n     3. [Main Disadvantages of N-Gram](#18)\n     4. [What Does Embeddings Provide For Text Mining](#19)\n3. [Take a Glance to Data](#20)\n     1. [WordCloud](#21)\n     2. [Target Distribution](#22)\n4. [Start Vectorizing by Simple CountVectorizer](#23)\n    1. [Make Your Hands Dirty;-) ](#24)\n    2. [Some Suggestions About Text Vectorizing ](#25)\n    3. [Classification](#26)\n        4. [Naive Bayes](#27)\n        5. [RandomForest](#28)\n        6. [MLP](#29)\n        7. [LogicRegression](#30)\n        8. [Classifiers Accuracy Comparisons on CountVectorizer](#31)\n        9. [Curse of Dimentionality](#32)\n5. [Embedding Methods_Training Our Word2Vec Model](#33)\n\t1. [Introduction](#34)\n\t2. [Python Tip. Using Generators To Avoid Memory Error](#35)\n\t3. [Texts and Vectors Relations](#36)\n\t4. [Training Using Pandas Piplines](#37)\n\t5. [Classifier Accuracy Comparison On Our Word2Vec Model](#38)\n6. [Embedding Methods_Glove](#39)\n\t1. [Introduction](#40)\n\t2. [Be careful about our method](#41)\n\t3. [GLOVE initiation](#42)\n7. [Embedding Methods_Program_300Vectors](#43)\n8. [Embedding Methods_GoogleNewsVectors](#44)\n9. [Embedding Methods_Wiki_Vectors](#45)\n10. [Conclusion on Traditonal Methods](#46)\n11. [Moving to DeepMoldes](#47)\n    1. [Introduction and Roadmap](#48)\n\t2. [Prepairing Data for Using in Keras](#49)\n12. [References](#100)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"<a id=\"1\"></a> <br>\n#  1-Introduction\n\nIn this competition we are dealing with text processing and embedding vectors. First of all, lets recap some basic and fundamental concepts.\n<a id=\"2\"></a> <br>\n* **A. Concept of Text Mining**\n\nIn principle, text is just another form of data, and text processing is just a special case of representation engineering. In reality, dealing with text requires dedicated pre-processing steps and sometimes specific expertise on the part of the data science In this introduction we can only scratch the surface, to give a basic overview of the techniques and issues involved in typical business applications. First, let’s discuss why text is so important and why it’s difficult.\n<a id=\"3\"></a> <br>\n* **B. Text Mining Importance and Difficulties**\n\nExploiting this vast amount of data requires converting it to a meaningful form. The Internet may be the home of “new media,” but much of it is the same form as old media. It contains a vast amount of text in the form of personal web pages, Quera, Twitter feeds, email, Facebook status updates, product descriptions, Reddit comments, blog postings—the list goes on. Underlying the search engines (Google and Bing) that we use everyday are massive amounts of text-oriented data science. Indeed, the thrust of Web 2.0 was about Internet sites allowing users to interact with one another as a community, and to generate much added content of a site. This user-generated content and interaction usually takes the form of text.\n\n<a id=\"4\"></a> <br>\n* **C. TEXT MINING IN BUSINESS**\n\nIn businesses such as our case study (Quera), understanding customer feedback often requires understanding text. This isn’t always the case; admittedly, some important consumer attitudes are represented explicitly as data or can be inferred through behavior, for example via five-star ratings, click-through patterns, conversion rates, and so on. We can also pay to have data collected and quantified through focus groups and online surveys. But in many cases if we want to “listen to the customer” we’ll actually have to read what she’s written—in product reviews, customer feedback forms, opinion pieces, and email messages.\n<a id=\"5\"></a> <br>\n* **D. DEALING WITH UNSTRUCTURED DATA**\n\nWhy Text Is Difficult Text is often referred to as “unstructured” data. This refers to the fact that text does not have the sort of structure that we normally expect for data: tables of records with fields having fixed meanings (essentially, collections of feature vectors), as well as links between the tables. Words can have varying lengths and text fields can have varying numbers of words.\n<a id=\"6\"></a> <br>\n* **E. DEALING WITH DIRTY SPACE**\n\nSometimes word order matters, sometimes not. As data, text is relatively dirty .People write ungrammatically, they misspell words, they run words together, they abbreviate unpredictably, and punctuate  randomly. Even when text is flawlessly expressed it may contain synonyms (multiple words with the same meaning) and homographs (one spelling shared among multiple words with different meanings).\n<a id=\"7\"></a> <br>\n* **F. WHY REPRESENTATION IS COMPLEX !**\n\nRepresentation Having discussed how difficult text can be, let’s go through the basic steps to transform a body of text into a set of data that can be fed into a data mining algorithm. The general strategy in text mining is to use the simplest (least expensive) or small. A document is composed of individual tokens or terms. For now, think of a token or term as just a word; as we go on we’ll show how they can be different from what are customarily thought of as words. A collection of documents is called a corpus.\n<a id=\"8\"></a> <br>\n* **G. WATH DOES BAG OF WORDS MAENS?!**\n\nBag of Words It is important to keep in mind the purpose of the text representation task. In essence, we are taking a set of documents—each of which is a relatively free-form sequence of words—and turning it into our familiar feature-vector form. Each document is one instance but we don’t know in advance what the features will be. The approach we introduce first is called “bag of words.” As the name implies, the approach is to treat every document as just a collection of individual words.\n<a id=\"9\"></a> <br>\n* **H. IGNORING GRAMMER AND NLP-BASED FEATURES**\n\nBOW approach ignores grammar, word order, sentence structure, and (usually) punctuation. It treats every word in a document as a potentially important keyword of the document. The representation is  straightforward and inexpensive to generate, and tends to work well for many tasks.\n<a id=\"10\"></a> <br>\n* **I. FEATURE EXTRACTION USING BOW**\n\nSo if every word is a possible feature, what will be the feature’s value in a given document? There are several approaches to this. In the most basic approach, each word is a token, and each document is represented by a one (if the token is present in the document) or a zero (the token is not present in the document). This approach simply reduces a document to the set of words contained in it.\n<a id=\"11\"></a> <br>\n* **J. TERM-FREQUENCY (TF)**\n\nThe next step up is to use the word count (frequency) in the document instead of just a zero or one. This allows us to differentiate between how many times a word is used; in some applications, the importance of a term in a document should increase with the number of times that term occurs. This is called the term frequency representation. It could simply defined by the following equation: \n*TF(t) = (Number of times term t appears in a document) / (Total number of terms in the document).*\n<a id=\"12\"></a> <br>\n* **K. INVERSE DOCUMENT FREQUENCY (IDF)**\n\nwhich measures how important a term is. While computing TF, all terms are considered equally important. However it is known that certain terms, such as \"is\", \"of\", and \"that\", may appear a lot of times but have little importance. Thus we need to weigh down the frequent terms while scale up the rare ones, by computing the following: \n*IDF(t) = log_e(Total number of documents / Number of documents with term t in it).*\nidf is on the best methods for Measuring Sparseness\n<a id=\"13\"></a> <br>\n* **L. Example FROM (tfidf.com)**\n\nConsider a document containing 100 words wherein the word cat appears 3 times. The term frequency (i.e., tf) for cat is then (3 / 100) = 0.03.\nNow, assume we have 10 million documents and the word cat appears in one thousand of these. Then, the inverse document frequency (i.e., idf) is\ncalculated as log(10,000,000 / 1,000) = 4. Thus, the Tf-idf weight is the product of these quantities: 0.03 * 4 = 0.12.\n<a id=\"14\"></a> <br>\n* **M. PRE-PROCESSING WHICH ARE CONVENTIONAL (SPECIALLY FOR TRADITIONAL MANNERS).**\n\nFirst, the case has been normalized: every term is in lowercase. This is so that words like Skype and SKYPE are counted as the same thing. Second, many words have been stemmed : their suffixes removed, so that verbs like announces , announced and announcing are all reduced to the term announce. Finally, stopwords have been removed. A stopword is a very common word in English (or whatever language is being parsed). The words the , and , of , and on are considered stopwords in English so they are typically removed.\n\n<a id=\"15\"></a> <br>\n#  2-EMBEDDING METHODS\n<a id=\"16\"></a> <br>\n* **A. WHAT DOES EMBEDDING MEANS ?!**\n\nAn embedding is a relatively low-dimensional space into which you can translate high-dimensional vectors. Embeddings make it easier to do machine learning on large inputs like sparse vectors representing words. Ideally, an embedding captures some of the semantics of the input by placing semantically similar inputs close together in the embedding space. An embedding can be learned and reused across models.\n<a id=\"17\"></a> <br>\n* **B. What Does N-GRAM Means?!**\n\nAs presented, the bag-of-words representation treats every individual word as a term, discarding word order entirely. In some cases, word order is important and you want to preserve some information about it in the representation. A next step up in complexity is to include sequences of adjacent words as terms. For example, we could include pairs of adjacent words so that if a document contained the sentence “The quick brown fox jumps.” it would be transformed into the set of its constitutent words { quick , brown , fox , jumps }, plus the tokens quick_brown , brown_fox , and fox_jumps . This general representation tactic is called n-grams . Adjacent pairs are commonly called bi-grams. If you hear a data scientist mention representing text as “bag of n-grams up to three” it simply means she’s representing each document using as\nfeatures its individual words, adjacent word pairs, and adjacent word triples. N-grams are useful when particular phrases are significant but their component words may not be. In a business news story, the appearance of the tri-gram exceed_analyst_expectation is more meaningful than simply knowing that the individual words analyst , expectation , and exceed appeared somewhere in a story. An advantage of using n-grams is that they are easy to generate; they require no linguistic knowledge or complex parsing algorithm.\n<a id=\"18\"></a> <br>\n* **C. MAIN DISADVANTAGE OF N-GRAM**\n\nThe main disadvantage of n-grams is that they greatly increase the size of the feature set. There are far more word pairs than individual words, and still more word triples. The number of features generated can quickly get out of hand. Data mining using n-grams almost always needs some special consideration for dealing with massive numbers of features, such as a feature selection stage or special consideration to computational storage space. \n<a id=\"19\"></a> <br>\n* **D. WHAT DOES EMBEDDING PROVIDE FOR US IN TEXT MINING ?!**\n\nIn the most simplest way, we can say embedding means replacing each word of our corpus with the correspondence feature space that pre-trained embedding models have provided. Some organizations have trained a deep learning models on popular datasets such as wikipedia. If you want to know how they have train the models check [this url](https://radimrehurek.com/gensim/tut1.html#from-strings-to-vectors). Training a deep learning model needs process and time. So, we cant do it by our personal laptops. Fortunately, some of these pretrained embeddings (i.e models which are existed in data directory of this competition) are published with opensource licences. We need only convert our text to the target feature spaces and do a classification task with wide range of classifier we have known.\n\n<a id=\"20\"></a> <br>\n#  3-Take a Glance To Data\n\nOk. lets do simple WordCloud for easily getting start ... "},{"metadata":{"trusted":true,"_uuid":"30e16ea6848c93565f236baef27904fe01205333"},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom wordcloud import WordCloud\nimport warnings\nimport collections\nimport sklearn as sklearn\nfrom sklearn.feature_extraction.text import CountVectorizer, TfidfTransformer,TfidfVectorizer\nfrom sklearn.model_selection import cross_val_score, KFold, train_test_split\nfrom sklearn.feature_extraction import text \nfrom sklearn.naive_bayes import MultinomialNB\nfrom sklearn.ensemble import RandomForestClassifier,VotingClassifier\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn import svm, tree\nfrom sklearn.svm import LinearSVC\nfrom sklearn.metrics import accuracy_score\nimport gensim\nfrom collections import defaultdict\nfrom itertools import islice\nfrom sklearn.ensemble import ExtraTreesClassifier\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.naive_bayes import GaussianNB\nfrom keras import Sequential\nfrom keras.layers import Bidirectional, GlobalMaxPool1D,Dense, Input, Embedding, Dropout, LSTM, CuDNNGRU\nfrom keras.models import Model\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.preprocessing.text import Tokenizer\n\n\nimport time\n%matplotlib inline\nwarnings.filterwarnings('ignore')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7484ec6027e908b9bf685f8a7d1f534bd056405e"},"cell_type":"code","source":"train_df = pd.DataFrame.from_csv(\"../input/train.csv\")\ntest_df =  pd.DataFrame.from_csv(\"../input/test.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"17f90343e5fa93152d1dbc27dfb9138480ffe966"},"cell_type":"code","source":"train_df.head(5)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fc1da97db569c7f7c1bac0fea64472d33e0a7581"},"cell_type":"markdown","source":"<a id=\"21\"></a> <br>\n**A. TEST WORDCLOUD**"},{"metadata":{"trusted":true,"_uuid":"b91a34214a407cb858d757a56d7a9a9e529c66b4"},"cell_type":"code","source":"text =\" \".join(train_df.question_text)\n # Create the wordcloud object\nwordcloud = WordCloud(width=1024, height=1024, margin=0).generate(text)\n \n# Display the generated image:\nfig,ax = plt.subplots(1,1,figsize=(10,10))\nax.imshow(wordcloud, interpolation='bilinear')\nax.axis(\"off\")\nax.margins(x=0, y=0)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1168f880cb5bfb70b671cc8661bd7e966a690b23"},"cell_type":"markdown","source":"<a id=\"22\"></a> <br>\n**B. TARGET DISTRIBUTION**"},{"metadata":{"trusted":true,"_uuid":"fbcace9fe350b2d3da2edfe80955c83423653c2f"},"cell_type":"code","source":"fig, ax = plt.subplots(1,1, figsize=(8,8))\nax.set_title(\"Target Status\")\nexplode=(0,0.1)\nlabels ='0','1'\nax.pie(list(dict(collections.Counter(list(train_df.target))).values()), explode=explode, labels=labels, autopct='%1.1f%%',shadow=True, startangle=90)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"929658291f04f2145bbd10040dee7e41b63aab7a"},"cell_type":"code","source":"del text","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"21f5cfa4e976aec619c36f1604c37a21dafeb2eb"},"cell_type":"markdown","source":"<a id=\"23\"></a> <br>\n#  4-Start Vectorizing by Simple CountVectorizer"},{"metadata":{"_uuid":"0113951cbdc519d99604287cd13eae5ed34fb045"},"cell_type":"markdown","source":"<a id=\"24\"></a> <br>\n**A. MAKE YOUR HANDS DIRTY ;-) **\n\nStarting with CountVectorizer.\n\nAs we mentioned in introduction, for converting string to feature spaces, we need do some transforms on string. As a result of this sort of transonsformations, the string will be converted to vectors space. One of teh most simplest ways for vectorizing the text and converting it to a formal dataset (which could be passed to classification algorithms) is countVectorizer. Simply, the count of each word will be calculated. In the second step, by using TermFrequency transformations, for any word and for any document, the relative proportions will be form a formal dataset.\n\nLets start with a countVectorizer..."},{"metadata":{"trusted":true,"_uuid":"cc98c768d1a6930c20285b3a2772e373620d68d6"},"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(train_df.question_text, train_df.target, test_size=0.33, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"92aaae58fbc868f2aa694912b92ad4efb7ad516c"},"cell_type":"code","source":"#initiation of countVectorizer\ncount_vectorizer = CountVectorizer(min_df=5, stop_words='english')\nvect = count_vectorizer.fit(train_df.question_text)\n\nX_vect_train = vect.transform(X_train) # documents-terms matrix of training set\nX_vect_test = vect.transform(X_test) # documents-terms matrix of testing set\n\ntf_train_transformer = TfidfTransformer(use_idf=False).fit(X_vect_train)\ntf_test_transformer =  TfidfTransformer(use_idf=False).fit(X_vect_test)\n\nxtrain_tf = tf_train_transformer.transform(X_vect_train)\nxtest_tf = tf_test_transformer.transform(X_vect_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"65af564dd47baddf242f4b7fda24cad89ed3c291"},"cell_type":"code","source":"type(xtrain_tf),xtrain_tf.shape, train_df.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"241a025c29baa3206c0fbea163ef03d1bb92f663"},"cell_type":"markdown","source":"Because of high dimentional problems we have encountered with, by default it returns the matrix in sparse mode. For representing in normal feature space, you can use todense method."},{"metadata":{"trusted":true,"_uuid":"6a512cf4f4a52c6e7202992435b8b7f92b25cbc7"},"cell_type":"code","source":"xtrain_tf[:5].todense()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3c98d7f2b19d4ba5ddf55be431e3ff16d64d8407"},"cell_type":"markdown","source":"Any of these rows, reveals the vector of each document. The columns are also correspondence to words existed in the dataset. You can access the feature names by 'get_feature_names()' method.\nLets look around some sections of features."},{"metadata":{"trusted":true,"_uuid":"79ed531b9b108e033d5b41d49cf8b7ddcfc406a5"},"cell_type":"code","source":"count_vectorizer.get_feature_names()[0:5]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a7b4b6758a7ca252a289fed1293037e087077dda"},"cell_type":"code","source":"count_vectorizer.get_feature_names()[30000:30005]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"37671e2cf96ddd86f5d6bd8a57470bc912ac21ee"},"cell_type":"code","source":"count_vectorizer.get_feature_names()[-5:]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"68d8877acd64fa96bc7a63b254e5753fda805d98"},"cell_type":"markdown","source":"There are also non-english featuers. These sort of features can reduce the accuracy."},{"metadata":{"trusted":true,"_uuid":"64341dc448b8969a23123d30b88867fd0d1e04f0"},"cell_type":"code","source":"[{k:count_vectorizer.vocabulary_[k]} for k in list(count_vectorizer.vocabulary_)[:10]]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cc27218c505490e0c7a3f16f0ceea37fe775aa4e"},"cell_type":"code","source":"list(sklearn.feature_extraction.text.ENGLISH_STOP_WORDS)[:10]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7422a655a179f8e8563f27d110f1bfb03ca9a862"},"cell_type":"markdown","source":"<a id=\"25\"></a> <br>\n**B. SOME SUGGESTIONS ABOUT TEXT VECTORIZING  **\n\nOk. Good.\n\nLets review some facts about text processing in this special case.\n\nAs it is obvious you can see there are various range of features (from foreign language to astonishing digits and numbers) in feature space.\nWe can do some modifications on the count_vectorizer. Some preprocessing could be done are:\n\n* Cleaning meaningless words from dataset.\n\n* Using threshold for repetition of words. for example discarding the words which are not repeated in whole records. it  could be setted by min_df in count_vectorizer\n\n* Defining threshold for minimum repetition of words. for example any word that has repetitioin fewer than 10 times, should be ignored.\n\n* Removing Stopwords. Stop words are common words that have low information for classification because they exist on any sentences. For example, words such as 'the','is', 'by' and this sort of words, have belonged to stop_words list. Removing them from dataset can increase the performance accuracy. Count_vectorizer supports using stop words by defatult. You can use it by CountVectorizer(stop_words=LIST_OF_STOP_WORDS). It is wise to see improvements in classifier performance by adding effect of stop_words to it."},{"metadata":{"_uuid":"f6fa1e1bc2abfb313bff0f9003912fb91f6b9912"},"cell_type":"markdown","source":"Now, We have the dataset which is ready for classification.\nLets do simple classifications for starting classification on first vectorized space (count vectorize)."},{"metadata":{"_uuid":"01818b3dfcd0490535ac28d1e8f0f9a6a66e0571"},"cell_type":"markdown","source":"<a id=\"26\"></a> <br>\n**C. CLASSIFICATION**"},{"metadata":{"trusted":true,"_uuid":"874a5f20b6214636deaa3d344b8fa0428b7e929d"},"cell_type":"code","source":"results_df = pd.DataFrame()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f9087b8869cf6fb931928cc3080f02775b3224ee"},"cell_type":"markdown","source":"<a id=\"27\"></a> <br>\n**D. Multinominal NB**"},{"metadata":{"trusted":true,"_uuid":"be6ad3411abf57023c0eebfe5dc4985d9fe72eef"},"cell_type":"code","source":"# MULTINOMINA_NAIVE_BAYES\nnb_ = MultinomialNB()\nnb_clf = nb_.fit(X=xtrain_tf, y=y_train)\nresults_df.set_value(\"NB\" , \"countVectorizer\" , accuracy_score(y_test,nb_clf.predict(xtest_tf)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c073aff6ae27da7127df39519355952838a7e2c1"},"cell_type":"markdown","source":"<a id=\"28\"></a> <br>\n**E. RF**"},{"metadata":{"trusted":true,"_uuid":"cb1d982ad33a894dcae3eca6ab7d912b0d1856a6","scrolled":false},"cell_type":"code","source":"# RANDOM_FORES_CLASSFIER\nrf_clf = RandomForestClassifier(n_estimators=25, max_depth=15,random_state=42)\nrf_clf.fit(X=xtrain_tf,y=y_train)\nresults_df.set_value(\"RF\" , \"countVectorizer\" , accuracy_score(y_test,rf_clf.predict(xtest_tf)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"64b3fba3a13816873b14c226d31b92ed0e613557"},"cell_type":"markdown","source":"<a id=\"29\"></a> <br>\n**F. MLP**"},{"metadata":{"trusted":true,"_uuid":"62b52f025042c0a255cbc267612f81d4aa5034c0"},"cell_type":"code","source":"#MLP_CLASSIFIER\nmlp_clf = MLPClassifier(solver='lbfgs', alpha=1e-4,hidden_layer_sizes=(20,10, 2), random_state=42)\nmlp_clf.fit(X=xtrain_tf, y=y_train)                         \nresults_df.set_value(\"MLP\" , \"countVectorizer\" , accuracy_score(y_test,mlp_clf.predict(xtest_tf)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a7bbba056618608d0155d65f2de30723edfc7a62"},"cell_type":"markdown","source":"<a id=\"30\"></a> <br>\n**G. LoggicReg**"},{"metadata":{"trusted":true,"_uuid":"46897c66dbe726205c4ae38240f8674d4841afec"},"cell_type":"code","source":"#LoggicRegression\nlreg_clf = LogisticRegression(solver='lbfgs', multi_class='multinomial',random_state=42)\nlreg_clf.fit(X=xtrain_tf, y=y_train)                         \nresults_df.set_value(\"LREG\" , \"countVectorizer\" , accuracy_score(y_test,lreg_clf.predict(xtest_tf)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0ccca7cb88f1b590de8ecceef07e603df2254a45"},"cell_type":"markdown","source":"<a id=\"31\"></a> <br>\n**H. CLASSIFIERS ACCURACY COMPARISON ON COUNT_VECTORIZER**\n\nCongratulation ! \n\nYou have tested large variety of classifiers.\nLets visualize the result of tested classifiers."},{"metadata":{"trusted":true,"_uuid":"84d76e0722a46a7d993985db3a53144fb2aadcf5"},"cell_type":"code","source":"results_df","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"016a9a65aebe675d2b675e367a1a2c2723d6ee13","trusted":true},"cell_type":"code","source":"fig,axes=plt.subplots(1,1,figsize=(8,8))\naxes.set_ylabel(\"Accuracy\")\nplt.ylim((.92,.97))\nresults_df.plot(kind=\"bar\",ax=axes)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3a18703924ec2e54c69a3a5c09343976cfcbc5ac"},"cell_type":"markdown","source":"Don't forget we have used limit number of layers and neurons. Results absolutely could be better by tuning parameters in more precise way.\n\nRandomForest also could revealed better results if it trains with more decision trees.\n\n**Simply, we can improve accuracy of these classifiers with more supervised training.**"},{"metadata":{"_uuid":"f1760a25a79298d04a837a4bb0fff32f817aa7cf"},"cell_type":"markdown","source":"<a id=\"32\"></a> <br>\n**I. CURSE OF DIMENTIONALITY**\n\n\nLook again at the dimension of train data. It has 49650 features !!! Dealing with such a huge feature space is really challenging.\nIf you do classification (similar to what we have done in previous part) you will get MemoryError. \nOne of the most challenging factors in data mining is dimensionality. Read about curse of dimensionality [here](https://en.wikipedia.org/wiki/Curse_of_dimensionality) if you want to know more about it.\n\nFor our example we have 49650 features. Absolutely classifiers such as DecisionTree, KNeighborsClassifier or SVM will return MemoryError.  \n"},{"metadata":{"_uuid":"1621310929eca4893ac5a9cff9171a7110c84ce7"},"cell_type":"markdown","source":"<a id=\"33\"></a> <br>\n#  5-EMBEDDING METHODS_WORD2VEC\n\n<a id=\"34\"></a> <br>\n**A. INTRODUCTION**\n\nNow you know how to deal with texts. At first step, we need to extract features from text; Then there is a need to investigate classifiers and tune them for getting better accuracy.  In previous sections you trained a classifier based on the count vectorizer. Now, lets move on to newer vectorizing method. The name is Word2Vec. As its name is simply representing, by using this manner any word will be replaced by a correspondence vectors. In this manner, training parameters are playing important role. lets try word2vec on our dataset to understand what does it means.\n\n**As these processes are time consuming, I have simulated these classifiers on my local machine. Committing these codes (In each commit waiting until finishing whole the  classifiers was really annoying for me) so I have commented the lines which are representing training process. You can easily uncomment them and try to  train classifier by yourself ;-). **\n\n<a id=\"35\"></a> <br>\n**B. PYTHON TIP. USE GENERATORS TO AVOID MEMORY ERROR**\n\nTo train a model on text, we need pass following steps: \n Reading text and keeping in memory,\n Tokenizing the text\n Iterating on specified window\n \n In use cases such as texts (especially large texts) keeping whole the text in memory is memory consumable. Using generators, can reduce the overhead of\n iterating. Python generators are a simple way of creating iterators. All the overhead we mentioned above are automatically handled by generators in\n Python. Simply speaking, a generator is a function that returns an object (iterator) which we can iterate over (one value at a time).\n"},{"metadata":{"trusted":true,"_uuid":"962e3efc2bc39b230fd85754b3cff0c4c24d1058"},"cell_type":"code","source":"text_ = train_df.question_text\ntargets_ = train_df.target\nclass GetSentences(object):\n    def __iter__(self):\n        counter = 0\n        for sentence_iter in text_:\n            tmp_sentence = sentence_iter\n            counter += 1\n            yield tmp_sentence.split()\nlen(text_)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"14481fcbe129d8cd5054b37ff7b69221a38413fd"},"cell_type":"markdown","source":"Training Word2Vec. \nWe will ignore words which have repetition less than 5 times and scrolling default window size equal to 5. Changing these parameters can change our model accuracy. We will check it in next steps."},{"metadata":{"trusted":true,"_uuid":"866674d87821edbb17af32d993e5895873ddd691"},"cell_type":"code","source":"num_features = 200  # Word vector dimensionality\nmin_word_count = 5  # Minimum word count\nnum_workers = 4  # Number of threads to run in parallel\ncontext = 10  # Context window size\ndownsampling = 1e-3  # Downsample setting for frequent words\nget_sentence = GetSentences()\nmodel = gensim.models.Word2Vec(sentences=get_sentence, min_count=min_word_count, size=num_features, workers=4)\nw2v = dict(zip(model.wv.index2word, model.wv.syn0))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"90c10cecd153ffa0c85e963f45e1bd0426cce03f","scrolled":true},"cell_type":"code","source":"next(iter(w2v.items()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"78c3259dec5d25b28e60a048efba354520e0fa1e"},"cell_type":"code","source":"list(islice(model.wv.vocab.items(), 5))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"495b3103c0edcc150c2bedc45e4963d53e411b9b"},"cell_type":"markdown","source":"As you can see, we have a vector for word 'the'. It means we can replace 'the' with correspondence vector. Lets do some interesting calculations with our word2vec model ;-)"},{"metadata":{"_uuid":"c78b73f2724b84da1909c322e181f812e6655fb5"},"cell_type":"markdown","source":"<a id=\"36\"></a> <br>\n**C. TEXTS and VECTORS RELATIONS**"},{"metadata":{"_uuid":"16ffe7251b0476943c0b44d4abb7b7cddd78d18f"},"cell_type":"markdown","source":"    **-. MOST SIMILAR WORDS TO ... **"},{"metadata":{"trusted":true,"_uuid":"1e50c056b8a9afd0f3e689e959f2a5361ee312f6"},"cell_type":"code","source":"\"MOST SIMILAR WORD TO MAN: {} AND MOST SIMILAR WORD TO QUERA:{} \".format(model.most_similar(['man'],topn=1),model.most_similar([\"quora\"],topn=3))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6bad322f14b0e769096478a990659bed77804f98"},"cell_type":"code","source":"model.most_similar(positive=[\"united\",\"state\"])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e6956bafe0aeacfa7c68e32b3d8111f330428ca5"},"cell_type":"markdown","source":"Or in more related way to competition challenges..."},{"metadata":{"trusted":true,"_uuid":"1bfce6c60120c7846266aa6521204b84be093f56"},"cell_type":"code","source":"# most similar words to 'sex'\nmodel.most_similar(positive=[\"sex\"])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"577b39a553190d47562456648508d705364434ad"},"cell_type":"code","source":"model.most_similar(positive=['gangbang', 'sex'], negative=['man'], topn=5) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b6ad0d0a6ecf9a0793186e870841f17ee374dbc6"},"cell_type":"markdown","source":"    **-.  [ X+Y ]  - [ Z ] ... **"},{"metadata":{"trusted":true,"_uuid":"ff7364685fa31dbd3d9df723f9fd6f377a485431"},"cell_type":"code","source":"model.most_similar(positive=['woman', 'king'], negative=['man'], topn=2) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"225ab064d32c95de7dff72fb1ee0a8b3ae2e2cf1"},"cell_type":"markdown","source":"As you can see there is vector for each word. The better trained models the higher accuracy we can get. Now you can understand what does the embedding models can do for us."},{"metadata":{"_uuid":"9727698ae407908cafd7644b0be426ff1f0596fc"},"cell_type":"markdown","source":"    **-. SAVING THE MODEL **"},{"metadata":{"trusted":true,"_uuid":"e4e2d501c560a6e7167992c485cb21e68443fc14"},"cell_type":"code","source":"model.save('word2vec.model')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1a3787f7f25e8e8b0bb879d1ee45b43e846f4b78"},"cell_type":"markdown","source":"<a id=\"37\"></a> <br>\n**D. TRAINING USING PANDAS PIPLINES**"},{"metadata":{"_uuid":"526d9999f49c62f4b0885fb512e9fa8628739f8c"},"cell_type":"markdown","source":"In this section we will use pipline facilities in training the data.\n\nAbout the pipelines there are two points should be mentioned. \n\nFrom the documentation:\n1. Pipelines can Sequentially apply a list of transforms and a final estimator. Intermediate steps of the pipeline must be ‘transforms’, that is, they must implement fit and transform methods. The final estimator only needs to implement fit.\n2. The purpose of the pipeline is to assemble several steps that can be cross-validated together while setting different parameters. \n"},{"metadata":{"_uuid":"88e2e5da10faf9d75b03bb237ff400ef5589f21d"},"cell_type":"markdown","source":"Until now, we have vector for each word. How we can convert this vectors to formal dataset which could be passed to traning algorithms ?!\n\nSimply manner could be mean of vectors in any dataset record. Another manner could be the same tfidf transformers. \n\nLets implement them for using in our pipeline.\n\nLets split Train and Test for training the classfieirs."},{"metadata":{"trusted":true,"_uuid":"b992752a6a851439acf42d68cbaae4c4715ababd"},"cell_type":"code","source":"class MeanEmbeddingVectorizer(object):\n    def __init__(self, word2vec):\n        self.word2vec = word2vec\n        # if a text is empty we should return a vector of zeros\n        # with the same dimensionality as all the other vectors\n        # self.dim = len(word2vec.itervalues().next())\n        self.dim = len(next(iter(self.word2vec.items()))[1])\n\n    def fit(self, X, y):\n        return self\n\n    def transform(self, X):\n        return np.array([np.mean([self.word2vec[w] for w in words if w in self.word2vec] or [np.zeros(self.dim)], axis=0) for words in X])\n\n\nclass TfidfEmbeddingVectorizer(object):\n    def __init__(self, word2vec):\n        self.word2vec = word2vec\n        self.word2weight = None\n        # self.dim = len(word2vec.itervalues().next())\n        self.dim = len(next(iter((word2vec.items()))))\n\n    def fit(self, X, y):\n        tfidf = TfidfVectorizer(analyzer=lambda x: x)\n        tfidf.fit(X)\n        # if a word was never seen - it must be at least as infrequent\n        # as any of the known words - so the default idf is the max of\n        # known idf's\n        max_idf = max(tfidf.idf_)\n        self.word2weight = defaultdict(\n            lambda: max_idf,\n            [(w, tfidf.idf_[i]) for w, i in tfidf.vocabulary_.items()])\n\n        return self\n\n    def transform(self, X):\n        return np.array([\n            np.mean([self.word2vec[w] * self.word2weight[w]\n                     for w in words if w in self.word2vec] or\n                    [np.zeros(self.dim)], axis=0)\n            for words in X\n        ])\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2a21034bf9f1354f62066dd79314f18f4d2572c1"},"cell_type":"code","source":"\"\"\"\nNOTE:\nAs these processes are time consuming, I have simulated these classifiers on my local machine. Committing these codes (In each commit waiting until finishing whole the \nclassifiers was really annoying for me) so I have commented the lines which are representing training process. You can easily uncomment them and try to\ntrain classifier by yourself ;-).\n\"\"\"\n\nprint(\"EXTRA TREE ...\")\n# 1- MEAN VECTORIZER\netree_w2v = Pipeline([(\"w2v mean vectorizer\", MeanEmbeddingVectorizer(w2v)), (\"extra trees\", ExtraTreesClassifier(n_estimators=25))])\n# etree_w2v.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"w2v_mean\", accuracy_score(y_test, etree_w2v.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\netree_w2v_tfidf = Pipeline([(\"w2v tfidf vectorizer\", TfidfEmbeddingVectorizer(w2v)), (\"extra trees\", ExtraTreesClassifier(n_estimators=25))])\n# etree_w2v_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"w2v_tfidf\", accuracy_score(y_test, etree_w2v_tfidf.predict(X_test)))\n\n####SVM####\nprint(\"SVM ... \")\n# 1- MAIN VECTORIZER\nsvm_w2v = Pipeline([(\"w2v mean vectorizer\", MeanEmbeddingVectorizer(w2v)), (\"SVM\", LinearSVC(random_state=0, tol=1e-4))])\n# svm_w2v.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"w2v_mean\", accuracy_score(y_test, etree_w2v_tfidf.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nsvm_w2v_tfidf = Pipeline([(\"word2vec vectorizer\", TfidfEmbeddingVectorizer(w2v)), (\"SVM\", LinearSVC(random_state=0, tol=1e-4))])\n# svm_w2v_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"w2v_tfidf\", accuracy_score(y_test, svm_w2v_tfidf.predict(X_test)))\n\n####MLP####\nprint(\"MLP ... \")\n# 1- MAIN VECTORIZER\nmlp_w2v = Pipeline(\n    [(\"w2v mean vectorizer\", MeanEmbeddingVectorizer(w2v)),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(20,10, 2), random_state=1))])\n# mlp_w2v.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"w2v_mean\", accuracy_score(y_test, mlp_w2v.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nmlp_w2v_tfidf = Pipeline(\n    [(\"word2vec vectorizer\", TfidfEmbeddingVectorizer(w2v)),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(20,10,2), random_state=1))])\n# mlp_w2v_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"w2v_tfidf\", accuracy_score(y_test, mlp_w2v_tfidf.predict(X_test)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"900ff0dba7b202f227b337fae6f640aec5dd20ba","_kg_hide-input":true},"cell_type":"code","source":"results_df.set_value(\"NB\",\"w2v_mean\",0.476575)\nresults_df.set_value(\"ExtraTree\",\"w2v_mean\",0.939119)\nresults_df.set_value(\"SVM\",\"w2v_mean\",0.939144)\nresults_df.set_value(\"MLP\",\"w2v_mean\",0.939035)\nresults_df.set_value(\"LREG\",\"w2v_mean\",0.938836)\n\nresults_df.set_value(\"NB\",\"w2v_tfidf\",0.260479)\nresults_df.set_value(\"ExtraTree\",\"w2v_tfidf\",0.939144)\nresults_df.set_value(\"SVM\",\"w2v_tfidf\",0.939033)\nresults_df.set_value(\"MLP\",\"w2v_tfidf\",0.939035)\nresults_df.set_value(\"LREG\",\"w2v_tfidf\",0.953311)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e5d3fb634a1d5693410fc2ba6bdb44b6d1ecb586"},"cell_type":"markdown","source":"<a id=\"38\"></a> <br>\n**E. CLASSIFIERS ACCURACY COMPARISON ON OUR WORD2VEC MODEL**\n"},{"metadata":{"trusted":true,"_uuid":"d3883b943db23eeaee9b210a1101b3eca94e6ea6"},"cell_type":"code","source":"results_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b314cba4b55b0b7b79580e41cf8101806bc24a6d"},"cell_type":"code","source":"fig,axes=plt.subplots(1,1,figsize=(8,8))\naxes.set_ylabel(\"Accuracy\")\naxes.set_title(\"word2vec results for 67% training and 23% testing\")\n# plt.ylim((.93,.95))\nresults_df[[\"w2v_mean\",\"w2v_tfidf\"]].dropna().plot(kind=\"bar\",ax=axes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c1a8f40ff33887732068fedd9c5061444823cecd"},"cell_type":"code","source":"del w2v\ndel model\nimport gc; gc.collect()\ntime.sleep(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"44ca32295fa53fb3bfbce5397d8b6903587cb356"},"cell_type":"markdown","source":"<a id=\"39\"></a> <br>\n#  6-EMBEDDING METHODS_GLOVE\n\n<a id=\"40\"></a> <br>\n**A. INTRODUCTION**\n\nIn prevous section we represented the results of own word2vec model. Now, kets do a classification with glov embedded models.\nIn the most simplest way, we can say for every embedding models we have a dictionary containing key (word) and values (list of double values which are correspondence to relation to other words which are exist on that dataset).\n\n\n<a id=\"41\"></a> <br>\n**B. BE CAREFUL ABOUT OUR MANNER.**\n\n**The point should be mentioned is that because of time and memory limitations, we have limit our training data to 67 percent of samples. The most reason for the low accuracies and a gap which is existed in most of the classifiers is related to type of our selection. In most of the classifiers, we also limit the parameters. for example in ExtraTree we use only 25 classifiers which is absolutely low. Or in the neural network MLP we have used two layer perceptron model which is approximately unsatisfied model. By the way, lets continue the process. We will try to extract all the information which is needed and use them in final step.**\n"},{"metadata":{"_uuid":"08efc6e7e4e4c409b7b99f3b0fea79f9b95748f6"},"cell_type":"markdown","source":"<a id=\"42\"></a> <br>\n**C. GLOVE INITIATIONS.**\n"},{"metadata":{"trusted":true,"_uuid":"12d0a48ed587d615b7ddda68f2831dc85ecfbf8e"},"cell_type":"code","source":"EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\nglov = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"a15b9848768ea2cd39ff26c623bfdd4155420ed4"},"cell_type":"code","source":"next(iter(glov.values()))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"761c0137c86163286b4eae3d3aff9cb5963dbf79"},"cell_type":"markdown","source":"As you can see, the process is completely similar to the previous section."},{"metadata":{"trusted":true,"_uuid":"fd5d100804784542989c1335970d6854b0b7d433"},"cell_type":"code","source":"\"\"\"\nNOTE:\nAs these processes are time consuming, I have simulated them on my local machine. Committing these codes (Waiting until finishing whole the trainin process for\nany commit really bothers me) so I have commented the lines which have belonged to training process. You can easily uncomment them and try to\ntrain classifier by yourself ;-).\n\"\"\"\n# GLOVE MODEL\n\n###NB####\nprint(\"Multinomina NB ...\")\nnb_glov = Pipeline([(\"glov mean vectorizer\", MeanEmbeddingVectorizer(glov)), (\"Guassian NB\", GaussianNB())])\n# nb_glov.fit(X=X_train, y=y_train)\n# results_df.set_value(\"GB\", \"glov_mean\", accuracy_score(y_test, nb_glov.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nnb_glov_tfidf = Pipeline([(\"glov tfidf vectorizer\", TfidfVectorizer(glov)), (\"transform\", TfidfTransformer()), (\"Guassian NB\", MultinomialNB())])\n# nb_glov_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"GB\", \"glov_tfidf\", accuracy_score(y_test, nb_glov_tfidf.predict(X_test)))\n\n###EXTRA TREE####\nprint(\"EXTRA TREE ...\")\n# 1- MEAN VECTORIZER\netree_glov = Pipeline([(\"glov mean vectorizer\", MeanEmbeddingVectorizer(glov)), (\"extra trees\", ExtraTreesClassifier(n_estimators=25))])\n# etree_glov.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"glov_mean\", accuracy_score(y_test, etree_glov.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\netree_glov_tfidf = Pipeline(\n    [(\"glov tfidf vectorizer\", TfidfVectorizer(glov)), (\"transform\", TfidfTransformer()), (\"extra trees\", ExtraTreesClassifier(n_estimators=25))])\n# etree_glov_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"glov_tfidf\", accuracy_score(y_test, etree_glov_tfidf.predict(X_test)))\n\n####SVM####\nprint(\"SVM ... \")\n# 1- MAIN VECTORIZER\nsvm_glov = Pipeline([(\"glov mean vectorizer\", MeanEmbeddingVectorizer(glov)), (\"SVM\", LinearSVC(random_state=42, tol=1e-5))])\n# svm_glov.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"glov_mean\", accuracy_score(y_test, svm_glov.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nsvm_glov_tfidf = Pipeline(\n    [(\"glov tfidf vectorizer\", TfidfVectorizer(glov)), (\"transform\", TfidfTransformer()), (\"SVM\", LinearSVC(random_state=0, tol=1e-5))])\n# svm_glov_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"glov_tfidf\", accuracy_score(y_test, svm_glov_tfidf.predict(X_test)))\n\n####MLP####\nprint(\"MLP ... \")\n# 1- MAIN VECTORIZER\nmlp_glov = Pipeline(\n    [(\"glov mean vectorizer\", MeanEmbeddingVectorizer(glov)),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(20, 10, 2), random_state=42))])\n# mlp_glov.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"glov_mean\", accuracy_score(y_test, mlp_glov.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nmlp_glov_tfidf = Pipeline(\n    [(\"glov tfidf vectorizer\", TfidfVectorizer(glov)), (\"transform\", TfidfTransformer()),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(20, 10, 2), random_state=42))])\n# mlp_glov_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"glov_tfidf\", accuracy_score(y_test, mlp_glov_tfidf.predict(X_test)))\n\n\n####LREG####\nprint(\"LREG ... \")\n# 1- MAIN VECTORIZER\nlreg_glov = Pipeline(\n    [(\"glov mean vectorizer\", MeanEmbeddingVectorizer(glov)),\n     (\"LREG\", LogisticRegression(solver='lbfgs', multi_class='multinomial', random_state=42))])\n# lreg_glov.fit(X=X_train, y=y_train)\n# results_df.set_value(\"LREG\", \"glov_mean\", accuracy_score(y_test, lreg_glov.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nlreg_glov_tfidf = Pipeline(\n    [(\"glov tfidf vectorizer\", TfidfVectorizer(glov)), (\"transform\", TfidfTransformer()),\n     (\"LREG\", LogisticRegression(solver='lbfgs', multi_class='multinomial', random_state=42))])\n# lreg_glov_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"LEREG\", \"glov_tfidf\", accuracy_score(y_test, lreg_glov_tfidf.predict(X_test)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5d89515a54f237d08d901a7286232d86a8b2a8d9","_kg_hide-input":true},"cell_type":"code","source":"results_df.set_value(\"NB\",\"glov_mean\",0.533364)\nresults_df.set_value(\"ExtraTree\",\"glov_mean\",0.939131)\nresults_df.set_value(\"SVM\",\"glov_mean\",0.939140)\nresults_df.set_value(\"MLP\",\"glov_mean\",0.938973)\nresults_df.set_value(\"LREG\",\"glov_mean\",0.938836)\n\nresults_df.set_value(\"NB\",\"glov_tfidf\",0.941720)\nresults_df.set_value(\"ExtraTree\",\"glov_tfidf\",0.945798)\nresults_df.set_value(\"SVM\",\"glov_tfidf\",0.953854)\nresults_df.set_value(\"MLP\",\"glov_tfidf\",0.939035)\nresults_df.set_value(\"LREG\",\"glov_tfidf\",0.953311)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fcd099c59f5ffe4832267f3ef010ac1922c7320c"},"cell_type":"code","source":"# fig,axes=plt.subplots(1,1,figsize=(15,8))\n# axes.set_ylabel(\"Accuracy\")\n# axes.set_title(\"GLOVE results for 67% training and 23% testing\")\n# results_df.plot(kind=\"bar\",ax=axes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8e880d0d4a61c81982f6a411ed32b776221b3617"},"cell_type":"code","source":"del glov\nimport gc; gc.collect()\ntime.sleep(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a6def61b7e77b8c37b9527d5d2eb8645f9f467fe"},"cell_type":"markdown","source":"<a id=\"43\"></a> <br>\n#  7-EMBEDDING METHODS_PROGRAM_300"},{"metadata":{"trusted":true,"_uuid":"e399785051b7c5dd5804110a0c53cf0c44ddc6d1"},"cell_type":"code","source":"EMBEDDING_FILE = '../input/embeddings/paragram_300_sl999/paragram_300_sl999.txt'\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\nprogram = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE,encoding=\"latin1\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5086a34e07db8d6a3e05813e1aa0fb3d950ff7b4"},"cell_type":"code","source":"next(iter(program.values()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"98cd30160962ef586a4bf78de707453e438643e4"},"cell_type":"code","source":"\"\"\"\nNOTE:\nAs these processes are time consuming, I have simulated them on my local machine. Committing these codes (Waiting until finishing whole the trainin process for\nany commit really bothers me) so I have commented the lines which have belonged to training process. You can easily uncomment them and try to\ntrain classifier by yourself ;-).\n\"\"\"\n\nprint(\"Multinomina NB ...\")\nnb_program = Pipeline([(\"program mean vectorizer\", MeanEmbeddingVectorizer(program)), (\"Guassian NB\", GaussianNB())])\n# nb_program.fit(X=X_train, y=y_train)\n# results_df.set_value(\"GB\", \"program_mean\", accuracy_score(y_test, nb_program.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nnb_program_tfidf = Pipeline([(\"program tfidf vectorizer\", TfidfVectorizer(program)), (\"transform\", TfidfTransformer()), (\"Guassian NB\", MultinomialNB())])\n# nb_program_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"GB\", \"program_tfidf\", accuracy_score(y_test, nb_program_tfidf.predict(X_test)))\n\n\n# PROGRAM-300 MODEL\n####EXTRA TREE####\nprint(\"EXTRA TREE ...\")\n# 1- MEAN VECTORIZER\netree_program = Pipeline([(\"program mean vectorizer\", MeanEmbeddingVectorizer(program)), (\"extra trees\", ExtraTreesClassifier(n_estimators=20))])\n# etree_program.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"program_mean\", accuracy_score(y_test, etree_program.predict(X_test)))\n\n\n# 2- TFIDF VECTORIZER\netree_program_tfidf = Pipeline([(\"program tfidf vectorizer\", TfidfEmbeddingVectorizer(program)), (\"extra trees\", ExtraTreesClassifier(n_estimators=20))])\n# etree_program_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"program_tfidf\", accuracy_score(y_test, etree_program_tfidf.predict(X_test)))\n\n####SVM####\nprint(\"SVM ... \")\n#1- MAIN VECTORIZER\nsvm_program = Pipeline([(\"program mean vectorizer\", MeanEmbeddingVectorizer(program)), (\"SVM\", LinearSVC(random_state=0, tol=1e-5))])\n# svm_program.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"program_mean\", accuracy_score(y_test, svm_program.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nsvm_program_tfidf = Pipeline([(\"word2vec vectorizer\", TfidfEmbeddingVectorizer(program)), (\"SVM\", LinearSVC(random_state=0, tol=1e-5))])\n# svm_program_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"program_tfidf\", accuracy_score(y_test, svm_program_tfidf.predict(X_test)))\n\n####MLP####\nprint(\"MLP ... \")\n# 1- MAIN VECTORIZER\nmlp_program = Pipeline(\n    [(\"program mean vectorizer\", MeanEmbeddingVectorizer(program)),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(20,10, 2), random_state=42))])\n# mlp_program.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"program_mean\", accuracy_score(y_test, mlp_program.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nmlp_program_tfidf = Pipeline(\n    [(\"word2vec vectorizer\", TfidfEmbeddingVectorizer(program)),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(20,10, 2), random_state=42))])\n# mlp_program_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"program_tfidf\", accuracy_score(y_test, mlp_program_tfidf.predict(X_test)))\n\n####LREG####\nprint(\"LREG ... \")\n# 1- MAIN VECTORIZER\nlreg_program = Pipeline(\n    [(\"program mean vectorizer\", MeanEmbeddingVectorizer(program)),\n     (\"LREG\", LogisticRegression(solver='lbfgs', multi_class='multinomial', random_state=42))])\n# lreg_program.fit(X=X_train, y=y_train)\n# results_df.set_value(\"LREG\", \"program_mean\", accuracy_score(y_test, lreg_program.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nlreg_program_tfidf = Pipeline(\n    [(\"program tfidf vectorizer\", TfidfVectorizer(program)), (\"transform\", TfidfTransformer()),\n     (\"LREG\", LogisticRegression(solver='lbfgs', multi_class='multinomial', random_state=42))])\n# lreg_program_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"LREG\", \"program_tfidf\", accuracy_score(y_test, lreg_program_tfidf.predict(X_test)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b9039fd990b06b6fb40fa81619bbd3c4a4f67add","_kg_hide-input":true},"cell_type":"code","source":"results_df.set_value(\"NB\",\"program_mean\",0.559782)\nresults_df.set_value(\"ExtraTree\",\"program_mean\",0.939124)\nresults_df.set_value(\"SVM\",\"program_mean\",0.939026)\nresults_df.set_value(\"MLP\",\"program_mean\",0.939035)\nresults_df.set_value(\"LREG\",\"program_mean\",0.938989)\n\nresults_df.set_value(\"NB\",\"program_tfidf\",0.941720)\nresults_df.set_value(\"ExtraTree\",\"program_tfidf\",0.945717)\nresults_df.set_value(\"SVM\",\"program_tfidf\",0.953854)\nresults_df.set_value(\"MLP\",\"program_tfidf\",0.939035)\nresults_df.set_value(\"LREG\",\"program_tfidf\",0.953311)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"52d94902f7e6f4b2ce69a84591aa533cf3967b40"},"cell_type":"code","source":"# fig,axes=plt.subplots(1,1,figsize=(15,8))\n# axes.set_ylabel(\"Accuracy\")\n# axes.set_title(\"PROGRAM results for 67% training and 23% testing\")\n# results_df.plot(kind=\"bar\",ax=axes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"26670c39125b5490d5625307b0b7a2245b46c734"},"cell_type":"code","source":"del program\nimport gc; gc.collect()\ntime.sleep(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aae841c29ea71577008465b0cfbbe9fd47bed798"},"cell_type":"markdown","source":"<a id=\"44\"></a> <br>\n#  8-EMBEDDING METHODS_GOOGLE_NEWS_VECTORS"},{"metadata":{"trusted":true,"_uuid":"66fe41c2b72252af1596b82ab464efbd27cc23e2"},"cell_type":"code","source":"model = gensim.models.KeyedVectors.load_word2vec_format(\n    fname='../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin',\n    binary=True)\ngoogle = dict(zip(model.wv.index2word, model.wv.syn0))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f0e062f2bf9faef0ff06a6282210aaf9e33e2680"},"cell_type":"code","source":"next(iter(google.values()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7f60694dce806c962be0d50fb232f8269c352270"},"cell_type":"code","source":"\"\"\"\nNOTE:\nAs these processes are time consuming, I have simulated them on my local machine. Committing these codes (Waiting until finishing whole the trainin process for\nany commit really bothers me) so I have commented the lines which have belonged to training process. You can easily uncomment them and try to\ntrain classifier by yourself ;-).\n\"\"\"\n\n\n\nprint(\"Multinomina NB ...\")\ngoogle_google = Pipeline([(\"google mean vectorizer\", MeanEmbeddingVectorizer(google)), (\"Guassian NB\", GaussianNB())])\n# google_google.fit(X=X_train, y=y_train)\n# results_df.set_value(\"GB\", \"google_mean\", accuracy_score(y_test, google_google.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\ngoogle_google_tfidf = Pipeline([(\"google tfidf vectorizer\", TfidfVectorizer(google)), (\"transform\", TfidfTransformer()), (\"Guassian NB\", MultinomialNB())])\n# google_google_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"GB\", \"google_tfidf\", accuracy_score(y_test, google_google_tfidf.predict(X_test)))\n\n\n# GOOGLE_NEWS_VEC MODEL\n####EXTRA TREE####\nprint(\"EXTRA TREE ...\")\n# 1- MEAN VECTORIZER\netree_google = Pipeline([(\"google mean vectorizer\", MeanEmbeddingVectorizer(google)), (\"extra trees\", ExtraTreesClassifier(n_estimators=20))])\n# etree_google.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"google_mean\", accuracy_score(y_test, etree_google.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\netree_google_tfidf = Pipeline([(\"google tfidf vectorizer\", TfidfEmbeddingVectorizer(google)), (\"extra trees\", ExtraTreesClassifier(n_estimators=20))])\n# etree_google_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"google_tfidf\", accuracy_score(y_test, etree_google_tfidf.predict(X_test)))\n\n####SVM####\nprint(\"SVM ... \")\n#1- MAIN VECTORIZER\nsvm_google = Pipeline([(\"google mean vectorizer\", MeanEmbeddingVectorizer(google)), (\"SVM\", LinearSVC(random_state=0, tol=1e-5))])\n# svm_google.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"google_mean\", accuracy_score(y_test, svm_google.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nsvm_google_tfidf = Pipeline([(\"word2vec vectorizer\", TfidfEmbeddingVectorizer(google)), (\"SVM\", LinearSVC(random_state=0, tol=1e-5))])\n# svm_google_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"google_tfidf\", accuracy_score(y_test, svm_google_tfidf.predict(X_test)))\n\n####MLP####\nprint(\"MLP ... \")\n# 1- MAIN VECTORIZER\nmlp_google = Pipeline(\n    [(\"google mean vectorizer\", MeanEmbeddingVectorizer(google)),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(5, 2), random_state=1))])\n# mlp_google.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"google_mean\", accuracy_score(y_test, mlp_google.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nmlp_google_tfidf = Pipeline(\n    [(\"word2vec vectorizer\", TfidfEmbeddingVectorizer(google)),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(5, 2), random_state=1))])\n# mlp_google_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"google_tfidf\", accuracy_score(y_test, mlp_google_tfidf.predict(X_test)))\n\n####LREG####\nprint(\"LREG ... \")\n# 1- MAIN VECTORIZER\nlreg_google = Pipeline(\n    [(\"google mean vectorizer\", MeanEmbeddingVectorizer(google)),\n     (\"LREG\", LogisticRegression(solver='lbfgs', multi_class='multinomial', random_state=42))])\n# lreg_google.fit(X=X_train, y=y_train)\n# results_df.set_value(\"LREG\", \"google_mean\", accuracy_score(y_test, lreg_google.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nlreg_google_tfidf = Pipeline(\n    [(\"google tfidf vectorizer\", TfidfVectorizer(google)), (\"transform\", TfidfTransformer()),\n     (\"LREG\", LogisticRegression(solver='lbfgs', multi_class='multinomial', random_state=42))])\n# lreg_google_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"LREG\", \"google_tfidf\", accuracy_score(y_test, lreg_google_tfidf.predict(X_test)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f136adbb6772a4fcc2da8b8d420f81cb0e4f74df","_kg_hide-input":true},"cell_type":"code","source":"results_df.set_value(\"NB\",\"google_mean\",0.527487)\nresults_df.set_value(\"ExtraTree\",\"google_mean\",0.939126)\nresults_df.set_value(\"SVM\",\"google_mean\",0.939038)\nresults_df.set_value(\"MLP\",\"google_mean\",0.939035)\nresults_df.set_value(\"LREG\",\"google_mean\",0.938806)\n\nresults_df.set_value(\"NB\",\"google_tfidf\",0.941720)\nresults_df.set_value(\"ExtraTree\",\"google_tfidf\",0.945448)\nresults_df.set_value(\"SVM\",\"google_tfidf\",0.953854)\nresults_df.set_value(\"MLP\",\"google_tfidf\",0.939035)\nresults_df.set_value(\"LREG\",\"google_tfidf\",0.953311)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8da0eb4e8b556fa0bac20f31abcf56d00a72cf39"},"cell_type":"code","source":"# fig,axes=plt.subplots(1,1,figsize=(15,8))\n# axes.set_ylabel(\"Accuracy\")\n# axes.set_title(\"GOOGLE_NEWS results for 67% training and 23% testing\")\n# results_df.plot(kind=\"bar\",ax=axes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"57503279f65ac2ab56fbcbc66634a6215be44576"},"cell_type":"code","source":"del google\ndel model\nimport gc; gc.collect()\ntime.sleep(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f00afcf3e30d8e302bfa4379ab075836725964c"},"cell_type":"markdown","source":"<a id=\"45\"></a> <br>\n#  9-EMBEDDING METHODS_WIKI_NEWS"},{"metadata":{"trusted":true,"_uuid":"5572457dba7bb6f11ab592b6408107b85a30ef4b"},"cell_type":"code","source":"model = gensim.models.KeyedVectors.load_word2vec_format(fname='../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec',encoding='utf-8')\nwiki = dict(zip(model.wv.index2word, model.wv.syn0))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"91bbff3c22415141018d479688eaaedd347e39e4"},"cell_type":"code","source":"next(iter(wiki.values()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dbbd5d12dc319c19aec2302bbf4e8151dbf77d82"},"cell_type":"code","source":"\"\"\"\nNOTE:\nAs these processes are time consuming, I have simulated them on my local machine. Committing these codes (Waiting until finishing whole the trainin process for\nany commit really bothers me) so I have commented the lines which have belonged to training process. You can easily uncomment them and try to\ntrain classifier by yourself ;-).\n\"\"\"\n\n\nprint(\"Multinomina NB ...\")\nwiki_wiki = Pipeline([(\"wiki mean vectorizer\", MeanEmbeddingVectorizer(wiki)), (\"Guassian NB\", GaussianNB())])\n# wiki_wiki.fit(X=X_train, y=y_train)\n# results_df.set_value(\"GB\", \"wiki_mean\", accuracy_score(y_test, wiki_wiki.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nwiki_wiki_tfidf = Pipeline([(\"wiki tfidf vectorizer\", TfidfVectorizer(wiki)), (\"transform\", TfidfTransformer()), (\"Guassian NB\", MultinomialNB())])\n# wiki_wiki_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"GB\", \"wiki_tfidf\", accuracy_score(y_test, wiki_wiki_tfidf.predict(X_test)))\n\n####EXTRA TREE####\nprint(\"EXTRA TREE ...\")\n# 1- MEAN VECTORIZER\netree_wiki = Pipeline([(\"wiki mean vectorizer\", MeanEmbeddingVectorizer(wiki)), (\"extra trees\", ExtraTreesClassifier(n_estimators=20))])\n# etree_wiki.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"wiki_mean\", accuracy_score(y_test, etree_wiki.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\netree_wiki_tfidf = Pipeline([(\"wiki tfidf vectorizer\", TfidfVectorizer(wiki)), (\"transform\", TfidfTransformer())\n                             , (\"extra trees\", ExtraTreesClassifier(n_estimators=20))])\n# etree_wiki_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"ExtraTree\", \"wiki_tfidf\", accuracy_score(y_test, etree_wiki_tfidf.predict(X_test)))\n\n####SVM####\nprint(\"SVM ... \")\n#1- MAIN VECTORIZER\nsvm_wiki = Pipeline([(\"wiki mean vectorizer\", MeanEmbeddingVectorizer(wiki)), (\"SVM\", LinearSVC(random_state=0, tol=1e-5))])\n# svm_wiki.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"wiki_mean\", accuracy_score(y_test, svm_wiki.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nsvm_wiki_tfidf = Pipeline([(\"wiki tfidf vectorizer\", TfidfVectorizer(wiki)), (\"transform\", TfidfTransformer()),\n                           (\"SVM\", LinearSVC(random_state=0, tol=1e-5))])\n# svm_wiki_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"SVM\", \"wiki_tfidf\", accuracy_score(y_test, svm_wiki_tfidf.predict(X_test)))\n\n####MLP####\nprint(\"MLP ... \")\n# 1- MAIN VECTORIZER\nmlp_wiki = Pipeline(\n    [(\"wiki mean vectorizer\", MeanEmbeddingVectorizer(wiki)),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(5, 2), random_state=1))])\n# mlp_wiki.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"wiki_mean\", accuracy_score(y_test, mlp_wiki.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nmlp_wiki_tfidf = Pipeline(\n    [(\"wiki tfidf vectorizer\", TfidfVectorizer(wiki)), (\"transform\", TfidfTransformer()),\n     (\"MLP\", MLPClassifier(solver='lbfgs', alpha=1e-5, hidden_layer_sizes=(5, 2), random_state=1))])\n# mlp_wiki_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"MLP\", \"wiki_tfidf\", accuracy_score(y_test, mlp_wiki_tfidf.predict(X_test)))\n\n####LREG####\nprint(\"LREG ... \")\n# 1- MAIN VECTORIZER\nlreg_wiki = Pipeline(\n    [(\"wiki mean vectorizer\", MeanEmbeddingVectorizer(wiki)),\n     (\"LREG\", LogisticRegression(solver='lbfgs', multi_class='multinomial', random_state=42))])\n# lreg_wiki.fit(X=X_train, y=y_train)\n# results_df.set_value(\"LREG\", \"wiki_mean\", accuracy_score(y_test, lreg_wiki.predict(X_test)))\n\n# 2- TFIDF VECTORIZER\nlreg_wiki_tfidf = Pipeline(\n    [(\"wiki tfidf vectorizer\", TfidfVectorizer(wiki)), (\"transform\", TfidfTransformer()),\n     (\"LREG\", LogisticRegression(solver='lbfgs', multi_class='multinomial', random_state=42))])\n# lreg_wiki_tfidf.fit(X=X_train, y=y_train)\n# results_df.set_value(\"LREG\", \"wiki_tfidf\", accuracy_score(y_test, lreg_wiki_tfidf.predict(X_test)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"34c8409eada07f679eea5e5e6e0adec28859144a","_kg_hide-input":true},"cell_type":"code","source":"results_df.set_value(\"NB\",\"wiki_mean\",0.939035)\nresults_df.set_value(\"ExtraTree\",\"wiki_mean\",0.939249)\nresults_df.set_value(\"SVM\",\"wiki_mean\",0.925018)\nresults_df.set_value(\"MLP\",\"wiki_mean\",0.074982)\nresults_df.set_value(\"LREG\",\"wiki_mean\",0.939035)\n\nresults_df.set_value(\"NB\",\"wiki_tfidf\",0.941720)\nresults_df.set_value(\"ExtraTree\",\"wiki_tfidf\",0.944973)\nresults_df.set_value(\"SVM\",\"wiki_tfidf\",0.953854)\nresults_df.set_value(\"MLP\",\"wiki_tfidf\",0.950935)\nresults_df.set_value(\"LREG\",\"wiki_tfidf\",0.953311)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2816741bbb08d3107a3929e9710857990bb60942"},"cell_type":"code","source":"results_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"83d63496d0cfaac7f1616af560148867295b4ebf"},"cell_type":"code","source":"del wiki\ndel model\nimport gc; gc.collect()\ntime.sleep(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f21e46ddcce5902c2bb55cbf29e72c94e2aa766e"},"cell_type":"markdown","source":"<a id=\"46\"></a> <br>\n#  10-CONCLUSION ON TRADITIONAL METHODS\n"},{"metadata":{"_uuid":"a823511534c063221a7496c201f770ae98bce4e5","trusted":true},"cell_type":"code","source":"fig,axes=plt.subplots(1,1,figsize=(15,8))\nplt.ylim((.5,1))\naxes.set_ylabel(\"Accuracy\")\naxes.set_title(\"Traditional Classfieris Results for 67% Training and 23% Testing with Two Types of Embedding\")\nresults_df[results_df.index != \"RF\"].plot(kind=\"bar\",ax=axes)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4d91d8714353c8b53edf052f41c396264a51ff70"},"cell_type":"markdown","source":"The point could be mentioed is approximately whole the classifiers and whole the embedding methods (except our trained w2v) have the similar accuracy another point could be mentioned is in overall range of accuracy in tfidf model is greater than mean-vectorizer models."},{"metadata":{"_uuid":"c35790ca2603175dfa0b3a683251640646532e96"},"cell_type":"markdown","source":"Now, Lets remove variables for having clean workspace for rest of the kernel."},{"metadata":{"trusted":true,"_uuid":"793c642379aba27888627edef3b7b77d17e93c2b"},"cell_type":"code","source":"for name in dir():\n    if not name.startswith('_'):\n        del globals()[name]\nfor name in dir():\n    if not name.startswith('_'):\n        del locals()[name]\nimport gc; gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2e681a78ab4cf9c0519bc48354a8fab047714e18"},"cell_type":"markdown","source":"<a id=\"47\"></a> <br>\n#  11-MOVING TO DEEP MODELS"},{"metadata":{"_uuid":"6ef10e6528a5f573eb764346c092fa4ee2738b3e"},"cell_type":"markdown","source":"<a id=\"48\"></a> <br>\n* **A. INTRODUCTION AND ROADMAP**\n\nCongratulations. Now you have enough information about \n        \n*What does text mining means ?!*\n\n*How to convert text to the apprehensible language for MachineLearning algorithms ?!*\n\n*What does countVectorizer and TfIdf vectorizer means ?! *\n\n*What embedding is ?!*\n\n*How to map embeddings to our data ?!*\n    \n"},{"metadata":{"_uuid":"1353ddc2b4778236932a8fcb545c1319584a54ad"},"cell_type":"markdown","source":"In this section, we will look at how we can do a similar procedure in deep learning world.  We will use Keras as our main deep learning interface. The back engine is also tensorflow.  In the next sections we will look at the embedding methods and how we can apply them in our manner using Keras api's.\n    "},{"metadata":{"_uuid":"5eaf4b65dfb183a0c82298ad676c2a2d94b2e4cb"},"cell_type":"markdown","source":"Keras offers an Embedding layer that can be used for neural networks on text data. It requires that the input data be integer encoded, so that each word is represented by a unique integer. This data preparation step can be performed using the Tokenizer API also provided with Keras. The Embedding layer is initialized with random weights and will learn an embedding for all of the words in the training dataset.\nIt is a flexible layer that can be used in a variety of ways, such as:\n\n1. It can be used alone to learn a word embedding that can be saved and used in another model later.\n\n2. It can be used as part of a deep learning model where the embedding is learned along with the model itself.\n\n3. It can be used to load a pre-trained word embedding model, a type of transfer learning.\n"},{"metadata":{"_uuid":"d72d74b9a4a55d38cee904918fc70036564b66ed"},"cell_type":"markdown","source":"The Embedding layer is defined as the first hidden layer of a network. It must specify 3 arguments:\n\n1. input_dim: This is the size of the vocabulary in the text data. For example, if your data is integer encoded to values between 0-10, then the size of the vocabulary would be 11 words.\n2. output_dim: This is the size of the vector space in which words will be embedded. It defines the size of the output vectors from this layer for each word. For example, it could be 32 or 100 or even larger. Test different values for your problem.\n3. input_length: This is the length of input sequences, as you would define for any input layer of a Keras model. For example, if all of your input documents are comprised of 1000 words, this would be 1000.\n\n"},{"metadata":{"_uuid":"87e14c797783df772d8345f1e0652b2e49efde27"},"cell_type":"markdown","source":"For example, we want have 5000 word and embedding size we want is 300. So, we have Embedding(input_dim=5000, output_dim=300). The output of the Embedding layer is a 2D vector with one embedding for each word in the input sequence of words (input document). If you wish to connect a Dense layer directly to an Embedding layer, you must first flatten the 2D output matrix to a 1D vector using the Flatten layer."},{"metadata":{"_uuid":"c9c7ffce0f03f969ddcbe49514411513aed53572"},"cell_type":"markdown","source":"<a id=\"49\"></a> <br>\n* **B. PREPARING DATA FOR USING IN KERAS**\n\nWe select X_train and X_test as the same way we had done in previous section.\n\nLets initialize the transformation process.\n"},{"metadata":{"trusted":true,"_uuid":"e9a23a0ce266dd5db4c5ceee077233c89693ed65"},"cell_type":"code","source":"import pandas as pd\nfrom keras import Sequential\nfrom keras.layers import Bidirectional, GlobalMaxPool1D\nfrom keras.layers import Bidirectional, GlobalMaxPool1D,Dense, Input, Embedding, Dropout, LSTM, CuDNNGRU\nfrom keras import Sequential\nfrom keras.models import Model\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.preprocessing.text import Tokenizer\nfrom sklearn.model_selection import train_test_split\ntrain_df = pd.DataFrame.from_csv(\"../input/train.csv\")\nX_train, X_test, y_train, y_test = train_test_split(train_df.question_text, train_df.target, test_size=0.33, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"02b50c172addac73136258a169260ecbaf7628a7"},"cell_type":"markdown","source":"Same to the process we had done in countVectorizer, we need to tokenize the text and convert them to list of integeres. But how ?!\n\nKeras provides the tokenizer method for doing it easily."},{"metadata":{"trusted":true,"_uuid":"8e1c522061e3a83a424fdb7ddee103c022d4191b"},"cell_type":"code","source":"num_words = 5000\nmaxlen = 100\ntokenizer = Tokenizer(num_words=num_words)\ntokenizer.fit_on_texts(list(X_train))\nX_train = tokenizer.texts_to_sequences(X_train)\nX_test = tokenizer.texts_to_sequences(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0f6831b28e3477d5301779e4462a304282ede4e4"},"cell_type":"code","source":"X_train[0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"90f0eb841be4a4b19fb0813a23b2bfc8e6220077"},"cell_type":"code","source":"len(X_train[0]),len(X_train[10])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1bc1b245ca7157824b2a1566f751d7bb495c41cb"},"cell_type":"markdown","source":"We need to have a standard dimentions for whole the X_train records. As you can see above, the dimantion of one vector is 14; and it is 22 for anothre record. we need to convert them to standard dimentoin.\nHopefully, like Tokenizer, keras also provide pad_sequence to solving this issue. Lets do it. \n\nLets consider the padding size 150. It means we ignore the words 151 and plus in sentences (if exist)"},{"metadata":{"_uuid":"dd77495a03b307795216077b187f14e84d99b2d2"},"cell_type":"markdown","source":""},{"metadata":{"trusted":true,"_uuid":"b7f11c34f72ed561c922b60f4f259a3dc9a62b33"},"cell_type":"code","source":"max_len = 150\nX_train = pad_sequences(X_train, maxlen=maxlen)\nX_test = pad_sequences(X_test, maxlen=maxlen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"950a26267e369ef25a28f8a2e580a6831f8c765d"},"cell_type":"code","source":"len(X_train[0]),len(X_train[10])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"83dd707677f226f217b5ebd942adac118fbb6d8a"},"cell_type":"markdown","source":"Now, you can see the issule has been solved. Now, whole the X_train has the same dimention. Now, it can be passed to the classfieirs. \n\nWe have used CuDNNGRU which is Fast GRU implementation backed by cuDNN. So, first of all turn on GPU for your kernel in the your kernel setting. If you are dealing with a system which is does not contain GPU, replace the CuDNNGRU with LSTM method. more information could be found [here.](https://stackoverflow.com/questions/49183538/simple-example-of-cudnngru-based-rnn-implementation-in-tensorflow)"},{"metadata":{"trusted":true,"_uuid":"c6d5d8627c0680a977d4484ba0e6a19bccaa6f90"},"cell_type":"code","source":"embedding_size = 300\nmodel = Sequential()\nmodel.add(Embedding(num_words, embedding_size))\n## Bidirectional wrapper for RNNs. It involves duplicating the first recurrent\n## layer in the network so that there are now two layers side-by-side, then \n## providing the input sequence as-is as input to the first layer and providing\n## a reversed copy of the input sequence to the second.\nmodel.add(Bidirectional(CuDNNGRU(64, return_sequences=True))) \nmodel.add(GlobalMaxPool1D())\nmodel.add(Dense(16, activation=\"relu\"))\nmodel.add(Dropout(0.1))\nmodel.add(Dense(1, activation=\"sigmoid\"))\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d5298c1be71abfe257490d8e87e434171f5ffa9e"},"cell_type":"code","source":"model.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"851db02583b7326811d1f5bdcf4d9ecd4b0b738a"},"cell_type":"code","source":"model.fit(x=X_train, y=y_train, batch_size=512, epochs=1, validation_data=(X_test, y_test))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f7cd20d57370c89b719a6f84b95392e19987737d"},"cell_type":"markdown","source":"Now we have created our model considering deep learning approach. \n\nAnother point could be touched on is, because of being time consuing procedures, I have limited the epoch numbers on 1. It is clear that the greater number of epochs is proportional to the more accuracy you can get.\n\nI will try to present some concise doctuments  about the difference existed in various types of deep learning models. and represent them in future commits."},{"metadata":{"_uuid":"2cf54836ccc4aba8e5088ed368787b30e9ff4024"},"cell_type":"markdown","source":"The kernel is under development and more analysis in both sides of concept descriptions and code developements have been queued to come in near future.\nBe in touch with it ;-)\n\nThe rest of tutorial will be completed  as soon as possible ;-).\n\nAny comment, idea or hint will be appreciated.\n\n**Your upvote will be motivation for me for continuing the kernel ;-)**\n"},{"metadata":{"_uuid":"141b3ae8ec42efd8f2a29d018b69bb7ee2bcb3ae"},"cell_type":"markdown","source":"<a id=\"100\"></a> <br>\n#  11-REFERENCES\n\n1- [Datascience for Buisiness.](https://www.amazon.com/Data-Science-Business-Data-Analytic-Thinking/dp/1449361323).\n\n2- [tf-idf website.](http://www.tfidf.com/)\n\n3- [Google Machine Learning Developement website.](https://developers.google.com/machine-learning/crash-course/embeddings/video-lecture)\n\n4- [Text Classification With Word2Vec.](http://nadbordrozd.github.io/blog/2016/05/20/text-classification-with-word2vec/)\n\n5-[How use word embedding layers for DeepLearning with Keras.](https://machinelearningmastery.com/use-word-embedding-layers-deep-learning-keras/)\n\n6-Personal experiences in similar projects.\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}