{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Beginner to Intermediate Natural Language Processing Guide\n==================\n---\n## Table of Contents\n\n---\n- [**0.0 Setup**](#0.0-setup)\n    + [**0.1 Python & Anaconda**](#01-python-&-anaconda)\n    + [**0.2 Libraries**](#02-libraries)\n    + [**0.3 Other**](#03-other)\n- [**1.0 Background**](#10-background)\n    + [**1.1 What is NLP?**](#11-what-is-nlp)\n    + [**1.2 Why is NLP Important?**](#12-why-is-nlp-importance)\n    + [**1.3 Why is NLP a \"hard\" problem?**](#13-why-is-nlp-a-hard-problem)\n    + [**1.4 Glossary**](#14-glossary)\n- [**2.0 Sentiment Analysis**](#20-sentiment-analysis)\n    + [**2.1 Preparing the Data**](#21-preparing-the-data)\n        * [**2.1.1 Training Data**](#211-training-data)\n        * [**2.1.2 Test Data**](#212-test-data)\n    + [**2.2 Building a Classifier**](#22-building-a-classifier)\n    + [**2.3 Classification**](#53-classification)\n    + [**2.4 Accuracy**](#24-accuracy)\n- [**3.0 Regular Expressions**](#30-regular-expressions)\n    + [**3.1 Simplest Form**](#31-simplest-form)\n    + [**3.2 Case Sensitivity**](#32-case-sensitivity)\n    + [**3.3 Disjunctions**](#33-disjunctions) \n    + [**3.4 Ranges**](#34-ranges) \n    + [**3.5 Exclusions**](#35-exclusions) \n    + [**3.6 Question Marks**](#36-question-marks) \n    + [**3.7 Kleene Star**](#37-kleene-star) \n    + [**3.8 Wildcards**](#38-wildcards) \n    + [**3.9 Kleene+**](#39-kleene) \n- [**4.0 Word Tagging and Models**](#40-word-tagging-and-models)\n    + [**4.1 NLTK Parts of Speech Tagger**](#41-nltk-parts-of-speech-tagger)\n        * [**4.1.1 Ambiguity**](#411-ambiguity)\n    + [**4.2 Unigram Models**](#42-unigram-models)\n    + [**4.3 Bigram Models**](#43-bigram-models)\n- [**5.0 Normalizing Text**](#40-normalizing-text)\n    + [**5.1 Stemming**](#51-stemming)\n        * [**5.1.1 What is Stemming?**](#511-what-is-stemming)\n        * [**5.1.2 Types of Stemmers**](#512-types-of-stemmers)\n    + [**5.2 Lemmatization**](#52-lemmatization)\n        * [**5.2.1 What is Lemmatization?**](#521-what-is-lemmatization)\n        * [**5.2.2 WordNetLemmatizer?**](#522-wordnetlemmatizer)\n- [**6.0 Final Words**](#60-final-words)\n    + [**6.1 Resources**](#61-resources)\n---","metadata":{"_uuid":"434834c5a28bdcd4ace40545ce05177e08ac724e"}},{"cell_type":"markdown","source":"## 0.0 Setup\n---\n\nThis NLP guide was written in Python 3.6.\n\n### 0.1 Python & Anaconda\n\nDownload [Python](https://www.python.org/downloads/) and [Pip](http://docs.continuum.io/anaconda/install).\n\n\n### 0.2 Libraries\n\n* Here, We are working with `re` library for **regular expressions** and `nltk` for **natural language processing** techniques, so make sure to install them! To install these libraries, enter the following commands into your terminal: \n\n``` \npip3 install re\npip3 install nltk\n```\n\n### 0.3 Other\n\nSince we'll be working on textual analysis, we'll be using datasets that are already well established and widely used. To gain access to these datasets, enter the following command into your command line: (Note that this might take a few minutes!)\n\n```\nsudo python3 -m nltk.downloader all\n```\n\nLastly, download the data we'll be working with in this example! \n\n[Positive Tweets](https://github.com/lesley2958/natural-language-processing/blob/master/pos_tweets.txt) <br>\n[Negative Tweets](https://github.com/lesley2958/natural-language-processing/blob/master/neg_tweets.txt)\n\nNow you're all set to begin!","metadata":{"_uuid":"d1528538b54a13e5090a0ae5baf6e51b7b830a04"}},{"cell_type":"markdown","source":"## 1.0 Background\n---\n\n### 1.1 What is NLP? \n\nNatural Language Processing, or NLP, is an area of computer science that focuses on developing techniques to produce machine-driven analyses of text.[**...Wikipedia**](https://www.google.co.in/url?sa=t&rct=j&q=&esrc=s&source=web&cd=29&cad=rja&uact=8&ved=2ahUKEwiakISXsoHfAhXUXysKHYQpAHIQmhMwHHoECAMQAg&url=https%3A%2F%2Fen.wikipedia.org%2Fwiki%2FNatural_language_processing&usg=AOvVaw1KDOCNufwMWaoNt_JWX4QL)\n\n### 1.2 Why is Natural Language Processing Important? \n\n* NLP expands the **unmitigated amount of data that can be used for getting insight for use**. Since so much of the data we have available is in the **major format of text,** this is exceedingly important to data science and for Industry !\n\n* A specific common application of **NLP** is each time you use a **language conversion tool.** This techniques used to **accurately convert text from one language to another**(Like **Google Translator**) very much falls under the umbrella of **\"natural language processing.\"**\n\n### 1.3 Why is NLP a \"hard\" problem? \n\n* The language is ambiguous. Once, the explanation of a person's sentence can be very different from another person's explanation. Because of this inability to be constantly clear, it is difficult to have a NLP technique that works perfectly.\n\n### 1.4 Glossary\n\nHere is some common terminology:\n\n<b>Corpus: </b> (Plural: Corpora) a collection of written texts that serve as our datasets.\n\n<b>nltk: </b> (Natural Language Toolkit) the python module we'll be using repeatedly; it has a lot of useful built-in NLP techniques.\n\n<b>Token: </b> a string of contiguous characters between two spaces, or between a space and punctuation marks. A token can also be an integer, real, or a number with a colon.\n","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a"}},{"cell_type":"markdown","source":"## 2.0 Sentiment Analysis  \n---\n\nSo you are wondering, what exactly is \"the Sentimental Analysis\"? \n\nWell, sentiment analysis involves building a system to collect and determine the emotional tone behind the words. This is important because you can understand the attitudes, opinions and emotions of people in your data. \n\nAt a high level, sentiment analysis involves the processing of **natural language and artificial intelligence** by assuming that the actual element of the text is converted to a format that can be read by a machine, and statistics are used to determine the actual sentiment,\n\n### 2.1 Preparing the Data \n\nTo accomplish sentiment analysis computationally, we have to use techniques that will allow us to learn from data that's already been labeled. \n\nSo what's the first step? Formatting the data so that we can actually apply NLP techniques. ","metadata":{"_uuid":"0423dc4a4511265f8d23d2e72af138f776ae0146"}},{"cell_type":"code","source":"import nltk\n\ndef format_sentence(sent):\n    return({word: True for word in nltk.word_tokenize(sent)})\n\nformat_sentence(\"Life is beautiful so Enjoy everymoment you have.\")","metadata":{"_uuid":"f8e90abefb5b156fa37b59800fb1a405b6046e18","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, `format_sentence` changes a piece of text, in this case a tweet, into a dictionary of words mapped to True booleans. Though not obvious from this function alone, this will eventually allow us to train  our prediction model by splitting the text into its tokens, i.e. <i>tokenizing</i> the text.","metadata":{"_uuid":"ad8066f3cee0d645435b66e1dd45ae4e4c70c5a6"}},{"cell_type":"code","source":"pos = []\nwith open(\"../input/sentimental-analysis-nlp/pos_tweets.txt\") as f:\n    for i in f: \n        pos.append([format_sentence(i), 'pos'])\n        \npos[0]","metadata":{"_uuid":"6f239b37d0be47bd6d7c3d832bb35ab411be6bdf","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"neg = []\nwith open(\"../input/sentimental-analysis-nlp/neg_tweets.txt\") as f:\n    for i in f: \n        neg.append([format_sentence(i), 'neg'])\n        \nneg[0]","metadata":{"_uuid":"9e96f0dd551eb8919da668227de5dcd161148ef7","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.1.1 Training Data\n\nNext, we'll split the labeled data we have into two pieces, one that can \"train\" data and the other to give us insight on how well our model is performing. The training data will inform our model on which features are most important.\n","metadata":{"_uuid":"afa04d65a38912bd7cf2c8ad1f2d0ff95864507a","trusted":true}},{"cell_type":"code","source":"training = pos[:int((.9)*len(pos))] + neg[:int((.9)*len(neg))]","metadata":{"_uuid":"b0bd6dd376a2487a0d1a5e919b6dee5662c5dbc9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.1.2 Test Data\n\nWe won't use the test data until the very end of this section, but nevertheless, we save the last 10% of the data to check the accuracy of our model. ","metadata":{"_uuid":"d6c9e067b7d1ccdcf8cb3e56fe2d8c4ace835eb0"}},{"cell_type":"code","source":"test = pos[int((.1)*len(pos)):] + neg[int((.1)*len(neg)):]","metadata":{"_uuid":"6d57bd0aa79210522df7e2c6cc3b0ac7df856d44","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.2 Building a Classifier\n\nAll NLTK classifiers work with feature structures, which can be simple dictionaries mapping a feature name to a feature value. In this example, we’ve used a simple bag of words model where every word is a feature name with a value of True.","metadata":{"_uuid":"dae4ac7594400aaae31a75cb994fddc15725e0f2"}},{"cell_type":"code","source":"from nltk.classify import NaiveBayesClassifier\n\nclassifier = NaiveBayesClassifier.train(training)","metadata":{"_uuid":"b26e66e3fc07fbeba07ea354ad40d9664efefa94","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To see which features informed our model the most, we can run this line of code:","metadata":{"_uuid":"b8b923adef57762b51818a9732c34b8cf12743d9"}},{"cell_type":"code","source":"classifier.show_most_informative_features()","metadata":{"_uuid":"09cd55068df0f43b1cbefe1480ecef2b82a16e74","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.3 Classification\n\nJust to see that our model works, let's try the classifier out with a positive example: ","metadata":{"_uuid":"f74e4f7a5608038dadabefc632d5fbd2cde62f54"}},{"cell_type":"code","source":"example1 = \"this workshop is awesome.\"\nexample2 = \"This workshop is not good\"\n\nprint(classifier.classify(format_sentence(example1)))\nprint(classifier.classify(format_sentence(example2)))","metadata":{"_uuid":"8e14183b19604e659182ab5c88e1b5506a42e356","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.4 Accuracy\n\nNow, there's no point in building a model if it doesn't work well. Luckily, once again, nltk comes to the rescue with a built in feature that allows us find the accuracy of our model.","metadata":{"_uuid":"ad31be3d133b715cd3c66ea43cfdec5c8177a846"}},{"cell_type":"code","source":"from nltk.classify.util import accuracy\n\nprint(accuracy(classifier, test))","metadata":{"_uuid":"b1c081504c628375a862cc46753a8ec79c0fe499","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Turns out it works decently well!\n\nBut it could be better! I think we can agree that the data is kind of messy - there are typos, abbreviations, grammatical errors of all sorts... So how do we handle that? Can we handle that? ","metadata":{"_uuid":"9b42be7f17ec33bf0a2d81a39ef55ca7b1d40d1c"}},{"cell_type":"markdown","source":"## 3.0 Regular Expressions\n---\n\nA regular expression is a sequence of characters that define a string.\n\n### 3.1 Simplest Form\n\nThe simplest form of a regular expression is a sequence of characters contained within <b>two backslashes</b>. For example, <i>python</i> would be  \n\n``` \n\\python\n```\n\n### 3.2 Case Sensitivity\n\nRegular Expressions are <b>case sensitive</b>, which means \n\n``` \n\\p and \\P\n```\nare distinguishable from eachother. This means <i>python</i> and <i>Python</i> would have to be represented differently, as follows: \n\n``` \n\\python and \\Python\n```\n\nWe can check these are different by running:\n","metadata":{"_uuid":"3e893b8bd975b4a8da5f1f547509c828d44c4876","trusted":true}},{"cell_type":"code","source":"import re\nre1 = re.compile('python')\nprint(bool(re1.match('Python')))","metadata":{"_uuid":"5345585c5047513a8e6720df90d3277ce14d6c48","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### 3.3 Disjunctions\n\nIf you want a regular expression to represent both <i>python</i> and <i>Python</i>, however, you can use <b>brackets</b> or the <b>pipe</b> symbol as the disjunction of the two forms. For example, \n\n``` \n[Pp]ython or \\Python|python\n```\n\ncould represent either <i>python</i> or <i>Python</i>. Likewise, \n\n``` \n[0123456789]\n```\n\nwould represent a single integer digit. The pipe symbols are typically used for interchangable strings, such as in the following example:\n\n```\n\\dog|cat\n```\n\n### 3.4 Ranges\n\nIf we want a regular expression to express the disjunction of a range of characters, we can use a <b>dash</b>. For example, instead of the previous example, we can write \n\n``` \n[0-9]\n```\nSimilarly, we can represent all characters of the alphabet with \n\n``` \n[a-z]\n```\n\n### 3.5 Exclusions\n\nBrackets can also be used to represent what an expression <b>cannot</b> be if you combine it with the <b>caret</b> sign. For example, the expression \n\n``` \n[^p]\n```\nrepresents any character, special characters included, but p.\n\n### 3.6 Question Marks \n\nQuestion marks can be used to represent the expressions containing zero or one instances of the previous character. For example, \n\n``` \n<i>\\colou?r\n```\nrepresents either <i>color</i> or <i>colour</i>. Question marks are often used in cases of plurality. For example, \n\n``` \n<i>\\computers?\n```\ncan be either <i>computers</i> or <i>computer</i>. If you want to extend this to more than one character, you can put the simple sequence within parenthesis, like this:\n\n```\n\\Feb(ruary)?\n```\nThis would evaluate to either <i>February</i> or <i>Feb</i>.\n\n### 3.7 Kleene Star\n\nTo represent the expressions containing zero or <b>more</b> instances of the previous character, we use an <b>asterisk</b> as the kleene star. To represent the set of strings containing <i>a, ab, abb, abbb, ...</i>, the following regular expression would be used:  \n```\n\\ab*\n```\n\n### 3.8 Wildcards\n\nWildcards are used to represent the possibility of any character and symbolized with a <b>period</b>. For example, \n\n```\n\\beg.n\n```\nFrom this regular expression, the strings <i>begun, begin, began,</i> etc., can be generated. \n\n### 3.9 Kleene+\n\nTo represent the expressions containing at <b>least</b> one or more instances of the previous character, we use a <b>plus</b> sign. To represent the set of strings containing <i>ab, abb, abbb, ...</i>, the following regular expression would be used:  \n\n```\n\\ab+\n```","metadata":{"_uuid":"9041ad546d3fe295c0100b7515c3c2c7d782ddb0","trusted":true}},{"cell_type":"markdown","source":"## 4.0 Word Tagging and Models\n---\n\nAny phrase, you can classify each word as a noun, a verb, a conjunction or any other type of words. When there are hundreds of thousands of prayers, even millions, it is obviously a huge and boring task. But it's not an impossible problem to solve on the IT front.\n\n### 4.1 NLTK Parts of Speech Tagger\n\nNLTK is a Python package that provides libraries for various word processing techniques, e.g. Classification, tokenization, stemming, parsing, but important for this example, tagging.","metadata":{"_uuid":"f576cfa4118f295af0cd0df5c9382f88b0701409"}},{"cell_type":"code","source":"import nltk \n\ntext = nltk.word_tokenize(\"Python is an awesome language!\")\nnltk.pos_tag(text)","metadata":{"_uuid":"bb425bc683a7daeb4a9e15f10bcd9854bc212ad4","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Not sure what DT, JJ, or any other tag is? Just try this in your python shell: ","metadata":{"_uuid":"ac3019ff568692f4207cc53d67d679582a6290ab"}},{"cell_type":"code","source":"nltk.help.upenn_tagset('JJ')","metadata":{"_uuid":"18c277e10ea4b30d144f7dcd8afc4b7d4438edb2","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 4.1.1 Ambiguity\n\nBut what happens if a word can be described as more than a part of speech? For example, the word \"sink\". Depending on the content of the sentence, it may be a noun or a verb.\n\nWhat happens when a text is a rhetorical instrument like sarcasm or irony? Of course, this can mislead the sentiment analyzer to misclassify a regular expression.\n\n### 4.2 Unigram Models\n\nDo you remember our word bag model from earlier? One of the features was that the word order was not taken into account. Therefore, the dictionaries can be used to map each word into true values.\n\nUnigram models are models where the order of our model makes no difference. You may be wondering why the unigram models interest us because they seem so simple without leaving their simplicity. They are a fundamental building block for many advanced NLP techniques.\n\n","metadata":{"_uuid":"e07d8cfea0fb20adabcc0fbce99153e4f9446dea"}},{"cell_type":"code","source":"from nltk.corpus import brown\n\nbrown_tagged_sents = brown.tagged_sents(categories='news')\nbrown_sents = brown.sents(categories='news')\nunigram_tagger = nltk.UnigramTagger(brown_tagged_sents)\nunigram_tagger.tag(brown_sents[2007])","metadata":{"_uuid":"d598484053f420ec94c9f750080fb22d03d8e05d","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.3 Bigram Models\n\nHere, ordering does matter. ","metadata":{"_uuid":"11b03321be04121915287ebc3b7c1333a1d021f3"}},{"cell_type":"code","source":"bigram_tagger = nltk.BigramTagger(brown_tagged_sents)\nbigram_tagger.tag(brown_sents[2007])","metadata":{"_uuid":"db8ffa0efc27f5d403bfebb1f5be8aad5a5adff3","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note the changes since the last time we tagged the words of this same sentence.","metadata":{"_uuid":"ab8f5dc71d0c893e0eb066a404fae986f038827f"}},{"cell_type":"markdown","source":"## 5.0 Normalizing Text\n---\nThe best data is consistent data, textual data usually not. But we can do it by normalizing it. We can do a number of things for that.\n\nAt least we can do all the text, so all lowercase. You may have done this before:\n\nGiven a piece of text,","metadata":{"_uuid":"b6b4713aab46e926703c28aeac61422b43a21854"}},{"cell_type":"code","source":"raw = \"OMG, Natural Language Processing is SO cool and I'm really enjoying this workshop!\"\ntokens = nltk.word_tokenize(raw)\ntokens = [i.lower() for i in tokens]\ntokens","metadata":{"_uuid":"5958dfdd22c8ffdcb6d8b4398d4300dff219feb4","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5.1 Stemming\n\nBut we can do more! \n\n#### 5.1.1 What is Stemming?\n\nStemming  is the process of converting the words of a sentence into non-editable parts. IIn the example of amusing, amusement, and amused above, the stem would be amus.\n\n#### 5.1.2 Types of Stemmers\n\n\nYou're probably wondering how to convert a bunch of words into their **stems**(***similar meaning words into same words***). Fortunately, NLTK has some integrated and well-established stemmers devices you can use! They work a little differently because they follow different rules - which you use depends on what you are currently working on.\n\nLet's start with Lancaster Stemmer:","metadata":{"_uuid":"99be0a8969d8d6751c0ac441823f7d619cabc965"}},{"cell_type":"code","source":"lancaster = nltk.LancasterStemmer()\nstems = [lancaster.stem(i) for i in tokens]\nstems","metadata":{"_uuid":"b44e01cd68c0805e51de51417da62d5897832c9c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Secondly, we try the Porter Stemmer:","metadata":{"_uuid":"fadc619741b27f79d259ba9963d19bf13b9b831e"}},{"cell_type":"code","source":"porter = nltk.PorterStemmer()\nstem = [porter.stem(i) for i in tokens]\nstem","metadata":{"_uuid":"514a1081c2ac952d4048a20681efaeb19036e7eb","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notice how \"natural\" maps to \"natur\" instead of \"nat\" and \"really\" maps to \"realli\" instead of \"real\" in the last stemmer. ","metadata":{"_uuid":"b7f52d0d2341011fbacf62354068a9ddb2e2d903"}},{"cell_type":"markdown","source":"### 5.2 Lemmatization\n\n#### 5.2.1 What is Lemmatization?\n\nLemmatization is the process of converting words from a sentence into a dictionary.  For example, given the words amusement, amusing, and amused, the lemma for each and all would be amuse.\n\n#### 5.2.2 WordNetLemmatizer\n\nOnce Again, NLTK is great and has an integrated Lemmatizer:","metadata":{"_uuid":"056e6a2d33ab085e2df79acb7a583de1ea110201"}},{"cell_type":"code","source":"from nltk import WordNetLemmatizer\n\nlemma = nltk.WordNetLemmatizer()\ntext = \"Women in technology are amazing at coding\"\nex = [i.lower() for i in text.split()]\nlemmas = [lemma.lemmatize(i) for i in ex]\nlemmas","metadata":{"_uuid":"7d50df04958dc7837d1839ff91218aceae721967","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notice that women is changed to \"woman\"! ","metadata":{"_uuid":"021e07cffa6d2f3e6f1f37a902bf4cddebd9c940"}},{"cell_type":"markdown","source":"## 6.0 Final Words \n---\n\nBack to our first sentiment analysis, we could have improved our model in many ways by applying some of the techniques we have just experienced. The data from Twitter are seemingly chaotic and inconsistent. If we really wanted an extremely accurate model, we could have pre-processed the tweets to clean it up.\n\nSecond, the way we built our classifier could have been improved. Our feature extraction was relatively simple and could have been improved with a bigram model instead of the word-bag template. We could have corrected our Bayes classifier to take only the most common words into account.\n\n### 6.1 Resources\n\n[Natural Language Processing With Python](http://bit.ly/nlp-w-python) <br>\n[Regular Expressions Cookbook](http://bit.ly/regular-expressions-cb)","metadata":{"_uuid":"a22603d8a240f1e0e6ee619800b37df0cb42847c"}},{"cell_type":"markdown","source":"Intermediate Natural Language Processing\n==================\n\n## Table of Contents\n\n- [0.0 Setup](#00)\n    + [0.1 Python and Pip](#01)\n    + [0.2 Libraries](#02)\n    + [0.3 Other](#03)\n    + [0.4 Virtual Environment](#04)\n- [1.0 Background](#10)\n+ [1.1 Polarity Flippers](#11)\n    * [1.1.1 Negation](#111)\n+ [1.2 Multiword Expressions](#12)\n+ [1.3 WordNet](#13)\n    * [1.3.1 Synsets](#131)\n    * [1.3.2 Negation](#132)\n+ [1.4 SentiWordNet](#14)\n+ [1.5 Stop Words](#15)\n+ [1.6 Testing](#16)\n    * [1.6.1 Cross Validation](#161)\n    * [1.6.2 Precision](#162)\n+ [1.7 Logistic Regression](#17)\n- [2.0 Information Extraction](#20)\n+ [2.1 Data Forms](#21)\n+ [2.2 What is Information Extraction?](#22)\n- [3.0 Chunking](#30)\n + [3.1 Noun Phrase Chunking](#31)\n- [4.0 Named Entity Extraction](#40)\n    + [4.1 spaCy](#41)\n    + [4.2 nltk](#42)\n- [5.0 Relation Extraction](#50)\n    + [5.1 Rule-Based Systems](#51)\n    + [5.2 Machine Learning](#52)\n- [6.0 Sentiment Analysis](#60)\n    + [6.1 Loading the Data](#61)\n    + [6.2 Preparing the Data](#62)\n    + [6.3 Linear Classifier](#63)","metadata":{"_uuid":"e1a3cc4a9cbcd5b79f4ab4a17408596b964e8332"}},{"cell_type":"markdown","source":"## 0.0 Setup <a id=\"00\"></a>\n\n---\nThis guide was written in Python 3.6.\n\n### 0.1 Python & Pip <a id=\"01\"></a>\n\nIf you haven't already, please download [Python](https://www.python.org/downloads/) and [Pip](https://pip.pypa.io/en/stable/installing/).\n\n### 0.2 Libraries <a id=\"02\"></a>\n\nWe'll be working with the re library for regular expressions and nltk for natural language processing techniques, so make sure to install them! To install these libraries, enter the following commands into your terminal: \n\n``` \npip3 install nltk==3.2.4\npip3 install spacy==1.8.2\npip3 install pandas==0.20.1\npip3 install scikit-learn==0.18.1\n```\n\n### 0.3 Other <a id=\"03\"></a>\n\nSentence boundary detection requires the dependency parse, which requires data to be installed, so enter the following command in your terminal. \n\n```\npython3 -m spacy.en.download all\n```\n\n### 0.4 Virtual Environment <a id=\"04\"></a>\n\nIf you'd like to work in a virtual environment, you can set it up as follows: \n```\npip3 install virtualenv\nvirtualenv your_env\n```\nAnd then launch it with: \n```\nsource your_env/bin/activate\n```\n\nTo execute the visualizations in matplotlib, do the following:\n\n```\ncd ~/.matplotlib\nvim matplotlibrc\n```\nAnd then, write `backend: TkAgg` in the file. Now you should be set up with your virtual environment!\n\nCool, now we're ready to start! ","metadata":{"_uuid":"33172683d7c66ff43a05ce18d008f06bba5013d8"}},{"cell_type":"markdown","source":"## 1.0 Background <a id=\"10\"></a>\n\n---\n\n### 1.1 Polarity Flippers<a id=\"11\"></a>\n\n\nPolarity flippers are words that change positive expressions into negative ones or vice versa. \n\n#### 1.1.1 Negation <a id=\"111\"></a>\n\nNegations directly change an expression's sentiment by preceding the word before it. An example would be\n\n```\nThe cat is not nice.\n```\n\n#### 1.1.2 Constructive Discourse Connectives\n\nConstructive Discourse Connectives are words which indirectly change an expression's meaning with words like \"but\". An example would be \n\n``` \nI usually like cats, but this cat is evil.\n```\n\n### 1.2 Multiword Expressions<a id=\"12\"></a>\n\nMultiword expressions are important because, depending on the context, can be considered positive or negative. For example, \n\n``` \nThis song is shit.\n```\nis definitely considered negative. Whereas\n\n``` \nThis song is the shit.\n```\nis actually considered positive, simply because of the addition of 'the' before the word 'shit'.\n\n### 1.3 WordNet<a id=\"13\"></a>\n\nWordNet is an English lexical database with emphasis on synonymy - sort of like a thesaurus. Specifically, nouns, verbs, adjectives and adjectives are grouped into synonym sets. \n\n#### 1.3.1 Synsets<a id=\"131\"></a>\n\nnltk has a built-in WordNet that we can use to find synonyms. We import it as such:\n","metadata":{"_uuid":"d0df80c5c9769761215b3843674861c42ab3b74f"}},{"cell_type":"code","source":"from nltk.corpus import wordnet as wn","metadata":{"_uuid":"20ea121cec1fbbec8e22a46957cc3b6710fb5c1b","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we feed a word to the **synsets() method**, the ***return value will be the class to which belongs***. For example, ***if we call the method on motorcycle,*** ","metadata":{"_uuid":"fa21f53abb41759e031a49bed26013d101b8658e"}},{"cell_type":"code","source":"print(wn.synsets('motorcar'))","metadata":{"_uuid":"56e385a3876b1fabde08bfca1545398a25975fa0","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Awesome stuff!*** But if we want to take it a step further, we can. We've previously learned what **lemmas are** - if you want to obtain the lemmas for a given **synonym set**, you can use the following method:","metadata":{"_uuid":"fb14ddc18c19daaa34f65efa248a723759a0a4ca","trusted":true}},{"cell_type":"code","source":"print(wn.synset('car.n.01').lemma_names())","metadata":{"_uuid":"1feb0503f0bc16babcd65b6ae938a527d86e60b7","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even more, you can do things like get the definition of a word: ","metadata":{"_uuid":"310b425a022cf69cbd44625ffdc96f16b7fcb4f4"}},{"cell_type":"code","source":"print(wn.synset('car.n.01').definition())","metadata":{"_uuid":"fb407c83da973a355af6788e5004a74372ed6d9e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.3.2 Negation<a id=\"132\"></a>\n\nWith **WordNet,** we can easily **detect negations**. This is great because ***it's not only fast, but it requires no training data and has a fairly good predictive accuracy.*** On the other hand, it's not able to handle context well or work with multiple word phrases. \n\n\n### 1.4 SentiWordNet<a id=\"14\"></a>\n\nBased on **WordNet synsets, SentiWordNet is a lexical resource for opinion mining**, where ***each synset is assigned three sentiment scores: positivity, negativity, and objectivity.***\n","metadata":{"_uuid":"17d124d6cafbf1abf547617b9bcecbdcfc93cb5c"}},{"cell_type":"code","source":"from nltk.corpus import sentiwordnet as swn\ncat = swn.senti_synset('cat.n.03')","metadata":{"_uuid":"c1fed37a6a666348cdd63b3336c5badb464181a7","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat.pos_score()","metadata":{"_uuid":"0717fae123a2f1467d4d694d192e9d1ad6950a8e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat.neg_score()","metadata":{"_uuid":"dd123f77a819402e458c72a4382124e8f2bd3270","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat.obj_score()","metadata":{"_uuid":"8488256b3a48a970918c11f039c520b86d6938b8","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat.unicode_repr()","metadata":{"_uuid":"bc9113c401a79d15c0f340218e308cb37df65547","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1.5 Stop Words<a id=\"15\"></a>\n\n**Stop words** are ***extremely common words that would be of little value in our analysis are often excluded from the vocabulary entirely***. Some common **examples** are determiners like the ***, a, an, another,*** but your list of stop words (or <b>stop list</b>) **depends on the context of the problem you're working on. **\n\n### 1.6 Testing<a id=\"16\"></a>\n\n\n#### 1.6.1 Cross Validation<a id=\"161\"></a>\n\n**Cross validation** is a ***model evaluation method that works by not using the entire data set when training the model***, i.e. some of the data is removed before training begins. ***Once training is completed, the removed data is used to test the performance of the learned model on this data.*** This is important because it prevents your model from over learning (or overfitting) your data.\n\n#### 1.6.2 Precision<a id=\"162\"></a>\n\n**Precision** is the ***percentage of retrieved instances that are relevant - it measures the exactness of a classifier***. A **higher precision means less false positives**, while **a lower precision means more false positives. **\n\n#### 1.6.3 Recall\n\n**Recall** is ***the percentage of relevant instances that are retrieved***. **Higher recall means less false negatives, while lower recall means more false negatives**. Improving recall can often decrease precision because it gets increasingly harder to be precise as the sample space increases.\n\n#### 1.6.4 F-measure \n\nThe **f1-score is a measure of a test's accuracy** that considers both the precision and the recall. \n\n### 1.7 Logistic Regression<a id=\"17\"></a>\n\n**Logistic regression** is a ***generalized linear model commonly used for classifying binary data. Its output is a continuous range of values between 0 and 1***, usually representing the **probability, and its input is some form of discrete predictor. **\n","metadata":{"_uuid":"10c1767b8216d9d6bfbf1246c15c82d952fb49bb"}},{"cell_type":"markdown","source":"## 2.0  Information Extraction<a id=\"20\"></a>\n\n---\nInformation Extraction is the process of acquiring meaning from text in a computational manner. \n\n### 2.1 Data Forms<a id=\"21\"></a>\n\n#### 2.1.1 Structured Data\n\nStructured Data is when there is a regular and predictable organization of entities and relationships.\n\n#### 2.1.2 Unstructured Data\n\nUnstructured data, as the name suggests, assumes no organization. This is the case with most written textual data. \n\n### 2.2 What is Information Extraction?<a id=\"22\"></a>\n\nWith that said, information extraction is the means by which you acquire structured data from a given unstructured dataset. There are a number of ways in which this can be done, but generally, information extraction consists of searching for specific types of entities and relationships between those entities. \n\nAn example is being given the following text, \n\n```\nMartin received a 98% on his math exam, whereas Jacob received a 84%. Eli, who also took the same test, received an 89%. Lastly, Ojas received a 72%.\n```\nThis is clearly unstructured. It requires reading for any logical relationships to be extracted. Through the use of information extraction techniques, however, we could output structured data such as the following: \n\n```\nName     Grade\nMartin   98\nJacob    84\nEli      89\nOjas     72\n```\n\n## 3.0 Chunking<a id=\"30\"></a>\n\n---\nChunking is used for entity recognition and segments and labels multitoken sequences. This typically involves segmenting multi-token sequences and labeling them with entity types, such as 'person', 'organization', or 'time'. \n\n### 3.1 Noun Phrase Chunking<a id=\"31\"></a>\n\nNoun Phrase Chunking, or NP-Chunking, is where we search for chunks corresponding to individual noun phrases.\n\nWe can use nltk, as is the case most of the time, to create a chunk parser. We begin with importing nltk and defining a sentence with its parts-of-speeches tagged (which we covered in the previous tutorial). ","metadata":{"_uuid":"dc40ed69f92848109875914fbdca96dd5747e406"}},{"cell_type":"code","source":"import nltk \nsentence = [(\"the\", \"DT\"), (\"little\", \"JJ\"), (\"yellow\", \"JJ\"), (\"dog\", \"NN\"), (\"barked\", \"VBD\"), (\"at\", \"IN\"), (\"the\", \"DT\"), (\"cat\", \"NN\")]","metadata":{"_uuid":"845f0699630870e5165f9b2efe60fb01f66fa9ea","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we define the tag pattern of an ***NP chunk.*** A tag pattern is a sequence of ***part-of-speech tags delimited using angle brackets,*** e.g. **`<DT>?<JJ>*<NN>`**. This is how the parse tree for a given sentence is acquired.  ","metadata":{"_uuid":"f16f8c26dce7c3f9484c704fa2867abc43e759a9"}},{"cell_type":"code","source":"pattern = \"NP: {<DT>?<JJ>*<NN>}\" ","metadata":{"_uuid":"d8971eb007e71c71bd0af4e436232ef5b3000ef6","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally we create the **chunk parser with the nltk `RegexpParser()` class. **","metadata":{"_uuid":"2d4dd813aa6ec00776f2bb8c295d78a44dd84cff"}},{"cell_type":"code","source":"NPChunker = nltk.RegexpParser(pattern)","metadata":{"_uuid":"2b602d11cb9911e428d359e033718e88e6810bed","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And lastly, we actually **parse** the ***example sentence and display its parse tree.***","metadata":{"_uuid":"e35df37525fde4de5e9ad7302c5b281ea8dae3cc"}},{"cell_type":"code","source":"result = NPChunker.parse(sentence) ","metadata":{"_uuid":"b05c6bae1941a0e1f881fe13ba78a7dd3829b6a2","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result.draw","metadata":{"_uuid":"fa0ce2172e8fbb2d3c7936d7f0a3867a1b0206f8","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.0 Named Entity Extraction <a id=\"40\"></a>\n\n---\n\nNamed entities are noun phrases that refer to specific types of individuals, such as organizations, people, dates, etc. Therefore, the purpose of a named entity recognition (NER) system is to identify all textual mentions of the named entities.\n\n### 4.1 spaCy <a id=\"41\"></a>\n\nIn the following exercise, we'll build our own named entity recognition system with the Python module `spaCy`, a Python module commonly used for Natural Language Processing in industry. ","metadata":{"_uuid":"1251c2f398d8e1f095fdba2c4c77f85c3cb3aac7"}},{"cell_type":"code","source":"import spacy\nimport pandas as pd","metadata":{"_uuid":"4c050287d2c05d50393f1240c54a8546548d282a","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using `spaCy`, we'll load the built-in **English tokenizer, tagger, parser, NER and word vectors.** We indicate this with the parameter `'en'`:","metadata":{"_uuid":"dcd1463ec927fe1ba122d7367d6f38b3e99b2070"}},{"cell_type":"code","source":"nlp = spacy.load('en')","metadata":{"_uuid":"48a5151e23d19c42ba80338337e2405ac8ee2e2e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We need an example to actually process, so below is some text from Columbia's website: ","metadata":{"_uuid":"83c60a02e547cc33a9c199c04b7b3ffcfebbfccb"}},{"cell_type":"code","source":"review = \"Columbia University was founded in 1754 as King's College by royal charter of King George II of England. It is the oldest institution of higher learning in the state of New York and the fifth oldest in the United States. Controversy preceded the founding of the College, with various groups competing to determine its location and religious affiliation. Advocates of New York City met with success on the first point, while the Anglicans prevailed on the latter. However, all constituencies agreed to commit themselves to principles of religious liberty in establishing the policies of the College. In July 1754, Samuel Johnson held the first classes in a new schoolhouse adjoining Trinity Church, located on what is now lower Broadway in Manhattan. There were eight students in the class. At King's College, the future leaders of colonial society could receive an education designed to 'enlarge the Mind, improve the Understanding, polish the whole Man, and qualify them to support the brightest Characters in all the elevated stations in life.'' One early manifestation of the institution's lofty goals was the establishment in 1767 of the first American medical school to grant the M.D. degree.\"","metadata":{"_uuid":"23c04a277e2ebfce88dfed928ecf09de87ec98d4","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With this example in mind, we feed it into the tokenizer.","metadata":{"_uuid":"c75b5de10ef7337ca4381d2cf742f8ce35f18c33"}},{"cell_type":"code","source":"doc = nlp(review)","metadata":{"_uuid":"06f0818578be524f48c7f13490086ce18ae9ae1a","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Going along ***the process of named entity extraction, we begin by segmenting the text, i.e. splitting it into a list of sentences. ***","metadata":{"_uuid":"720afb6762242538a57420e6acbad559f25d12ed"}},{"cell_type":"code","source":"sentences = [sentence.orth_ for sentence in doc.sents] # list of sentences\nprint(\"There were {} sentences found.\".format(len(sentences)))","metadata":{"_uuid":"acd2966bbf1423dac505e69aeea40e715c402456","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we go a step further, and **count the number of nounphrases** by **taking advantage of chunk properties.**","metadata":{"_uuid":"a02a3685e9e718d47325c67aa77ad336720bbe2c"}},{"cell_type":"code","source":"nounphrases = [[np.orth_, np.root.head.orth_] for np in doc.noun_chunks]\nprint(\"There were {} noun phrases found.\".format(len(nounphrases)))","metadata":{"_uuid":"dad908964764c54d94c0c22b111c8777f0b95b25","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Lastly, we achieve our final goal: entity extraction.","metadata":{"_uuid":"a1df8be6f0732957e6b04eb4f79b232923ec5c44"}},{"cell_type":"code","source":"entities = list(doc.ents) # converts entities into a list\nprint(\"There were {} entities found\".format(len(entities)))","metadata":{"_uuid":"e6eec36f0fba1d9125e5c4716b2bcc3ab34b84a5","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### So now, we can turn this into a DataFrame for better visualization: ","metadata":{"_uuid":"12ceb01dbe0e298128a28294c5a0a1a40e230a16"}},{"cell_type":"code","source":"orgs_and_people = [entity.orth_ for entity in entities if entity.label_ in ['ORG','PERSON']]\npd.DataFrame(orgs_and_people)","metadata":{"_uuid":"4d89e34538de97895b61e0c98b26fdc340bce7a4","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### In summary, named entity extraction typically follows the process of sentence segmentation, noun phrase chunking, and, finally, entity extraction.","metadata":{"_uuid":"1a5054f8c158c20f87bb66f865f882fe92043009"}},{"cell_type":"markdown","source":"\n### 4.2 nltk <a id=\"42\"></a>\n\nNext, we'll work through a **similar example as before, this time using the nltk module to extract the named entities through the use of chunk parsing.** As always, we begin by importing our needed modules and example:  ","metadata":{"_uuid":"d3c4e36b574cfa76dcd279972057fb4dfe78990e"}},{"cell_type":"code","source":"import nltk\nimport re\ncontent = \"Starbucks has not been doing well lately\"","metadata":{"_uuid":"3d6dc799360eeac87405ec4a38008f03a2ea5bef","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Then, as always, we tokenize the sentence and follow up with parts-of-speech tagging.","metadata":{"_uuid":"3a9d69b42b4441dc924cce79024512aa049c7135"}},{"cell_type":"code","source":"tokenized = nltk.word_tokenize(content)\ntagged = nltk.pos_tag(tokenized)\nprint(tagged)","metadata":{"_uuid":"3d3b10b3c1bfff15ed7c1e655460de0fd1c179fe","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### So we take this POS tagged sentence and feed it to the `nltk.ne_chunk()` method. This method returns a nested Tree object, so we display the content with namedEnt.draw(). ","metadata":{"_uuid":"c17b29fe6577b1532b0317ee70680d2c2610f4a7"}},{"cell_type":"code","source":"namedEnt = nltk.ne_chunk(tagged)\nnamedEnt.draw","metadata":{"_uuid":"20f8f0119fb1922e741b78c38ebb23fb5275cdb9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, if you wanted to simply get the named entities from the namedEnt object we created, how do you think you would go about doing so?\n\n## 5.0 Relation Extraction <a id=\"50\"></a>\n\n---\n\nOnce we have identified named entities in a text, we then want to analyze for the relations that exist between them. This can be performed using either rule-based systems, which typically look for specific patterns in the text that connect entities and the intervening words, or using machine learning systems that typically attempt to learn such patterns automatically from a training corpus.\n\n### 5.1 Rule-Based Systems<a id=\"51\"></a>\n\nIn the rule-based systems approach, we look for all triples of the form (X, a, Y), where X and Y are named entities and a is the string of words that indicates the relationship between X and Y. Using regular expressions, we can pull out those instances of a that express the relation that we are looking for. \n\nIn the following code, we search for strings that contain the word \"in\". The special regular expression `(?!\\b.+ing\\b)` allows us to disregard strings such as `success in supervising the transition of`, where \"in\" is followed by a gerund. \n","metadata":{"_uuid":"2c8ca325450412d98075eab56de12475cc1c5389","trusted":true}},{"cell_type":"code","source":"IN = re.compile(r'.*\\bin\\b(?!\\b.+ing)')\nfor doc in nltk.corpus.ieer.parsed_docs('NYT_19980315'):\n    for rel in nltk.sem.relextract.extract_rels('ORG', 'LOC', doc,corpus='ieer', pattern = IN):\n        print (nltk.sem.relextract.rtuple(rel))","metadata":{"_uuid":"45493934d662e54ae163a3c37358feedf495a01b","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that the X and Y named entitities types all match with one another! Object type matching is an important and required part of this process. \n\n### 5.2 Machine Learning <a id=\"52\"> </a>\n\nWe won't be going through an example of a machine learning based entity extraction algorithm, but it's important to note the different machine learning algorithms that can be implemented to accomplish this task of relation extraction. \n\nMost simply, Logistic Regression can be used to classify the objects that relate to one another. But additionally, algorithms like Suport Vector Machines and Random Forest could also accomplish the job. Which algorithm you ultimately choose depends on which outperforms in terms of speed and accuracy.\n\nIn summary, it's important to note that while these algorithms will likely have high accurate rates, labeling thousands of relations (and entities!) is incredibly expensive.   \n\n\n## 6.0 Sentiment Analysis<a id=\"60\"></a>\n\n---\n\nAs we saw in the previous tutorial, sentiment analysis refers to the use of text analysis and statistical learning to identify and extract subjective information in textual data. For our last exercise in this tutorial, we'll introduce and use linear models in the context of a sentiment analysis problem.\n\n### 6.1 Loading the Data<a id=\"61\"></a>\n\nFirst, we begin by loading the data. Since we'll be using data available online, we'll use the urllib module to avoid having to manually download any data.","metadata":{"_uuid":"b6d52ea5ef7cf6dfcd85e6b173d47321114bd57d","trusted":true}},{"cell_type":"code","source":"import urllib.request","metadata":{"_uuid":"8b9a72393f72540c93e3dcd2ece56534a9fa1337","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So then we'll define the ***test and training data URLs to variables, as well as filenames for each of those datasets.***","metadata":{"_uuid":"2c0e3a59eb9e522507cd7a13ddb1bfc3f2108c19","trusted":true}},{"cell_type":"code","source":"test_file = '../input/sentimental-analysis-nlp/test_data.csv'\ntrain_file = '../input/sentimental-analysis-nlp/train_data.csv'","metadata":{"_uuid":"c8cc2a885ca34fbb2e06b6d43cedafd9b7ff7c03","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using the **links and filenames** from above, we'll officially download the data using the **urlib.request.urlretrieve** method.","metadata":{"_uuid":"214c42a1de6bfda71b13cd9a751532faad3fd292","trusted":true}},{"cell_type":"code","source":"# test_data_f = urllib.request.urlretrieve(test_file)\n# train_data_f = urllib.request.urlretrieve(train_file)","metadata":{"_uuid":"535b24c7588dfea483d58a29f96e3cd9adbcd45e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Now that we've downloaded our datasets, we can load them into pandas dataframes. First for the test data:","metadata":{"_uuid":"c51947a9095d5e1ebf6166f4e5564e04118371ce"}},{"cell_type":"code","source":"import pandas as pd\n\ntest_data_df = pd.read_csv(test_file, header=None, delimiter=\"\\t\", quoting=3)\ntest_data_df.columns = [\"Text\"]","metadata":{"_uuid":"94955033515f36e301f3d7bd298c189727b5abee","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Next for training data: ","metadata":{"_uuid":"b581ac53227582e6940d6b2aba16b78cfbe5f4b0","trusted":true}},{"cell_type":"code","source":"train_data_df = pd.read_csv(train_file, header=None, delimiter=\"\\t\", quoting=3)\ntrain_data_df.columns = [\"Sentiment\",\"Text\"]","metadata":{"_uuid":"e36626ab1468f21aad6c53371a072f8dba57ef6e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Just to see how the dataframe looks, let's call the .head() method on both dataframes. ","metadata":{"_uuid":"0b9ebbca96b50c74e38e9153531fd464f7945acf","trusted":true}},{"cell_type":"code","source":"test_data_df.head()","metadata":{"_uuid":"8885915d1de8dead70e383b88aa345752b6d993e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data_df.head()","metadata":{"_uuid":"f937d6afde65f02fbae5fdb65746986c806cce54","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6.2 Preparing the Data<a id=\"62\"></a>\n\nTo implement our bag-of-words linear classifier, we need our data in a format that allows us to feed it in to the classifer. Using sklearn.feature_extraction.text.CountVectorizer in the Python scikit learn module, we can convert the text documents to a matrix of token counts. So first, we import all the needed modules: \n","metadata":{"_uuid":"53dd214610a1868c6ce4afe3934cf156dba2ee26","trusted":true}},{"cell_type":"code","source":"import re\nimport nltk\nfrom sklearn.feature_extraction.text import CountVectorizer        \nfrom nltk.stem.porter import PorterStemmer","metadata":{"_uuid":"c9463df3339e17018773872987211b7867425c03","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We need to **remove punctuations, lowercase, remove stop words, and stem words.** All these steps can be **directly performed by CountVectorizer** if we pass the right parameter values. We can do this as follows. \n\nWe first create a **stemmer, using the Porter Stemmer implementation.**","metadata":{"_uuid":"a85becf625271a8ec1ccdd5d3c5a98728805ca10","trusted":true}},{"cell_type":"code","source":"stemmer = PorterStemmer()\ndef stem_tokens(tokens, stemmer):\n    stemmed = [stemmer.stem(item) for item in tokens]\n    return(stemmed)","metadata":{"_uuid":"12951aaf20d2856a87afd18342c4a6bee452028c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Here, we have our tokenizer, which removes non-letters and stems:","metadata":{"_uuid":"4c45276fc9b8b172c7e22d9383511b76d2141049","trusted":true}},{"cell_type":"code","source":"def tokenize(text):\n    text = re.sub(\"[^a-zA-Z]\", \" \", text)\n    tokens = nltk.word_tokenize(text)\n    stems = stem_tokens(tokens, stemmer)\n    return(stems)","metadata":{"_uuid":"e19cf60022a85b6f332daeca5a77261a8f833e18","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Here we init the vectoriser with the CountVectorizer class, making sure to pass our tokenizer and stemmers as parameters, remove stop words, and lowercase all characters.","metadata":{"_uuid":"ac90436d57744aae754ceaaa8efd35e17f62e6b8","trusted":true}},{"cell_type":"code","source":"vectorizer = CountVectorizer(\nanalyzer = 'word',\ntokenizer = tokenize,\nlowercase = True,\nstop_words = 'english',\nmax_features = 85\n)","metadata":{"_uuid":"a2316944b9aa4c2061aab96d1bd347c5817bfa7e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we use the ***fit_transform()*** method to **transform our corpus data into feature vectors.** Since the input needed is a **list of strings, we concatenate all of our training and test data.**","metadata":{"_uuid":"93b39010a3548170ea77ec001651dd96576ffeb7"}},{"cell_type":"code","source":"features = vectorizer.fit_transform(\ntrain_data_df.Text.tolist() + test_data_df.Text.tolist())","metadata":{"_uuid":"95d3f7c8a75f0193464d456d09e07229f2ec0a9d","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we're simply converting the features to an array for easier use.","metadata":{"_uuid":"5980a923a310bbc7cea9175d24bb65592a40bede"}},{"cell_type":"code","source":"features_nd = features.toarray()","metadata":{"_uuid":"706abda9a085d249cf8ed8fcc45b67a071a3d7eb","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6.3 Linear Classifier<a id=\"63\"></a>\nFinally, we begin building our classifier. Earlier we learned what a bag-of-words model. Here, we'll be using a similar model, but with some modifications. To refresh your mind, this kind of model simplifies text to a multi-set of terms frequencies.\n\nSo first we'll split our training data to get an evaluation set. As we mentioned before, we'll use cross validation to split the data. sklearn has a built-in method that will do this for us. All we need to do is provide the data and assign a training percentage (in this case, 75%).","metadata":{"_uuid":"d5da6e2110d9e04a269aea448dee8fe28b004df1"}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train, X_test, y_train, y_test  = train_test_split(features_nd[0:len(train_data_df)], train_data_df.Sentiment,train_size=0.85, random_state=1234)","metadata":{"_uuid":"34b0ae63ffbec89aa26e5f90be83a9e836b56c4f","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we're ready to **train our classifier.** We'll be using **Logistic Regression to model this data.** Once again, **sklearn has a built-in model for you to use, so we begin by importing the needed modules and calling the class.**","metadata":{"_uuid":"d8bf3a8e266bdc35d10c559f950bfc742f43a2ca"}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nlog_model = LogisticRegression()","metadata":{"_uuid":"ba8fbf2d8e3ac3f82d8501acb3dd6edab0bdf92c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And as always, we need actually do the training, so we call the **`.fit()`** method on our data.","metadata":{"_uuid":"2e670259c444b2805c9bc49c695cde5d40d143d7"}},{"cell_type":"code","source":"log_model = log_model.fit(X=X_train, y=y_train)","metadata":{"_uuid":"9847a615d0d89aef61e97e565a88d459fba29fbb","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we use the classifier to label the evaluation set we created earlier:","metadata":{"_uuid":"ef90ee50d2b67dbe03530e7bbc7e23319363c3e2"}},{"cell_type":"code","source":"y_pred = log_model.predict(X_test)","metadata":{"_uuid":"c6784756bef87afbaba4ea1eeb25a04004180231","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred","metadata":{"_uuid":"eb5556f7fe6ec52e7e83bbc2eb04efec674f076b","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6.4 Accuracy <a id=\"64\"></a>\n\nIn sklearn, there is a function called sklearn.metrics.classification_report which calculates several types of predictive scores on a classification model. So here we check out how exactly our model is performing:\n","metadata":{"_uuid":"6f6dba1c205e8a5b972fc24915dd24bce3f5a84e"}},{"cell_type":"code","source":"from sklearn.metrics import classification_report\nprint(classification_report(y_test, y_pred))","metadata":{"_uuid":"41fa853456c3108ccf606571fd101c527234da88","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"where precision, recall, and f1-score are the accuracy values discussed in the section 1.6. Support is the number of occurrences of each class in y_true and x_true.\n\n\n### 6.5 Retraining <a id=\"65\"></a>\n\nFinally, we can re-train our model with all the training data and use it for sentiment classification with the original unlabeled test set. \n\nSo we repeat the process from earlier, this time with different data:","metadata":{"_uuid":"d1ccbb3e9590910d18c0c8287b5dcf983b1149c2"}},{"cell_type":"code","source":"log_model = LogisticRegression()\nlog_model = log_model.fit(X=features_nd[0:len(train_data_df)], y=train_data_df.Sentiment)\ntest_pred = log_model.predict(features_nd[len(train_data_df):])","metadata":{"_uuid":"cce0ff8c49a00b8c0f67a8bf273df5c43211bd67","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So again, we can see what the predictions look:","metadata":{"_uuid":"2c26bacf7d87cb938094f03d19c933c84c413c3d"}},{"cell_type":"code","source":"test_pred","metadata":{"_uuid":"65f147e36ae6a3a4dde08ee5b5bae75cc7dff8e6","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And lastly, let's actually look at our predictions! Using the random module to select a random sliver of the data we predicted on, we'll print the results.  ","metadata":{"_uuid":"cec9152ac7678fb724a74bfe6d9a5d74b94aa373"}},{"cell_type":"code","source":"import random\nspl = random.sample(range(len(test_pred)), 10)\nfor text, sentiment in zip(test_data_df.Text[spl], test_pred[spl]):\n    print(sentiment, text)","metadata":{"_uuid":"d2937df80a100543867a19072bf18b22e2b98fe5","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Thanks for reading...Happy learning!!!","metadata":{"_uuid":"0d19762b93b268770b531fed93bc688dfe8a9880","trusted":true}}]}