{"cells":[{"metadata":{},"cell_type":"markdown","source":"### If you liked and / or this content was helpful, I appreciate your upvote :)"},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"3"},"deletable":false},"cell_type":"markdown","source":"## 1. Tools for text processing\n<p><img style=\"float: right ; margin: 5px 20px 5px 10px; width: 45%\" src=\"https://s3.amazonaws.com/assets.datacamp.com/production/project_38/img/Moby_Dick_p510_illustration.jpg\"> </p>\n<p>What are the most frequent words in Herman Melville's novel, Moby Dick, and how often do they occur?</p>\n<p>In this notebook, we'll scrape the novel <em>Moby Dick</em> from the website <a href=\"https://www.gutenberg.org/\">Project Gutenberg</a> (which contains a large corpus of books) using the Python package <code>requests</code>. Then we'll extract words from this web data using <code>BeautifulSoup</code>. Finally, we'll dive into analyzing the distribution of words using the Natural Language ToolKit (<code>nltk</code>). </p>\n<p>The <em>Data Science pipeline</em> we'll build in this notebook can be used to visualize the word frequency distributions of any novel that you can find on Project Gutenberg. The natural language processing tools used here apply to much of the data that data scientists encounter as a vast proportion of the world's data is unstructured data and includes a great deal of text.</p>\n<p>Let's start by loading in the three main Python packages we are going to use.</p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"3"}},"cell_type":"code","source":"# Importing requests, BeautifulSoup and nltk\nimport requests\nfrom bs4 import BeautifulSoup\nimport nltk\nfrom nltk.stem import WordNetLemmatizer \nprint('Done!')","execution_count":null,"outputs":[]},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"10"},"deletable":false},"cell_type":"markdown","source":"## 2. Request Moby Dick\n<p>To analyze Moby Dick, we need to get the contents of Moby Dick from <em>somewhere</em>. Luckily, the text is freely available online at Project Gutenberg as an HTML file: https://www.gutenberg.org/files/2701/2701-h/2701-h.htm .</p>\n<p><strong>Note</strong> that HTML stands for Hypertext Markup Language and is the standard markup language for the web.</p>\n<p>To fetch the HTML file with Moby Dick we're going to use the <code>request</code> package to make a <code>GET</code> request for the website, which means we're <em>getting</em> data from it. This is what you're doing through a browser when visiting a webpage, but now we're getting the requested page directly into Python instead. </p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"10"}},"cell_type":"code","source":"import codecs\n\nhtml = codecs.open(\"../input/2701-h.htm\", 'r', 'utf-8')\nprint('Done!')","execution_count":null,"outputs":[]},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"17"},"deletable":false},"cell_type":"markdown","source":"## 3. Get the text from the HTML\n<p>This HTML is not quite what we want. However, it does <em>contain</em> what we want: the text of <em>Moby Dick</em>. What we need to do now is <em>wrangle</em> this HTML to extract the text of the novel. For this we'll use the package <code>BeautifulSoup</code>.</p>\n<p>Firstly, a word on the name of the package: Beautiful Soup? In web development, the term \"tag soup\" refers to structurally or syntactically incorrect HTML code written for a web page. What Beautiful Soup does best is to make tag soup beautiful again and to extract information from it with ease! In fact, the main object created and queried when using this package is called <code>BeautifulSoup</code>. After creating the soup, we can use its <code>.get_text()</code> method to extract the text.</p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"17"}},"cell_type":"code","source":"# Creating a BeautifulSoup object from the HTML\nsoup = BeautifulSoup(html)\n\n# Getting the text out of the soup\ntext = soup.get_text()\n\n# Printing out text between characters 32000 and 34000\nprint(text[32000:34000])","execution_count":null,"outputs":[]},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"24"},"deletable":false},"cell_type":"markdown","source":"## 4. Extract the words\n<p>We now have the text of the novel! There is some unwanted stuff at the start and some unwanted stuff at the end. We could remove it, but this content is so much smaller in amount than the text of Moby Dick that, to a first approximation, it is okay to leave it in.</p>\n<p>Now that we have the text of interest, it's time to count how many times each word appears, and for this we'll use <code>nltk</code> – the Natural Language Toolkit. We'll start by tokenizing the text, that is, remove everything that isn't a word (whitespace, punctuation, etc.) and then split the text into a list of words.</p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"24"}},"cell_type":"code","source":"# Creating a tokenizer\ntokenizer = nltk.tokenize.RegexpTokenizer('\\w+')\n\n# Tokenizing the text\ntokens = tokenizer.tokenize(text)\n\n# Printing out the first 8 words / tokens \nprint(tokens[:8])","execution_count":null,"outputs":[]},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"31"},"deletable":false},"cell_type":"markdown","source":"## 5. Make the words lowercase\n<p>OK! We're nearly there. Note that in the above 'Or' has a capital 'O' and that in other places it may not, but both 'Or' and 'or' should be counted as the same word. For this reason, we should build a list of all words in <em>Moby Dick</em> in which all capital letters have been made lower case.</p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"31"}},"cell_type":"code","source":"# A new list to hold the lowercased words\n# Looping through the tokens and make them lower case\nwords = [word.lower() for word in tokens]\n\n# Printing out the first 8 words / tokens \nprint(words[:8])","execution_count":null,"outputs":[]},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"38"},"deletable":false},"cell_type":"markdown","source":"## 6. Load in stop words\n<p>It is common practice to remove words that appear a lot in the English language such as 'the', 'of' and 'a' because they're not so interesting. Such words are known as <em>stop words</em>. The package <code>nltk</code> includes a good list of stop words in English that we can use.</p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"38"}},"cell_type":"code","source":"# Getting the English stop words from nltk\nsw = nltk.corpus.stopwords.words('english')\n\n# Printing out the first eight stop words\nprint(sw[:8])","execution_count":null,"outputs":[]},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"45"},"deletable":false},"cell_type":"markdown","source":"## 7. Remove stop words in Moby Dick\n<p>We now want to create a new list with all <code>words</code> in Moby Dick, except those that are stop words (that is, those words listed in <code>sw</code>). One way to get this list is to loop over all elements of <code>words</code> and add each word to a new list if they are <em>not</em> in <code>sw</code>.</p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"45"}},"cell_type":"code","source":"# A new list to hold Moby Dick with No Stop words\n# Appending to words_ns all words that are in words but not in sw\n\nwords_ns = [word for word in words if word not in sw]\n\n# Printing the first 5 words_ns to check that stop words are gone\nprint(words_ns[:10])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 8. Lemmatization with NLTK\nLemmatization is the process of grouping together the different inflected forms of a word so they can be analysed as a single item. Lemmatization is similar to stemming but it brings context to the words. So it links words with similar meaning to one word.\n\nText preprocessing includes both Stemming as well as Lemmatization. Many times people find these two terms confusing. Some treat these two as same. Actually, lemmatization is preferred over Stemming because lemmatization does morphological analysis of the words.\n\nExample:"},{"metadata":{"trusted":true},"cell_type":"code","source":"lemmatizer = WordNetLemmatizer() \n\nprint(\"rocks :\", lemmatizer.lemmatize(\"rocks\")) \nprint(\"corpora :\", lemmatizer.lemmatize(\"corpora\")) \n  \n# a denotes adjective in \"pos\" \nprint(\"better :\", lemmatizer.lemmatize(\"better\", pos =\"a\")) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Lemmatization in the words of the book."},{"metadata":{"trusted":true},"cell_type":"code","source":"words_lem = [lemmatizer.lemmatize(word) for word in words_ns]\n\nprint(words_lem[:10])","execution_count":null,"outputs":[]},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"52"},"deletable":false},"cell_type":"markdown","source":"## 8. We have the answer\n<p>Our original question was:</p>\n<blockquote>\n  <p>What are the most frequent words in Herman Melville's novel Moby Dick and how often do they occur?</p>\n</blockquote>\n<p>We are now ready to answer that! Let's create a word frequency distribution plot using <code>nltk</code>. </p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"52"}},"cell_type":"code","source":"# This command display figures inline\nfrom matplotlib.pyplot import figure\n%matplotlib inline\n\n# Creating the word frequency distribution\nfreqdist = nltk.FreqDist(words_ns)\n\n# Plotting the word frequency distribution\nfigure(figsize=(10,5))\nfreqdist.plot(25)","execution_count":null,"outputs":[]},{"metadata":{"editable":false,"tags":["context"],"run_control":{"frozen":true},"dc":{"key":"59"},"deletable":false},"cell_type":"markdown","source":"## 9. The most common word\n<p>Nice! The frequency distribution plot above is the answer to our question. </p>\n<p>The natural language processing skills we used in this notebook are also applicable to much of the data that Data Scientists encounter as the vast proportion of the world's data is unstructured data and includes a great deal of text. </p>\n<p>So, what word turned out to (<em>not surprisingly</em>) be the most common word in Moby Dick?</p>"},{"metadata":{"trusted":true,"tags":["sample_code"],"dc":{"key":"59"}},"cell_type":"code","source":"# What's the most common word in Moby Dick?\nmost_common_word = 'whale'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### If you liked and / or this content was helpful, I appreciate your upvote :)"}],"metadata":{"language_info":{"version":"3.5.2","mimetype":"text/x-python","file_extension":".py","nbconvert_exporter":"python","name":"python","pygments_lexer":"ipython3","codemirror_mode":{"version":3,"name":"ipython"}},"kernelspec":{"display_name":"Python 3","name":"python3","language":"python"}},"nbformat":4,"nbformat_minor":1}