{"cells":[{"metadata":{"_uuid":"bf0b3ae77e0643e5ef50ca46cff8160ff8dc91fb","_execution_state":"idle","trusted":false},"cell_type":"markdown","source":"This kernel is based on the practice done over the book \"Natural Language Processing with Python\" by Steven Bird, Ewan Klein, Edward Loper.\n\nIn this part, we will focus on \"Computing with Language : Texts & Words\". It gives us interest and better understanding of basic statistics around language processing. "},{"metadata":{"trusted":true,"_uuid":"ae09ff4372a710e052ea516dc6c93fde6e41d301"},"cell_type":"code","source":"# Checking System Version\nimport sys\nprint(sys.version)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ebf45be4747612184634494864176513ab0a3b2b"},"cell_type":"markdown","source":"Downloading the **NLTK Book Collection** from 'nltk' package. It consists of about **30 compressed files** requiring about 100Mb disk space."},{"metadata":{"trusted":true,"_uuid":"67460421552dfdae8886e7dc4eb049a0f7abd397"},"cell_type":"code","source":"from nltk.book import *","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ceb1014ce116d64fd571220f0ebafc81be57dc14","scrolled":true},"cell_type":"code","source":"# Recall text by typing their name only\ntext7","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"19a1ec53c863a31d9b575f50da126816c6b0ea33"},"cell_type":"code","source":"# It gives you list as per the nltk book import\ntexts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c52c7d63dfce69ac1a86ac26d0ed2d93a76236e5"},"cell_type":"markdown","source":"**SEARCHING TEXT**    \n\nA **concordance** permits us to see **words in context** and **search for a specific word**. For example, we saw that monstrous occurred in contexts such as the **___ pictures** and the **___ size**."},{"metadata":{"trusted":true,"_uuid":"66342d5fb24a1d89b3c01dfa60ed650fa532a77e","scrolled":true},"cell_type":"code","source":"# Searching text\ntext1.concordance(\"monstrous\") ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c7762e4093d87cdd6a1b3b0307e29c200e7eeb46"},"cell_type":"code","source":"text2.concordance(\"affection\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e6e752224f38fa472a969acc112197405aa99bdc"},"cell_type":"markdown","source":"'**similar**'is used to get **similar words in the given range of context provided by 'concordance'**. Observe that **we get different results for different texts**. Austen uses this word quite differently from Melville; for her, monstrous has positive connotations, and sometimes functions as an intensifier like the word very."},{"metadata":{"trusted":true,"_uuid":"e1321e6298e43ccbf7cc39de51c93f680c6586c3"},"cell_type":"code","source":"text1.similar(\"monstrous\") ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4a1fde9ff18896ea0fd7d38ca3ad47199e96e1d2"},"cell_type":"code","source":"text2.similar(\"monstrous\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c0e2a89f6b6b1b311a420ed6a46c0217af70d04a"},"cell_type":"markdown","source":"The term **common_contexts** allows us to examine just the* **contexts that are shared by\ntwo or more words***, such as monstrous and very. We have to enclose these words by\nsquare brackets as well as parentheses, and separate them with a comma:"},{"metadata":{"trusted":true,"_uuid":"9c30e7fb67ad4b9850a277062c688965e5b70a84"},"cell_type":"code","source":"text2.common_contexts([\"monstrous\",\"very\"])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c1c6b2d3e77dc323e87e157cceff6c6cb2ea9e27"},"cell_type":"markdown","source":"It is one thing to automatically detect that a particular word occurs in a text, and to\ndisplay some words that appear in the same context. However, we can also determine\nthe location of a word in the text: how many words from the beginning it appears. This\npositional information can be displayed using a **dispersion plot**. Each stripe represents\nan instance of a word, and each row represents the entire text."},{"metadata":{"trusted":true,"_uuid":"c4e8ef5839e1b2f066d635430b06dce7af9ffda5"},"cell_type":"code","source":"text4.dispersion_plot([\"citizens\",\"democracy\",\"freedom\",\"duties\",\"America\"])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ff9fa09049e5d3a260b1a45482548b703ec08333"},"cell_type":"markdown","source":"**'len'** can be used to count vocabulary in a text line. It will count** words, punctuations, symbols, white spaces** and return you total count for that. So Genesis(text3) has 44,764 words and punctuation symbols, or “tokens.” A** token** is the technical name for **a sequence of characters**—such as hairy, his, or :)—that we want to treat as a group. When we count the number of tokens in a text, say, the phrase to be or not to be, we are **counting occurrences of these sequences**."},{"metadata":{"trusted":true,"_uuid":"60bb70d9b554dcf5371033cf30ac7fb2ed3d1e5c"},"cell_type":"code","source":"print(len(\"nishant;., is here\"))\n\nList_N1 = ['Call', 'me', 'Nishant', '.']\nprint(len(List_N1))\n\nprint(len(text3))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"567c5e61da9e142e54cf5d03e48b622e15d3e75e"},"cell_type":"markdown","source":"**The vocabulary of a text** is just the **set of (unique)tokens** that it uses, since in a set, **all duplicates are collapsed together**. In Python we can obtain the vocabulary items of text3 with the command: **set(text3)**. When you do this, many screens of words will fly past. "},{"metadata":{"trusted":true,"_uuid":"f658f425aaf486bc7b1f569db1966367ebe54231","scrolled":false},"cell_type":"code","source":"set(text3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"debac81ba36baee97f34f4908ddaa75c353ca3cb"},"cell_type":"code","source":"len(set(text3))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"34e5d851c6e5861c7cf0b57be2b90a8f5ffd49ec"},"cell_type":"markdown","source":"**By wrapping sorted()** around the Python expression set(text3) , we obtain** a sorted list of vocabulary items**, beginning with various punctuation symbols and continuing with words starting with A. All capitalized words precede lowercase words. We discover the size of the vocabulary indirectly, by asking for the number of items in the set, and again we can use len to obtain this number . Although it has 44,764 tokens, this book has only **2,789 distinct words, or “word types.”** A word type is the form or spelling of the word independently of its specific occurrences in a text—that is, the word considered as a unique item of vocabulary. **Our count of 2,789 items will include punctuation symbols, so we will generally call these unique items types instead of word types.**"},{"metadata":{"trusted":true,"_uuid":"a764855953964732339cc1eca6f17594547d204e"},"cell_type":"code","source":"sorted(set(text3))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"12f36ff42040c6de544004ced8f638541de2c1c6"},"cell_type":"markdown","source":"**Lexical Richness** : In computational linguistics, lexical richness is a measure of how many tokens (individual words and punctuation) occur in a given text, divided by how many types (**unique words and punctuation**) occur in that same text.\n\nlet’s calculate a measure of the lexical richness of the text. The next example shows us that each word is used 16 times on average."},{"metadata":{"trusted":true,"_uuid":"4bd77c98fa76678e2f0200600e4f24fe9aa1cdff"},"cell_type":"code","source":"len(text3)/len(set(text3))  #tokens (individual words and punctuation) occur in a given text, divided by how many types (unique words and punctuation)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"60e95faf8b3af61cd4754d8e52776038ffe4ba44"},"cell_type":"markdown","source":"We can count how often a word occurs in a text using '**count**' and so on it's percentage..."},{"metadata":{"trusted":true,"_uuid":"ba3028665cc41cda5ad758fd2db22d7213e0559b"},"cell_type":"code","source":"print(text5.count('lol')) # count of a specific word.\n100*(text5.count('lol'))/len(text5) # percentage against total word count.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"79a762c04f8850362d0183f9095860f902d60517"},"cell_type":"markdown","source":"**Make Functions** for lexical diversity/richness and percentage as follows: "},{"metadata":{"trusted":true,"_uuid":"340df1fa2e6da2e572266a3c922e12b7c41488e2"},"cell_type":"code","source":"#tokens (individual words and punctuation) occur in a given text, divided by how many types (unique words and punctuation)\ndef lex_dive(text):\n    return len(text)/len(set(text))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e1386df54858affe0c364517f8782b3e038ff951"},"cell_type":"code","source":"def percentage(count, total):\n    return 100*count/total","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"721444a4680fd8e5b8f1b0977ccfa24e1627c425"},"cell_type":"code","source":"lex_dive(text3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b114e25aa0fc9b8671641757ccc102e1ca6d08e1"},"cell_type":"code","source":"percentage(text5.count('lol'),len(text5))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a4638bc2d765da4e4a877121d0359d20eda2a5d8"},"cell_type":"markdown","source":"**Frequency Distribution**  tells us the frequency of each vocabulary item in the text. (In general, it could count any kind of\nobservable event.) It is a “distribution” since it tells us how the total number of word tokens in the text are distributed across the vocabulary items."},{"metadata":{"trusted":true,"_uuid":"e75814a2dfe1393e0445e200d8b22d7f948a708f"},"cell_type":"code","source":"# We can inspect the total number of words (“outcomes”) that have been counted up using FreqDist\nfdist1 = FreqDist(text1)\nfdist1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"48ff20e15e8b929f04362fe3d3aaae927d1535da"},"cell_type":"code","source":"#The expression keys() gives us a list of all the distinct types in the text\nvocab1 = fdist1.keys()\nvocab1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6984750418219127fbd5baac51be938bda4af85c"},"cell_type":"code","source":"# Count of word 'whale' in the frequency distribution list\nfdist1['whale']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"569a404ce49340c27c36a047a80334198cc24600"},"cell_type":"markdown","source":"**50 MOST FREQUENTLY USED WORDS:**\n\nOnly one word, **whale, is slightly informative!** It occurs over 900 times. The rest of the words tell us nothing about the text; they’re just English “plumbing.” \nWhat proportion of the text is taken up with such words? We can generate a **cumulative frequency plot** for these words, using **fdist1.plot(50, cumulative=True)**, to produce the graph (as given below). **These 50 words account for nearly half the book!**"},{"metadata":{"trusted":true,"_uuid":"8dff7f2d5c37dba90b8340975a76b746797668a1"},"cell_type":"code","source":"# To set figure size of the plot\nfrom matplotlib.pyplot import figure\nfigure(num=None, figsize=(10, 5), dpi=80, facecolor='w', edgecolor='k')\n\n# Cumulative Frequency Plot for 50 words\nfdist1.plot(50,cumulative = True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"23afc143a80564d4c4381ef29f7dafa08cd5c2e6"},"cell_type":"markdown","source":"**INFREQUENT/RARE WORDS (OCCURS ONLY ONCE):**\n\nThe **words that occur once** only, the socalled **hapaxes**. View them by typing **fdist1.hapaxes()**. This list contains lexicographer, cetological, contraband, expostulations, and about 9,000 others. It seems that there are too many rare words, and without seeing the context we probably can’t guess what half of the hapaxes mean in any case! Since neither frequent nor infrequent words help, we need to try something else."},{"metadata":{"trusted":true,"_uuid":"c52978e96659d5f1154f4e7dd4d7cffe9b1777de","scrolled":false},"cell_type":"code","source":"fdist1.hapaxes()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2c369d791a5ff24b9449e7445acbc36b9f008cce"},"cell_type":"markdown","source":"**Fine-grained Selection of Words**\n\nLet's find the words from the vocabulary of the text that are more than 15 characters long."},{"metadata":{"trusted":true,"_uuid":"066f3ec67fd5ac4b10ca3059774d33b3432f5f6a"},"cell_type":"code","source":"V = set(text1)\nlong_words = [w for w in V if len(w) > 15]\nsorted(long_words)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"86fe9191b585f63fbf6fe04a350228ef2c2ad316"},"cell_type":"markdown","source":"Here are all words from the chat corpus that are longer than seven characters, that occur more than seven times:"},{"metadata":{"trusted":true,"_uuid":"41a307cf82b6ad5e3424713eba3b2eadcef085b0"},"cell_type":"code","source":"fdist5 = FreqDist(text5)\nsorted(w for w in set(text5) if len(w) > 7 and fdist5[w] > 7)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"487344dde836b8a0195970e3dc3b941b52106453"},"cell_type":"markdown","source":"**Collocations and Bigrams**\n\nA collocation is a sequence of words that occur together unusually often. Thus **red wine is a collocation**, whereas **the wine is not**. To get a handle on collocations, we start off by extracting from a text a list of word pairs, also known as bigrams."},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"458348985aeb6136b8e72e121b5097475389b740"},"cell_type":"code","source":"from nltk import bigrams\nlist(bigrams(['more', 'is', 'said', 'than', 'done']))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6e4a26f470bd90a45c9f4ea8768b566b9c10e735"},"cell_type":"markdown","source":"**Note**\n*If you omitted list() above, and just typed bigrams(['more', ...]), you would have seen output of the form <generator object bigrams at 0x10fb8b3a8>. This is Python's way of saying that it is ready to compute a sequence of items, in this case, bigrams. For now, you just need to know to tell Python to convert it into a list, using list().*"},{"metadata":{"trusted":true,"_uuid":"c21cf8c399ddebcf63f58bd8ad76ec32afc92cfc"},"cell_type":"code","source":"bigrams(['more', 'is', 'said', 'than', 'done'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"84878e5f9db9656f4a021aef1b5a43c104f71af4"},"cell_type":"markdown","source":"**Collocations** are essentially **just frequent bigrams**, except that we want to pay more attention to the cases that involve rare words. In particular, we want to find **bigrams that occur more often** than we would expect based on the frequency of the individual words. The collocations() function does this for us."},{"metadata":{"trusted":true,"_uuid":"00a46a3a5269414bde937ea1f6f7b791daeab421"},"cell_type":"code","source":"text4.collocations()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"501af23707bdc96e7aa8e0ac81946b7ddaed6fdb"},"cell_type":"markdown","source":"We can look at the distribution of word lengths in a text, by creating a FreqDist out of a long list of numbers, where each number is the length of the corresponding word in the text:"},{"metadata":{"trusted":true,"_uuid":"ba1f4c6194d74d52c47488a90348d581f892bd35"},"cell_type":"code","source":"[len(w) for w in text1]  #  A list of the lengths of words in text1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"750ef6cffbacbb5499e41127e51b512b53ce61d4","scrolled":true},"cell_type":"code","source":"fdist = FreqDist(len(w) for w in text1) # Frequency distribution of lengths of words in text1\nprint(fdist)\nfdist","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"732646d7049470be427ca28e0112e4bb4428622a"},"cell_type":"code","source":"print(fdist.most_common()) # A list of frequency disctribution of lengths of words\n\nprint(fdist.max()) # Most frequent word length\n\nprint(fdist[3]) # Frequncy of word length 3\n\nfdist.freq(3) # Percent Frequency of word length 3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b260b621b6d71ceae5a28c2693104faa3733c033"},"cell_type":"markdown","source":"**Functions Defined for NLTK's Frequency Distributions:**\n\nfdist = FreqDist(samples)\t-->   create a frequency distribution containing the given samples <br>\nfdist[sample] += 1\t-->  increment the count for this sample <br>\nfdist['monstrous']\t-->  count of the number of times a given sample occurred <br>\nfdist.freq('monstrous')\t-->  frequency of a given sample <br>\nfdist.N()\t-->  total number of samples <br>\nfdist.most_common(n)\t-->  the n most common samples and their frequencies <br>\nfor sample in fdist:\t-->  iterate over the samples <br>\nfdist.max()\t-->  sample with the greatest count <br>\nfdist.tabulate()\t-->  tabulate the frequency distribution <br>\nfdist.plot()\t-->  graphical plot of the frequency distribution <br>\nfdist.plot(cumulative=True)\t-->  cumulative plot of the frequency distribution <br>\nfdist1 |= fdist2\t-->  update fdist1 with counts from fdist2 <br>\nfdist1 < fdist2\t-->  test if samples in fdist1 occur less frequently than in fdist2 <br>"},{"metadata":{"_uuid":"cc9c9ba0b898e1dfb3eecf3ea2509fb56227302a"},"cell_type":"markdown","source":"**Word Comparison Operator**\n\ns.startswith(t)\t -->  test if s starts with t <br>\ns.endswith(t)\t-->  test if s ends with t <br>\nt in s\t              -->  test if t is a substring of s <br>\ns.islower()\t      -->  test if s contains cased characters and all are lowercase <br>\ns.isupper()\t     -->  test if s contains cased characters and all are uppercase <br>\ns.isalpha()  \t -->  test if s is non-empty and all characters in s are alphabetic <br>\ns.isalnum() \t-->  test if s is non-empty and all characters in s are alphanumeric <br>\ns.isdigit()\t       -->  test if s is non-empty and all characters in s are digits <br>\ns.istitle()     \t-->  test if s contains cased characters and is titlecased (i.e. all words in s have initial capitals) <br>"},{"metadata":{"trusted":true,"_uuid":"0756a7e398479c0178cb5054faa21d7407643156"},"cell_type":"code","source":"sorted(w for w in set(text1) if w.endswith('ableness')) # Operator endswith","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bdeaaca0a0e5a0d85cca53ebc450c61349c6384d"},"cell_type":"code","source":"sorted(term for term in set(text4) if 'gnt' in term) # words with substring 'gnt'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ec1b7ac4fbb2541ce44a8b6717fa2bc8ca5b7001"},"cell_type":"code","source":"sorted(item for item in set(text6) if item.istitle()) # titlecased words","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"01e3fcce958b016769ae4a9cff93798094093fab"},"cell_type":"code","source":"sorted(item for item in set(sent7) if item.isdigit()) #test if s is non-empty and all characters in s are digits","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d6e1d2ff5a9a279bce910c4776870745c0556f5a"},"cell_type":"code","source":"sorted(w for w in set(text7) if '-' in w and 'index' in w)  # Words with two substring(s)/character(s)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2556c8e696bc911d078f70a44e5089730d5480e1"},"cell_type":"code","source":"sorted(wd for wd in set(text3) if wd.istitle() and len(wd) > 10)  # Titlecased words with more than 10 characters","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"83514b362f8377839900713e1c1e47f1573e40be"},"cell_type":"code","source":"sorted(w for w in set(sent7) if not w.islower()) # test if s contains cased characters and all are lowercase. Here with not","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1bca2aa299e0fb25d4a5061b1ba15a51d7a71121"},"cell_type":"code","source":"[len(w) for w in text1]\n[w.upper() for w in text1]  # test if s contains cased characters and all are uppercase","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"991bcbe7016fe5fc4b1560e145386a8538f864fe"},"cell_type":"code","source":"sent1 = ['Call', 'me', 'Nishant', '.']\n\n# Looping with Condition using 'for' & 'if'\nfor xyzzy in sent1:\n    if xyzzy.endswith('t'):\n        print(xyzzy)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}