{"cells":[{"metadata":{"_uuid":"46da1853d8bdd2e761e4f50d4537bb622fa3da0a"},"cell_type":"markdown","source":"## Before we begin:"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"I would like to take the opportunityI appreciate the help of **[Bukun](https://www.kaggle.com/ambarish), [Andrew Lukyanenko](https://www.kaggle.com/artgor/),  [SRK](https://www.kaggle.com/sudalairajkumar) & [Kxx](https://www.kaggle.com/kailex)  ** fot their wonderful kernels :)\n\n** Problem Statement: ** In this competition, we will develop models that identify and flag insincere questions. It is a binary classification problem as our target variable is dichotomoous [0,1] i.e. whether the questions on  quora is insincere quiestion or NOT!\n\nThis Notebook will go through the indepth Extensive Exploratory Data Analysis+ Feature Engg in NLP in R. Afterwards,  we will incorpoarte the FE with binary classification modeling. \n\nIf you like my work, then please don't forget to upvote this Tutorial since it will keep me motivating to perform more in-depth reserach towards NLP and Text Mining.I am hoping you will enjoy the deep exploration into this dataset. \n\nLet the **Insincerity** begin. :P\n"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"# Table of Contents:\n\n\n## 1. [LOAD LIBRARIES](#section1)\n\n## 2. [Basic Preprocessing for Stats](#section2)\n\n## 3. [Overview of Framework](#section3)\n\n## 4. [TOKENIZATION](#section4)\n\n## 5. [Quora questions which contain \"Bangalore\"](#section5)\n\n## 6. [Top 20 words in question Text in Bangalore](#section6)\n\n## 7. [Top 20 Most Common Words in Question Texts](#section7)\n>    ## 7.1. [Top 20 Most Common Words in questions(Uncleaned)](#section7.1)\n>   ## 7.2. [Top 20 Most Common Words in questions(Cleaned with custom stopwords)](#section7.2)\n>    ## 7.3. [WordCloud of the Common Words:( Filtered Custom Stopwords)](#section7.3)\n   \n   \n\n## 8. [Parts of Speech](#section8)\n>    ## 8.1 [World Cloud of Adjective](#section8.1)\n>   ## 8.2 [Transitive Verb Word Cloud](#section8.2)\n>   ## 8.3 [InTransitive Verb Word Cloud](#section8.3)\n   \n   \n   \n## 9. [TF-IDF](#section9)\n>   ## 9.1 [  TF-IDF based on QuestionId](#section9.1)\n\n   \n \n  \n\n## 10. [Sentiment Analysis](#section10)\n>   ## 10.1 [What is Sentiment Analysis-Framework](#section10.1)\n>   ## 10.2 [How does it work](#section10.2)\n>   ## 10.3 [Explore Sentiment Lexicons](#section10.3)\n>   ## 10.4 [Top Contributing words and their correponding NRC sentiment score based on QuestionId](#section10.4)\n>   ## 10.5 [Top Contributing words and their correponding AFINN sentiment score based on QuestionId](#section10.5)\n>   ## 10.6 [Get the sentiment from the first text](#section10.6)\n>   ## 10.7 [BING / NRC /AFINN SENTIMENTS based on Questions:](#section10.7)\n>   ## 10.8 [Overall NRC Sentiment by Questions](#section10.8)\n>   ## 10.9 [Overall BING Sentiment by Questions](#section10.9)\n>   ## 10.10 [Overall AFINN Sentiment by Questions](#section10.10)\n   \n  \n   \n   \n   \n\n## 11.  [ Mood Ring : Relationship b/w Mood & Question Texts based on NRC Sentiment](#section11)\n\n## 12. [Polar Melting](#section12)\n\n  ## WORK IN PROGRESS from below##\n\n\n## 13. [Most common positive and negative words(Cleaned but not Stemmed)](#section13)\n>  ## 13.1 [Word Clouds  based on Bing Sentiments](#section13.1)\n>    ## 13.2[Word Clouds  based on NRC Sentiments ](#section13.2)\n\n\n## 14. [Real-Time Sentiment](#section14)\n>  ## 14. 1[Top Word Classification based on NRC Sentiment of Question Texts](#section14.1)  \n> ## 14. 2[Top Word Classification based on AFINN Sentiment of Question Texts](#section14.2)\n> ## 14. 3[Top Word Classification based on BING Sentiment of Question Texts](#section14.3)\n\n## 15. [Relationships between words: n-grams and correlations](#section15)\n>   ## 15. 1[Tokenizing by n-gram](#section15.1)  \n> ## 15. 2[Analyzing bigrams based on Question Texts](#section15.2)  \n> ## 15. 3[Romance in Question Texts](#section15.3)  \n>  ## 15.4 [Bigram TF-IDF based on Question Texts](#section15.4)  \n>  ## 15.5 [Using bigrams to provide context in sentiment analysis based on Dialogues](#section15.5)  \n>  ## 15.6 [Visualizing a network of bigrams with ggraph](#section15.6)  \n>  ## 15.7 [Directed graph:To determine how the negative conversation in Insecurity Question Texts are moving from one to the other](#section15.7)  \n\n\n## 16. [Topic Modelling](#section16)\n\n## 17. [References](#section17)"},{"metadata":{"_uuid":"72963347d6cd68d08f4edca8c6cd07957554c393"},"cell_type":"markdown","source":"<a id=\"section1\"></a>\n# 1. LOAD LIBRARIES"},{"metadata":{"trusted":true,"_uuid":"3cf4f613090bb0fc760d1e2b18d5ff9c6d0cf6b2"},"cell_type":"code","source":"suppressMessages(library(stringr))\nsuppressMessages(library(DT))\nsuppressMessages(library(igraph))\nsuppressMessages(library(ggraph))\nsuppressMessages(library(tm))\nsuppressMessages(library(wordcloud2))\nsuppressMessages(library(wordcloud))\nsuppressMessages(library(caret))\nsuppressMessages(library(dplyr))\nsuppressMessages(library(magrittr))\nsuppressMessages(library(plyr))\nrm(list=ls())\n\nfillColor = \"#FFA07A\"\nfillColor2 = \"#F1C40F\"\n\nsuppressMessages(library(tidyverse)) # general utility & workflow functions\nsuppressMessages(library(tidytext)) # tidy implimentation of NLP methods\nsuppressMessages(library(topicmodels)) # for LDA topic modelling \nsuppressMessages(library(tm)) # general text mining functions, making document term matrixes\nsuppressMessages(library(SnowballC)) # for stemming\nsuppressMessages(library(glue)) ##for pasting strings\nset.seed(2018)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"134748c67d43e10fe3400c667186284e141e5953"},"cell_type":"code","source":"Quora_IQS_train = suppressMessages(read_csv(\"../input/train.csv\"))\nhead(Quora_IQS_train,2)\n\n\nQuora_IQS_test = suppressMessages(read_csv(\"../input/test.csv\"))\nhead(Quora_IQS_test,2)\n\nQuora_IQS=Quora_IQS_train %>% bind_rows(Quora_IQS_test)\ntail(Quora_IQS_test,2)\n\ntable(Quora_IQS_train$target)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d8f232574fa3c0c99aa5bdf0555dcb79c492ed4c"},"cell_type":"markdown","source":"<a id=\"section2\"></a>\n# 2.Basic Preprocessing for Stats:"},{"metadata":{"_uuid":"e707b3e5da04124a42ece1da397337fd365704e5","trusted":true},"cell_type":"code","source":"#---------------------------\ncat(\"Basic preprocessing & stats of Phrase...\\n\")\nQuora_IQS <- Quora_IQS %>% \n  mutate(length = str_length(question_text),\n         ncap = str_count(question_text, \"[A-Z]\"),\n         ncap_len = ncap / length,\n         nexcl = str_count(question_text, fixed(\"!\")),\n         nquest = str_count(question_text, fixed(\"?\")),\n         npunct = str_count(question_text, \"[[:punct:]]\"),\n         nword = str_count(question_text, \"\\\\w+\"),\n         nsymb = str_count(question_text, \"&|@|#|\\\\$|%|\\\\*|\\\\^\"),\n         nsmile = str_count(question_text, \"((?::|;|=)(?:-)?(?:\\\\)|D|P))\")) \n\nhead(Quora_IQS,2)\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"68984420c7864a80d3efc7faad2a02851bfd665e"},"cell_type":"markdown","source":"<a id=\"section3\"></a>\n# 3.  Overview of Framework:"},{"metadata":{"_uuid":"babd897293beb3ba896c4cbd4241f3fbcfdfb621"},"cell_type":"markdown","source":"<img src=\"https://www.tidytextmining.com/images/tidyflow-ch-5.png\" width=100%>"},{"metadata":{"_uuid":"b167075462e56663d91884f3658c32a8411bfb2a"},"cell_type":"markdown","source":"<a id=\"section4\"></a>\n# 4. TOKENIZATION:\nOne common task in NLP (Natural Language Processing) is tokenization. \"Tokens\" are usually individual words (at least in languages like English) and \"tokenization\" is taking a text or set of text and breaking it up into individual its words. These tokens are then used as the input for other types of analysis or tasks, like parsing (automatically tagging the syntactic relationship between words).\n\nIn this tutorial you'll learn how to:\n\nRead text into R Select only certain lines Tokenize text using the tidytext package Calculate token frequency (how often each token shows up in the dataset) Write reusable functions to do all of the above and make your work reproducible For this tutorial we'll be using a corpus of transcribed speech from bi-lingual children speaking in English.\n\nThis dataset of kid's speech is really cool, but it's in a bit of a weird file format. These files were generated by CLAN, a specialized program for transcribing children's speech. Under the hood, however, they're just text files with some additional formatting. With a little text processing we can just treat them like raw text files.\n\nLet's do that, and find out if there's a relationship between how often different children use disfluencies (words like \"um\" or \"uh\") and how long they've been exposed to English.\n\nwe have a tibble of sentences(dialogues) that the each charcater said.\n\nLet's start by making our data tidy. Tidy data has three qualities:\n\nEach variable forms a column.\n\nEach observation forms a row.\n\nEach type of observational unit forms a table.\n\nFortunately, we don't have to start tidying from scratch, we can use the tidytext package!"},{"metadata":{"_uuid":"7576a109bb3e2ba90a17dfb9f7949f1474ee62cd"},"cell_type":"markdown","source":"<img src=\"https://www.tidytextmining.com/images/tidyflow-ch-1.png\" width=100%>"},{"metadata":{"trusted":true,"_uuid":"1868c02ba8c0a37d1bcbb46ac5d4ec21b6292669"},"cell_type":"code","source":"# use the unnest_tokens function to get the words from the \"value\" column of \"child\nQIQS_Tokens <- Quora_IQS %>% unnest_tokens(word, question_text)\nhead(QIQS_Tokens,3)\n \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b08c234cd0e0db5abc3b051bbd8672bc507f5077"},"cell_type":"markdown","source":"<a id=\"section5\"></a>\n# 5. Quora Questions which contain \"Bangalore\":\n\n    P.S: I am based out of Bangalore,India. Just interested to see the Quora Questions from there :)"},{"metadata":{"trusted":true,"_uuid":"339248be01e66e66df0c7d40fb18eec92155e9af"},"cell_type":"code","source":"Bangalore= QIQS_Tokens%>% filter(word %in% \"bangalore\")\nBangalore <- Bangalore[,1] %>% left_join(Quora_IQS,\"qid\")\nhead(Bangalore)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"00c403e02deb95cd8d2758a8a8c4584140b83f69"},"cell_type":"markdown","source":"<a id=\"section6\"></a>\n# 6. Top 20 words in question Text in Bangalore"},{"metadata":{"trusted":true,"_uuid":"5f5022410ef53a60688d5d02cfa646b93aac77b0"},"cell_type":"code","source":"custom_stop_words <- tibble(word = c(\"bangalore\", \"Bangalore\"))\nBangalore %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n  filter(!word %in% custom_stop_words$word) %>%\n  dplyr:: count(word,sort = TRUE) %>%\n  ungroup() %>%\n  mutate(word = factor(word, levels = rev(unique(word)))) %>%\n  head(20) %>%\n  ggplot(aes(x = word,y = n)) +\n  geom_bar(stat='identity',colour=\"white\", fill =fillColor) +\n  geom_text(aes(x = word, y = 1, label = paste0(\"(\",n,\")\",sep=\"\")),\n            hjust=0, vjust=.5, size = 4, colour = 'black',\n            fontface = 'bold') +\n  labs(x = 'Word', y = 'Word Count', \n       title = 'Top 20 most Common Words in Bangalore question text') +\n  coord_flip() + \n  theme_bw()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bd9c2b51c7f14846227bafa905ffa1caa0ce6d68"},"cell_type":"code","source":"Bangalore %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n  filter(!word %in% custom_stop_words$word) %>%\n  dplyr::count(word,sort = TRUE) %>%\n  ungroup()  %>%\n  head(50) %>%\n  \n  with(wordcloud(word, n, max.words = 50,colors=brewer.pal(8, \"Dark2\")))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"edb6872a8bef14a7d07bec292de870a55eb3ed4d"},"cell_type":"markdown","source":"<a id=\"section7\"></a>\n# 7. Top 20 Most Common Words in Question Texts"},{"metadata":{"_uuid":"c879f75dd08326563390a55cfd37c11e39876f7c"},"cell_type":"markdown","source":"<a id=\"section7.1\"></a>\n# 7.1 Top 20 most Common Words in Question Texts (Uncleaned):"},{"metadata":{"trusted":true,"_uuid":"643f0676cf88224206069a6cc5db22f83f534821"},"cell_type":"code","source":"Quora_IQS %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n  dplyr:: count(word,sort = TRUE) %>%\n  ungroup() %>%\n  mutate(word = factor(word, levels = rev(unique(word)))) %>%\n  head(20) %>%\n  ggplot(aes(x = word,y = n)) +\n  geom_bar(stat='identity',colour=\"white\", fill =fillColor) +\n  geom_text(aes(x = word, y = 1, label = paste0(\"(\",n,\")\",sep=\"\")),\n            hjust=0, vjust=.5, size = 4, colour = 'black',\n            fontface = 'bold') +\n  labs(x = 'Word', y = 'Word Count', \n       title = 'Top 20 most Common Words') +\n  coord_flip() + \n  theme_bw()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2ee1fb8ae6c7e05fa172721ef85c183295f55a6e"},"cell_type":"markdown","source":"**Since it's a Quora Insincere Qustions classification, therefore it's expected that 'quora', word   will be used in most qustions[It comes in Top 8th]. But it isn't  an important words in classification. We need to remove those unnecssary words from our phrase ti get a cleaned document.**"},{"metadata":{"_uuid":"1a9bd6bd0555598ed85f90c2437d9f3eb4481b84"},"cell_type":"markdown","source":"<a id=\"section7.2\"></a>\n# 7.2Top 20 most Common Words in Question Texts (Cleaned with custom stopwords):"},{"metadata":{"trusted":true,"_uuid":"e00c3e64ae8ea669cfd7c92d962f86296e41be8a"},"cell_type":"code","source":"custom_stop_words <- tibble(word = c(\"Quora\", \"quora\",'2'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"93de660a62d642aee14a93a66dec06c228ecc482"},"cell_type":"code","source":"Quora_IQS %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n  filter(!word %in% custom_stop_words$word) %>%\n  dplyr:: count(word,sort = TRUE) %>%\n  ungroup() %>%\n  mutate(word = factor(word, levels = rev(unique(word)))) %>%\n  head(20) %>%\n  ggplot(aes(x = word,y = n)) +\n  geom_bar(stat='identity',colour=\"white\", fill =fillColor) +\n  geom_text(aes(x = word, y = 1, label = paste0(\"(\",n,\")\",sep=\"\")),\n            hjust=0, vjust=.5, size = 4, colour = 'black',\n            fontface = 'bold') +\n  labs(x = 'Word', y = 'Word Count', \n       title = 'Top 20 most Common Words') +\n  coord_flip() + \n  theme_bw()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b81b6e8b0873af04bf082f6035a5b8266057d13"},"cell_type":"markdown","source":"<a id=\"section7.3\"></a>\n# 7.3WordCloud of the Common Words:( Filtered Custom Stopwords)"},{"metadata":{"trusted":true,"_uuid":"13052cd43b1833f39dc291f2edf32ff31df1ccbb"},"cell_type":"code","source":"Quora_IQS %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n  filter(!word %in% custom_stop_words$word) %>%\n  dplyr::count(word,sort = TRUE) %>%\n  ungroup()  %>%\n  head(50) %>%\n  \n  with(wordcloud(word, n, max.words = 50,colors=brewer.pal(8, \"Dark2\")))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d7aee8cc4df111a0e1dd3c84d037572c02b19cc2"},"cell_type":"markdown","source":"<a id=\"section8\"></a>\n# 8.Parts of Speech"},{"metadata":{"trusted":true,"_uuid":"b712b858aaee717e8bb3393f31f86aecc47dc948"},"cell_type":"code","source":"glimpse(parts_of_speech)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"672f4c7857f40aee60165d43dcc348896ac6af1b"},"cell_type":"markdown","source":"<a id=\"section8.1\"></a>\n# Different Parts of Speech\n\n# 8.1World Cloud of Adjective:"},{"metadata":{"trusted":true,"_uuid":"08ccf607ba1f45c909007557bbdff8541d1e1ea2"},"cell_type":"code","source":"Quora_IQS %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n   filter(!word %in% custom_stop_words$word) %>%\n  left_join(parts_of_speech) %>%\n  filter(pos == \"Adjective\") %>%\n  \n  dplyr:: count(word,sort = TRUE) %>%\n  ungroup()  %>%\n  head(50) %>%\n  \n  with(wordcloud(word, n, max.words = 50,colors=brewer.pal(8, \"Dark2\")))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"929bf3b1922d4b623ce7f155dca5644f8b5d872f"},"cell_type":"markdown","source":"<a id=\"section8.2\"></a>\n# 8.2Transitive Verb Word Cloud:"},{"metadata":{"trusted":true,"_uuid":"719c159d164bde6646fbec9e53b2cbd4f91429d3"},"cell_type":"code","source":"Quora_IQS %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n   filter(!word %in% custom_stop_words$word) %>%\n  left_join(parts_of_speech) %>%\n  filter(pos == \"Verb (transitive)\") %>%\n  \n  dplyr:: count(word,sort = TRUE) %>%\n  ungroup()  %>%\n  head(50) %>%\n  \n  with(wordcloud(word, n, max.words = 50,colors=brewer.pal(8, \"Dark2\")))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"29b509e3cfc47cd9d820edac40ad7f9b63a49753"},"cell_type":"markdown","source":"<a id=\"section8.3\"></a>\n# 8.3 Intransitive Verb Word Cloud:"},{"metadata":{"trusted":true,"_uuid":"1e4b2e4b28356b45fbf7058e0e77f4e73e4fc20b"},"cell_type":"code","source":"Quora_IQS %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n   filter(!word %in% custom_stop_words$word) %>%\n  left_join(parts_of_speech) %>%\n  filter(pos == \"Verb (intransitive)\") %>%\n  \n  dplyr:: count(word,sort = TRUE) %>%\n  ungroup()  %>%\n  head(50) %>%\n  \n  with(wordcloud(word, n, max.words = 50,colors=brewer.pal(8, \"Dark2\")))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9d9e0185b9c47c6374de498924588dc18f1ad51c"},"cell_type":"markdown","source":"<a id=\"section9\"></a>\n# 9.TF-IDF\nWe wish to find out the important words which are spoken by the characters. Example for your young child , the most important word is mom. Example for a bar tender , important words would be related to drinks.\n\nWe would explore this using a fascinating concept known as Term Frequency - Inverse Document Frequency. Quite a mouthful, but we will unpack it and clarify each and every term.\n\nA document in this case is the set of lines spoken by a character.\n\nE.g. The words spoken by Vader is a single document.The words spoken by Luke is a another document.\n\nTherefore we have different documents for each Character.\n\n**TF-IDF computes a weight which represents the importance of a term inside a document**.It does this by comparing the frequency of usage inside an individual document as opposed to the entire data set (a collection of documents). The importance increases proportionally to the number of times a word appears in the individual document itself–this is called Term Frequency. However, if multiple documents contain the same word many times then you run into a problem. That’s why TF-IDF also offsets this value by the frequency of the term in the entire document set, a value called Inverse Document Frequency.\n\nTF(t) = (Number of times term t appears in a document) / (Total number of terms in the document) IDF(t) = log_e(Total number of documents / Number of documents with term t in it).\n\n# Value = TF * IDF\nTwenty Most Important words for the Twenty Most Active Characters Here using TF-IDF , we investigate the Twenty Most Important words for the Twenty Most Active Question Texts."},{"metadata":{"_uuid":"eba840818d625a1695f5d13edcff5ee154b1e0ae"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"981720eb79435d2c384d5fd7b7dce5e2d452aee9"},"cell_type":"markdown","source":"<a id=\"section9.1\"></a>\n# 9.1TF-IDF based on QuestionId:"},{"metadata":{"trusted":true,"_uuid":"d12ba1a46a200e924f7a1cd8ba1128d77c27be79"},"cell_type":"code","source":"Topqids = Quora_IQS %>%\n  group_by(qid) %>%\n  tally(sort = TRUE) \n#Get the Top 20 qids\nTop20qids = Topqids[1:10,]$qid\n\n##################################################################################\n\n# Prepare for the bind_tf_idf function\n\n##################################################################################\n\n\nSW_Words <- Quora_IQS %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n  filter(!word %in% custom_stop_words$word) %>%\n  dplyr:: count(qid, word, sort = TRUE) %>%\n  ungroup()\n\ntotal_words <- SW_Words %>% \n  group_by(qid) %>% \n  dplyr::summarize(total = sum(n))\n\nSW_Words <- left_join(SW_Words, total_words)\n\nSW_WordsFull <- SW_Words %>%\n  filter(!is.na(qid)) %>%\n  bind_tf_idf(word, qid, n)\n\n#Now we are ready to use the bind_tf_idf which computes the tf-idf for each term.    \n\nSW_Words <- SW_Words %>% filter( qid %in% Top20qids) %>%\n  bind_tf_idf(word, qid, n)\n\nplot_Sw_Words <- SW_Words %>%\n  arrange(desc(tf_idf)) %>%\n  mutate(word = factor(word, levels = rev(unique(word))))\n\nplot_Sw_Words %>% \n  top_n(20) %>%\n  ggplot(aes(word, tf_idf, fill = qid)) +\n  geom_col() +\n  labs(x = NULL, y = \"tf-idf\") +\n  coord_flip() +\n  theme_bw()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f844bab74b7fa84b77e8193bf9a77446efe17c4e"},"cell_type":"code","source":"#Choose words with low IDF\nSW_Words_2 <- SW_Words %>%\n  bind_tf_idf(word, qid, n)\n\nLowIDF = SW_Words_2 %>%\n  arrange((idf)) %>%\n  select(word,idf)\n\n#Get the Unique Words with LowIDF\nUniqueLowIDF = unique(LowIDF$word)\nhead(UniqueLowIDF)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6f896050b89931975082696f0cb6bc1092f40851"},"cell_type":"markdown","source":"Therefore **'people', 'affect', 'velocity', 'dressing', '1960s' ,'nation'** -- these are the words with UniqueLowIDF based on PhraseIds i.e. they appears least in different documents(**questiontexts**)"},{"metadata":{"_uuid":"078bd1cb5c53262736b13f96bc1878c79ddfb699"},"cell_type":"markdown","source":"<a id=\"section10\"></a>\n# 10. Sentiment Analysis"},{"metadata":{"trusted":true,"_uuid":"e4039b211f68b7b11756b0bdf01b676b5675c7e4"},"cell_type":"markdown","source":"<a id=\"section10.1\"></a>\n# 10.1What is sentiment analysis?\n\nSentiment analysis is the computational task of automatically determining what feelings a writer is expressing in text. Sentiment is often framed as a binary distinction (positive vs. negative), but it can also be a more fine-grained, like identifying the specific emotion an author is expressing (like fear, joy or anger).\n\nSentiment analysis is used for many applications, especially in business intelligence. Some examples of applications for sentiment analysis include:\n\nAnalyzing the social media discussion around a certain topic Evaluating survey responses Determining whether product reviews are positive or negative Sentiment analysis is not perfect, and as with any automatic analysis of language, you will have errors in your results. It also cannot tell you why a writer is feeling a certain way. However, it can be useful to quickly summarize some qualities of text, especially if you have so much text that a human reader cannot analyze all of it."},{"metadata":{"trusted":true,"_uuid":"4d2932dc79ec4ac6176650318ad4b55d9958faee"},"cell_type":"markdown","source":"<img src=\"https://www.tidytextmining.com/images/tidyflow-ch-2.png\" width=100%>"},{"metadata":{"trusted":true,"_uuid":"5cc23b6210fcf8d55951d5e0a33ada987cdd8d9a"},"cell_type":"markdown","source":"<a id=\"section10.2\"></a>\n# 10.2 How does it work?\n\nThere are many ways to do sentiment analysis (if you're interested, you can see many of them here). Many approches use the same general idea, however:\n\nCreate or find a list of words associated with strongly positive or negative sentiment. Count the number of positive and negative words in the text. Analyze the mix of positive to negative words. Many positive words and few negative words indicates positive sentiment, while many negative words and few positive words indicates negative sentiment. The first step, creating or finding a word list (also called a lexicon), is generally the most time-consuming. While you can often use a lexicon that already exists, if your text is discussing a specific topic you may need to add to or modify it.\n\n\"Sick\" is an example of a word that can have positive or negative sentiment depending on what it's used to refer to. If you're discussing a pet store that sells a lot of sick animals, the sentiment is probably negative. On the other hand, if you're talking about a skateboarding instructor who taught you how to do a lot of sick flips, the sentiment is probably very positive."},{"metadata":{"_uuid":"932cb44e5fa98aea5be593e9ed2f42e83f5ee103"},"cell_type":"markdown","source":"<a id=\"section10.3\"></a>\n# 10.3 Explore Sentiment Lexicons\nThe tidytext package includes a dataset called sentiments which provides several distinct lexicons. These lexicons are dictionaries of words with an assigned sentiment category or value. tidytext provides three general purpose lexicons:\n\n**AFINN**: assigns words with a score that runs between** -5 and 5**, with negative scores indicating negative sentiment and positive scores indicating positive sentiment\n\n**Bing**: assigns words into **positive and negative **categories\n\n**NRC**: assigns words into one or more of the following ten categories: **positive, negative, anger, anticipation, disgust, fear, joy, sadness, surprise, and trust**\n\nIn order to examine the lexicons, create a data frame called new_sentiments. Filter out a financial lexicon, create a binary (also described as polar) sentiment field for the AFINN lexicon by converting the numerical score to positive or negative, and add a field that holds the distinct word count for each lexicon."},{"metadata":{"_uuid":"7d0b2b0b5222e1c51a4a71ec1730da18923224e4"},"cell_type":"markdown","source":"<a id=\"section10.4\"></a>\n# 10.4 Top Contributing words and their correponding NRC sentiment score based on QuestionId:"},{"metadata":{"trusted":true,"_uuid":"0be934f71953a7d4c78d0321e4c884b829d3c69a"},"cell_type":"code","source":"contributions <- Quora_IQS %>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n  filter(!word %in% custom_stop_words$word) %>%\n  filter(qid != \"NA\") %>%\n  dplyr::count(qid, word, sort = TRUE) %>%\n  ungroup() %>%\n  \n  inner_join(get_sentiments(\"bing\"), by = \"word\") %>%\n  group_by(word) %>%\n  dplyr::summarize(occurences = n(),\n                   contribution = sum(n))\n\ncontributions %>%\n  top_n(20, abs(contribution)) %>%\n  mutate(word = reorder(word, contribution)) %>%\n  ggplot(aes(word, contribution, fill = contribution > 0)) +\n  geom_col(show.legend = FALSE) +\n  coord_flip() + theme_bw()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"42e0eb7e6e83468cebf4c26c6058b6b7addd87d5"},"cell_type":"markdown","source":"Therefore, we know that contribution of** \"trump\" and \"love\" are the highest & 2nd highest sentiment score**. But from NRC sentiment iverall contribution we dont know which one is having +ve impact(+ve score ) and which one is  having -ve impact(-ve score) on reviews. **To distiguish b/w +ve and -ve score we need AFINN Sentiment score,**"},{"metadata":{"trusted":true,"_uuid":"caa476fe60a3b82a1006ee70b2191c2e0b0bcd3c"},"cell_type":"markdown","source":"<a id=\"section10.5\"></a>\n# 10.5Top Contributing words and their correponding AFINN sentiment score based on QuestionId\n\nWe investigate how often positive and negative words occurred in these episodes. Which Reviews were the most positive or negative overall?\n\nWe will use the AFINN sentiment lexicon, which provides numeric positivity scores for each word, and visualize it with a bar plot.We limit the sentiment analysis to the Top 20 characters who have spoken the most in the episodes."},{"metadata":{"_uuid":"e2e7e03289f682700b840b70f5e1fe059cee3ca3","trusted":true},"cell_type":"code","source":"contributions <- Quora_IQS%>%\n  unnest_tokens(word, question_text) %>%\n  filter(!word %in% stop_words$word) %>%\n  filter(!word %in% custom_stop_words$word) %>%\n  filter(qid != \"NA\") %>%\n  dplyr::count(qid, word, sort = TRUE) %>%\n  ungroup() %>%\n  \n  inner_join(get_sentiments(\"afinn\"), by = \"word\") %>%\n  group_by(word) %>%\n  dplyr::summarize(occurences = n(),\n                   contribution = sum(score))\n\ncontributions %>%\n  top_n(20, abs(contribution)) %>%\n  mutate(word = reorder(word, contribution)) %>%\n  ggplot(aes(word, contribution, fill = contribution > 0)) +\n  geom_col(show.legend = FALSE) +\n  coord_flip() + theme_bw()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2936d28c62cf9b70044b12429c93779a6e1b89c7"},"cell_type":"markdown","source":"Now we can clearly distinguish the +ve & -ve impact of words on insecurity questions.\n\n**Blue** marked words are having **+ve sentiment scores** & **red** marked are having **-ve scores**. \n\n\"**love**\" is having the most +ve score whereas **bad** is most -ve."},{"metadata":{"_uuid":"524fe6b6ba240212ec5e30ec29925b1e8ede2a8a"},"cell_type":"markdown","source":"<a id=\"section10.6\"></a>\n#  10.6 Get the sentiment from the first text: "},{"metadata":{"trusted":true,"_uuid":"573382adf0c74643c9f87a42967ebb8c2f69c73d"},"cell_type":"code","source":"\ncontributions %>%\n  filter(!word %in% stop_words$word) %>%\n  filter(!word %in% custom_stop_words$word) %>%\n  inner_join(get_sentiments(\"bing\")) %>% # pull out only sentiment words\n  dplyr::count(sentiment) %>% # count the # of positive & negative words\n  spread(sentiment, n, fill = 0) %>% # made data wide rather than narrow\n  mutate(sentiment = positive - negative) # # of positive words - # of negative owrds","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"86a135279fafe4f3d1c6380b1d1667ce51550300"},"cell_type":"markdown","source":" So this text has **847 negative polarity words and  407 positive polarity words**. This means that there are -440 more negative than positive words in this text.\n\n **Wohh!! Negative sentiments are humongous in Quora Insecurity question texts. So, we can come into one conclusion that, insincerity leads us to negative sentiments i.e. psychologically insincerity creates a negative impact. Suprising isn't it!!**\n \n Now that we know how to get the sentiment for a given text, let's write a function to do this more quickly and easily and then apply that function to every text in our dataset."},{"metadata":{"trusted":true,"_uuid":"16a05e04aab949074a70ff73ec9d78ce1fc0311b"},"cell_type":"code","source":"suppressMessages(library(widyr)) #Use for pairwise correlation\n\n#Visualizations!\nsuppressMessages(library(ggplot2)) #Visualizations (also included in the tidyverse package)\nsuppressMessages(library(ggrepel)) #`geom_label_repel`\nsuppressMessages(library(gridExtra)) #`grid.arrange()` for multi-graphs\nsuppressMessages(library(knitr)) #Create nicely formatted output tables\nsuppressMessages(library(kableExtra)) #Create nicely formatted output tables\nsuppressMessages(library(formattable)) #For the color_tile function\nsuppressMessages(library(circlize)) #Visualizations - chord diagram\nsuppressMessages(library(memery)) #Memes - images with plots\nsuppressMessages(library(magick)) #Memes - images with plots (image_read)\nsuppressMessages(library(yarrr))  #Pirate plot\nsuppressMessages(library(radarchart)) #Visualizations\nsuppressMessages(library(igraph)) #ngram network diagrams\nsuppressMessages(library(ggraph)) #ngram network diagrams\n\n#Define some colors to use throughout\nmy_colors <- c(\"#E69F00\", \"#56B4E9\", \"#009E73\", \"#CC79A7\", \"#D55E00\", \"#D65E00\")\n\n#Customize ggplot2's default theme settings\n#This tutorial doesn't actually pass any parameters, but you may use it again in future tutorials so it's nice to have the options\ntheme_lyrics <- function(aticks = element_blank(),\n                         pgminor = element_blank(),\n                         lt = element_blank(),\n                         lp = \"none\")\n{\n  theme(plot.title = element_text(hjust = 0.5), #Center the title\n        axis.ticks = aticks, #Set axis ticks to on or off\n        panel.grid.minor = pgminor, #Turn the minor grid lines on or off\n        legend.title = lt, #Turn the legend title on or off\n        legend.position = lp) #Turn the legend on or off\n}\n\n#Customize the text tables for consistency using HTML formatting\nmy_kable_styling <- function(dat, caption) {\n  kable(dat, \"html\", escape = FALSE, caption = caption) %>%\n  kable_styling(bootstrap_options = c(\"striped\", \"condensed\", \"bordered\"),\n                full_width = FALSE)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3cd4b494ca2e5e6bfdd78b1ab8916a63fe00b4a4"},"cell_type":"markdown","source":"<a id=\"section10.7\"></a>\n# 10.7 BING / NRC /AFINN SENTIMENTS based on Question Text:"},{"metadata":{"trusted":true,"_uuid":"436657a72e42ec6ee685daacf25f0bda47fff621"},"cell_type":"code","source":"SW_tidy <- Quora_IQS %>%\n  unnest_tokens(word, question_text) %>% #Break the Phrase into individual words\n  filter(!nchar(word) < 3) %>% \n  anti_join(stop_words) %>%\n  filter(!word %in% custom_stop_words$word)\n\nSW_bing <- SW_tidy %>%\n  inner_join(get_sentiments(\"bing\"))\nprint(\"BING SENTIMETNTS by Question Texts:\")\nhead(SW_bing,2)\ntail(SW_bing,2)\n\n\nSW_nrc <- SW_tidy %>%\n  inner_join(get_sentiments(\"nrc\"))\nprint(\"Overall NRC SENTIMETNTS by Question Texts:\")\nhead(SW_nrc,2)\n\nSW_nrc_sub <- SW_tidy %>%\n  inner_join(get_sentiments(\"nrc\")) %>%\n  filter(!sentiment %in% c(\"positive\", \"negative\"))\nprint(\"Neutral NRC SENTIMETNTS by Question Texts:\")\ntail(SW_nrc_sub,2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3e0d794dad7b955687f00be472294b1f2f8daf33"},"cell_type":"markdown","source":"<a id=\"section10.8\"></a>\n# 10.8 Overall NRC Sentiment by Question Texts:"},{"metadata":{"trusted":true,"_uuid":"910bb985045d49ad96e8484d0715dda0752f0abb"},"cell_type":"code","source":"SW_tidy <- Quora_IQS %>%\n  unnest_tokens(word, question_text) %>% #Break the Phrase into individual words\n  filter(!nchar(word) < 3) %>% \n  anti_join(stop_words) %>%\n  filter(!word %in% custom_stop_words$word)\n\nSW_nrc <- SW_tidy %>%\n  inner_join(get_sentiments(\"nrc\"))\n\n\nnrc_plot <- SW_nrc %>%\n  group_by(sentiment) %>%\n  dplyr::summarise(word_count = n()) %>%\n  ungroup() %>%\n  mutate(sentiment = reorder(sentiment, word_count)) %>%\n  #Use `fill = -word_count` to make the larger bars darker\n  ggplot(aes(sentiment, word_count, fill = -word_count)) +\n  geom_col() +\n  guides(fill = FALSE) + #Turn off the legend\n  theme_lyrics() +\n  labs(x = NULL, y = \"Word Count\") +\n  #scale_y_continuous(limits = c(0, 5000)) + #Hard code the axis limit\n  ggtitle(\"Overall NRC Sentiment by Question Texts\") +\n  coord_flip()\n\nnrc_plot","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b7a080589cec77f9007cfd0506e322af40e9ac3c"},"cell_type":"markdown","source":"****Therefore, we can can conclude that positive and tustworthy questions have surpassed the negative sentiments although the question texts are related to the insincerity from different ids(people).****"},{"metadata":{"_uuid":"b47fabf7ac955b6b009562e574acec2fe743e94f"},"cell_type":"markdown","source":"<a id=\"section10.9\"></a>\n# 10.9 Overall BING Sentiment by Question Texts:"},{"metadata":{"trusted":true,"_uuid":"249ab836c5063cf3f938f5d25c9e3f0014796161"},"cell_type":"code","source":"SW_tidy <- Quora_IQS %>%\n  unnest_tokens(word, question_text) %>% #Break the Phrase into individual words\n  filter(!nchar(word) < 3) %>% \n  anti_join(stop_words) %>%\n  filter(!word %in% custom_stop_words$word)\n\nSW_nrc <- SW_tidy %>%\n  inner_join(get_sentiments(\"bing\"))\n\n\nbing_plot <- SW_nrc %>%\n  group_by(sentiment) %>%\n  dplyr::summarise(word_count = n()) %>%\n  ungroup() %>%\n  mutate(sentiment = reorder(sentiment, word_count)) %>%\n  #Use `fill = -word_count` to make the larger bars darker\n  ggplot(aes(sentiment, word_count, fill = -word_count)) +\n  geom_col() +\n  guides(fill = FALSE) + #Turn off the legend\n  theme_lyrics() +\n  labs(x = NULL, y = \"Word Count\") +\n  #scale_y_continuous(limits = c(0, 5000)) + #Hard code the axis limit\n  ggtitle(\"Overall BING Sentiment by Question Texts\") +\n  coord_flip()\n\nbing_plot","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"15cfad74e8097141f87bf0d43b1536d5d14832f6"},"cell_type":"markdown","source":"**But from the Bing sentiments it is quite clear that overall the sentiments are mostly negative while clubbing all  the insecured texts as +ve and -ve only**"},{"metadata":{"_uuid":"63fa21acea721395c279ffc84f9a6a099adddf01"},"cell_type":"markdown","source":"<a id=\"section10.10\"></a>\n# 10.10 Overall AFINN Sentiment by Question Texts:"},{"metadata":{"trusted":true,"_uuid":"6f979fabef83bdb781f500a094699a71587ce002"},"cell_type":"code","source":"SW_tidy <- Quora_IQS %>%\n  unnest_tokens(word, question_text) %>% #Break the Phrase into individual words\n  filter(!nchar(word) < 3) %>% \n  anti_join(stop_words) %>%\n  filter(!word %in% custom_stop_words$word)\n\nSW_nrc <- SW_tidy %>%\n  inner_join(get_sentiments(\"afinn\"))\n\n\nafinn_plot <- SW_nrc %>%\n  group_by(score) %>%\n  dplyr::summarise(word_count = n()) %>%\n  ungroup() %>%\n  mutate(sentiment = reorder(score, word_count)) %>%\n  #Use `fill = -word_count` to make the larger bars darker\n  ggplot(aes(score, word_count, fill = -word_count)) +\n  geom_col() +\n  guides(fill = FALSE) + #Turn off the legend\n  theme_lyrics() +\n  labs(x = NULL, y = \"Word Count\") +\n  #scale_y_continuous(limits = c(0, 5000)) + #Hard code the axis limit\n  ggtitle(\"Overall AFINN Sentiment by Question Texts\") +\n  coord_flip()\n\nafinn_plot","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a083eb7e1c8b7ebb5608048a1d9a82ac9bea0339"},"cell_type":"markdown","source":"**AFINN Sentiment Score lies within [-5,5]**.\n\n**+2 is the most positive score based on occurences of words in different sentences and different phrases, whereas -2 is the most negative score.**"},{"metadata":{"_uuid":"444c1669de2cec875784b5055e3cf93029cd6a0f"},"cell_type":"markdown","source":"<a id=\"section11\"></a>\n# 11. Mood Ring : Relationship b/w Mood & Insincere Questions based on NRC Sentiment: "},{"metadata":{"trusted":true,"_uuid":"0e62171c98f80d6bf9e0789fc015eac44552040c"},"cell_type":"code","source":"new_sentiments <- sentiments %>% #From the tidytext package\n  filter(lexicon != \"loughran\") %>% #Remove the finance lexicon\n  mutate( sentiment = ifelse(lexicon == \"AFINN\" & score >= 0, \"positive\",\n                             ifelse(lexicon == \"AFINN\" & score < 0,\n                                    \"negative\", sentiment))) %>%\n  group_by(lexicon) %>%\n  mutate(words_in_lexicon = n_distinct(word)) %>%\n  ungroup()\n\nhead(new_sentiments)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d26911892763d895b4ed132e4a3459b27abc8a51"},"cell_type":"code","source":"grid.col = c(\"IV\" = my_colors[1], \"V\" = my_colors[2], \"VI\" = my_colors[3], \"anger\" = \"red\", \"anticipation\" = \"blue\", \"disgust\" = \"black\", \"fear\" = \"yellow\", \"joy\" = \"green\", \"sadness\" = \"grey\", \"surprise\" = \"violet\", \"trust\" = \"pink\")\nSW_tidy <- Quora_IQS %>%\n  unnest_tokens(word, question_text) %>% #Break the Phrase into individual words\n  filter(!nchar(word) < 3) %>% \n  anti_join(stop_words) %>%\n  anti_join(custom_stop_words)\n\nSW_nrc <- SW_tidy %>%\n  inner_join(get_sentiments(\"nrc\"))\n\n\n\nmood <-  SW_nrc %>%\n  dplyr:: count(sentiment) %>%\n  group_by( sentiment) %>%\n  #summarise(sentiment_sum = sum(n)) %>%\n  ungroup()\n\ncircos.clear()\n#Set the gap size\ncircos.par(gap.after = c(rep(5, length(unique(mood[[1]])) - 1), 15,\n                         rep(5, length(unique(mood[[2]])) - 1), 15))\nchordDiagram(mood, grid.col = grid.col, transparency = .2)\ntitle(\"Mood Ring : Relationship b/w Mood & Insincere Questions on NRC Sentiment\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8252a9534124d1407f1f143c6d29330a5d7b03f1"},"cell_type":"markdown","source":"<a id=\"section12\"></a>\n# 12.Polar Melting: \n\nSince you're looking at sentiment from a polar perspective, you might want to see weather or not the **lexicon sentiment **changes over the **Actual Sentiment(tagged)  of train data******. \n\nThis time use geom_smooth() with the loess method for a smoother curve and another geom_smooth() with method = lm for a linear smooth curve."},{"metadata":{"trusted":true,"_uuid":"c480a8cf861f569b673f9e5114d6afa63b00c2f4"},"cell_type":"code","source":"#unique(SW_bing$target)\n\nSW_bing<-SW_bing %>% filter(SW_bing$target %in% c(0,1))\n head(SW_bing)   \n\nSW_polarity_Sentiment <- SW_bing %>%\n  dplyr::count(sentiment, target) %>%\n  spread(sentiment, n, fill = 0) %>%\n  mutate(polarity = positive - negative,\n    percent_positive = positive / (positive + negative) * 100)\n\npolarity_over_Sentiment <- SW_polarity_Sentiment %>%\n  ggplot(aes(target, polarity, color = ifelse(polarity >= 0,my_colors[5],my_colors[4]))) +\n  geom_col() +\n  geom_smooth(method = \"loess\", se = FALSE) +\n  geom_smooth(method = \"lm\", se = FALSE, aes(color = my_colors[1])) +\n  theme_lyrics() + theme(plot.title = element_text(size = 11)) +\n  xlab(NULL) + ylab(NULL) +\n  ggtitle(\"Polarity Over Actual Sentiment of Train\")\n\n\n\nrelative_polarity_over_Sentiment <- SW_polarity_Sentiment %>%\n  ggplot(aes(target, percent_positive , color = ifelse(polarity >= 0,my_colors[5],my_colors[4]))) +\n  geom_col() +\n  geom_smooth(method = \"loess\", se = FALSE) +\n  geom_smooth(method = \"lm\", se = FALSE, aes(color = my_colors[1])) +\n  theme_lyrics() + theme(plot.title = element_text(size = 11)) +\n  xlab(NULL) + ylab(NULL) +\n  ggtitle(\"Percent Positive Over Actual Sentiment of Train\")\n\nsuppressMessages(grid.arrange(polarity_over_Sentiment, relative_polarity_over_Sentiment, ncol = 2))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ead4c6349329c904708a44cd452d635b2f9e715f"},"cell_type":"markdown","source":"**A few extremes were adjusted in the second graph above, but the overall polarity trend over Actual(Tagged) Sentiment is postive in both cases.** [positive slope]"},{"metadata":{"_uuid":"e2051dd60f4f9171ce2118e709fdb85d8b06eb72"},"cell_type":"markdown","source":"# To be continued,Stay Tuned!! :)"}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}