{"cells":[{"metadata":{},"cell_type":"markdown","source":"Exploratory Text Analysis and feature addition using **udpipe** and **unsupervised ML**:-\n**UDPipe** — R package provides language-agnostic tokenization, tagging, lemmatization and dependency parsing of raw text, which is an essential part in natural language processing."},{"metadata":{"trusted":true},"cell_type":"code","source":"library(tidyverse) \nlibrary(udpipe)\nlibrary(lattice)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"\n#Loading the train data and test data .\nGQ_DATA_train<-read.csv(\"../input/google-quest-challenge/train.csv\",header=TRUE,sep=\",\")\nGQ_DATA_test<-read.csv(\"../input/google-quest-challenge/test.csv\",header=TRUE,sep=\",\")\n\n#head(GQ_DATA_1)\n\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Distribution of data based on `category`"},{"metadata":{"trusted":true},"cell_type":"code","source":"GQ_DATA_train %>% group_by(category) %>% count() %>% ggplot() + geom_line(aes(category,n, group = 1))\nGQ_DATA_test %>% group_by(category) %>% count() %>% ggplot() + geom_line(aes(category,n, group = 1))\n# The below graph shows the test and train are fairly balanced in terms of category. (with `technology` the major stake holder)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now loading the pretrained `udpipe` text models for `english` language."},{"metadata":{"trusted":true},"cell_type":"code","source":"x <- udpipe_download_model(language = \"english\")\nx$file_model\nud_english <- udpipe_load_model(x$file_model)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Analysis of all text features one by one using the `udpipe` models: (Plotting Part-of-speech tags from the given text)"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n######## TRAIN DATA\n\ns <- udpipe_annotate(ud_english, GQ_DATA_train$question_title)\nx <- data.frame(s)\n\n\n\n########## \nstats <- txt_freq(x$upos)\nstats$key <- factor(stats$key, levels = rev(stats$key))\nbarchart(key ~ freq, data = stats, col = \"yellow\", \n         main = \"UPOS TRAIN DATA(Universal Parts of Speech)\\n frequency of occurrence\", \n         xlab = \"Freq\")\n\n## NOUNS\nstats <- subset(x, upos %in% c(\"NOUN\")) \nstats <- txt_freq(stats$token)\nstats$key <- factor(stats$key, levels = rev(stats$key))\nbarchart(key ~ freq, data = head(stats, 20), col = \"cadetblue\", \n         main = \"Most occurring nouns TRAIN DATA\", xlab = \"Freq\")\n\n\n## ADJECTIVES\nstats <- subset(x, upos %in% c(\"ADJ\")) \nstats <- txt_freq(stats$token)\nstats$key <- factor(stats$key, levels = rev(stats$key))\nbarchart(key ~ freq, data = head(stats, 20), col = \"purple\", \n         main = \"Most occurring adjectives TRAIN DATA\", xlab = \"Freq\")\n         \n         \n## VERBS\nstats <- subset(x, upos %in% c(\"VERB\")) \nstats <- txt_freq(stats$token)\nstats$key <- factor(stats$key, levels = rev(stats$key))\nbarchart(key ~ freq, data = head(stats, 20), col = \"gold\", \n         main = \"Most occurring Verbs TRAIN DATA\", xlab = \"Freq\")\n         \n\n         \n         \n######## TEST DATA\n         \ns <- udpipe_annotate(ud_english, GQ_DATA_test$question_title)\nx <- data.frame(s)\n\n\n\n########## \nstats <- txt_freq(x$upos)\nstats$key <- factor(stats$key, levels = rev(stats$key))\nbarchart(key ~ freq, data = stats, col = \"yellow\", \n         main = \"UPOS TEST DATA(Universal Parts of Speech)\\n frequency of occurrence\", \n         xlab = \"Freq\")\n\n## NOUNS\nstats <- subset(x, upos %in% c(\"NOUN\")) \nstats <- txt_freq(stats$token)\nstats$key <- factor(stats$key, levels = rev(stats$key))\nbarchart(key ~ freq, data = head(stats, 20), col = \"cadetblue\", \n         main = \"Most occurring nouns TEST DATA\", xlab = \"Freq\")\n\n\n## ADJECTIVES\nstats <- subset(x, upos %in% c(\"ADJ\")) \nstats <- txt_freq(stats$token)\nstats$key <- factor(stats$key, levels = rev(stats$key))\nbarchart(key ~ freq, data = head(stats, 20), col = \"purple\", \n         main = \"Most occurring adjectives TEST DATA\", xlab = \"Freq\")\n         \n         \n## VERBS\nstats <- subset(x, upos %in% c(\"VERB\")) \nstats <- txt_freq(stats$token)\nstats$key <- factor(stats$key, levels = rev(stats$key))\nbarchart(key ~ freq, data = head(stats, 20), col = \"gold\", \n         main = \"Most occurring Verbs TEST DATA\", xlab = \"Freq\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we can see there is sufficient correlation in the `question title` between train and test data. (top 5 in `noun`, `adjective` and `verb` category are almost similar.) Similar analysis can be done for other text fields and can also be done at a `category` level."},{"metadata":{},"cell_type":"markdown","source":"Now using some unsupervised Machine Learning for further analysis. Below I have used `RAKE` for analysis.\n\nRAKE is one of the most popular (unsupervised) algorithms for extracting keywords in Information retrieval. RAKE short for Rapid Automatic Keyword Extraction algorithm, is a domain independent keyword extraction algorithm which tries to determine key phrases in a body of text by analyzing the frequency of word appearance and its co-occurrence with other words in the text."},{"metadata":{"trusted":true},"cell_type":"code","source":"# For train data.\ns <- udpipe_annotate(ud_english, GQ_DATA_train$question_title)\nx <- data.frame(s)\n\nstats <- keywords_rake(x = x, term = \"lemma\", group = \"doc_id\", \n                       relevant = x$upos %in% c(\"NOUN\", \"ADJ\"))\nstats$key <- factor(stats$keyword, levels = rev(stats$keyword))\nbarchart(key ~ rake, data = head(subset(stats, freq > 3), 20), col = \"red\", \n         main = \"Keywords identified by RAKE-TRAIN DATA\", \n         xlab = \"Rake\")\n\n\n# For test data.\ns <- udpipe_annotate(ud_english, GQ_DATA_test$question_title)\nx <- data.frame(s)\n\nstats <- keywords_rake(x = x, term = \"lemma\", group = \"doc_id\", \n                       relevant = x$upos %in% c(\"NOUN\", \"ADJ\"))\nstats$key <- factor(stats$keyword, levels = rev(stats$keyword))\nbarchart(key ~ rake, data = head(subset(stats, freq > 3), 20), col = \"red\", \n         main = \"Keywords identified by RAKE-TEST DATA\", \n         xlab = \"Rake\")\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"BY using **`RAKE`*** we see there is some difference in the key words identifies which is primarily due to small test data."},{"metadata":{},"cell_type":"markdown","source":"Analysing further more we can add more value to the above training data by adding above features, moreover we can also look for other logical word phrases like TOP NOUN — VERB Pairs as Keyword pairs.\n\nBelow is an example of the above."},{"metadata":{"trusted":true},"cell_type":"code","source":"## sequence of POS tags (noun phrases / verb phrases)\n\n\n## Train data\n\n\ns <- udpipe_annotate(ud_english, GQ_DATA_train$question_title)\nx <- data.frame(s)\n\n\nx$phrase_tag <- as_phrasemachine(x$upos, type = \"upos\")\nstats <- keywords_phrases(x = x$phrase_tag, term = tolower(x$token), \n                          pattern = \"(A|N)*N(P+D*(A|N)*N)*\", \n                          is_regex = TRUE, detailed = FALSE)\nstats <- subset(stats, ngram > 1 & freq > 3)\nstats$key <- factor(stats$keyword, levels = rev(stats$keyword))\nbarchart(key ~ freq, data = head(stats, 20), col = \"magenta\", \n         main = \"Keywords - simple noun phrases-TRAIN DATA\", xlab = \"Frequency\")\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Similar more logical phrases can be added which can add significant value when added as features during training an XgBoost or other tree based models."}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}