---
title: "Quora EDA"
date: '`r Sys.Date()`'
output:
  html_document:
    number_sections: true
    fig_caption: true
    toc: true
    fig_width: 7
    fig_height: 4.5
    theme: cosmo
    highlight: tango
    code_folding: hide
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE)
```

The following kernel will be worked on over the next few days to build the foundation of submission models. Please feel free to critique it as I go along to help me build my skillset.

Quora is a knowledge sharing platform that brings those seeking answers to questions and those with knowledge of the topic together. This competition is aimed at identifying "insincere" questions (*Quora is a Q&A platform that empowers people to share and grow the world’s knowledge. People come to Quora to ask questions about any subject, read high quality knowledge that's personalized and relevant to them, and share their own knowledge with others. Quora is a place to share knowledge and better understand the world.*).


# Versions

1. Simple EDA of target variable against question metadata
2. Wordcloud text analysis

```{r libraries, warning=FALSE, message=FALSE, include=FALSE}
library(tidyverse)
library(stringr)
library(data.table)
library(scales)
library(quanteda)
library(RColorBrewer)
```

```{r data, warning=FALSE, message=FALSE, include=FALSE}
train <- fread("../input/train.csv", stringsAsFactors = F)
```

```{r format}
eda_colours <- c("orange", "steelblue")

```

# Inspect the data

```{r inspect}
glimpse(train)
summary(train)
```

# What do "Insincere" questions look like
```{r}
train %>% filter(target == 1) %>% select (question_text) %>% head(n=10)
```

# Target Feature

In the training set provided, only 6.2% of observations are flagged as insincere (target = 1).

```{r target}
train %>%
  group_by(target) %>%
  summarise(n = n()) %>%
  mutate(percentage = n / sum(n)) %>%
  ggplot(aes(x=factor(target), y= n)) +
  geom_bar(stat = "identity", fill = eda_colours) +
  geom_text(aes(label = percent(percentage)), vjust = -0.5) +
  theme_minimal() +
  labs(x= "Target", y= "Count")
```

# Feature Engineering

Some preliminary feature engineering has been undertaken to see if there are patterns that can halp identify insincere questions.

These will be built upon over subsequent versions.
```{r features}
train <- train %>%
  mutate(word_count = str_count(question_text, " ") +1,
         question_length = nchar(question_text),
         punctuation_count = str_count(question_text, "[[:punct:]]"),
         average_word_length = question_length / word_count,
         question_mark_count = str_count(question_text, "\\?"),
         numbers_in_question = str_detect(question_text, "[[:digit:]]"),
         number_count = str_count(question_text, "[[:digit:]]"))
```


# Features vs Target Variable Visualisations

The density plot below shows that insincere questions tend to have slightly higher word counts
```{r p1}
ggplot(data = train, aes(x= word_count, fill = factor(target))) +
  geom_density(alpha = 0.5, adjust = 2) +
  scale_fill_manual(values = eda_colours, name = "Target") +
  labs(x="Word Count", y="", title = "Do insincere questions have more or less words in them?") +
  theme_minimal()
```

This is confirmed by the summary statistics below.

```{r s1}
by(train$word_count, train$target, summary)
```

Similarly, insincere questions tend to have more characters as well. This is to be expected as the word count is also higher in insincere questions.

```{r p2}
ggplot(data = train, aes(x= question_length, fill = factor(target))) +
  geom_density(alpha = 0.5, adjust = 2) +
  scale_fill_manual(values = eda_colours, name = "Target") +
  labs(x="Question Length", y="", title = "Do insincere questions have more or less characters in them?") +
  theme_minimal()
```

There doesn't appear to be a difference in the average word length between insincere and sincere questions. This variable won't be used in modelling.

```{r p3}
ggplot(data = train, aes(x= average_word_length, fill = factor(target))) +
  geom_density(alpha = 0.5, adjust = 2) +
  scale_fill_manual(values = eda_colours, name = "Target") +
  labs(x="Average Word Length", y="", title = "Do insincere questions contain longer words?") +
  theme_minimal()
```

Punctuation counts between insincere and sincere questions thend to be slightly higher between sincere and insincere questions. Interestingly, there is a question with 339 puncuation symbols used?!

```{r s2}
by(train$punctuation_count, train$target, summary)
```

Ah, they appear to be formatted equations that are being read in as raw text. This will need to be examined in more detail later.

```{r long_punc}
train[train$punctuation_count > 100,]
```

I was curious to see if nested questions within the question were used in insincere questions. The summary statistics appears to indicate there wasn't.

```{r s3}
by(train$question_mark_count, train$target, summary)
```

There appears to be a very mionor relationship between the presence of numeric characters and the comment sincerity - if there is a numerical character in the question posed, a greater proportion of those questions are classed sincere than if there is no numeric character present. 

```{r p4}
train %>%
  group_by(numbers_in_question, target) %>%
  summarise(n = n()) %>%
  mutate(percent_numbers = n / sum(n)) %>%
  ggplot(aes(x= numbers_in_question, y= percent_numbers, fill = factor(target))) +
  geom_bar(stat = "identity") +
  scale_fill_manual(values = eda_colours, name = "Target") +
  scale_y_continuous(labels = percent) +
  labs(x= "Numbers in question?", y= "Percent Insincere", title = "Is the presence of numbers in a question an indicator of sincerity?") +
  theme_classic()
```

# Text Analysis

```{r text_setup}
question_corpus <- corpus(train$question_text)

docvars(question_corpus) <- train$target


word.colour <- brewer.pal(10, "RdBu") 
```

There is a clear difference in the most frequent words used in sincere vs insincere Quora questions. Race-based terms, vulgarities and of course Trump appear more frequently in the questions tagged Insincere.

Importantly, the most popular words in *insincere* questions **do not** appear in *sincere* questions.

These will be examined in further before building a predictive model to classify the questions.

## Insincere Wordcloud

```{r insincerewordcloud}
#subsetting only the insincere messages
insincere.plot<-corpus_subset(question_corpus,docvar1==1)

#now creating a document-feature matrix using dfm()
insincere.plot<-dfm(insincere.plot, tolower = TRUE, remove_punct = TRUE, remove_twitter = TRUE, remove_numbers = TRUE, remove=stopwords(source = "smart"))

textplot_wordcloud(insincere.plot, min_count = 300, color = word.colour)  
#title("Insincere Words", col.main = "steelblue", cex.main = 2)

```

## Sincere Wordcloud

```{r sincerewordcloud}
#subsetting only the insincere messages
sincere.plot<-corpus_subset(question_corpus,docvar1==0)

#now creating a document-feature matrix using dfm()
sincere.plot<-dfm(sincere.plot, tolower = TRUE, remove_punct = TRUE, remove_twitter = TRUE, remove_numbers = TRUE, remove=stopwords(source = "smart"))

textplot_wordcloud(sincere.plot, min_count = 500, color = word.colour)  
#title("Sincere Words", col.main = "orange", cex.main = 2)
```
