---
title: "Exploring Toxic Comments"
date: "23/03/2020"
output:
  html_document:
    code_folding: hide
    theme: journal
    highlight: tango
    df_print: paged
    fig_height: 8
    fig_width: 11
    number_sections: true
    toc: yes
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE)
```

```{r, warning=FALSE, message=FALSE}
# set up plotting theme
theme_jason <- function(legend_pos="top", base_size=12, font=NA){
  
  # come up with some default text details
  txt <- element_text(size = base_size+3, colour = "black", face = "plain")
  bold_txt <- element_text(size = base_size+3, colour = "black", face = "bold")
  
  # use the theme_minimal() theme as a baseline
  theme_minimal(base_size = base_size, base_family = font)+
    theme(text = txt,
          # axis title and text
          axis.title.x = element_text(size = 15, hjust = 1),
          axis.title.y = element_text(size = 15),
          # gridlines on plot
          panel.grid.major = element_line(linetype = 2),
          panel.grid.minor = element_line(linetype = 2),
          # title and subtitle text
          plot.title = element_text(size = 18, colour = "grey25", face = "bold"),
          plot.subtitle = element_text(size = 16, colour = "grey44"),
          ###### clean up!
          legend.key = element_blank(),
          # the strip.* arguments are for faceted plots
          strip.background = element_blank(),
          strip.text = element_text(face = "bold", size = 13, colour = "grey35")) +
    #----- AXIS -----#
    theme(
      #### remove Tick marks
      axis.ticks=element_blank(),
      ### legend depends on argument in function and no title
      legend.position = legend_pos,
      legend.title = element_blank(),
      legend.background = element_rect(fill = NULL, size = 0.5,linetype = 2)
    )
}
plot_cols <- c("#498972", "#3E8193", "#BC6E2E", "#A09D3C", "#E06E77", "#7589BC", "#A57BAF", "#4D4D4D")
```


# Introduction

This competition asks competitors to try to identify toxicity in online conversations, where toxicity is defined as anything rude, disrespectful or otherwise likely to make someone leave a discussion. If these toxic contributions can be identified, we could have a safer, more collaborative internet.

This is the third competition of its kind. The first competition in 2018 [Toxic Comment Classification Challenge](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge), Kagglers built multi-headed models to recognize toxicity and several subtypes of toxicity. In 2019, in the [Unintended Bias in Toxicity Classification Challenge](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification), you worked to build toxicity models that operate fairly across a diverse range of conversations. This year, we're able to take advantage of Kaggle's new TPU support and have been challenged with building multilingual models with English-only training data.

## What am I predicting?
Competitors are to predict the probability that a comment is toxic. A toxic comment would receive a `1.0`. A benign, non-toxic comment would receive a `0.0`. In the test set, all comments are classified as either a `1.0` or a `0.0`.

## Evaluation Metric
Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.


This analysis is only intended to be an EDA on the toxic comments themselves. A modelling perspective notebook may come somewhere later down the track.

```{r, warning=FALSE, message=FALSE}
# load libraries
library(tidyverse)
library(data.table)
library(scales)
library(patchwork)
```


# The Data

```{r, warning=FALSE, message=FALSE}
train18 <- fread("../input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train.csv")
train19 <- fread("../input/jigsaw-multilingual-toxic-comment-classification/jigsaw-unintended-bias-train.csv")

test <- fread("../input/jigsaw-multilingual-toxic-comment-classification/test.csv")
validation <- fread("../input/jigsaw-multilingual-toxic-comment-classification/validation.csv")

sample_sub <- fread("../input/jigsaw-multilingual-toxic-comment-classification/sample_submission.csv")
```


## Training Data

The number of observations in our training sets combined is `r scales::comma(nrow(train18) + nrow(train19))`, (`r scales::comma(nrow(train18))` rows in the 2018 data and `r scales::comma(nrow(train19))` in 2019).


From the below tables, we can see that there is a discrepancy between the variables between the two training sets.

The 2018 data has `r ncol(train18)`, while the 2019 data has `r ncol(train19)`.

The 2018 set has the folowing header names:

`r names(train18)`

While the 2019 data has:

`r names(train19)`

```{r, warning=FALSE, message=FALSE}
DT::datatable(head(train18))
DT::datatable(head(train19))
```


## Test and Validation sets

```{r, warning=FALSE, message=FALSE}
DT::datatable(head(test))
DT::datatable(head(validation))
```

The test and validation sets only have three and four columns respectively.


# Target Variable Analysis

The target variable to predict is `toxic`.

When plotting the distribution of it in the training datasets, it can be seen that in 2018, the variable was a binary one (1 or 0), however it's a continuous variable between 0 and 1 in 2019.

The validation set is binary.

```{r, warning=FALSE, message=FALSE}
tr18 <- train18 %>% 
  ggplot(aes(toxic)) +
  geom_histogram(fill="wheat3", colour="grey40") +
  scale_y_continuous(labels = comma) +
  ggtitle("Train 2018", subtitle = "Binary") +
  theme_jason()

tr19 <- train19 %>% 
  ggplot(aes(toxic)) +
  geom_histogram(fill="wheat3", colour="grey40") +
  scale_y_continuous(labels = comma) +
  ggtitle("Train 2019", subtitle = "Continuous") +
  theme_jason()


val <- validation %>% 
  ggplot(aes(toxic)) +
  geom_histogram(fill="wheat3", colour="grey40") +
  scale_y_continuous(labels = comma) +
  ggtitle("Validation", subtitle = "Binary") +
  theme_jason()

(tr18 + tr19 + val)
```

# Target Languages

While our training sets are entirely English, The competition is about using English text to predict the toxicity of text in another language.

Languages are displayed in their abbreviated versions. The abbreviations are expanded below:

|Abbreviation|Language|
|------------|--------|
|es|Spanish|
|fr|French|
|it|Italian|
|pt|Portuguese|
|ru|Russian|
|tr|Turkish|


The count of each language in the test and validation sets are displayed below. Some observations:

* There are six languages in the test set, wiht Turkish being the most frequent language
* Only three languages in the validation set; Turkish, Italian and Spanish
* Portuguese, Russian and French comments are the next highest aftre Turkish in the test set, yet don't appear in the validation set

```{r, warning=FALSE, message=FALSE}
test_lan <- test %>% 
  count(lang) %>% 
  ggplot(aes(x=reorder(lang,n), y=n)) +
  geom_col(fill="wheat3", colour="grey40") +
  geom_text(aes(label = comma(n)), hjust=1) +
  ggtitle("Test data set") +
  coord_flip() +
  theme_jason() +
  theme(axis.title.x = element_blank(), axis.text.x = element_blank(), axis.title.y = element_blank(),
        panel.grid.major.x = element_blank(), panel.grid.minor.x = element_blank(), panel.grid.major.y = element_blank(), panel.grid.minor.y = element_blank())

val_lan <- validation %>% 
  count(lang) %>% 
  ggplot(aes(x=reorder(lang,n), y=n)) +
  geom_col(fill="wheat3", colour="grey40") +
  geom_text(aes(label = comma(n)), hjust=1) +
  ggtitle("Validation data set") +
  coord_flip() +
  theme_jason() +
  theme(axis.title.x = element_blank(), axis.text.x = element_blank(), axis.title.y = element_blank(),
        panel.grid.major.x = element_blank(), panel.grid.minor.x = element_blank(), panel.grid.major.y = element_blank(), panel.grid.minor.y = element_blank())

(test_lan + val_lan)
```

**To be continued**