# Note

Code inlining is not working, but this is due to the usage of "system" in rmarkdown... Can't do much about that.

It seems Kaggle does not allow to tar files even if it could save bandwidth! (they delete your kernel without any prior warning - so don't do that!)

# Quick look at the data!

Source: https://www.kaggle.com/c/outbrain-click-prediction/data

# The dataset

The dataset for this challenge contains a sample of users’ page views and clicks, as observed on multiple publisher sites in the United States between 14-June-2016 and 28-June-2016. Each viewed page or clicked recommendation is further accompanied by some semantic attributes of those documents. For full details, see data specifications below.

The dataset contains numerous sets of content recommendations served to a specific user in a specific context. Each context (i.e. a set of recommendations) is given a display_id. In each such set, the user has clicked on at least one recommendation. The identities of the clicked recommendations in the test set are not revealed. Your task is to rank the recommendations in each group by decreasing predicted likelihood of being clicked.

As a warning, this is a very large relational dataset. While most of the tables are small enough to fit in memory, the page views log (page_views.csv) is over 2 billion rows and 100GB uncompressed. We have also uploaded a sample version of this file with the first 10,000,000 rows. The MD5 checksum of page_views.csv.zip is 3742c116bab4030e0a7ea1c0be623bd9.

# Data Fields

Each user in the dataset is represented by a unique id (uuid). A person can view a document (document_id), which is simply a web page with content (e.g.  a news article). On each document, a set of ads (ad_id) are displayed. Each ad belongs to a campaign (campaign_id) run by an advertiser (advertiser_id). You are also provided metadata about the document, such as which entities are mentioned, a taxonomy of categories, the topics mentioned, and the publisher.

# Privacy Reminder

Outbrain is releasing 2 Billion page views and 16,900,000 clicks of 700 Million unique users, across 560 sites. The data is anonymized. Please remember that participants are prohibited from de-anonymizing or reverse engineering data or combining the data with other publicly available information. Outbrain does not collect or hold PII (personally identifiable information), and the user identifiers we are releasing here are obscured. To protect its publisher partners, Outbrain is not releasing URLs of viewed or clicked stories, but rather anonymized document and site identifiers. The task at hand is click prediction, and by downloading the dataset, participants agree to use the data for that task alone, and will not attempt to reverse engineer the mapping from document, site, and user identifiers to URLs, site names or actual users.

***

# The data

Note: page_views.csv is NOT available in Kaggle kernels.

```{r}
library(data.table)
```

We are provided many files, namely:

* clicks_train.zip: 389.75 MB
* clicks_test.zip: 135.43 MB
* documents_meta.zip: 15.51 MB
* documents_categories.zip: 32.34 MB
* documents_entities.zip: 125.67 MB
* documents_topics.zip: 120.91 MB
* promoted_content.zip: 2.52 MB
* events.zip: 477.74 MB
* page_views.zip: 29.71 GB
* page_views_sample.zip: 148.51 MB
* sample_submission.zip: 99.57 MB

## How do they look uncompressed?

Using Unix shell:

```{r, echo=TRUE, eval=FALSE}
#cat(system("ls -sh ../input/*", intern = TRUE))
cat(system("ls -sh ../input/clicks_train.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/clicks_test.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_meta.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_categories.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_entities.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_topics.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/promoted_content.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/events.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/page_views_sample.csv", intern = TRUE), " (page_views.csv is not provided for Kaggle kernels)\n", sep = "")
cat(system("ls -sh ../input/sample_submission.csv", intern = TRUE), "\n", sep = "")
```

```{r, echo=FALSE, eval=TRUE}
cat(system("ls -sh ../input/clicks_train.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/clicks_test.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_meta.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_categories.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_entities.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_topics.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/promoted_content.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/events.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/page_views_sample.csv", intern = TRUE), " (page_views.csv is not provided for Kaggle kernels)\n", sep = "")
cat(system("ls -sh ../input/sample_submission.csv", intern = TRUE), "\n", sep = "")
```

## How many lines in each file?

```{r, echo=TRUE}
cat(system("cat ../input/clicks_train.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/clicks_test.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/documents_meta.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/documents_categories.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/documents_entities.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/documents_topics.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/promoted_content.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/events.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/page_views_sample.csv | wc -l", intern = TRUE), "\n", sep = "")
cat(system("cat ../input/sample_submission.csv | wc -l", intern = TRUE), "\n", sep = "")
```

## MD5 hashes:

```{r, echo=TRUE}
cat(system("md5sum ../input/clicks_train.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/clicks_test.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/documents_meta.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/documents_categories.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/documents_entities.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/documents_topics.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/promoted_content.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/events.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/page_views_sample.csv", intern = TRUE), "\n", sep = "")
cat(system("md5sum ../input/sample_submission.csv", intern = TRUE), "\n", sep = "")
```

## TARing files (with file size output!)

It seems Kaggle does not allow to tar files even if it could save bandwidth! (they delete your kernel without any prior warning - so don't do that!)

```{r, echo=TRUE}

#system("tar -jcvf clicks_train.tar.bz2 ../input/clicks_train.csv", intern = TRUE)
#cat(system("ls -sh clicks_train.tar.bz2", intern = TRUE), " --- vs 389.75 MB\n", sep = "")

#system("tar -jcvf clicks_test.tar.bz2 ../input/clicks_test.csv", intern = TRUE)
#cat(system("ls -sh clicks_test.tar.bz2", intern = TRUE), " --- vs 135.43 MB\n", sep = "")

#system("tar -jcvf documents_meta.tar.bz2 ../input/documents_meta.csv", intern = TRUE)
#cat(system("ls -sh documents_meta.tar.bz2", intern = TRUE), " --- vs 15.51 MB\n", sep = "")

#system("tar -jcvf documents_categories.tar.bz2 ../input/documents_categories.csv", intern = TRUE)
#cat(system("ls -sh documents_categories.tar.bz2", intern = TRUE), " --- vs 32.34 MB\n", sep = "")

#system("tar -jcvf documents_entities.tar.bz2 ../input/documents_entities.csv", intern = TRUE)
#cat(system("ls -sh documents_entities.tar.bz2", intern = TRUE), " --- vs 125.67 MB\n", sep = "")

#system("tar -jcvf documents_topics.tar.bz2 ../input/documents_topics.csv", intern = TRUE)
#cat(system("ls -sh documents_topics.tar.bz2", intern = TRUE), " --- vs 120.91 MB\n", sep = "")

#system("tar -jcvf promoted_content.tar.bz2 ../input/promoted_content.csv", intern = TRUE)
#cat(system("ls -sh promoted_content.tar.bz2", intern = TRUE), " --- vs 2.52 MB\n", sep = "")

#system("tar -jcvf events.tar.bz2 ../input/events.csv", intern = TRUE)
#cat(system("ls -sh events.tar.bz2", intern = TRUE), " --- vs 477.74 MB\n", sep = "")

#system("tar -jcvf page_views_sample.tar.bz2 ../input/page_views_sample.csv", intern = TRUE)
#cat(system("ls -sh page_views_sample.tar.bz2", intern = TRUE), " --- vs 148.51 MB\n", sep = "")

#system("tar -jcvf sample_submission.tar.bz2 ../input/sample_submission.csv", intern = TRUE)
#cat(system("ls -sh sample_submission.tar.bz2", intern = TRUE), " --- vs 99.57 MB\n", sep = "")
```

***

# The files!

We will print 25 rows of each file!

```{r, echo=TRUE}
N <- 25
```

***

## Setting up "pretty print"

```{r, echo=TRUE}
pprint <- function(data) {
    cat(pprint_helper(data), sep = "\n")
}

pprint_helper <- function(data) {
    out <- paste(names(data), collapse = " | ")
    out <- c(out, paste(rep("---", ncol(data)), collapse = " | "))
    invisible(apply(data, 1, function(x) {
        out <<- c(out, paste(x, collapse = " | "))
    }))
    return(out)
}
```

***

## Label proportions

Let's look at the label proportions.

```{r, echo=TRUE, results='asis'}
Y <- fread("../input/clicks_train.csv", header = TRUE, select = 3, showProgress = FALSE)
print(table(Y))
```

What about the ratios?

```{r, echo=TRUE, results='asis'}
cat("Absolute proportions of positives (vs 87141732): ", 16874593/87141732, "\n --- Proportions of positives vs negatives: ", table(Y)[2]/table(Y)[1], "\n --- Inverted proportions (negatives vs positives): ", table(Y)[1]/table(Y)[2], sep = "")
```

P.S: can't believe I had to hardcode the ratio in Rmarkdown... (really no reason, but it was not printing properly...)

***

## clicks_train.csv

clicks_train.csv is the training set, showing which of a set of ads was clicked.

* display_id
* ad_id
* clicked (1 if clicked, 0 otherwise)

```{r, echo=TRUE, results='asis'}
data <- fread("../input/clicks_train.csv", header = TRUE, nrows = N)
pprint(data)
Y <- fread("../input/clicks_train.csv", header = TRUE, select = 3, showProgress = FALSE)
print(table(Y))
```

***

## clicks_test.csv

clicks_test.csv is the same as clicks_train.csv, except it does not have the clicked ad. This is the file you should use to predict. Each display_id has only one clicked ad. Note that test set contains display_ids from the entire dataset timeframe. Additionally, the public/private sampling for the competition is uniformly random, not based on time. These sampling choices were intentional, in spite of the possibility that participants can look ahead in time.

```{r, echo=TRUE, results='asis'}
data <- fread("../input/clicks_test.csv", header = TRUE, nrows = N)
pprint(data)
```

***

## documents_meta.csv

documents_meta.csv provides details on the documents.

* document_id
* source_id (the part of the site on which the document is displayed, e.g. edition.cnn.com)
* publisher_id
* publish_time

```{r, echo=TRUE, results='asis'}
data <- fread("../input/documents_meta.csv", header = TRUE, nrows = N)
pprint(data)
```

***

## document_categories.csv

**Warning: there is some Machine Learning done as there seems to be a confidence level with unnatural numbers!!! Therefore, we are applying Machine Learning over Machine Learning! =)**

documents_topics.csv, documents_entities.csv, and documents_categories.csv all provide information about the content in a document, as well as Outbrain's confidence in each respective relationship. For example, an entity_id can represent a person, organization, or location. The rows in documents_entities.csv give the confidence that the given entity was referred to in the document.

```{r, echo=TRUE, results='asis'}
data <- fread("../input/documents_categories.csv", header = TRUE, nrows = N)
pprint(data)
```

***

## documents_entities.csv

**Warning: there is some Machine Learning done as there seems to be a confidence level with unnatural numbers!!! Therefore, we are applying Machine Learning over Machine Learning! =)**

documents_topics.csv, documents_entities.csv, and documents_categories.csv all provide information about the content in a document, as well as Outbrain's confidence in each respective relationship. For example, an entity_id can represent a person, organization, or location. The rows in documents_entities.csv give the confidence that the given entity was referred to in the document.

```{r, echo=TRUE, results='asis'}
data <- fread("../input/documents_entities.csv", header = TRUE, nrows = N)
pprint(data)
```

***

## documents_topics.csv

**Warning: there is some Machine Learning done as there seems to be a confidence level with unnatural numbers!!! Therefore, we are applying Machine Learning over Machine Learning! =)**

documents_topics.csv, documents_entities.csv, and documents_categories.csv all provide information about the content in a document, as well as Outbrain's confidence in each respective relationship. For example, an entity_id can represent a person, organization, or location. The rows in documents_entities.csv give the confidence that the given entity was referred to in the document.

```{r, echo=TRUE, results='asis'}
data <- fread("../input/documents_topics.csv", header = TRUE, nrows = N)
pprint(data)
```

***

## promoted_content.csv

promoted_content.csv provides details on the ads.

* ad_id
* document_id
* campaign_id
* advertiser_id

```{r, echo=TRUE, results='asis'}
data <- fread("../input/promoted_content.csv", header = TRUE, nrows = N)
pprint(data)
```

***

## events.csv

events.csv provides information on the display_id context. It covers both the train and test set.

* display_id
* uuid
* document_id
* timestamp
* platform
* geo_location

```{r, echo=TRUE, results='asis'}
data <- fread("../input/events.csv", header = TRUE, nrows = N)
pprint(data)
```

***

## page_views.csv (warning: original not available, only the sample)

page_views.csv is a the log of users visiting documents. To save disk space, the timestamps in the entire dataset are relative to the first time in the dataset. If you wish to recover the actual epoch time of the visit, add 1465876799998 to the timestamp.

* uuid
* document_id
* timestamp (ms since 1970-01-01 - 1465876799998)
* platform (desktop = 1, mobile = 2, tablet =3)
* geo_location (country>state>DMA)
* traffic_source (internal = 1, search = 2, social = 3)

```{r, echo=TRUE, results='asis'}
data <- fread("../input/page_views_sample.csv", header = TRUE, nrows = N)
pprint(data)
```

***

## sample_submission.csv

sample_submission.csv shows the correct submission format.

```{r, echo=TRUE, results='asis'}
data <- fread("../input/sample_submission.csv", header = TRUE, nrows = N)
pprint(data)
```

