# Quick look at the data!

Source: https://www.kaggle.com/c/outbrain-click-prediction/data

# The dataset

The dataset for this challenge contains a sample of users’ page views and clicks, as observed on multiple publisher sites in the United States between 14-June-2016 and 28-June-2016. Each viewed page or clicked recommendation is further accompanied by some semantic attributes of those documents. For full details, see data specifications below.

The dataset contains numerous sets of content recommendations served to a specific user in a specific context. Each context (i.e. a set of recommendations) is given a display_id. In each such set, the user has clicked on at least one recommendation. The identities of the clicked recommendations in the test set are not revealed. Your task is to rank the recommendations in each group by decreasing predicted likelihood of being clicked.

As a warning, this is a very large relational dataset. While most of the tables are small enough to fit in memory, the page views log (page_views.csv) is over 2 billion rows and 100GB uncompressed. We have also uploaded a sample version of this file with the first 10,000,000 rows. The MD5 checksum of page_views.csv.zip is 3742c116bab4030e0a7ea1c0be623bd9.

# Data Fields

Each user in the dataset is represented by a unique id (uuid). A person can view a document (document_id), which is simply a web page with content (e.g.  a news article). On each document, a set of ads (ad_id) are displayed. Each ad belongs to a campaign (campaign_id) run by an advertiser (advertiser_id). You are also provided metadata about the document, such as which entities are mentioned, a taxonomy of categories, the topics mentioned, and the publisher.

# Privacy Reminder

Outbrain is releasing 2 Billion page views and 16,900,000 clicks of 700 Million unique users, across 560 sites. The data is anonymized. Please remember that participants are prohibited from de-anonymizing or reverse engineering data or combining the data with other publicly available information. Outbrain does not collect or hold PII (personally identifiable information), and the user identifiers we are releasing here are obscured. To protect its publisher partners, Outbrain is not releasing URLs of viewed or clicked stories, but rather anonymized document and site identifiers. The task at hand is click prediction, and by downloading the dataset, participants agree to use the data for that task alone, and will not attempt to reverse engineer the mapping from document, site, and user identifiers to URLs, site names or actual users.

# The data

Note: page_views.csv is NOT available in Kaggle kernels.

```{r}
library(data.table)
```

We are provided many files, namely:

* clicks_train.zip: 389.75 MB
* clicks_test.zip: 135.43 MB
* documents_meta.zip: 15.51 MB
* documents_categories.zip: 32.34 MB
* documents_entities.zip: 125.67 MB
* documents_topics.zip: 120.91 MB
* promoted_content.zip: 2.52 MB
* events.zip: 477.74 MB
* page_views.zip: 29.71 GB
* page_views_sample.zip: 148.51 MB
* sample_submission.zip: 99.57 MB

How do they look uncompressed?

Using Unix shell:

```{r, echo=TRUE, eval=FALSE}
#cat(system("ls -sh ../input/*", intern = TRUE))
cat(system("ls -sh ../input/clicks_train.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/clicks_test.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_meta.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_categories.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_entities.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_topics.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/promoted_content.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/events.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/page_views_sample.csv", intern = TRUE), " (page_views.csv is not provided for Kaggle kernels)\n", sep = "")
cat(system("ls -sh ../input/sample_submission.csv", intern = TRUE), "\n", sep = "")
```

```{r, echo=FALSE, eval=TRUE}
cat(system("ls -sh ../input/clicks_train.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/clicks_test.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_meta.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_categories.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_entities.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/documents_topics.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/promoted_content.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/events.csv", intern = TRUE), "\n", sep = "")
cat(system("ls -sh ../input/page_views_sample.csv", intern = TRUE), " (page_views.csv is not provided for Kaggle kernels)\n", sep = "")
cat(system("ls -sh ../input/sample_submission.csv", intern = TRUE), "\n", sep = "")
```

# TARing files:

```{r, echo=TRUE}
system("tar -jcvf documents_categories.tar.bz2 ../input/documents_categories.csv", intern = TRUE)
cat(system("ls -sh documents_categories.tar.bz2", intern = TRUE), "\n", sep = "")
```

# The files!

## clicks_train.csv

clicks_train.csv is the training set, showing which of a set of ads was clicked.

* display_id
* ad_id
* clicked (1 if clicked, 0 otherwise)

```{r, echo=TRUE}
data <- fread("../input/clicks_train.csv", header = TRUE, nrows = 15)
print(data)
```

## clicks_test.csv

clicks_test.csv is the same as clicks_train.csv, except it does not have the clicked ad. This is the file you should use to predict. Each display_id has only one clicked ad. Note that test set contains display_ids from the entire dataset timeframe. Additionally, the public/private sampling for the competition is uniformly random, not based on time. These sampling choices were intentional, in spite of the possibility that participants can look ahead in time.

```{r, echo=TRUE}
data <- fread("../input/clicks_test.csv", header = TRUE, nrows = 15)
print(data)
```

## documents_meta.csv

documents_meta.csv provides details on the documents.

* document_id
* source_id (the part of the site on which the document is displayed, e.g. edition.cnn.com)
* publisher_id
* publish_time

```{r, echo=TRUE}
data <- fread("../input/documents_meta.csv", header = TRUE, nrows = 15)
print(data)
```

## document_categories.csv

documents_topics.csv, documents_entities.csv, and documents_categories.csv all provide information about the content in a document, as well as Outbrain's confidence in each respective relationship. For example, an entity_id can represent a person, organization, or location. The rows in documents_entities.csv give the confidence that the given entity was referred to in the document.

```{r, echo=TRUE}
data <- fread("../input/documents_categories.csv", header = TRUE, nrows = 15)
print(data)
```

## documents_entities.csv

documents_topics.csv, documents_entities.csv, and documents_categories.csv all provide information about the content in a document, as well as Outbrain's confidence in each respective relationship. For example, an entity_id can represent a person, organization, or location. The rows in documents_entities.csv give the confidence that the given entity was referred to in the document.

```{r, echo=TRUE}
data <- fread("../input/documents_entities.csv", header = TRUE, nrows = 15)
print(data)
```

## documents_topics.csv

documents_topics.csv, documents_entities.csv, and documents_categories.csv all provide information about the content in a document, as well as Outbrain's confidence in each respective relationship. For example, an entity_id can represent a person, organization, or location. The rows in documents_entities.csv give the confidence that the given entity was referred to in the document.

```{r, echo=TRUE}
data <- fread("../input/documents_topics.csv", header = TRUE, nrows = 15)
print(data)
```

## promoted_content.csv

promoted_content.csv provides details on the ads.

* ad_id
* document_id
* campaign_id
* advertiser_id

```{r, echo=TRUE}
data <- fread("../input/promoted_content.csv", header = TRUE, nrows = 15)
print(data)
```

## events.csv

events.csv provides information on the display_id context. It covers both the train and test set.

* display_id
* uuid
* document_id
* timestamp
* platform
* geo_location

```{r, echo=TRUE}
data <- fread("../input/events.csv", header = TRUE, nrows = 15)
print(data)
```

## page_views.csv (warning: opening the sample as it is too large to be loaded)

page_views.csv is a the log of users visiting documents. To save disk space, the timestamps in the entire dataset are relative to the first time in the dataset. If you wish to recover the actual epoch time of the visit, add 1465876799998 to the timestamp.

* uuid
* document_id
* timestamp (ms since 1970-01-01 - 1465876799998)
* platform (desktop = 1, mobile = 2, tablet =3)
* geo_location (country>state>DMA)
* traffic_source (internal = 1, search = 2, social = 3)

```{r, echo=TRUE}
data <- fread("../input/page_views_sample.csv", header = TRUE, nrows = 15)
print(data)
```

## sample_submission.csv

sample_submission.csv shows the correct submission format.

```{r, echo=TRUE}
data <- fread("../input/sample_submission.csv", header = TRUE, nrows = 15)
print(data)
```

