{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "e9c55ae0-847d-a717-ebea-a2f2940ce219"
      },
      "source": [
        "##What does Outbrain do?\n",
        "\n",
        "When you are serving the web (say you are on CNN website), after you finish reading your news there will be rows of links at the bottom of the article, showing relevant stories. Those recommendations are delivered by Ourbrain. Outbrain promoted these stories from advertiser on publisher websites."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "7ed44335-6796-88f0-c2b8-1d7d20e5a5c4"
      },
      "source": [
        "## Looking at data files one by one.\n",
        "\n",
        "### Information about each document\n",
        "\n",
        "* promoted_content.csv\t.zip (2.52 mb)\n",
        "* documents_meta.csv\t.zip (15.51 mb)\n",
        "* documents_categories.csv\t.zip (32.34 mb)\n",
        "* documents_entities.csv\t.zip (125.67 mb)\n",
        "* documents_topics.csv\t.zip (120.91 mb)\n",
        "\n",
        "### Information about clicks\n",
        "* clicks_train.csv\t.zip (389.75 mb)\n",
        "* events.csv\t.zip (477.74 mb)\n",
        "* page_views_sample.csv\n",
        "\n",
        "\n",
        "\n",
        "*Redundant information*\n",
        "\n",
        "* sample_submission.csv\t.zip (99.57 mb)\n",
        "* clicks_test.csv\t.zip (135.43 mb)\n",
        "* page_views.csv\t.zip (29.71 gb)\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "da5efee3-943d-a7bb-5221-345992246eb0"
      },
      "outputs": [],
      "source": [
        "import pandas as pd\n",
        "import numpy as np"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "6b31ec67-5427-e78c-f7d4-14309268dae2"
      },
      "outputs": [],
      "source": [
        "doc_ad_df = pd.read_csv('../input/promoted_content.csv', nrows=5)\n",
        "doc_ad_df.head()\n",
        "\n",
        "# Here we have the information about each promoted content (ad). Its id, document, campaign id, and advertiser id."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "40d93403-f596-d24a-9a0a-d1498fab16e5"
      },
      "outputs": [],
      "source": [
        "doc_meta_df = pd.read_csv('../input/documents_meta.csv', nrows=5)\n",
        "doc_meta_df.head()\n",
        "\n",
        "# Here we have the document id, source id, publisher id, and the time the promoted contents are published."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "0ef7249b-1eb1-751d-97cf-c7d96e875b0b"
      },
      "outputs": [],
      "source": [
        "doc_cat_df = pd.read_csv('../input/documents_categories.csv', nrows=5)\n",
        "doc_cat_df.head()\n",
        "\n",
        "# Here's the categories of the documents (aka. promoted contents). And Outbrain's confidence for category assignments. "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "856e20a0-16d5-4da0-d2db-ed8a3e67f969"
      },
      "outputs": [],
      "source": [
        "doc_en_df = pd.read_csv('../input/documents_entities.csv', nrows=5)\n",
        "doc_en_df.head()\n",
        "\n",
        "# Kaggle said \"an entity_id can represent a person, organization, or location.\" So this is person, organization, location this document is referring to along with Kaggle's NER confidence."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "d0e671a6-3369-f35c-95dc-5a2e5f8cb08d"
      },
      "outputs": [],
      "source": [
        "doc_top_df = pd.read_csv('../input/documents_topics.csv', nrows=5)\n",
        "doc_top_df.head()\n",
        "\n",
        "# Okay, this is the topic of the article."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c65ef686-7a48-72a9-c1de-44cd21568412"
      },
      "outputs": [],
      "source": [
        "cli_df = pd.read_csv('../input/clicks_train.csv', nrows=5)\n",
        "cli_df.head()\n",
        "\n",
        "# They collect click label based on ad_id. This is what we are predicting.\n",
        "# Display id is the set of recommendations during the click"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e3e7075d-5df4-22f3-fdd1-3d3a3c148457"
      },
      "outputs": [],
      "source": [
        "events_df = pd.read_csv('../input/events.csv', nrows=5)\n",
        "events_df.head()\n",
        "\n",
        "# each event here is a click event? display_id is the set of recommendations shown during the click. \n",
        "# this table shows what promoted content was clicked, the time, platform, geolocation."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "89962bc9-3499-0cb8-e63d-0bde50e7738d"
      },
      "outputs": [],
      "source": [
        "pv_df = pd.read_csv('../input/page_views_sample.csv', nrows=5)\n",
        "pv_df.head()\n",
        "\n",
        "# This is all the page views that Outbrain is able to track, I guess? uuid is practically unique id of the event. The time, platform, location, and traffic source of the page view."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "41f5749c-0c23-c5e4-8cd6-b176d2d80e2d"
      },
      "source": [
        "## Now f"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.5.2"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}