{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4506f6d5-fc74-da3a-6c85-1bf13917e675"
      },
      "source": [
        "Let's see what I can get working with data. How good parameters I can retrieve from the data to train a good model."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b56246b5-2d55-b35a-589a-6d8a1990e7df"
      },
      "outputs": [],
      "source": [
        "# This Python 3 environment comes with many helpful analytics libraries installed\n",
        "# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n",
        "# For example, here's several helpful packages to load in \n",
        "\n",
        "import numpy as np # linear algebra\n",
        "import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n",
        "import matplotlib.pyplot as plt\n",
        "import warnings\n",
        "warnings.filterwarnings('ignore')\n",
        "\n",
        "# Input data files are available in the \"../input/\" directory.\n",
        "# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n",
        "\n",
        "from subprocess import check_output\n",
        "print(check_output([\"ls\", \"../input\"]).decode(\"utf8\"))\n",
        "\n",
        "# Any results you write to the current directory are saved as output."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "363b53d0-5aaf-008e-66ed-a0d2132e9a8f"
      },
      "source": [
        "First, let's load data stuff!"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c6584a72-9b1d-9e10-8378-a8ba3786343a"
      },
      "outputs": [],
      "source": [
        "documents_categories_df = pd.read_csv(\"../input/documents_categories.csv\")\n",
        "documents_entities_df = pd.read_csv(\"../input/documents_entities.csv\")\n",
        "documents_meta_df = pd.read_csv(\"../input/documents_meta.csv\")\n",
        "documents_topics_df = pd.read_csv(\"../input/documents_topics.csv\")\n",
        "events_df = pd.read_csv(\"../input/events.csv\")\n",
        "page_views_df = pd.read_csv(\"../input/page_views_sample.csv\", nrows=50000)\n",
        "promoted_content_df = pd.read_csv(\"../input/promoted_content.csv\")\n",
        "clicks_train_df = pd.read_csv(\"../input/clicks_train.csv\", nrows=1000)\n",
        "\n",
        "print(\"Dataframes count:\")\n",
        "print(\"documents_categories - {0}\".format(len(documents_categories_df)))\n",
        "print(\"documents_entities - {0}\".format(len(documents_entities_df)))\n",
        "print(\"documents_meta - {0}\".format(len(documents_meta_df)))\n",
        "print(\"documents_topics - {0}\".format(len(documents_topics_df)))\n",
        "print(\"events - {0}\".format(len(events_df)))\n",
        "print(\"page_views - {0}\".format(len(page_views_df)))\n",
        "print(\"promoted_content - {0}\".format(len(promoted_content_df)))\n",
        "print(\"clicks_train - {0}\".format(len(clicks_train_df)))"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "e1cc9d54-4321-584f-0b10-1ea50396c71c"
      },
      "source": [
        "Now, we will explore data a little. First let me see what I have as train dataset. My intentions here are getting the most information from documents so we can train a model. Let's check our clicks_train dataframe first and see what we get."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "12e0b5fe-f10e-1d50-1163-a3d00caab383"
      },
      "outputs": [],
      "source": [
        "clicks_train_df.head(10)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b671a327-d3cc-03b7-c9a0-87c367b4a5ef"
      },
      "source": [
        "We have only our display_id (which leads us to the events dataset), ad_id (provided by promoted_content dataset) and a binary clicked state. All right, let's analyze these two datasets, events and promoted_content."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "8c025b00-80ff-f865-2588-d5111470496b"
      },
      "outputs": [],
      "source": [
        "events_df.head(10)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ee7c2a1e-b557-680a-e8b2-a8ae4386c5b3"
      },
      "outputs": [],
      "source": [
        "promoted_content_df.head(10)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "e00bdd64-fa8c-70c7-4b94-bbc1c5a69cb5"
      },
      "source": [
        "Wow, we have a lot of information here! On events dataset, we can get which user access which document, when the access was made. Now we can get some insights about documents users accessed. The first thing is understand a little about ads, so we look if we have more than one document per ad. "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "18099106-2f99-63dc-af5a-fc7e4f525ac2"
      },
      "outputs": [],
      "source": [
        "promoted_content_df.groupby([\"ad_id\"])[\"document_id\"].count()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "540e7167-e65e-e565-1e50-8d7faacc258e"
      },
      "source": [
        "Every ad appears only in one document, so we can get info based on document.  Doing this, we can later see what is the user preferences about content, so we can use this as a parameter for ad click decision. Now, let's check document data."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9cbaeae1-2e8b-fd7b-87d2-d715384ea06c"
      },
      "outputs": [],
      "source": [
        "documents_meta_df.head(10)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "dc1035d9-cb31-b9b5-52de-78a9dd826934"
      },
      "outputs": [],
      "source": [
        "documents_categories_df.head(10)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "38b229d4-8dc4-c28a-1185-43b743e77795"
      },
      "outputs": [],
      "source": [
        "documents_entities_df.head(10)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "fcef22cb-76fe-0af4-73ab-ce9700e3b074"
      },
      "outputs": [],
      "source": [
        "documents_topics_df.head(10)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "7ab7c477-55d7-205b-f67d-dc8dfa1062d8"
      },
      "source": [
        "Well, all document data are based on confidence level. So we don't have certainty about info since documents was classified by some machine learning process. We need to take account all the topics, categories and entities because we don't have a good estimation to just believe in the highest confidence level. One thing we can do is consider all items in a list, sorted by their confidence level. The problem about it is our machine learning model will need to know how to process this kind of parameter. But, before we make a decision about this thing, let's analyze the page_views dataset and see what insights we have."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "25e1c7aa-d2a0-0e46-259b-74d3e5df8da4"
      },
      "outputs": [],
      "source": [
        "page_views_df.head(10)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "794fe013-14d0-c252-f960-8002241e4980"
      },
      "source": [
        "Let's try some good things here. First, let's see some user behavior. We will search documents access."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c013e19a-bf11-e9a9-18e7-075e5b8509b7"
      },
      "outputs": [],
      "source": [
        "page_views_df.groupby(['document_id'])['uuid'].count().sort_values(ascending=False)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "04307746-0364-614b-3319-19bd4a578327"
      },
      "outputs": [],
      "source": ""
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.5.2"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}