{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "909b0967-ea67-344d-aab0-eec82bcef501"
      },
      "source": [
        "# What's in the data?\n",
        "## First thing is first:  I will get out the measuring tape and look at the data\n",
        "     1. Read in data, then check the heads and count them all up\n",
        "     2. See what it looks like, because hell, IDK what I am doing right now.\n",
        "\n",
        "\n",
        "## Second thing is second: Plot some of the data points\n",
        "     1. Plot the data with confidence level's (2)\n",
        "     2. Eat pizza (brain food for me)\n",
        "     3. Pizza gives powers, so I will google some thought provoking questions to answer here\n",
        "     4. Had the pizza, got really bloated. Damn. Feel really full, gonna netflix until next idea comes to me\n",
        "\n",
        "## Third thing is ...\n",
        "     1. So this is where I start  to ask my self what I know about the data?\n",
        "     2. Well there is a ton of it to say the least, but the question is rather a deeper one...\n",
        "     3. If I am on website 'x' what am I doing on it, and what am I going to look at, let alone click on\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9c93e13d-5bd5-84df-0d02-275638469828"
      },
      "outputs": [],
      "source": [
        "\n",
        "import pandas as pd\n",
        "import numpy as np\n",
        "%matplotlib inline  "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "318e4beb-8c4c-b01f-c8af-a70ff1e152c1"
      },
      "outputs": [],
      "source": [
        "# Yeah, I could have made a for loop. Some men would rather see the world burn. \n",
        "\n",
        "#df = pd.read_csv('../input/documents_categories.csv') #['document_id, category_id, confidence_level']\n",
        "#df1 = pd.read_csv('../input/clicks_test.csv') # ['Display_id, ad_id']\n",
        "#df2 = pd.read_csv('../input/documents_meta.csv') #['document_id, source_id, publisher_id, publish_time']\n",
        "#df3 = pd.read_csv('../input/documents_entities.csv') #['document_id, entity_id, confidence_level']\n",
        "#df4 = pd.read_csv('../input/promoted_content.csv') #['ad_id, document_id, campaign_id, advertister_id']\n",
        "#df5 = pd.read_csv('../input/sample_submission.csv') #['display_id, ad_id']\n",
        "#df6 = pd.read_csv('../input/documents_topics.csv') #['document_id, topic_id, confidence_level']\n",
        "#df7 = pd.read_csv('../input/clicks_train.csv') # ['display_id, ad_id, clicked']\n",
        "#df8 = pd.read_csv('../input/events.csv')# ['Display_id, uuid, document_id, timestamp, platform, geo_location']\n",
        "###df9 = pd.read_csv('../input/page_views.csv') \n",
        "#df10 = pd.read_csv('../input/page_views_sample.csv') #['uuid, document_id, timestamp, platform, geo_location, traffic_source' ]"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "08359009-7058-ca65-803a-e628453a7d86"
      },
      "source": [
        "Hmm weird, I can't read in the page_views.csv... Mehhh"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "16f5824c-470a-956a-c66c-9dbb73670f19"
      },
      "source": [
        "# Statistical Analysis #1 Machine Learning\n",
        "## **Documents_categories.csv**\n",
        "From the looks of it, there has been some pre-analysis done. \n",
        "\n",
        "More precisely, it has been done on **document_id** & **category_id**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "19df8ab4-83d1-be4d-449a-33830119a95a"
      },
      "outputs": [],
      "source": [
        "df = pd.read_csv('../input/documents_categories.csv') #['document_id, category_id, confidence_level']\n",
        "print (df.count()) # 5481475  (int64)\n",
        "print (df.head(10))"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "1fea1b0d-5d02-a9cd-69ee-27d1b835ba4c"
      },
      "source": [
        "#clicks_test.csv\n",
        "\n",
        " 1. List item"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "8d63f392-7843-1ae1-e6a9-fd4aa8b1ddbb"
      },
      "outputs": [],
      "source": [
        "\n",
        "df1 = pd.read_csv('../input/clicks_test.csv') # ['Display_id, ad_id']\n",
        "print (df1.count()) # 32225162\n",
        "print (df1.head(10))"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "890f9df8-613a-9a88-7bed-48066a8424e9"
      },
      "source": [
        "# documents_meta.csv"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "7781036c-926a-7206-c15d-586e2aafcfa5"
      },
      "outputs": [],
      "source": [
        "df2 = pd.read_csv('../input/documents_meta.csv') \n",
        "#['document_id, source_id, publisher_id, publish_time']\n",
        "print (df2.head())\n",
        "print (df2.count()) # Each id has a differing amount of data as a whole \n",
        "#dunno why but it's prob insignificant."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "d74246d1-3ec9-77a9-d80e-c61b1eb096f2"
      },
      "source": [
        "# Statistical Analysis #2 (Machine learning)\n",
        "Already baked into **documents_entities.csv** given the **\"confidence_level\"** var\n",
        "\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "24dd37f2-87a3-67bd-f0c6-8ff1bd4b9454"
      },
      "outputs": [],
      "source": [
        "df3 = pd.read_csv('../input/documents_entities.csv') #['document_id, entity_id, confidence_level']\n",
        "df3.count() #5537552"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "bd764663-609d-b1c9-ebf1-03f9483e38e9"
      },
      "source": [
        "# promoted_content.csv"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "646703b7-f9b0-735e-5f9f-563c410233b0"
      },
      "outputs": [],
      "source": [
        "df4 = pd.read_csv('../input/promoted_content.csv') #['ad_id, document_id, campaign_id, advertister_id']\n",
        "df.count() #5481475 \n",
        "# This is the same as documents_categories.csv"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4e35e7f2-3d4c-2c0b-feef-206c84514c57"
      },
      "source": [
        "# sample_submission.csv"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "7ae9cad5-7646-07d9-e7c7-b07bd4336722"
      },
      "outputs": [],
      "source": [
        "df5 = pd.read_csv('../input/sample_submission.csv') #['display_id, ad_id']\n",
        "print (df5.count())\n",
        "print (df5.head())\n",
        "#df5[1:3]\n",
        "#seems like the ad_id uses really big numbers, probably like a barcode for iding"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4456d3be-0592-99e6-4100-e1d21a8d6e47"
      },
      "source": [
        "# documents_topics.csv"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "10848b74-54ac-bf89-fada-275aef74c620"
      },
      "outputs": [],
      "source": [
        "df6 = pd.read_csv('../input/documents_topics.csv') #['document_id, topic_id, confidence_level']\n",
        "df6.count() # 11325960\n",
        "df6.hist()\n",
        "df6.plot(x='confidence_level', y='document_id', kind='kde')\n",
        "#hmmm that's very interst...\n",
        "# Changed x to be confid_lvl shit got effed up a bit\n",
        "# "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4fc9ff22-ef7a-7185-d962-8f85d4fbaae6"
      },
      "source": [
        "#clicks_train.csv "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a7f3705d-477e-7096-7c5e-c3626475c8a2"
      },
      "outputs": [],
      "source": [
        "df7 = pd.read_csv('../input/clicks_train.csv') # ['display_id, ad_id, clicked']\n",
        "df7.count() #87141731"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "509722ae-3c25-1f61-b621-e7720aa3c1f1"
      },
      "source": [
        "# events.csv"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c757fcd1-df4b-0458-0166-86b0dce4dda0"
      },
      "outputs": [],
      "source": [
        "df8 = pd.read_csv('../input/events.csv')# ['Display_id, uuid, document_id, timestamp, platform, geo_location']\n",
        "df8.count()"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b589ce59-29c1-a861-6fa3-2e0ac18a0f16"
      },
      "outputs": [],
      "source": [
        ""
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9c3489be-bfcb-4b50-03a9-c22eaa367347"
      },
      "outputs": [],
      "source": [
        ""
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.5.2"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}