{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "653a43bc-7d18-717a-02eb-9dfd3510a4af"
      },
      "source": [
        "## Starter Code ##\n",
        "This code is designed to provide a few much smaller basic files as well as a few files to get you started on the project."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "eebdfb7b-cbf5-c1fd-0f1c-5a0d6e51a7f7"
      },
      "source": [
        "This code creates a directory to collect all the reduced files."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c5c2b7b2-598a-44f7-295a-6e130374c7a1"
      },
      "outputs": [],
      "source": [
        "from os import mkdir\n",
        "from os.path import isdir\n",
        "import shutil\n",
        "\n",
        "if isdir('./events_100'):\n",
        "    shutil.rmtree('./events_100')\n",
        "mkdir('./events_100')"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "531545d2-2def-bcfa-3161-f5a1ea7ff584"
      },
      "source": [
        "The following loads a library (pandas) so that I can use it to perform analysis more easily. It also loads a function that randomly selects items from a list so that we can chose which 10 documents we are interested in.\n",
        "\n",
        "This loads the events file (and includes a few options) then selects 10 document_ids at random from it. After that, it filters the events file to only include events containing one of those documents. Finally, it writes the new smaller events data to a file."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b0fe0278-4f21-5b39-5574-c977b0d7ab8e"
      },
      "outputs": [],
      "source": [
        "import pandas as pd\n",
        "from numpy.random import choice\n",
        "\n",
        "with open('../input/events.csv') as in_file:\n",
        "    df = pd.read_csv(in_file,\n",
        "                     header=0,\n",
        "                     index_col=False,\n",
        "                     usecols=('display_id', 'document_id', 'platform', 'geo_location', 'timestamp'),\n",
        "                     dtype={'display_id': int, 'document_id': int, 'platform': str, 'geo_location': str, 'timestamp': int}\n",
        "                    )\n",
        "docs = choice(df['document_id'].unique(),100,False)\n",
        "df = df[df['document_id'].isin(docs)]\n",
        "x = df['geo_location'].apply(lambda x: (x.split('>') + [None, ]*3)[:3])\n",
        "df['geo_0'] = x.apply(lambda x:x[0])\n",
        "df['geo_1'] = x.apply(lambda x:x[1])\n",
        "df['geo_2'] = x.apply(lambda x:x[2])\n",
        "df.drop('geo_location', axis=1, inplace=True)\n",
        "df['platform'] = df['platform'].astype('category')\n",
        "df['timestamp'] = pd.to_datetime(df['timestamp']+1465876799998, unit='ms')\n",
        "df.to_csv('./events_100/events.csv')"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "d7c90451-08e3-857d-52bb-c78fbe52fbfd"
      },
      "source": [
        "Now that documents and corresponding display_ids are selected, we will load the document_id or display_ids from it as appropriate to join onto other tables.\n",
        "\n",
        "This starts by loading the file we just created, then we make lists of all distinct document_ids and display_ids in that file. Then we set df to null in order to free up memory."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "359f79a0-0b73-a167-ded8-87d99a428096"
      },
      "outputs": [],
      "source": [
        "with open('./events_100/events.csv') as in_file:\n",
        "    df = pd.read_csv(in_file,\n",
        "                     header=0,\n",
        "                     index_col=False,\n",
        "                     usecols=('display_id', 'document_id'),\n",
        "                     dtype={'display_id': int, 'document_id': int}\n",
        "                    )\n",
        "document_ids = df['document_id'].unique()\n",
        "display_ids = df['display_id'].unique()\n",
        "df = None"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "42a156e4-e04b-e75f-429b-8e18a3fa9814"
      },
      "source": [
        "Since we know the files that we will be loading ahead of time, we will be making a list of some basic information about them then iterating through that list. Specifically, we will be specifying the file name, which columns we care about, and whether it joins on display_id or document_id"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "d2c1351b-89bf-1fcc-f467-30afcc4aff46"
      },
      "outputs": [],
      "source": [
        "files = [\n",
        "    {'filename': 'documents_categories', 'join_on': 'document_id', 'columns': {'category_id': int, 'confidence_level': float}},\n",
        "    {'filename': 'documents_topics', 'join_on': 'document_id', 'columns': {'topic_id': int, 'confidence_level': float}},\n",
        "    {'filename': 'documents_entities', 'join_on': 'document_id', 'columns': {'entity_id': str, 'confidence': float}},\n",
        "    {'filename': 'documents_meta', 'join_on': 'document_id', 'columns': {'source_id': float, 'publisher_id': float, 'publish_time': str}},\n",
        "    {'filename': 'promoted_content', 'join_on': 'document_id', 'columns': {'campaign_id': int, 'advertiser_id': int, 'ad_id': int}},\n",
        "    {'filename': 'clicks_train', 'join_on': 'display_id', 'columns': {'ad_id': int, 'clicked': bool}},\n",
        "]"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b1338f5c-8281-6e47-cbed-d245bc500c67"
      },
      "source": [
        "Now that we have a list of properties for each file we are going to look at, we will create smaller files for each of them."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "eec34ee2-1beb-c777-c997-1e7359a36ade"
      },
      "outputs": [],
      "source": [
        "for f in files:\n",
        "    df = pd.read_csv(\"../input/%s.csv\" % f['filename'],\n",
        "                     header=0,\n",
        "                     index_col=False,\n",
        "                     usecols=list(f['columns'].keys()).extend(f['join_on']),\n",
        "                     dtype=f['columns']\n",
        "                    )\n",
        "    if f['join_on'] == 'document_id':\n",
        "        df = df[df['document_id'].isin(document_ids)]\n",
        "    else:\n",
        "        df = df[df['display_id'].isin(display_ids)]\n",
        "    df.to_csv(\"./events_100/%s.csv\" % f['filename'])\n",
        "    df = None\n",
        "    print(\"File %s has been created\" % f['filename'])"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.5.2"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}