{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "2f903da0-a9ed-9ab4-6e58-eea1e5e09d6f"
      },
      "source": [
        "As I checked if access information for landing page of ad clicks is in page_views.csv for clicks_test.csv in [this script](https://www.kaggle.com/its7171/outbrain-click-prediction/leakage-solution/discussion), I tryed to check for clicks_train.csv to estimate this efects."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a59fcf76-6900-80e6-6511-0de2654dfc3d"
      },
      "outputs": [],
      "source": [
        "import numpy as np\n",
        "import pandas as pd\n",
        "\n",
        "# set None for full data\n",
        "#nrows=None\n",
        "nrows=10000000\n",
        "\n",
        "# Is access information for landing page of ad click in page_views.csv?\n",
        "df_train = pd.read_csv('../input/clicks_train.csv', nrows=nrows)\n",
        "df_ad = pd.read_csv('../input/promoted_content.csv', nrows=nrows,\n",
        "    usecols  = ('ad_id','document_id'),\n",
        "    dtype={'ad_id': np.int, 'uuid': np.str, 'document_id': np.str})\n",
        "df_events = pd.read_csv('../input/events.csv',nrows=nrows,\n",
        "    usecols  = ('display_id','uuid','timestamp'),\n",
        "    dtype={'display_id': np.int, 'uuid': np.str, 'timestamp': np.int})\n",
        "df_train = pd.merge(df_train, df_ad, on='ad_id', how='left')\n",
        "df_train = pd.merge(df_train, df_events, on='display_id', how='left')\n",
        "df_train['usr_doc'] = df_train['uuid'] + '_' + df_train['document_id']\n",
        "df_train = df_train.set_index('usr_doc')\n",
        "time_dict = df_train[['timestamp']].to_dict()['timestamp']\n",
        "# set page_views.csv for full data\n",
        "# f = open(\"../input/page_views.csv\", \"r\")\n",
        "f = open(\"../input/page_views_sample.csv\",\"r\")\n",
        "line = f.readline().strip()\n",
        "head_arr = line.split(\",\")\n",
        "fld_index = dict(zip(head_arr,range(0,len(head_arr))))\n",
        "total = 0\n",
        "while 1:\n",
        "    line = f.readline().strip()\n",
        "    if nrows is not None and total == nrows:\n",
        "        break\n",
        "    total += 1\n",
        "    if line == '':\n",
        "        break\n",
        "    arr = line.split(\",\")\n",
        "    usr_doc = arr[fld_index['uuid']] + '_' + arr[fld_index['document_id']]\n",
        "    if usr_doc in time_dict:\n",
        "        #don't use timestamp yet.\n",
        "        #time_diff = time_dict[usr_doc] - int(arr[fld_index['timestamp']])\n",
        "        #if abs(time_diff) < 600:\n",
        "            # set -1 if found that this user sow this document\n",
        "            time_dict[usr_doc] = -1\n",
        "\n",
        "df_train=df_train.reset_index()\n",
        "df_train['fixed_timestamp'] = df_train['usr_doc'].apply(lambda x: time_dict[x])\n",
        "found_in_page_views = set(df_train[df_train['fixed_timestamp'] < 0].index)\n",
        "clicked = set(df_train[df_train['clicked'] == 1].index)\n",
        "all_ids = set(df_train.index)\n",
        "\n",
        "TP = len(clicked & found_in_page_views)\n",
        "FP = len(found_in_page_views - clicked)\n",
        "FN = len(clicked - found_in_page_views)\n",
        "recall = TP/float(TP+FN)\n",
        "precision = TP/float(TP+FP)\n",
        "print('TP:{}'.format(TP))\n",
        "print('FP:{}'.format(FP))\n",
        "print('FN:{}'.format(FN))\n",
        "print('recall:{0:.1f}%'.format(recall*100))\n",
        "print('precision:{0:.1f}%'.format(precision*100))"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "19e25aa7-c3c9-1f1d-e745-1e6bedd5a70d"
      },
      "source": [
        "For full data, this output would be:\n",
        "\n",
        "<pre>\n",
        "TP:724749\n",
        "FP:31813\n",
        "FN:16149844\n",
        "recall:4.3%\n",
        "precision:95.8%\n",
        "</pre>"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b8601b9e-cdc5-e75c-4fcd-810e36873375"
      },
      "source": [
        "Only 4.3% of Click data is found in page_views.csv.\n",
        "This percentage is far smaller than I thought.\n",
        "I thought that landing page log should be in page_views.csv, if some ad link is clicked.\n",
        "Where is remaining 95.7% access?\n",
        "Does page_views.csv includes only 4.3% sampling data?\n",
        "Or Outbrain does not have all page view data for landing page of ads?\n",
        "\n",
        "update: I guess that page_views.csv includes access logs for all the page which has ads in it. If ad landing page dose not have ads, the access for the landing page will not be recorded in page_views.csv. So only 4.3% of ad may have ads in the landing page.\n",
        " \n",
        "On the other hand if access information for landing page of ad clicks is found in page_views.csv, 95.8% of them are clicked.\n",
        "Suppose test data has same high precision, this feature would be useful."
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.5.2"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}