{
  "id": 24110,
  "title": "Difference between events.csv and page_views.csv",
  "url": "/competitions/outbrain-click-prediction/discussion/24110",
  "author_name": "",
  "post_date": "2016-10-05T20:06:57.107Z",
  "votes": null,
  "comment_count": 13,
  "views": 2123,
  "content": "<p>Hi all! Excited to be starting this competition.</p>\n\n<p>I was wondering what the difference was between the events.csv and page_views.csv files. They both seem to include many of the same variables, although the events.csv seems to include the display_id while page_views.csv includes the traffic_source. Can anyone shed some light on this? Thanks!</p>",
  "messages": [
    {
      "id": "137875",
      "postDate": "10/05/2016 20:06:57",
      "content": "<p>Hi all! Excited to be starting this competition.</p>\n\n<p>I was wondering what the difference was between the events.csv and page_views.csv files. They both seem to include many of the same variables, although the events.csv seems to include the display_id while page_views.csv includes the traffic_source. Can anyone shed some light on this? Thanks!</p>",
      "rawMarkdown": "Hi all! Excited to be starting this competition.\r\n\r\nI was wondering what the difference was between the events.csv and page_views.csv files. They both seem to include many of the same variables, although the events.csv seems to include the display_id while page_views.csv includes the traffic_source. Can anyone shed some light on this? Thanks!",
      "votes": null
    },
    {
      "id": "137895",
      "postDate": "10/05/2016 21:35:20",
      "content": "<p>I'm pretty sure the difference is that page_views.csv  has all of the visits that outbrain was able to track, while events.csv only describes page views that led to a click.</p>",
      "rawMarkdown": "I'm pretty sure the difference is that page_views.csv  has all of the visits that outbrain was able to track, while events.csv only describes page views that led to a click.",
      "votes": null
    },
    {
      "id": "137901",
      "postDate": "10/05/2016 22:00:12",
      "content": "<p>Every entry in events.csv is a different click event organized by the unique event id (technically the 'display_id'). Every entry is page_views.csv also represents a different click event, but is organized by 'document_id'. Unfortunately each entry in page_views.csv does not correspond to a unique user, so if you want to join the two tables you ought to join on 'timestamp'. I think...</p>",
      "rawMarkdown": "Every entry in events.csv is a different click event organized by the unique event id (technically the 'display_id'). Every entry is page_views.csv also represents a different click event, but is organized by 'document_id'. Unfortunately each entry in page_views.csv does not correspond to a unique user, so if you want to join the two tables you ought to join on 'timestamp'. I think...",
      "votes": null
    },
    {
      "id": "138019",
      "postDate": "10/06/2016 16:38:49",
      "content": "<p>Ok, but how can we sure about timestamp for both datasets? They didn't say if events timestamp is also relative to the first time in dataset. Ideally, timestamps should be comparable, but I don't know if I need to convert the page_views timestamp first and than compare it with events timestamp or do the comparison right away. </p>\n\n<p>Also, for me, page_views is just another source of data, like the other datasets, I don't think just a direct join of the two datasets will help, but I might be wrong about this. Maybe we need to group some data first, get statistical data regarding documents from page_views, but this is just an idea.</p>",
      "rawMarkdown": "Ok, but how can we sure about timestamp for both datasets? They didn't say if events timestamp is also relative to the first time in dataset. Ideally, timestamps should be comparable, but I don't know if I need to convert the page_views timestamp first and than compare it with events timestamp or do the comparison right away. \r\n\r\nAlso, for me, page_views is just another source of data, like the other datasets, I don't think just a direct join of the two datasets will help, but I might be wrong about this. Maybe we need to group some data first, get statistical data regarding documents from page_views, but this is just an idea.",
      "votes": null
    },
    {
      "id": "138694",
      "postDate": "10/10/2016 12:43:23",
      "content": "<p>Well it can't be the same data organised by different attributes because page_views is huge and events is (relatively) tiny.  My inference is that:</p>\n\n<ol>\n<li>An event represents a page view where ads were at least displayed and possibly (probably) one was clicked.  Navigating events.display_id -&gt; clicks_test.display_id will return a set of all the ads that were displayed on the page.  Navigating to promoted_content on ad_id will give the details of the ads.  Navigating to the documents_... tables on document_id will give you the details of the page that the ads were displayed on.</li>\n<li>A page view represents a page that was viewed by a tracked user.  You can navigate to the documents_... tables on document_id to collect information about the page.</li>\n</ol>\n\n<p>It looks to me like you could take an event, investigate the details of the page the ads were on, and the details of the ad.  You could then use the uuid to build up a profile of other page views for the user.</p>\n\n<p>Note to Kaggle... I think it's pretty poor that this question has been on the forum for five days now and no response!  -1 from me.</p>",
      "rawMarkdown": "Well it can't be the same data organised by different attributes because page_views is huge and events is (relatively) tiny.  My inference is that:\r\n\r\n 1. An event represents a page view where ads were at least displayed and possibly (probably) one was clicked.  Navigating events.display_id -> clicks_test.display_id will return a set of all the ads that were displayed on the page.  Navigating to promoted_content on ad_id will give the details of the ads.  Navigating to the documents_... tables on document_id will give you the details of the page that the ads were displayed on.\r\n 2. A page view represents a page that was viewed by a tracked user.  You can navigate to the documents_... tables on document_id to collect information about the page.\r\n\r\nIt looks to me like you could take an event, investigate the details of the page the ads were on, and the details of the ad.  You could then use the uuid to build up a profile of other page views for the user.\r\n\r\nNote to Kaggle... I think it's pretty poor that this question has been on the forum for five days now and no response!  -1 from me.",
      "votes": null
    },
    {
      "id": "138696",
      "postDate": "10/10/2016 12:53:26",
      "content": "<p>@Terrible Tadpole, it is the same thing I am discovering so far, but in my preliminary data exploration, it seems users accessed just one document, so page views dataset is useless to build this kind of profile about user. The only useful info I get from the dataset so far is document access count and distribution of that access across geo location.</p>",
      "rawMarkdown": "Terrible Tadpole, it is the same thing I am discovering so far, but in my preliminary data exploration, it seems users accessed just one document, so page views dataset is useless to build this kind of profile about user. The only useful info I get from the dataset so far is document access count and distribution of that access across geo location.",
      "votes": null
    },
    {
      "id": "138700",
      "postDate": "10/10/2016 13:47:25",
      "content": "<p>@TerribleTadpole your interpretation is correct. events.csv represents ad displays (in clicks_test.csv and clicks_train.csv) and there is always a click. page_views.csv is just users viewing documents.</p>",
      "rawMarkdown": "TerribleTadpole your interpretation is correct. events.csv represents ad displays (in clicks_test.csv and clicks_train.csv) and there is always a click. page_views.csv is just users viewing documents.",
      "votes": null
    },
    {
      "id": "138751",
      "postDate": "10/10/2016 19:03:09",
      "content": "<p>Think of an EVENT as a distinct interaction of the reader with the main document page. Every time you visit the page, an &quot;event&quot; is generated. </p>\n\n<p>For instance, this is the main document of interest (to you, as a reader):</p>\n\n<p><a href=\"http://www.cnn.com/2016/10/10/politics/paul-ryan-said-he-wont-defend-donald-trump/index.html\">http://www.cnn.com/2016/10/10/politics/paul-ryan-said-he-wont-defend-donald-trump/index.html</a></p>\n\n<p>Everytime you visit this page, you will see Outbrain's actual paid content section recommending different &quot;SET OF ADS&quot; to you. Try it out and see in action. </p>\n\n<p>The page view creates the following info:</p>\n\n<pre><code>    - Who (user)\n    - What (document)\n    - When (day and time)\n    - Where (geo)\n    - Platform (desktop, tablet, mobile)\n    - Source (user source)\n</code></pre>\n\n<p>Notice it does not have anything to do with Ads (paid/promoted content). </p>\n\n<p>Event however captures that info//</p>",
      "rawMarkdown": "Think of an EVENT as a distinct interaction of the reader with the main document page. Every time you visit the page, an \"event\" is generated. \r\n\r\nFor instance, this is the main document of interest (to you, as a reader):\r\n\r\nhttp://www.cnn.com/2016/10/10/politics/paul-ryan-said-he-wont-defend-donald-trump/index.html\r\n\r\nEverytime you visit this page, you will see Outbrain's actual paid content section recommending different \"SET OF ADS\" to you. Try it out and see in action. \r\n\r\nThe page view creates the following info:\r\n\r\n        - Who (user)\r\n        - What (document)\r\n        - When (day and time)\r\n        - Where (geo)\r\n        - Platform (desktop, tablet, mobile)\r\n        - Source (user source)\r\n\r\nNotice it does not have anything to do with Ads (paid/promoted content). \r\n\r\nEvent however captures that info//",
      "votes": null
    },
    {
      "id": "138827",
      "postDate": "10/11/2016 02:11:20",
      "content": "<p>If we go from clicks_test/train  -&gt; events to get a list of display_id, uuid, ad_id.  This gives you a list of which user saw the ads in a particular display_id.</p>\n\n<p>Now suppose you take this and join to promoted_content to get display_id, uuid, ad_id, promoted_content.document_id.   This gives you a list of the potential documents the user could have routed to after clicking the ad.</p>\n\n<p>If the uuid clicked on an ad_id which was promoting a particular document_id, shouldn't page_views have a record for that uuid visiting the promoted document_id?  I am finding this to not be the case.</p>\n\n<p>Am I misunderstanding the relationships here?</p>",
      "rawMarkdown": "If we go from clicks_test/train  -> events to get a list of display_id, uuid, ad_id.  This gives you a list of which user saw the ads in a particular display_id.\r\n\r\nNow suppose you take this and join to promoted_content to get display_id, uuid, ad_id, promoted_content.document_id.   This gives you a list of the potential documents the user could have routed to after clicking the ad.\r\n\r\nIf the uuid clicked on an ad_id which was promoting a particular document_id, shouldn't page_views have a record for that uuid visiting the promoted document_id?  I am finding this to not be the case.\r\n\r\nAm I misunderstanding the relationships here?",
      "votes": null
    },
    {
      "id": "138853",
      "postDate": "10/11/2016 08:06:32",
      "content": "<p>Okay... in the absence of anything authoritative (still) from Kaggle or Outbrain, here's my take:</p>\n\n<p>clicks_test gives you all the ads shown for a given event (display_id).\nSo select * from clicks_test where display_id = 16874594; gives you 6 rows, which means six ads were shown on that occasion.</p>\n\n<p>The events data will tell you what page those ads were embedded in, and the other attributes of how the page was accessed and by whom:\nselect * from events where display_id = 16874594; will give you one row with the uuid and the document_id</p>\n\n<p>You can use that document_id to find more information about the page from the documents_... tables (meta, entities, topics, categories).  Note that we're still talking about the page in which the ad was embedded.</p>\n\n<p>To look at the page that the ad links to, you need the promoted_content data:\nselect * from promoted_content where ad_id = 162754; will return 1 row, from which I infer that an ad always links to the same web page.  However, if you query you will find that many ads may link to the same web page:\nselect * from promoted_content where document_id = 1292723; will return 38 ads, all distinct, reinforcing my inference that an ad always links to the same page, but that page may be the target of many ads.</p>\n\n<p>The page_views data appears to be intended as a cloud of other, unrelated, non-promoted content that the user is believed to have visited.  I believe that the essence of this competition is to analyse this cloud to ascertain attributes of the user's viewing behaviour that are predictive of their ad-clicking behaviour.  Putting the views of promoted material would give our algorithms data that a real-life application would not have; namely that in real life the user would not yet have clicked an ad. </p>\n\n<p>All this is just my opinion of course... not gospel!</p>",
      "rawMarkdown": "Okay... in the absence of anything authoritative (still) from Kaggle or Outbrain, here's my take:\r\n\r\nclicks_test gives you all the ads shown for a given event (display_id).\r\nSo select * from clicks_test where display_id = 16874594; gives you 6 rows, which means six ads were shown on that occasion.\r\n\r\nThe events data will tell you what page those ads were embedded in, and the other attributes of how the page was accessed and by whom:\r\nselect * from events where display_id = 16874594; will give you one row with the uuid and the document_id\r\n\r\nYou can use that document_id to find more information about the page from the documents_... tables (meta, entities, topics, categories).  Note that we're still talking about the page in which the ad was embedded.\r\n\r\nTo look at the page that the ad links to, you need the promoted_content data:\r\nselect * from promoted_content where ad_id = 162754; will return 1 row, from which I infer that an ad always links to the same web page.  However, if you query you will find that many ads may link to the same web page:\r\nselect * from promoted_content where document_id = 1292723; will return 38 ads, all distinct, reinforcing my inference that an ad always links to the same page, but that page may be the target of many ads.\r\n\r\nThe page_views data appears to be intended as a cloud of other, unrelated, non-promoted content that the user is believed to have visited.  I believe that the essence of this competition is to analyse this cloud to ascertain attributes of the user's viewing behaviour that are predictive of their ad-clicking behaviour.  Putting the views of promoted material would give our algorithms data that a real-life application would not have; namely that in real life the user would not yet have clicked an ad. \r\n\r\nAll this is just my opinion of course... not gospel!",
      "votes": null
    },
    {
      "id": "139404",
      "postDate": "10/13/2016 22:39:04",
      "content": "<p>Terrible Tadpole is up to something. Good comments! </p>",
      "rawMarkdown": "Terrible Tadpole is up to something. Good comments!",
      "votes": null
    },
    {
      "id": "139465",
      "postDate": "10/14/2016 09:37:12",
      "content": "<p>@Bryan</p>\n\n<p>I believe that the ads can lead to pages that are outside of Outbrain control, think pages like your personal blog for example. So those clicks will not be present in page_views.</p>\n\n<p>The promoted_content give us info about the ad and on which docs the ad was visible, but there is no indication about the &quot;routing&quot; after the ad is clicked.</p>\n\n<p>[quote=Bryan Johnson;138827]</p>\n\n<p>If we go from clicks_test/train  -&gt; events to get a list of display_id, uuid, ad_id.  This gives you a list of which user saw the ads in a particular display_id.</p>\n\n<p>Now suppose you take this and join to promoted_content to get display_id, uuid, ad_id, promoted_content.document_id.   This gives you a list of the potential documents the user could have routed to after clicking the ad.</p>\n\n<p>If the uuid clicked on an ad_id which was promoting a particular document_id, shouldn't page_views have a record for that uuid visiting the promoted document_id?  I am finding this to not be the case.</p>\n\n<p>Am I misunderstanding the relationships here?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Bryan\r\n\r\nI believe that the ads can lead to pages that are outside of Outbrain control, think pages like your personal blog for example. So those clicks will not be present in page_views.\r\n\r\nThe promoted_content give us info about the ad and on which docs the ad was visible, but there is no indication about the \"routing\" after the ad is clicked.\r\n\r\n\r\n[quote=Bryan Johnson;138827]\r\n\r\nIf we go from clicks_test/train  -> events to get a list of display_id, uuid, ad_id.  This gives you a list of which user saw the ads in a particular display_id.\r\n\r\nNow suppose you take this and join to promoted_content to get display_id, uuid, ad_id, promoted_content.document_id.   This gives you a list of the potential documents the user could have routed to after clicking the ad.\r\n\r\nIf the uuid clicked on an ad_id which was promoting a particular document_id, shouldn't page_views have a record for that uuid visiting the promoted document_id?  I am finding this to not be the case.\r\n\r\nAm I misunderstanding the relationships here?\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "139790",
      "postDate": "10/16/2016 16:32:07",
      "content": "<p>The fact that distinct display_ids from clicks_train and clicks_test are exactly the display_ids of events is a very strong evidence that events are just attributes of clicks (<a href=\"https://www.kaggle.com/andris/outbrain-click-prediction/schema-inference-1\">Script1</a>).</p>\n\n<p>Therefore, I interpret page_views as logs of when visitors visit documents, while both clicks left joined to events represent logs of visitors clicking on ads displayed in the documents, and views include clicks/events (<a href=\"https://www.kaggle.com/andris/outbrain-click-prediction/schema-inference-2\">Script2</a>).</p>\n\n<p>Too bad we do not get to see which ads were displayed on those visits when visitors did NOT click (this is a hint to Outbrain ;) ).</p>",
      "rawMarkdown": "The fact that distinct display_ids from clicks_train and clicks_test are exactly the display_ids of events is a very strong evidence that events are just attributes of clicks ([Script1][1]).\r\n\r\nTherefore, I interpret page_views as logs of when visitors visit documents, while both clicks left joined to events represent logs of visitors clicking on ads displayed in the documents, and views include clicks/events ([Script2][2]).\r\n\r\nToo bad we do not get to see which ads were displayed on those visits when visitors did NOT click (this is a hint to Outbrain ;) ).\r\n\r\n\r\n  [1]: https://www.kaggle.com/andris/outbrain-click-prediction/schema-inference-1\r\n  [2]: https://www.kaggle.com/andris/outbrain-click-prediction/schema-inference-2",
      "votes": null
    },
    {
      "id": "139989",
      "postDate": "10/17/2016 19:33:43",
      "content": "<p>I've just uploaded the <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/\">&quot;Unveiling page_views.csv with PySpark&quot; notebook</a> with full analytics of page_views.csv and its relationship with events.csv. </p>\n\n<p>After some hours of processing in a Spark Cluster, we've got answer to questions like:  </p>\n\n<ul>\n<li><strong>How to join page_views.csv and events.csv?</strong></li>\n<li><strong>Is events.csv a subset of page_views.csv?</strong></li>\n<li><strong>Are there additional page views for users in events.csv?</strong></li>\n</ul>\n\n<p>Enjoy, and please upvote if you like ;)</p>",
      "rawMarkdown": "I've just uploaded the [\"Unveiling page_views.csv with PySpark\" notebook](https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/) with full analytics of page_views.csv and its relationship with events.csv. \r\n\r\nAfter some hours of processing in a Spark Cluster, we've got answer to questions like:  \r\n\r\n- **How to join page_views.csv and events.csv?**\r\n- **Is events.csv a subset of page_views.csv?**\r\n- **Are there additional page views for users in events.csv?**\r\n\r\n\r\nEnjoy, and please upvote if you like ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 137895,
      "author_name": "midikiman",
      "author_url": "",
      "post_date": "10/05/2016 21:35:20",
      "content": "<p>I'm pretty sure the difference is that page_views.csv  has all of the visits that outbrain was able to track, while events.csv only describes page views that led to a click.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137901,
      "author_name": "austincap",
      "author_url": "",
      "post_date": "10/05/2016 22:00:12",
      "content": "<p>Every entry in events.csv is a different click event organized by the unique event id (technically the 'display_id'). Every entry is page_views.csv also represents a different click event, but is organized by 'document_id'. Unfortunately each entry in page_views.csv does not correspond to a unique user, so if you want to join the two tables you ought to join on 'timestamp'. I think...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138019,
      "author_name": "geekfox",
      "author_url": "",
      "post_date": "10/06/2016 16:38:49",
      "content": "<p>Ok, but how can we sure about timestamp for both datasets? They didn't say if events timestamp is also relative to the first time in dataset. Ideally, timestamps should be comparable, but I don't know if I need to convert the page_views timestamp first and than compare it with events timestamp or do the comparison right away. </p>\n\n<p>Also, for me, page_views is just another source of data, like the other datasets, I don't think just a direct join of the two datasets will help, but I might be wrong about this. Maybe we need to group some data first, get statistical data regarding documents from page_views, but this is just an idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138694,
      "author_name": "terribletadpole",
      "author_url": "",
      "post_date": "10/10/2016 12:43:23",
      "content": "<p>Well it can't be the same data organised by different attributes because page_views is huge and events is (relatively) tiny.  My inference is that:</p>\n\n<ol>\n<li>An event represents a page view where ads were at least displayed and possibly (probably) one was clicked.  Navigating events.display_id -&gt; clicks_test.display_id will return a set of all the ads that were displayed on the page.  Navigating to promoted_content on ad_id will give the details of the ads.  Navigating to the documents_... tables on document_id will give you the details of the page that the ads were displayed on.</li>\n<li>A page view represents a page that was viewed by a tracked user.  You can navigate to the documents_... tables on document_id to collect information about the page.</li>\n</ol>\n\n<p>It looks to me like you could take an event, investigate the details of the page the ads were on, and the details of the ad.  You could then use the uuid to build up a profile of other page views for the user.</p>\n\n<p>Note to Kaggle... I think it's pretty poor that this question has been on the forum for five days now and no response!  -1 from me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138696,
      "author_name": "geekfox",
      "author_url": "",
      "post_date": "10/10/2016 12:53:26",
      "content": "<p>@Terrible Tadpole, it is the same thing I am discovering so far, but in my preliminary data exploration, it seems users accessed just one document, so page views dataset is useless to build this kind of profile about user. The only useful info I get from the dataset so far is document access count and distribution of that access across geo location.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138700,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "10/10/2016 13:47:25",
      "content": "<p>@TerribleTadpole your interpretation is correct. events.csv represents ad displays (in clicks_test.csv and clicks_train.csv) and there is always a click. page_views.csv is just users viewing documents.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138751,
      "author_name": "ceeker",
      "author_url": "",
      "post_date": "10/10/2016 19:03:09",
      "content": "<p>Think of an EVENT as a distinct interaction of the reader with the main document page. Every time you visit the page, an &quot;event&quot; is generated. </p>\n\n<p>For instance, this is the main document of interest (to you, as a reader):</p>\n\n<p><a href=\"http://www.cnn.com/2016/10/10/politics/paul-ryan-said-he-wont-defend-donald-trump/index.html\">http://www.cnn.com/2016/10/10/politics/paul-ryan-said-he-wont-defend-donald-trump/index.html</a></p>\n\n<p>Everytime you visit this page, you will see Outbrain's actual paid content section recommending different &quot;SET OF ADS&quot; to you. Try it out and see in action. </p>\n\n<p>The page view creates the following info:</p>\n\n<pre><code>    - Who (user)\n    - What (document)\n    - When (day and time)\n    - Where (geo)\n    - Platform (desktop, tablet, mobile)\n    - Source (user source)\n</code></pre>\n\n<p>Notice it does not have anything to do with Ads (paid/promoted content). </p>\n\n<p>Event however captures that info//</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138827,
      "author_name": "bryanpjohnson",
      "author_url": "",
      "post_date": "10/11/2016 02:11:20",
      "content": "<p>If we go from clicks_test/train  -&gt; events to get a list of display_id, uuid, ad_id.  This gives you a list of which user saw the ads in a particular display_id.</p>\n\n<p>Now suppose you take this and join to promoted_content to get display_id, uuid, ad_id, promoted_content.document_id.   This gives you a list of the potential documents the user could have routed to after clicking the ad.</p>\n\n<p>If the uuid clicked on an ad_id which was promoting a particular document_id, shouldn't page_views have a record for that uuid visiting the promoted document_id?  I am finding this to not be the case.</p>\n\n<p>Am I misunderstanding the relationships here?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138853,
      "author_name": "terribletadpole",
      "author_url": "",
      "post_date": "10/11/2016 08:06:32",
      "content": "<p>Okay... in the absence of anything authoritative (still) from Kaggle or Outbrain, here's my take:</p>\n\n<p>clicks_test gives you all the ads shown for a given event (display_id).\nSo select * from clicks_test where display_id = 16874594; gives you 6 rows, which means six ads were shown on that occasion.</p>\n\n<p>The events data will tell you what page those ads were embedded in, and the other attributes of how the page was accessed and by whom:\nselect * from events where display_id = 16874594; will give you one row with the uuid and the document_id</p>\n\n<p>You can use that document_id to find more information about the page from the documents_... tables (meta, entities, topics, categories).  Note that we're still talking about the page in which the ad was embedded.</p>\n\n<p>To look at the page that the ad links to, you need the promoted_content data:\nselect * from promoted_content where ad_id = 162754; will return 1 row, from which I infer that an ad always links to the same web page.  However, if you query you will find that many ads may link to the same web page:\nselect * from promoted_content where document_id = 1292723; will return 38 ads, all distinct, reinforcing my inference that an ad always links to the same page, but that page may be the target of many ads.</p>\n\n<p>The page_views data appears to be intended as a cloud of other, unrelated, non-promoted content that the user is believed to have visited.  I believe that the essence of this competition is to analyse this cloud to ascertain attributes of the user's viewing behaviour that are predictive of their ad-clicking behaviour.  Putting the views of promoted material would give our algorithms data that a real-life application would not have; namely that in real life the user would not yet have clicked an ad. </p>\n\n<p>All this is just my opinion of course... not gospel!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139404,
      "author_name": "henrique1977",
      "author_url": "",
      "post_date": "10/13/2016 22:39:04",
      "content": "<p>Terrible Tadpole is up to something. Good comments! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139465,
      "author_name": "mdbmilkov",
      "author_url": "",
      "post_date": "10/14/2016 09:37:12",
      "content": "<p>@Bryan</p>\n\n<p>I believe that the ads can lead to pages that are outside of Outbrain control, think pages like your personal blog for example. So those clicks will not be present in page_views.</p>\n\n<p>The promoted_content give us info about the ad and on which docs the ad was visible, but there is no indication about the &quot;routing&quot; after the ad is clicked.</p>\n\n<p>[quote=Bryan Johnson;138827]</p>\n\n<p>If we go from clicks_test/train  -&gt; events to get a list of display_id, uuid, ad_id.  This gives you a list of which user saw the ads in a particular display_id.</p>\n\n<p>Now suppose you take this and join to promoted_content to get display_id, uuid, ad_id, promoted_content.document_id.   This gives you a list of the potential documents the user could have routed to after clicking the ad.</p>\n\n<p>If the uuid clicked on an ad_id which was promoting a particular document_id, shouldn't page_views have a record for that uuid visiting the promoted document_id?  I am finding this to not be the case.</p>\n\n<p>Am I misunderstanding the relationships here?</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139790,
      "author_name": "andris",
      "author_url": "",
      "post_date": "10/16/2016 16:32:07",
      "content": "<p>The fact that distinct display_ids from clicks_train and clicks_test are exactly the display_ids of events is a very strong evidence that events are just attributes of clicks (<a href=\"https://www.kaggle.com/andris/outbrain-click-prediction/schema-inference-1\">Script1</a>).</p>\n\n<p>Therefore, I interpret page_views as logs of when visitors visit documents, while both clicks left joined to events represent logs of visitors clicking on ads displayed in the documents, and views include clicks/events (<a href=\"https://www.kaggle.com/andris/outbrain-click-prediction/schema-inference-2\">Script2</a>).</p>\n\n<p>Too bad we do not get to see which ads were displayed on those visits when visitors did NOT click (this is a hint to Outbrain ;) ).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139989,
      "author_name": "gspmoreira",
      "author_url": "",
      "post_date": "10/17/2016 19:33:43",
      "content": "<p>I've just uploaded the <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/\">&quot;Unveiling page_views.csv with PySpark&quot; notebook</a> with full analytics of page_views.csv and its relationship with events.csv. </p>\n\n<p>After some hours of processing in a Spark Cluster, we've got answer to questions like:  </p>\n\n<ul>\n<li><strong>How to join page_views.csv and events.csv?</strong></li>\n<li><strong>Is events.csv a subset of page_views.csv?</strong></li>\n<li><strong>Are there additional page views for users in events.csv?</strong></li>\n</ul>\n\n<p>Enjoy, and please upvote if you like ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "137875": "Hi all! Excited to be starting this competition.\r\n\r\nI was wondering what the difference was between the events.csv and page_views.csv files. They both seem to include many of the same variables, although the events.csv seems to include the display_id while page_views.csv includes the traffic_source. Can anyone shed some light on this? Thanks!",
    "137895": "I'm pretty sure the difference is that page_views.csv  has all of the visits that outbrain was able to track, while events.csv only describes page views that led to a click.",
    "137901": "Every entry in events.csv is a different click event organized by the unique event id (technically the 'display_id'). Every entry is page_views.csv also represents a different click event, but is organized by 'document_id'. Unfortunately each entry in page_views.csv does not correspond to a unique user, so if you want to join the two tables you ought to join on 'timestamp'. I think...",
    "138019": "Ok, but how can we sure about timestamp for both datasets? They didn't say if events timestamp is also relative to the first time in dataset. Ideally, timestamps should be comparable, but I don't know if I need to convert the page_views timestamp first and than compare it with events timestamp or do the comparison right away. \r\n\r\nAlso, for me, page_views is just another source of data, like the other datasets, I don't think just a direct join of the two datasets will help, but I might be wrong about this. Maybe we need to group some data first, get statistical data regarding documents from page_views, but this is just an idea.",
    "138694": "Well it can't be the same data organised by different attributes because page_views is huge and events is (relatively) tiny.  My inference is that:\r\n\r\n 1. An event represents a page view where ads were at least displayed and possibly (probably) one was clicked.  Navigating events.display_id -> clicks_test.display_id will return a set of all the ads that were displayed on the page.  Navigating to promoted_content on ad_id will give the details of the ads.  Navigating to the documents_... tables on document_id will give you the details of the page that the ads were displayed on.\r\n 2. A page view represents a page that was viewed by a tracked user.  You can navigate to the documents_... tables on document_id to collect information about the page.\r\n\r\nIt looks to me like you could take an event, investigate the details of the page the ads were on, and the details of the ad.  You could then use the uuid to build up a profile of other page views for the user.\r\n\r\nNote to Kaggle... I think it's pretty poor that this question has been on the forum for five days now and no response!  -1 from me.",
    "138696": "Terrible Tadpole, it is the same thing I am discovering so far, but in my preliminary data exploration, it seems users accessed just one document, so page views dataset is useless to build this kind of profile about user. The only useful info I get from the dataset so far is document access count and distribution of that access across geo location.",
    "138700": "TerribleTadpole your interpretation is correct. events.csv represents ad displays (in clicks_test.csv and clicks_train.csv) and there is always a click. page_views.csv is just users viewing documents.",
    "138751": "Think of an EVENT as a distinct interaction of the reader with the main document page. Every time you visit the page, an \"event\" is generated. \r\n\r\nFor instance, this is the main document of interest (to you, as a reader):\r\n\r\nhttp://www.cnn.com/2016/10/10/politics/paul-ryan-said-he-wont-defend-donald-trump/index.html\r\n\r\nEverytime you visit this page, you will see Outbrain's actual paid content section recommending different \"SET OF ADS\" to you. Try it out and see in action. \r\n\r\nThe page view creates the following info:\r\n\r\n        - Who (user)\r\n        - What (document)\r\n        - When (day and time)\r\n        - Where (geo)\r\n        - Platform (desktop, tablet, mobile)\r\n        - Source (user source)\r\n\r\nNotice it does not have anything to do with Ads (paid/promoted content). \r\n\r\nEvent however captures that info//",
    "138827": "If we go from clicks_test/train  -> events to get a list of display_id, uuid, ad_id.  This gives you a list of which user saw the ads in a particular display_id.\r\n\r\nNow suppose you take this and join to promoted_content to get display_id, uuid, ad_id, promoted_content.document_id.   This gives you a list of the potential documents the user could have routed to after clicking the ad.\r\n\r\nIf the uuid clicked on an ad_id which was promoting a particular document_id, shouldn't page_views have a record for that uuid visiting the promoted document_id?  I am finding this to not be the case.\r\n\r\nAm I misunderstanding the relationships here?",
    "138853": "Okay... in the absence of anything authoritative (still) from Kaggle or Outbrain, here's my take:\r\n\r\nclicks_test gives you all the ads shown for a given event (display_id).\r\nSo select * from clicks_test where display_id = 16874594; gives you 6 rows, which means six ads were shown on that occasion.\r\n\r\nThe events data will tell you what page those ads were embedded in, and the other attributes of how the page was accessed and by whom:\r\nselect * from events where display_id = 16874594; will give you one row with the uuid and the document_id\r\n\r\nYou can use that document_id to find more information about the page from the documents_... tables (meta, entities, topics, categories).  Note that we're still talking about the page in which the ad was embedded.\r\n\r\nTo look at the page that the ad links to, you need the promoted_content data:\r\nselect * from promoted_content where ad_id = 162754; will return 1 row, from which I infer that an ad always links to the same web page.  However, if you query you will find that many ads may link to the same web page:\r\nselect * from promoted_content where document_id = 1292723; will return 38 ads, all distinct, reinforcing my inference that an ad always links to the same page, but that page may be the target of many ads.\r\n\r\nThe page_views data appears to be intended as a cloud of other, unrelated, non-promoted content that the user is believed to have visited.  I believe that the essence of this competition is to analyse this cloud to ascertain attributes of the user's viewing behaviour that are predictive of their ad-clicking behaviour.  Putting the views of promoted material would give our algorithms data that a real-life application would not have; namely that in real life the user would not yet have clicked an ad. \r\n\r\nAll this is just my opinion of course... not gospel!",
    "139404": "Terrible Tadpole is up to something. Good comments!",
    "139465": "Bryan\r\n\r\nI believe that the ads can lead to pages that are outside of Outbrain control, think pages like your personal blog for example. So those clicks will not be present in page_views.\r\n\r\nThe promoted_content give us info about the ad and on which docs the ad was visible, but there is no indication about the \"routing\" after the ad is clicked.\r\n\r\n\r\n[quote=Bryan Johnson;138827]\r\n\r\nIf we go from clicks_test/train  -> events to get a list of display_id, uuid, ad_id.  This gives you a list of which user saw the ads in a particular display_id.\r\n\r\nNow suppose you take this and join to promoted_content to get display_id, uuid, ad_id, promoted_content.document_id.   This gives you a list of the potential documents the user could have routed to after clicking the ad.\r\n\r\nIf the uuid clicked on an ad_id which was promoting a particular document_id, shouldn't page_views have a record for that uuid visiting the promoted document_id?  I am finding this to not be the case.\r\n\r\nAm I misunderstanding the relationships here?\r\n\r\n[/quote]",
    "139790": "The fact that distinct display_ids from clicks_train and clicks_test are exactly the display_ids of events is a very strong evidence that events are just attributes of clicks ([Script1][1]).\r\n\r\nTherefore, I interpret page_views as logs of when visitors visit documents, while both clicks left joined to events represent logs of visitors clicking on ads displayed in the documents, and views include clicks/events ([Script2][2]).\r\n\r\nToo bad we do not get to see which ads were displayed on those visits when visitors did NOT click (this is a hint to Outbrain ;) ).\r\n\r\n\r\n  [1]: https://www.kaggle.com/andris/outbrain-click-prediction/schema-inference-1\r\n  [2]: https://www.kaggle.com/andris/outbrain-click-prediction/schema-inference-2",
    "139989": "I've just uploaded the [\"Unveiling page_views.csv with PySpark\" notebook](https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/) with full analytics of page_views.csv and its relationship with events.csv. \r\n\r\nAfter some hours of processing in a Spark Cluster, we've got answer to questions like:  \r\n\r\n- **How to join page_views.csv and events.csv?**\r\n- **Is events.csv a subset of page_views.csv?**\r\n- **Are there additional page views for users in events.csv?**\r\n\r\n\r\nEnjoy, and please upvote if you like ;)"
  },
  "source": "meta"
}