{
  "id": 1656,
  "title": "Data aggregation and primary key",
  "url": "/competitions/kddcup2012-track2/discussion/1656",
  "author_name": "",
  "post_date": "2012-03-30T16:35:42.363Z",
  "votes": null,
  "comment_count": 2,
  "views": 2549,
  "content": "<p>Hi,</p>\r\n<p>As I understand from:</p>\r\n<blockquote>\r\n<p><em><span style=\"font-family:'arial black','avant garde'\">We divide each session into multiple instances, where each instance describes an impressed ad under a certain setting&nbsp; (i.e., with certain depth and position values). &nbsp;We aggregate instances with\r\n the same user id, ad id, query, and setting in order to reduce the dataset size</span>.</em></p>\r\n</blockquote>\r\n<p><em></em>in &quot;training.txt&quot; userid&#43;adid&#43;queryid&#43;depth&#43;position is a primary key. But that is not true, aproximately 1% of the instances are duplicated in that sense. For example, the following one is duplicated</p>\r\n<pre>$ cat training.txt | grep 21258213 | grep 356694<br>0\t1\t1658343530815135762\t21258213\t2298\t3\t2\t7246\t4307\t7183\t16230\t356694<br>0\t1\t1658343530815135762\t21258213\t2298\t3\t3\t7246\t4307\t7183\t16137\t356694<br>0\t1\t1658343530815135762\t21258213\t2298\t3\t3\t7246\t4307\t7183\t16230\t356694</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>It's a mistake or I am misunderstanding the problem statement?</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>Thank you.</pre>\r\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "9857",
      "postDate": "03/30/2012 16:35:42",
      "content": "<p>Hi,</p>\r\n<p>As I understand from:</p>\r\n<blockquote>\r\n<p><em><span style=\"font-family:'arial black','avant garde'\">We divide each session into multiple instances, where each instance describes an impressed ad under a certain setting&nbsp; (i.e., with certain depth and position values). &nbsp;We aggregate instances with\r\n the same user id, ad id, query, and setting in order to reduce the dataset size</span>.</em></p>\r\n</blockquote>\r\n<p><em></em>in &quot;training.txt&quot; userid&#43;adid&#43;queryid&#43;depth&#43;position is a primary key. But that is not true, aproximately 1% of the instances are duplicated in that sense. For example, the following one is duplicated</p>\r\n<pre>$ cat training.txt | grep 21258213 | grep 356694<br>0\t1\t1658343530815135762\t21258213\t2298\t3\t2\t7246\t4307\t7183\t16230\t356694<br>0\t1\t1658343530815135762\t21258213\t2298\t3\t3\t7246\t4307\t7183\t16137\t356694<br>0\t1\t1658343530815135762\t21258213\t2298\t3\t3\t7246\t4307\t7183\t16230\t356694</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>It's a mistake or I am misunderstanding the problem statement?</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>Thank you.</pre>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "9873",
      "postDate": "03/31/2012 06:14:49",
      "content": "<p>Hi, LARCA,</p>\r\n<p>Thank you for the question. My colleague, Liubin Wang, helped checking the data. We think there is no mistake in the data set.\r\n</p>\r\n<p>The 1st and the 2nd instances cannot be aggregated due to different position (on the 7th field).<br>\r\nThe 1st and the 3rd instances cannot be aggregated due to different position.<br>\r\nThe 2nd and the 3rd cannot be aggregated due to different description id (on the 11th field).</p>\r\n<p>As you quoted:</p>\r\n<p>We aggregate instances with the same user id, ad id, query, and setting in order to reduce the dataset size.</p>\r\n<p>The &quot;position&quot; and &quot;description id&quot; are parts of &quot;setting&quot;.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "9880",
      "postDate": "03/31/2012 18:49:16",
      "content": "<p>Thnk you Yi Wang for the reply. </p>\r\n<p>So what makes unique each instance is userid&#43;adid&#43;queryid&#43;depth&#43;position&#43;descriptionid. Any other feature to consider as part of &quot;setting&quot;?</p>\r\n<p>Thanks again!</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 9873,
      "author_name": "yiwang0",
      "author_url": "",
      "post_date": "03/31/2012 06:14:49",
      "content": "<p>Hi, LARCA,</p>\r\n<p>Thank you for the question. My colleague, Liubin Wang, helped checking the data. We think there is no mistake in the data set.\r\n</p>\r\n<p>The 1st and the 2nd instances cannot be aggregated due to different position (on the 7th field).<br>\r\nThe 1st and the 3rd instances cannot be aggregated due to different position.<br>\r\nThe 2nd and the 3rd cannot be aggregated due to different description id (on the 11th field).</p>\r\n<p>As you quoted:</p>\r\n<p>We aggregate instances with the same user id, ad id, query, and setting in order to reduce the dataset size.</p>\r\n<p>The &quot;position&quot; and &quot;description id&quot; are parts of &quot;setting&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 9880,
      "author_name": "mingot",
      "author_url": "",
      "post_date": "03/31/2012 18:49:16",
      "content": "<p>Thnk you Yi Wang for the reply. </p>\r\n<p>So what makes unique each instance is userid&#43;adid&#43;queryid&#43;depth&#43;position&#43;descriptionid. Any other feature to consider as part of &quot;setting&quot;?</p>\r\n<p>Thanks again!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "9857": "",
    "9873": "",
    "9880": ""
  },
  "source": "meta"
}