{
  "id": 6257,
  "title": "Duplicate TermIds in a single query",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6257",
  "author_name": "",
  "post_date": "2013-11-08T03:54:46.063Z",
  "votes": null,
  "comment_count": 2,
  "views": 806,
  "content": "<p>There are cases where, within a single query, the ListOfTerms contains the same TermID more than once (e.g., training data, session 64, serpid 0, TermID 2926659 appears twice).</p>\n<p>This doesn't really make sense if we think of a query as a list of keywords (e.g., I wouldn't usually type &quot;black cat cat&quot; into Google, I would just type &quot;black cat&quot;).</p>\n<p>On the other hand, if we think of a query as a natural language string (e.g., &quot;a black cat on a hot tin roof&quot;) this makes sense (the letter 'a' shows up twice).</p>\n<p>Can someone clarify this? Or are we not supposed to know?</p>\n<p>&nbsp;</p>\n<p>Thank you.</p>\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "33398",
      "postDate": "11/08/2013 03:54:46",
      "content": "<p>There are cases where, within a single query, the ListOfTerms contains the same TermID more than once (e.g., training data, session 64, serpid 0, TermID 2926659 appears twice).</p>\n<p>This doesn't really make sense if we think of a query as a list of keywords (e.g., I wouldn't usually type &quot;black cat cat&quot; into Google, I would just type &quot;black cat&quot;).</p>\n<p>On the other hand, if we think of a query as a natural language string (e.g., &quot;a black cat on a hot tin roof&quot;) this makes sense (the letter 'a' shows up twice).</p>\n<p>Can someone clarify this? Or are we not supposed to know?</p>\n<p>&nbsp;</p>\n<p>Thank you.</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33401",
      "postDate": "11/08/2013 05:21:41",
      "content": "<p><span style=\"line-height: 1.4\">Yes, usually users wouldn't type 'black cat cat'. However, &nbsp;queries are not just lists of keywords: they might contain prepositions, articles*, special and&nbsp;</span><span style=\"line-height: 1.4\">punctuation symbols. Apart from that,&nbsp;we can expect various types of copy-pasted pieces of text and typos.</span></p>\n<p>We didn't apply any normalization on our side (duplicate removal, sorting of termids, etc) so that the participants can find their way of working with this information which is the most suitable for them.</p>\n<p>&nbsp;</p>\n<p>*(actually, there are no articles in Russian, but some queries can be non-Russian)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33404",
      "postDate": "11/08/2013 06:39:17",
      "content": "<p>That helps a lot. Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 33401,
      "author_name": "eugene1751",
      "author_url": "",
      "post_date": "11/08/2013 05:21:41",
      "content": "<p><span style=\"line-height: 1.4\">Yes, usually users wouldn't type 'black cat cat'. However, &nbsp;queries are not just lists of keywords: they might contain prepositions, articles*, special and&nbsp;</span><span style=\"line-height: 1.4\">punctuation symbols. Apart from that,&nbsp;we can expect various types of copy-pasted pieces of text and typos.</span></p>\n<p>We didn't apply any normalization on our side (duplicate removal, sorting of termids, etc) so that the participants can find their way of working with this information which is the most suitable for them.</p>\n<p>&nbsp;</p>\n<p>*(actually, there are no articles in Russian, but some queries can be non-Russian)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33404,
      "author_name": "fellows",
      "author_url": "",
      "post_date": "11/08/2013 06:39:17",
      "content": "<p>That helps a lot. Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "33398": "",
    "33401": "",
    "33404": ""
  },
  "source": "meta"
}