{
  "id": 6310,
  "title": "Clicks at the same unit of time",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6310",
  "author_name": "",
  "post_date": "2013-11-14T12:22:51.607Z",
  "votes": null,
  "comment_count": 5,
  "views": 1378,
  "content": "<p>In&nbsp;https://www.kaggle.com/c/yandex-personalized-web-search-challenge/data how should the following sentence be interpreted :</p>\n<p>&quot;From this set of queries we filter out all queries<br>with clicks performed at the same unit of time&quot;</p>\n<p>&nbsp;</p>\n<p>Does that mean two clicks done with the same unit of time, or does that mean a click performed at the same unit of time as the query?&nbsp;</p>\n<p>In both case, when does that happen?</p>",
  "messages": [
    {
      "id": "33698",
      "postDate": "11/14/2013 12:22:51",
      "content": "<p>In&nbsp;https://www.kaggle.com/c/yandex-personalized-web-search-challenge/data how should the following sentence be interpreted :</p>\n<p>&quot;From this set of queries we filter out all queries<br>with clicks performed at the same unit of time&quot;</p>\n<p>&nbsp;</p>\n<p>Does that mean two clicks done with the same unit of time, or does that mean a click performed at the same unit of time as the query?&nbsp;</p>\n<p>In both case, when does that happen?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33709",
      "postDate": "11/14/2013 16:09:37",
      "content": "<p>Hi Paul,</p>\n<p>Yes, it is not very clear. Suppose that have some internal &quot;unit&quot; of time that they use (just a guess, say 10 seconds).&nbsp; If the user clicked twice within that 10 second interval, it gets filtered out.&nbsp; (or possibly the click and the query were in the same time window).&nbsp;</p>\n<p>But does it matter?&nbsp; I understood that entire page to be processing that they did to create the train/test sets. and I don't have any code to address any of these issues myself (with the exception of relevance).&nbsp; Of course that may change in the future.&nbsp; </p>\n<p>In other contests, kagglers have benefited from data leaks inherent in the data due to the described pre-processing.&nbsp;&nbsp; Is this what you are looking for?</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33710",
      "postDate": "11/14/2013 16:25:58",
      "content": "<p>Thanks for the answer.</p>\n<p>I'm trying to reproduce a test set as close as possible to their own test set.<br>To do so, I'd like to be as accurate as possible in the selection of the test queries.</p>\n<p>I'm not exactly sure you answered my question about this very specific point, but hopefully, somebody from Yandex will come here and enlight me. :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33716",
      "postDate": "11/14/2013 18:55:05",
      "content": "<p>Here <a href=\"http://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6181/difficulties-understand-data-format/32968\">http://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6181/difficulties-understand-data-format/32968</a> Yandex gives clues about the &quot;spirit&quot; of removing &quot;clicks performed at the same unit of time&quot;:</p>\n<p>&nbsp;</p>\n<p>This is to be sure that &quot;dwell<em> time is uniquely defined (this can be problematic in case of two clicks at the same time unit)&quot;</em></p>\n<p>&nbsp;</p>\n<p>Not sure what it exactly means, but they speak of &quot;two clicks done with the same unit of time&quot; and not explicitely of &quot;click performed at the same unit of time as the query&quot;.</p>\n<p>I agree with Paul, precisions could be useful</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33744",
      "postDate": "11/15/2013 07:54:59",
      "content": "<p>Sorry for the ambiguity.</p>\n<p>What we wanted to say is that in some (very rare) sessions there are two clicks performed within the same unit of time. The corresponding queries are excluded from the list of test query candidates to simplify calculating the results' relevance values. Otherwise, the relevance labels will be dependent on the order of the lines in the file.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33774",
      "postDate": "11/15/2013 21:08:05",
      "content": "<p>Thank you very for your answer !</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 33709,
      "author_name": "geringer",
      "author_url": "",
      "post_date": "11/14/2013 16:09:37",
      "content": "<p>Hi Paul,</p>\n<p>Yes, it is not very clear. Suppose that have some internal &quot;unit&quot; of time that they use (just a guess, say 10 seconds).&nbsp; If the user clicked twice within that 10 second interval, it gets filtered out.&nbsp; (or possibly the click and the query were in the same time window).&nbsp;</p>\n<p>But does it matter?&nbsp; I understood that entire page to be processing that they did to create the train/test sets. and I don't have any code to address any of these issues myself (with the exception of relevance).&nbsp; Of course that may change in the future.&nbsp; </p>\n<p>In other contests, kagglers have benefited from data leaks inherent in the data due to the described pre-processing.&nbsp;&nbsp; Is this what you are looking for?</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33710,
      "author_name": "",
      "author_url": "",
      "post_date": "11/14/2013 16:25:58",
      "content": "<p>Thanks for the answer.</p>\n<p>I'm trying to reproduce a test set as close as possible to their own test set.<br>To do so, I'd like to be as accurate as possible in the selection of the test queries.</p>\n<p>I'm not exactly sure you answered my question about this very specific point, but hopefully, somebody from Yandex will come here and enlight me. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33716,
      "author_name": "cbourguignat",
      "author_url": "",
      "post_date": "11/14/2013 18:55:05",
      "content": "<p>Here <a href=\"http://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6181/difficulties-understand-data-format/32968\">http://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6181/difficulties-understand-data-format/32968</a> Yandex gives clues about the &quot;spirit&quot; of removing &quot;clicks performed at the same unit of time&quot;:</p>\n<p>&nbsp;</p>\n<p>This is to be sure that &quot;dwell<em> time is uniquely defined (this can be problematic in case of two clicks at the same time unit)&quot;</em></p>\n<p>&nbsp;</p>\n<p>Not sure what it exactly means, but they speak of &quot;two clicks done with the same unit of time&quot; and not explicitely of &quot;click performed at the same unit of time as the query&quot;.</p>\n<p>I agree with Paul, precisions could be useful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33744,
      "author_name": "eugene1751",
      "author_url": "",
      "post_date": "11/15/2013 07:54:59",
      "content": "<p>Sorry for the ambiguity.</p>\n<p>What we wanted to say is that in some (very rare) sessions there are two clicks performed within the same unit of time. The corresponding queries are excluded from the list of test query candidates to simplify calculating the results' relevance values. Otherwise, the relevance labels will be dependent on the order of the lines in the file.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33774,
      "author_name": "",
      "author_url": "",
      "post_date": "11/15/2013 21:08:05",
      "content": "<p>Thank you very for your answer !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "33698": "",
    "33709": "",
    "33710": "",
    "33716": "",
    "33744": "",
    "33774": ""
  },
  "source": "meta"
}