{
  "id": 6181,
  "title": "Difficulties understand data format",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6181",
  "author_name": "",
  "post_date": "2013-10-31T19:52:50.677Z",
  "votes": null,
  "comment_count": 5,
  "views": 3462,
  "content": "<p>Hi,</p>\n<p>I have trouble understanding the following&nbsp; statement (excerpt from <a href=\"http://www.kaggle.com/c/yandex-personalized-web-search-challenge/data\" target=\"_blank\">this page</a>):</p>\n<p><em>For each user from </em><em>the test period, we take all her queries from the test period&nbsp;with at </em><em>least one click with the dwell time not less than 50 time units (so, the </em><em>clicked document is relevant or highly relevant according to our </em><em>definition of personal relevance, see Evaluation).</em><br><em> From this set of queries we filter out all queries with clicks&nbsp; </em><em>performed at the same unit of time. Finally, from the resulting set of </em><em>queries we uniformly sample only one query and consider it to be a test&nbsp; </em><em>query.&nbsp;</em><br><em>If the sampled </em><em>query does not have any short-term context (it is the first one in the </em><em>session) and the user that asked this query has no search sessions in </em><em>he training period, we remove this query from the test set (since, it </em><em>has neither short nor long-term context useful for personalization). </em></p>\n<p><em>We </em><em>do not disclose any user actions made after the test query. However, the </em><em>user's actions performed in the same session before the test query are </em><em>provided (if any).</em></p>\n<p>Could you please enlighten me about this?</p>\n<p>Thank you so much!</p>\n<p>X</p>",
  "messages": [
    {
      "id": "32968",
      "postDate": "10/31/2013 19:52:50",
      "content": "<p>Hi,</p>\n<p>I have trouble understanding the following&nbsp; statement (excerpt from <a href=\"http://www.kaggle.com/c/yandex-personalized-web-search-challenge/data\" target=\"_blank\">this page</a>):</p>\n<p><em>For each user from </em><em>the test period, we take all her queries from the test period&nbsp;with at </em><em>least one click with the dwell time not less than 50 time units (so, the </em><em>clicked document is relevant or highly relevant according to our </em><em>definition of personal relevance, see Evaluation).</em><br><em> From this set of queries we filter out all queries with clicks&nbsp; </em><em>performed at the same unit of time. Finally, from the resulting set of </em><em>queries we uniformly sample only one query and consider it to be a test&nbsp; </em><em>query.&nbsp;</em><br><em>If the sampled </em><em>query does not have any short-term context (it is the first one in the </em><em>session) and the user that asked this query has no search sessions in </em><em>he training period, we remove this query from the test set (since, it </em><em>has neither short nor long-term context useful for personalization). </em></p>\n<p><em>We </em><em>do not disclose any user actions made after the test query. However, the </em><em>user's actions performed in the same session before the test query are </em><em>provided (if any).</em></p>\n<p>Could you please enlighten me about this?</p>\n<p>Thank you so much!</p>\n<p>X</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33190",
      "postDate": "11/04/2013 17:07:18",
      "content": "<p>Hello,</p>\n<p>Sorry for the late reply.</p>\n<p>This text describes how the test set is generated. This algorithm can be useful for the participants who want to build their own validation/evaluation sets from the training data. Informally, the described algorithm achieves the following requirements for each test query:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">at least one non-zero label is available;</span></li>\n<li><span style=\"line-height: 1.4\">dwell time is uniquely defined (this can be problematic in case of two clicks at the same time unit);</span></li>\n<li><span style=\"line-height: 1.4\">one test query per user;</span></li>\n<li><span style=\"line-height: 1.4\">the test query can sampled from any part of the session (e.g., it can be the first, the second, ..., the last);</span></li>\n<li><span style=\"line-height: 1.4\">we want to test the personalization algorithms only on the queries with some context available;</span></li>\n<li><span style=\"line-height: 1.4\">no information about any events occurred after the test query is provided (as it is unrealistic in real-life scenario).</span></li>\n</ol>\n<p>Hope this helps.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33502",
      "postDate": "11/10/2013 09:22:16",
      "content": "<p>I still don't understand that &quot;If the sampled query does not have any short-term context (it is the first one in the session)...&quot;</p>\n<p>What does means &quot;sort-term context&quot;? Is it the first query in a session?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33544",
      "postDate": "11/11/2013 07:01:03",
      "content": "<p>We adopt the following terminology.&nbsp;The test query's&nbsp;<em>short-term context&nbsp;</em>includes all actions performed in the same session, but before the test query itself.&nbsp;All sessions of the same user that took place before the session with the test query are called the&nbsp;<em>long-term contex</em>t.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37624",
      "postDate": "01/20/2014 05:23:24",
      "content": "<p><br>To,<br>Guocong,<br>I am student of computer science. I want to learn how to apply machine learning techniques to<br>solve practical problems. I think Kaggle is great platform for it.<br>I am complete newbie to this field. Can you please suggest courses and other references to start<br>with.<br>Regards,<br>Adwait</p>\n<p>p.s. sorry for unrelated post to this thread, i didnt find any other way to contact specific kaggler</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "37646",
      "postDate": "01/20/2014 17:30:23",
      "content": "<p>If you click username on the profile there is a Contact tab where you can email a kaggler. Also, under each post there is the 'Email user' link.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 33190,
      "author_name": "eugene1751",
      "author_url": "",
      "post_date": "11/04/2013 17:07:18",
      "content": "<p>Hello,</p>\n<p>Sorry for the late reply.</p>\n<p>This text describes how the test set is generated. This algorithm can be useful for the participants who want to build their own validation/evaluation sets from the training data. Informally, the described algorithm achieves the following requirements for each test query:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">at least one non-zero label is available;</span></li>\n<li><span style=\"line-height: 1.4\">dwell time is uniquely defined (this can be problematic in case of two clicks at the same time unit);</span></li>\n<li><span style=\"line-height: 1.4\">one test query per user;</span></li>\n<li><span style=\"line-height: 1.4\">the test query can sampled from any part of the session (e.g., it can be the first, the second, ..., the last);</span></li>\n<li><span style=\"line-height: 1.4\">we want to test the personalization algorithms only on the queries with some context available;</span></li>\n<li><span style=\"line-height: 1.4\">no information about any events occurred after the test query is provided (as it is unrealistic in real-life scenario).</span></li>\n</ol>\n<p>Hope this helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33502,
      "author_name": "songgc",
      "author_url": "",
      "post_date": "11/10/2013 09:22:16",
      "content": "<p>I still don't understand that &quot;If the sampled query does not have any short-term context (it is the first one in the session)...&quot;</p>\n<p>What does means &quot;sort-term context&quot;? Is it the first query in a session?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33544,
      "author_name": "eugene1751",
      "author_url": "",
      "post_date": "11/11/2013 07:01:03",
      "content": "<p>We adopt the following terminology.&nbsp;The test query's&nbsp;<em>short-term context&nbsp;</em>includes all actions performed in the same session, but before the test query itself.&nbsp;All sessions of the same user that took place before the session with the test query are called the&nbsp;<em>long-term contex</em>t.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37624,
      "author_name": "adwaitpathak",
      "author_url": "",
      "post_date": "01/20/2014 05:23:24",
      "content": "<p><br>To,<br>Guocong,<br>I am student of computer science. I want to learn how to apply machine learning techniques to<br>solve practical problems. I think Kaggle is great platform for it.<br>I am complete newbie to this field. Can you please suggest courses and other references to start<br>with.<br>Regards,<br>Adwait</p>\n<p>p.s. sorry for unrelated post to this thread, i didnt find any other way to contact specific kaggler</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 37646,
      "author_name": "denissavenkov",
      "author_url": "",
      "post_date": "01/20/2014 17:30:23",
      "content": "<p>If you click username on the profile there is a Contact tab where you can email a kaggler. Also, under each post there is the 'Email user' link.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "32968": "",
    "33190": "",
    "33502": "",
    "33544": "",
    "37624": "",
    "37646": ""
  },
  "source": "meta"
}