{
  "id": 6725,
  "title": "Understanding Train Dataset",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6725",
  "author_name": "",
  "post_date": "2013-12-31T09:06:22.073Z",
  "votes": null,
  "comment_count": 2,
  "views": 1818,
  "content": "<p>Hi</p>\n<p>&nbsp;When I open the train data set,I really don't understand what are these numbers and digits.&nbsp;please help me.</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "36862",
      "postDate": "12/31/2013 09:06:22",
      "content": "<p>Hi</p>\n<p>&nbsp;When I open the train data set,I really don't understand what are these numbers and digits.&nbsp;please help me.</p>\n<p>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36910",
      "postDate": "01/02/2014 10:55:33",
      "content": "<p>It's not easy indeed.</p>\n<p><span style=\"line-height: 1.4\">The log format description is described here.<br></span><span style=\"line-height: 1.4\">http://www.kaggle.com/c/yandex-personalized-web-search-challenge/details/logs-format</span></p>\n<p>You probably noticed that lines have different length.<br>As explained in the above link, there are different kind of lines, namingly</p>\n<p>SESSION META (Marking the start of a new session),&nbsp;QUERY (A&nbsp;query done by the user), &nbsp;CLICKS, and finally &nbsp;TEST QUERIES (Which are the one you need to re-rank</p>\n<p>Many people (including us), have supplied python scripts to help you parse the log format.<br>You could save time by checking them out.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36929",
      "postDate": "01/02/2014 22:58:07",
      "content": "<p>This is the definitive forum thread on parsing that Paul refers to&nbsp;https://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6489/python-code-for-parsing-data</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 36910,
      "author_name": "",
      "author_url": "",
      "post_date": "01/02/2014 10:55:33",
      "content": "<p>It's not easy indeed.</p>\n<p><span style=\"line-height: 1.4\">The log format description is described here.<br></span><span style=\"line-height: 1.4\">http://www.kaggle.com/c/yandex-personalized-web-search-challenge/details/logs-format</span></p>\n<p>You probably noticed that lines have different length.<br>As explained in the above link, there are different kind of lines, namingly</p>\n<p>SESSION META (Marking the start of a new session),&nbsp;QUERY (A&nbsp;query done by the user), &nbsp;CLICKS, and finally &nbsp;TEST QUERIES (Which are the one you need to re-rank</p>\n<p>Many people (including us), have supplied python scripts to help you parse the log format.<br>You could save time by checking them out.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36929,
      "author_name": "jzchew",
      "author_url": "",
      "post_date": "01/02/2014 22:58:07",
      "content": "<p>This is the definitive forum thread on parsing that Paul refers to&nbsp;https://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6489/python-code-for-parsing-data</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "36862": "",
    "36910": "",
    "36929": ""
  },
  "source": "meta"
}