{
  "id": 6254,
  "title": "Data issues",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6254",
  "author_name": "",
  "post_date": "2013-11-07T15:27:39.280Z",
  "votes": null,
  "comment_count": 19,
  "views": 3643,
  "content": "<p>While checking the data consistency, I've found some uncertainty.</p>\n<p>Here is the beginning of &quot;train&quot;:</p>\n<p>0&nbsp;&nbsp; &nbsp;M&nbsp;&nbsp; &nbsp;4&nbsp;&nbsp; &nbsp;0&nbsp;</p>\n<p>0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;Q&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;10047345&nbsp;&nbsp;&nbsp; ... &nbsp;&nbsp; 68893581,5149883&nbsp;&nbsp;&nbsp; ...</p>\n<p>0&nbsp;&nbsp; &nbsp;108&nbsp;&nbsp; &nbsp;C&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;50628761</p>\n<p>0&nbsp;&nbsp; &nbsp;1080&nbsp;&nbsp; &nbsp;C&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;50628761</p>\n<p>We see two different click on the same url.</p>\n<p>Is this the last click?</p>\n<p>Suppose we have one more click in this session:</p>\n<p>0&nbsp;&nbsp;&nbsp; 2000&nbsp;&nbsp; C&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp;&nbsp; 68893581</p>\n<p>What is the dwell time: 1080-108, or 2000-1080, or 2000-108 ?</p>\n<p>We need to reproduce exactly your algorithm of relevance calculation, so how should we operate with repeated clicking?</p>",
  "messages": [
    {
      "id": "33375",
      "postDate": "11/07/2013 15:27:39",
      "content": "<p>While checking the data consistency, I've found some uncertainty.</p>\n<p>Here is the beginning of &quot;train&quot;:</p>\n<p>0&nbsp;&nbsp; &nbsp;M&nbsp;&nbsp; &nbsp;4&nbsp;&nbsp; &nbsp;0&nbsp;</p>\n<p>0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;Q&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;10047345&nbsp;&nbsp;&nbsp; ... &nbsp;&nbsp; 68893581,5149883&nbsp;&nbsp;&nbsp; ...</p>\n<p>0&nbsp;&nbsp; &nbsp;108&nbsp;&nbsp; &nbsp;C&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;50628761</p>\n<p>0&nbsp;&nbsp; &nbsp;1080&nbsp;&nbsp; &nbsp;C&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;50628761</p>\n<p>We see two different click on the same url.</p>\n<p>Is this the last click?</p>\n<p>Suppose we have one more click in this session:</p>\n<p>0&nbsp;&nbsp;&nbsp; 2000&nbsp;&nbsp; C&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp;&nbsp; 68893581</p>\n<p>What is the dwell time: 1080-108, or 2000-1080, or 2000-108 ?</p>\n<p>We need to reproduce exactly your algorithm of relevance calculation, so how should we operate with repeated clicking?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33379",
      "postDate": "11/07/2013 16:07:18",
      "content": "<p><span style=\"line-height: 1.4\">The last click is the click with the latest timestamp in the session, so the second &nbsp;click on&nbsp;50628761 in train.gz is the latest in the session with its SessionId equal to 0.</span></p>\n<p>&nbsp;</p>\n<p>In your example, two clicks have the following dwell times: 1080 - 108 and 2000 - 1080. When labelling a URLID's relevance, we consider all the clicks associated with the URLID and use the maximum of the dwell times (1080 - 108 &gt; 2000 - 1080 in your example).</p>\n<p>You can find an additional discussion&nbsp;<a href=\"https://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6170/definition-of-relevance\">in another topic</a>&nbsp;.</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33437",
      "postDate": "11/08/2013 18:31:18",
      "content": "<p>file</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33439",
      "postDate": "11/08/2013 18:42:34",
      "content": "<p>The file above contains the relevances for the first 1000 queries.</p>\n<p>Calculation of relevances for the training sample is purely technical task, but one may easily make mistake on it. So I think it is a good idea to share such results for comparison and checking.</p>\n<p>It would be nice to have such file in competition data...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33505",
      "postDate": "11/10/2013 13:16:22",
      "content": "<p>my first diff:</p>\n<p>6th line (session 4, serp 0), your result 0000000002</p>\n<p>data: 1st dwell time 86-27&gt;=50, 3rd dwell time 216-121&gt;=50, so i guess it should be 0001010002</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33507",
      "postDate": "11/10/2013 14:52:23",
      "content": "<p>Thank you very much. I've found a mistake in my program.</p>\n<p>Here is the corrected file.</p>\n<p>I could post the full file, but I am not sure that this is allowed by the rules.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34867",
      "postDate": "11/18/2013 22:24:24",
      "content": "<p>I don't think you're interpreting the last click in a SERP properly, you're treating it as &quot;highly relevant,&quot; even when there's another SERP in the same session.</p>\n<p>For example, session 13 SERP 2 has a query then a single click at time 5658.&nbsp; In the same session, the user re-submits their search at time 5753, so the dwell time on that click should be 5753 - 5658 = 95, a &quot;1&quot; grade.&nbsp; But you have it marked as a &quot;2&quot; grade.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34883",
      "postDate": "11/19/2013 08:16:13",
      "content": "<p>[quote=Martin C. Martin;34867]</p>\n<p>I don't think you're interpreting the last click in a SERP properly, you're treating it as &quot;highly relevant,&quot; even when there's another SERP in the same session.</p>\n<p>[/quote]</p>\n<p>Thank you for the comment.</p>\n<p>But here is citation from &quot;Evaluation&quot;: &quot;2 (highly relevant) grade corresponds to documents with clicks and dwell time not shorter than 400 time units and to the <br>document with the last click in the search session&quot;.</p>\n<p>So, the click in session 13 SERP 2 should be highly relevant because it is really the last click in the session, since there were only queries but no clicks later!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34893",
      "postDate": "11/19/2013 11:42:29",
      "content": "<p>Thanks.&nbsp; Eugene, can you confirm that the last click counts as &quot;highly relevant,&quot; even when the user comes back and does another search but doesn't click on anything?&nbsp; Doing another search suggests they didn't find what they wanted with their click, and are looking for something else.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34915",
      "postDate": "11/19/2013 17:05:29",
      "content": "<p>Sorry for the ambiguity in the description. We'll fix it asap.</p>\n<p>The relevance level of 2 is assigned to the clicks which have dwell time not smaller than 400 time units and to the clicks that are <strong>the last actions in the corresponding sessions</strong>.</p>\n<p>In other words, if the dwell time can be calculated then the relevance is assigned according to it. If the click is the last action in the session (in other words, the next action is performed after a long time window separating two sessions) then the relevance is set to 2.</p>\n<p>We believe this definition is more consistent.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34940",
      "postDate": "11/19/2013 22:44:17",
      "content": "<p>Victor, any chance you can update your file?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34973",
      "postDate": "11/20/2013 19:38:05",
      "content": "<p>Hi</p>\n<p>I download the Data, unziped it , but I can not open Them. which software should I use to open them. files do not have&nbsp; any extension !!!</p>\n<p>Please help me.</p>\n<p>&nbsp;</p>\n<p>TNX</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34974",
      "postDate": "11/20/2013 21:15:59",
      "content": "<p>You can open such files with good Text Editors like Notepad++. (except the train file ;))</p>\n<p>But for further editings you need to get line for line I think. Your RAM is not that big the store the data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34984",
      "postDate": "11/21/2013 05:09:48",
      "content": "<p>[quote=Martin C. Martin;34940]</p>\n<p>Victor, any chance you can update your file?</p>\n<p>[/quote]</p>\n<p>Here it is.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34985",
      "postDate": "11/21/2013 05:14:05",
      "content": "<p>[quote=Ali Reza Honarvar;34973]</p>\n<p>I download the Data, unziped it , but I can not open Them. which software should I use to open them. files do not have&nbsp; any extension !!!</p>\n<p>[/quote]</p>\n<p>I use the F3-viewer in FAR-manager.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35093",
      "postDate": "11/23/2013 15:03:02",
      "content": "<p>[quote=Eugene;34915]</p>\n<p>Sorry for the ambiguity in the description. We'll fix it asap.</p>\n<p>[/quote]</p>\n<p>May be is it possible to provide a file with calculated relevances?</p>\n<p>There is still some uncertainty with the cases when some clicks occurs later than the next query. And with simultaneous events too.</p>\n<p>Here is my implementation of relevance (grade) calculation: &nbsp;&nbsp;&nbsp; </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35150",
      "postDate": "11/25/2013 08:55:34",
      "content": "<p>[quote=Ali Reza Honarvar;34973]</p>\n<p>Hi</p>\n<p>I download the Data, unziped it , but I can not open Them. which software should I use to open them. files do not have&nbsp; any extension !!!</p>\n<p>Please help me.</p>\n<p>[/quote]</p>\n<p><span style=\"line-height: 1.4\">I don't think you would want to open the complete train file at once. Under Linux, you can use 'more' to view the files. I advice you to develop your code on a small sample. You can create a sample without unzipping using the following command: 'gzip -cd train.gz | head -1000 &gt; train.sample'. Or on the unzipped file: 'cat train |&nbsp; head -1000 &gt; train.sample'</span></p>\n<p>If you are using java, then you can read the gz-files without unzipping them using the class&nbsp;GZIPInputStream.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35204",
      "postDate": "11/25/2013 18:12:04",
      "content": "<p>Hi</p>\n<p>in windows you can use 'Large Texte File Viewer' to open the file.</p>\n<p>also you can use 'HJSplit' to split train file in many parts.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35483",
      "postDate": "11/29/2013 08:42:05",
      "content": "<p>I found this strange piece of data:</p>\n<p>22 &nbsp; &nbsp; &nbsp;0 &nbsp; &nbsp; &nbsp; Q &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; 16185818 &nbsp; &nbsp; &nbsp; &nbsp;4078517,3143668,1988399 35767635,3364600 &nbsp; &nbsp; &nbsp; &nbsp;57421304,4536568 &nbsp; &nbsp; &nbsp; &nbsp;35296159,3323845 &nbsp; &nbsp; &nbsp; &nbsp;49235419,4140722 &nbsp; &nbsp; &nbsp; &nbsp;68222659,5116467 &nbsp; &nbsp; &nbsp; &nbsp;55142634,4429180 &nbsp; &nbsp; &nbsp; &nbsp;36577088,3407786 &nbsp; &nbsp; &nbsp; &nbsp;61943983,4767047 &nbsp; &nbsp; &nbsp; &nbsp;65157339,4954478 &nbsp; &nbsp; &nbsp; &nbsp;42867421,3782572</p>\n<p>22 &nbsp; &nbsp; &nbsp;25 &nbsp; &nbsp; &nbsp;Q &nbsp; &nbsp; &nbsp; 2 &nbsp; &nbsp; &nbsp; 16185820 &nbsp; &nbsp; &nbsp; &nbsp;4078517,3143668,1988399,2133749 30325446,2938326 &nbsp; &nbsp; &nbsp; &nbsp;28846733,2833948 &nbsp; &nbsp; &nbsp; &nbsp;35769151,3364600&nbsp; &nbsp; &nbsp; &nbsp; 57008295,4511361 &nbsp; &nbsp; &nbsp; &nbsp;28175892,2768660 &nbsp; &nbsp; &nbsp; &nbsp;28864909,2833991 &nbsp; &nbsp; &nbsp; &nbsp;60703868,4697323 &nbsp; &nbsp; &nbsp; &nbsp;22790363,2294429 &nbsp; &nbsp; &nbsp; &nbsp;57692687,4551821&nbsp; &nbsp; &nbsp; &nbsp; 6473264,800608</p>\n<p>22 &nbsp; &nbsp; &nbsp;25 &nbsp; &nbsp; &nbsp;Q &nbsp; &nbsp; &nbsp; 1 &nbsp; &nbsp; &nbsp; 16185818 &nbsp; &nbsp; &nbsp; &nbsp;4078517,3143668,1988399 35767635,3364600 &nbsp; &nbsp; &nbsp; &nbsp;57421304,4536568 &nbsp; &nbsp; &nbsp; &nbsp;35296159,3323845 &nbsp; &nbsp; &nbsp; &nbsp;49235419,4140722 &nbsp; &nbsp; &nbsp; &nbsp;68222659,5116467 &nbsp; &nbsp; &nbsp; &nbsp;55142634,4429180 &nbsp; &nbsp; &nbsp; &nbsp;36577088,3407786 &nbsp; &nbsp; &nbsp; &nbsp;61943983,4767047 &nbsp; &nbsp; &nbsp; &nbsp;65157339,4954478 &nbsp; &nbsp; &nbsp; &nbsp;42867421,3782572</p>\n<p>22 &nbsp; &nbsp; &nbsp;31 &nbsp; &nbsp; &nbsp;C &nbsp; &nbsp; &nbsp; 2 &nbsp; &nbsp; &nbsp; 30325446</p>\n<p>22 &nbsp; &nbsp; &nbsp;112 &nbsp; &nbsp; C &nbsp; &nbsp; &nbsp; 2 &nbsp; &nbsp; &nbsp; 28175892</p>\n<p>Notice two queries at the same time!&nbsp;</p>\n<p>This corresponds to strings #51-53 in your file :</p>\n<p>0000000000 &nbsp;(serpid = 0)</p>\n<p>1000100000 &nbsp;(serpid = 2)</p>\n<p>0000000000 &nbsp;(serpid = 1)</p>\n<p>But don't you think we should sort them by SERPID?</p>\n<p>In the first 1000 queries this trouble (with two queries in one unit of time written in reverse order of SERPIDs) appears one another time in session #458.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35486",
      "postDate": "11/29/2013 10:04:54",
      "content": "<p>...and also I found mistake in your file:</p>\n<p>309 &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; Q &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; 11466723 &nbsp; &nbsp; &nbsp; &nbsp;3332767,4291604 29370045,2863645 &nbsp; &nbsp; &nbsp; &nbsp;55303432,4433251 &nbsp; &nbsp; &nbsp; &nbsp;3276363,443265 &nbsp;29322278,2863497&nbsp; &nbsp; &nbsp; &nbsp; 29371141,2863645 &nbsp; &nbsp; &nbsp; &nbsp;36829352,3418807 &nbsp; &nbsp; &nbsp; &nbsp;48935893,4127107 &nbsp; &nbsp; &nbsp; &nbsp;27402385,2655130 &nbsp; &nbsp; &nbsp; &nbsp;28137920,2761328 &nbsp; &nbsp; &nbsp; &nbsp;62315277,4791822</p>\n<p>309 &nbsp; &nbsp; 13 &nbsp; &nbsp; &nbsp;Q &nbsp; &nbsp; &nbsp; 1 &nbsp; &nbsp; &nbsp; 11466723 &nbsp; &nbsp; &nbsp; &nbsp;3332767,4291604 29370045,2863645 &nbsp; &nbsp; &nbsp; &nbsp;55303432,4433251 &nbsp; &nbsp; &nbsp; &nbsp;3276363,443265 &nbsp;29322278,2863497&nbsp; &nbsp; &nbsp; &nbsp; 29371141,2863645 &nbsp; &nbsp; &nbsp; &nbsp;36829352,3418807 &nbsp; &nbsp; &nbsp; &nbsp;48935893,4127107 &nbsp; &nbsp; &nbsp; &nbsp;27402385,2655130 &nbsp; &nbsp; &nbsp; &nbsp;28137920,2761328 &nbsp; &nbsp; &nbsp; &nbsp;62315277,4791822</p>\n<p>309 &nbsp; &nbsp; 13 &nbsp; &nbsp; &nbsp;C &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; 29370045</p>\n<p>310 &nbsp; &nbsp; M &nbsp; &nbsp; &nbsp; 2 &nbsp; &nbsp; &nbsp; 52 &nbsp; &nbsp; &nbsp;</p>\n<p>Here should be &quot;2000000000&quot; instead of &quot;0000000000&quot; in your file (line 557) because the click is the last in this session.&nbsp;I suppose, this may be associated with mixed SERPIDs.</p>\n<p>Excluding this two things I have the same results for relevances.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 33379,
      "author_name": "eugene1751",
      "author_url": "",
      "post_date": "11/07/2013 16:07:18",
      "content": "<p><span style=\"line-height: 1.4\">The last click is the click with the latest timestamp in the session, so the second &nbsp;click on&nbsp;50628761 in train.gz is the latest in the session with its SessionId equal to 0.</span></p>\n<p>&nbsp;</p>\n<p>In your example, two clicks have the following dwell times: 1080 - 108 and 2000 - 1080. When labelling a URLID's relevance, we consider all the clicks associated with the URLID and use the maximum of the dwell times (1080 - 108 &gt; 2000 - 1080 in your example).</p>\n<p>You can find an additional discussion&nbsp;<a href=\"https://www.kaggle.com/c/yandex-personalized-web-search-challenge/forums/t/6170/definition-of-relevance\">in another topic</a>&nbsp;.</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33437,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "11/08/2013 18:31:18",
      "content": "<p>file</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33439,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "11/08/2013 18:42:34",
      "content": "<p>The file above contains the relevances for the first 1000 queries.</p>\n<p>Calculation of relevances for the training sample is purely technical task, but one may easily make mistake on it. So I think it is a good idea to share such results for comparison and checking.</p>\n<p>It would be nice to have such file in competition data...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33505,
      "author_name": "corrision",
      "author_url": "",
      "post_date": "11/10/2013 13:16:22",
      "content": "<p>my first diff:</p>\n<p>6th line (session 4, serp 0), your result 0000000002</p>\n<p>data: 1st dwell time 86-27&gt;=50, 3rd dwell time 216-121&gt;=50, so i guess it should be 0001010002</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33507,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "11/10/2013 14:52:23",
      "content": "<p>Thank you very much. I've found a mistake in my program.</p>\n<p>Here is the corrected file.</p>\n<p>I could post the full file, but I am not sure that this is allowed by the rules.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34867,
      "author_name": "martincmartin",
      "author_url": "",
      "post_date": "11/18/2013 22:24:24",
      "content": "<p>I don't think you're interpreting the last click in a SERP properly, you're treating it as &quot;highly relevant,&quot; even when there's another SERP in the same session.</p>\n<p>For example, session 13 SERP 2 has a query then a single click at time 5658.&nbsp; In the same session, the user re-submits their search at time 5753, so the dwell time on that click should be 5753 - 5658 = 95, a &quot;1&quot; grade.&nbsp; But you have it marked as a &quot;2&quot; grade.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34883,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "11/19/2013 08:16:13",
      "content": "<p>[quote=Martin C. Martin;34867]</p>\n<p>I don't think you're interpreting the last click in a SERP properly, you're treating it as &quot;highly relevant,&quot; even when there's another SERP in the same session.</p>\n<p>[/quote]</p>\n<p>Thank you for the comment.</p>\n<p>But here is citation from &quot;Evaluation&quot;: &quot;2 (highly relevant) grade corresponds to documents with clicks and dwell time not shorter than 400 time units and to the <br>document with the last click in the search session&quot;.</p>\n<p>So, the click in session 13 SERP 2 should be highly relevant because it is really the last click in the session, since there were only queries but no clicks later!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34893,
      "author_name": "martincmartin",
      "author_url": "",
      "post_date": "11/19/2013 11:42:29",
      "content": "<p>Thanks.&nbsp; Eugene, can you confirm that the last click counts as &quot;highly relevant,&quot; even when the user comes back and does another search but doesn't click on anything?&nbsp; Doing another search suggests they didn't find what they wanted with their click, and are looking for something else.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34915,
      "author_name": "eugene1751",
      "author_url": "",
      "post_date": "11/19/2013 17:05:29",
      "content": "<p>Sorry for the ambiguity in the description. We'll fix it asap.</p>\n<p>The relevance level of 2 is assigned to the clicks which have dwell time not smaller than 400 time units and to the clicks that are <strong>the last actions in the corresponding sessions</strong>.</p>\n<p>In other words, if the dwell time can be calculated then the relevance is assigned according to it. If the click is the last action in the session (in other words, the next action is performed after a long time window separating two sessions) then the relevance is set to 2.</p>\n<p>We believe this definition is more consistent.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34940,
      "author_name": "martincmartin",
      "author_url": "",
      "post_date": "11/19/2013 22:44:17",
      "content": "<p>Victor, any chance you can update your file?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34973,
      "author_name": "alirezahonarvar",
      "author_url": "",
      "post_date": "11/20/2013 19:38:05",
      "content": "<p>Hi</p>\n<p>I download the Data, unziped it , but I can not open Them. which software should I use to open them. files do not have&nbsp; any extension !!!</p>\n<p>Please help me.</p>\n<p>&nbsp;</p>\n<p>TNX</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34974,
      "author_name": "curtis0",
      "author_url": "",
      "post_date": "11/20/2013 21:15:59",
      "content": "<p>You can open such files with good Text Editors like Notepad++. (except the train file ;))</p>\n<p>But for further editings you need to get line for line I think. Your RAM is not that big the store the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34984,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "11/21/2013 05:09:48",
      "content": "<p>[quote=Martin C. Martin;34940]</p>\n<p>Victor, any chance you can update your file?</p>\n<p>[/quote]</p>\n<p>Here it is.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34985,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "11/21/2013 05:14:05",
      "content": "<p>[quote=Ali Reza Honarvar;34973]</p>\n<p>I download the Data, unziped it , but I can not open Them. which software should I use to open them. files do not have&nbsp; any extension !!!</p>\n<p>[/quote]</p>\n<p>I use the F3-viewer in FAR-manager.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35093,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "11/23/2013 15:03:02",
      "content": "<p>[quote=Eugene;34915]</p>\n<p>Sorry for the ambiguity in the description. We'll fix it asap.</p>\n<p>[/quote]</p>\n<p>May be is it possible to provide a file with calculated relevances?</p>\n<p>There is still some uncertainty with the cases when some clicks occurs later than the next query. And with simultaneous events too.</p>\n<p>Here is my implementation of relevance (grade) calculation: &nbsp;&nbsp;&nbsp; </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35150,
      "author_name": "suzanv",
      "author_url": "",
      "post_date": "11/25/2013 08:55:34",
      "content": "<p>[quote=Ali Reza Honarvar;34973]</p>\n<p>Hi</p>\n<p>I download the Data, unziped it , but I can not open Them. which software should I use to open them. files do not have&nbsp; any extension !!!</p>\n<p>Please help me.</p>\n<p>[/quote]</p>\n<p><span style=\"line-height: 1.4\">I don't think you would want to open the complete train file at once. Under Linux, you can use 'more' to view the files. I advice you to develop your code on a small sample. You can create a sample without unzipping using the following command: 'gzip -cd train.gz | head -1000 &gt; train.sample'. Or on the unzipped file: 'cat train |&nbsp; head -1000 &gt; train.sample'</span></p>\n<p>If you are using java, then you can read the gz-files without unzipping them using the class&nbsp;GZIPInputStream.</p>\n<p>Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35204,
      "author_name": "habib128987",
      "author_url": "",
      "post_date": "11/25/2013 18:12:04",
      "content": "<p>Hi</p>\n<p>in windows you can use 'Large Texte File Viewer' to open the file.</p>\n<p>also you can use 'HJSplit' to split train file in many parts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35483,
      "author_name": "chervinskii",
      "author_url": "",
      "post_date": "11/29/2013 08:42:05",
      "content": "<p>I found this strange piece of data:</p>\n<p>22 &nbsp; &nbsp; &nbsp;0 &nbsp; &nbsp; &nbsp; Q &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; 16185818 &nbsp; &nbsp; &nbsp; &nbsp;4078517,3143668,1988399 35767635,3364600 &nbsp; &nbsp; &nbsp; &nbsp;57421304,4536568 &nbsp; &nbsp; &nbsp; &nbsp;35296159,3323845 &nbsp; &nbsp; &nbsp; &nbsp;49235419,4140722 &nbsp; &nbsp; &nbsp; &nbsp;68222659,5116467 &nbsp; &nbsp; &nbsp; &nbsp;55142634,4429180 &nbsp; &nbsp; &nbsp; &nbsp;36577088,3407786 &nbsp; &nbsp; &nbsp; &nbsp;61943983,4767047 &nbsp; &nbsp; &nbsp; &nbsp;65157339,4954478 &nbsp; &nbsp; &nbsp; &nbsp;42867421,3782572</p>\n<p>22 &nbsp; &nbsp; &nbsp;25 &nbsp; &nbsp; &nbsp;Q &nbsp; &nbsp; &nbsp; 2 &nbsp; &nbsp; &nbsp; 16185820 &nbsp; &nbsp; &nbsp; &nbsp;4078517,3143668,1988399,2133749 30325446,2938326 &nbsp; &nbsp; &nbsp; &nbsp;28846733,2833948 &nbsp; &nbsp; &nbsp; &nbsp;35769151,3364600&nbsp; &nbsp; &nbsp; &nbsp; 57008295,4511361 &nbsp; &nbsp; &nbsp; &nbsp;28175892,2768660 &nbsp; &nbsp; &nbsp; &nbsp;28864909,2833991 &nbsp; &nbsp; &nbsp; &nbsp;60703868,4697323 &nbsp; &nbsp; &nbsp; &nbsp;22790363,2294429 &nbsp; &nbsp; &nbsp; &nbsp;57692687,4551821&nbsp; &nbsp; &nbsp; &nbsp; 6473264,800608</p>\n<p>22 &nbsp; &nbsp; &nbsp;25 &nbsp; &nbsp; &nbsp;Q &nbsp; &nbsp; &nbsp; 1 &nbsp; &nbsp; &nbsp; 16185818 &nbsp; &nbsp; &nbsp; &nbsp;4078517,3143668,1988399 35767635,3364600 &nbsp; &nbsp; &nbsp; &nbsp;57421304,4536568 &nbsp; &nbsp; &nbsp; &nbsp;35296159,3323845 &nbsp; &nbsp; &nbsp; &nbsp;49235419,4140722 &nbsp; &nbsp; &nbsp; &nbsp;68222659,5116467 &nbsp; &nbsp; &nbsp; &nbsp;55142634,4429180 &nbsp; &nbsp; &nbsp; &nbsp;36577088,3407786 &nbsp; &nbsp; &nbsp; &nbsp;61943983,4767047 &nbsp; &nbsp; &nbsp; &nbsp;65157339,4954478 &nbsp; &nbsp; &nbsp; &nbsp;42867421,3782572</p>\n<p>22 &nbsp; &nbsp; &nbsp;31 &nbsp; &nbsp; &nbsp;C &nbsp; &nbsp; &nbsp; 2 &nbsp; &nbsp; &nbsp; 30325446</p>\n<p>22 &nbsp; &nbsp; &nbsp;112 &nbsp; &nbsp; C &nbsp; &nbsp; &nbsp; 2 &nbsp; &nbsp; &nbsp; 28175892</p>\n<p>Notice two queries at the same time!&nbsp;</p>\n<p>This corresponds to strings #51-53 in your file :</p>\n<p>0000000000 &nbsp;(serpid = 0)</p>\n<p>1000100000 &nbsp;(serpid = 2)</p>\n<p>0000000000 &nbsp;(serpid = 1)</p>\n<p>But don't you think we should sort them by SERPID?</p>\n<p>In the first 1000 queries this trouble (with two queries in one unit of time written in reverse order of SERPIDs) appears one another time in session #458.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35486,
      "author_name": "chervinskii",
      "author_url": "",
      "post_date": "11/29/2013 10:04:54",
      "content": "<p>...and also I found mistake in your file:</p>\n<p>309 &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; Q &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; 11466723 &nbsp; &nbsp; &nbsp; &nbsp;3332767,4291604 29370045,2863645 &nbsp; &nbsp; &nbsp; &nbsp;55303432,4433251 &nbsp; &nbsp; &nbsp; &nbsp;3276363,443265 &nbsp;29322278,2863497&nbsp; &nbsp; &nbsp; &nbsp; 29371141,2863645 &nbsp; &nbsp; &nbsp; &nbsp;36829352,3418807 &nbsp; &nbsp; &nbsp; &nbsp;48935893,4127107 &nbsp; &nbsp; &nbsp; &nbsp;27402385,2655130 &nbsp; &nbsp; &nbsp; &nbsp;28137920,2761328 &nbsp; &nbsp; &nbsp; &nbsp;62315277,4791822</p>\n<p>309 &nbsp; &nbsp; 13 &nbsp; &nbsp; &nbsp;Q &nbsp; &nbsp; &nbsp; 1 &nbsp; &nbsp; &nbsp; 11466723 &nbsp; &nbsp; &nbsp; &nbsp;3332767,4291604 29370045,2863645 &nbsp; &nbsp; &nbsp; &nbsp;55303432,4433251 &nbsp; &nbsp; &nbsp; &nbsp;3276363,443265 &nbsp;29322278,2863497&nbsp; &nbsp; &nbsp; &nbsp; 29371141,2863645 &nbsp; &nbsp; &nbsp; &nbsp;36829352,3418807 &nbsp; &nbsp; &nbsp; &nbsp;48935893,4127107 &nbsp; &nbsp; &nbsp; &nbsp;27402385,2655130 &nbsp; &nbsp; &nbsp; &nbsp;28137920,2761328 &nbsp; &nbsp; &nbsp; &nbsp;62315277,4791822</p>\n<p>309 &nbsp; &nbsp; 13 &nbsp; &nbsp; &nbsp;C &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; 29370045</p>\n<p>310 &nbsp; &nbsp; M &nbsp; &nbsp; &nbsp; 2 &nbsp; &nbsp; &nbsp; 52 &nbsp; &nbsp; &nbsp;</p>\n<p>Here should be &quot;2000000000&quot; instead of &quot;0000000000&quot; in your file (line 557) because the click is the last in this session.&nbsp;I suppose, this may be associated with mixed SERPIDs.</p>\n<p>Excluding this two things I have the same results for relevances.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "33375": "",
    "33379": "",
    "33437": "",
    "33439": "",
    "33505": "",
    "33507": "",
    "34867": "",
    "34883": "",
    "34893": "",
    "34915": "",
    "34940": "",
    "34973": "",
    "34974": "",
    "34984": "",
    "34985": "",
    "35093": "",
    "35150": "",
    "35204": "",
    "35483": "",
    "35486": ""
  },
  "source": "meta"
}