{
  "id": 6489,
  "title": "Python code for parsing data",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6489",
  "author_name": "",
  "post_date": "2013-12-02T00:15:44.837Z",
  "votes": 14,
  "comment_count": 23,
  "views": 8630,
  "content": "<p>In an effort to get more people involved in this competition and to selfishly&nbsp;have more eyes to sanity check my code I'm attaching the code I am using for parsing the training and test data for this competition.&nbsp;<br><br>This could be used as a first step for more complicated preprocessing methods.&nbsp;<br><br>I've made some assumptions when generating this code, so please let me know if any of these assumptions are violated during your investigations in the data. I'll try to quickly update the code to fix any errors.&nbsp;<br><br>Assumptions:&nbsp;</p>\n<p>1. All logs have the form in this strict order:<br>- session metadata<br>- one or more queries<br>- zero or more clicks&nbsp;<br>2. sessions are self-contained in that order, so if new session metadata is seen then it is assumed that the previous session being parsed will not appear in the data again<br>3. Clicks are directly associated with the last query observed in logs<br><br>This is a first attempt at parsing big data in python, so if you notice that I'm doing something very wrong in my code or have suggestions on how I could clean up the code or do things more efficiently, please let me know.&nbsp;<br><br>The pastebin for the code is:&nbsp;http://pastebin.com/gaPnVwNH<br><br>Furthermore, here is a snippet for how to use the parser code to extract the first 100 sessions from the training file:<br><br>import gzip, parser<br>f = gzip.open('data/train.gz', 'rb') #assuming train.gz is located in data/<br>sp = parser.parse_sessions(f)<br>sessions = [sp.next() for i in range(100)]<br><br>Hope this helps :)<br><br>UPDATE:&nbsp;<br>I've modified the code to fix the bug noticed by&nbsp;kinnskogr. Two changes are that a session can now have multiple queries, and each query is associated with its own clicks.&nbsp;</p>",
  "messages": [
    {
      "id": "35608",
      "postDate": "12/02/2013 00:15:44",
      "content": "<p>In an effort to get more people involved in this competition and to selfishly&nbsp;have more eyes to sanity check my code I'm attaching the code I am using for parsing the training and test data for this competition.&nbsp;<br><br>This could be used as a first step for more complicated preprocessing methods.&nbsp;<br><br>I've made some assumptions when generating this code, so please let me know if any of these assumptions are violated during your investigations in the data. I'll try to quickly update the code to fix any errors.&nbsp;<br><br>Assumptions:&nbsp;</p>\n<p>1. All logs have the form in this strict order:<br>- session metadata<br>- one or more queries<br>- zero or more clicks&nbsp;<br>2. sessions are self-contained in that order, so if new session metadata is seen then it is assumed that the previous session being parsed will not appear in the data again<br>3. Clicks are directly associated with the last query observed in logs<br><br>This is a first attempt at parsing big data in python, so if you notice that I'm doing something very wrong in my code or have suggestions on how I could clean up the code or do things more efficiently, please let me know.&nbsp;<br><br>The pastebin for the code is:&nbsp;http://pastebin.com/gaPnVwNH<br><br>Furthermore, here is a snippet for how to use the parser code to extract the first 100 sessions from the training file:<br><br>import gzip, parser<br>f = gzip.open('data/train.gz', 'rb') #assuming train.gz is located in data/<br>sp = parser.parse_sessions(f)<br>sessions = [sp.next() for i in range(100)]<br><br>Hope this helps :)<br><br>UPDATE:&nbsp;<br>I've modified the code to fix the bug noticed by&nbsp;kinnskogr. Two changes are that a session can now have multiple queries, and each query is associated with its own clicks.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35708",
      "postDate": "12/02/2013 22:29:47",
      "content": "<p>Hi Miroslaw,</p>\n<p>&nbsp; &nbsp;I think you've got a bug in your parser. It assumes that there's one query per sessions ID. You'll end up with an output which uses the last query performed during one sessions, but the clicks from all the queries in the session. See SessionID 5 for an example where there are 5 clicks spread across 3 queries. For a minimal fix, you should initialize the Query field as a list, and append to it as you discover query lines.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35711",
      "postDate": "12/02/2013 23:09:54",
      "content": "<p>Thanks for the feedback kinnskogr! You're right, that is a bug.&nbsp;<br><br>That minimal fix would work in maintaining queries but it would be difficult to associate clicks directly with a specific query. How about instead we assign clicks directly to query objects? That should be more robust. Would it be appropriate to assume that all clicks in between two queries in a session are associated with the first query?&nbsp;<br><br>I've updated the code to apply the fix. A session can contain multiple queries, and each query is associated with its own clicks.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35721",
      "postDate": "12/03/2013 02:06:16",
      "content": "<p>That would work as well. It's basically just a design decision, and depends on what kind of steps you want to do while pre-processing. I prefer two separate lists, then you can treat them as tables in a relational database and do joins to generate sets of features.</p>\n<p>I think you can make the assumption that all clicks in between two queries are associated with the first (I didn't see any exceptions to that rule), but if you want to be more robust, you can check the&nbsp;SERPID, since the key is shared by the queries and the clicks and uniquely identifies each query in a given session.</p>\n<p>By the way, thanks for making your code public. I'd taken a similar approach, but was trying to nest pandas datastructures. Your code using the native python types is about an order of magnitude faster!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35733",
      "postDate": "12/03/2013 08:40:16",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35737",
      "postDate": "12/03/2013 10:28:45",
      "content": "<p>Also, thanks for making the code public it helps to understand things a lot more and like you said, this will help more people to want to be part of this competition. Your code is very simple and straightforward and so that is something that is appreciated by many.&nbsp;</p>\n<p>I had tried the code myself would like to see the changes that you will apply to fix the 1 query -&gt; many click problem. At least this serves as a baseline for people that are not too good with programming.&nbsp;</p>\n<p>Thank you.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35739",
      "postDate": "12/03/2013 11:57:43",
      "content": "<p>[quote=kinnskogr;35721]</p>\n<p>... but if you want to be more robust, you can check the SERPID, since the key is shared by the queries and the clicks and uniquely identifies each query in a given session.</p>\n<p>[/quote]</p>\n<p>Haha, I've been scratching my head for a while trying to figure out what the SERPID was for. Now it just seems obvious. Doh!</p>\n<p>Thanks for clearing that up for me :)</p>\n<p>[quote=Ibelmopan Belizean;35737]</p>\n<p>Also, thanks for making the code public it helps to understand things a lot more and like you said, this will help more people to want to be part of this competition. Your code is very simple and straightforward and so that is something that is appreciated by many.&nbsp;</p>\n<p>I had tried the code myself would like to see the changes that you will apply to fix the 1 query -&gt; many click problem. At least this serves as a baseline for people that are not too good with programming.&nbsp;</p>\n<p>Thank you.</p>\n<p>[/quote]<br>The code was updated about 12h ago, so if you followed the link in the first post since then, then you should have gotten the code with the applied fix. I added a note on line 5 of the file describing the changes I made. To summarize: I ended&nbsp;up assigning clicks directly to each query object. So as an example, if you wanted to get the clicks from the first query in a session you could do something like this:<br><br>session['Query'][0]['Clicks']<br><br>If you want to see how many queries are in a session you could do:<br><br>len(session['Query'])<br><br>kinnskogr's solution to have clicks owned directly by the session object is also equally valid, just a matter of personal preference. So if you prefer kinnskogr's approach you can modify the code as an exercise. It should only be three minor modifications to the updated code.&nbsp;<br><br></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35855",
      "postDate": "12/05/2013 07:11:43",
      "content": "<p>Hi Miroslaw,</p>\n<p>Thanks for the code! &nbsp;I think I will use it to get started :)</p>\n<p>&gt;&nbsp;I've updated the code to apply the fix. A session can contain multiple queries, and each query is associated with its own clicks.&nbsp;</p>\n<p>Perhaps you can update the comment in the code to reflect this? &nbsp;The current comment looks as if it can only contain one query:</p>\n<p>Query: { TimePassed: int,<br> SERPID: int,<br> QueryID: int,<br> ListOfTerms: [TermID_1, ...],<br> Clicks: [{ TimePassed: int, SERPID: int, URLID: int }, ...],<br> URL_DOMAIN: [(URLID_1, DomainID_1), ...] } }</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35856",
      "postDate": "12/05/2013 07:39:02",
      "content": "<p>[quote=shan;35855]<span style=\"line-height: 1.4\">&nbsp;</span></p>\n<p>Perhaps you can update the comment in the code to reflect this? &nbsp;The current comment looks as if it can only contain one query ...<span style=\"line-height: 1.4\">[/quote]<br><br>Good point. I've updated the docstring.&nbsp;</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35935",
      "postDate": "12/07/2013 03:47:21",
      "content": "<p>On top of Miroslaw's parser, I made a parser to generate user objects instead of session objects.</p>\n<p>Sessions belong to users, and queries belong to sessions.</p>\n<p>http://pastebin.com/0fSMCrU9</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36058",
      "postDate": "12/11/2013 12:32:26",
      "content": "<p>Here is our parsing script.<br>https://gist.github.com/poulejapon/7909562</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36086",
      "postDate": "12/11/2013 19:22:37",
      "content": "<p>i am getting error as</p>\n<p>AttributeError: 'module' object has no attribute 'parse_sessions'</p>\n<p>Though i have used customised parser module given by you</p>\n<p>Where i am doing wrong</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36087",
      "postDate": "12/11/2013 19:23:55",
      "content": "<p>[quote=Parthiban Gowthaman;36086]</p>\n<p>i am getting error as</p>\n<p>AttributeError: 'module' object has no attribute 'parse_sessions'</p>\n<p>Though i have used customised parser module given by you</p>\n<p>Where i am doing wrong</p>\n<p>[/quote]</p>\n<p>I am using the first parser code given in this forum</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36174",
      "postDate": "12/14/2013 20:19:18",
      "content": "<p>&nbsp;</p>\n<p><strong>Gowthaman</strong> - I *think* it was a plain typo on Miroslaw's part .</p>\n<p>Following, is my best guess - may be <strong>@Miroslaw</strong> can confirm ?</p>\n<p># dir(parse) has no methods/functions which take a generator object - see this for yourself. &nbsp;</p>\n<p>sp=parse_sessions(f) &nbsp;# and <strong>NOT</strong> parse.parse_sessions(f)&nbsp;<br>sessions = [sp.next() for i in range(100)]</p>\n<p>Post modifying this I could get the session objects as&nbsp;</p>\n<p>[{'Query': [{'URL_DOMAIN': [(50504886, 4217515), (9848058, 1084315), (50534229, 4217515), (50591618, 4217515), (26242582, 2597528), (34623075, 3279130), (68893581, 5149883), (50628761, 4217517), (32262001, 3142702), (35443881, 3339757)], 'SERPID': 0, 'QueryID': 10047345, 'TimePassed': 0, 'ListOfTerms': [3080290, 4098689], 'Clicks': [{'URLID': 50628761, 'TimePassed': 108, 'SERPID': 0}, {'URLID': 50628761, 'TimePassed': 1080, 'SERPID': 0}]}], 'SessionID': 0, 'USERID': 0, 'Day': 4}]</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36180",
      "postDate": "12/15/2013 06:46:44",
      "content": "<p>[quote=ekta1007;36174]</p>\n<p>&nbsp;</p>\n<p><strong>Gowthaman</strong> - I *think* it was a plain typo on Miroslaw's part .</p>\n<p>Following, is my best guess - may be <strong>@Miroslaw</strong> can confirm ?</p>\n<p># dir(pa<span style=\"line-height: 1.4\">rse) has no methods/functions which take a generator object - see this for yourself. &nbsp;</span></p>\n<p><span style=\"line-height: 1.4\">sp=parse_sessions(f) &nbsp;# and </span><strong style=\"line-height: 1.4\">NOT</strong><span style=\"line-height: 1.4\"> parse.parse_sessions(f)&nbsp;</span></p>\n<p>sessions = [sp.next() for i in range(100)]</p>\n<p>Post modifying this I could get the session objects as&nbsp;</p>\n<p>[{'Query': [{'URL_DOMAIN': [(50504886, 4217515), (9848058, 1084315), (50534229, 4217515), (50591618, 4217515), (26242582, 2597528), (34623075, 3279130), (68893581, 5149883), (50628761, 4217517), (32262001, 3142702), (35443881, 3339757)], 'SERPID': 0, 'QueryID': 10047345, 'TimePassed': 0, 'ListOfTerms': [3080290, 4098689], 'Clicks': [{'URLID': 50628761, 'TimePassed': 108, 'SERPID': 0}, {'URLID': 50628761, 'TimePassed': 1080, 'SERPID': 0}]}], 'SessionID': 0, 'USERID': 0, 'Day': 4}]</p>\n<p>&nbsp;</p>\n<p>[/quote]</p>\n<p>Thanks i got it.When i import parser it is importing default parser in python.So i named parser customized module given in forum as parthi &amp; called&nbsp;</p>\n<p>sp=parthi.parse_sessions(f)</p>\n<p>its working.</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36217",
      "postDate": "12/16/2013 13:27:20",
      "content": "<p>Hi Miroslaw,</p>\n<p>&nbsp;</p>\n<p>your scripst is great, thanks a lot!</p>\n<p>I think I found a small bug in there.</p>\n<p>Your script misses to parse the last session in the file. You yield a session when a new meta tag shows up. But for the last session, there is no following meta tag.</p>\n<p>An additional &quot;yield s&quot; directly following the for loop solved it.</p>\n<p>&nbsp;</p>\n<p>I think it's not a problem with the train file, but when you parse the test file, you miss to predict a whole session.</p>\n<p>&nbsp;</p>\n<p>Sebastian</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36397",
      "postDate": "12/20/2013 08:25:49",
      "content": "<p>Can I discusss a dataset from another US company. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36511",
      "postDate": "12/22/2013 16:26:15",
      "content": "<p>How much time did it take to parse the train file completely? I have written my parser in c++ and it is running for 4 hours.</p>\n<p>&nbsp;</p>\n<p>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36512",
      "postDate": "12/22/2013 16:42:38",
      "content": "<p>It took 7 hours for me (including writing it to a mongoDB).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36518",
      "postDate": "12/22/2013 17:39:19",
      "content": "<p>[quote=Sebastian Butterweck;36217]</p>\n<p>Hi Miroslaw,</p>\n<p>&nbsp;</p>\n<p>your scripst is great, thanks a lot!</p>\n<p>I think I found a small bug in there.</p>\n<p>Your script misses to parse the last session in the file. You yield a session when a new meta tag shows up. But for the last session, there is no following meta tag.</p>\n<p>An additional &quot;yield s&quot; directly following the for loop solved it.</p>\n<p>&nbsp;</p>\n<p>I think it's not a problem with the train file, but when you parse the test file, you miss to predict a whole session.</p>\n<p>&nbsp;</p>\n<p>Sebastian</p>\n<p>[/quote]</p>\n<p>Thanks for posting this on the forums, you're right it's a bug. Shan actually found that bug a few weeks ago and sent me an email describing the error so I'd also like to give him some credit for that :) <br><br>I've updated the code on pastebin for future reference.</p>\n<p>[quote=DerivedByData;36511]</p>\n<p>How much time did it take to parse the train file completely? I have written my parser in c++ and it is running for 4 hours.</p>\n<p><span style=\"line-height: 1.4\">[/quote]</span><br><br><span style=\"line-height: 1.4\">My script runs in about an hour for the training file and under 5 minutes for the test file. I write all my parsing results to&nbsp;</span>separate<span style=\"line-height: 1.4\">&nbsp;files for later processing.&nbsp;</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36748",
      "postDate": "12/28/2013 10:55:27",
      "content": "<p>Hey, I guess in first parse you had just collected the data and put it in structured form in you DB. After that you must have extracted features from that crude data. How much time did it take for generating the features?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36749",
      "postDate": "12/28/2013 11:05:31",
      "content": "<p>From start to end, including learning we take probably a day or so on a 12 core computer here...&nbsp;:)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "807955",
      "postDate": "04/15/2020 04:40:38",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks",
      "votes": null
    },
    {
      "id": "994779",
      "postDate": "09/02/2020 00:06:30",
      "content": "<p><a href=\"https://www.kaggle.com/shan4104543\" target=\"_blank\">@shan4104543</a> <br>\nThanks for your codes.<br>\nI am sorry that the link is not available, could you please update a new link for the script?</p>\n<p>Thankyou</p>",
      "rawMarkdown": "shan4104543 \nThanks for your codes.\nI am sorry that the link is not available, could you please update a new link for the script?\n\nThankyou",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 994779,
      "author_name": "jiaqiwang1107",
      "author_url": "",
      "post_date": "09/02/2020 00:06:30",
      "content": "<p><a href=\"https://www.kaggle.com/shan4104543\" target=\"_blank\">@shan4104543</a> <br>\nThanks for your codes.<br>\nI am sorry that the link is not available, could you please update a new link for the script?</p>\n<p>Thankyou</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35708,
      "author_name": "kinnskogr",
      "author_url": "",
      "post_date": "12/02/2013 22:29:47",
      "content": "<p>Hi Miroslaw,</p>\n<p>&nbsp; &nbsp;I think you've got a bug in your parser. It assumes that there's one query per sessions ID. You'll end up with an output which uses the last query performed during one sessions, but the clicks from all the queries in the session. See SessionID 5 for an example where there are 5 clicks spread across 3 queries. For a minimal fix, you should initialize the Query field as a list, and append to it as you discover query lines.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35711,
      "author_name": "miroslaw",
      "author_url": "",
      "post_date": "12/02/2013 23:09:54",
      "content": "<p>Thanks for the feedback kinnskogr! You're right, that is a bug.&nbsp;<br><br>That minimal fix would work in maintaining queries but it would be difficult to associate clicks directly with a specific query. How about instead we assign clicks directly to query objects? That should be more robust. Would it be appropriate to assume that all clicks in between two queries in a session are associated with the first query?&nbsp;<br><br>I've updated the code to apply the fix. A session can contain multiple queries, and each query is associated with its own clicks.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35721,
      "author_name": "kinnskogr",
      "author_url": "",
      "post_date": "12/03/2013 02:06:16",
      "content": "<p>That would work as well. It's basically just a design decision, and depends on what kind of steps you want to do while pre-processing. I prefer two separate lists, then you can treat them as tables in a relational database and do joins to generate sets of features.</p>\n<p>I think you can make the assumption that all clicks in between two queries are associated with the first (I didn't see any exceptions to that rule), but if you want to be more robust, you can check the&nbsp;SERPID, since the key is shared by the queries and the clicks and uniquely identifies each query in a given session.</p>\n<p>By the way, thanks for making your code public. I'd taken a similar approach, but was trying to nest pandas datastructures. Your code using the native python types is about an order of magnitude faster!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35733,
      "author_name": "ibelmopanbelizean",
      "author_url": "",
      "post_date": "12/03/2013 08:40:16",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 35737,
      "author_name": "ibelmopanbelizean",
      "author_url": "",
      "post_date": "12/03/2013 10:28:45",
      "content": "<p>Also, thanks for making the code public it helps to understand things a lot more and like you said, this will help more people to want to be part of this competition. Your code is very simple and straightforward and so that is something that is appreciated by many.&nbsp;</p>\n<p>I had tried the code myself would like to see the changes that you will apply to fix the 1 query -&gt; many click problem. At least this serves as a baseline for people that are not too good with programming.&nbsp;</p>\n<p>Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35739,
      "author_name": "miroslaw",
      "author_url": "",
      "post_date": "12/03/2013 11:57:43",
      "content": "<p>[quote=kinnskogr;35721]</p>\n<p>... but if you want to be more robust, you can check the SERPID, since the key is shared by the queries and the clicks and uniquely identifies each query in a given session.</p>\n<p>[/quote]</p>\n<p>Haha, I've been scratching my head for a while trying to figure out what the SERPID was for. Now it just seems obvious. Doh!</p>\n<p>Thanks for clearing that up for me :)</p>\n<p>[quote=Ibelmopan Belizean;35737]</p>\n<p>Also, thanks for making the code public it helps to understand things a lot more and like you said, this will help more people to want to be part of this competition. Your code is very simple and straightforward and so that is something that is appreciated by many.&nbsp;</p>\n<p>I had tried the code myself would like to see the changes that you will apply to fix the 1 query -&gt; many click problem. At least this serves as a baseline for people that are not too good with programming.&nbsp;</p>\n<p>Thank you.</p>\n<p>[/quote]<br>The code was updated about 12h ago, so if you followed the link in the first post since then, then you should have gotten the code with the applied fix. I added a note on line 5 of the file describing the changes I made. To summarize: I ended&nbsp;up assigning clicks directly to each query object. So as an example, if you wanted to get the clicks from the first query in a session you could do something like this:<br><br>session['Query'][0]['Clicks']<br><br>If you want to see how many queries are in a session you could do:<br><br>len(session['Query'])<br><br>kinnskogr's solution to have clicks owned directly by the session object is also equally valid, just a matter of personal preference. So if you prefer kinnskogr's approach you can modify the code as an exercise. It should only be three minor modifications to the updated code.&nbsp;<br><br></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35855,
      "author_name": "shan4104543",
      "author_url": "",
      "post_date": "12/05/2013 07:11:43",
      "content": "<p>Hi Miroslaw,</p>\n<p>Thanks for the code! &nbsp;I think I will use it to get started :)</p>\n<p>&gt;&nbsp;I've updated the code to apply the fix. A session can contain multiple queries, and each query is associated with its own clicks.&nbsp;</p>\n<p>Perhaps you can update the comment in the code to reflect this? &nbsp;The current comment looks as if it can only contain one query:</p>\n<p>Query: { TimePassed: int,<br> SERPID: int,<br> QueryID: int,<br> ListOfTerms: [TermID_1, ...],<br> Clicks: [{ TimePassed: int, SERPID: int, URLID: int }, ...],<br> URL_DOMAIN: [(URLID_1, DomainID_1), ...] } }</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35856,
      "author_name": "miroslaw",
      "author_url": "",
      "post_date": "12/05/2013 07:39:02",
      "content": "<p>[quote=shan;35855]<span style=\"line-height: 1.4\">&nbsp;</span></p>\n<p>Perhaps you can update the comment in the code to reflect this? &nbsp;The current comment looks as if it can only contain one query ...<span style=\"line-height: 1.4\">[/quote]<br><br>Good point. I've updated the docstring.&nbsp;</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35935,
      "author_name": "shan4104543",
      "author_url": "",
      "post_date": "12/07/2013 03:47:21",
      "content": "<p>On top of Miroslaw's parser, I made a parser to generate user objects instead of session objects.</p>\n<p>Sessions belong to users, and queries belong to sessions.</p>\n<p>http://pastebin.com/0fSMCrU9</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36058,
      "author_name": "",
      "author_url": "",
      "post_date": "12/11/2013 12:32:26",
      "content": "<p>Here is our parsing script.<br>https://gist.github.com/poulejapon/7909562</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36086,
      "author_name": "parthibangowthaman",
      "author_url": "",
      "post_date": "12/11/2013 19:22:37",
      "content": "<p>i am getting error as</p>\n<p>AttributeError: 'module' object has no attribute 'parse_sessions'</p>\n<p>Though i have used customised parser module given by you</p>\n<p>Where i am doing wrong</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36087,
      "author_name": "parthibangowthaman",
      "author_url": "",
      "post_date": "12/11/2013 19:23:55",
      "content": "<p>[quote=Parthiban Gowthaman;36086]</p>\n<p>i am getting error as</p>\n<p>AttributeError: 'module' object has no attribute 'parse_sessions'</p>\n<p>Though i have used customised parser module given by you</p>\n<p>Where i am doing wrong</p>\n<p>[/quote]</p>\n<p>I am using the first parser code given in this forum</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36174,
      "author_name": "ektagrover1",
      "author_url": "",
      "post_date": "12/14/2013 20:19:18",
      "content": "<p>&nbsp;</p>\n<p><strong>Gowthaman</strong> - I *think* it was a plain typo on Miroslaw's part .</p>\n<p>Following, is my best guess - may be <strong>@Miroslaw</strong> can confirm ?</p>\n<p># dir(parse) has no methods/functions which take a generator object - see this for yourself. &nbsp;</p>\n<p>sp=parse_sessions(f) &nbsp;# and <strong>NOT</strong> parse.parse_sessions(f)&nbsp;<br>sessions = [sp.next() for i in range(100)]</p>\n<p>Post modifying this I could get the session objects as&nbsp;</p>\n<p>[{'Query': [{'URL_DOMAIN': [(50504886, 4217515), (9848058, 1084315), (50534229, 4217515), (50591618, 4217515), (26242582, 2597528), (34623075, 3279130), (68893581, 5149883), (50628761, 4217517), (32262001, 3142702), (35443881, 3339757)], 'SERPID': 0, 'QueryID': 10047345, 'TimePassed': 0, 'ListOfTerms': [3080290, 4098689], 'Clicks': [{'URLID': 50628761, 'TimePassed': 108, 'SERPID': 0}, {'URLID': 50628761, 'TimePassed': 1080, 'SERPID': 0}]}], 'SessionID': 0, 'USERID': 0, 'Day': 4}]</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36180,
      "author_name": "parthibangowthaman",
      "author_url": "",
      "post_date": "12/15/2013 06:46:44",
      "content": "<p>[quote=ekta1007;36174]</p>\n<p>&nbsp;</p>\n<p><strong>Gowthaman</strong> - I *think* it was a plain typo on Miroslaw's part .</p>\n<p>Following, is my best guess - may be <strong>@Miroslaw</strong> can confirm ?</p>\n<p># dir(pa<span style=\"line-height: 1.4\">rse) has no methods/functions which take a generator object - see this for yourself. &nbsp;</span></p>\n<p><span style=\"line-height: 1.4\">sp=parse_sessions(f) &nbsp;# and </span><strong style=\"line-height: 1.4\">NOT</strong><span style=\"line-height: 1.4\"> parse.parse_sessions(f)&nbsp;</span></p>\n<p>sessions = [sp.next() for i in range(100)]</p>\n<p>Post modifying this I could get the session objects as&nbsp;</p>\n<p>[{'Query': [{'URL_DOMAIN': [(50504886, 4217515), (9848058, 1084315), (50534229, 4217515), (50591618, 4217515), (26242582, 2597528), (34623075, 3279130), (68893581, 5149883), (50628761, 4217517), (32262001, 3142702), (35443881, 3339757)], 'SERPID': 0, 'QueryID': 10047345, 'TimePassed': 0, 'ListOfTerms': [3080290, 4098689], 'Clicks': [{'URLID': 50628761, 'TimePassed': 108, 'SERPID': 0}, {'URLID': 50628761, 'TimePassed': 1080, 'SERPID': 0}]}], 'SessionID': 0, 'USERID': 0, 'Day': 4}]</p>\n<p>&nbsp;</p>\n<p>[/quote]</p>\n<p>Thanks i got it.When i import parser it is importing default parser in python.So i named parser customized module given in forum as parthi &amp; called&nbsp;</p>\n<p>sp=parthi.parse_sessions(f)</p>\n<p>its working.</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36217,
      "author_name": "butterweck",
      "author_url": "",
      "post_date": "12/16/2013 13:27:20",
      "content": "<p>Hi Miroslaw,</p>\n<p>&nbsp;</p>\n<p>your scripst is great, thanks a lot!</p>\n<p>I think I found a small bug in there.</p>\n<p>Your script misses to parse the last session in the file. You yield a session when a new meta tag shows up. But for the last session, there is no following meta tag.</p>\n<p>An additional &quot;yield s&quot; directly following the for loop solved it.</p>\n<p>&nbsp;</p>\n<p>I think it's not a problem with the train file, but when you parse the test file, you miss to predict a whole session.</p>\n<p>&nbsp;</p>\n<p>Sebastian</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36397,
      "author_name": "sheikhadnanahmedusmani",
      "author_url": "",
      "post_date": "12/20/2013 08:25:49",
      "content": "<p>Can I discusss a dataset from another US company. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36511,
      "author_name": "abhayprakash",
      "author_url": "",
      "post_date": "12/22/2013 16:26:15",
      "content": "<p>How much time did it take to parse the train file completely? I have written my parser in c++ and it is running for 4 hours.</p>\n<p>&nbsp;</p>\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36512,
      "author_name": "butterweck",
      "author_url": "",
      "post_date": "12/22/2013 16:42:38",
      "content": "<p>It took 7 hours for me (including writing it to a mongoDB).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36518,
      "author_name": "miroslaw",
      "author_url": "",
      "post_date": "12/22/2013 17:39:19",
      "content": "<p>[quote=Sebastian Butterweck;36217]</p>\n<p>Hi Miroslaw,</p>\n<p>&nbsp;</p>\n<p>your scripst is great, thanks a lot!</p>\n<p>I think I found a small bug in there.</p>\n<p>Your script misses to parse the last session in the file. You yield a session when a new meta tag shows up. But for the last session, there is no following meta tag.</p>\n<p>An additional &quot;yield s&quot; directly following the for loop solved it.</p>\n<p>&nbsp;</p>\n<p>I think it's not a problem with the train file, but when you parse the test file, you miss to predict a whole session.</p>\n<p>&nbsp;</p>\n<p>Sebastian</p>\n<p>[/quote]</p>\n<p>Thanks for posting this on the forums, you're right it's a bug. Shan actually found that bug a few weeks ago and sent me an email describing the error so I'd also like to give him some credit for that :) <br><br>I've updated the code on pastebin for future reference.</p>\n<p>[quote=DerivedByData;36511]</p>\n<p>How much time did it take to parse the train file completely? I have written my parser in c++ and it is running for 4 hours.</p>\n<p><span style=\"line-height: 1.4\">[/quote]</span><br><br><span style=\"line-height: 1.4\">My script runs in about an hour for the training file and under 5 minutes for the test file. I write all my parsing results to&nbsp;</span>separate<span style=\"line-height: 1.4\">&nbsp;files for later processing.&nbsp;</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36748,
      "author_name": "abhayprakash",
      "author_url": "",
      "post_date": "12/28/2013 10:55:27",
      "content": "<p>Hey, I guess in first parse you had just collected the data and put it in structured form in you DB. After that you must have extracted features from that crude data. How much time did it take for generating the features?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36749,
      "author_name": "",
      "author_url": "",
      "post_date": "12/28/2013 11:05:31",
      "content": "<p>From start to end, including learning we take probably a day or so on a 12 core computer here...&nbsp;:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 807955,
      "author_name": "venkateshindia",
      "author_url": "",
      "post_date": "04/15/2020 04:40:38",
      "content": "<p>thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "35608": "",
    "35708": "",
    "35711": "",
    "35721": "",
    "35733": "",
    "35737": "",
    "35739": "",
    "35855": "",
    "35856": "",
    "35935": "",
    "36058": "",
    "36086": "",
    "36087": "",
    "36174": "",
    "36180": "",
    "36217": "",
    "36397": "",
    "36511": "",
    "36512": "",
    "36518": "",
    "36748": "",
    "36749": "",
    "807955": "thanks",
    "994779": "shan4104543 \nThanks for your codes.\nI am sorry that the link is not available, could you please update a new link for the script?\n\nThankyou"
  },
  "source": "meta"
}