{
  "id": 2720,
  "title": "Problem parsing train.cvs with basic_benchmark.py ",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2720",
  "author_name": "",
  "post_date": "2012-09-20T04:58:29.657Z",
  "votes": null,
  "comment_count": 4,
  "views": 2024,
  "content": "<p>I'm having a trouble using the basic_benchmark.py with the trann.cvs file, the problem is that when I parse the cvs file some rows have nan values to the column &quot;BodyMarkdown&quot;.&nbsp;</p>\r\n<p>Some one had already this problem? I think this needs to be fixed and posted on github, mostly because new users will use this code.</p>\r\n<p>&nbsp;The problem is that one questions has no body text, but the real question on the site has some text. If you want to 'clean' the file here is my solution:</p>\r\n<p>&nbsp;</p>\r\n<pre>import csv<br>import sys<br>from datetime import datetime<br>import time<br><br>d = {}<br><br>def parse_line(PostId, PostCreationDate, OwnerUserId, OwnerCreationDate, ReputationAtPostCreation, OwnerUndeletedAnswerCountAtPostTime, Title, BodyMarkdown, Tag1, Tag2, Tag3, Tag4, Tag5, PostClosedDate, OpenStatus):<br>    print Title<br>    print BodyMarkdown<br>ifile = open('../train.csv', 'rt')<br>ofile = open('../clean-train.csv', 'wt')<br>f = csv.reader(ifile, delimiter=',')<br>w = csv.writer(ofile, delimiter=',', quotechar='&quot;')<br>for row in f:<br>    if (len(row) == 15):<br>        if (len(row[7]) == 0):<br>            print row<br>        else:<br>            w.writerow(row) <br>    else:<br>        print row<br>ifile.close()<br>ofile.close()</pre>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "14623",
      "postDate": "09/20/2012 04:58:29",
      "content": "<p>I'm having a trouble using the basic_benchmark.py with the trann.cvs file, the problem is that when I parse the cvs file some rows have nan values to the column &quot;BodyMarkdown&quot;.&nbsp;</p>\r\n<p>Some one had already this problem? I think this needs to be fixed and posted on github, mostly because new users will use this code.</p>\r\n<p>&nbsp;The problem is that one questions has no body text, but the real question on the site has some text. If you want to 'clean' the file here is my solution:</p>\r\n<p>&nbsp;</p>\r\n<pre>import csv<br>import sys<br>from datetime import datetime<br>import time<br><br>d = {}<br><br>def parse_line(PostId, PostCreationDate, OwnerUserId, OwnerCreationDate, ReputationAtPostCreation, OwnerUndeletedAnswerCountAtPostTime, Title, BodyMarkdown, Tag1, Tag2, Tag3, Tag4, Tag5, PostClosedDate, OpenStatus):<br>    print Title<br>    print BodyMarkdown<br>ifile = open('../train.csv', 'rt')<br>ofile = open('../clean-train.csv', 'wt')<br>f = csv.reader(ifile, delimiter=',')<br>w = csv.writer(ofile, delimiter=',', quotechar='&quot;')<br>for row in f:<br>    if (len(row) == 15):<br>        if (len(row[7]) == 0):<br>            print row<br>        else:<br>            w.writerow(row) <br>    else:<br>        print row<br>ifile.close()<br>ofile.close()</pre>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14735",
      "postDate": "09/24/2012 03:25:40",
      "content": "<p>I tried the script you provided, but still get</p>\r\n<p>['2967852', '06/03/2010 16:21:14', '357697', '06/03/2010 16:21:14', '1', '0', 'Working with NSNumberFormatter.', '', 'objective-c', 'nsnumberformatter', '', '', '', '', 'open']</p>\r\n<p>&nbsp;&nbsp;&nbsp; for row in f:<br>\r\n_csv.Error: newline inside string</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14821",
      "postDate": "09/25/2012 22:15:41",
      "content": "<p>The error happens whit my script or when you try to parse the train.cvs in another code?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14833",
      "postDate": "09/26/2012 13:39:03",
      "content": "<p>I didn't manage to run the basic benchmark because of some Pandas/Numpy problem.</p>\r\n<p>However in my own code, csv.reader parses the file without problems.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14843",
      "postDate": "09/26/2012 18:31:55",
      "content": "<p>[quote=Foxtrot;14833]</p>\r\n<p>I didn't manage to run the basic benchmark because of some Pandas/Numpy problem.</p>\r\n<p>However in my own code, csv.reader parses the file without problems.</p>\r\n<p>[/quote]</p>\r\n<p>You updated&nbsp;Numpy to install Pandas? In some older Ubuntu versions when you do that you broke some dependencies, I use Uubuntu 11.10 and 12.04 and works like a charm ;D</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 14735,
      "author_name": "darkoram",
      "author_url": "",
      "post_date": "09/24/2012 03:25:40",
      "content": "<p>I tried the script you provided, but still get</p>\r\n<p>['2967852', '06/03/2010 16:21:14', '357697', '06/03/2010 16:21:14', '1', '0', 'Working with NSNumberFormatter.', '', 'objective-c', 'nsnumberformatter', '', '', '', '', 'open']</p>\r\n<p>&nbsp;&nbsp;&nbsp; for row in f:<br>\r\n_csv.Error: newline inside string</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14821,
      "author_name": "alessandrosena",
      "author_url": "",
      "post_date": "09/25/2012 22:15:41",
      "content": "<p>The error happens whit my script or when you try to parse the train.cvs in another code?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14833,
      "author_name": "zygmunt",
      "author_url": "",
      "post_date": "09/26/2012 13:39:03",
      "content": "<p>I didn't manage to run the basic benchmark because of some Pandas/Numpy problem.</p>\r\n<p>However in my own code, csv.reader parses the file without problems.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14843,
      "author_name": "alessandrosena",
      "author_url": "",
      "post_date": "09/26/2012 18:31:55",
      "content": "<p>[quote=Foxtrot;14833]</p>\r\n<p>I didn't manage to run the basic benchmark because of some Pandas/Numpy problem.</p>\r\n<p>However in my own code, csv.reader parses the file without problems.</p>\r\n<p>[/quote]</p>\r\n<p>You updated&nbsp;Numpy to install Pandas? In some older Ubuntu versions when you do that you broke some dependencies, I use Uubuntu 11.10 and 12.04 and works like a charm ;D</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "14623": "",
    "14735": "",
    "14821": "",
    "14833": "",
    "14843": ""
  },
  "source": "meta"
}