{
  "id": 2455,
  "title": "training-sample.csv Parsing Problems",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2455",
  "author_name": "",
  "post_date": "2012-08-27T13:08:00.807Z",
  "votes": null,
  "comment_count": 6,
  "views": 3243,
  "content": "<p>I've tried unsuccessfully importing in MySQL directly and also read/writing with OpenCSV in Java. Anyone else tried and succeeded? Should I just stop with those approaches and munge it with Python's csv module instead? My first two approaches both choke\r\n on BodyMarkdown. I'm working on OSX and ran the file through dos2unix.&nbsp;</p>",
  "messages": [
    {
      "id": "13503",
      "postDate": "08/27/2012 13:08:00",
      "content": "<p>I've tried unsuccessfully importing in MySQL directly and also read/writing with OpenCSV in Java. Anyone else tried and succeeded? Should I just stop with those approaches and munge it with Python's csv module instead? My first two approaches both choke\r\n on BodyMarkdown. I'm working on OSX and ran the file through dos2unix.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13504",
      "postDate": "08/27/2012 13:46:12",
      "content": "<p>I've got the same problem using OpenCSV. If I rig it to catch and ignore anything that doesn't parse properly I get:</p>\r\n<p>1,652,370 open questions<br>\r\n35,760 closed questions<br>\r\n33,216,433 unparseable lines</p>\r\n<p>Interestingly, it seems that OpenCSV has more problems handling the closed questions than the open ones. Perhaps a valid solution to the challenge is &quot;see if OpenCSV can parse the question's CSV data, if it can't then the question should be closed&quot;.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13506",
      "postDate": "08/27/2012 15:04:09",
      "content": "<p>Wow, that made me laugh inquisitively.&nbsp; I think that would conform to Norvig's approach more than Chomsky's.</p>\r\n<p>&nbsp;</p>\r\n<p>http://arnoldit.com/wordpress/2012/08/22/machine-learning-paragons-square-off/</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13555",
      "postDate": "08/28/2012 12:06:23",
      "content": "<p>If anyone else runs into this issue, I ended up writing a custom routine to parse the file. &nbsp;Here's the code:</p>\r\n<pre>//using a BufferedInputStream is strongly recommended\r\npublic static String[] readEntryFromStream(InputStream in) throws Exception {\r\n\tList buffer = new ArrayList();\r\n\t\r\n\tchar next = (char)in.read();\r\n\twhile (next == '\\r' || next == '\\n') {\r\n\t\t//strip leading newlines\r\n\t\tnext = (char)in.read();\r\n\t}\r\n\tif (next == Character.MAX_VALUE) {\r\n\t\t//EOF\r\n\t\treturn null;\r\n\t}\r\n\t\r\n\tboolean emptyFinalField = false;\r\n\twhile (next != '\\r' &amp;&amp; next != '\\n' &amp;&amp; next != Character.MAX_VALUE) {\r\n\t\temptyFinalField = false;\r\n\t\tStringBuffer current = new StringBuffer();\r\n\t\t\r\n\t\tif (next != '&quot;') {\r\n\t\t\t//simple case\r\n\t\t\tif (next != ',') {\r\n\t\t\t\tcurrent.append(next);\r\n\t\t\t}\r\n\t\t\twhile (next != ',' &amp;&amp; next != '\\n' &amp;&amp; next != '\\r') {\r\n\t\t\t\tnext = (char)in.read();\r\n\t\t\t\tif (next == Character.MAX_VALUE) {\r\n\t\t\t\t\tbreak;\r\n\t\t\t\t}\r\n\t\t\t\tif (next != ',' &amp;&amp; next != '\\n' &amp;&amp; next != '\\r'){\r\n\t\t\t\t\tcurrent.append(next);\r\n\t\t\t\t}\r\n\t\t\t}\r\n\t\t}\r\n\t\telse {\r\n\t\t\t//harder case\r\n\t\t\tboolean ignoresNextChar = false;\r\n\t\t\tboolean suspectEndOfSection = false;\r\n\t\t\twhile (next != Character.MAX_VALUE) {\r\n\t\t\t\tnext = (char)in.read();\r\n\t\t\t\tif (next == Character.MAX_VALUE) {\r\n\t\t\t\t\tbreak;\r\n\t\t\t\t}\r\n\t\t\t\t\r\n\t\t\t\tif ((next == ',') &amp;&amp; suspectEndOfSection) {\r\n\t\t\t\t\tbreak;//endOfSection = true;\r\n\t\t\t\t}\r\n\t\t\t\telse if (suspectEndOfSection) {\r\n\t\t\t\t\t//wasn't really the end\r\n\t\t\t\t\tcurrent.append('&quot;');\r\n\t\t\t\t\tsuspectEndOfSection = false;\r\n\t\t\t\t\tif (next == '&quot;') {\r\n\t\t\t\t\t\t//two quotes in a row; do not treat the current quote as a delimiting quote\r\n\t\t\t\t\t\tignoresNextChar = true;\r\n\t\t\t\t\t}\r\n\t\t\t\t}\r\n\t\t\t\t\r\n\t\t\t\tif (next == '&quot;' &amp;&amp; ! ignoresNextChar) {\r\n\t\t\t\t\tsuspectEndOfSection = true;\r\n\t\t\t\t}\r\n\t\t\t\telse if (ignoresNextChar) {\r\n\t\t\t\t\t//current.append(next);\r\n\t\t\t\t\tignoresNextChar = false;\r\n\t\t\t\t}\r\n\t\t\t\telse {\r\n\t\t\t\t\tcurrent.append(next);\r\n\t\t\t\t}\r\n\t\t\t}\r\n\t\t}\r\n\t\t\r\n\t\tbuffer.add(current.toString());\r\n\t\t//System.out.println(&quot;Ended iteration; parsed=&quot; &#43; current.toString() &#43; &quot;, next=&quot; &#43; next);\r\n\t\tif (next == ',') {\r\n\t\t\temptyFinalField = true;\r\n\t\t}\r\n\t\tnext = (char)in.read();\r\n\t}\r\n\t\r\n\tif (emptyFinalField) {\r\n\t\tbuffer.add(&quot;&quot;);\r\n\t}\r\n\t\r\n\t//Thread.sleep(5000);\r\n\treturn buffer.toArray(new String[]{});\r\n}</pre>\r\n<p>Not saying this is the best CSV parser, or even that it's technically to spec as far as CSV parsers go, but it parses the dataset.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13561",
      "postDate": "08/28/2012 14:22:13",
      "content": "<p>Since you mentioned python, it works without any issue using python csv.reader.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13797",
      "postDate": "09/02/2012 13:07:52",
      "content": "<p>For C# I use Lumenworks CSV Reader, free, opensource, works in Mono too. So far, no problems with it (or I don't know something). http://www.codeproject.com/Articles/9258/A-Fast-CSV-Reader</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14305",
      "postDate": "09/13/2012 21:44:22",
      "content": "<p>For JVM world one case try http://ostermiller.org/utils/CSV.html, it seems to be working.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13504,
      "author_name": "adamroth",
      "author_url": "",
      "post_date": "08/27/2012 13:46:12",
      "content": "<p>I've got the same problem using OpenCSV. If I rig it to catch and ignore anything that doesn't parse properly I get:</p>\r\n<p>1,652,370 open questions<br>\r\n35,760 closed questions<br>\r\n33,216,433 unparseable lines</p>\r\n<p>Interestingly, it seems that OpenCSV has more problems handling the closed questions than the open ones. Perhaps a valid solution to the challenge is &quot;see if OpenCSV can parse the question's CSV data, if it can't then the question should be closed&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13506,
      "author_name": "mcstar",
      "author_url": "",
      "post_date": "08/27/2012 15:04:09",
      "content": "<p>Wow, that made me laugh inquisitively.&nbsp; I think that would conform to Norvig's approach more than Chomsky's.</p>\r\n<p>&nbsp;</p>\r\n<p>http://arnoldit.com/wordpress/2012/08/22/machine-learning-paragons-square-off/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13555,
      "author_name": "adamroth",
      "author_url": "",
      "post_date": "08/28/2012 12:06:23",
      "content": "<p>If anyone else runs into this issue, I ended up writing a custom routine to parse the file. &nbsp;Here's the code:</p>\r\n<pre>//using a BufferedInputStream is strongly recommended\r\npublic static String[] readEntryFromStream(InputStream in) throws Exception {\r\n\tList buffer = new ArrayList();\r\n\t\r\n\tchar next = (char)in.read();\r\n\twhile (next == '\\r' || next == '\\n') {\r\n\t\t//strip leading newlines\r\n\t\tnext = (char)in.read();\r\n\t}\r\n\tif (next == Character.MAX_VALUE) {\r\n\t\t//EOF\r\n\t\treturn null;\r\n\t}\r\n\t\r\n\tboolean emptyFinalField = false;\r\n\twhile (next != '\\r' &amp;&amp; next != '\\n' &amp;&amp; next != Character.MAX_VALUE) {\r\n\t\temptyFinalField = false;\r\n\t\tStringBuffer current = new StringBuffer();\r\n\t\t\r\n\t\tif (next != '&quot;') {\r\n\t\t\t//simple case\r\n\t\t\tif (next != ',') {\r\n\t\t\t\tcurrent.append(next);\r\n\t\t\t}\r\n\t\t\twhile (next != ',' &amp;&amp; next != '\\n' &amp;&amp; next != '\\r') {\r\n\t\t\t\tnext = (char)in.read();\r\n\t\t\t\tif (next == Character.MAX_VALUE) {\r\n\t\t\t\t\tbreak;\r\n\t\t\t\t}\r\n\t\t\t\tif (next != ',' &amp;&amp; next != '\\n' &amp;&amp; next != '\\r'){\r\n\t\t\t\t\tcurrent.append(next);\r\n\t\t\t\t}\r\n\t\t\t}\r\n\t\t}\r\n\t\telse {\r\n\t\t\t//harder case\r\n\t\t\tboolean ignoresNextChar = false;\r\n\t\t\tboolean suspectEndOfSection = false;\r\n\t\t\twhile (next != Character.MAX_VALUE) {\r\n\t\t\t\tnext = (char)in.read();\r\n\t\t\t\tif (next == Character.MAX_VALUE) {\r\n\t\t\t\t\tbreak;\r\n\t\t\t\t}\r\n\t\t\t\t\r\n\t\t\t\tif ((next == ',') &amp;&amp; suspectEndOfSection) {\r\n\t\t\t\t\tbreak;//endOfSection = true;\r\n\t\t\t\t}\r\n\t\t\t\telse if (suspectEndOfSection) {\r\n\t\t\t\t\t//wasn't really the end\r\n\t\t\t\t\tcurrent.append('&quot;');\r\n\t\t\t\t\tsuspectEndOfSection = false;\r\n\t\t\t\t\tif (next == '&quot;') {\r\n\t\t\t\t\t\t//two quotes in a row; do not treat the current quote as a delimiting quote\r\n\t\t\t\t\t\tignoresNextChar = true;\r\n\t\t\t\t\t}\r\n\t\t\t\t}\r\n\t\t\t\t\r\n\t\t\t\tif (next == '&quot;' &amp;&amp; ! ignoresNextChar) {\r\n\t\t\t\t\tsuspectEndOfSection = true;\r\n\t\t\t\t}\r\n\t\t\t\telse if (ignoresNextChar) {\r\n\t\t\t\t\t//current.append(next);\r\n\t\t\t\t\tignoresNextChar = false;\r\n\t\t\t\t}\r\n\t\t\t\telse {\r\n\t\t\t\t\tcurrent.append(next);\r\n\t\t\t\t}\r\n\t\t\t}\r\n\t\t}\r\n\t\t\r\n\t\tbuffer.add(current.toString());\r\n\t\t//System.out.println(&quot;Ended iteration; parsed=&quot; &#43; current.toString() &#43; &quot;, next=&quot; &#43; next);\r\n\t\tif (next == ',') {\r\n\t\t\temptyFinalField = true;\r\n\t\t}\r\n\t\tnext = (char)in.read();\r\n\t}\r\n\t\r\n\tif (emptyFinalField) {\r\n\t\tbuffer.add(&quot;&quot;);\r\n\t}\r\n\t\r\n\t//Thread.sleep(5000);\r\n\treturn buffer.toArray(new String[]{});\r\n}</pre>\r\n<p>Not saying this is the best CSV parser, or even that it's technically to spec as far as CSV parsers go, but it parses the dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13561,
      "author_name": "hackinghabits",
      "author_url": "",
      "post_date": "08/28/2012 14:22:13",
      "content": "<p>Since you mentioned python, it works without any issue using python csv.reader.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13797,
      "author_name": "macias",
      "author_url": "",
      "post_date": "09/02/2012 13:07:52",
      "content": "<p>For C# I use Lumenworks CSV Reader, free, opensource, works in Mono too. So far, no problems with it (or I don't know something). http://www.codeproject.com/Articles/9258/A-Fast-CSV-Reader</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14305,
      "author_name": "macias",
      "author_url": "",
      "post_date": "09/13/2012 21:44:22",
      "content": "<p>For JVM world one case try http://ostermiller.org/utils/CSV.html, it seems to be working.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13503": "",
    "13504": "",
    "13506": "",
    "13555": "",
    "13561": "",
    "13797": "",
    "14305": ""
  },
  "source": "meta"
}