{
  "id": 5178,
  "title": "Efficienct *.csv access (Python)",
  "url": "/competitions/belkin-energy-disaggregation-competition/discussion/5178",
  "author_name": "",
  "post_date": "2013-07-23T12:52:11.627Z",
  "votes": null,
  "comment_count": 7,
  "views": 2166,
  "content": "<p>I am curious if anyone has come across a better way to grab lines out of the middle of the HF.csv file. I did a test of different approaches to pull lines 30,017 - 30,020 as an arbitrary test and the best I've gotten so far is 13.9 seconds, not exactly what I'd call interactive.</p>\n<p>As a helpful hint, HF.csv is the only file where it's worth using seek to skip data in the file. It requires a little more analysis on line-widths, but it's better than the 23.4 seconds I got from the standard file iterator.</p>\n<p>- Jacob</p>",
  "messages": [
    {
      "id": "27545",
      "postDate": "07/23/2013 12:52:11",
      "content": "<p>I am curious if anyone has come across a better way to grab lines out of the middle of the HF.csv file. I did a test of different approaches to pull lines 30,017 - 30,020 as an arbitrary test and the best I've gotten so far is 13.9 seconds, not exactly what I'd call interactive.</p>\n<p>As a helpful hint, HF.csv is the only file where it's worth using seek to skip data in the file. It requires a little more analysis on line-widths, but it's better than the 23.4 seconds I got from the standard file iterator.</p>\n<p>- Jacob</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27550",
      "postDate": "07/23/2013 14:01:47",
      "content": "<p>Parse the data once and store it an a more suitable format. What language/tools are you using?</p>\n<p>--Beau</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27565",
      "postDate": "07/23/2013 19:46:35",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27567",
      "postDate": "07/23/2013 19:47:15",
      "content": "<p>As Beau suggested, read with a ODBC driver from a database if you can set it up quickly.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27570",
      "postDate": "07/23/2013 20:19:25",
      "content": "<p>For Python, pickle (cPickle) is easy and works with everything. If that's insufficient, have a look at pandas. It supports HDF5 which is very fast.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27571",
      "postDate": "07/23/2013 21:05:52",
      "content": "<p>grep? head? sed? http://stackoverflow.com/questions/191364/quick-unix-command-to-display-specific-lines-in-the-middle-of-a-file</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27590",
      "postDate": "07/24/2013 12:08:37",
      "content": "<p>As the title suggests, I'm doing everything from Python 2.7. I'm quite familiar with the Linux commands, but I think they would have to deal with the same problems of variable line widths (i.e. finding newlines). I had considered making a fixed-width version of the HF.csv files, but I think the right way to go is PyTables or Pandas. Both seem to have a straight-forward interface and if this <a href=\"https://ep2013.europython.eu/media/conference/slides/fast-data-mining-with-pytables-and-pandas.pdf\" target=\"_blank\">EuroPython presentation</a> is correct, it looks like I should use PyTables to do some initial processing out of the files (at least the HF.csv files) and Pandas for the heavier duty processing once I've narrowed down the data set.</p>\n<p>Since both have compression built in, then I can comment out my system calls to 7-zip :)</p>\n<p>It would have been nice if the MATLAB variables hadn't been nested under that Buffer structure (although I know it's easier that way in MATLAB), although I'm still not sure my PC could have decompressed the HF data by itself. Well if it was easy, it wouldn't be interesting!</p>\n<p>Just for reference, here are the numbers I got on my none-too-new PC for reading lines 30017-30020 out of the different *.csv files:</p>\n<p><code><span style=\"text-decoration: underline\">TimeTicksHF</span>&nbsp; <span style=\"text-decoration: underline\">LF1I</span>&nbsp;&nbsp;&nbsp;&nbsp; <span style=\"text-decoration: underline\">HF</span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <span style=\"text-decoration: underline\">Approach</span></code></p>\n<p><code>13.4 ms&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 66.9 ms&nbsp; 23.4 s&nbsp; for line in filepointer:<br></code></p>\n<p><code>6.0 ms&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 58.5 ms&nbsp; 25.5 s&nbsp; Iterating through the file using itertools.islice</code></p>\n<p><code>149.0 ms&nbsp;&nbsp;&nbsp;&nbsp; 1.24 s&nbsp;&nbsp; 13.9 s&nbsp; Using seek to skip as much data as possible and readlines to keep in sync<br></code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27602",
      "postDate": "07/24/2013 16:33:59",
      "content": "<p>Build yourself an index, storing the byte offsets of e.g. every 10th line.</p>\n<p>This takes fairly little extra memory, and allows fast seeking.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 27550,
      "author_name": "beaupiccart",
      "author_url": "",
      "post_date": "07/23/2013 14:01:47",
      "content": "<p>Parse the data once and store it an a more suitable format. What language/tools are you using?</p>\n<p>--Beau</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27565,
      "author_name": "",
      "author_url": "",
      "post_date": "07/23/2013 19:46:35",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 27567,
      "author_name": "",
      "author_url": "",
      "post_date": "07/23/2013 19:47:15",
      "content": "<p>As Beau suggested, read with a ODBC driver from a database if you can set it up quickly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27570,
      "author_name": "beaupiccart",
      "author_url": "",
      "post_date": "07/23/2013 20:19:25",
      "content": "<p>For Python, pickle (cPickle) is easy and works with everything. If that's insufficient, have a look at pandas. It supports HDF5 which is very fast.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27571,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "07/23/2013 21:05:52",
      "content": "<p>grep? head? sed? http://stackoverflow.com/questions/191364/quick-unix-command-to-display-specific-lines-in-the-middle-of-a-file</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27590,
      "author_name": "jacob10",
      "author_url": "",
      "post_date": "07/24/2013 12:08:37",
      "content": "<p>As the title suggests, I'm doing everything from Python 2.7. I'm quite familiar with the Linux commands, but I think they would have to deal with the same problems of variable line widths (i.e. finding newlines). I had considered making a fixed-width version of the HF.csv files, but I think the right way to go is PyTables or Pandas. Both seem to have a straight-forward interface and if this <a href=\"https://ep2013.europython.eu/media/conference/slides/fast-data-mining-with-pytables-and-pandas.pdf\" target=\"_blank\">EuroPython presentation</a> is correct, it looks like I should use PyTables to do some initial processing out of the files (at least the HF.csv files) and Pandas for the heavier duty processing once I've narrowed down the data set.</p>\n<p>Since both have compression built in, then I can comment out my system calls to 7-zip :)</p>\n<p>It would have been nice if the MATLAB variables hadn't been nested under that Buffer structure (although I know it's easier that way in MATLAB), although I'm still not sure my PC could have decompressed the HF data by itself. Well if it was easy, it wouldn't be interesting!</p>\n<p>Just for reference, here are the numbers I got on my none-too-new PC for reading lines 30017-30020 out of the different *.csv files:</p>\n<p><code><span style=\"text-decoration: underline\">TimeTicksHF</span>&nbsp; <span style=\"text-decoration: underline\">LF1I</span>&nbsp;&nbsp;&nbsp;&nbsp; <span style=\"text-decoration: underline\">HF</span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <span style=\"text-decoration: underline\">Approach</span></code></p>\n<p><code>13.4 ms&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 66.9 ms&nbsp; 23.4 s&nbsp; for line in filepointer:<br></code></p>\n<p><code>6.0 ms&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 58.5 ms&nbsp; 25.5 s&nbsp; Iterating through the file using itertools.islice</code></p>\n<p><code>149.0 ms&nbsp;&nbsp;&nbsp;&nbsp; 1.24 s&nbsp;&nbsp; 13.9 s&nbsp; Using seek to skip as much data as possible and readlines to keep in sync<br></code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27602,
      "author_name": "fuhang",
      "author_url": "",
      "post_date": "07/24/2013 16:33:59",
      "content": "<p>Build yourself an index, storing the byte offsets of e.g. every 10th line.</p>\n<p>This takes fairly little extra memory, and allows fast seeking.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "27545": "",
    "27550": "",
    "27565": "",
    "27567": "",
    "27570": "",
    "27571": "",
    "27590": "",
    "27602": ""
  },
  "source": "meta"
}