{
  "id": 4392,
  "title": "extra_unsupervised_data.csv file",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4392",
  "author_name": "",
  "post_date": "2013-04-20T19:22:40.560Z",
  "votes": null,
  "comment_count": 7,
  "views": 2202,
  "content": "<p>The&nbsp;extra_unsupervised_data.csv file is large enough that I am not able to load it in my 12 GB of RAM (at least in R and Excel).</p>\r\n<p>My workaround would be to take a random subset of that file, like 50%, but I am not sure how to do it since I am not able to open it. &nbsp;Any ideas?</p>",
  "messages": [
    {
      "id": "23240",
      "postDate": "04/20/2013 19:22:40",
      "content": "<p>The&nbsp;extra_unsupervised_data.csv file is large enough that I am not able to load it in my 12 GB of RAM (at least in R and Excel).</p>\r\n<p>My workaround would be to take a random subset of that file, like 50%, but I am not sure how to do it since I am not able to open it. &nbsp;Any ideas?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23241",
      "postDate": "04/20/2013 19:30:39",
      "content": "<p>You could go through the file one line at a time. In R:</p>\r\n<p style=\"padding-left:30px\"><span><span class=\"x_il\">infile</span>&nbsp;&lt;-&nbsp;<span class=\"x_il\">file</span>(&quot;in.csv&quot;, open='r')</span></p>\r\n<p style=\"padding-left:30px\"><span>outfile &lt;-&nbsp;<span class=\"x_il\">file</span>(&quot;out.csv&quot;, open=&quot;w&quot;)</span></p>\r\n<p style=\"padding-left:30px\"><span>while (length(line &lt;- readLines(<span class=\"x_il\">fileConnection</span>, n=1)) &gt; 0) {</span></p>\r\n<p style=\"padding-left:30px\"><span>&nbsp; &nbsp; &nbsp; &nbsp; if (rbinom(1,1,0.2)==1) writeLines(line, outfile)</span></p>\r\n<p style=\"padding-left:30px\"><span>}</span></p>\r\n<p>This will make a new file with about 20% of the lines. (Oh, probably be more careful so you still have a header row.)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23242",
      "postDate": "04/20/2013 19:33:11",
      "content": "<p>Thanks a lot for the help!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23245",
      "postDate": "04/20/2013 19:51:34",
      "content": "<p>It should not take up that much memory. I don't know what R and Excel are doing, but here is the memory consumption for a few data structures:</p>\r\n<p>dense 32 bit matrix: 0.9 GB</p>\r\n<p>dense 64 bit matrix: 1.9 GB</p>\r\n<p>sparse matrix, using 64 bits to store the row, 64 bits to store the column, and 64 bits to store the value of each element: 5.7 GB</p>\r\n<p>To take up 12 GB, it'd have to somehow use an average of 400 bits per entry in the matrix.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23855",
      "postDate": "05/03/2013 04:11:54",
      "content": "<p>Ian is correct in that in R the data is exactly 1.9 GB (you can get that info with object.size)</p>\r\n<p>However R's read.csv performs terribly on large datasets and consumes a lot of unnecessary memory and time in the process.</p>\r\n<p>For a data set like this you can get much better results with 'scan':</p>\r\n<p>&nbsp; &nbsp; extra.data &lt;- scan(file=&quot;extra_unsupervised_data.csv&quot;,sep=',')</p>\r\n<p>And to make this data into a matrix you simply do this:</p>\r\n<p>&nbsp; &nbsp; extra.data.m&nbsp;&lt;- matrix(extra.data,ncol=1875)</p>\r\n<p>This should definitely work with 12GB of ram and very likely work with less. It's also a very useful trick for working with reasonably large data sets in R</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23883",
      "postDate": "05/04/2013 00:12:34",
      "content": "<p>I'm pretty sure that you'll want to include byrow=TRUE in your conversation to a matrix.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23884",
      "postDate": "05/04/2013 01:42:21",
      "content": "<p>A couple other options from here:</p>\r\n<p><a href=\"http://stackoverflow.com/questions/1727772/quickly-reading-very-large-tables-as-dataframes-in-r\">http://stackoverflow.com/questions/1727772/quickly-reading-very-large-tables-as-dataframes-in-r</a></p>\r\n<p>library(sqldf)<br>\r\nf &lt;- file('extra_unsupervised_data/extra_unsupervised_data.csv')<br>\r\nsystem.time(bigdf &lt;- sqldf(&quot;select * from f&quot;, dbname = tempfile(), file.format = list(header = T, row.names = F)))<br>\r\n<br>\r\n&nbsp; user&nbsp; system elapsed <br>\r\n320.92&nbsp;&nbsp; 28.30&nbsp; 350.48</p>\r\n<p>&gt; require(data.table)<br>\r\nLoading required package: data.table<br>\r\ndata.table 1.8.8&nbsp; For help type: help(&quot;data.table&quot;)<br>\r\n&gt; system.time(DT &lt;- fread('extra_unsupervised_data/extra_unsupervised_data.csv'))<br>\r\n&nbsp;&nbsp; user&nbsp; system elapsed <br>\r\n&nbsp;204.51&nbsp;&nbsp;&nbsp; 0.99&nbsp; 206.61</p>\r\n<p>Plus data.table gives you a progress indicator.</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23886",
      "postDate": "05/04/2013 04:22:03",
      "content": "<p>David normally you would be correct but it looks as though the .csv is actually in column order. Since it's so large I just did this quick python sanity check:</p>\r\n<pre>f = open('/extra_unsupervised_data.csv')<br><span style=\"line-height:1.4em\">b = f.readline()</span><br><span style=\"line-height:1.4em\">bs = b.split(',')</span><br><span style=\"line-height:1.4em\">len(bs)</span><br>&gt;&gt; 135735</pre>\r\n<pre>So each line in the .csv file seems to represent an entire column of data.</pre>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 23241,
      "author_name": "ajschumacher",
      "author_url": "",
      "post_date": "04/20/2013 19:30:39",
      "content": "<p>You could go through the file one line at a time. In R:</p>\r\n<p style=\"padding-left:30px\"><span><span class=\"x_il\">infile</span>&nbsp;&lt;-&nbsp;<span class=\"x_il\">file</span>(&quot;in.csv&quot;, open='r')</span></p>\r\n<p style=\"padding-left:30px\"><span>outfile &lt;-&nbsp;<span class=\"x_il\">file</span>(&quot;out.csv&quot;, open=&quot;w&quot;)</span></p>\r\n<p style=\"padding-left:30px\"><span>while (length(line &lt;- readLines(<span class=\"x_il\">fileConnection</span>, n=1)) &gt; 0) {</span></p>\r\n<p style=\"padding-left:30px\"><span>&nbsp; &nbsp; &nbsp; &nbsp; if (rbinom(1,1,0.2)==1) writeLines(line, outfile)</span></p>\r\n<p style=\"padding-left:30px\"><span>}</span></p>\r\n<p>This will make a new file with about 20% of the lines. (Oh, probably be more careful so you still have a header row.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23242,
      "author_name": "benoitplante",
      "author_url": "",
      "post_date": "04/20/2013 19:33:11",
      "content": "<p>Thanks a lot for the help!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23245,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/20/2013 19:51:34",
      "content": "<p>It should not take up that much memory. I don't know what R and Excel are doing, but here is the memory consumption for a few data structures:</p>\r\n<p>dense 32 bit matrix: 0.9 GB</p>\r\n<p>dense 64 bit matrix: 1.9 GB</p>\r\n<p>sparse matrix, using 64 bits to store the row, 64 bits to store the column, and 64 bits to store the value of each element: 5.7 GB</p>\r\n<p>To take up 12 GB, it'd have to somehow use an average of 400 bits per entry in the matrix.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23855,
      "author_name": "willkurt",
      "author_url": "",
      "post_date": "05/03/2013 04:11:54",
      "content": "<p>Ian is correct in that in R the data is exactly 1.9 GB (you can get that info with object.size)</p>\r\n<p>However R's read.csv performs terribly on large datasets and consumes a lot of unnecessary memory and time in the process.</p>\r\n<p>For a data set like this you can get much better results with 'scan':</p>\r\n<p>&nbsp; &nbsp; extra.data &lt;- scan(file=&quot;extra_unsupervised_data.csv&quot;,sep=',')</p>\r\n<p>And to make this data into a matrix you simply do this:</p>\r\n<p>&nbsp; &nbsp; extra.data.m&nbsp;&lt;- matrix(extra.data,ncol=1875)</p>\r\n<p>This should definitely work with 12GB of ram and very likely work with less. It's also a very useful trick for working with reasonably large data sets in R</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23883,
      "author_name": "dmcgarry",
      "author_url": "",
      "post_date": "05/04/2013 00:12:34",
      "content": "<p>I'm pretty sure that you'll want to include byrow=TRUE in your conversation to a matrix.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23884,
      "author_name": "barrenwuffet",
      "author_url": "",
      "post_date": "05/04/2013 01:42:21",
      "content": "<p>A couple other options from here:</p>\r\n<p><a href=\"http://stackoverflow.com/questions/1727772/quickly-reading-very-large-tables-as-dataframes-in-r\">http://stackoverflow.com/questions/1727772/quickly-reading-very-large-tables-as-dataframes-in-r</a></p>\r\n<p>library(sqldf)<br>\r\nf &lt;- file('extra_unsupervised_data/extra_unsupervised_data.csv')<br>\r\nsystem.time(bigdf &lt;- sqldf(&quot;select * from f&quot;, dbname = tempfile(), file.format = list(header = T, row.names = F)))<br>\r\n<br>\r\n&nbsp; user&nbsp; system elapsed <br>\r\n320.92&nbsp;&nbsp; 28.30&nbsp; 350.48</p>\r\n<p>&gt; require(data.table)<br>\r\nLoading required package: data.table<br>\r\ndata.table 1.8.8&nbsp; For help type: help(&quot;data.table&quot;)<br>\r\n&gt; system.time(DT &lt;- fread('extra_unsupervised_data/extra_unsupervised_data.csv'))<br>\r\n&nbsp;&nbsp; user&nbsp; system elapsed <br>\r\n&nbsp;204.51&nbsp;&nbsp;&nbsp; 0.99&nbsp; 206.61</p>\r\n<p>Plus data.table gives you a progress indicator.</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23886,
      "author_name": "willkurt",
      "author_url": "",
      "post_date": "05/04/2013 04:22:03",
      "content": "<p>David normally you would be correct but it looks as though the .csv is actually in column order. Since it's so large I just did this quick python sanity check:</p>\r\n<pre>f = open('/extra_unsupervised_data.csv')<br><span style=\"line-height:1.4em\">b = f.readline()</span><br><span style=\"line-height:1.4em\">bs = b.split(',')</span><br><span style=\"line-height:1.4em\">len(bs)</span><br>&gt;&gt; 135735</pre>\r\n<pre>So each line in the .csv file seems to represent an entire column of data.</pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "23240": "",
    "23241": "",
    "23242": "",
    "23245": "",
    "23855": "",
    "23883": "",
    "23884": "",
    "23886": ""
  },
  "source": "meta"
}