{
  "id": 6411,
  "title": "Tips for Handling Large Data Files",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6411",
  "author_name": "",
  "post_date": "2013-11-25T17:04:26.873Z",
  "votes": 3,
  "comment_count": 3,
  "views": 2956,
  "content": "<p>&nbsp;</p>\n<p><em>I had previously posted some of this information in another thread, but I thought it deserved its own thread so anyone can find and also contribute to this information.</em> <br><br>For those of you still struggling to open and process the files for this competition, here are some ideas that I used to successfully process the files:<br><br>1. You can view the first 1000 lines of the file by using the linux command:&nbsp;&nbsp;</p>\n<p>&nbsp;&nbsp;&nbsp; head -n 1000 train &gt; trainhead.txt<br><br>2. There is a lot of data in those files.&nbsp; Figure out which pieces of data you are going to work with (at least for your first iteration).&nbsp; Then write a separate parser to read the file, strip out the info that you need and write that to a separate file(s).&nbsp; In my first pass, I wanted global clicks per URL and personalization data based on dwell time.&nbsp; You can always go back, re-parse and get more data as needed, such as Domain vs URL<br><br>3. Your parser should be written in a very fast compiled language such as C, C++ OR make sure you use a language with fast JIT compiler such as Java, Pypy, Julia, etc.. (I don't know if R does JIT) &nbsp; My parser runs in about 10 minutes, my submission process takes about 2-3 hours (which I want to rewrite to get down to about 10 minutes). &nbsp;<br><br>4. In your parser, you can read the train data in session chunks.&nbsp; Extract what you need, save that data away, delete that session's memory and then read the next session.&nbsp; This cuts down tremendously on the RAM required to parse the file.&nbsp; You should be able to parse the file with a 4GB machine.<br><br>5. When writing out the intermediate file, you can save a lot of RAM if instead of building up the entire output file in storage, just write it out to the outFile in chunks. More specifically, write your output lines to an array, when the array length hits 1000 rows of output, write it to the file, clear the array and continue with parsing.<br><br>6. If you really want to use a database, make sure it is up to the task.&nbsp; The transactions need to be be completed asynchronously, you don't want your write transactions to slow down your parser.&nbsp;</p>\n<p>7. You can speed things up if you have two separate (physical) hard drives.&nbsp; Reading and writing to the same hard drive will really slow things down.&nbsp; Read from one drive, write to the other if you can. <br><br>8. If your database supports bulk load try to use that.&nbsp; In this case, you would parse the file, generate the bulk load format and write that out to a flat file.&nbsp; Afterward, run the the bulk loader with this file. <br><br>9. If you plan to use SQL, make sure you are using static not dynamic SQL.&nbsp; (BTW its been a while so the terminology may have changed).&nbsp; The &quot;access plan&quot; for the SQL statements needs to be bound to the database before your start, at compile time.&nbsp; If you write dynamic SQL, then each and every Insert statement needs to go through the access planning process during runtime, therefore much slower.<br><br>10. If you can afford it an SSD will really speed things up.&nbsp; I initially considered upgrading my RAM, and do the whole thing in core, but I was out of memory slots which meant I needed to buy ALL new RAM.&nbsp; Not worth it, so I just got smarter in my parser. <br><br>Hopefully these ideas should get you started.&nbsp; On the other hand if what I am saying is totally foreign to you, you may not have sufficient computer science background to compete in this one.<br><br>Good Luck !</p>",
  "messages": [
    {
      "id": "35200",
      "postDate": "11/25/2013 17:04:26",
      "content": "<p>&nbsp;</p>\n<p><em>I had previously posted some of this information in another thread, but I thought it deserved its own thread so anyone can find and also contribute to this information.</em> <br><br>For those of you still struggling to open and process the files for this competition, here are some ideas that I used to successfully process the files:<br><br>1. You can view the first 1000 lines of the file by using the linux command:&nbsp;&nbsp;</p>\n<p>&nbsp;&nbsp;&nbsp; head -n 1000 train &gt; trainhead.txt<br><br>2. There is a lot of data in those files.&nbsp; Figure out which pieces of data you are going to work with (at least for your first iteration).&nbsp; Then write a separate parser to read the file, strip out the info that you need and write that to a separate file(s).&nbsp; In my first pass, I wanted global clicks per URL and personalization data based on dwell time.&nbsp; You can always go back, re-parse and get more data as needed, such as Domain vs URL<br><br>3. Your parser should be written in a very fast compiled language such as C, C++ OR make sure you use a language with fast JIT compiler such as Java, Pypy, Julia, etc.. (I don't know if R does JIT) &nbsp; My parser runs in about 10 minutes, my submission process takes about 2-3 hours (which I want to rewrite to get down to about 10 minutes). &nbsp;<br><br>4. In your parser, you can read the train data in session chunks.&nbsp; Extract what you need, save that data away, delete that session's memory and then read the next session.&nbsp; This cuts down tremendously on the RAM required to parse the file.&nbsp; You should be able to parse the file with a 4GB machine.<br><br>5. When writing out the intermediate file, you can save a lot of RAM if instead of building up the entire output file in storage, just write it out to the outFile in chunks. More specifically, write your output lines to an array, when the array length hits 1000 rows of output, write it to the file, clear the array and continue with parsing.<br><br>6. If you really want to use a database, make sure it is up to the task.&nbsp; The transactions need to be be completed asynchronously, you don't want your write transactions to slow down your parser.&nbsp;</p>\n<p>7. You can speed things up if you have two separate (physical) hard drives.&nbsp; Reading and writing to the same hard drive will really slow things down.&nbsp; Read from one drive, write to the other if you can. <br><br>8. If your database supports bulk load try to use that.&nbsp; In this case, you would parse the file, generate the bulk load format and write that out to a flat file.&nbsp; Afterward, run the the bulk loader with this file. <br><br>9. If you plan to use SQL, make sure you are using static not dynamic SQL.&nbsp; (BTW its been a while so the terminology may have changed).&nbsp; The &quot;access plan&quot; for the SQL statements needs to be bound to the database before your start, at compile time.&nbsp; If you write dynamic SQL, then each and every Insert statement needs to go through the access planning process during runtime, therefore much slower.<br><br>10. If you can afford it an SSD will really speed things up.&nbsp; I initially considered upgrading my RAM, and do the whole thing in core, but I was out of memory slots which meant I needed to buy ALL new RAM.&nbsp; Not worth it, so I just got smarter in my parser. <br><br>Hopefully these ideas should get you started.&nbsp; On the other hand if what I am saying is totally foreign to you, you may not have sufficient computer science background to compete in this one.<br><br>Good Luck !</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35209",
      "postDate": "11/25/2013 18:35:25",
      "content": "<p>Here is a blog post I found about speeding up R code. It may be as simple putting all your important code in a function and compiling that function:&nbsp;</p>\n<p>http://lookingatdata.blogspot.com/2012/04/speeding-up-r-computations-pt-ii.html</p>\n<p>&nbsp;</p>\n<p>Enjoy!</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35231",
      "postDate": "11/25/2013 21:27:39",
      "content": "<p>11. You can spend 1-2 evenings for implement binary serialization/deserialization. I use this approach and load train file for 3 minutes instead of 18 minutes (speedup 6 times). Additionally you deal with smaller file (7.52Gb vs 15.3Gb).</p>\n<p>12. Use a data properties. Note that datasets are sorted by USERID and Day fields - it would help.</p>\n<p>13. If you use C# or Python look toward &quot;yield&quot; keyword. You can read train and test files interlaced (remember tip 12?) and save a lot of memory.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35441",
      "postDate": "11/28/2013 13:16:25",
      "content": "<p>To read large files try reading them in chunks. Pandas.read_table has native support for chunking (setting chunk_size) which is useful if you want to chunk multi-line CSV's. If all data is structured to be on a single line (like in this Kaggle Challenge), you can also chunk with Python native code like so:</p>\n<p><code>i = 0<br>with open(input_file) as f:<br>&nbsp; for next_n_lines in izip_longest(*[f] * 1000): #1k chunk size<br>&nbsp; &nbsp; for line in list(next_n_lines):<br>&nbsp; &nbsp; &nbsp; i += 1</code></p>\n<p><span style=\"line-height: 1.4\">If you insert data into a database like sqlite3 you can use transactions or batch commits.&nbsp;</span></p>\n<p><code>if i % 5000000 == 0:<br>&nbsp; conn.commit()</code></p>\n<p>This commits the inserts at 5 million line count intervals. You can get 50k+ linereads, line manipulation and inserts per second this way.</p>\n<p>Once you have the data in a database look at optimizing your database with indexed columns so a select query will take a few milliseconds.</p>\n<p>This will allow you to make a submission in a reasonable time frame (around an hour) on a budget laptop with limited memory.</p>\n<p>I can second the OP's post that using SSD will speed-up this process a lot (by a factor 3-4) as write read is the bottleneck here.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 35209,
      "author_name": "geringer",
      "author_url": "",
      "post_date": "11/25/2013 18:35:25",
      "content": "<p>Here is a blog post I found about speeding up R code. It may be as simple putting all your important code in a function and compiling that function:&nbsp;</p>\n<p>http://lookingatdata.blogspot.com/2012/04/speeding-up-r-computations-pt-ii.html</p>\n<p>&nbsp;</p>\n<p>Enjoy!</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35231,
      "author_name": "radium1",
      "author_url": "",
      "post_date": "11/25/2013 21:27:39",
      "content": "<p>11. You can spend 1-2 evenings for implement binary serialization/deserialization. I use this approach and load train file for 3 minutes instead of 18 minutes (speedup 6 times). Additionally you deal with smaller file (7.52Gb vs 15.3Gb).</p>\n<p>12. Use a data properties. Note that datasets are sorted by USERID and Day fields - it would help.</p>\n<p>13. If you use C# or Python look toward &quot;yield&quot; keyword. You can read train and test files interlaced (remember tip 12?) and save a lot of memory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35441,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "11/28/2013 13:16:25",
      "content": "<p>To read large files try reading them in chunks. Pandas.read_table has native support for chunking (setting chunk_size) which is useful if you want to chunk multi-line CSV's. If all data is structured to be on a single line (like in this Kaggle Challenge), you can also chunk with Python native code like so:</p>\n<p><code>i = 0<br>with open(input_file) as f:<br>&nbsp; for next_n_lines in izip_longest(*[f] * 1000): #1k chunk size<br>&nbsp; &nbsp; for line in list(next_n_lines):<br>&nbsp; &nbsp; &nbsp; i += 1</code></p>\n<p><span style=\"line-height: 1.4\">If you insert data into a database like sqlite3 you can use transactions or batch commits.&nbsp;</span></p>\n<p><code>if i % 5000000 == 0:<br>&nbsp; conn.commit()</code></p>\n<p>This commits the inserts at 5 million line count intervals. You can get 50k+ linereads, line manipulation and inserts per second this way.</p>\n<p>Once you have the data in a database look at optimizing your database with indexed columns so a select query will take a few milliseconds.</p>\n<p>This will allow you to make a submission in a reasonable time frame (around an hour) on a budget laptop with limited memory.</p>\n<p>I can second the OP's post that using SSD will speed-up this process a lot (by a factor 3-4) as write read is the bottleneck here.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "35200": "",
    "35209": "",
    "35231": "",
    "35441": ""
  },
  "source": "meta"
}