{
  "id": 6353,
  "title": "Someone else using (R)MySQL as Database? Need some Feedback",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6353",
  "author_name": "",
  "post_date": "2013-11-20T16:18:09.950Z",
  "votes": null,
  "comment_count": 3,
  "views": 1492,
  "content": "<p>Hey my friends,</p>\n<p>I'm just a little bit frustrated how &quot;<em>bad</em>&quot; my idea was to load the data into a MySQL database. But I also don't want to change this idea because I don't have that much experience about other comparable technologies. So it would be very helpful for me if someone else - using MySQL with the <strong>RMySQL</strong> package - could give me some advise.</p>\n<p>I've read already <strong>100k</strong> lines from the train file and it took me about <strong>32 minutes</strong> without any errors. I'm not sure that this is a good time for such less lines of code. So this way would take me about approx. 40 days to read the whole train file - without any errors.</p>\n<p>I don't want any sounrce code because I'm taking this as a challange but I also don't wanna give up because this project is pretty cool indeed. So just give me advise on how to configure my RMySQL packge, MySQL database or whatever. I already changed so many things in the MySQL.cfg like how many data can be load etc. It did not change anything on my computing time. Btw I'm not working on a Linux :D</p>\n<p>Edit: Maybe I need to ask you if it's the RMySQL package or my XAMPP(MySQL) that limits the conncetions/speed?</p>\n<p>Thanks for this great project!</p>\n<p>Cheers</p>",
  "messages": [
    {
      "id": "34964",
      "postDate": "11/20/2013 16:18:09",
      "content": "<p>Hey my friends,</p>\n<p>I'm just a little bit frustrated how &quot;<em>bad</em>&quot; my idea was to load the data into a MySQL database. But I also don't want to change this idea because I don't have that much experience about other comparable technologies. So it would be very helpful for me if someone else - using MySQL with the <strong>RMySQL</strong> package - could give me some advise.</p>\n<p>I've read already <strong>100k</strong> lines from the train file and it took me about <strong>32 minutes</strong> without any errors. I'm not sure that this is a good time for such less lines of code. So this way would take me about approx. 40 days to read the whole train file - without any errors.</p>\n<p>I don't want any sounrce code because I'm taking this as a challange but I also don't wanna give up because this project is pretty cool indeed. So just give me advise on how to configure my RMySQL packge, MySQL database or whatever. I already changed so many things in the MySQL.cfg like how many data can be load etc. It did not change anything on my computing time. Btw I'm not working on a Linux :D</p>\n<p>Edit: Maybe I need to ask you if it's the RMySQL package or my XAMPP(MySQL) that limits the conncetions/speed?</p>\n<p>Thanks for this great project!</p>\n<p>Cheers</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34966",
      "postDate": "11/20/2013 17:03:15",
      "content": "<p>Hi Curtis,</p>\n<p>I am not using R or MySQL but I had a false start with MongoDB/PyMongo that I abandoned due to the same issue you are having.&nbsp; Here are some ideas that I used to successfully process the files:</p>\n<p>1. There is a lot of data in those files.&nbsp; Figure out which pieces of data you are going to work with (at least for your first iteration).&nbsp; The write a separate parser to read the file, strip out the info that you need and write that to a separate file(s).&nbsp; First off,&nbsp; I wanted global clicks per URL and personalization data based on dwell time.&nbsp;&nbsp;</p>\n<p>2. Your parser should be written in a very fast language such as C. There may be some versions of R with a JIT compiler but I don't know for sure.&nbsp; It is likely the version of R you are using is a slow interpreted version. (my parser runs in about 10 minutes, my submission process takes about 2-3 hours).</p>\n<p>3. In your parser, you can read the train data in session chunks .&nbsp; Extract what you need and then throw that data away. Then read the next session's data etc.&nbsp; This cuts down tremendously on the RAM required to parse the file. &nbsp;</p>\n<p>3. The later stages of your solution could be written in R, since the amount data will be more manageable and R has great built in packages.</p>\n<p>4. If you really want to use a database, make sure it is up to the task.&nbsp; The transactions need to be be completed asynchronously, you don't want your write transactions to slow down your parser.&nbsp; Also, you can speed things up if you have two separate hard drives.&nbsp; Reading and writing to the same hard drive will really slow things down.&nbsp; Read from one drive, write to the other if you can.&nbsp;</p>\n<p>5. If you still want to use SQL, make sure you are using static not dynamic SQL.&nbsp; (BTW its been a while so the terminology may have changed).&nbsp; The &quot;access plan&quot; for the SQL statements needs to be bound to the database before your start, at compile time.&nbsp; If you write dynamic SQL, then each and every Insert statement needs to go through the access planning process during runtime, therefore much slower.</p>\n<p>6. If you can afford it an SSD will really speed things up.&nbsp; I initially considered upgrading my RAM, and do the whole thing in core, but I was out of memory slots which meant I needed to buy ALL new RAM.&nbsp; Not worth it, so I just got smarter in my parser.&nbsp;</p>\n<p>These ideas should get you started...</p>\n<p>Good Luck !</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34978",
      "postDate": "11/20/2013 23:01:15",
      "content": "<p>MySql can accept bulk inserts.&nbsp; Maybe that would minimize the traffic between the parser and the db.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "34996",
      "postDate": "11/21/2013 13:25:14",
      "content": "<p>[quote=zero zero;34978]</p>\n<p>MySql can accept bulk inserts.&nbsp; Maybe that would minimize the traffic between the parser and the db.</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>Sure but also especially this has been tested by myself. It works on a small set of about 10k rows. But working on a huge amount it just blows up the MySQL. I think I have tested many different ways and none of them worked well with MySQL.</p>\n<p>Edit: I think I just found a way which isn't that great but should work and will be executed in about 30 hours. Will let it run this weekend.</p>\n<p>@geringer: Well I like your post but it does not help in any way. I can't choose tools I have never tested or worked with before. Maybe they allow you to be faster but at which costs. I will give my MySQL another try. =)</p>\n<p>A great weekend to all of you!</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 34966,
      "author_name": "geringer",
      "author_url": "",
      "post_date": "11/20/2013 17:03:15",
      "content": "<p>Hi Curtis,</p>\n<p>I am not using R or MySQL but I had a false start with MongoDB/PyMongo that I abandoned due to the same issue you are having.&nbsp; Here are some ideas that I used to successfully process the files:</p>\n<p>1. There is a lot of data in those files.&nbsp; Figure out which pieces of data you are going to work with (at least for your first iteration).&nbsp; The write a separate parser to read the file, strip out the info that you need and write that to a separate file(s).&nbsp; First off,&nbsp; I wanted global clicks per URL and personalization data based on dwell time.&nbsp;&nbsp;</p>\n<p>2. Your parser should be written in a very fast language such as C. There may be some versions of R with a JIT compiler but I don't know for sure.&nbsp; It is likely the version of R you are using is a slow interpreted version. (my parser runs in about 10 minutes, my submission process takes about 2-3 hours).</p>\n<p>3. In your parser, you can read the train data in session chunks .&nbsp; Extract what you need and then throw that data away. Then read the next session's data etc.&nbsp; This cuts down tremendously on the RAM required to parse the file. &nbsp;</p>\n<p>3. The later stages of your solution could be written in R, since the amount data will be more manageable and R has great built in packages.</p>\n<p>4. If you really want to use a database, make sure it is up to the task.&nbsp; The transactions need to be be completed asynchronously, you don't want your write transactions to slow down your parser.&nbsp; Also, you can speed things up if you have two separate hard drives.&nbsp; Reading and writing to the same hard drive will really slow things down.&nbsp; Read from one drive, write to the other if you can.&nbsp;</p>\n<p>5. If you still want to use SQL, make sure you are using static not dynamic SQL.&nbsp; (BTW its been a while so the terminology may have changed).&nbsp; The &quot;access plan&quot; for the SQL statements needs to be bound to the database before your start, at compile time.&nbsp; If you write dynamic SQL, then each and every Insert statement needs to go through the access planning process during runtime, therefore much slower.</p>\n<p>6. If you can afford it an SSD will really speed things up.&nbsp; I initially considered upgrading my RAM, and do the whole thing in core, but I was out of memory slots which meant I needed to buy ALL new RAM.&nbsp; Not worth it, so I just got smarter in my parser.&nbsp;</p>\n<p>These ideas should get you started...</p>\n<p>Good Luck !</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34978,
      "author_name": "zerozero",
      "author_url": "",
      "post_date": "11/20/2013 23:01:15",
      "content": "<p>MySql can accept bulk inserts.&nbsp; Maybe that would minimize the traffic between the parser and the db.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 34996,
      "author_name": "curtis0",
      "author_url": "",
      "post_date": "11/21/2013 13:25:14",
      "content": "<p>[quote=zero zero;34978]</p>\n<p>MySql can accept bulk inserts.&nbsp; Maybe that would minimize the traffic between the parser and the db.</p>\n<p>[/quote]</p>\n<p>&nbsp;</p>\n<p>Sure but also especially this has been tested by myself. It works on a small set of about 10k rows. But working on a huge amount it just blows up the MySQL. I think I have tested many different ways and none of them worked well with MySQL.</p>\n<p>Edit: I think I just found a way which isn't that great but should work and will be executed in about 30 hours. Will let it run this weekend.</p>\n<p>@geringer: Well I like your post but it does not help in any way. I can't choose tools I have never tested or worked with before. Maybe they allow you to be faster but at which costs. I will give my MySQL another try. =)</p>\n<p>A great weekend to all of you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "34964": "",
    "34966": "",
    "34978": "",
    "34996": ""
  },
  "source": "meta"
}