{
  "id": 16030,
  "title": "Benchmark really slow. Is it normal?",
  "url": "/competitions/dato-native/discussion/16030",
  "author_name": "",
  "post_date": "2015-08-19T19:16:25.370Z",
  "votes": null,
  "comment_count": 2,
  "views": 758,
  "content": "<p>I'm trying to run <code>process_html.py</code> on the data, but at the current speed it would take ~8 hours. Now, I know that my computer is a bit old and that the parsing operation can be optimized. I am just wondering whether it's worth to try, given the circumstances. Could it be that with my hardware it is not possible to work with such a big dataset at all? How fast does it work on your pc?</p>\n\n<p>Other posts in the forum suggest that a decent speed would be 200k files processed in half an hour, whereas I'm stuck at 20k.</p>\n\n<p>If the specs are important, I have a 4GiB RAM, dual core @ 3 GHz.</p>",
  "messages": [
    {
      "id": "89837",
      "postDate": "08/19/2015 19:16:25",
      "content": "<p>I'm trying to run <code>process_html.py</code> on the data, but at the current speed it would take ~8 hours. Now, I know that my computer is a bit old and that the parsing operation can be optimized. I am just wondering whether it's worth to try, given the circumstances. Could it be that with my hardware it is not possible to work with such a big dataset at all? How fast does it work on your pc?</p>\n\n<p>Other posts in the forum suggest that a decent speed would be 200k files processed in half an hour, whereas I'm stuck at 20k.</p>\n\n<p>If the specs are important, I have a 4GiB RAM, dual core @ 3 GHz.</p>",
      "rawMarkdown": "I'm trying to run `process_html.py` on the data, but at the current speed it would take ~8 hours. Now, I know that my computer is a bit old and that the parsing operation can be optimized. I am just wondering whether it's worth to try, given the circumstances. Could it be that with my hardware it is not possible to work with such a big dataset at all? How fast does it work on your pc?\r\n\r\nOther posts in the forum suggest that a decent speed would be 200k files processed in half an hour, whereas I'm stuck at 20k.\r\n\r\nIf the specs are important, I have a 4GiB RAM, dual core @ 3 GHz.",
      "votes": null
    },
    {
      "id": "89980",
      "postDate": "08/21/2015 05:09:35",
      "content": "<p>Yes, should probably be refactored using multiprocessing.Pool. It seems 100% CPU bound.\nYou might also check which HTML/XML parser BeautifulSoup is using. Some are faster than others.</p>",
      "rawMarkdown": "Yes, should probably be refactored using multiprocessing.Pool. It seems 100% CPU bound.\r\nYou might also check which HTML/XML parser BeautifulSoup is using. Some are faster than others.",
      "votes": null
    },
    {
      "id": "89986",
      "postDate": "08/21/2015 06:43:34",
      "content": "<p>Try reworking the parsing function and use lxml instead of BeautifulSoup (this gave a massive speed boost). A full read (+parse + hash) takes about 25 minutes (SSD + i5).\nAlso, if RAM is your issue, try streaming the data instead of holding it all in memory.\nYou can do the parsing+hashing+modelling that way with less than 300MB memory.</p>",
      "rawMarkdown": "Try reworking the parsing function and use lxml instead of BeautifulSoup (this gave a massive speed boost). A full read (+parse + hash) takes about 25 minutes (SSD + i5).\r\nAlso, if RAM is your issue, try streaming the data instead of holding it all in memory.\r\nYou can do the parsing+hashing+modelling that way with less than 300MB memory.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 89980,
      "author_name": "iamcarbonatedmilk",
      "author_url": "",
      "post_date": "08/21/2015 05:09:35",
      "content": "<p>Yes, should probably be refactored using multiprocessing.Pool. It seems 100% CPU bound.\nYou might also check which HTML/XML parser BeautifulSoup is using. Some are faster than others.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 89986,
      "author_name": "yekinci",
      "author_url": "",
      "post_date": "08/21/2015 06:43:34",
      "content": "<p>Try reworking the parsing function and use lxml instead of BeautifulSoup (this gave a massive speed boost). A full read (+parse + hash) takes about 25 minutes (SSD + i5).\nAlso, if RAM is your issue, try streaming the data instead of holding it all in memory.\nYou can do the parsing+hashing+modelling that way with less than 300MB memory.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "89837": "I'm trying to run `process_html.py` on the data, but at the current speed it would take ~8 hours. Now, I know that my computer is a bit old and that the parsing operation can be optimized. I am just wondering whether it's worth to try, given the circumstances. Could it be that with my hardware it is not possible to work with such a big dataset at all? How fast does it work on your pc?\r\n\r\nOther posts in the forum suggest that a decent speed would be 200k files processed in half an hour, whereas I'm stuck at 20k.\r\n\r\nIf the specs are important, I have a 4GiB RAM, dual core @ 3 GHz.",
    "89980": "Yes, should probably be refactored using multiprocessing.Pool. It seems 100% CPU bound.\r\nYou might also check which HTML/XML parser BeautifulSoup is using. Some are faster than others.",
    "89986": "Try reworking the parsing function and use lxml instead of BeautifulSoup (this gave a massive speed boost). A full read (+parse + hash) takes about 25 minutes (SSD + i5).\r\nAlso, if RAM is your issue, try streaming the data instead of holding it all in memory.\r\nYou can do the parsing+hashing+modelling that way with less than 300MB memory."
  },
  "source": "meta"
}