{
  "id": 15849,
  "title": "Amounts of ram",
  "url": "/competitions/dato-native/discussion/15849",
  "author_name": "",
  "post_date": "2015-08-08T15:24:54.717Z",
  "votes": 4,
  "comment_count": 6,
  "views": 1800,
  "content": "<p>Can anyone load the entire test bag of words dataset into main memory, or...?</p>\n\n<ul>\n<li>Are you predicting in batches.</li>\n<li>Are you predicting 1 observation at a time.</li>\n</ul>\n\n<p>8 GiBs don't seem to be cutting it for me this time :\\</p>",
  "messages": [
    {
      "id": "88903",
      "postDate": "08/08/2015 15:24:54",
      "content": "<p>Can anyone load the entire test bag of words dataset into main memory, or...?</p>\n\n<ul>\n<li>Are you predicting in batches.</li>\n<li>Are you predicting 1 observation at a time.</li>\n</ul>\n\n<p>8 GiBs don't seem to be cutting it for me this time :\\</p>",
      "rawMarkdown": "Can anyone load the entire test bag of words dataset into main memory, or...?\r\n\r\n- Are you predicting in batches.\r\n- Are you predicting 1 observation at a time.\r\n\r\n8 GiBs don't seem to be cutting it for me this time :\\",
      "votes": null
    },
    {
      "id": "88961",
      "postDate": "08/09/2015 16:41:29",
      "content": "<p>I stream the files from inside the archives with ZipFile in Python. Unfortunately no tfidf then, but still, a nice score for purely online learning.</p>\n\n<p>8GB is then enough to run through 35GB of unzipped data and predict 200k+ samples in about half an hour. Memory usage is negligible.</p>\n\n<pre><code>def yield_data_from_zip_by_id(archive_path,ids):\n    # ids is a dictionary with the labels for values. Use dummy labels for test.\n    # yields data, id and label of a sample.\n    archive = zipfile.ZipFile(archive_path, 'r')\n    file_paths = zipfile.ZipFile.namelist(archive)\n    for file_path in file_paths:\n            if file_path[2:] in ids:\n                    data = archive.read(file_path)\n                    yield data, file_path[2:], ids[file_path[2:]]\n\ndef yield_data_from_zip(archive_path):\n    archive = zipfile.ZipFile(archive_path, 'r')\n    file_paths = zipfile.ZipFile.namelist(archive)\n    for file_path in file_paths:\n            data = archive.read(file_path)\n            yield data, file_path[2:]\n</code></pre>",
      "rawMarkdown": "I stream the files from inside the archives with ZipFile in Python. Unfortunately no tfidf then, but still, a nice score for purely online learning.\r\n\r\n8GB is then enough to run through 35GB of unzipped data and predict 200k+ samples in about half an hour. Memory usage is negligible.\r\n\r\n    def yield_data_from_zip_by_id(archive_path,ids):\r\n        # ids is a dictionary with the labels for values. Use dummy labels for test.\r\n        # yields data, id and label of a sample.\r\n        archive = zipfile.ZipFile(archive_path, 'r')\r\n        file_paths = zipfile.ZipFile.namelist(archive)\r\n        for file_path in file_paths:\r\n                if file_path[2:] in ids:\r\n                        data = archive.read(file_path)\r\n                        yield data, file_path[2:], ids[file_path[2:]]\r\n\r\n    def yield_data_from_zip(archive_path):\r\n        archive = zipfile.ZipFile(archive_path, 'r')\r\n        file_paths = zipfile.ZipFile.namelist(archive)\r\n        for file_path in file_paths:\r\n                data = archive.read(file_path)\r\n                yield data, file_path[2:]",
      "votes": null
    },
    {
      "id": "88962",
      "postDate": "08/09/2015 17:52:14",
      "content": "<p>8GB is definitely enough for performing tfidf on training data. And I think it's not necessary to perform tfidf based on words in test data. So just devide test set into several parts, and then make prediction by part</p>",
      "rawMarkdown": "8GB is definitely enough for performing tfidf on training data. And I think it's not necessary to perform tfidf based on words in test data. So just devide test set into several parts, and then make prediction by part",
      "votes": null
    },
    {
      "id": "89023",
      "postDate": "08/10/2015 12:52:04",
      "content": "<p>What do you mean &quot;test bag of words?&quot;  I haven't even started working on this yet, but I suspect with a little pre-processing, I'll be able to fit my final (sparse) document-term matrix into memory, especially after dropping extremely rare terms.</p>",
      "rawMarkdown": "What do you mean \"test bag of words?\"  I haven't even started working on this yet, but I suspect with a little pre-processing, I'll be able to fit my final (sparse) document-term matrix into memory, especially after dropping extremely rare terms.",
      "votes": null
    },
    {
      "id": "89392",
      "postDate": "08/14/2015 19:44:35",
      "content": "<p>Depends how you parse your page. \nI Tried a brute force experiments to load everything in the memory and treating HTML as Text files. then i used Scikit-learn vectorizers (which uses SciPy sparse matrices which are memory friendly ) and it needed around 8-10 GB of Rams. \nTaking into consideration that i don't load all files in the memory, so only the vectors space is loaded.\nSo if you do basic pre-processing by removing all tags..etc or at least parsing for the webpages probably 8 gb will be more than enough.</p>",
      "rawMarkdown": "Depends how you parse your page. \r\nI Tried a brute force experiments to load everything in the memory and treating HTML as Text files. then i used Scikit-learn vectorizers (which uses SciPy sparse matrices which are memory friendly ) and it needed around 8-10 GB of Rams. \r\nTaking into consideration that i don't load all files in the memory, so only the vectors space is loaded.\r\nSo if you do basic pre-processing by removing all tags..etc or at least parsing for the webpages probably 8 gb will be more than enough.",
      "votes": null
    },
    {
      "id": "89894",
      "postDate": "08/20/2015 13:34:58",
      "content": "<p>I loaded all the dataset on memory, and it took 27.1 GB</p>",
      "rawMarkdown": "I loaded all the dataset on memory, and it took 27.1 GB",
      "votes": null
    },
    {
      "id": "89895",
      "postDate": "08/20/2015 13:48:18",
      "content": "<p>Takes me 1 GB!</p>",
      "rawMarkdown": "Takes me 1 GB!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 88961,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "08/09/2015 16:41:29",
      "content": "<p>I stream the files from inside the archives with ZipFile in Python. Unfortunately no tfidf then, but still, a nice score for purely online learning.</p>\n\n<p>8GB is then enough to run through 35GB of unzipped data and predict 200k+ samples in about half an hour. Memory usage is negligible.</p>\n\n<pre><code>def yield_data_from_zip_by_id(archive_path,ids):\n    # ids is a dictionary with the labels for values. Use dummy labels for test.\n    # yields data, id and label of a sample.\n    archive = zipfile.ZipFile(archive_path, 'r')\n    file_paths = zipfile.ZipFile.namelist(archive)\n    for file_path in file_paths:\n            if file_path[2:] in ids:\n                    data = archive.read(file_path)\n                    yield data, file_path[2:], ids[file_path[2:]]\n\ndef yield_data_from_zip(archive_path):\n    archive = zipfile.ZipFile(archive_path, 'r')\n    file_paths = zipfile.ZipFile.namelist(archive)\n    for file_path in file_paths:\n            data = archive.read(file_path)\n            yield data, file_path[2:]\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88962,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "08/09/2015 17:52:14",
      "content": "<p>8GB is definitely enough for performing tfidf on training data. And I think it's not necessary to perform tfidf based on words in test data. So just devide test set into several parts, and then make prediction by part</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 89023,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "08/10/2015 12:52:04",
      "content": "<p>What do you mean &quot;test bag of words?&quot;  I haven't even started working on this yet, but I suspect with a little pre-processing, I'll be able to fit my final (sparse) document-term matrix into memory, especially after dropping extremely rare terms.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 89392,
      "author_name": "hadyelsahar",
      "author_url": "",
      "post_date": "08/14/2015 19:44:35",
      "content": "<p>Depends how you parse your page. \nI Tried a brute force experiments to load everything in the memory and treating HTML as Text files. then i used Scikit-learn vectorizers (which uses SciPy sparse matrices which are memory friendly ) and it needed around 8-10 GB of Rams. \nTaking into consideration that i don't load all files in the memory, so only the vectors space is loaded.\nSo if you do basic pre-processing by removing all tags..etc or at least parsing for the webpages probably 8 gb will be more than enough.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 89894,
      "author_name": "insikk",
      "author_url": "",
      "post_date": "08/20/2015 13:34:58",
      "content": "<p>I loaded all the dataset on memory, and it took 27.1 GB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 89895,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "08/20/2015 13:48:18",
      "content": "<p>Takes me 1 GB!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "88903": "Can anyone load the entire test bag of words dataset into main memory, or...?\r\n\r\n- Are you predicting in batches.\r\n- Are you predicting 1 observation at a time.\r\n\r\n8 GiBs don't seem to be cutting it for me this time :\\",
    "88961": "I stream the files from inside the archives with ZipFile in Python. Unfortunately no tfidf then, but still, a nice score for purely online learning.\r\n\r\n8GB is then enough to run through 35GB of unzipped data and predict 200k+ samples in about half an hour. Memory usage is negligible.\r\n\r\n    def yield_data_from_zip_by_id(archive_path,ids):\r\n        # ids is a dictionary with the labels for values. Use dummy labels for test.\r\n        # yields data, id and label of a sample.\r\n        archive = zipfile.ZipFile(archive_path, 'r')\r\n        file_paths = zipfile.ZipFile.namelist(archive)\r\n        for file_path in file_paths:\r\n                if file_path[2:] in ids:\r\n                        data = archive.read(file_path)\r\n                        yield data, file_path[2:], ids[file_path[2:]]\r\n\r\n    def yield_data_from_zip(archive_path):\r\n        archive = zipfile.ZipFile(archive_path, 'r')\r\n        file_paths = zipfile.ZipFile.namelist(archive)\r\n        for file_path in file_paths:\r\n                data = archive.read(file_path)\r\n                yield data, file_path[2:]",
    "88962": "8GB is definitely enough for performing tfidf on training data. And I think it's not necessary to perform tfidf based on words in test data. So just devide test set into several parts, and then make prediction by part",
    "89023": "What do you mean \"test bag of words?\"  I haven't even started working on this yet, but I suspect with a little pre-processing, I'll be able to fit my final (sparse) document-term matrix into memory, especially after dropping extremely rare terms.",
    "89392": "Depends how you parse your page. \r\nI Tried a brute force experiments to load everything in the memory and treating HTML as Text files. then i used Scikit-learn vectorizers (which uses SciPy sparse matrices which are memory friendly ) and it needed around 8-10 GB of Rams. \r\nTaking into consideration that i don't load all files in the memory, so only the vectors space is loaded.\r\nSo if you do basic pre-processing by removing all tags..etc or at least parsing for the webpages probably 8 gb will be more than enough.",
    "89894": "I loaded all the dataset on memory, and it took 27.1 GB",
    "89895": "Takes me 1 GB!"
  },
  "source": "meta"
}