{
  "id": 12644,
  "title": "Code for parallel feature extraction without unzipping the data",
  "url": "/competitions/malware-classification/discussion/12644",
  "author_name": "",
  "post_date": "2015-03-01T08:46:29.827Z",
  "votes": 13,
  "comment_count": 9,
  "views": 3631,
  "content": "<p>Hello, I've written a sample code in Python how to traverse .7z archive in parallel and extract byte counts.</p>\n<p>You will need&nbsp;<a href=\"https://github.com/libarchive/libarchive\">libarchive</a><strong>,</strong>&nbsp;compile it from source, do not use the version from ubuntu repo. And then install <a href=\"https://pypi.python.org/pypi/libarchive/0.4.3\">python wrapper for libarchive</a>&nbsp;and you are done if you use Anaconda.&nbsp;</p>\n<p>Hope it will help&nbsp;you to get to the top :)</p>",
  "messages": [
    {
      "id": "65172",
      "postDate": "03/01/2015 08:46:29",
      "content": "<p>Hello, I've written a sample code in Python how to traverse .7z archive in parallel and extract byte counts.</p>\n<p>You will need&nbsp;<a href=\"https://github.com/libarchive/libarchive\">libarchive</a><strong>,</strong>&nbsp;compile it from source, do not use the version from ubuntu repo. And then install <a href=\"https://pypi.python.org/pypi/libarchive/0.4.3\">python wrapper for libarchive</a>&nbsp;and you are done if you use Anaconda.&nbsp;</p>\n<p>Hope it will help&nbsp;you to get to the top :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65598",
      "postDate": "03/06/2015 12:23:35",
      "content": "<p>I wonder why there is no reply here .... I am anyways going to test your code, thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65600",
      "postDate": "03/06/2015 13:22:54",
      "content": "<p>Just one question, can you give me a clue of how long it did take you to run the code with only byte count as feature.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65671",
      "postDate": "03/07/2015 10:15:01",
      "content": "<p>[quote=AnotherFloyd;65600]</p>\n<p>Just one question, can you give me a clue of how long it did take you to run the code with only byte count as feature.</p>\n<p>[/quote]</p>\n<p>Have no idea to be true, I just wanted to show where you can insert your feature extractor. I have not run exactly this one for testing the performance and as long as you extract something simple unzipping the data is the bottleneck, so this code will not be 6 times faster&nbsp;for pool of 6 processes than without multiprocessing. But when doing regex multiprocessing helps a lot.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68787",
      "postDate": "03/29/2015 00:29:46",
      "content": "<p>Hi Dmitry,</p>\n<p>Could you please let me know how to properly install lib-archive on Windows. When ever i run your code I always get this error in python shell.</p>\n<p>Traceback (most recent call last):<br> File &quot;E:\\MMCL TT\\feat_extract.py&quot;, line 1, in &lt;module&gt;<br> import libarchive.public<br> File &quot;C:\\Python27\\lib\\site-packages\\libarchive\\public.py&quot;, line 1, in &lt;module&gt;<br> from libarchive.adapters.archive_read import \\<br> File &quot;C:\\Python27\\lib\\site-packages\\libarchive\\adapters\\archive_read.py&quot;, line 7, in &lt;module&gt;<br> import libarchive.calls.archive_read<br> File &quot;C:\\Python27\\lib\\site-packages\\libarchive\\calls\\archive_read.py&quot;, line 3, in &lt;module&gt;<br> from libarchive.library import libarchive<br> File &quot;C:\\Python27\\lib\\site-packages\\libarchive\\library.py&quot;, line 16, in &lt;module&gt;<br> libarchive = ctypes.cdll.LoadLibrary(_FILEPATH)<br> File &quot;C:\\Python27\\lib\\ctypes\\__init__.py&quot;, line 443, in LoadLibrary<br> return self._dlltype(name)<br> File &quot;C:\\Python27\\lib\\ctypes\\__init__.py&quot;, line 365, in __init__<br> self._handle = _dlopen(self._name, mode)<br>WindowsError: [Error 126] The specified module could not be found</p>\n\n<p>It would of great help if you can provide me a solution for this.</p>\n<p>Thanks for your code :)</p>\n<p>Kaushik</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68788",
      "postDate": "03/29/2015 00:31:43",
      "content": "<p>The error code is truncated.. I'm attaching it in a text file</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68820",
      "postDate": "03/29/2015 06:29:28",
      "content": "<p>[quote=Kaushik;68788]</p>\n<p>The error code is truncated.. I'm attaching it in a text file</p>\n<p>[/quote]</p>\n<p>I have no idea how to install libraries in Windows. It is much easier in linux that's why I use Ubuntu.&nbsp;</p>\n<p>I assume pylibarchive cannot find libarchive dll, there's probably a path variable in python for dll search dirs, you need to append your libarchive build dir there (only a guess).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68990",
      "postDate": "03/30/2015 09:21:19",
      "content": "<p>Hi Dmitry,&nbsp;</p>\n<p>Thank you for sharing the code. I tried to run it for train.7z but the code returns following exception:&nbsp;<br>&quot;ValueError: Archive iteration (read_next_header) returned error: (-30) [Damaged 7-Zip archive] &quot;<br><br>I am pretty sure that archive is fine and is not damaged (check sum is fine). &nbsp;BTW the code works fine for the sampleData.7z . Could you please give any hints on it?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69003",
      "postDate": "03/30/2015 12:42:59",
      "content": "<p>Hi,</p>\n<p>I faced with the same error. To be able to work with the archive without extracting all the data I developed an utility that converts 7z to ordinary zip archive with which it is possible to work with Python's zipfile module.</p>\n<p>This utility unlike other tools extracts and zips files one by one. Thus, you do not need ~200 Gb of free space on your hard drive. If you have this amount of space, it is better (faster) to use tools like &quot;atool&quot; to convert.&nbsp;</p>\n<p>The code of the utility is here:</p>\n<p><a href=\"https://github.com/zyrikby/7z2zip_converter/\">https://github.com/zyrikby/7z2zip_converter/</a></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69097",
      "postDate": "03/30/2015 22:13:53",
      "content": "<p>[quote=ITankoyeu;68990]</p>\n<p>Hi Dmitry,&nbsp;</p>\n<p>Thank you for sharing the code. I tried to run it for train.7z but the code returns following exception:&nbsp;<br>&quot;ValueError: Archive iteration (read_next_header) returned error: (-30) [Damaged 7-Zip archive] &quot;<br><br>I am pretty sure that archive is fine and is not damaged (check sum is fine). &nbsp;BTW the code works fine for the sampleData.7z . Could you please give any hints on it?</p>\n<p>[/quote]</p>\n<p>Yep, I faced a lot of strange libarchve exceptions, and this one too. Have no solution for this one, for me restarting the terminal with ipython notebook helped, but that is just magic.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 65598,
      "author_name": "shubhabrata",
      "author_url": "",
      "post_date": "03/06/2015 12:23:35",
      "content": "<p>I wonder why there is no reply here .... I am anyways going to test your code, thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65600,
      "author_name": "shubhabrata",
      "author_url": "",
      "post_date": "03/06/2015 13:22:54",
      "content": "<p>Just one question, can you give me a clue of how long it did take you to run the code with only byte count as feature.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65671,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "03/07/2015 10:15:01",
      "content": "<p>[quote=AnotherFloyd;65600]</p>\n<p>Just one question, can you give me a clue of how long it did take you to run the code with only byte count as feature.</p>\n<p>[/quote]</p>\n<p>Have no idea to be true, I just wanted to show where you can insert your feature extractor. I have not run exactly this one for testing the performance and as long as you extract something simple unzipping the data is the bottleneck, so this code will not be 6 times faster&nbsp;for pool of 6 processes than without multiprocessing. But when doing regex multiprocessing helps a lot.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68787,
      "author_name": "kaushik256147",
      "author_url": "",
      "post_date": "03/29/2015 00:29:46",
      "content": "<p>Hi Dmitry,</p>\n<p>Could you please let me know how to properly install lib-archive on Windows. When ever i run your code I always get this error in python shell.</p>\n<p>Traceback (most recent call last):<br> File &quot;E:\\MMCL TT\\feat_extract.py&quot;, line 1, in &lt;module&gt;<br> import libarchive.public<br> File &quot;C:\\Python27\\lib\\site-packages\\libarchive\\public.py&quot;, line 1, in &lt;module&gt;<br> from libarchive.adapters.archive_read import \\<br> File &quot;C:\\Python27\\lib\\site-packages\\libarchive\\adapters\\archive_read.py&quot;, line 7, in &lt;module&gt;<br> import libarchive.calls.archive_read<br> File &quot;C:\\Python27\\lib\\site-packages\\libarchive\\calls\\archive_read.py&quot;, line 3, in &lt;module&gt;<br> from libarchive.library import libarchive<br> File &quot;C:\\Python27\\lib\\site-packages\\libarchive\\library.py&quot;, line 16, in &lt;module&gt;<br> libarchive = ctypes.cdll.LoadLibrary(_FILEPATH)<br> File &quot;C:\\Python27\\lib\\ctypes\\__init__.py&quot;, line 443, in LoadLibrary<br> return self._dlltype(name)<br> File &quot;C:\\Python27\\lib\\ctypes\\__init__.py&quot;, line 365, in __init__<br> self._handle = _dlopen(self._name, mode)<br>WindowsError: [Error 126] The specified module could not be found</p>\n\n<p>It would of great help if you can provide me a solution for this.</p>\n<p>Thanks for your code :)</p>\n<p>Kaushik</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68788,
      "author_name": "kaushik256147",
      "author_url": "",
      "post_date": "03/29/2015 00:31:43",
      "content": "<p>The error code is truncated.. I'm attaching it in a text file</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68820,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "03/29/2015 06:29:28",
      "content": "<p>[quote=Kaushik;68788]</p>\n<p>The error code is truncated.. I'm attaching it in a text file</p>\n<p>[/quote]</p>\n<p>I have no idea how to install libraries in Windows. It is much easier in linux that's why I use Ubuntu.&nbsp;</p>\n<p>I assume pylibarchive cannot find libarchive dll, there's probably a path variable in python for dll search dirs, you need to append your libarchive build dir there (only a guess).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68990,
      "author_name": "itankoyeu",
      "author_url": "",
      "post_date": "03/30/2015 09:21:19",
      "content": "<p>Hi Dmitry,&nbsp;</p>\n<p>Thank you for sharing the code. I tried to run it for train.7z but the code returns following exception:&nbsp;<br>&quot;ValueError: Archive iteration (read_next_header) returned error: (-30) [Damaged 7-Zip archive] &quot;<br><br>I am pretty sure that archive is fine and is not damaged (check sum is fine). &nbsp;BTW the code works fine for the sampleData.7z . Could you please give any hints on it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69003,
      "author_name": "zyrikby",
      "author_url": "",
      "post_date": "03/30/2015 12:42:59",
      "content": "<p>Hi,</p>\n<p>I faced with the same error. To be able to work with the archive without extracting all the data I developed an utility that converts 7z to ordinary zip archive with which it is possible to work with Python's zipfile module.</p>\n<p>This utility unlike other tools extracts and zips files one by one. Thus, you do not need ~200 Gb of free space on your hard drive. If you have this amount of space, it is better (faster) to use tools like &quot;atool&quot; to convert.&nbsp;</p>\n<p>The code of the utility is here:</p>\n<p><a href=\"https://github.com/zyrikby/7z2zip_converter/\">https://github.com/zyrikby/7z2zip_converter/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69097,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "03/30/2015 22:13:53",
      "content": "<p>[quote=ITankoyeu;68990]</p>\n<p>Hi Dmitry,&nbsp;</p>\n<p>Thank you for sharing the code. I tried to run it for train.7z but the code returns following exception:&nbsp;<br>&quot;ValueError: Archive iteration (read_next_header) returned error: (-30) [Damaged 7-Zip archive] &quot;<br><br>I am pretty sure that archive is fine and is not damaged (check sum is fine). &nbsp;BTW the code works fine for the sampleData.7z . Could you please give any hints on it?</p>\n<p>[/quote]</p>\n<p>Yep, I faced a lot of strange libarchve exceptions, and this one too. Have no solution for this one, for me restarting the terminal with ipython notebook helped, but that is just magic.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "65172": "",
    "65598": "",
    "65600": "",
    "65671": "",
    "68787": "",
    "68788": "",
    "68820": "",
    "68990": "",
    "69003": "",
    "69097": ""
  },
  "source": "meta"
}