{
  "id": 13219,
  "title": "Hardware configuration for this challenge",
  "url": "/competitions/malware-classification/discussion/13219",
  "author_name": "",
  "post_date": "2015-04-03T09:21:34.483Z",
  "votes": null,
  "comment_count": 4,
  "views": 1468,
  "content": "<p>Hello everyone,</p>\n<p>What are the hardware configuration that&nbsp;you use for this challenge ?</p>\n<p>I ask this question because I tried to extract features using n-grams of 4 bytes but my 8 gigs of RAM are not enough.</p>",
  "messages": [
    {
      "id": "69512",
      "postDate": "04/03/2015 09:21:34",
      "content": "<p>Hello everyone,</p>\n<p>What are the hardware configuration that&nbsp;you use for this challenge ?</p>\n<p>I ask this question because I tried to extract features using n-grams of 4 bytes but my 8 gigs of RAM are not enough.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69528",
      "postDate": "04/03/2015 14:24:11",
      "content": "<p>Hi Thomas,</p>\n<p>I think 8G RAM is quite enough for most of the computation. However, I think it also depends on the operating system and the hardware brand.</p>\n<p>I am using Macbook with 8G RAM and it is okay for handling 2-Grams.</p>\n<p>If you mean 4-Grams, I am sure it needs a lot of memory and also time. How do you want to train them ? Even feature selection is hardly possible.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69552",
      "postDate": "04/03/2015 18:51:39",
      "content": "<p>Hi,</p>\n\n<p>Yes, I mean 4-Grams. Using c++ and storing unique n-grams in an std::unordered_map, I use 4 GB of RAM to store 40M n-grams extracted from 300 files.</p>\n<p>I want to collect all unique n-gram from each file in each class and do the same as described in the following pdf file : http://www.jmlr.org/papers/volume7/kolter06a/kolter06a.pdf</p>\n\n<p>I tried another technique to classify malware without using n-grams, but I failed. I only get 25% accuracy. The idea is simple: I convert all .bytes files into png images, then I use gabor filter bank to extract features. Finally I train a k-nearest Neighbors or a random forest to classify unknown malwares. I think the problem comes from gabor filters. How can I diagnose these filters? How can I extract features from images?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69577",
      "postDate": "04/03/2015 23:17:18",
      "content": "<p>Hi Thomas_S -</p>\n<p>If it helps, I'm on a laptop and C++ as well, and approached the memory constraint by extracting 4-grams in chunks (e.g. do all 4-grams that start with 0x00xxxxxx, and save them to file for later processing, then do 0x01xxxxxx etc). Takes a bit longer since you do more loops, but you make up for it with the C++.</p>\n<p>Once you're done, you'll have a number of 4-gram files to use to then try various&nbsp;importance algorithm on</p>\n<p>ash</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71257",
      "postDate": "04/11/2015 01:05:05",
      "content": "<p>I was able to process the 4-grams using python and R on a laptop with 8 gig of ram, but I can't handle them with any memory-intensive fitting methods. I was, however, able to use pca or other pre-processing methods to pick out the important 4-grams and use them in my more memory-intensive fitting procedures.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 69528,
      "author_name": "mahmadi",
      "author_url": "",
      "post_date": "04/03/2015 14:24:11",
      "content": "<p>Hi Thomas,</p>\n<p>I think 8G RAM is quite enough for most of the computation. However, I think it also depends on the operating system and the hardware brand.</p>\n<p>I am using Macbook with 8G RAM and it is okay for handling 2-Grams.</p>\n<p>If you mean 4-Grams, I am sure it needs a lot of memory and also time. How do you want to train them ? Even feature selection is hardly possible.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69552,
      "author_name": "thomasseleck",
      "author_url": "",
      "post_date": "04/03/2015 18:51:39",
      "content": "<p>Hi,</p>\n\n<p>Yes, I mean 4-Grams. Using c++ and storing unique n-grams in an std::unordered_map, I use 4 GB of RAM to store 40M n-grams extracted from 300 files.</p>\n<p>I want to collect all unique n-gram from each file in each class and do the same as described in the following pdf file : http://www.jmlr.org/papers/volume7/kolter06a/kolter06a.pdf</p>\n\n<p>I tried another technique to classify malware without using n-grams, but I failed. I only get 25% accuracy. The idea is simple: I convert all .bytes files into png images, then I use gabor filter bank to extract features. Finally I train a k-nearest Neighbors or a random forest to classify unknown malwares. I think the problem comes from gabor filters. How can I diagnose these filters? How can I extract features from images?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69577,
      "author_name": "ashhafez",
      "author_url": "",
      "post_date": "04/03/2015 23:17:18",
      "content": "<p>Hi Thomas_S -</p>\n<p>If it helps, I'm on a laptop and C++ as well, and approached the memory constraint by extracting 4-grams in chunks (e.g. do all 4-grams that start with 0x00xxxxxx, and save them to file for later processing, then do 0x01xxxxxx etc). Takes a bit longer since you do more loops, but you make up for it with the C++.</p>\n<p>Once you're done, you'll have a number of 4-gram files to use to then try various&nbsp;importance algorithm on</p>\n<p>ash</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71257,
      "author_name": "thekannman",
      "author_url": "",
      "post_date": "04/11/2015 01:05:05",
      "content": "<p>I was able to process the 4-grams using python and R on a laptop with 8 gig of ram, but I can't handle them with any memory-intensive fitting methods. I was, however, able to use pca or other pre-processing methods to pick out the important 4-grams and use them in my more memory-intensive fitting procedures.&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "69512": "",
    "69528": "",
    "69552": "",
    "69577": "",
    "71257": ""
  },
  "source": "meta"
}