{
  "id": 12956,
  "title": "Some useful reference",
  "url": "/competitions/malware-classification/discussion/12956",
  "author_name": "",
  "post_date": "2015-03-21T11:13:51.837Z",
  "votes": 6,
  "comment_count": 10,
  "views": 3377,
  "content": "<p>Some papers I've found particularly relevant to this competition:</p>\n\n<p>Learning to detect malicious executables in the wild</p>\n<p>http://www.jmlr.org/papers/volume7/kolter06a/kolter06a.pdf</p>\n\n<p>Malware Analysis and Classification: A Survey</p>\n<p>Journal of Information Security, 2014, 5, 56-64 Published Online April 2014 in SciRes. http://www.scirp.org/journal/jis http://dx.doi.org/10.4236/jis.2014.52006</p>\n<p>Hopes this give some new ideas to try</p>",
  "messages": [
    {
      "id": "67515",
      "postDate": "03/21/2015 11:13:51",
      "content": "<p>Some papers I've found particularly relevant to this competition:</p>\n\n<p>Learning to detect malicious executables in the wild</p>\n<p>http://www.jmlr.org/papers/volume7/kolter06a/kolter06a.pdf</p>\n\n<p>Malware Analysis and Classification: A Survey</p>\n<p>Journal of Information Security, 2014, 5, 56-64 Published Online April 2014 in SciRes. http://www.scirp.org/journal/jis http://dx.doi.org/10.4236/jis.2014.52006</p>\n<p>Hopes this give some new ideas to try</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67606",
      "postDate": "03/22/2015 12:37:30",
      "content": "<p>Thanks !</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67629",
      "postDate": "03/22/2015 18:51:31",
      "content": "<p>Thank You :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67680",
      "postDate": "03/23/2015 05:18:51",
      "content": "<p>&quot;We then produced n-grams, by combining each four-byte sequence into a<br>single term. For instance, for the byte sequence ff 00 ab 3e 12 b3, the corresponding n-grams<br>would be ff00ab3e, 00ab3e12, and ab3e12b3. This processing resulted in 255;904;403 distinct<br>n-grams. One could also compute n-grams from words, something we explored and discuss further<br>in Section 5.2. Using the n-grams from all of the executables, we applied techniques from text<br>classification, which we discuss further in the next section.&quot;</p>\n\n<p>Just wondering whether anyone has tried this idea with generating 255M n-grams :-)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67700",
      "postDate": "03/23/2015 12:29:07",
      "content": "<p>[quote=Vinh Nguyen;67680]</p>\n<p>Just wondering whether anyone has tried this idea with generating 255M n-grams :-)</p>\n<p>[/quote]</p>\n<p>yes!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67779",
      "postDate": "03/23/2015 20:30:20",
      "content": "<p>[quote=Vinh Nguyen;67680]</p>\n<p>&quot;We then produced n-grams, by combining each four-byte sequence into a<br>single term. For instance, for the byte sequence ff 00 ab 3e 12 b3, the corresponding n-grams<br>would be ff00ab3e, 00ab3e12, and ab3e12b3. This processing resulted in 255;904;403 distinct<br>n-grams. One could also compute n-grams from words, something we explored and discuss further<br>in Section 5.2. Using the n-grams from all of the executables, we applied techniques from text<br>classification, which we discuss further in the next section.&quot;</p>\n<p>Just wondering whether anyone has tried this idea with generating 255M n-grams :-)</p>\n<p>[/quote]</p>\n<p>I don't get it...4 bytes = 32 bits should give 2^32, which is ~4.29 billion distinct 'words'.</p>\n<p>What i'm missing?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67787",
      "postDate": "03/23/2015 20:53:53",
      "content": "<p>[quote=Bats &amp; Robots;67779]</p>\n<p>I don't get it...4 bytes = 32 bits should give 2^32, which is ~4.29 billion distinct 'words'.</p>\n<p>What i'm missing?</p>\n<p>[/quote]</p>\n<p>I think they are looking a certain sequence of bytes, not all possibilities, which is 4 billion and you are right in that case.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67803",
      "postDate": "03/23/2015 22:22:20",
      "content": "<p>@rcarson</p>\n<p>Humm..now it makes sense, thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67857",
      "postDate": "03/24/2015 01:21:56",
      "content": "<p>[quote=Abhishek;67700]</p>\n<p>[quote=Vinh Nguyen;67680]</p>\n<p>Just wondering whether anyone has tried this idea with generating 255M n-grams :-)</p>\n<p>[/quote]</p>\n<p>yes!</p>\n<p>[/quote]</p>\n\n<p>interesting. Must be a very laborious job. I wonder how we can get rid of the line number (which breaks the continuity of the byte stream) so that &nbsp;off the shelf tools like grep can be used to count the n-gram</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68167",
      "postDate": "03/25/2015 06:47:31",
      "content": "<p>[quote=Vinh Nguyen;67857]</p>\n<p>interesting. Must be a very laborious job. I wonder how we can get rid of the line number (which breaks the continuity of the byte stream) so that &nbsp;off the shelf tools like grep can be used to count the n-gram</p>\n<p>[/quote]</p>\n\n<p>Here you go:&nbsp;</p>\n<p>&quot; &quot;.join([x[9:].strip() for x in open(byte_file).readlines()])</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68764",
      "postDate": "03/28/2015 22:06:37",
      "content": "<p>[quote=Vinh Nguyen;67515]</p>\n<p>Some papers I've found particularly relevant to this competition:</p>\n<p>Learning to detect malicious executables in the wild</p>\n<p>http://www.jmlr.org/papers/volume7/kolter06a/kolter06a.pdf</p>\n<p>Malware Analysis and Classification: A Survey</p>\n<p>Journal of Information Security, 2014, 5, 56-64 Published Online April 2014 in SciRes. http://www.scirp.org/journal/jis http://dx.doi.org/10.4236/jis.2014.52006</p>\n<p>Hopes this give some new ideas to try</p>\n<p>[/quote]</p>\n<p>Do you know of papers that have used the data in the .asm files?</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 67606,
      "author_name": "bamine",
      "author_url": "",
      "post_date": "03/22/2015 12:37:30",
      "content": "<p>Thanks !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67629,
      "author_name": "kaushik256147",
      "author_url": "",
      "post_date": "03/22/2015 18:51:31",
      "content": "<p>Thank You :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67680,
      "author_name": "vinhnguyen",
      "author_url": "",
      "post_date": "03/23/2015 05:18:51",
      "content": "<p>&quot;We then produced n-grams, by combining each four-byte sequence into a<br>single term. For instance, for the byte sequence ff 00 ab 3e 12 b3, the corresponding n-grams<br>would be ff00ab3e, 00ab3e12, and ab3e12b3. This processing resulted in 255;904;403 distinct<br>n-grams. One could also compute n-grams from words, something we explored and discuss further<br>in Section 5.2. Using the n-grams from all of the executables, we applied techniques from text<br>classification, which we discuss further in the next section.&quot;</p>\n\n<p>Just wondering whether anyone has tried this idea with generating 255M n-grams :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67700,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "03/23/2015 12:29:07",
      "content": "<p>[quote=Vinh Nguyen;67680]</p>\n<p>Just wondering whether anyone has tried this idea with generating 255M n-grams :-)</p>\n<p>[/quote]</p>\n<p>yes!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67779,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "03/23/2015 20:30:20",
      "content": "<p>[quote=Vinh Nguyen;67680]</p>\n<p>&quot;We then produced n-grams, by combining each four-byte sequence into a<br>single term. For instance, for the byte sequence ff 00 ab 3e 12 b3, the corresponding n-grams<br>would be ff00ab3e, 00ab3e12, and ab3e12b3. This processing resulted in 255;904;403 distinct<br>n-grams. One could also compute n-grams from words, something we explored and discuss further<br>in Section 5.2. Using the n-grams from all of the executables, we applied techniques from text<br>classification, which we discuss further in the next section.&quot;</p>\n<p>Just wondering whether anyone has tried this idea with generating 255M n-grams :-)</p>\n<p>[/quote]</p>\n<p>I don't get it...4 bytes = 32 bits should give 2^32, which is ~4.29 billion distinct 'words'.</p>\n<p>What i'm missing?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67787,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/23/2015 20:53:53",
      "content": "<p>[quote=Bats &amp; Robots;67779]</p>\n<p>I don't get it...4 bytes = 32 bits should give 2^32, which is ~4.29 billion distinct 'words'.</p>\n<p>What i'm missing?</p>\n<p>[/quote]</p>\n<p>I think they are looking a certain sequence of bytes, not all possibilities, which is 4 billion and you are right in that case.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67803,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "03/23/2015 22:22:20",
      "content": "<p>@rcarson</p>\n<p>Humm..now it makes sense, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67857,
      "author_name": "vinhnguyen",
      "author_url": "",
      "post_date": "03/24/2015 01:21:56",
      "content": "<p>[quote=Abhishek;67700]</p>\n<p>[quote=Vinh Nguyen;67680]</p>\n<p>Just wondering whether anyone has tried this idea with generating 255M n-grams :-)</p>\n<p>[/quote]</p>\n<p>yes!</p>\n<p>[/quote]</p>\n\n<p>interesting. Must be a very laborious job. I wonder how we can get rid of the line number (which breaks the continuity of the byte stream) so that &nbsp;off the shelf tools like grep can be used to count the n-gram</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68167,
      "author_name": "nthanhtam",
      "author_url": "",
      "post_date": "03/25/2015 06:47:31",
      "content": "<p>[quote=Vinh Nguyen;67857]</p>\n<p>interesting. Must be a very laborious job. I wonder how we can get rid of the line number (which breaks the continuity of the byte stream) so that &nbsp;off the shelf tools like grep can be used to count the n-gram</p>\n<p>[/quote]</p>\n\n<p>Here you go:&nbsp;</p>\n<p>&quot; &quot;.join([x[9:].strip() for x in open(byte_file).readlines()])</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68764,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "03/28/2015 22:06:37",
      "content": "<p>[quote=Vinh Nguyen;67515]</p>\n<p>Some papers I've found particularly relevant to this competition:</p>\n<p>Learning to detect malicious executables in the wild</p>\n<p>http://www.jmlr.org/papers/volume7/kolter06a/kolter06a.pdf</p>\n<p>Malware Analysis and Classification: A Survey</p>\n<p>Journal of Information Security, 2014, 5, 56-64 Published Online April 2014 in SciRes. http://www.scirp.org/journal/jis http://dx.doi.org/10.4236/jis.2014.52006</p>\n<p>Hopes this give some new ideas to try</p>\n<p>[/quote]</p>\n<p>Do you know of papers that have used the data in the .asm files?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "67515": "",
    "67606": "",
    "67629": "",
    "67680": "",
    "67700": "",
    "67779": "",
    "67787": "",
    "67803": "",
    "67857": "",
    "68167": "",
    "68764": ""
  },
  "source": "meta"
}