{
  "id": 12596,
  "title": "Data Representation... another naive question.",
  "url": "/competitions/malware-classification/discussion/12596",
  "author_name": "",
  "post_date": "2015-02-23T19:19:27.103Z",
  "votes": null,
  "comment_count": 4,
  "views": 1384,
  "content": "<p>Been trying to make this really simple for myself and ignore the size of the data....</p>\n<p>A Simple Feature:</p>\n<p>Unique string signatures</p>\n<p>256 possible character in a byte :)</p>\n<p>x</p>\n<p>xx</p>\n<p>xxx</p>\n<p>xxxx</p>\n<p>xxxxx</p>\n<p>1) Each file uniquely represents itself with a series of bytes the length of the file.</p>\n<p>2) Within the set&nbsp;that is all the files there are N sets of unique ordered byte strings (ignoring location for a moment)</p>\n<p>This set could be represented as a really really really big matrix whose location represents the byte string and</p>\n<p>hypothesis&nbsp;=&nbsp;number of times it occurs in the file / number of times it occurs in the set.</p>\n<p>-----</p>\n<p>Matrix 2: relative location of unique strings:</p>\n<p>there is another insanely big&nbsp;Matrix that could represent relative locations of the unique strings in the byte files. (not exactly sure how this could be applied in the real world but it would be meaningful within the data set.</p>\n<p>----</p>\n<p>ordered sets and location interact with each other in even bigger permutations</p>\n<p>---</p>\n<p>Training&nbsp;data could be averaged out to establish the hypothesis for each category:</p>\n<p>A*b = C</p>\n<p>Then with an infinite amount of compute time we could solve the equation via normalization OR given the things so darn big I could&nbsp;use some kind of learning algorithm &nbsp;to approximate a solution.</p>\n<p>Ignoring the details.. of mind numbing size and dimensionality of the matrices am I starting to think about this the right way?</p>\n<p>My head hurts.</p>\n<p>... Edit... Some kind of sparse matrix to get rid of the huge number of zeros.... &nbsp;</p>\n<p><span style=\"line-height: 1.4\">&nbsp;</span></p>\n<p><span style=\"line-height: 1.4\">.&nbsp;</span></p>",
  "messages": [
    {
      "id": "64778",
      "postDate": "02/23/2015 19:19:27",
      "content": "<p>Been trying to make this really simple for myself and ignore the size of the data....</p>\n<p>A Simple Feature:</p>\n<p>Unique string signatures</p>\n<p>256 possible character in a byte :)</p>\n<p>x</p>\n<p>xx</p>\n<p>xxx</p>\n<p>xxxx</p>\n<p>xxxxx</p>\n<p>1) Each file uniquely represents itself with a series of bytes the length of the file.</p>\n<p>2) Within the set&nbsp;that is all the files there are N sets of unique ordered byte strings (ignoring location for a moment)</p>\n<p>This set could be represented as a really really really big matrix whose location represents the byte string and</p>\n<p>hypothesis&nbsp;=&nbsp;number of times it occurs in the file / number of times it occurs in the set.</p>\n<p>-----</p>\n<p>Matrix 2: relative location of unique strings:</p>\n<p>there is another insanely big&nbsp;Matrix that could represent relative locations of the unique strings in the byte files. (not exactly sure how this could be applied in the real world but it would be meaningful within the data set.</p>\n<p>----</p>\n<p>ordered sets and location interact with each other in even bigger permutations</p>\n<p>---</p>\n<p>Training&nbsp;data could be averaged out to establish the hypothesis for each category:</p>\n<p>A*b = C</p>\n<p>Then with an infinite amount of compute time we could solve the equation via normalization OR given the things so darn big I could&nbsp;use some kind of learning algorithm &nbsp;to approximate a solution.</p>\n<p>Ignoring the details.. of mind numbing size and dimensionality of the matrices am I starting to think about this the right way?</p>\n<p>My head hurts.</p>\n<p>... Edit... Some kind of sparse matrix to get rid of the huge number of zeros.... &nbsp;</p>\n<p><span style=\"line-height: 1.4\">&nbsp;</span></p>\n<p><span style=\"line-height: 1.4\">.&nbsp;</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64789",
      "postDate": "02/23/2015 22:46:06",
      "content": "<p>Interesting ideas, but I would start with something much simpler. How about byte frequencies? There are only 256 different bytes, so for every file you can simply record how many times each byte appears. Ignore any location information. That way you can reduce the training/test set from half a terabyte to a few MB. Even with such a drastically reduced data set you can achieve a log loss of less than 0.2, I believe.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64792",
      "postDate": "02/23/2015 23:07:19",
      "content": "<p>[quote=David Tran;64789]</p>\n<p>Interesting ideas, but I would start with something much simpler. How about byte frequencies? There are only 256 different bytes, so for every file you can simply record how many times each byte appears. Ignore any location information. That way you can reduce the training/test set from half a terabyte to a few MB. Even with such a drastically reduced data set you can achieve a log loss of less than 0.2, I believe.</p>\n<p>[/quote]</p>\n<p>&lt; 0.02!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64793",
      "postDate": "02/23/2015 23:32:25",
      "content": "<p>The only problem I see with single byte frequency other than as a student exercise is that I cannot intuit that the results would have any meaningful value.&nbsp;&nbsp;&nbsp; Meaning&nbsp; of the data is in a sequence of instructions.&nbsp;&nbsp; Any single byte has no inherent meaning.&nbsp;&nbsp;&nbsp; Identifying a binary is going to make sense in terms of combination of functions/data used to leverage user rights,&nbsp; spam people, leverage ip ports...&nbsp; There should be common threads that can be recognized in the signal.&nbsp;&nbsp; Their location, relation to each other and incidence are the things that would provide nice signatures for viruses/trojans.&nbsp;&nbsp; Or at least that is my intuition....</p>\n\n<p>I'm confused why log loss is relevant here.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64796",
      "postDate": "02/24/2015 01:04:10",
      "content": "<p>The log loss is the metric by which the quality of your algorithm is evaluated:</p>\n<p><a href=\"http://www.kaggle.com/c/malware-classification/details/evaluation\">http://www.kaggle.com/c/malware-classification/details/evaluation</a></p>\n<p>The lower the log loss, the better your classification scheme. The fact that you can achieve a low log loss with just byte counts (for comparison, random guessing yields a loss of 2.2) proves that raw frequency distributions actually carry a lot of information about the type of malware. Individual bytes don't mean much, but their statistical distribution does. If you can add semantic information to that, great, but anything that carries information is valuable.</p>\n<p>If this is your first competition, simply processing the data set and generating a valid submission can be highly instructive. You typically start with a simple solution and then improve on it iteratively as you incorporate more advanced ideas.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 64789,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "02/23/2015 22:46:06",
      "content": "<p>Interesting ideas, but I would start with something much simpler. How about byte frequencies? There are only 256 different bytes, so for every file you can simply record how many times each byte appears. Ignore any location information. That way you can reduce the training/test set from half a terabyte to a few MB. Even with such a drastically reduced data set you can achieve a log loss of less than 0.2, I believe.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64792,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/23/2015 23:07:19",
      "content": "<p>[quote=David Tran;64789]</p>\n<p>Interesting ideas, but I would start with something much simpler. How about byte frequencies? There are only 256 different bytes, so for every file you can simply record how many times each byte appears. Ignore any location information. That way you can reduce the training/test set from half a terabyte to a few MB. Even with such a drastically reduced data set you can achieve a log loss of less than 0.2, I believe.</p>\n<p>[/quote]</p>\n<p>&lt; 0.02!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64793,
      "author_name": "robertfontaine",
      "author_url": "",
      "post_date": "02/23/2015 23:32:25",
      "content": "<p>The only problem I see with single byte frequency other than as a student exercise is that I cannot intuit that the results would have any meaningful value.&nbsp;&nbsp;&nbsp; Meaning&nbsp; of the data is in a sequence of instructions.&nbsp;&nbsp; Any single byte has no inherent meaning.&nbsp;&nbsp;&nbsp; Identifying a binary is going to make sense in terms of combination of functions/data used to leverage user rights,&nbsp; spam people, leverage ip ports...&nbsp; There should be common threads that can be recognized in the signal.&nbsp;&nbsp; Their location, relation to each other and incidence are the things that would provide nice signatures for viruses/trojans.&nbsp;&nbsp; Or at least that is my intuition....</p>\n\n<p>I'm confused why log loss is relevant here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64796,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "02/24/2015 01:04:10",
      "content": "<p>The log loss is the metric by which the quality of your algorithm is evaluated:</p>\n<p><a href=\"http://www.kaggle.com/c/malware-classification/details/evaluation\">http://www.kaggle.com/c/malware-classification/details/evaluation</a></p>\n<p>The lower the log loss, the better your classification scheme. The fact that you can achieve a low log loss with just byte counts (for comparison, random guessing yields a loss of 2.2) proves that raw frequency distributions actually carry a lot of information about the type of malware. Individual bytes don't mean much, but their statistical distribution does. If you can add semantic information to that, great, but anything that carries information is valuable.</p>\n<p>If this is your first competition, simply processing the data set and generating a valid submission can be highly instructive. You typically start with a simple solution and then improve on it iteratively as you incorporate more advanced ideas.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "64778": "",
    "64789": "",
    "64792": "",
    "64793": "",
    "64796": ""
  },
  "source": "meta"
}