{
  "id": 13515,
  "title": "Brief Description of 10th Solution",
  "url": "/competitions/malware-classification/writeups/gilberto-titericz-junior-brief-description-of-10th",
  "author_name": "",
  "post_date": "2015-04-19T19:25:38.030Z",
  "votes": 11,
  "comment_count": 2,
  "views": 1376,
  "content": "<p>Unfortunately I didn't have much time for this competition, but here it is my solution. It can still have a LOT&nbsp;of improvements.</p>\n<p>Basically I created 4 features groups.</p>\n<p>1) - Convert each .bytes file into an image matrix and extracted 320 features acording SARVAM method. Then used kmeans and other clustering algorithms like Affinity propagation, Spectral clustering, etc, to extract groups with different number of classes of that 320 features. It generated me 10 categorical features. It don't have great performance, but together with other features can improve a little the final results.</p>\n<p>2) - Count of many manual choosen words from the .asm files by section (.header, .text, .data, etc), like: &quot;import, sub, loc, call, push, xor, add, jmp, security, dll, ret, unk, extrn, re, LZX, Decode, bf, obfuscator, rwe, COLLAPSED, dd, dw, db, nop, export, dllentry, reg, crypt, CWinApp, dllimport, Cookie, Thread, bb, jumptable, File, Delphi, Microsoft&quot;. In&nbsp;my tests that features have a good performance. It generated me&nbsp;407&nbsp;count features. interesting&nbsp;I found more that 1 file with word &quot;obfuscator&quot; inside, and it is an obfuscator.</p>\n<p>3) - Most interesting feature, mean line length by section (.asm files). I gives me 11 features from all sections&nbsp;(.header, .text, .data, etc). Performed good. Probably using stardard deviation of line length can also give good performance, but I didn't tested it. To build that features I suposed Mean Line Length by section could work as a malware signature.</p>\n<p>4) - Files sizes compressed and uncompressed and .bytes and .asm and first byte of each .bytes files. 5&nbsp;features.</p>\n<p>My last step was join all those features and run an multiclass random forest with many trees to find the feature importances. Then I just pick around 300 more important features and trainned it using crossvalidation with xgboost. High number of trees and low eta gives a more stable response in gbdt algorithms. My local CV was 0.0103.</p>\n<p>Then I aplied a trim technique to that results testing in my crossvalidated trainset. For each instance if the maximum prob was higher than 0.9998 I turn it to 1 and the other 8 classes to 0. If any class have prob below 0.0002 I set it to 0 and sum the prob value to the class with the high prob in that instance. It improved my CV trained dataset performance from 0.0103 to ~0.009. And private 0.00684.</p>\n<p>As the trim technique is very risk and only 1 mislabeled instance can destroy your score, I trusted in my local CV, so I choose those two submission as my final models. It seems my trimmed model worked well.</p>\n<p>Also I sent a lot of&nbsp;email (via Kaggle) to other competitors (more than&nbsp;10) to team up and nobody returned... anybody knows if that email system is working?</p>",
  "messages": [
    {
      "id": "72503",
      "postDate": "04/19/2015 19:25:38",
      "content": "<p>Unfortunately I didn't have much time for this competition, but here it is my solution. It can still have a LOT&nbsp;of improvements.</p>\n<p>Basically I created 4 features groups.</p>\n<p>1) - Convert each .bytes file into an image matrix and extracted 320 features acording SARVAM method. Then used kmeans and other clustering algorithms like Affinity propagation, Spectral clustering, etc, to extract groups with different number of classes of that 320 features. It generated me 10 categorical features. It don't have great performance, but together with other features can improve a little the final results.</p>\n<p>2) - Count of many manual choosen words from the .asm files by section (.header, .text, .data, etc), like: &quot;import, sub, loc, call, push, xor, add, jmp, security, dll, ret, unk, extrn, re, LZX, Decode, bf, obfuscator, rwe, COLLAPSED, dd, dw, db, nop, export, dllentry, reg, crypt, CWinApp, dllimport, Cookie, Thread, bb, jumptable, File, Delphi, Microsoft&quot;. In&nbsp;my tests that features have a good performance. It generated me&nbsp;407&nbsp;count features. interesting&nbsp;I found more that 1 file with word &quot;obfuscator&quot; inside, and it is an obfuscator.</p>\n<p>3) - Most interesting feature, mean line length by section (.asm files). I gives me 11 features from all sections&nbsp;(.header, .text, .data, etc). Performed good. Probably using stardard deviation of line length can also give good performance, but I didn't tested it. To build that features I suposed Mean Line Length by section could work as a malware signature.</p>\n<p>4) - Files sizes compressed and uncompressed and .bytes and .asm and first byte of each .bytes files. 5&nbsp;features.</p>\n<p>My last step was join all those features and run an multiclass random forest with many trees to find the feature importances. Then I just pick around 300 more important features and trainned it using crossvalidation with xgboost. High number of trees and low eta gives a more stable response in gbdt algorithms. My local CV was 0.0103.</p>\n<p>Then I aplied a trim technique to that results testing in my crossvalidated trainset. For each instance if the maximum prob was higher than 0.9998 I turn it to 1 and the other 8 classes to 0. If any class have prob below 0.0002 I set it to 0 and sum the prob value to the class with the high prob in that instance. It improved my CV trained dataset performance from 0.0103 to ~0.009. And private 0.00684.</p>\n<p>As the trim technique is very risk and only 1 mislabeled instance can destroy your score, I trusted in my local CV, so I choose those two submission as my final models. It seems my trimmed model worked well.</p>\n<p>Also I sent a lot of&nbsp;email (via Kaggle) to other competitors (more than&nbsp;10) to team up and nobody returned... anybody knows if that email system is working?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72717",
      "postDate": "04/20/2015 18:50:54",
      "content": "<p>Muito bom encontrar um Brasileiro aqui!</p>\n<p>Cara, sua solu&#231;&#227;o &#233; uma das mais robustas pelo que vi, com mais tunning provavelmente voce estaria no top 3!</p>\n<p>Quem sabe a gente n&#227;o team-up em competi&#231;&#245;es futuras?</p>\n<p>------</p>\n<p>It's nice to see a Brazilian here! Your solution is pretty robust, i think you could have gotten top 3 with more tunning!</p>\n<p>Maybe we could team-up in future comps?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72750",
      "postDate": "04/20/2015 21:25:45",
      "content": "<p>Can you give your code to extract the 320 features from images?&nbsp;I had the same idea but I failed to acheive it.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 72717,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "04/20/2015 18:50:54",
      "content": "<p>Muito bom encontrar um Brasileiro aqui!</p>\n<p>Cara, sua solu&#231;&#227;o &#233; uma das mais robustas pelo que vi, com mais tunning provavelmente voce estaria no top 3!</p>\n<p>Quem sabe a gente n&#227;o team-up em competi&#231;&#245;es futuras?</p>\n<p>------</p>\n<p>It's nice to see a Brazilian here! Your solution is pretty robust, i think you could have gotten top 3 with more tunning!</p>\n<p>Maybe we could team-up in future comps?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72750,
      "author_name": "thomasseleck",
      "author_url": "",
      "post_date": "04/20/2015 21:25:45",
      "content": "<p>Can you give your code to extract the 320 features from images?&nbsp;I had the same idea but I failed to acheive it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "72503": "",
    "72717": "",
    "72750": ""
  },
  "source": "meta"
}