{
  "id": 13511,
  "title": "Brief Description of 3rd Solution",
  "url": "/competitions/malware-classification/writeups/mikhail-dmitry-stanislav-brief-description-of-3rd-",
  "author_name": "",
  "post_date": "2015-04-21T14:19:33.373Z",
  "votes": 11,
  "comment_count": 5,
  "views": 2495,
  "content": "<p>Hi everyone, thanks for this competition! Special&nbsp;thanks to my teammates - you are amazing!</p>\n<p>In this post I will briefly describe my&nbsp;features and semi-supervised trick.</p>\n<p>The features we used consists of ~200 features from me and 800 features from Dmitry. <br>My features are:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">count of lines for every sections ('.data', '.rsrc', '.text' and so on) in *.asm and entropy of this distribution. I found this kind of features are very useful!</span></li>\n<li><span style=\"line-height: 1.4\">asm instruction count, like &#8216;jmp&#8217;, 'mov', &#8230; I use 30 most frequent instructions.</span></li>\n<li><span style=\"line-height: 1.4\">sizes of *.asm and *.bytes and it&#8217;s ratio.</span></li>\n<li><span style=\"line-height: 1.4\">sys calls (grepped by &#8220;__stdcall&#8221; from *.asm) + TF-IDF + NMF(n_components=10).&nbsp;</span><span style=\"line-height: 1.4\">I</span>&nbsp;found it that TF-IDF works well and NMF performs better than PCA for dimentionallity reduction (at least for this task)</li>\n<li><span style=\"line-height: 1.4\">like previous, but for functions (grepped by &#8220;FUNCTION&#8221;)</span></li>\n<li><span style=\"line-height: 1.4\">4-gramms. We perform feature selection by freqeuncy analisys + LinearSVC(penalty=&#8217;l1&#8217;).transform() + RF.feature_importance(). It&#8217;s allow as to select ~130 4-gramms which solely gave 0.01193/0.00958 on public/private LB using single model!</span></li>\n<li><span style=\"line-height: 1.4\">10-gramms. The process was similar to 4-gramms extraction and selecting, but we use hashing and some magic(see below) on last stage. Finally we use only </span><strong style=\"line-height: 1.4\">14</strong><span style=\"line-height: 1.4\"> 10-gramms -- and they improved our scores!</span></li>\n</ol>\n<p>Using this features (without 10-gramms), xgboost and some kind of bagging-magic (thanks Stanislav for this!) we got score public/private: 0.00847 / 0.00559</p>\n<p>Similarly to <a href=\"http://www.kaggle.com/c/malware-classification/forums/t/13490/say-no-to-overfitting-approaches-sharing/72311#post72311\">LittleBoat&amp;Co's approach</a>, we&#8217;ve done semi-supervised trick, but instead choosing label by the max probability, we sampled labels based on test predictions, and averaged several runs. This trick improves our score from 0.0055 to 0.0045 on private LB (single model)</p>\n<p>We were able to go down from 0.0045 to 0.0041 on private LB by adding 14 10-gram feats. We selected them as follows: we generated out of fold predictions for train and sorted objects by true class probability, labeled as &#8216;1&#8217; top 100 worse, and run RF on 500 preselected 10-grams. We selected the most important features, and it turned out we can have 1.0 accuracy for this binary task just with 14 feats. Since the feats were naturally 10 gramms, the objects of class &#8216;1&#8217; were rather different, the number of feats is 10 times less than number of objects, we wanted to think that we will not overfit using these feats (but you will in general).</p>\n<p>Dmitry will describe his approach and some mixing stuff soon.</p>",
  "messages": [
    {
      "id": "72467",
      "postDate": "04/19/2015 16:48:54",
      "content": "<p>Hi everyone, thanks for this competition! Special&nbsp;thanks to my teammates - you are amazing!</p>\n<p>In this post I will briefly describe my&nbsp;features and semi-supervised trick.</p>\n<p>The features we used consists of ~200 features from me and 800 features from Dmitry. <br>My features are:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">count of lines for every sections ('.data', '.rsrc', '.text' and so on) in *.asm and entropy of this distribution. I found this kind of features are very useful!</span></li>\n<li><span style=\"line-height: 1.4\">asm instruction count, like &#8216;jmp&#8217;, 'mov', &#8230; I use 30 most frequent instructions.</span></li>\n<li><span style=\"line-height: 1.4\">sizes of *.asm and *.bytes and it&#8217;s ratio.</span></li>\n<li><span style=\"line-height: 1.4\">sys calls (grepped by &#8220;__stdcall&#8221; from *.asm) + TF-IDF + NMF(n_components=10).&nbsp;</span><span style=\"line-height: 1.4\">I</span>&nbsp;found it that TF-IDF works well and NMF performs better than PCA for dimentionallity reduction (at least for this task)</li>\n<li><span style=\"line-height: 1.4\">like previous, but for functions (grepped by &#8220;FUNCTION&#8221;)</span></li>\n<li><span style=\"line-height: 1.4\">4-gramms. We perform feature selection by freqeuncy analisys + LinearSVC(penalty=&#8217;l1&#8217;).transform() + RF.feature_importance(). It&#8217;s allow as to select ~130 4-gramms which solely gave 0.01193/0.00958 on public/private LB using single model!</span></li>\n<li><span style=\"line-height: 1.4\">10-gramms. The process was similar to 4-gramms extraction and selecting, but we use hashing and some magic(see below) on last stage. Finally we use only </span><strong style=\"line-height: 1.4\">14</strong><span style=\"line-height: 1.4\"> 10-gramms -- and they improved our scores!</span></li>\n</ol>\n<p>Using this features (without 10-gramms), xgboost and some kind of bagging-magic (thanks Stanislav for this!) we got score public/private: 0.00847 / 0.00559</p>\n<p>Similarly to <a href=\"http://www.kaggle.com/c/malware-classification/forums/t/13490/say-no-to-overfitting-approaches-sharing/72311#post72311\">LittleBoat&amp;Co's approach</a>, we&#8217;ve done semi-supervised trick, but instead choosing label by the max probability, we sampled labels based on test predictions, and averaged several runs. This trick improves our score from 0.0055 to 0.0045 on private LB (single model)</p>\n<p>We were able to go down from 0.0045 to 0.0041 on private LB by adding 14 10-gram feats. We selected them as follows: we generated out of fold predictions for train and sorted objects by true class probability, labeled as &#8216;1&#8217; top 100 worse, and run RF on 500 preselected 10-grams. We selected the most important features, and it turned out we can have 1.0 accuracy for this binary task just with 14 feats. Since the feats were naturally 10 gramms, the objects of class &#8216;1&#8217; were rather different, the number of feats is 10 times less than number of objects, we wanted to think that we will not overfit using these feats (but you will in general).</p>\n<p>Dmitry will describe his approach and some mixing stuff soon.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72468",
      "postDate": "04/19/2015 16:51:36",
      "content": "<p>Nice</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72470",
      "postDate": "04/19/2015 16:56:52",
      "content": "<p>Hello, it was amazing competition. Thanks our tuning master Stanislav for introducing incredible training techuniques for our classifiers.</p>\n<ol>\n<li><span style=\"line-height: 1.4\">I extracted byte counts, bigrams, i divided byte file into 4 parts and extracted byte counts from parts.&nbsp;</span></li>\n<li><span style=\"line-height: 1.4\">From .asm file I extracted counts for asm commands (about 100), registers, section names, segments, &#8216;_stdcall *&#8217; names, out types of stcalls &#8216;* __ stdcall&#8217;, &#8216;Import from *&#8217;, &#8216;extern *&#8217;, &#8216;loc_*&#8217;, &#8216;*.dll&#8217; (a lot of noise). Some keywords, that I saw in files like &#8216;microsoft&#8217;, &#8216;installdir&#8217;, &#8216;cookie&#8217;, &#8216;HKEY_LOCAL_MACHINE&#8217;, &#8216;sp-analysis failed&#8217; and so on. I aslo extracted reglike strings (e.g {0231-.. -1401}).</span></li>\n<li><span style=\"line-height: 1.4\">I extracted strings, but end up not using them directly. I just computed some statistics of strings length distribution, and it helped a lot.&nbsp;</span></li>\n<li><span style=\"line-height: 1.4\">Another good features I had was entropy features -- I computed it for sliding windows over the .bytes sequence and used some statistics of its distibution (quantiles, percentiles, mean, max, std; all of these with diff(entropy)), ended up with more than 200 entropy feats.&nbsp;</span></li>\n<li><span style=\"line-height: 1.4\">Line counts, size, etc.</span></li>\n</ol>\n<p>I did not want to use counts from 1) and any features Mikhail used directly, so I used NMF(X), NMF (log (X+1)) (they differ a lot!). All the counts from 2) were fed into linearSVC with L1 penalty in order to select features.</p>\n<p>All of the resulting features were used in one XGBoost model, averaged over several runs.</p>\n<p>Mixing: <br>We found, that our classifiers were making mistakes at different classes differently, so we mixed them for every class separately. Our overall mixing scheme was extremely sophisticated and we had mix of all our temporary solutions throughout&nbsp;the competition in the final submission. But all we did is per-class mixing and simple linear mixing with using 2d and 3rd order interaction between predictions.</p>\n<p>We were bold enough to trim the predictions to 0 and 1 for 1,3 and 7th classes. We got lucky and got 0.0041 -&gt; 0.0039 private score(but we had some thoughts why it is ok to do it with these classes). The second submission had no trimming, so we were expecting not to fall down too much.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72480",
      "postDate": "04/19/2015 17:07:40",
      "content": "<p>cool!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "73054",
      "postDate": "04/22/2015 00:35:05",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74130",
      "postDate": "04/22/2015 23:19:20",
      "content": "<p>We are going to share our code later</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 72468,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "04/19/2015 16:51:36",
      "content": "<p>Nice</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72470,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "04/19/2015 16:56:52",
      "content": "<p>Hello, it was amazing competition. Thanks our tuning master Stanislav for introducing incredible training techuniques for our classifiers.</p>\n<ol>\n<li><span style=\"line-height: 1.4\">I extracted byte counts, bigrams, i divided byte file into 4 parts and extracted byte counts from parts.&nbsp;</span></li>\n<li><span style=\"line-height: 1.4\">From .asm file I extracted counts for asm commands (about 100), registers, section names, segments, &#8216;_stdcall *&#8217; names, out types of stcalls &#8216;* __ stdcall&#8217;, &#8216;Import from *&#8217;, &#8216;extern *&#8217;, &#8216;loc_*&#8217;, &#8216;*.dll&#8217; (a lot of noise). Some keywords, that I saw in files like &#8216;microsoft&#8217;, &#8216;installdir&#8217;, &#8216;cookie&#8217;, &#8216;HKEY_LOCAL_MACHINE&#8217;, &#8216;sp-analysis failed&#8217; and so on. I aslo extracted reglike strings (e.g {0231-.. -1401}).</span></li>\n<li><span style=\"line-height: 1.4\">I extracted strings, but end up not using them directly. I just computed some statistics of strings length distribution, and it helped a lot.&nbsp;</span></li>\n<li><span style=\"line-height: 1.4\">Another good features I had was entropy features -- I computed it for sliding windows over the .bytes sequence and used some statistics of its distibution (quantiles, percentiles, mean, max, std; all of these with diff(entropy)), ended up with more than 200 entropy feats.&nbsp;</span></li>\n<li><span style=\"line-height: 1.4\">Line counts, size, etc.</span></li>\n</ol>\n<p>I did not want to use counts from 1) and any features Mikhail used directly, so I used NMF(X), NMF (log (X+1)) (they differ a lot!). All the counts from 2) were fed into linearSVC with L1 penalty in order to select features.</p>\n<p>All of the resulting features were used in one XGBoost model, averaged over several runs.</p>\n<p>Mixing: <br>We found, that our classifiers were making mistakes at different classes differently, so we mixed them for every class separately. Our overall mixing scheme was extremely sophisticated and we had mix of all our temporary solutions throughout&nbsp;the competition in the final submission. But all we did is per-class mixing and simple linear mixing with using 2d and 3rd order interaction between predictions.</p>\n<p>We were bold enough to trim the predictions to 0 and 1 for 1,3 and 7th classes. We got lucky and got 0.0041 -&gt; 0.0039 private score(but we had some thoughts why it is ok to do it with these classes). The second submission had no trimming, so we were expecting not to fall down too much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72480,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/19/2015 17:07:40",
      "content": "<p>cool!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 73054,
      "author_name": "soumyasmruti",
      "author_url": "",
      "post_date": "04/22/2015 00:35:05",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 74130,
      "author_name": "mikhailtrofimov",
      "author_url": "",
      "post_date": "04/22/2015 23:19:20",
      "content": "<p>We are going to share our code later</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "72467": "",
    "72468": "",
    "72470": "",
    "72480": "",
    "73054": "",
    "74130": ""
  },
  "source": "meta"
}