{
  "id": 13509,
  "title": "Brief Description of 7th-Place Solution",
  "url": "/competitions/malware-classification/writeups/silogram-brief-description-of-7th-place-solution",
  "author_name": "",
  "post_date": "2015-04-19T15:53:14.903Z",
  "votes": 30,
  "comment_count": 12,
  "views": 3602,
  "content": "<p>First, congratulations to the winners. This was a fun competition with an interesting dataset and great communication in the forums. I&#8217;m posting a description of my solution because it seems to use many fewer features (25) than the winning solutions and includes an interesting calibration technique for RF prediction probabilities.</p>\n<p><strong>General Approach</strong></p>\n<p>1. Generate features from the .asm and .byte files.&nbsp;<span style=\"line-height: 1.4\">Features included:</span></p>\n<p style=\"padding-left: 30px\">&#8226; General characteristics, such as the length of the files, start address of the asm file, etc.</p>\n<p style=\"padding-left: 30px\">&#8226; Counts of various objects such as hex codes, hex 2grams, assembly language commands, etc.</p>\n<p style=\"padding-left: 30px\">&#8226; Hashes of specific parts of the asm file, such as each of the first 50 lines and specific properties of the asm file (e.g., segment type, segment permissions, etc.)</p>\n<p>2. Reduce the feature set (which started at about 2000 features) with greedy feature selection. CV was done with 10-fold cross-validation repeated 10 times with re-randomized folds, so 100 folds altogether.</p>\n<p>3. Generate prediction probabilities with random forest (RF) algorithm (sklearn implementation, estimators ranging from 500 &#8211; 1,500)</p>\n<p>4. Arguably the most important step, calibrate the prediction probabilities by raising them to a power selected via grid search (actually 9 different powers, one for each class). See below for details of this technique.</p>\n<p><strong>Performance </strong></p>\n<p>Public LB Performance was consistently a little better than CV performance but almost always moved in the same direction. The final CV and LB log losses were .005 and .004, respectively. The Private LB result was .006.</p>\n<p><strong>Model </strong></p>\n<p>The final solution was a single RF model consisting of the mean of 10 forests. Each forest consisted of 1,500 trees. The purpose of producing 10 different forests was to randomly change the hash seed for each forest to discourage RF from treating the hash features as numerical values. In approximate order of importance, the features were:</p>\n<p>- hash of 11th line of asm file (this is the line is most of the asm files that identifies the target processor)</p>\n<p>- length of HEADER section in asm file</p>\n<p>- count of '??' 2grams in bytes file</p>\n<p>- count of '00' 2grams in bytes file</p>\n<p>- length of .rsrc section in asm file</p>\n<p>- count of &#8216;A6&#8217; 2grams</p>\n<p>- count of &#8216;DE&#8217; 2grams</p>\n<p>- sum of all hex values in bytes files (excluding &#8216;?&#8217; chars)</p>\n<p>- count of &#8216;E8&#8217; 2grams</p>\n<p>- count of movzx commands</p>\n<p>- count of align commands</p>\n<p>- length of .idat section</p>\n<p>- hash of segment type</p>\n<p>- length of .data section</p>\n<p>- count of proc commands</p>\n<p>- first address of asm file</p>\n<p>- hash of Segment Permissions</p>\n<p>- count of jb commands</p>\n<p>- hash of line 10 in asm file</p>\n<p>- count of xor commands</p>\n<p>- count of dll calls</p>\n<p>- length of DATA section</p>\n<p>- count of sbb commands</p>\n<p>- count of fstp commands</p>\n<p>- hash of 34th line in DATA section</p>\n<p><strong>Probability Calibration</strong></p>\n<p>While RF did a very good job of separating the classes, it did a poor job of predicting probabilities. This is a well-known problem with the RF algorithm, as described in [1]. Because each tree contains only a subset of features, even a very good model will include some trees with poor predictions, which will skew the probabilities towards the center. I now know that there are a number of existing techniques to re-calibrate the probabilities produced by an RF classifier, but I was unaware of this research at the time I encountered the issue, so I devised a heuristic to solve the problem.</p>\n<p>Since the prediction probabilities produced by RF will be too conservative (e.g., too close to .5), it seemed reasonable that raising the probabilities to a power <em>n</em> and then re-normalizing might improve the results. For example, in a two-class problem, suppose the probabilities are .95 and .05. After raising both to the power of 2 and then renormalizing, the probabilities would be .997 and .003.</p>\n<p>Applying this technique to the results of the Malware model dramatically improved the log loss score. Before applying the calibration, CV scores from my best model were about .02. Applying a standard power of&nbsp;2.3 to all probabilities reduced the CV score to about .006. The LB scores fell by a similar amount.&nbsp;</p>\n<p><span style=\"line-height: 1.4\">I found further that I could improve the log loss score even more by applying a different calibration power to each class. This makes sense since the model is&nbsp;better at identifying some classes than others, so some classes need more calibration than others. Interpreting the calibration powers as tuning parameters and using a CV grid search to find the best ones resulted in the following list of calibration powers:</span></p>\n<p>Class1: 3.1</p>\n<p>Class2: 1.8</p>\n<p>Class3: 4.4</p>\n<p>Class4: 2.8</p>\n<p>Class5: 1.9</p>\n<p>Class6: 2.9</p>\n<p>Class7: 2.7</p>\n<p>Class8: 2.4</p>\n<p>Class9: 1.8</p>\n<p>For example, all predictions for Class 1 were raised to the power of 3.1, and all predictions for Class 2 were raised to the power of 1.8, etc.</p>\n<p><span style=\"line-height: 1.4\">Applying this heuristic to the raw probability scores produced by the RF model reduced the CV score to about 0.005 and the public LB score to about .004.&nbsp;Further, the calibration powers were relatively stable. Changing them a little had correspondingly&nbsp;small effect on the log loss value. This makes&nbsp;makes the technique far superior to threshholding, which can have dramatic and unexpected effects on the logloss score.&nbsp;</span></p>\n<p><span style=\"line-height: 1.4\">Since my feature set was relatively small and not especially creative, I think the calibration technique had the biggest impact on producing a top-10 result. I suspect that if others who used RF (or other ensemble methods) had applied this technique, they could have easily surpassed my score.</span></p>\n<p>[1] H. Bostr&#246;m. Estimating class probabilities in random forests. In Proc. of the International Conference on Machine Learning and Applications, pages 211&#8211;216, 2007.</p>",
  "messages": [
    {
      "id": "72451",
      "postDate": "04/19/2015 15:53:14",
      "content": "<p>First, congratulations to the winners. This was a fun competition with an interesting dataset and great communication in the forums. I&#8217;m posting a description of my solution because it seems to use many fewer features (25) than the winning solutions and includes an interesting calibration technique for RF prediction probabilities.</p>\n<p><strong>General Approach</strong></p>\n<p>1. Generate features from the .asm and .byte files.&nbsp;<span style=\"line-height: 1.4\">Features included:</span></p>\n<p style=\"padding-left: 30px\">&#8226; General characteristics, such as the length of the files, start address of the asm file, etc.</p>\n<p style=\"padding-left: 30px\">&#8226; Counts of various objects such as hex codes, hex 2grams, assembly language commands, etc.</p>\n<p style=\"padding-left: 30px\">&#8226; Hashes of specific parts of the asm file, such as each of the first 50 lines and specific properties of the asm file (e.g., segment type, segment permissions, etc.)</p>\n<p>2. Reduce the feature set (which started at about 2000 features) with greedy feature selection. CV was done with 10-fold cross-validation repeated 10 times with re-randomized folds, so 100 folds altogether.</p>\n<p>3. Generate prediction probabilities with random forest (RF) algorithm (sklearn implementation, estimators ranging from 500 &#8211; 1,500)</p>\n<p>4. Arguably the most important step, calibrate the prediction probabilities by raising them to a power selected via grid search (actually 9 different powers, one for each class). See below for details of this technique.</p>\n<p><strong>Performance </strong></p>\n<p>Public LB Performance was consistently a little better than CV performance but almost always moved in the same direction. The final CV and LB log losses were .005 and .004, respectively. The Private LB result was .006.</p>\n<p><strong>Model </strong></p>\n<p>The final solution was a single RF model consisting of the mean of 10 forests. Each forest consisted of 1,500 trees. The purpose of producing 10 different forests was to randomly change the hash seed for each forest to discourage RF from treating the hash features as numerical values. In approximate order of importance, the features were:</p>\n<p>- hash of 11th line of asm file (this is the line is most of the asm files that identifies the target processor)</p>\n<p>- length of HEADER section in asm file</p>\n<p>- count of '??' 2grams in bytes file</p>\n<p>- count of '00' 2grams in bytes file</p>\n<p>- length of .rsrc section in asm file</p>\n<p>- count of &#8216;A6&#8217; 2grams</p>\n<p>- count of &#8216;DE&#8217; 2grams</p>\n<p>- sum of all hex values in bytes files (excluding &#8216;?&#8217; chars)</p>\n<p>- count of &#8216;E8&#8217; 2grams</p>\n<p>- count of movzx commands</p>\n<p>- count of align commands</p>\n<p>- length of .idat section</p>\n<p>- hash of segment type</p>\n<p>- length of .data section</p>\n<p>- count of proc commands</p>\n<p>- first address of asm file</p>\n<p>- hash of Segment Permissions</p>\n<p>- count of jb commands</p>\n<p>- hash of line 10 in asm file</p>\n<p>- count of xor commands</p>\n<p>- count of dll calls</p>\n<p>- length of DATA section</p>\n<p>- count of sbb commands</p>\n<p>- count of fstp commands</p>\n<p>- hash of 34th line in DATA section</p>\n<p><strong>Probability Calibration</strong></p>\n<p>While RF did a very good job of separating the classes, it did a poor job of predicting probabilities. This is a well-known problem with the RF algorithm, as described in [1]. Because each tree contains only a subset of features, even a very good model will include some trees with poor predictions, which will skew the probabilities towards the center. I now know that there are a number of existing techniques to re-calibrate the probabilities produced by an RF classifier, but I was unaware of this research at the time I encountered the issue, so I devised a heuristic to solve the problem.</p>\n<p>Since the prediction probabilities produced by RF will be too conservative (e.g., too close to .5), it seemed reasonable that raising the probabilities to a power <em>n</em> and then re-normalizing might improve the results. For example, in a two-class problem, suppose the probabilities are .95 and .05. After raising both to the power of 2 and then renormalizing, the probabilities would be .997 and .003.</p>\n<p>Applying this technique to the results of the Malware model dramatically improved the log loss score. Before applying the calibration, CV scores from my best model were about .02. Applying a standard power of&nbsp;2.3 to all probabilities reduced the CV score to about .006. The LB scores fell by a similar amount.&nbsp;</p>\n<p><span style=\"line-height: 1.4\">I found further that I could improve the log loss score even more by applying a different calibration power to each class. This makes sense since the model is&nbsp;better at identifying some classes than others, so some classes need more calibration than others. Interpreting the calibration powers as tuning parameters and using a CV grid search to find the best ones resulted in the following list of calibration powers:</span></p>\n<p>Class1: 3.1</p>\n<p>Class2: 1.8</p>\n<p>Class3: 4.4</p>\n<p>Class4: 2.8</p>\n<p>Class5: 1.9</p>\n<p>Class6: 2.9</p>\n<p>Class7: 2.7</p>\n<p>Class8: 2.4</p>\n<p>Class9: 1.8</p>\n<p>For example, all predictions for Class 1 were raised to the power of 3.1, and all predictions for Class 2 were raised to the power of 1.8, etc.</p>\n<p><span style=\"line-height: 1.4\">Applying this heuristic to the raw probability scores produced by the RF model reduced the CV score to about 0.005 and the public LB score to about .004.&nbsp;Further, the calibration powers were relatively stable. Changing them a little had correspondingly&nbsp;small effect on the log loss value. This makes&nbsp;makes the technique far superior to threshholding, which can have dramatic and unexpected effects on the logloss score.&nbsp;</span></p>\n<p><span style=\"line-height: 1.4\">Since my feature set was relatively small and not especially creative, I think the calibration technique had the biggest impact on producing a top-10 result. I suspect that if others who used RF (or other ensemble methods) had applied this technique, they could have easily surpassed my score.</span></p>\n<p>[1] H. Bostr&#246;m. Estimating class probabilities in random forests. In Proc. of the International Conference on Machine Learning and Applications, pages 211&#8211;216, 2007.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72473",
      "postDate": "04/19/2015 16:58:47",
      "content": "<p>Thank you so much for posting this.</p>\n<p>Its funny because I remember looking at my probabilities and thinking that I had a pretty good feature set, but was leaving too much logloss &quot;on the table&quot; in my high confidence predictions, and wouldn't it be sweet if I could tweak those <strong>elegantly..&nbsp;</strong>It was nagging at me the whole time</p>\n\n<p>I decided (wrongly) to press on digging for features...</p>\n<p>I tried hand trimming them but it was too messy. &nbsp; I just love how you can apply this very slick technique for free almost and just see what it does to the CV. &nbsp;</p>\n<p>really really appreciate you taking the time to post</p>\n\n<p>awesome stuff</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72477",
      "postDate": "04/19/2015 17:04:59",
      "content": "<p>Impressive tricks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72481",
      "postDate": "04/19/2015 17:08:49",
      "content": "<p>Thanks for sharing!</p>\n<p>At the beginning I used RF too and I used squared predictions (in your terms - power 2 for all classes). It helps a lot!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72485",
      "postDate": "04/19/2015 17:36:07",
      "content": "<p>I ended doing a similar calibration power but with an one parameter function for the power of the class &nbsp;</p>\n\n<p>powerCls = x + (1.0-clsAccuracy)</p>\n\n<p>and then did a grid a search for the parameter x</p>\n<p>Now I am curious what I left on the table by not grid searching the powers for each class.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72492",
      "postDate": "04/19/2015 18:41:39",
      "content": "<p>I did the exact same thing in terms of probabilities calibration by raising classes to powers. I've been using this method since Tradeshift. It's a black art. For me it looked more like:</p>\n<p>mat1 &lt;- fread(&quot;/home/mikeskim/sub1.csv&quot;,header=TRUE,data.table=F)<br> powerVec = c(1.45, 0.9, 1.3, 3.6, 3.1, 1.1, 1.8, 1.2, 1.1)<br> for (x in 2:10) { mat1[,x] = mat1[,x]^powerVec[x-1] } <br>That last power of 1.1 was not optimized but the rest are close to optimal.</p>\n<p>It's scary how multiple Kagglers will all converge independently on the same tricks. I bet we're not the only people using this method of calibration.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72851",
      "postDate": "04/21/2015 05:10:03",
      "content": "<p>Congratulations and thanks for revealing all these tricks! I have some questions:</p>\n<p>1) Could you elaborate a little more on how you combined the 10 RFs?</p>\n<p>2) Did you use 2 bytes (or more) hex counts as features? (things like &quot;FF 00&quot; or &quot;FC 00 BC&quot;). It seems that the answer is no, but im curious because they are strong features.</p>\n<p>3) Which kind of greedy selection you used?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72867",
      "postDate": "04/21/2015 07:06:50",
      "content": "<p>@Bats &amp; Robots</p>\n<p>1) Just took the mean of the predictions from all 10 forests</p>\n<p>2) Yes, I did use some 2gram counts -- ??, 00, etc. See list above.</p>\n<p>3) I used a sort of funny system for greedy selection. I started by adding features based on CV scores, but used a relatively low number of trees (500) and folds (10) for this step so it tended to add more features than were optimal. Then I removed features using cv with more trees (1000) and more folds (100). I found that this method was much more efficient than starting with all potential features.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "73057",
      "postDate": "04/22/2015 00:38:27",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74521",
      "postDate": "04/24/2015 16:41:17",
      "content": "<p>Thanks so much for sharing your solution and trick! I also used RF but did not know that you can do such probability calibration. Nice work ~</p>\n<p>It is good to know that you can achieve this performance with so few number of features, and am looking forward to&nbsp;reading your code &amp; paper.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75231",
      "postDate": "04/28/2015 12:30:33",
      "content": "<p>Thanks for posting your solution. Regarding calibration, why not directly turn probs to 0-1 values taking the max score/prob over all classes, turn to 1, all rest to 0? The perfect log-loss would work like this.</p>\n<p>I am curious of effect of turning probs to 0-1&nbsp;versus raising to power.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75234",
      "postDate": "04/28/2015 12:34:01",
      "content": "<p>[quote=Georgia;75231]</p>\n<p>Thanks for posting your solution. Regarding calibration, why not directly turn probs to 0-1 values taking the max score/prob over all classes, turn to 1, all rest to 0? The perfect log-loss would work like this.</p>\n<p>I am curious of effect of turning probs to 0-1&nbsp;versus raising to power.</p>\n<p>[/quote]</p>\n<p>LogLoss is very sensetive to marginal predictions, so turning probs to 0-1 will probably lead to awful score.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "1807848",
      "postDate": "06/01/2022 11:40:36",
      "content": "<p>ohh.., this trick is awesome.</p>",
      "rawMarkdown": "ohh.., this trick is awesome.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1807848,
      "author_name": "klokeshsai",
      "author_url": "",
      "post_date": "06/01/2022 11:40:36",
      "content": "<p>ohh.., this trick is awesome.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72473,
      "author_name": "scharf",
      "author_url": "",
      "post_date": "04/19/2015 16:58:47",
      "content": "<p>Thank you so much for posting this.</p>\n<p>Its funny because I remember looking at my probabilities and thinking that I had a pretty good feature set, but was leaving too much logloss &quot;on the table&quot; in my high confidence predictions, and wouldn't it be sweet if I could tweak those <strong>elegantly..&nbsp;</strong>It was nagging at me the whole time</p>\n\n<p>I decided (wrongly) to press on digging for features...</p>\n<p>I tried hand trimming them but it was too messy. &nbsp; I just love how you can apply this very slick technique for free almost and just see what it does to the CV. &nbsp;</p>\n<p>really really appreciate you taking the time to post</p>\n\n<p>awesome stuff</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72477,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/19/2015 17:04:59",
      "content": "<p>Impressive tricks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72481,
      "author_name": "mikhailtrofimov",
      "author_url": "",
      "post_date": "04/19/2015 17:08:49",
      "content": "<p>Thanks for sharing!</p>\n<p>At the beginning I used RF too and I used squared predictions (in your terms - power 2 for all classes). It helps a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72485,
      "author_name": "jbsilva",
      "author_url": "",
      "post_date": "04/19/2015 17:36:07",
      "content": "<p>I ended doing a similar calibration power but with an one parameter function for the power of the class &nbsp;</p>\n\n<p>powerCls = x + (1.0-clsAccuracy)</p>\n\n<p>and then did a grid a search for the parameter x</p>\n<p>Now I am curious what I left on the table by not grid searching the powers for each class.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72492,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "04/19/2015 18:41:39",
      "content": "<p>I did the exact same thing in terms of probabilities calibration by raising classes to powers. I've been using this method since Tradeshift. It's a black art. For me it looked more like:</p>\n<p>mat1 &lt;- fread(&quot;/home/mikeskim/sub1.csv&quot;,header=TRUE,data.table=F)<br> powerVec = c(1.45, 0.9, 1.3, 3.6, 3.1, 1.1, 1.8, 1.2, 1.1)<br> for (x in 2:10) { mat1[,x] = mat1[,x]^powerVec[x-1] } <br>That last power of 1.1 was not optimized but the rest are close to optimal.</p>\n<p>It's scary how multiple Kagglers will all converge independently on the same tricks. I bet we're not the only people using this method of calibration.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72851,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "04/21/2015 05:10:03",
      "content": "<p>Congratulations and thanks for revealing all these tricks! I have some questions:</p>\n<p>1) Could you elaborate a little more on how you combined the 10 RFs?</p>\n<p>2) Did you use 2 bytes (or more) hex counts as features? (things like &quot;FF 00&quot; or &quot;FC 00 BC&quot;). It seems that the answer is no, but im curious because they are strong features.</p>\n<p>3) Which kind of greedy selection you used?</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72867,
      "author_name": "psilogram",
      "author_url": "",
      "post_date": "04/21/2015 07:06:50",
      "content": "<p>@Bats &amp; Robots</p>\n<p>1) Just took the mean of the predictions from all 10 forests</p>\n<p>2) Yes, I did use some 2gram counts -- ??, 00, etc. See list above.</p>\n<p>3) I used a sort of funny system for greedy selection. I started by adding features based on CV scores, but used a relatively low number of trees (500) and folds (10) for this step so it tended to add more features than were optimal. Then I removed features using cv with more trees (1000) and more folds (100). I found that this method was much more efficient than starting with all potential features.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 73057,
      "author_name": "soumyasmruti",
      "author_url": "",
      "post_date": "04/22/2015 00:38:27",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 74521,
      "author_name": "jieqchen",
      "author_url": "",
      "post_date": "04/24/2015 16:41:17",
      "content": "<p>Thanks so much for sharing your solution and trick! I also used RF but did not know that you can do such probability calibration. Nice work ~</p>\n<p>It is good to know that you can achieve this performance with so few number of features, and am looking forward to&nbsp;reading your code &amp; paper.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75231,
      "author_name": "georgia",
      "author_url": "",
      "post_date": "04/28/2015 12:30:33",
      "content": "<p>Thanks for posting your solution. Regarding calibration, why not directly turn probs to 0-1 values taking the max score/prob over all classes, turn to 1, all rest to 0? The perfect log-loss would work like this.</p>\n<p>I am curious of effect of turning probs to 0-1&nbsp;versus raising to power.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75234,
      "author_name": "mikhailtrofimov",
      "author_url": "",
      "post_date": "04/28/2015 12:34:01",
      "content": "<p>[quote=Georgia;75231]</p>\n<p>Thanks for posting your solution. Regarding calibration, why not directly turn probs to 0-1 values taking the max score/prob over all classes, turn to 1, all rest to 0? The perfect log-loss would work like this.</p>\n<p>I am curious of effect of turning probs to 0-1&nbsp;versus raising to power.</p>\n<p>[/quote]</p>\n<p>LogLoss is very sensetive to marginal predictions, so turning probs to 0-1 will probably lead to awful score.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "72451": "",
    "72473": "",
    "72477": "",
    "72481": "",
    "72485": "",
    "72492": "",
    "72851": "",
    "72867": "",
    "73057": "",
    "74521": "",
    "75231": "",
    "75234": "",
    "1807848": "ohh.., this trick is awesome."
  },
  "source": "meta"
}