{
  "id": 13490,
  "title": "Say no to Overfitting approaches sharing",
  "url": "/competitions/malware-classification/discussion/13490",
  "author_name": "",
  "post_date": "2015-04-18T17:20:37.413Z",
  "votes": 19,
  "comment_count": 31,
  "views": 12739,
  "content": "<p>Hi everyone, thanks for the great game! I will briefly describe what we did here. A more detailed description and code will be presented later after some cleaning.</p>\n<p>The features we used consists of rcarson's 4k features and 3k features from me. I will just describe mine and let rcarson reveal his magic. :)</p>\n<p>1. &nbsp;single byte count. Inspired by the beat benchmark code. Thanks!</p>\n<p>2. &nbsp;instruction count, like 'mov', 'jp', ...</p>\n<p>3. &nbsp;4 gram byte features. I used info gain to select the best 500 features for each class. So there are 500 * 9 of them.</p>\n<p>4. DAF features. Like the one described in this paper.</p>\n<p>https://www.utdallas.edu/~lkhan/papers/A%20Knowledge-based%20Approach%20to%20Detect%20New%20Malicious%20Executables.pdf</p>\n<p>Each feature is a standard representation of asm instructions, following the format: &nbsp;name.para1.para2, such as or.memory.register.</p>\n<p>5. DLL functions. The original idea is to get the dll features described in paper from 4. But I had no idea what that is, so I just gathered the function names instead.&nbsp;</p>\n<p>The ways I gathered features of 3,4,5 are very greedy, so I did a random forest on them and selected about 2k features out of all.</p>\n<p>6. asm image features. Convert each asm file into 'image array' like the step 0 in this blog ( Thanks SARVAM team! )</p>\n<p>http://sarvamblog.blogspot.ca/2014/08/supervised-classification-with-k-fold.html</p>\n<p>And then only took the first 800 values from the array.</p>\n<p>And they all make more than 3k features. Combining rcarson's 4k features, a xgboost gives about 0.0052 cv score.</p>\n<p>And rcarson uses these features and he magically generated 2 other xgboost models.</p>\n<p>A geo mean of these 3 models gives about 0.0042 cv score, and 0.0031 in private board.</p>\n<p>Another trick:</p>\n<p>Take the result from the geo mean of the 3 models described above, I generated the label of test set by choosing the max probability.</p>\n<p>And then, I combine them into training and generate a 'semi learned' result. The key here is not including the data for training and testing. Say, you can divide the data into 2 parts A and B. So when you train A, you only include B and the entire training set. Otherwise you will overfit your misclassified points and the score will be just terrible.</p>\n<p>It achieves 0.00316 cv score by some playing and tuning, and 0.0024 in private board.</p>\n<p>The score in the current private board is a geo mean combination of 'semi learned' model and other models above.</p>\n<p>Hope it helps! :)</p>",
  "messages": [
    {
      "id": "72311",
      "postDate": "04/18/2015 17:20:37",
      "content": "<p>Hi everyone, thanks for the great game! I will briefly describe what we did here. A more detailed description and code will be presented later after some cleaning.</p>\n<p>The features we used consists of rcarson's 4k features and 3k features from me. I will just describe mine and let rcarson reveal his magic. :)</p>\n<p>1. &nbsp;single byte count. Inspired by the beat benchmark code. Thanks!</p>\n<p>2. &nbsp;instruction count, like 'mov', 'jp', ...</p>\n<p>3. &nbsp;4 gram byte features. I used info gain to select the best 500 features for each class. So there are 500 * 9 of them.</p>\n<p>4. DAF features. Like the one described in this paper.</p>\n<p>https://www.utdallas.edu/~lkhan/papers/A%20Knowledge-based%20Approach%20to%20Detect%20New%20Malicious%20Executables.pdf</p>\n<p>Each feature is a standard representation of asm instructions, following the format: &nbsp;name.para1.para2, such as or.memory.register.</p>\n<p>5. DLL functions. The original idea is to get the dll features described in paper from 4. But I had no idea what that is, so I just gathered the function names instead.&nbsp;</p>\n<p>The ways I gathered features of 3,4,5 are very greedy, so I did a random forest on them and selected about 2k features out of all.</p>\n<p>6. asm image features. Convert each asm file into 'image array' like the step 0 in this blog ( Thanks SARVAM team! )</p>\n<p>http://sarvamblog.blogspot.ca/2014/08/supervised-classification-with-k-fold.html</p>\n<p>And then only took the first 800 values from the array.</p>\n<p>And they all make more than 3k features. Combining rcarson's 4k features, a xgboost gives about 0.0052 cv score.</p>\n<p>And rcarson uses these features and he magically generated 2 other xgboost models.</p>\n<p>A geo mean of these 3 models gives about 0.0042 cv score, and 0.0031 in private board.</p>\n<p>Another trick:</p>\n<p>Take the result from the geo mean of the 3 models described above, I generated the label of test set by choosing the max probability.</p>\n<p>And then, I combine them into training and generate a 'semi learned' result. The key here is not including the data for training and testing. Say, you can divide the data into 2 parts A and B. So when you train A, you only include B and the entire training set. Otherwise you will overfit your misclassified points and the score will be just terrible.</p>\n<p>It achieves 0.00316 cv score by some playing and tuning, and 0.0024 in private board.</p>\n<p>The score in the current private board is a geo mean combination of 'semi learned' model and other models above.</p>\n<p>Hope it helps! :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72312",
      "postDate": "04/18/2015 17:43:23",
      "content": "<p>A big congratulations for hitting the jackpot!</p>\n<p>Thanks for sharing your approach and glad that you found our blog useful.&nbsp;</p>\n<p>We fell pretty badly and still trying to recover. &nbsp;We believe it was partly dup to choosing bad submissions and partly due to going for hard decosions. Guess it's a hard learning lesson for us.&nbsp;</p>\n<p>Could you also share on how you choose the final decisions? &nbsp;Were they all soft?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72313",
      "postDate": "04/18/2015 17:52:04",
      "content": "<p>[quote=Lakshman Nataraj;72312]</p>\n<p>A big congratulations for hitting the jackpot!</p>\n<p>Thanks for sharing your approach and glad that you found our blog useful.&nbsp;</p>\n<p>We fell pretty badly and still trying to recover. &nbsp;We believe it was partly dup to choosing bad submissions and partly due to going for hard decosions. Guess it's a hard learning lesson for us.&nbsp;</p>\n<p>Could you also share on how you choose the final decisions? &nbsp;Were they all soft?</p>\n<p>[/quote]</p>\n<p>Thanks again for the blog!&nbsp;</p>\n<p>Yes, they are both soft. I think for logloss, it is not a good idea to do&nbsp;hard calibrations, because even one totally wrong prediction can be nightmare.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72314",
      "postDate": "04/18/2015 17:57:42",
      "content": "<p>[quote=Little Boat;72313]</p>\n<p>Thanks again for the blog!&nbsp;</p>\n<p>Yes, they are both soft. I think for logloss, it is not a good idea to do&nbsp;hard calibrations, because even one totally wrong prediction can be nightmare.</p>\n<p>[/quote]</p>\n<p>Thanks, &nbsp;i realize now. Great work again!&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72315",
      "postDate": "04/18/2015 18:13:46",
      "content": "<p>Ok, I'll talk about the rest '4k' features. It has two parts: opcode count and header count</p>\n<p>1) opcode count. We don't use any assembly instruction dictionary to find proper candidates beforehand. We basically count every word in asm files. The criterion is simple: the 'interesting' word has to be frequent in at least one asm files. We only select the words which appear more than 200 times in at least one asm file. Threshold 200 is chosen manually. We also throw away any bytes or address in asm files. In the end we have 165 such words. and they happen to be opcodes mostly.</p>\n<p>Another thing is if an opcode lives inside a for ever loop, its count x10. Forever loop is detected based on unconditional jump instruction such as jmp.</p>\n<p>We find simple&nbsp;count is more useful than frequency or tfidf in this case. Based on 165 'interesting' opcodes, we extract 2gram, 3gram and 4gram of them.</p>\n<p>2) As pointed out by&nbsp;gmilosev, header count such as .text, .code, .upx is a good feature. We just don't use frequency. Maybe we should try that as well.</p>\n<p>So 1) + 2) are 70k features. We use random forest to select useful ones and we end up with 4k features.</p>\n<p>This 4k features alone with xgb, we get 0.01 cv + 0.0089 public lb</p>\n<p>3) we generate some 'asm image' features like what Little Boat does above. The difference is only the 'Pure code' part in asm files are kept and addresses in each line are removed. The first 800 pixels of the image are used. I call them 'code' feature.</p>\n<p>So our final two solutions are : 1)&nbsp;sub1:&nbsp;ensemble solution + calibration&nbsp;2) sub2:ensemble solution + semi-supervised learning.&nbsp;</p>\n<p>We just found ensemble alone is better in private LB but worse in public LB, compared to ensemble+calibration</p>\n<p>The ensemble include three models:</p>\n<p>model1: 4k feature:&nbsp;opcode&nbsp;2g,3g,4g+ header cound + Little boat's asm image feature + Little Boat's frequency+byte4g+daf features.</p>\n<p>use shallow xgb, eta=0.2, min_child_weight=1, depth=1, num_round=200, we get 0.0055 cv</p>\n<p>model2: 4k feature + Little boat's asm image feature&nbsp;+&nbsp;'code' feature</p>\n<p>use deep xgb, eta=0.25,min_child_weight=1,depth=30, num_round=50, we get 0.0059 cv</p>\n<p>model 3: all Little Boat's features + 4k features, use shallow xgb we get 0.0051 cv</p>\n<p>Our ensemble=model1^0.1*model2^0.4*model3^0.5</p>\n<p>The weights are obtained through grid search in cross validation.</p>\n<p>it gets 0.0046 cv, 0.0039 public LB, 0.0026 private LB</p>\n<p>We did some calibration, it becomes, 0.0041 cv, 0.0035 public LB and 0.0031 private LB</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72329",
      "postDate": "04/18/2015 19:32:13",
      "content": "<p>Briefly our approach:</p>\n<p>Our cvs were 80-20 in training and 50-50 in meta.</p>\n<p>datasets:</p>\n<p>1) 1 gram bytes (256)</p>\n<p>2) 2 gram Bytes (65k)</p>\n<p>3) 3 gram &nbsp;Bytes (top 60 K most popular)</p>\n<p>4) 4 gram Bytes (top 60 K most popular)</p>\n<p>5) top 100K most popular per subject full-line bytes&nbsp;</p>\n<p>6) .zip , gzip ratios vs normal sizes along with other files's stats</p>\n<p>7) top 400 k most popular features in .asm (1 gram) exlduing words of length 2 and 8</p>\n<p>8) selection of features form .asm that made sense (based on section, counts non-alphanumeric)</p>\n<p>all these modeled with 60% xgboost and 40% extratreesclassifier</p>\n<p>Generated a couple of datasets with different combos from all of these (e.g. 1grams bytes with .asm datasets or 2gram bytes with .asm datasets etc) around 20 in number.</p>\n<p>Then meta modelling to combine all again with xgboost and extra 50-50</p>\n<p>we only submitted if at least 9 out of 10 (50%-50%) folds improved our score</p>\n<p>our scores follow linearly the public leader board all along - We never over fitted.</p>\n<p>Stuff that did not work for us:</p>\n<p>-Linear models</p>\n<p>-clipping predictions (I tested internally different thresholds and all failed miserably - never bothered submitting to see what happens)</p>\n<p>-more than 5&gt; grams in bytes (while all other datasets are present)&nbsp;</p>\n<p>-2 grams or more in .asm</p>\n<p>I would presume you need a couple of days (if not weeks to fully reproduce our solution :/)</p>\n<p>One last:</p>\n<p>Big credit to my teammate Gert for creating the best data sets with&nbsp;exceptional feature generation all along. It was again an honor playing and exchanging ideas.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72345",
      "postDate": "04/18/2015 20:45:58",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;72329]</p>\n<p>Briefly our approach:</p>\n<p>Our cvs were 80-20 in training and 50-50 in meta.</p>\n<p>datasets:</p>\n<p>1) 1 gram bytes (256)</p>\n<p>2) 2 gram Bytes (65k)</p>\n<p>3) 3 gram &nbsp;Bytes (top 60 K most popular)</p>\n<p>4) 4 gram Bytes (top 60 K most popular)</p>\n<p>5) top 100K most popular per subject full-line bytes&nbsp;</p>\n<p>6) .zip , gzip ratios vs normal sizes along with other files's stats</p>\n<p>7) top 400 k most popular features in .asm (1 gram) exlduing words of length 2 and 8</p>\n<p>8) selection of features form .asm that made sense (based on section, counts non-alphanumeric)</p>\n<p>all these modeled with 60% xgboost and 40% extratreesclassifier</p>\n<p>Generated a couple of datasets with different combos from all of these (e.g. 1grams bytes with .asm datasets or 2gram bytes with .asm datasets etc) around 20 in number.</p>\n<p>Then meta modelling to combine all again with xgboost and extra 50-50</p>\n<p>we only submitted if at least 9 out of 10 (50%-50%) folds improved our score</p>\n<p>our scores follow linearly the public leader board all along - We never over fitted.</p>\n<p>Stuff that did not work for us:</p>\n<p>-Linear models</p>\n<p>-clipping predictions (I tested internally different thresholds and all failed miserably - never bothered submitting to see what happens)</p>\n<p>-more than 5&gt; grams in bytes (while all other datasets are present)&nbsp;</p>\n<p>-2 grams or more in .asm</p>\n<p>I would presume you need a couple of days (if not weeks to fully reproduce our solution :/)</p>\n<p>One last:</p>\n<p>Big credit to my teammate Gert for creating the best data sets with&nbsp;exceptional feature generation all along. It was again an honor playing and exchanging ideas.</p>\n<p>[/quote]</p>\n<p>Great job! Does it require lots of computer power to reproduce what you did? :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72346",
      "postDate": "04/18/2015 20:47:42",
      "content": "<p>@rcarson, any particular reason why you used geom mean ? Just tryed both and arithmetic&nbsp;and geom and this one was better ? Any ideas why ?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72347",
      "postDate": "04/18/2015 20:51:01",
      "content": "<p>[quote=Dmitry Ulyanov;72346]</p>\n<p>@rcarson, any particular reason why you used geom mean ? Just tryed both and arithmetic&nbsp;and geom and this one was better ? Any ideas why ?&nbsp;</p>\n<p>[/quote]</p>\n<p>We did some math and the geo mean tends to beat average with high probability in terms of multi logloss, assuming every model does well.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72348",
      "postDate": "04/18/2015 20:54:51",
      "content": "<p>Yeah, I tried both. In this case, geo mean is better. My general impression was that geo mean is better for averaging 'different' models and arithmetic mean is better for averaging 'similar' models. Correct me if I am wrong :P</p>\n\n<p>In our case, the three models are similar on most samples but very different on outliers. Since outliers are dominant at this point, geo mean is better. This is how I justify it : )</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72351",
      "postDate": "04/18/2015 21:05:31",
      "content": "<p>@Little Boat, @rcarson thanks! I myself found funny thing on mixing. We had algos X, Y, Z and we built final decision as mean(Z, X , Y , XY*4 , XYZ*4 ) , so the funny thing is that &quot;4&quot;, you can grid search it, I basically thought we need it&nbsp;because XY &lt; X, XY &lt; Y and will not sum up to 1, so we scale predictions somehow, I tried to use XY / sum (XY) as fair probabilities, but never beaten the combination from above. This data science sometimes involves no science :)&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72354",
      "postDate": "04/18/2015 21:14:31",
      "content": "<p>@Dmitry Ulyanov&nbsp;</p>\n<p>Did you see a big improvement doing the mixing? I have no idea this kind of blending can actually work.. Thanks for sharing!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72357",
      "postDate": "04/18/2015 21:24:46",
      "content": "<p>@Little Boat &nbsp;</p>\n<p>Yes, I had a huge local boost by introducing XY and XYZ in combination. Overall, mixing helped a lot, I guess our strongest model with a lot of tweeks got 0.005 CV, but we argued could we belive it or not since it used semisupervised technique (slightly different from yours). Mikhail will post our solution soon.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72358",
      "postDate": "04/18/2015 21:49:25",
      "content": "<p>[quote=Little Boat;72345]</p>\n<p>Great job! Does it require lots of computer power to reproduce what you did? :)</p>\n<p>[/quote]</p>\n\n<p>It does not require huge power...However extra threads could heavily facilitate the process! Most feature engineering is online (printing in files) so it does not take much RAM. It takes time though..</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72365",
      "postDate": "04/18/2015 23:50:25",
      "content": "<p>My solution:</p>\n<p>Feature group 1) - Convert each .bytes&nbsp;file into an image matrix&nbsp;and extract 320 features acording SARVAM method. Then use kmeans and other clustering algorithms to extract groups with different number of classes of that 320 features. &nbsp;It generated 10 features to me.</p>\n<p>Feature group 2) - Count of many manual choosen words from the .asm files by section (.header, .text, .data, etc), like: instructions, coments, dll names, imports, jmp, unk, conditional jumps, some names, etc. It my tests that features have a good performance. It generated about 550 features.</p>\n<p>Feature group 3) - &nbsp;Most interesting feature, .asm files mean line length by section. I gives me 11 features from all sections. Performed good. Probably using stardard deviation of line length can also give good performance, but I didn't tested it. To build that features I suposed Mean Line Length by section could work as a malware signature.&nbsp;</p>\n<p>Feature group 4) - Files sizes, .bytes file first byte. 3 features.</p>\n<p>My last step was join all those features and run an multiclass random forest with many trees to find the feature importances. Then I just pick around 350 more important features and trainned it using crossvalidation with xgboost. High number of trees&nbsp;and low eta gives a more stable response in gbm algorithms. My local CV was 0.0103. Then I aplied a trim technique to that results testing in my crossvalidated trainset. For each instance if the maximum prob was higher than 0.9998 I turn it to 1 and the other 8 classes to 0. If any class have prob below 0.0002 I set it to 0 and sum the prob value to the class with the high prob in that instance. It improved my CV trained dataset performance from 0.0103 to ~0.009. And private 0.00684.</p>\n<p>As the trim technique is very risk and only 1 mislabeled instance can destroy you results,&nbsp;I trusted in my local CV, so I choose those two submission as my final models. It seems my trimmed model worked well.</p>\n<p>Giba</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72407",
      "postDate": "04/19/2015 09:37:58",
      "content": "<p>Congratulations to the &quot;say noooo to overfitting&quot; team (I rooted for you as soon as I saw your name)- Can you please elaborate on the hardware specs you used, both for feature extraction and modeling etc? (ram, processors etc.)</p>\n\n<p>Thanks dearly and again congratulations,</p>\n<p>WC</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72408",
      "postDate": "04/19/2015 09:49:02",
      "content": "<p>Thanks for sharing the approach ..... Ha!, Ha!, Ha!.....</p>\n<p>They say that there is not such thing as a new idea.... I thought I was the one <strong>who invented this approach</strong>....I was actually<strong> patting myself on the back</strong> about coming up with the idea.... The vastness of the internet surly does indicate original ideas are very rare.</p>\n<p>As I said to you &quot;Little Boat&quot; in another topic in the Forum ...all I need was to get my training parameters correct right so I could get out of the rut of my 98.8% and move into the the 99.9% plus and I think I could have won this completion easily ....</p>\n<p>Let this be a lesson to anyone starting a challenge late (Monday the 13th my first submission);</p>\n<p>such is life ;)</p>\n<p>Congratulation for win.....&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72419",
      "postDate": "04/19/2015 12:00:03",
      "content": "<p>[quote=Gilberto Titericz Junior;72365]</p>\n<p>My solution:</p>\n<p>Feature group 1) - Convert each .bytes&nbsp;file into an image matrix&nbsp;and extract 320 features acording SARVAM method. Then use kmeans and other clustering algorithms to extract groups with different number of classes of that 320 features. &nbsp;It generated 10 features to me.</p>\n<p>[/quote]</p>\n<p>As you did in 1), I tried the SARVAM method, but I didn't manage to get good results. I failed to use the leargist Library in my Python code. How did you use it? Is it compatible with Windows?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72424",
      "postDate": "04/19/2015 12:25:08",
      "content": "<p>To Small Boat and his team .... Again Congratulations!!!</p>\n<p>But, I see this discussion is about devolve into a conversation about techniques and tools....</p>\n<p>So before it goes any further I ask all:</p>\n<p>Do you understand underlying biophysical/neurological/physiological (I can't find the proper word right now) why this approach works and why it is fundamental to offering a solution to all static pattern, classification problems, however diverse they may seem from reading a pattern recognitions challenge? ---It is the reason I was patting myself on the back for thinking I came up with an original idea in machine learning/intelligence.</p>\n<p>Techniques are wonderful for brute force getting an answer to a problem/question. However, it really is necessary to fully understand why a technique works...... and from my view of&nbsp; the world that inspiration comes from natures examples in spending billions of years perfecting....</p>\n<p>In passing ... Just because you know how to add two numbers together doesn't mean you understand why you can add them and the underlying structure that supports the ability to add two numbers together.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72434",
      "postDate": "04/19/2015 14:41:42",
      "content": "<p>[quote=Wild Cherry;72407]</p>\n<p>Congratulations to the &quot;say noooo to overfitting&quot; team (I rooted for you as soon as I saw your name)- Can you please elaborate on the hardware specs you used, both for feature extraction and modeling etc? (ram, processors etc.)</p>\n<p>Thanks dearly and again congratulations,</p>\n<p>WC</p>\n<p>[/quote]</p>\n<p>For rcarson, all the feature extraction can be done by a laptop, say, with 4GB memory, I believe.</p>\n<p>For me, since I have a server with 104G memory, I trade code optimization with memory. But only the 4 gram part eats a lot of memory, the others can be done with 16G or maybe 32G. But I am sure even 4 gram you can build the feature with 4GB memory.</p>\n\n<p>For modeling, xgboost is so good that any laptop can do it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72437",
      "postDate": "04/19/2015 14:50:22",
      "content": "<p>[quote=Little Boat;72434]</p>\n<p>[quote=Wild Cherry;72407]</p>\n<p>Can you please elaborate on the hardware specs you used, both for feature extraction and modeling etc? (ram, processors etc.)</p>\n<p>Thanks dearly and again congratulations,</p>\n<p>WC</p>\n<p>[/quote]</p>\n<p>For rcarson, all the feature extraction can be done by a laptop, say, with 4GB memory, I believe.</p>\n<p>[/quote]</p>\n<p>Yeah, in reality I used 16 GB memory, it's more than enough :D</p>\n<p>Processor wise we used a 4th gen i7. After features are ready, it took 20 mins at most to run a 4 fold cv or 10 mins to generate a solution with xgboost. But as Little Boat points out, feature engineering is the most demanding part, it takes a day or two to generate all the features we used in the final solution.</p>\n<p>In my part, the most critical resource is disk, to do all kinds of experiments I used up 1.7 TB in the end, although most of them are not useful :P</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72438",
      "postDate": "04/19/2015 14:52:37",
      "content": "<p>[quote=Michael George Hart;72424]</p>\n<p>To Small Boat and his team .... Again Congratulations!!!</p>\n<p>But, I see this discussion is about devolve into a conversation about techniques and tools....</p>\n<p>So before it goes any further I ask all:</p>\n<p>Do you understand underlying biophysical/neurological/physiological (I can't find the proper word right now) why this approach works and why it is fundamental to offering a solution to all static pattern, classification problems, however diverse they may seem from reading a pattern recognitions challenge? ---It is the reason I was patting myself on the back for thinking I came up with an original idea in machine learning/intelligence.</p>\n<p>Techniques are wonderful for brute force getting an answer to a problem/question. However, it really is necessary to fully understand why a technique works...... and from my view of&nbsp; the world that inspiration comes from natures examples in spending billions of years perfecting....</p>\n<p>In passing ... Just because you know how to add two numbers together doesn't mean you understand why you can add them and the underlying structure that supports the ability to add two numbers together.</p>\n<p>[/quote]</p>\n<p>Thanks&nbsp;Michael. And the answer is no. I don't know why some features work but some don't. We just kept trying, and if it worked, we try to understand it, if it didn't work, we move on to another idea.&nbsp;</p>\n<p>That said, after trying lots of ideas, we do have a feeling for what kind of feature to find. For example, we realized that if there are golden features, they should hide somewhere in the asm files. And that was the motivation to try the asm array features, and it worked very well.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72454",
      "postDate": "04/19/2015 16:00:30",
      "content": "<p>[quote=Michael George Hart;72408]</p>\n<p>&nbsp;I think I could have won this completion easily ....</p>\n<p>[/quote]</p>\n<p>There are many other kaggle competitions going on at the moment for you to <em>easily</em> win.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72458",
      "postDate": "04/19/2015 16:18:11",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;72454]</p>\n<p>[quote=Michael George Hart;72408]</p>\n<p>&nbsp;I think I could have won this completion easily ....</p>\n<p>[/quote]</p>\n<p>There are many other kaggle competitions going on at the moment for you to <em>easily</em> win.</p>\n<p>[/quote]</p>\n<p>If only it was so easy. ;)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72460",
      "postDate": "04/19/2015 16:21:31",
      "content": "<p>[quote=Little Boat;72438]</p>\n<p>I don't know why some features work but some don't. We just kept trying, and if it worked, we try to understand it, if it didn't work, we move on to another idea.&nbsp;</p>\n<p>[/quote]</p>\n<p>I have also&nbsp;traditionally found this brute approach to be better than heavily spending time to understand why something works or not. I also&nbsp;believe this is the power of machine learning versus traditional analysis (aka we rely more on the machine algorithms) . I guess the latter is still useful when you try to sell your stuff to clients !</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72462",
      "postDate": "04/19/2015 16:27:58",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;72460]</p>\n<p>[quote=Little Boat;72438]</p>\n<p>I don't know why some features work but some don't. We just kept trying, and if it worked, we try to understand it, if it didn't work, we move on to another idea.&nbsp;</p>\n<p>[/quote]</p>\n<p>I have also&nbsp;traditionally found this brute approach to be better than heavily spending time to understand why something works or not. I also&nbsp;believe this is the power of machine learning versus traditional analysis (aka we rely more on the machine algorithms) . I guess the latter is still useful when you try to sell your stuff to clients !</p>\n<p>[/quote]</p>\n<p>Can't agree more. I think just like bias variance trade off, there is also interpretability and accuracy trade off. And for kaggle competitions, I think interpretability doesn't matter that much.</p>\n<p><span style=\"line-height: 1.4\">And not only clients, but also your non tech boss. :)</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72466",
      "postDate": "04/19/2015 16:45:57",
      "content": "<p>@Little Boat and others ...</p>\n<p>&nbsp; Here are the foundation to how I came up with my little theory, I jokingly called, &quot;The Fundamental Theory of Machine Classification&quot;</p>\n<p>1. Machine Intelligence does not necessarily need to human like or even mammalian like --we may very well not be able to recognize other &quot;intelligences&quot; that did not necessarily follow the DNA path of our planet's evolution.</p>\n<p>2. Most intelligence (fast understanding of an object) seems to operate off of a GIST... from my personal experience, which I believe is true for most humans, All GIST get translated into our visual system (verb, sound, etc...)</p>\n<p>Mix those two idea along with an old paper, I read sometime ago,&nbsp; &quot;<a href=\"http://cvcl.mit.edu/Papers/Oliva04.pdf\">Gist of the Scene</a>&quot; published in a Neurobiology&nbsp; and you get any dataset that can be transformed into an image can be use to at least learn and classify anything .... couple that with the Softmax function at the back-end of any learning system your get a good sense of probability how much something belongs to a class.</p>\n<p>Just because the human visual recognition system, which happens to carries with it a lot of evolutionary environmental baggage, does not see the GIST of a Scene, does not necessarily mean that an artificially created&nbsp; neurological system of an artificial environment will not be able see things that a human would never be able to see.... hopefully you can see the power of this... eg what a human sees and what a machine sees in the night sky could be radically different and useful to astronomer&nbsp;</p>\n<p>So, that is the biological foundation how I came up with the theory ..... I have tested it out on many toy problems with very good results... This particular contest is the fist time I have ever had real-world dataset this big to push the concepts to the limits... it held up great unfortunately it took more that 1100 epoch's and 1 1/2 days of training to get above the 98.87% accuracy; getting into the 99.9% which obviously occurred after the contest ended.</p>\n<p>That business you did&nbsp; n-grams and features selection/extraction what really was brilliant ... that is going to be very useful to me as I continue to pursue the idea....</p>\n<p>I am very curious how long did it take to train your system?</p>\n<p>Thanks</p>\n<p>Michael</p>\n<p>PS... Anyone using this idea please give me are least an honourable mention or get someone to sponsor me so I can spend full-time flushing the theory out ;) :p :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72482",
      "postDate": "04/19/2015 17:13:12",
      "content": "<p>@Little Boat</p>\n<p>Congratulations! Great work!</p>\n<p>I got one question. Can you elaborate your final trick about&nbsp;'semi learned' results?&nbsp;I think I get the gist of it, but still a bit puzzled.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72488",
      "postDate": "04/19/2015 18:01:09",
      "content": "<p>[quote=Zhao Yilong;72482]</p>\n<p>@Little Boat</p>\n<p>Congratulations! Great work!</p>\n<p>I got one question. Can you elaborate your final trick about&nbsp;'semi learned' results?&nbsp;I think I get the gist of it, but still a bit puzzled.</p>\n<p>[/quote]</p>\n<p>Sure. So say you trained a model on the entire training set and got a predicted distribution for each test point.</p>\n<p>Now, you used the max probability as the label for that point. Say point 1 got 0.999 on class 5, then we assign the label 5 to point 1. And you have the labels for the test set. The good news is, you can easily reach 98,99% accuracy for this problem. So your labels are very close to the truth.</p>\n<p>Now, you divide your test set into N fold, say 4 fold, A, B, C, D.</p>\n<p>And then you use the entire training set + A,B,C to train a model and predict on D.</p>\n<p>And then entire training set + A, B, D to predict C. etc..</p>\n<p>The reason it works is, I think, we have more points to train. A lot of papers talk about introducing pseudo data to boost the result, and with this high accuracy, test set with only a few wrong labels should be a great choice for that.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72489",
      "postDate": "04/19/2015 18:02:36",
      "content": "<p>Congratulations to all others, well played in this very interesting and harddisk-challenging competition!</p>\n<p>Also&nbsp;many kudos for my teammate Marios (better known as Kazanova)&nbsp;and&nbsp;his excellent modelling skills. At some point (quite soon) I had to no choice but to give up matching the performance of&nbsp;his ingenious mix of xgboost models - which made me&nbsp;entirely focus on feature generation. Much of that&nbsp;was done on my old 32bits desktop with a slow disk, but plenty of time.</p>\n<p>The best features came from the asm files, and an&nbsp;important insight for us was that splitting features out by section (the first word of each line, like: HEADER, .text, .data) improves results a lot. To our surprise, meta-features about interpunction [' ','?','.',',',':',';','+','-','=','[','(','_','*','!','\\\\','/','\\''] worked better than contents of the text [A-Z0-9] parts.</p>\n<p>Apart from the byte ngrams, we also calculated the compressed sizes of each 4 KB block and generated statistics about their distribution over each file.</p>\n<p>Finally thanks to&nbsp;ThierryS for sharing the 'stupid solution' (that soon made us think about compression) and&nbsp;Lakshman for sharing his paper (that inspired the 4KB block distribution statistics).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "73056",
      "postDate": "04/22/2015 00:37:07",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "1361082",
      "postDate": "06/22/2021 14:35:52",
      "content": "<p>That's awesome work, and thank you for sharing.<br>\nDiscussions are also helpful. Learn a lot!</p>",
      "rawMarkdown": "That's awesome work, and thank you for sharing.\nDiscussions are also helpful. Learn a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1361082,
      "author_name": "ss8651twtw",
      "author_url": "",
      "post_date": "06/22/2021 14:35:52",
      "content": "<p>That's awesome work, and thank you for sharing.<br>\nDiscussions are also helpful. Learn a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72312,
      "author_name": "",
      "author_url": "",
      "post_date": "04/18/2015 17:43:23",
      "content": "<p>A big congratulations for hitting the jackpot!</p>\n<p>Thanks for sharing your approach and glad that you found our blog useful.&nbsp;</p>\n<p>We fell pretty badly and still trying to recover. &nbsp;We believe it was partly dup to choosing bad submissions and partly due to going for hard decosions. Guess it's a hard learning lesson for us.&nbsp;</p>\n<p>Could you also share on how you choose the final decisions? &nbsp;Were they all soft?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72313,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/18/2015 17:52:04",
      "content": "<p>[quote=Lakshman Nataraj;72312]</p>\n<p>A big congratulations for hitting the jackpot!</p>\n<p>Thanks for sharing your approach and glad that you found our blog useful.&nbsp;</p>\n<p>We fell pretty badly and still trying to recover. &nbsp;We believe it was partly dup to choosing bad submissions and partly due to going for hard decosions. Guess it's a hard learning lesson for us.&nbsp;</p>\n<p>Could you also share on how you choose the final decisions? &nbsp;Were they all soft?</p>\n<p>[/quote]</p>\n<p>Thanks again for the blog!&nbsp;</p>\n<p>Yes, they are both soft. I think for logloss, it is not a good idea to do&nbsp;hard calibrations, because even one totally wrong prediction can be nightmare.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72314,
      "author_name": "",
      "author_url": "",
      "post_date": "04/18/2015 17:57:42",
      "content": "<p>[quote=Little Boat;72313]</p>\n<p>Thanks again for the blog!&nbsp;</p>\n<p>Yes, they are both soft. I think for logloss, it is not a good idea to do&nbsp;hard calibrations, because even one totally wrong prediction can be nightmare.</p>\n<p>[/quote]</p>\n<p>Thanks, &nbsp;i realize now. Great work again!&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72315,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/18/2015 18:13:46",
      "content": "<p>Ok, I'll talk about the rest '4k' features. It has two parts: opcode count and header count</p>\n<p>1) opcode count. We don't use any assembly instruction dictionary to find proper candidates beforehand. We basically count every word in asm files. The criterion is simple: the 'interesting' word has to be frequent in at least one asm files. We only select the words which appear more than 200 times in at least one asm file. Threshold 200 is chosen manually. We also throw away any bytes or address in asm files. In the end we have 165 such words. and they happen to be opcodes mostly.</p>\n<p>Another thing is if an opcode lives inside a for ever loop, its count x10. Forever loop is detected based on unconditional jump instruction such as jmp.</p>\n<p>We find simple&nbsp;count is more useful than frequency or tfidf in this case. Based on 165 'interesting' opcodes, we extract 2gram, 3gram and 4gram of them.</p>\n<p>2) As pointed out by&nbsp;gmilosev, header count such as .text, .code, .upx is a good feature. We just don't use frequency. Maybe we should try that as well.</p>\n<p>So 1) + 2) are 70k features. We use random forest to select useful ones and we end up with 4k features.</p>\n<p>This 4k features alone with xgb, we get 0.01 cv + 0.0089 public lb</p>\n<p>3) we generate some 'asm image' features like what Little Boat does above. The difference is only the 'Pure code' part in asm files are kept and addresses in each line are removed. The first 800 pixels of the image are used. I call them 'code' feature.</p>\n<p>So our final two solutions are : 1)&nbsp;sub1:&nbsp;ensemble solution + calibration&nbsp;2) sub2:ensemble solution + semi-supervised learning.&nbsp;</p>\n<p>We just found ensemble alone is better in private LB but worse in public LB, compared to ensemble+calibration</p>\n<p>The ensemble include three models:</p>\n<p>model1: 4k feature:&nbsp;opcode&nbsp;2g,3g,4g+ header cound + Little boat's asm image feature + Little Boat's frequency+byte4g+daf features.</p>\n<p>use shallow xgb, eta=0.2, min_child_weight=1, depth=1, num_round=200, we get 0.0055 cv</p>\n<p>model2: 4k feature + Little boat's asm image feature&nbsp;+&nbsp;'code' feature</p>\n<p>use deep xgb, eta=0.25,min_child_weight=1,depth=30, num_round=50, we get 0.0059 cv</p>\n<p>model 3: all Little Boat's features + 4k features, use shallow xgb we get 0.0051 cv</p>\n<p>Our ensemble=model1^0.1*model2^0.4*model3^0.5</p>\n<p>The weights are obtained through grid search in cross validation.</p>\n<p>it gets 0.0046 cv, 0.0039 public LB, 0.0026 private LB</p>\n<p>We did some calibration, it becomes, 0.0041 cv, 0.0035 public LB and 0.0031 private LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72329,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "04/18/2015 19:32:13",
      "content": "<p>Briefly our approach:</p>\n<p>Our cvs were 80-20 in training and 50-50 in meta.</p>\n<p>datasets:</p>\n<p>1) 1 gram bytes (256)</p>\n<p>2) 2 gram Bytes (65k)</p>\n<p>3) 3 gram &nbsp;Bytes (top 60 K most popular)</p>\n<p>4) 4 gram Bytes (top 60 K most popular)</p>\n<p>5) top 100K most popular per subject full-line bytes&nbsp;</p>\n<p>6) .zip , gzip ratios vs normal sizes along with other files's stats</p>\n<p>7) top 400 k most popular features in .asm (1 gram) exlduing words of length 2 and 8</p>\n<p>8) selection of features form .asm that made sense (based on section, counts non-alphanumeric)</p>\n<p>all these modeled with 60% xgboost and 40% extratreesclassifier</p>\n<p>Generated a couple of datasets with different combos from all of these (e.g. 1grams bytes with .asm datasets or 2gram bytes with .asm datasets etc) around 20 in number.</p>\n<p>Then meta modelling to combine all again with xgboost and extra 50-50</p>\n<p>we only submitted if at least 9 out of 10 (50%-50%) folds improved our score</p>\n<p>our scores follow linearly the public leader board all along - We never over fitted.</p>\n<p>Stuff that did not work for us:</p>\n<p>-Linear models</p>\n<p>-clipping predictions (I tested internally different thresholds and all failed miserably - never bothered submitting to see what happens)</p>\n<p>-more than 5&gt; grams in bytes (while all other datasets are present)&nbsp;</p>\n<p>-2 grams or more in .asm</p>\n<p>I would presume you need a couple of days (if not weeks to fully reproduce our solution :/)</p>\n<p>One last:</p>\n<p>Big credit to my teammate Gert for creating the best data sets with&nbsp;exceptional feature generation all along. It was again an honor playing and exchanging ideas.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72345,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/18/2015 20:45:58",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;72329]</p>\n<p>Briefly our approach:</p>\n<p>Our cvs were 80-20 in training and 50-50 in meta.</p>\n<p>datasets:</p>\n<p>1) 1 gram bytes (256)</p>\n<p>2) 2 gram Bytes (65k)</p>\n<p>3) 3 gram &nbsp;Bytes (top 60 K most popular)</p>\n<p>4) 4 gram Bytes (top 60 K most popular)</p>\n<p>5) top 100K most popular per subject full-line bytes&nbsp;</p>\n<p>6) .zip , gzip ratios vs normal sizes along with other files's stats</p>\n<p>7) top 400 k most popular features in .asm (1 gram) exlduing words of length 2 and 8</p>\n<p>8) selection of features form .asm that made sense (based on section, counts non-alphanumeric)</p>\n<p>all these modeled with 60% xgboost and 40% extratreesclassifier</p>\n<p>Generated a couple of datasets with different combos from all of these (e.g. 1grams bytes with .asm datasets or 2gram bytes with .asm datasets etc) around 20 in number.</p>\n<p>Then meta modelling to combine all again with xgboost and extra 50-50</p>\n<p>we only submitted if at least 9 out of 10 (50%-50%) folds improved our score</p>\n<p>our scores follow linearly the public leader board all along - We never over fitted.</p>\n<p>Stuff that did not work for us:</p>\n<p>-Linear models</p>\n<p>-clipping predictions (I tested internally different thresholds and all failed miserably - never bothered submitting to see what happens)</p>\n<p>-more than 5&gt; grams in bytes (while all other datasets are present)&nbsp;</p>\n<p>-2 grams or more in .asm</p>\n<p>I would presume you need a couple of days (if not weeks to fully reproduce our solution :/)</p>\n<p>One last:</p>\n<p>Big credit to my teammate Gert for creating the best data sets with&nbsp;exceptional feature generation all along. It was again an honor playing and exchanging ideas.</p>\n<p>[/quote]</p>\n<p>Great job! Does it require lots of computer power to reproduce what you did? :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72346,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "04/18/2015 20:47:42",
      "content": "<p>@rcarson, any particular reason why you used geom mean ? Just tryed both and arithmetic&nbsp;and geom and this one was better ? Any ideas why ?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72347,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/18/2015 20:51:01",
      "content": "<p>[quote=Dmitry Ulyanov;72346]</p>\n<p>@rcarson, any particular reason why you used geom mean ? Just tryed both and arithmetic&nbsp;and geom and this one was better ? Any ideas why ?&nbsp;</p>\n<p>[/quote]</p>\n<p>We did some math and the geo mean tends to beat average with high probability in terms of multi logloss, assuming every model does well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72348,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/18/2015 20:54:51",
      "content": "<p>Yeah, I tried both. In this case, geo mean is better. My general impression was that geo mean is better for averaging 'different' models and arithmetic mean is better for averaging 'similar' models. Correct me if I am wrong :P</p>\n\n<p>In our case, the three models are similar on most samples but very different on outliers. Since outliers are dominant at this point, geo mean is better. This is how I justify it : )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72351,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "04/18/2015 21:05:31",
      "content": "<p>@Little Boat, @rcarson thanks! I myself found funny thing on mixing. We had algos X, Y, Z and we built final decision as mean(Z, X , Y , XY*4 , XYZ*4 ) , so the funny thing is that &quot;4&quot;, you can grid search it, I basically thought we need it&nbsp;because XY &lt; X, XY &lt; Y and will not sum up to 1, so we scale predictions somehow, I tried to use XY / sum (XY) as fair probabilities, but never beaten the combination from above. This data science sometimes involves no science :)&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72354,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/18/2015 21:14:31",
      "content": "<p>@Dmitry Ulyanov&nbsp;</p>\n<p>Did you see a big improvement doing the mixing? I have no idea this kind of blending can actually work.. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72357,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "04/18/2015 21:24:46",
      "content": "<p>@Little Boat &nbsp;</p>\n<p>Yes, I had a huge local boost by introducing XY and XYZ in combination. Overall, mixing helped a lot, I guess our strongest model with a lot of tweeks got 0.005 CV, but we argued could we belive it or not since it used semisupervised technique (slightly different from yours). Mikhail will post our solution soon.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72358,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "04/18/2015 21:49:25",
      "content": "<p>[quote=Little Boat;72345]</p>\n<p>Great job! Does it require lots of computer power to reproduce what you did? :)</p>\n<p>[/quote]</p>\n\n<p>It does not require huge power...However extra threads could heavily facilitate the process! Most feature engineering is online (printing in files) so it does not take much RAM. It takes time though..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72365,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "04/18/2015 23:50:25",
      "content": "<p>My solution:</p>\n<p>Feature group 1) - Convert each .bytes&nbsp;file into an image matrix&nbsp;and extract 320 features acording SARVAM method. Then use kmeans and other clustering algorithms to extract groups with different number of classes of that 320 features. &nbsp;It generated 10 features to me.</p>\n<p>Feature group 2) - Count of many manual choosen words from the .asm files by section (.header, .text, .data, etc), like: instructions, coments, dll names, imports, jmp, unk, conditional jumps, some names, etc. It my tests that features have a good performance. It generated about 550 features.</p>\n<p>Feature group 3) - &nbsp;Most interesting feature, .asm files mean line length by section. I gives me 11 features from all sections. Performed good. Probably using stardard deviation of line length can also give good performance, but I didn't tested it. To build that features I suposed Mean Line Length by section could work as a malware signature.&nbsp;</p>\n<p>Feature group 4) - Files sizes, .bytes file first byte. 3 features.</p>\n<p>My last step was join all those features and run an multiclass random forest with many trees to find the feature importances. Then I just pick around 350 more important features and trainned it using crossvalidation with xgboost. High number of trees&nbsp;and low eta gives a more stable response in gbm algorithms. My local CV was 0.0103. Then I aplied a trim technique to that results testing in my crossvalidated trainset. For each instance if the maximum prob was higher than 0.9998 I turn it to 1 and the other 8 classes to 0. If any class have prob below 0.0002 I set it to 0 and sum the prob value to the class with the high prob in that instance. It improved my CV trained dataset performance from 0.0103 to ~0.009. And private 0.00684.</p>\n<p>As the trim technique is very risk and only 1 mislabeled instance can destroy you results,&nbsp;I trusted in my local CV, so I choose those two submission as my final models. It seems my trimmed model worked well.</p>\n<p>Giba</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72407,
      "author_name": "yahalom",
      "author_url": "",
      "post_date": "04/19/2015 09:37:58",
      "content": "<p>Congratulations to the &quot;say noooo to overfitting&quot; team (I rooted for you as soon as I saw your name)- Can you please elaborate on the hardware specs you used, both for feature extraction and modeling etc? (ram, processors etc.)</p>\n\n<p>Thanks dearly and again congratulations,</p>\n<p>WC</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72408,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/19/2015 09:49:02",
      "content": "<p>Thanks for sharing the approach ..... Ha!, Ha!, Ha!.....</p>\n<p>They say that there is not such thing as a new idea.... I thought I was the one <strong>who invented this approach</strong>....I was actually<strong> patting myself on the back</strong> about coming up with the idea.... The vastness of the internet surly does indicate original ideas are very rare.</p>\n<p>As I said to you &quot;Little Boat&quot; in another topic in the Forum ...all I need was to get my training parameters correct right so I could get out of the rut of my 98.8% and move into the the 99.9% plus and I think I could have won this completion easily ....</p>\n<p>Let this be a lesson to anyone starting a challenge late (Monday the 13th my first submission);</p>\n<p>such is life ;)</p>\n<p>Congratulation for win.....&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72419,
      "author_name": "thomasseleck",
      "author_url": "",
      "post_date": "04/19/2015 12:00:03",
      "content": "<p>[quote=Gilberto Titericz Junior;72365]</p>\n<p>My solution:</p>\n<p>Feature group 1) - Convert each .bytes&nbsp;file into an image matrix&nbsp;and extract 320 features acording SARVAM method. Then use kmeans and other clustering algorithms to extract groups with different number of classes of that 320 features. &nbsp;It generated 10 features to me.</p>\n<p>[/quote]</p>\n<p>As you did in 1), I tried the SARVAM method, but I didn't manage to get good results. I failed to use the leargist Library in my Python code. How did you use it? Is it compatible with Windows?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72424,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/19/2015 12:25:08",
      "content": "<p>To Small Boat and his team .... Again Congratulations!!!</p>\n<p>But, I see this discussion is about devolve into a conversation about techniques and tools....</p>\n<p>So before it goes any further I ask all:</p>\n<p>Do you understand underlying biophysical/neurological/physiological (I can't find the proper word right now) why this approach works and why it is fundamental to offering a solution to all static pattern, classification problems, however diverse they may seem from reading a pattern recognitions challenge? ---It is the reason I was patting myself on the back for thinking I came up with an original idea in machine learning/intelligence.</p>\n<p>Techniques are wonderful for brute force getting an answer to a problem/question. However, it really is necessary to fully understand why a technique works...... and from my view of&nbsp; the world that inspiration comes from natures examples in spending billions of years perfecting....</p>\n<p>In passing ... Just because you know how to add two numbers together doesn't mean you understand why you can add them and the underlying structure that supports the ability to add two numbers together.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72434,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/19/2015 14:41:42",
      "content": "<p>[quote=Wild Cherry;72407]</p>\n<p>Congratulations to the &quot;say noooo to overfitting&quot; team (I rooted for you as soon as I saw your name)- Can you please elaborate on the hardware specs you used, both for feature extraction and modeling etc? (ram, processors etc.)</p>\n<p>Thanks dearly and again congratulations,</p>\n<p>WC</p>\n<p>[/quote]</p>\n<p>For rcarson, all the feature extraction can be done by a laptop, say, with 4GB memory, I believe.</p>\n<p>For me, since I have a server with 104G memory, I trade code optimization with memory. But only the 4 gram part eats a lot of memory, the others can be done with 16G or maybe 32G. But I am sure even 4 gram you can build the feature with 4GB memory.</p>\n\n<p>For modeling, xgboost is so good that any laptop can do it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72437,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/19/2015 14:50:22",
      "content": "<p>[quote=Little Boat;72434]</p>\n<p>[quote=Wild Cherry;72407]</p>\n<p>Can you please elaborate on the hardware specs you used, both for feature extraction and modeling etc? (ram, processors etc.)</p>\n<p>Thanks dearly and again congratulations,</p>\n<p>WC</p>\n<p>[/quote]</p>\n<p>For rcarson, all the feature extraction can be done by a laptop, say, with 4GB memory, I believe.</p>\n<p>[/quote]</p>\n<p>Yeah, in reality I used 16 GB memory, it's more than enough :D</p>\n<p>Processor wise we used a 4th gen i7. After features are ready, it took 20 mins at most to run a 4 fold cv or 10 mins to generate a solution with xgboost. But as Little Boat points out, feature engineering is the most demanding part, it takes a day or two to generate all the features we used in the final solution.</p>\n<p>In my part, the most critical resource is disk, to do all kinds of experiments I used up 1.7 TB in the end, although most of them are not useful :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72438,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/19/2015 14:52:37",
      "content": "<p>[quote=Michael George Hart;72424]</p>\n<p>To Small Boat and his team .... Again Congratulations!!!</p>\n<p>But, I see this discussion is about devolve into a conversation about techniques and tools....</p>\n<p>So before it goes any further I ask all:</p>\n<p>Do you understand underlying biophysical/neurological/physiological (I can't find the proper word right now) why this approach works and why it is fundamental to offering a solution to all static pattern, classification problems, however diverse they may seem from reading a pattern recognitions challenge? ---It is the reason I was patting myself on the back for thinking I came up with an original idea in machine learning/intelligence.</p>\n<p>Techniques are wonderful for brute force getting an answer to a problem/question. However, it really is necessary to fully understand why a technique works...... and from my view of&nbsp; the world that inspiration comes from natures examples in spending billions of years perfecting....</p>\n<p>In passing ... Just because you know how to add two numbers together doesn't mean you understand why you can add them and the underlying structure that supports the ability to add two numbers together.</p>\n<p>[/quote]</p>\n<p>Thanks&nbsp;Michael. And the answer is no. I don't know why some features work but some don't. We just kept trying, and if it worked, we try to understand it, if it didn't work, we move on to another idea.&nbsp;</p>\n<p>That said, after trying lots of ideas, we do have a feeling for what kind of feature to find. For example, we realized that if there are golden features, they should hide somewhere in the asm files. And that was the motivation to try the asm array features, and it worked very well.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72454,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "04/19/2015 16:00:30",
      "content": "<p>[quote=Michael George Hart;72408]</p>\n<p>&nbsp;I think I could have won this completion easily ....</p>\n<p>[/quote]</p>\n<p>There are many other kaggle competitions going on at the moment for you to <em>easily</em> win.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72458,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/19/2015 16:18:11",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;72454]</p>\n<p>[quote=Michael George Hart;72408]</p>\n<p>&nbsp;I think I could have won this completion easily ....</p>\n<p>[/quote]</p>\n<p>There are many other kaggle competitions going on at the moment for you to <em>easily</em> win.</p>\n<p>[/quote]</p>\n<p>If only it was so easy. ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72460,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "04/19/2015 16:21:31",
      "content": "<p>[quote=Little Boat;72438]</p>\n<p>I don't know why some features work but some don't. We just kept trying, and if it worked, we try to understand it, if it didn't work, we move on to another idea.&nbsp;</p>\n<p>[/quote]</p>\n<p>I have also&nbsp;traditionally found this brute approach to be better than heavily spending time to understand why something works or not. I also&nbsp;believe this is the power of machine learning versus traditional analysis (aka we rely more on the machine algorithms) . I guess the latter is still useful when you try to sell your stuff to clients !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72462,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/19/2015 16:27:58",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;72460]</p>\n<p>[quote=Little Boat;72438]</p>\n<p>I don't know why some features work but some don't. We just kept trying, and if it worked, we try to understand it, if it didn't work, we move on to another idea.&nbsp;</p>\n<p>[/quote]</p>\n<p>I have also&nbsp;traditionally found this brute approach to be better than heavily spending time to understand why something works or not. I also&nbsp;believe this is the power of machine learning versus traditional analysis (aka we rely more on the machine algorithms) . I guess the latter is still useful when you try to sell your stuff to clients !</p>\n<p>[/quote]</p>\n<p>Can't agree more. I think just like bias variance trade off, there is also interpretability and accuracy trade off. And for kaggle competitions, I think interpretability doesn't matter that much.</p>\n<p><span style=\"line-height: 1.4\">And not only clients, but also your non tech boss. :)</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72466,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/19/2015 16:45:57",
      "content": "<p>@Little Boat and others ...</p>\n<p>&nbsp; Here are the foundation to how I came up with my little theory, I jokingly called, &quot;The Fundamental Theory of Machine Classification&quot;</p>\n<p>1. Machine Intelligence does not necessarily need to human like or even mammalian like --we may very well not be able to recognize other &quot;intelligences&quot; that did not necessarily follow the DNA path of our planet's evolution.</p>\n<p>2. Most intelligence (fast understanding of an object) seems to operate off of a GIST... from my personal experience, which I believe is true for most humans, All GIST get translated into our visual system (verb, sound, etc...)</p>\n<p>Mix those two idea along with an old paper, I read sometime ago,&nbsp; &quot;<a href=\"http://cvcl.mit.edu/Papers/Oliva04.pdf\">Gist of the Scene</a>&quot; published in a Neurobiology&nbsp; and you get any dataset that can be transformed into an image can be use to at least learn and classify anything .... couple that with the Softmax function at the back-end of any learning system your get a good sense of probability how much something belongs to a class.</p>\n<p>Just because the human visual recognition system, which happens to carries with it a lot of evolutionary environmental baggage, does not see the GIST of a Scene, does not necessarily mean that an artificially created&nbsp; neurological system of an artificial environment will not be able see things that a human would never be able to see.... hopefully you can see the power of this... eg what a human sees and what a machine sees in the night sky could be radically different and useful to astronomer&nbsp;</p>\n<p>So, that is the biological foundation how I came up with the theory ..... I have tested it out on many toy problems with very good results... This particular contest is the fist time I have ever had real-world dataset this big to push the concepts to the limits... it held up great unfortunately it took more that 1100 epoch's and 1 1/2 days of training to get above the 98.87% accuracy; getting into the 99.9% which obviously occurred after the contest ended.</p>\n<p>That business you did&nbsp; n-grams and features selection/extraction what really was brilliant ... that is going to be very useful to me as I continue to pursue the idea....</p>\n<p>I am very curious how long did it take to train your system?</p>\n<p>Thanks</p>\n<p>Michael</p>\n<p>PS... Anyone using this idea please give me are least an honourable mention or get someone to sponsor me so I can spend full-time flushing the theory out ;) :p :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72482,
      "author_name": "yilongzhao",
      "author_url": "",
      "post_date": "04/19/2015 17:13:12",
      "content": "<p>@Little Boat</p>\n<p>Congratulations! Great work!</p>\n<p>I got one question. Can you elaborate your final trick about&nbsp;'semi learned' results?&nbsp;I think I get the gist of it, but still a bit puzzled.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72488,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/19/2015 18:01:09",
      "content": "<p>[quote=Zhao Yilong;72482]</p>\n<p>@Little Boat</p>\n<p>Congratulations! Great work!</p>\n<p>I got one question. Can you elaborate your final trick about&nbsp;'semi learned' results?&nbsp;I think I get the gist of it, but still a bit puzzled.</p>\n<p>[/quote]</p>\n<p>Sure. So say you trained a model on the entire training set and got a predicted distribution for each test point.</p>\n<p>Now, you used the max probability as the label for that point. Say point 1 got 0.999 on class 5, then we assign the label 5 to point 1. And you have the labels for the test set. The good news is, you can easily reach 98,99% accuracy for this problem. So your labels are very close to the truth.</p>\n<p>Now, you divide your test set into N fold, say 4 fold, A, B, C, D.</p>\n<p>And then you use the entire training set + A,B,C to train a model and predict on D.</p>\n<p>And then entire training set + A, B, D to predict C. etc..</p>\n<p>The reason it works is, I think, we have more points to train. A lot of papers talk about introducing pseudo data to boost the result, and with this high accuracy, test set with only a few wrong labels should be a great choice for that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72489,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "04/19/2015 18:02:36",
      "content": "<p>Congratulations to all others, well played in this very interesting and harddisk-challenging competition!</p>\n<p>Also&nbsp;many kudos for my teammate Marios (better known as Kazanova)&nbsp;and&nbsp;his excellent modelling skills. At some point (quite soon) I had to no choice but to give up matching the performance of&nbsp;his ingenious mix of xgboost models - which made me&nbsp;entirely focus on feature generation. Much of that&nbsp;was done on my old 32bits desktop with a slow disk, but plenty of time.</p>\n<p>The best features came from the asm files, and an&nbsp;important insight for us was that splitting features out by section (the first word of each line, like: HEADER, .text, .data) improves results a lot. To our surprise, meta-features about interpunction [' ','?','.',',',':',';','+','-','=','[','(','_','*','!','\\\\','/','\\''] worked better than contents of the text [A-Z0-9] parts.</p>\n<p>Apart from the byte ngrams, we also calculated the compressed sizes of each 4 KB block and generated statistics about their distribution over each file.</p>\n<p>Finally thanks to&nbsp;ThierryS for sharing the 'stupid solution' (that soon made us think about compression) and&nbsp;Lakshman for sharing his paper (that inspired the 4KB block distribution statistics).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 73056,
      "author_name": "soumyasmruti",
      "author_url": "",
      "post_date": "04/22/2015 00:37:07",
      "content": "",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "72311": "",
    "72312": "",
    "72313": "",
    "72314": "",
    "72315": "",
    "72329": "",
    "72345": "",
    "72346": "",
    "72347": "",
    "72348": "",
    "72351": "",
    "72354": "",
    "72357": "",
    "72358": "",
    "72365": "",
    "72407": "",
    "72408": "",
    "72419": "",
    "72424": "",
    "72434": "",
    "72437": "",
    "72438": "",
    "72454": "",
    "72458": "",
    "72460": "",
    "72462": "",
    "72466": "",
    "72482": "",
    "72488": "",
    "72489": "",
    "73056": "",
    "1361082": "That's awesome work, and thank you for sharing.\nDiscussions are also helpful. Learn a lot!"
  },
  "source": "meta"
}