{
  "id": 13463,
  "title": "Anyone wants to share your local performance?",
  "url": "/competitions/malware-classification/discussion/13463",
  "author_name": "",
  "post_date": "2015-04-17T16:25:38.663Z",
  "votes": 2,
  "comment_count": 12,
  "views": 1869,
  "content": "<p>Now that competition is reaching to its end, lets share our local validation scores?</p>\n<p>My local 20 fold CV accuracy is 0.9974236.</p>\n<p>My local 20 fold CV MLogLoss is&nbsp; 0.0103</p>\n<p>Confusion Matrix:</p>\n<p><code>&nbsp;&nbsp;&nbsp;&nbsp; 0&nbsp; &nbsp; 1&nbsp;&nbsp;&nbsp; 2&nbsp;&nbsp; 3&nbsp; 4&nbsp;&nbsp; 5&nbsp;&nbsp; 6 &nbsp;&nbsp; 7&nbsp;&nbsp;&nbsp; 8<br> 0 1539&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 2&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br> 1&nbsp;&nbsp;&nbsp; 1 2476&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0<br> 2&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 2941&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 0&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br> 3&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp; 473 0&nbsp;&nbsp; 1&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0<br> 4&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0 40&nbsp;&nbsp; 1&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 1<br> 5&nbsp;&nbsp;&nbsp; 3&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0 747&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br> 6&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 0 397&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br> 7&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp; 1&nbsp; 0&nbsp;&nbsp; 0&nbsp;&nbsp; 1 1221&nbsp;&nbsp;&nbsp; 2<br> 8&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 1&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 5 1006</code></p>\n<p>obs. [row = target classes; col = prediction]</p>",
  "messages": [
    {
      "id": "72105",
      "postDate": "04/17/2015 16:25:38",
      "content": "<p>Now that competition is reaching to its end, lets share our local validation scores?</p>\n<p>My local 20 fold CV accuracy is 0.9974236.</p>\n<p>My local 20 fold CV MLogLoss is&nbsp; 0.0103</p>\n<p>Confusion Matrix:</p>\n<p><code>&nbsp;&nbsp;&nbsp;&nbsp; 0&nbsp; &nbsp; 1&nbsp;&nbsp;&nbsp; 2&nbsp;&nbsp; 3&nbsp; 4&nbsp;&nbsp; 5&nbsp;&nbsp; 6 &nbsp;&nbsp; 7&nbsp;&nbsp;&nbsp; 8<br> 0 1539&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 2&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br> 1&nbsp;&nbsp;&nbsp; 1 2476&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0<br> 2&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 2941&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 0&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br> 3&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp; 473 0&nbsp;&nbsp; 1&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0<br> 4&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0 40&nbsp;&nbsp; 1&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 1<br> 5&nbsp;&nbsp;&nbsp; 3&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0 747&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br> 6&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 0 397&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br> 7&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp; 1&nbsp; 0&nbsp;&nbsp; 0&nbsp;&nbsp; 1 1221&nbsp;&nbsp;&nbsp; 2<br> 8&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp; 0&nbsp; 0&nbsp;&nbsp; 1&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp; 5 1006</code></p>\n<p>obs. [row = target classes; col = prediction]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72114",
      "postDate": "04/17/2015 17:30:22",
      "content": "<p>Would you please tell us about the number of features you used ?</p>\n<p>10 CV : 0.9968 , 0.0130, Number of features : Around 1200</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72115",
      "postDate": "04/17/2015 17:54:32",
      "content": "<p>We used about 7K features. and the best single model has cv 0.0035, and LB score 0.0037.</p>\n\n<p>Edit: the best single model has&nbsp;0.9992639 accuracy.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72138",
      "postDate": "04/17/2015 19:52:36",
      "content": "<p>I build more than 1K features from the .asm files, but used only about ~400 useful features in my last model.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72140",
      "postDate": "04/17/2015 20:01:51",
      "content": "<p>with FFFFFFFF + 9000 features, I think we are the ones who created most number of features ;)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72142",
      "postDate": "04/17/2015 20:06:46",
      "content": "<p>Ok..... something must be wrong with me haha, but I used 1.4M features, and my CV 0.0077 and LB 0.0073, this is without any parameter tuning (<a href=\"https://www.youtube.com/watch?v=8cT_Ulmcrys\">Ain't Nobody Got Time For That</a>) and single model.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72144",
      "postDate": "04/17/2015 20:20:24",
      "content": "<p>[quote=NxGTR;72142]</p>\n<p>Ok..... something must be wrong with me haha, but I used 1.4M features, and my CV 0.0077 and LB 0.0073, this is without any parameter tuning (<a href=\"https://www.youtube.com/watch?v=8cT_Ulmcrys\">Ain't Nobody Got Time For That</a>) and single model.</p>\n<p>[/quote]</p>\n<p>Wow, how did you train your 1.4M features? Using VW?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72147",
      "postDate": "04/17/2015 20:32:23",
      "content": "<p>Our best so far has been an accuracy of 99.926 (8 samples misclassified) with 60+ features.</p>\n<p>But the log loss is very tricky. If we take what the classifier gives, we get a log loss of&nbsp;0.0048. If we threshold it, we get .002.</p>\n<p>If we give hard scores (only 8 are wrong), our log loss jumps to 0.02</p>\n<p>So I guess this competition is mainly about finding a good threshold for the classifier outputs. Depending on the threshold, anything can happen.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72150",
      "postDate": "04/17/2015 21:04:52",
      "content": "<p>[quote=Little Boat;72144]</p>\n<p>Wow, how did you train your 1.4M features? Using VW?</p>\n<p>[/quote]</p>\n<p>Well, I pretty much can train them with anything I want right now as I built the features to maximize the sparsity between classes, so the final dataset is extremely sparse and I can train it on my personal laptop.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72153",
      "postDate": "04/17/2015 21:17:25",
      "content": "<p>[quote=NxGTR;72150]</p>\n<p>[quote=Little Boat;72144]</p>\n<p>Wow, how did you train your 1.4M features? Using VW?</p>\n<p>[/quote]</p>\n<p>Well, I pretty much can train them with anything I want right now as I built the features to maximize the sparsity between classes, so the final dataset is extremely sparse and I can train it on my personal laptop.</p>\n<p>[/quote]</p>\n<p>I see. Why don't you tune your model and maybe do some averaging... It helps ours. And you have time to do it!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72161",
      "postDate": "04/17/2015 21:34:00",
      "content": "<p>[quote=Abhishek;72140]</p>\n<p>with FFFFFFFF + 9000 features, I think we are the ones who created most number of features ;)</p>\n<p>[/quote]</p>\n<p>I have lost count but , but in our ensemble we have a couple of models with more than 400 K features, If I had to count, I would say we should have generated over 1 Million features.</p>\n<p>Our 10-fold 50/50 cv</p>\n<p>Loglikelihood (fold 1/10):0.0038178<br>Loglikelihood (fold 2/10):0.004158945<br>Loglikelihood (fold 3/10):0.003210165<br>Loglikelihood (fold 4/10):0.005466825<br>Loglikelihood (fold 5/10):0.00578151<br>Loglikelihood (fold 6/10):0.003658095<br>Loglikelihood (fold 7/10):0.004555845<br>Loglikelihood (fold 8/10):0.005502735<br>Loglikelihood (fold 9/10):0.007954065<br>Loglikelihood (fold 10/10):0.00649593<br>Average M loglikelihood:0.005060475</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72169",
      "postDate": "04/17/2015 21:44:00",
      "content": "<p>[quote]</p>\n<p>I see. Why don't you tune your model and maybe do some averaging... It helps ours. And you have time to do it!</p>\n<p>[/quote]</p>\n<p>Thanks for cheering up, that was the plan like 3 weeks ago, but I had to work overseas and just came back (still jet lagged), it seems that teaming up could have been a good choice XD, anyway, as gamers say: gl, hf and gg</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72208",
      "postDate": "04/18/2015 00:23:20",
      "content": "<p>I did not use this model within my submissions since it did not ensemble well with the *secret sauce*, but here's it anyways.</p>\n<p><img src=\"http://www.kaggle.com/blobs/download/forum-message-attachment-files/2338/44.png\" alt width=\"795\" height=\"311\"></p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 72114,
      "author_name": "mahmadi",
      "author_url": "",
      "post_date": "04/17/2015 17:30:22",
      "content": "<p>Would you please tell us about the number of features you used ?</p>\n<p>10 CV : 0.9968 , 0.0130, Number of features : Around 1200</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72115,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/17/2015 17:54:32",
      "content": "<p>We used about 7K features. and the best single model has cv 0.0035, and LB score 0.0037.</p>\n\n<p>Edit: the best single model has&nbsp;0.9992639 accuracy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72138,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "04/17/2015 19:52:36",
      "content": "<p>I build more than 1K features from the .asm files, but used only about ~400 useful features in my last model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72140,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/17/2015 20:01:51",
      "content": "<p>with FFFFFFFF + 9000 features, I think we are the ones who created most number of features ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72142,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "04/17/2015 20:06:46",
      "content": "<p>Ok..... something must be wrong with me haha, but I used 1.4M features, and my CV 0.0077 and LB 0.0073, this is without any parameter tuning (<a href=\"https://www.youtube.com/watch?v=8cT_Ulmcrys\">Ain't Nobody Got Time For That</a>) and single model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72144,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/17/2015 20:20:24",
      "content": "<p>[quote=NxGTR;72142]</p>\n<p>Ok..... something must be wrong with me haha, but I used 1.4M features, and my CV 0.0077 and LB 0.0073, this is without any parameter tuning (<a href=\"https://www.youtube.com/watch?v=8cT_Ulmcrys\">Ain't Nobody Got Time For That</a>) and single model.</p>\n<p>[/quote]</p>\n<p>Wow, how did you train your 1.4M features? Using VW?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72147,
      "author_name": "",
      "author_url": "",
      "post_date": "04/17/2015 20:32:23",
      "content": "<p>Our best so far has been an accuracy of 99.926 (8 samples misclassified) with 60+ features.</p>\n<p>But the log loss is very tricky. If we take what the classifier gives, we get a log loss of&nbsp;0.0048. If we threshold it, we get .002.</p>\n<p>If we give hard scores (only 8 are wrong), our log loss jumps to 0.02</p>\n<p>So I guess this competition is mainly about finding a good threshold for the classifier outputs. Depending on the threshold, anything can happen.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72150,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "04/17/2015 21:04:52",
      "content": "<p>[quote=Little Boat;72144]</p>\n<p>Wow, how did you train your 1.4M features? Using VW?</p>\n<p>[/quote]</p>\n<p>Well, I pretty much can train them with anything I want right now as I built the features to maximize the sparsity between classes, so the final dataset is extremely sparse and I can train it on my personal laptop.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72153,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/17/2015 21:17:25",
      "content": "<p>[quote=NxGTR;72150]</p>\n<p>[quote=Little Boat;72144]</p>\n<p>Wow, how did you train your 1.4M features? Using VW?</p>\n<p>[/quote]</p>\n<p>Well, I pretty much can train them with anything I want right now as I built the features to maximize the sparsity between classes, so the final dataset is extremely sparse and I can train it on my personal laptop.</p>\n<p>[/quote]</p>\n<p>I see. Why don't you tune your model and maybe do some averaging... It helps ours. And you have time to do it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72161,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "04/17/2015 21:34:00",
      "content": "<p>[quote=Abhishek;72140]</p>\n<p>with FFFFFFFF + 9000 features, I think we are the ones who created most number of features ;)</p>\n<p>[/quote]</p>\n<p>I have lost count but , but in our ensemble we have a couple of models with more than 400 K features, If I had to count, I would say we should have generated over 1 Million features.</p>\n<p>Our 10-fold 50/50 cv</p>\n<p>Loglikelihood (fold 1/10):0.0038178<br>Loglikelihood (fold 2/10):0.004158945<br>Loglikelihood (fold 3/10):0.003210165<br>Loglikelihood (fold 4/10):0.005466825<br>Loglikelihood (fold 5/10):0.00578151<br>Loglikelihood (fold 6/10):0.003658095<br>Loglikelihood (fold 7/10):0.004555845<br>Loglikelihood (fold 8/10):0.005502735<br>Loglikelihood (fold 9/10):0.007954065<br>Loglikelihood (fold 10/10):0.00649593<br>Average M loglikelihood:0.005060475</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72169,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "04/17/2015 21:44:00",
      "content": "<p>[quote]</p>\n<p>I see. Why don't you tune your model and maybe do some averaging... It helps ours. And you have time to do it!</p>\n<p>[/quote]</p>\n<p>Thanks for cheering up, that was the plan like 3 weeks ago, but I had to work overseas and just came back (still jet lagged), it seems that teaming up could have been a good choice XD, anyway, as gamers say: gl, hf and gg</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72208,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "04/18/2015 00:23:20",
      "content": "<p>I did not use this model within my submissions since it did not ensemble well with the *secret sauce*, but here's it anyways.</p>\n<p><img src=\"http://www.kaggle.com/blobs/download/forum-message-attachment-files/2338/44.png\" alt width=\"795\" height=\"311\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "72105": "",
    "72114": "",
    "72115": "",
    "72138": "",
    "72140": "",
    "72142": "",
    "72144": "",
    "72147": "",
    "72150": "",
    "72153": "",
    "72161": "",
    "72169": "",
    "72208": ""
  },
  "source": "meta"
}