{
  "id": 13291,
  "title": "Amazing performance",
  "url": "/competitions/malware-classification/discussion/13291",
  "author_name": "",
  "post_date": "2015-04-09T05:06:05.527Z",
  "votes": null,
  "comment_count": 35,
  "views": 7245,
  "content": "<p>Although private LB may tell a different story, it is still amazing that top 2 teams (almost) get all answers correct in public LB. I have only seen&nbsp;0 loss or 100% accuracy in Knowledge contests. &nbsp;Great work!</p>",
  "messages": [
    {
      "id": "70087",
      "postDate": "04/09/2015 05:06:05",
      "content": "<p>Although private LB may tell a different story, it is still amazing that top 2 teams (almost) get all answers correct in public LB. I have only seen&nbsp;0 loss or 100% accuracy in Knowledge contests. &nbsp;Great work!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70091",
      "postDate": "04/09/2015 06:48:31",
      "content": "<p>0.0000000 log loss. That's something you don't see every day.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70092",
      "postDate": "04/09/2015 06:59:40",
      "content": "<p style=\"text-align: left\">Hand labelling, I bet!!!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70093",
      "postDate": "04/09/2015 07:04:19",
      "content": "<p>Can we have the hand labelling classification rules please :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70094",
      "postDate": "04/09/2015 07:06:39",
      "content": "<p>With so few outliers like 15 or 20, hand labeling would&nbsp;be no different from&nbsp;calibration. So I don't think it is against any rules. or I can not tell the difference.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70095",
      "postDate": "04/09/2015 07:13:01",
      "content": "<p>the rules do say:</p>\n<p>&quot;Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.&quot;</p>\n<p>I don't know why it says 'may' instead of 'should/shall'&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70096",
      "postDate": "04/09/2015 07:21:59",
      "content": "<p>how is hand labeling defined? If, for example, I look through the dataset and see that&nbsp;some class contains a string in the asm files that is unique to it. Is labeling all files that contain that string as belonging to that class considered hand labeling or not?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70098",
      "postDate": "04/09/2015 07:29:08",
      "content": "<p>Think this is called &quot;signature based detection&quot; not hand labeling :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70113",
      "postDate": "04/09/2015 13:40:57",
      "content": "<p>I'm going to bet there will be a massive shift from public to private.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70125",
      "postDate": "04/09/2015 15:20:55",
      "content": "<p>I bet the top 2 will remain their super performance in private, and they will tie in 0.000001!&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70129",
      "postDate": "04/09/2015 17:38:01",
      "content": "<p>what about the 2 after top2 ? :P</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70142",
      "postDate": "04/09/2015 19:12:18",
      "content": "<p>[quote=Abhishek;70129]</p>\n<p>what about the 2 after top2 ? :P</p>\n<p>[/quote]</p>\n<p>I have no idea. Maybe we could benefit from a new round of CV sharing. Our best signal model achieves 4-fold cv of 0.0051 and 0.0044 public LB. What about yours? :P</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70145",
      "postDate": "04/09/2015 19:46:46",
      "content": "<p>I'm in the same ballpark -- .0052 10-fold CV, .0041 LB.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70146",
      "postDate": "04/09/2015 19:51:33",
      "content": "<p>Thank you all for sharing :D</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70175",
      "postDate": "04/10/2015 06:13:51",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70193",
      "postDate": "04/10/2015 09:25:21",
      "content": "<p>Lol</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70194",
      "postDate": "04/10/2015 09:38:38",
      "content": "<p>Don't be disappointed soon Sachin, Go on.</p>\n<p>There are still 7 days to the end ;-)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71271",
      "postDate": "04/11/2015 07:37:37",
      "content": "<p>[quote=Abhishek;70193]</p>\n<p>Lol</p>\n<p>[/quote]<br><br></p>\n<p>Hi Abhishek,</p>\n<p>Can you please explain if someone has got a score of 0 on the public leaderboard and he chooses 2 best entries with scores 0 and 0 how can private leaderboard score be different. Since private data is never revealed to the participants and we are not asked to submit our code, how can they determine the results on private data if it is not a subset of public data with a different mix. Or am I totally missing something?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71273",
      "postDate": "04/11/2015 08:17:43",
      "content": "<p>@Sachin: you submit 10k predictions. Some 3k (fixed on kaggle side) used for calculate Public LB. Other 7k will be used for calculating Private score.</p>\n<p>So, there is no problem.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71275",
      "postDate": "04/11/2015 08:38:32",
      "content": "<p>[quote=Mikhail Trofimov;71273]</p>\n<p>@Sachin: you submit 10k predictions. Some 3k (fixed on kaggle side) used for calculate Public LB. Other 7k will be used for calculating Private score.</p>\n<p>So, there is no problem.</p>\n<p>[/quote]</p>\n<p>Thanks Mikhail</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71679",
      "postDate": "04/15/2015 11:39:11",
      "content": "<p>[quote=Abhishek;70092]</p>\n<p style=\"text-align: left\">Hand labelling, I bet!!!</p>\n<p>[/quote]</p>\n\n<p>Do you still bet on that one? :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71680",
      "postDate": "04/15/2015 11:52:42",
      "content": "<p>[quote=Stergios;71679]</p>\n<p>Do you still bet on that one? :)</p>\n<p>[/quote]</p>\n<p>Abhishek would never do something like that :p . It is not in his mentality .&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71681",
      "postDate": "04/15/2015 11:55:37",
      "content": "<p>[quote=Stergios;71679]</p>\n<p>[quote=Abhishek;70092]</p>\n<p style=\"text-align: left\">Hand labelling, I bet!!!</p>\n<p>[/quote]</p>\n<p>Do you still bet on that one? :)</p>\n<p>[/quote]</p>\n<p>$10 bucks says its handlabeling for all teams above us.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71684",
      "postDate": "04/15/2015 12:29:02",
      "content": "<p>Not necessarily (---with the exception of&nbsp; two submission I did my accident and one that was in the incorrect format--) every submission I have had a significant jump in my ranking ...</p>\n<p>eg.&nbsp; my initial post, I had a 98.34% accuracy on my classifier and was place 9th from last place.</p>\n<p>I was able to move to 98.74% accuracy on my classifier and jump only an additional 2 levels in my ranking...</p>\n<p>My most recent submission I had a little better than 98.8% on my classifier and at my submission I was surprised to find out I had jumped in ranking by, if I recall correctly, 97 point in ranking....</p>\n<p>So my guess that the people/group in the top ranking is in the 99% in whatever classification method they are using...</p>\n<p>I have been using an unproven theory about ML that seems to be working for me.</p>\n<p>A few hours ago I realized, in my rush, to join the competition, I had introduced a lot of noise into my testing and training datasets.... It could be the reason I could not seem to get very far beyond 98.8% on my classification system. --I am in the process of removing the noise and hopefully push beyond the 98.8% accuracy</p>\n<p>This is my first competition ever, I am blown away and humbled at the fact that all the people above my ranking probably have classification system doing better that 98.8% and closely packed together if someone was to actually look at the raw numbers... </p>\n<p>I think if I or anyone else can get to the 99.9% they maybe able to take over the top spot...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71686",
      "postDate": "04/15/2015 12:33:31",
      "content": "<p>I think everyone in top 10 already achieved 99.9% accuracy.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71687",
      "postDate": "04/15/2015 12:44:07",
      "content": "<p>Ok, so our numbers are, don't think it hurts to share:</p>\n<p>CV 10 folds, overall accuracy 0.9993558,&nbsp; avg log loss 0.0015 ... 0.00291786719343801 depending on the seed...</p>\n<p>Rounded high probs (&gt;0.9999) to 1 not to loose on log loss.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71693",
      "postDate": "04/15/2015 14:45:55",
      "content": "<p>@Michael, the ranking has everything to do with how the evaluation metric, the average log loss, is being calculated. If I'm doing my math right, even if you have 100.00% accuracy, you only get &nbsp;a LB score of 1.0 of you only put 0.367 probabilities on the correct class. &nbsp;To get a LB score of 0.01, you would need to put 0.99 probabilities on the correct class. The trick then is hedging against the misclassifications, because you're heavily penalized for putting too little on the correct class.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71717",
      "postDate": "04/15/2015 15:54:07",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;71680]</p>\n<p>[quote=Stergios;71679]</p>\n<p>Do you still bet on that one? :)</p>\n<p>[/quote]</p>\n<p>Abhishek would never do something like that :p . It is not in his mentality .&nbsp;</p>\n<p>[/quote]</p>\n<p>Here's a question: is hand labeling or non algorithmic ex-post processing of&nbsp;submission files inherently cheating? Or is it only if you end up choosing that submission as your final model? I mean, would it be wrong to hand label some observations, get LB feedback, and use that information to legitimately (aka algorithmically) tune your model? Furthermore, how is hand labeling of a specific observation different from hand labeling of clusters of observations?&nbsp;Take for example competitions where the test set is chronologically following the train set. There is usually a trend component that is difficult to model. In&nbsp;KDD 2014 for example, many teams, including mine,&nbsp;multiplied chunks of predictions by certain factors to account for time trends that would have otherwise been impossible to model.&nbsp;That is actually a very common Kaggle trick, used in many competitions, including SeeClick, KDD, Influenza,... In those cases the output of a base model is hand crafted before submission. Is that cheating? It hasn't been in the past.</p>\n<p>I think hand labeling in this competition is a touchy topic. The accuracy is so high (and models so good) that you can legitimately hone down the problematic observations by pure modelling. It is not like Cats &amp; Dogs where you could score a perfect submission w/o having any model. So, with probably a handful of uncertain observations, would it be considered cheating to just try and manually change predictions to 1's and 0's and gather LB feedback?</p>\n<p>Don't want to seem arrogant or anything, especially because I'm far from the top. But if I had a model that allowed me to hone down the remainder logloss to a handful of observations, I would hand label those for the sake of observing LB feedback. Would I choose one of those submissions as my final 2? Heck no. But would I use that information to improve my models if possible? Upvote this post if you want the answer (but you should know the answer :-) ).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71765",
      "postDate": "04/15/2015 18:43:13",
      "content": "<p>[quote=Giulio;71717]</p>\n<p>Don't want to seem arrogant or anything, especially because I'm far from the top. But if I had a model that allowed me to hone down the remainder logloss to a handful of observations, I would hand label those for the sake of observing LB feedback. Would I choose one of those submissions as my final 2? Heck no. But would I use that information to improve my models if possible? Upvote this post if you want the answer (but you should know the answer :-) ).</p>\n<p>[/quote]</p>\n<p>For me, that is cheating (aka I would never do that). Given the rather small size of the data set, including 30-40 difficult-to-score hand-labelled observation can make some impact. Anyway, I believe there will be shake-up due to this kind of over fitting- we'll see . &nbsp;</p>\n<p>I think there is a clear distinction in getting leaderboard feedback (even probing) and (-while having a relatively small training set-) to increase your set of known labels artificially like that.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71779",
      "postDate": "04/15/2015 20:12:42",
      "content": "<p>For me, this is a grey area. &nbsp;The top 3 teams do have an upper hand now. Predicting outliers largely depend on outliers we know. Now they know outliers in public LB and we don't yet. I expect them will benefit from this information.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72093",
      "postDate": "04/17/2015 15:22:36",
      "content": "<p>@David Shinn</p>\n<p>I am using the softmax function across the output of my classifier so those .999 in the right place you mentioned is automatically done for me.</p>\n<p>My one problem is that I did not stick with what I know and used gradient decent algorithm to optimize my system ..... I am really having a hard time getting beyond 98.8.....</p>\n<p>I truly believe when I finally properly tune the gradient decent optimizer and get into the 99.9xyz I will surely blow the top of the leaders....</p>\n<p>Hopefully I can do it before the end of the competition that occurs in about 8 hours.....</p>\n<p>It was only yesterday I was able to get the noise out the the datasets; something I introduced by accident :(</p>\n<p>All I need is time .... It is really unfortunately be at Kaggle for two week and just one week ago decide to participate.....</p>\n<p>Here is to pulling something out of my hat and get beyond the 98.8% on the cleaned up datasets</p>\n<p>;)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72106",
      "postDate": "04/17/2015 17:01:09",
      "content": "<p>@Michael,</p>\n<p>Good luck to you. &nbsp;Even if you achieved 99.9% accuracy on the Public LB, if you're putting too little probabilities on your misclassifications, you still won't break the top 10. &nbsp;I estimate that at best one&nbsp;can achieve an average log_loss of around 0.0069 on the Public LB if one has a 99.9% accuracy &nbsp;with 1.00 probabilities on the correct classes but put 0.001 probabilities on the&nbsp;misclassifications. &nbsp;See my calcs below. &nbsp;I've learned a lot about the log_loss metric this competition.</p>\n<p><code>In [22]: import math</code></p>\n<p><code></code><code>In [23]: (1 - 0.999)*3000 # Get estimated counts of misclassified at 99.9% accuracy<br>Out[23]: 3.0000000000000027</code></p>\n<p><code>In [24]: -math.log(0.001)*3 / 3000 # Get estimated impact on score if put 0.001 prob on those misclassifications<br>Out[24]: 0.006907755278982137</code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72126",
      "postDate": "04/17/2015 18:37:41",
      "content": "<p>Okay, your calculated probability seems justified...</p>\n<p>That said, do you think it is by pure luck the leader have such good scores....</p>\n<p>Actually I just notice that half the top 10 have more than 128 submission and realized once you get to certain level of of classification score say 98.8% accuracy in classification on the training data you could actually use the submission process to weed out incorrect classification within each submission (classic probability/statistical analysis) .... I think seem like a possible way to boost ones score beyond brute force training of ones system....</p>\n<p>I think this maybe a reasonable thing I/we/competitors need to invest some time in automating. Does such a system as describe exist?</p>\n<p>Is this what people mean by labelling?&nbsp; It is the first time I have heard of it. Is/are there any papers on the subject?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72141",
      "postDate": "04/17/2015 20:03:46",
      "content": "<p>[quote=Michael George Hart;72126]</p>\n<p>Okay, your calculated probability seems justified...</p>\n<p>That said, do you think it is by pure luck the leader have such good scores....</p>\n<p>Actually I just notice that half the top 10 have more than 128 submission and realized once you get to certain level of of classification score say 98.8% accuracy in classification on the training data you could actually use the submission process to weed out incorrect classification within each submission (classic probability/statistical analysis) .... I think seem like a possible way to boost ones score beyond brute force training of ones system....</p>\n<p>I think this maybe a reasonable thing I/we/competitors need to invest some time in automating. Does such a system as describe exist?</p>\n<p>Is this what people mean by labelling?&nbsp; It is the first time I have heard of it. Is/are there any papers on the subject?</p>\n<p>[/quote]</p>\n<p>You can easily get to 99% accuracy without any calibration or hand labelling. Just a single model with some fair features can do it. The problem here is not accuracy, but how to optimize the logloss with great generalization. And the reason we have many submissions is mainly because we entered early and felt very comfortable to see even tiny improvement on the LB, even though we know it most likely may not show in the private board.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72211",
      "postDate": "04/18/2015 00:31:51",
      "content": "<p>I knew when I saw the 0.0000 logloss there would be massive shift between public and private. It's not arrogance. It's just very easy to get one mistake and be pounded hard (not infinitely but very hard) for it. Calibration is hard. After Otto I'd like to release some secret sauce with respect to this.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72217",
      "postDate": "04/18/2015 00:45:09",
      "content": "<p style=\"text-align: left\">Its not about 0.00. We never used 0.0. It was done only to show that its possible. There were a lot of other reasons for the drop. And we will release everything soom</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 70091,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "04/09/2015 06:48:31",
      "content": "<p>0.0000000 log loss. That's something you don't see every day.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70092,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/09/2015 06:59:40",
      "content": "<p style=\"text-align: left\">Hand labelling, I bet!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70093,
      "author_name": "drpatrickchan",
      "author_url": "",
      "post_date": "04/09/2015 07:04:19",
      "content": "<p>Can we have the hand labelling classification rules please :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70094,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/09/2015 07:06:39",
      "content": "<p>With so few outliers like 15 or 20, hand labeling would&nbsp;be no different from&nbsp;calibration. So I don't think it is against any rules. or I can not tell the difference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70095,
      "author_name": "",
      "author_url": "",
      "post_date": "04/09/2015 07:13:01",
      "content": "<p>the rules do say:</p>\n<p>&quot;Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.&quot;</p>\n<p>I don't know why it says 'may' instead of 'should/shall'&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70096,
      "author_name": "izikgo",
      "author_url": "",
      "post_date": "04/09/2015 07:21:59",
      "content": "<p>how is hand labeling defined? If, for example, I look through the dataset and see that&nbsp;some class contains a string in the asm files that is unique to it. Is labeling all files that contain that string as belonging to that class considered hand labeling or not?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70098,
      "author_name": "gmilosev",
      "author_url": "",
      "post_date": "04/09/2015 07:29:08",
      "content": "<p>Think this is called &quot;signature based detection&quot; not hand labeling :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70113,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "04/09/2015 13:40:57",
      "content": "<p>I'm going to bet there will be a massive shift from public to private.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70125,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/09/2015 15:20:55",
      "content": "<p>I bet the top 2 will remain their super performance in private, and they will tie in 0.000001!&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70129,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/09/2015 17:38:01",
      "content": "<p>what about the 2 after top2 ? :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70142,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/09/2015 19:12:18",
      "content": "<p>[quote=Abhishek;70129]</p>\n<p>what about the 2 after top2 ? :P</p>\n<p>[/quote]</p>\n<p>I have no idea. Maybe we could benefit from a new round of CV sharing. Our best signal model achieves 4-fold cv of 0.0051 and 0.0044 public LB. What about yours? :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70145,
      "author_name": "psilogram",
      "author_url": "",
      "post_date": "04/09/2015 19:46:46",
      "content": "<p>I'm in the same ballpark -- .0052 10-fold CV, .0041 LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70146,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/09/2015 19:51:33",
      "content": "<p>Thank you all for sharing :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70175,
      "author_name": "sachinvernekar",
      "author_url": "",
      "post_date": "04/10/2015 06:13:51",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 70193,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/10/2015 09:25:21",
      "content": "<p>Lol</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70194,
      "author_name": "mahmadi",
      "author_url": "",
      "post_date": "04/10/2015 09:38:38",
      "content": "<p>Don't be disappointed soon Sachin, Go on.</p>\n<p>There are still 7 days to the end ;-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71271,
      "author_name": "sachinvernekar",
      "author_url": "",
      "post_date": "04/11/2015 07:37:37",
      "content": "<p>[quote=Abhishek;70193]</p>\n<p>Lol</p>\n<p>[/quote]<br><br></p>\n<p>Hi Abhishek,</p>\n<p>Can you please explain if someone has got a score of 0 on the public leaderboard and he chooses 2 best entries with scores 0 and 0 how can private leaderboard score be different. Since private data is never revealed to the participants and we are not asked to submit our code, how can they determine the results on private data if it is not a subset of public data with a different mix. Or am I totally missing something?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71273,
      "author_name": "mikhailtrofimov",
      "author_url": "",
      "post_date": "04/11/2015 08:17:43",
      "content": "<p>@Sachin: you submit 10k predictions. Some 3k (fixed on kaggle side) used for calculate Public LB. Other 7k will be used for calculating Private score.</p>\n<p>So, there is no problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71275,
      "author_name": "sachinvernekar",
      "author_url": "",
      "post_date": "04/11/2015 08:38:32",
      "content": "<p>[quote=Mikhail Trofimov;71273]</p>\n<p>@Sachin: you submit 10k predictions. Some 3k (fixed on kaggle side) used for calculate Public LB. Other 7k will be used for calculating Private score.</p>\n<p>So, there is no problem.</p>\n<p>[/quote]</p>\n<p>Thanks Mikhail</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71679,
      "author_name": "asterios",
      "author_url": "",
      "post_date": "04/15/2015 11:39:11",
      "content": "<p>[quote=Abhishek;70092]</p>\n<p style=\"text-align: left\">Hand labelling, I bet!!!</p>\n<p>[/quote]</p>\n\n<p>Do you still bet on that one? :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71680,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "04/15/2015 11:52:42",
      "content": "<p>[quote=Stergios;71679]</p>\n<p>Do you still bet on that one? :)</p>\n<p>[/quote]</p>\n<p>Abhishek would never do something like that :p . It is not in his mentality .&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71681,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/15/2015 11:55:37",
      "content": "<p>[quote=Stergios;71679]</p>\n<p>[quote=Abhishek;70092]</p>\n<p style=\"text-align: left\">Hand labelling, I bet!!!</p>\n<p>[/quote]</p>\n<p>Do you still bet on that one? :)</p>\n<p>[/quote]</p>\n<p>$10 bucks says its handlabeling for all teams above us.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71684,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/15/2015 12:29:02",
      "content": "<p>Not necessarily (---with the exception of&nbsp; two submission I did my accident and one that was in the incorrect format--) every submission I have had a significant jump in my ranking ...</p>\n<p>eg.&nbsp; my initial post, I had a 98.34% accuracy on my classifier and was place 9th from last place.</p>\n<p>I was able to move to 98.74% accuracy on my classifier and jump only an additional 2 levels in my ranking...</p>\n<p>My most recent submission I had a little better than 98.8% on my classifier and at my submission I was surprised to find out I had jumped in ranking by, if I recall correctly, 97 point in ranking....</p>\n<p>So my guess that the people/group in the top ranking is in the 99% in whatever classification method they are using...</p>\n<p>I have been using an unproven theory about ML that seems to be working for me.</p>\n<p>A few hours ago I realized, in my rush, to join the competition, I had introduced a lot of noise into my testing and training datasets.... It could be the reason I could not seem to get very far beyond 98.8% on my classification system. --I am in the process of removing the noise and hopefully push beyond the 98.8% accuracy</p>\n<p>This is my first competition ever, I am blown away and humbled at the fact that all the people above my ranking probably have classification system doing better that 98.8% and closely packed together if someone was to actually look at the raw numbers... </p>\n<p>I think if I or anyone else can get to the 99.9% they maybe able to take over the top spot...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71686,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/15/2015 12:33:31",
      "content": "<p>I think everyone in top 10 already achieved 99.9% accuracy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71687,
      "author_name": "gmilosev",
      "author_url": "",
      "post_date": "04/15/2015 12:44:07",
      "content": "<p>Ok, so our numbers are, don't think it hurts to share:</p>\n<p>CV 10 folds, overall accuracy 0.9993558,&nbsp; avg log loss 0.0015 ... 0.00291786719343801 depending on the seed...</p>\n<p>Rounded high probs (&gt;0.9999) to 1 not to loose on log loss.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71693,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "04/15/2015 14:45:55",
      "content": "<p>@Michael, the ranking has everything to do with how the evaluation metric, the average log loss, is being calculated. If I'm doing my math right, even if you have 100.00% accuracy, you only get &nbsp;a LB score of 1.0 of you only put 0.367 probabilities on the correct class. &nbsp;To get a LB score of 0.01, you would need to put 0.99 probabilities on the correct class. The trick then is hedging against the misclassifications, because you're heavily penalized for putting too little on the correct class.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71717,
      "author_name": "adjgiulio",
      "author_url": "",
      "post_date": "04/15/2015 15:54:07",
      "content": "<p>[quote=&#924;&#945;&#961;&#953;&#959;&#962; &#924;&#953;&#967;&#945;&#951;&#955;&#953;&#948;&#951;&#962; KazAnova;71680]</p>\n<p>[quote=Stergios;71679]</p>\n<p>Do you still bet on that one? :)</p>\n<p>[/quote]</p>\n<p>Abhishek would never do something like that :p . It is not in his mentality .&nbsp;</p>\n<p>[/quote]</p>\n<p>Here's a question: is hand labeling or non algorithmic ex-post processing of&nbsp;submission files inherently cheating? Or is it only if you end up choosing that submission as your final model? I mean, would it be wrong to hand label some observations, get LB feedback, and use that information to legitimately (aka algorithmically) tune your model? Furthermore, how is hand labeling of a specific observation different from hand labeling of clusters of observations?&nbsp;Take for example competitions where the test set is chronologically following the train set. There is usually a trend component that is difficult to model. In&nbsp;KDD 2014 for example, many teams, including mine,&nbsp;multiplied chunks of predictions by certain factors to account for time trends that would have otherwise been impossible to model.&nbsp;That is actually a very common Kaggle trick, used in many competitions, including SeeClick, KDD, Influenza,... In those cases the output of a base model is hand crafted before submission. Is that cheating? It hasn't been in the past.</p>\n<p>I think hand labeling in this competition is a touchy topic. The accuracy is so high (and models so good) that you can legitimately hone down the problematic observations by pure modelling. It is not like Cats &amp; Dogs where you could score a perfect submission w/o having any model. So, with probably a handful of uncertain observations, would it be considered cheating to just try and manually change predictions to 1's and 0's and gather LB feedback?</p>\n<p>Don't want to seem arrogant or anything, especially because I'm far from the top. But if I had a model that allowed me to hone down the remainder logloss to a handful of observations, I would hand label those for the sake of observing LB feedback. Would I choose one of those submissions as my final 2? Heck no. But would I use that information to improve my models if possible? Upvote this post if you want the answer (but you should know the answer :-) ).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71765,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "04/15/2015 18:43:13",
      "content": "<p>[quote=Giulio;71717]</p>\n<p>Don't want to seem arrogant or anything, especially because I'm far from the top. But if I had a model that allowed me to hone down the remainder logloss to a handful of observations, I would hand label those for the sake of observing LB feedback. Would I choose one of those submissions as my final 2? Heck no. But would I use that information to improve my models if possible? Upvote this post if you want the answer (but you should know the answer :-) ).</p>\n<p>[/quote]</p>\n<p>For me, that is cheating (aka I would never do that). Given the rather small size of the data set, including 30-40 difficult-to-score hand-labelled observation can make some impact. Anyway, I believe there will be shake-up due to this kind of over fitting- we'll see . &nbsp;</p>\n<p>I think there is a clear distinction in getting leaderboard feedback (even probing) and (-while having a relatively small training set-) to increase your set of known labels artificially like that.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71779,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/15/2015 20:12:42",
      "content": "<p>For me, this is a grey area. &nbsp;The top 3 teams do have an upper hand now. Predicting outliers largely depend on outliers we know. Now they know outliers in public LB and we don't yet. I expect them will benefit from this information.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72093,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/17/2015 15:22:36",
      "content": "<p>@David Shinn</p>\n<p>I am using the softmax function across the output of my classifier so those .999 in the right place you mentioned is automatically done for me.</p>\n<p>My one problem is that I did not stick with what I know and used gradient decent algorithm to optimize my system ..... I am really having a hard time getting beyond 98.8.....</p>\n<p>I truly believe when I finally properly tune the gradient decent optimizer and get into the 99.9xyz I will surely blow the top of the leaders....</p>\n<p>Hopefully I can do it before the end of the competition that occurs in about 8 hours.....</p>\n<p>It was only yesterday I was able to get the noise out the the datasets; something I introduced by accident :(</p>\n<p>All I need is time .... It is really unfortunately be at Kaggle for two week and just one week ago decide to participate.....</p>\n<p>Here is to pulling something out of my hat and get beyond the 98.8% on the cleaned up datasets</p>\n<p>;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72106,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "04/17/2015 17:01:09",
      "content": "<p>@Michael,</p>\n<p>Good luck to you. &nbsp;Even if you achieved 99.9% accuracy on the Public LB, if you're putting too little probabilities on your misclassifications, you still won't break the top 10. &nbsp;I estimate that at best one&nbsp;can achieve an average log_loss of around 0.0069 on the Public LB if one has a 99.9% accuracy &nbsp;with 1.00 probabilities on the correct classes but put 0.001 probabilities on the&nbsp;misclassifications. &nbsp;See my calcs below. &nbsp;I've learned a lot about the log_loss metric this competition.</p>\n<p><code>In [22]: import math</code></p>\n<p><code></code><code>In [23]: (1 - 0.999)*3000 # Get estimated counts of misclassified at 99.9% accuracy<br>Out[23]: 3.0000000000000027</code></p>\n<p><code>In [24]: -math.log(0.001)*3 / 3000 # Get estimated impact on score if put 0.001 prob on those misclassifications<br>Out[24]: 0.006907755278982137</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72126,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/17/2015 18:37:41",
      "content": "<p>Okay, your calculated probability seems justified...</p>\n<p>That said, do you think it is by pure luck the leader have such good scores....</p>\n<p>Actually I just notice that half the top 10 have more than 128 submission and realized once you get to certain level of of classification score say 98.8% accuracy in classification on the training data you could actually use the submission process to weed out incorrect classification within each submission (classic probability/statistical analysis) .... I think seem like a possible way to boost ones score beyond brute force training of ones system....</p>\n<p>I think this maybe a reasonable thing I/we/competitors need to invest some time in automating. Does such a system as describe exist?</p>\n<p>Is this what people mean by labelling?&nbsp; It is the first time I have heard of it. Is/are there any papers on the subject?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72141,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "04/17/2015 20:03:46",
      "content": "<p>[quote=Michael George Hart;72126]</p>\n<p>Okay, your calculated probability seems justified...</p>\n<p>That said, do you think it is by pure luck the leader have such good scores....</p>\n<p>Actually I just notice that half the top 10 have more than 128 submission and realized once you get to certain level of of classification score say 98.8% accuracy in classification on the training data you could actually use the submission process to weed out incorrect classification within each submission (classic probability/statistical analysis) .... I think seem like a possible way to boost ones score beyond brute force training of ones system....</p>\n<p>I think this maybe a reasonable thing I/we/competitors need to invest some time in automating. Does such a system as describe exist?</p>\n<p>Is this what people mean by labelling?&nbsp; It is the first time I have heard of it. Is/are there any papers on the subject?</p>\n<p>[/quote]</p>\n<p>You can easily get to 99% accuracy without any calibration or hand labelling. Just a single model with some fair features can do it. The problem here is not accuracy, but how to optimize the logloss with great generalization. And the reason we have many submissions is mainly because we entered early and felt very comfortable to see even tiny improvement on the LB, even though we know it most likely may not show in the private board.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72211,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "04/18/2015 00:31:51",
      "content": "<p>I knew when I saw the 0.0000 logloss there would be massive shift between public and private. It's not arrogance. It's just very easy to get one mistake and be pounded hard (not infinitely but very hard) for it. Calibration is hard. After Otto I'd like to release some secret sauce with respect to this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72217,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "04/18/2015 00:45:09",
      "content": "<p style=\"text-align: left\">Its not about 0.00. We never used 0.0. It was done only to show that its possible. There were a lot of other reasons for the drop. And we will release everything soom</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "70087": "",
    "70091": "",
    "70092": "",
    "70093": "",
    "70094": "",
    "70095": "",
    "70096": "",
    "70098": "",
    "70113": "",
    "70125": "",
    "70129": "",
    "70142": "",
    "70145": "",
    "70146": "",
    "70175": "",
    "70193": "",
    "70194": "",
    "71271": "",
    "71273": "",
    "71275": "",
    "71679": "",
    "71680": "",
    "71681": "",
    "71684": "",
    "71686": "",
    "71687": "",
    "71693": "",
    "71717": "",
    "71765": "",
    "71779": "",
    "72093": "",
    "72106": "",
    "72126": "",
    "72141": "",
    "72211": "",
    "72217": ""
  },
  "source": "meta"
}