{
  "id": 4347,
  "title": "Meta Questions",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4347",
  "author_name": "",
  "post_date": "2013-04-17T03:23:09.377Z",
  "votes": null,
  "comment_count": 13,
  "views": 3279,
  "content": "<p>Seems interesting enough to get drawn into. I want to force myself to learn python, but I'll probably fall back on R a lot. I did have a few questions:</p>\r\n<p>Why use accuracy as the evaluation metric? It's got to be the least efficient evaluation metric out there. I'm always partial to (multinomial-) log-loss provided the testing data is sampled from the same distribution as the training data, but I'm sure there's\r\n some multinomial extension of AUC as well.</p>\r\n<p>What is up with the &quot;#.0&quot; class labels? Are we just trying to give the newbies a hard time? Kaggle's parses are usually written to be highly forgiving.</p>\r\n<p>Any restrictions on what we do with the test data? Can I toss it in the unsupervised work etc.?</p>\r\n<p>Why does this forum code still strip all my whitespace when using an up-to-date Chrome on Windows?</p>",
  "messages": [
    {
      "id": "22937",
      "postDate": "04/17/2013 03:23:09",
      "content": "<p>Seems interesting enough to get drawn into. I want to force myself to learn python, but I'll probably fall back on R a lot. I did have a few questions:</p>\r\n<p>Why use accuracy as the evaluation metric? It's got to be the least efficient evaluation metric out there. I'm always partial to (multinomial-) log-loss provided the testing data is sampled from the same distribution as the training data, but I'm sure there's\r\n some multinomial extension of AUC as well.</p>\r\n<p>What is up with the &quot;#.0&quot; class labels? Are we just trying to give the newbies a hard time? Kaggle's parses are usually written to be highly forgiving.</p>\r\n<p>Any restrictions on what we do with the test data? Can I toss it in the unsupervised work etc.?</p>\r\n<p>Why does this forum code still strip all my whitespace when using an up-to-date Chrome on Windows?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22940",
      "postDate": "04/17/2013 03:54:50",
      "content": "<p>I would also like to know if we can include test set data in the unsupervised learning.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22941",
      "postDate": "04/17/2013 04:18:26",
      "content": "<ol>\r\n<li><span style=\"line-height:1.4em\">You don't need to learn any Python for this, obviously :)</span>\r\n</li><li><span style=\"line-height:1.4em\">Accuracy on the underlying task (which we'll disclose at the end of the competition) is what we really care about to be honest, at the end of the day. For this task, accuracy and multi-class log-loss correlate pretty well,\r\n though.</span> </li><li>There are other reasons to prefer accuracy to logprob(correct_class) for evaluation, namely that the loss for a given example is actually bounded (as opposed to log-loss, which can be arbitrarily large for a particular example).\r\n</li><li><span style=\"line-height:1.4em\">AUC for multi-class problems is a little awkard (see http://users.dsic.upv.es/grupos/elp/cferri/vus-ecml03-camera-ready4.pdf for a treatment). Not even sure kaggle implements this.</span>\r\n</li><li><span style=\"line-height:1.4em\">As far as using the test data for unsupervised learning... I guess we don't cover this by the rules. I'll chat with the other organizers and get back to you.</span>\r\n</li></ol>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22959",
      "postDate": "04/17/2013 12:38:48",
      "content": "<p>Follow up to Dumitru's post:</p>\r\n<p>3. Following up on what Dumitru said, in my experience this effect leads the rankings given by log loss to not be very statistically robust. Because a model can lose an unbounded amount of log likelihood based on its output for a single example, the ranking\r\n of two models is often driven by outliers.</p>\r\n<p>5. Yoshua and I are both fine with allowing the test data for unsupervised learning.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22965",
      "postDate": "04/17/2013 13:36:58",
      "content": "Thanks for the responses. As for the robustness of log-loss, I've never ran into that issue. If a competitor is stupid enough to estimate something is 1e-8 likely to happen and it does, then they deserve to lose. In practical competitions and real work,\r\n I seldom see class predictions below 0.5%. &quot;Strange events permit themselves the luxury of occurring.&quot; I also understand that log-loss and accuracy are correlated. My actual complaint was about the efficiency of accuracy. Basically, I argue that the determination\r\n of the winner would be more consistent under log-loss given the limited amount of testing data. Said another way, the winner would be more likely to remain the winner using log-loss even if you expanding the testing sample to 100k observations. This also brings\r\n up the question of why did you choose a 50/50 split of public/private? I would think the accuracy of the private set would be more important than the accuracy of the public set.",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22967",
      "postDate": "04/17/2013 14:31:01",
      "content": "<p>The public and private set are the same size: http://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/data</p>\r\n<p><span style=\"line-height:1.4em\">I actually see this all the time, and it's not driven by wildly overconfident models assigning epsilon probability to black swan events. Algorithms that get 15% error on the CIFAR-10 dataset often get better log likelihood\r\n scores than algorithms that get under 10% error. The issue isn't that the more accurate model assigns a probability of 1e-8 to a single example. The issue is that the less accurate model has low confidence in a large number of mistakes while the more accurate\r\n model has medium confidence in a small number of mistakes.</span></p>\r\n<p>Both metrics have their uses for different applications. In my opinion, log likelihood is usually the better metric for a sub-component of a system, like the observation model of an HMM. Overconfidence is a more serious flaw in subcomponents because they\r\n make the larger system unaware of multiple possibilities. Accuracy is usually the better metric for the final output of a system. In most practical applications, you have to use your system to commit to a single action (do you give the patient a C-section\r\n or not?), and you don't get bonus points for doing the wrong thing with low confidence.</p>\r\n<p>As for which one generalizes better to a larger test set, I'd like to see a theorem or some empirical work. Log likelihood gives you a real number per example, instead of just one bit, but it's also prone to being driven by outliers. I suspect that if you\r\n have extremely few labels the first effect dominates but if you have a medium amount of labels like we do, the second dominates.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22975",
      "postDate": "04/17/2013 15:18:34",
      "content": "I did realize the public and private were the same size; I asked why they were the same size. I would have assumed the accuracy of the private set would be more important than the public set. I also agree that there are times you might want to report/maximize\r\n accuracy. However, given the reality of limited testing, I'd want to pick a metric that stayed consistent were the testing volume expanded. I don't have evidence at hand supporting the efficiency advantage of log-loss, but I will look for some in the near\r\n future. Additionally, I guess my own training/work is focused on &quot;quantifying uncertainty&quot; so I just can't fathom not wanting to know probabilistic predictions. I also don't put credit in the outlier argument because even you agree that nobody should assign\r\n 1e-8 probability to any outcome. I don't think log-loss is perfect by any means; I'm a big proponent of rank-based evaluation metrics when the testing data is not sampled from the same distribution as the training data. Speaking of which, will you disclose\r\n if the test and train data are sampled from the same population for this contest?",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22976",
      "postDate": "04/17/2013 15:22:10",
      "content": "A brief Google search at least turns up this stepping stone into loss-metric efficiency: http://cling.csd.uwo.ca/papers/ijcai03.pdf",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22979",
      "postDate": "04/17/2013 15:25:17",
      "content": "<p>Yes, that's AUC, not log loss, and it only applies to the single class case. Dumitru already explained why we don't want to use AUC for a multiclass problem.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22984",
      "postDate": "04/17/2013 16:55:51",
      "content": "<p>I was not proposing to use AUC (or some multinomial extension) I was merely putting forward an article that discussed loss-metric efficiency in formal terms.&nbsp; On an empricial front, I present this simple bit of R code.&nbsp; From it we can see that a simple ridged\r\n linear model produces a smooth bias-variance curve when considering log-loss, but a bumpy roller-coaster ride for accuracy.&nbsp; We also see that a classic linear model seems to be a horrible idea (as your benchmark showed).&nbsp; I'd be interested to see a benchmark\r\n that just predicted the most prevelent class.&nbsp; I played a bit with the cv options below and didn't see much sensitivity.&nbsp; Doing grouped=TRUE changes how the folds are aggregated for example.&nbsp; I attached the output of running this once.</p>\r\n<p>df.train &lt;- read.csv('train.csv')<br>\r\ntable(df.train$label)<br>\r\n<br>\r\nrequire(glmnet)<br>\r\n<br>\r\nTestEff &lt;- function(i.measure) {<br>\r\n&nbsp; return(cv.glmnet(<br>\r\n&nbsp;&nbsp;&nbsp; x=as.matrix(df.train[,-1])<br>\r\n&nbsp;&nbsp;&nbsp; ,y=factor(df.train$label)<br>\r\n&nbsp;&nbsp;&nbsp; ,family='multinomial'<br>\r\n&nbsp;&nbsp;&nbsp; ,standardize=TRUE<br>\r\n&nbsp;&nbsp;&nbsp; ,alpha=0.5<br>\r\n&nbsp;&nbsp;&nbsp; ,nfolds=10L<br>\r\n&nbsp;&nbsp;&nbsp; ,type.measure=i.measure<br>\r\n&nbsp;&nbsp;&nbsp; ,lambda.min.ratio=0.1<br>\r\n&nbsp;&nbsp;&nbsp; ,nlambda=25<br>\r\n&nbsp; ))}<br>\r\n<br>\r\ncv.acc &lt;- TestEff('class')<br>\r\ncv.logloss &lt;- TestEff('deviance')<br>\r\npar(mfrow=c(2,1))<br>\r\nplot(cv.acc)<br>\r\nplot(cv.logloss)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22987",
      "postDate": "04/17/2013 17:21:41",
      "content": "<p>The smoothness of the plot isn't what matters. What matters is how often the two winners change if you change the test set. You can see from your plot that the log loss is actually worse in this respect--the error bars around the model with the lowest log\r\n deviance actually completely overlap with several other points on the curve. The winner of the misclassification curve also has a lot of overlap with its neighbors but not as much.</p>\r\n<p>You can also see from these curves that the best likelihood doesn't always correspond to the best classification. Depending n your application, one or the other might matter more. We, the contest organizers, told you that for this task, what matters is classification\r\n accuracy. You don't know what the task is so you just have to take our word for it. Pretend it's picking which drug a patient should be prescribed. If you pick the wrong drug and the patient has a bad reaction, you don't get bonus points for saying you weren't\r\n very confident in your choice. This consideration ovverrides your concern about the statistical robustness of the rankings. It doesn't matter if your rankings are robust if they're ranking based on the wrong property of the model.</p>\r\n<p>There's not much point in debating it further--the contest is already launched, and it's not fair to change the evaluation after it's launched.\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22988",
      "postDate": "04/17/2013 17:28:50",
      "content": "I don't expect you to change this contest. I do realistically hope to influence the choice of loss metric in future contests however. Thanks for the discussion.",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22989",
      "postDate": "04/17/2013 17:34:34",
      "content": "<p>OK. Note that we did use AUC for the multimodal learning contest, since that involved a binary classification. I agree with you that AUC is nearly always better than accuracy when it is available.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23027",
      "postDate": "04/18/2013 01:10:16",
      "content": "<p>Using log loss or cross entropy would unfairly bias against many machine learning techniques such as adaboost and svms.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 22940,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "04/17/2013 03:54:50",
      "content": "<p>I would also like to know if we can include test set data in the unsupervised learning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22941,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "04/17/2013 04:18:26",
      "content": "<ol>\r\n<li><span style=\"line-height:1.4em\">You don't need to learn any Python for this, obviously :)</span>\r\n</li><li><span style=\"line-height:1.4em\">Accuracy on the underlying task (which we'll disclose at the end of the competition) is what we really care about to be honest, at the end of the day. For this task, accuracy and multi-class log-loss correlate pretty well,\r\n though.</span> </li><li>There are other reasons to prefer accuracy to logprob(correct_class) for evaluation, namely that the loss for a given example is actually bounded (as opposed to log-loss, which can be arbitrarily large for a particular example).\r\n</li><li><span style=\"line-height:1.4em\">AUC for multi-class problems is a little awkard (see http://users.dsic.upv.es/grupos/elp/cferri/vus-ecml03-camera-ready4.pdf for a treatment). Not even sure kaggle implements this.</span>\r\n</li><li><span style=\"line-height:1.4em\">As far as using the test data for unsupervised learning... I guess we don't cover this by the rules. I'll chat with the other organizers and get back to you.</span>\r\n</li></ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22959,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/17/2013 12:38:48",
      "content": "<p>Follow up to Dumitru's post:</p>\r\n<p>3. Following up on what Dumitru said, in my experience this effect leads the rankings given by log loss to not be very statistically robust. Because a model can lose an unbounded amount of log likelihood based on its output for a single example, the ranking\r\n of two models is often driven by outliers.</p>\r\n<p>5. Yoshua and I are both fine with allowing the test data for unsupervised learning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22965,
      "author_name": "biasvariance",
      "author_url": "",
      "post_date": "04/17/2013 13:36:58",
      "content": "Thanks for the responses. As for the robustness of log-loss, I've never ran into that issue. If a competitor is stupid enough to estimate something is 1e-8 likely to happen and it does, then they deserve to lose. In practical competitions and real work,\r\n I seldom see class predictions below 0.5%. &quot;Strange events permit themselves the luxury of occurring.&quot; I also understand that log-loss and accuracy are correlated. My actual complaint was about the efficiency of accuracy. Basically, I argue that the determination\r\n of the winner would be more consistent under log-loss given the limited amount of testing data. Said another way, the winner would be more likely to remain the winner using log-loss even if you expanding the testing sample to 100k observations. This also brings\r\n up the question of why did you choose a 50/50 split of public/private? I would think the accuracy of the private set would be more important than the accuracy of the public set.",
      "votes": null,
      "replies": []
    },
    {
      "id": 22967,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/17/2013 14:31:01",
      "content": "<p>The public and private set are the same size: http://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/data</p>\r\n<p><span style=\"line-height:1.4em\">I actually see this all the time, and it's not driven by wildly overconfident models assigning epsilon probability to black swan events. Algorithms that get 15% error on the CIFAR-10 dataset often get better log likelihood\r\n scores than algorithms that get under 10% error. The issue isn't that the more accurate model assigns a probability of 1e-8 to a single example. The issue is that the less accurate model has low confidence in a large number of mistakes while the more accurate\r\n model has medium confidence in a small number of mistakes.</span></p>\r\n<p>Both metrics have their uses for different applications. In my opinion, log likelihood is usually the better metric for a sub-component of a system, like the observation model of an HMM. Overconfidence is a more serious flaw in subcomponents because they\r\n make the larger system unaware of multiple possibilities. Accuracy is usually the better metric for the final output of a system. In most practical applications, you have to use your system to commit to a single action (do you give the patient a C-section\r\n or not?), and you don't get bonus points for doing the wrong thing with low confidence.</p>\r\n<p>As for which one generalizes better to a larger test set, I'd like to see a theorem or some empirical work. Log likelihood gives you a real number per example, instead of just one bit, but it's also prone to being driven by outliers. I suspect that if you\r\n have extremely few labels the first effect dominates but if you have a medium amount of labels like we do, the second dominates.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22975,
      "author_name": "biasvariance",
      "author_url": "",
      "post_date": "04/17/2013 15:18:34",
      "content": "I did realize the public and private were the same size; I asked why they were the same size. I would have assumed the accuracy of the private set would be more important than the public set. I also agree that there are times you might want to report/maximize\r\n accuracy. However, given the reality of limited testing, I'd want to pick a metric that stayed consistent were the testing volume expanded. I don't have evidence at hand supporting the efficiency advantage of log-loss, but I will look for some in the near\r\n future. Additionally, I guess my own training/work is focused on &quot;quantifying uncertainty&quot; so I just can't fathom not wanting to know probabilistic predictions. I also don't put credit in the outlier argument because even you agree that nobody should assign\r\n 1e-8 probability to any outcome. I don't think log-loss is perfect by any means; I'm a big proponent of rank-based evaluation metrics when the testing data is not sampled from the same distribution as the training data. Speaking of which, will you disclose\r\n if the test and train data are sampled from the same population for this contest?",
      "votes": null,
      "replies": []
    },
    {
      "id": 22976,
      "author_name": "biasvariance",
      "author_url": "",
      "post_date": "04/17/2013 15:22:10",
      "content": "A brief Google search at least turns up this stepping stone into loss-metric efficiency: http://cling.csd.uwo.ca/papers/ijcai03.pdf",
      "votes": null,
      "replies": []
    },
    {
      "id": 22979,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/17/2013 15:25:17",
      "content": "<p>Yes, that's AUC, not log loss, and it only applies to the single class case. Dumitru already explained why we don't want to use AUC for a multiclass problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22984,
      "author_name": "biasvariance",
      "author_url": "",
      "post_date": "04/17/2013 16:55:51",
      "content": "<p>I was not proposing to use AUC (or some multinomial extension) I was merely putting forward an article that discussed loss-metric efficiency in formal terms.&nbsp; On an empricial front, I present this simple bit of R code.&nbsp; From it we can see that a simple ridged\r\n linear model produces a smooth bias-variance curve when considering log-loss, but a bumpy roller-coaster ride for accuracy.&nbsp; We also see that a classic linear model seems to be a horrible idea (as your benchmark showed).&nbsp; I'd be interested to see a benchmark\r\n that just predicted the most prevelent class.&nbsp; I played a bit with the cv options below and didn't see much sensitivity.&nbsp; Doing grouped=TRUE changes how the folds are aggregated for example.&nbsp; I attached the output of running this once.</p>\r\n<p>df.train &lt;- read.csv('train.csv')<br>\r\ntable(df.train$label)<br>\r\n<br>\r\nrequire(glmnet)<br>\r\n<br>\r\nTestEff &lt;- function(i.measure) {<br>\r\n&nbsp; return(cv.glmnet(<br>\r\n&nbsp;&nbsp;&nbsp; x=as.matrix(df.train[,-1])<br>\r\n&nbsp;&nbsp;&nbsp; ,y=factor(df.train$label)<br>\r\n&nbsp;&nbsp;&nbsp; ,family='multinomial'<br>\r\n&nbsp;&nbsp;&nbsp; ,standardize=TRUE<br>\r\n&nbsp;&nbsp;&nbsp; ,alpha=0.5<br>\r\n&nbsp;&nbsp;&nbsp; ,nfolds=10L<br>\r\n&nbsp;&nbsp;&nbsp; ,type.measure=i.measure<br>\r\n&nbsp;&nbsp;&nbsp; ,lambda.min.ratio=0.1<br>\r\n&nbsp;&nbsp;&nbsp; ,nlambda=25<br>\r\n&nbsp; ))}<br>\r\n<br>\r\ncv.acc &lt;- TestEff('class')<br>\r\ncv.logloss &lt;- TestEff('deviance')<br>\r\npar(mfrow=c(2,1))<br>\r\nplot(cv.acc)<br>\r\nplot(cv.logloss)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22987,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/17/2013 17:21:41",
      "content": "<p>The smoothness of the plot isn't what matters. What matters is how often the two winners change if you change the test set. You can see from your plot that the log loss is actually worse in this respect--the error bars around the model with the lowest log\r\n deviance actually completely overlap with several other points on the curve. The winner of the misclassification curve also has a lot of overlap with its neighbors but not as much.</p>\r\n<p>You can also see from these curves that the best likelihood doesn't always correspond to the best classification. Depending n your application, one or the other might matter more. We, the contest organizers, told you that for this task, what matters is classification\r\n accuracy. You don't know what the task is so you just have to take our word for it. Pretend it's picking which drug a patient should be prescribed. If you pick the wrong drug and the patient has a bad reaction, you don't get bonus points for saying you weren't\r\n very confident in your choice. This consideration ovverrides your concern about the statistical robustness of the rankings. It doesn't matter if your rankings are robust if they're ranking based on the wrong property of the model.</p>\r\n<p>There's not much point in debating it further--the contest is already launched, and it's not fair to change the evaluation after it's launched.\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22988,
      "author_name": "biasvariance",
      "author_url": "",
      "post_date": "04/17/2013 17:28:50",
      "content": "I don't expect you to change this contest. I do realistically hope to influence the choice of loss metric in future contests however. Thanks for the discussion.",
      "votes": null,
      "replies": []
    },
    {
      "id": 22989,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/17/2013 17:34:34",
      "content": "<p>OK. Note that we did use AUC for the multimodal learning contest, since that involved a binary classification. I agree with you that AUC is nearly always better than accuracy when it is available.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23027,
      "author_name": "neocortex",
      "author_url": "",
      "post_date": "04/18/2013 01:10:16",
      "content": "<p>Using log loss or cross entropy would unfairly bias against many machine learning techniques such as adaboost and svms.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "22937": "",
    "22940": "",
    "22941": "",
    "22959": "",
    "22965": "",
    "22967": "",
    "22975": "",
    "22976": "",
    "22979": "",
    "22984": "",
    "22987": "",
    "22988": "",
    "22989": "",
    "23027": ""
  },
  "source": "meta"
}