{
  "id": 2499,
  "title": "Multiclass Logarithmic Loss",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2499",
  "author_name": "",
  "post_date": "2012-08-29T23:27:37.730Z",
  "votes": 5,
  "comment_count": 3,
  "views": 4603,
  "content": "<p>I've read the <a href=\"https://www.kaggle.com/wiki/MultiClassLogLoss\" target=\"_blank\">\r\nofficial overview</a>&nbsp;and have a reasonable understanding of how this value is derived. &nbsp;However it's still somewhat unclear to me what the significance of this value is/what it actually measures/represents, in plain terms. &nbsp;For example, If one algorithm scores\r\n a .2 and another scores a .5, what can be said about them beyond the simple fact that &quot;algorithm 1 performs better than algorithm 2&quot;? &nbsp;From those numbers, is it possible to quantify exactly how much better algorithm 1 is, or how much algorithm 2 would need\r\n to be improved to match algorithm 1 (like in terms of number/frequency of incorrect predictions)?</p>\r\n<p>It seems clear that a value of 0 indicates a perfect run, but what does a value of 1 correlate to, if anything? &nbsp;Random guessing? &nbsp;And does a value above 1 mean anything in particular? &nbsp;Like maybe that the approach does even worse than random guessing?</p>\r\n<p>Pretend that you're talking to someone who gets bored and wanders off when people start invoking complex mathmatics to explain relatively simple concepts. &nbsp;Thanks.</p>",
  "messages": [
    {
      "id": "13633",
      "postDate": "08/29/2012 23:27:37",
      "content": "<p>I've read the <a href=\"https://www.kaggle.com/wiki/MultiClassLogLoss\" target=\"_blank\">\r\nofficial overview</a>&nbsp;and have a reasonable understanding of how this value is derived. &nbsp;However it's still somewhat unclear to me what the significance of this value is/what it actually measures/represents, in plain terms. &nbsp;For example, If one algorithm scores\r\n a .2 and another scores a .5, what can be said about them beyond the simple fact that &quot;algorithm 1 performs better than algorithm 2&quot;? &nbsp;From those numbers, is it possible to quantify exactly how much better algorithm 1 is, or how much algorithm 2 would need\r\n to be improved to match algorithm 1 (like in terms of number/frequency of incorrect predictions)?</p>\r\n<p>It seems clear that a value of 0 indicates a perfect run, but what does a value of 1 correlate to, if anything? &nbsp;Random guessing? &nbsp;And does a value above 1 mean anything in particular? &nbsp;Like maybe that the approach does even worse than random guessing?</p>\r\n<p>Pretend that you're talking to someone who gets bored and wanders off when people start invoking complex mathmatics to explain relatively simple concepts. &nbsp;Thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13635",
      "postDate": "08/30/2012 05:56:47",
      "content": "<p>There are several ways of interpreting the log loss metric - the most usual is in information theoretic terms, as the cross-entropy of two distributions. However, this is a little bit abstract, and if you're not already familiar with information theory,\r\n it probably won't be very illuminating.</p>\r\n<p>There's another way of looking at it though, which might be more intuitive. Imagine that you're a gambler, betting on whether each Stack Overflow question will be closed or not. If you're wrong, you lose your stake, and if you're right, you get \\\\(X\\\\) times\r\n your stake back (let's assume \\\\(X\\\\) is fairly large, so that it's always worth you betting). In a one-off situation, your best bet is to put all your money on the most likely outcome, but if you repeat this strategy, you're likely to end up broke pretty\r\n quickly. You want to hedge your bets, so that even if an unlikely outcome occurs, you've still got some money left to keep betting with. Rather than maximising your one-off profits, you're more interested in the long-term growth rate of your balance. In particular,\r\n over the long run you care about the log of your return (e.g. halving your money is as bad as doubling your money is good).</p>\r\n<p>If the probability of outcome \\\\(i\\\\) is \\\\(y_i\\\\), and you bet a proportion \\\\(p_i\\\\) of your money on that outcome, then after each round your expected log return is \\\\(\\sum_i y_i \\log (X p_i) = \\log X &#43; \\sum_i y_i \\log p_i\\\\). The first term (\\\\(\\log\r\n X\\\\)) is the return rate of a perfect gambler, who always puts his money on the right outcome. The second term (which you'll probably recognise as the log-loss metric) is how much worse off you are if you bet according to the distribution \\\\(y_i\\\\).</p>\r\n<p>This may not be the application that Stack Overflow have in mind. I doubt they run a pool internally on which questions will be closed. However, it does show a simple way in which this error metric arises. In reality though, this metric was probably chosen\r\n because it's a well-understood commonly used metric for this kind of problem, which penalises excessive claims of certainty.</p>\r\n<p>To answer your other questions, a value of 1 is fairly meaningless here. A better benchmark might be the value we get if we predict uniform probabilities for all classes. In this case, with 5 classes, this is equal to \\\\(\\log 5 \\simeq 1.60944\\ldots\\\\). Anyone\r\n scoring less than this is probably doing something wrong.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13646",
      "postDate": "08/30/2012 15:37:39",
      "content": "<blockquote>\r\n<p><span>This may not be the application that Stack Overflow have in mind. I doubt they run a pool internally on which questions will be closed.</span></p>\r\n</blockquote>\r\n<p>&nbsp;</p>\r\n<p>Martin! You just revealed the dark underbelly of SO!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "758040",
      "postDate": "02/27/2020 11:46:59",
      "content": "<p>Log loss, aka logistic loss or cross-entropy loss.</p>\n\n<p>This is the loss function used in (multinomial) logistic regression and extensions of it such as neural networks, defined as the negative log-likelihood of the true labels given a probabilistic classifier’s predictions. The log loss is only defined for two or more labels. For a single sample with true label yt in {0,1} and estimated probability yp that yt = 1, the log loss is</p>\n\n<p>-log P(yt|yp) = -(yt log(yp) + (1 - yt) log(1 - yp))</p>\n\n<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html\">https://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html</a></p>",
      "rawMarkdown": "Log loss, aka logistic loss or cross-entropy loss.\n\nThis is the loss function used in (multinomial) logistic regression and extensions of it such as neural networks, defined as the negative log-likelihood of the true labels given a probabilistic classifier’s predictions. The log loss is only defined for two or more labels. For a single sample with true label yt in {0,1} and estimated probability yp that yt = 1, the log loss is\n\n-log P(yt|yp) = -(yt log(yp) + (1 - yt) log(1 - yp))\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13635,
      "author_name": "martinoleary",
      "author_url": "",
      "post_date": "08/30/2012 05:56:47",
      "content": "<p>There are several ways of interpreting the log loss metric - the most usual is in information theoretic terms, as the cross-entropy of two distributions. However, this is a little bit abstract, and if you're not already familiar with information theory,\r\n it probably won't be very illuminating.</p>\r\n<p>There's another way of looking at it though, which might be more intuitive. Imagine that you're a gambler, betting on whether each Stack Overflow question will be closed or not. If you're wrong, you lose your stake, and if you're right, you get \\\\(X\\\\) times\r\n your stake back (let's assume \\\\(X\\\\) is fairly large, so that it's always worth you betting). In a one-off situation, your best bet is to put all your money on the most likely outcome, but if you repeat this strategy, you're likely to end up broke pretty\r\n quickly. You want to hedge your bets, so that even if an unlikely outcome occurs, you've still got some money left to keep betting with. Rather than maximising your one-off profits, you're more interested in the long-term growth rate of your balance. In particular,\r\n over the long run you care about the log of your return (e.g. halving your money is as bad as doubling your money is good).</p>\r\n<p>If the probability of outcome \\\\(i\\\\) is \\\\(y_i\\\\), and you bet a proportion \\\\(p_i\\\\) of your money on that outcome, then after each round your expected log return is \\\\(\\sum_i y_i \\log (X p_i) = \\log X &#43; \\sum_i y_i \\log p_i\\\\). The first term (\\\\(\\log\r\n X\\\\)) is the return rate of a perfect gambler, who always puts his money on the right outcome. The second term (which you'll probably recognise as the log-loss metric) is how much worse off you are if you bet according to the distribution \\\\(y_i\\\\).</p>\r\n<p>This may not be the application that Stack Overflow have in mind. I doubt they run a pool internally on which questions will be closed. However, it does show a simple way in which this error metric arises. In reality though, this metric was probably chosen\r\n because it's a well-understood commonly used metric for this kind of problem, which penalises excessive claims of certainty.</p>\r\n<p>To answer your other questions, a value of 1 is fairly meaningless here. A better benchmark might be the value we get if we predict uniform probabilities for all classes. In this case, with 5 classes, this is equal to \\\\(\\log 5 \\simeq 1.60944\\ldots\\\\). Anyone\r\n scoring less than this is probably doing something wrong.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13646,
      "author_name": "thellf",
      "author_url": "",
      "post_date": "08/30/2012 15:37:39",
      "content": "<blockquote>\r\n<p><span>This may not be the application that Stack Overflow have in mind. I doubt they run a pool internally on which questions will be closed.</span></p>\r\n</blockquote>\r\n<p>&nbsp;</p>\r\n<p>Martin! You just revealed the dark underbelly of SO!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 758040,
      "author_name": "srinivasav22",
      "author_url": "",
      "post_date": "02/27/2020 11:46:59",
      "content": "<p>Log loss, aka logistic loss or cross-entropy loss.</p>\n\n<p>This is the loss function used in (multinomial) logistic regression and extensions of it such as neural networks, defined as the negative log-likelihood of the true labels given a probabilistic classifier’s predictions. The log loss is only defined for two or more labels. For a single sample with true label yt in {0,1} and estimated probability yp that yt = 1, the log loss is</p>\n\n<p>-log P(yt|yp) = -(yt log(yp) + (1 - yt) log(1 - yp))</p>\n\n<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html\">https://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13633": "",
    "13635": "",
    "13646": "",
    "758040": "Log loss, aka logistic loss or cross-entropy loss.\n\nThis is the loss function used in (multinomial) logistic regression and extensions of it such as neural networks, defined as the negative log-likelihood of the true labels given a probabilistic classifier’s predictions. The log loss is only defined for two or more labels. For a single sample with true label yt in {0,1} and estimated probability yp that yt = 1, the log loss is\n\n-log P(yt|yp) = -(yt log(yp) + (1 - yt) log(1 - yp))\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.metrics.log_loss.html"
  },
  "source": "meta"
}