{
  "id": 2748,
  "title": "Basis of Multi-Class Log Loss in Literature or Journals?",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2748",
  "author_name": "",
  "post_date": "2012-09-25T03:28:56.330Z",
  "votes": 1,
  "comment_count": 2,
  "views": 3356,
  "content": "<p>Good evening.</p>\r\n<p>I have noticed the evaluation technique on this competition and was curious if this technique is used to evaluate multiclass probabalistic models in academic works?&nbsp; If so, is this a new technique or a technique that has been around for a while.</p>\r\n<p>Thank you for any guidance.</p>\r\n<p>Mike</p>",
  "messages": [
    {
      "id": "14767",
      "postDate": "09/25/2012 03:28:56",
      "content": "<p>Good evening.</p>\r\n<p>I have noticed the evaluation technique on this competition and was curious if this technique is used to evaluate multiclass probabalistic models in academic works?&nbsp; If so, is this a new technique or a technique that has been around for a while.</p>\r\n<p>Thank you for any guidance.</p>\r\n<p>Mike</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14770",
      "postDate": "09/25/2012 04:33:28",
      "content": "<p>It makes sense for this contest because the labels are so noisy. A theoretically perfect classifier which doesn't specify probabilities, but instead just picks the most likely label, would do terrible for this task, because a great many of the posts which\r\n<em>should</em> be e.g. off-topic are actually in the open state. Thus an entry with a worse score may actually be a better classifier, in a pure precision/recall sense, on a hand-labeled dataset.</p>\r\n<p>So instead, the name of the game is to not be overconfident about predictions and hedge your bets as well as possible by assigning proper probabilities. You want to maximize the joint likelihood of the labels given the data, which would be a product of all\r\n the probabilities you've assigned to the &quot;true&quot; labels, but that would be a very, very small number. So instead, you sum up the log-probabilities, which is equivalent but computable.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "15115",
      "postDate": "10/03/2012 04:57:46",
      "content": "<p>LogLoss will mean that whether you classify an open question wrongly or one of the sparse classes wrongly - they will be penalized in the same manner.\r\n<br>\r\nSome other metric might have been apt for this competition</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 14770,
      "author_name": "andysloane",
      "author_url": "",
      "post_date": "09/25/2012 04:33:28",
      "content": "<p>It makes sense for this contest because the labels are so noisy. A theoretically perfect classifier which doesn't specify probabilities, but instead just picks the most likely label, would do terrible for this task, because a great many of the posts which\r\n<em>should</em> be e.g. off-topic are actually in the open state. Thus an entry with a worse score may actually be a better classifier, in a pure precision/recall sense, on a hand-labeled dataset.</p>\r\n<p>So instead, the name of the game is to not be overconfident about predictions and hedge your bets as well as possible by assigning proper probabilities. You want to maximize the joint likelihood of the labels given the data, which would be a product of all\r\n the probabilities you've assigned to the &quot;true&quot; labels, but that would be a very, very small number. So instead, you sum up the log-probabilities, which is equivalent but computable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 15115,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "10/03/2012 04:57:46",
      "content": "<p>LogLoss will mean that whether you classify an open question wrongly or one of the sparse classes wrongly - they will be penalized in the same manner.\r\n<br>\r\nSome other metric might have been apt for this competition</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "14767": "",
    "14770": "",
    "15115": ""
  },
  "source": "meta"
}