{
  "id": 12940,
  "title": "Real world log loss score meaning",
  "url": "/competitions/malware-classification/discussion/12940",
  "author_name": "",
  "post_date": "2015-03-20T15:21:47.390Z",
  "votes": null,
  "comment_count": 4,
  "views": 2213,
  "content": "<p>I was wondering, with this metric there are multiple very different way to get the same score, e.g, few very bad predictions vs lots almost correct predictions, both, given the number of &quot;few&quot; and &quot;lots&quot; could lead to exactly same score, however in real world they behave very different.</p>\n<p>Given this premise, my question is then...what really... in real world does this metric tell us, is really a 0.1 score twice as good as a 0.2 ?</p>",
  "messages": [
    {
      "id": "67418",
      "postDate": "03/20/2015 15:21:47",
      "content": "<p>I was wondering, with this metric there are multiple very different way to get the same score, e.g, few very bad predictions vs lots almost correct predictions, both, given the number of &quot;few&quot; and &quot;lots&quot; could lead to exactly same score, however in real world they behave very different.</p>\n<p>Given this premise, my question is then...what really... in real world does this metric tell us, is really a 0.1 score twice as good as a 0.2 ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67430",
      "postDate": "03/20/2015 17:24:16",
      "content": "<p>It just makes the competition more interesting. There is always a hope to improve. But for real world, my experience is usually several metrics are used instead of just one. I don't know why kaggle chooses log loss, cause I think a weighted score on multiple metrics seems better.</p>\n<p>And no, logloss is not linear so 0.1 is not twice as good as 0.2.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67477",
      "postDate": "03/21/2015 00:07:56",
      "content": "<p>The generated error-per-instance is not linear, but the &quot;accumulative error score&quot;*, it is, isn't it?</p>\n<p>* e.g. you get twice the error score if you have the same error twice.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67489",
      "postDate": "03/21/2015 03:25:26",
      "content": "<p>[quote=NxGTR;67477]</p>\n<p>The generated error-per-instance is not linear, but the &quot;accumulative error score&quot;*, it is, isn't it?</p>\n<p>* e.g. you get twice the error score if you have the same error twice.</p>\n<p>[/quote]</p>\n<p>Well, but you have one shot for each instance right? You either predict it as 0.1 or 0.05. And&nbsp;all constant 0.05 is not twice as good/bad as 0.1.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67528",
      "postDate": "03/21/2015 15:28:37",
      "content": "<p>Got it.<br>Along the competition I have noticed groups of people stuck for a while in 0.011x now in 0.008x, but the Top 2 in 0.004 seems like unreachable for me, so I was just wondering if there was a way to estimate (lets discard potential overfitting by now) how much better the model could be. Now I know that &quot;twice as good&quot; it is not.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 67430,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "03/20/2015 17:24:16",
      "content": "<p>It just makes the competition more interesting. There is always a hope to improve. But for real world, my experience is usually several metrics are used instead of just one. I don't know why kaggle chooses log loss, cause I think a weighted score on multiple metrics seems better.</p>\n<p>And no, logloss is not linear so 0.1 is not twice as good as 0.2.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67477,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "03/21/2015 00:07:56",
      "content": "<p>The generated error-per-instance is not linear, but the &quot;accumulative error score&quot;*, it is, isn't it?</p>\n<p>* e.g. you get twice the error score if you have the same error twice.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67489,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "03/21/2015 03:25:26",
      "content": "<p>[quote=NxGTR;67477]</p>\n<p>The generated error-per-instance is not linear, but the &quot;accumulative error score&quot;*, it is, isn't it?</p>\n<p>* e.g. you get twice the error score if you have the same error twice.</p>\n<p>[/quote]</p>\n<p>Well, but you have one shot for each instance right? You either predict it as 0.1 or 0.05. And&nbsp;all constant 0.05 is not twice as good/bad as 0.1.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67528,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "03/21/2015 15:28:37",
      "content": "<p>Got it.<br>Along the competition I have noticed groups of people stuck for a while in 0.011x now in 0.008x, but the Top 2 in 0.004 seems like unreachable for me, so I was just wondering if there was a way to estimate (lets discard potential overfitting by now) how much better the model could be. Now I know that &quot;twice as good&quot; it is not.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "67418": "",
    "67430": "",
    "67477": "",
    "67489": "",
    "67528": ""
  },
  "source": "meta"
}