{
  "id": 20996,
  "title": "Human prediction accuracy benchmark",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20996",
  "author_name": "",
  "post_date": "2016-05-16T15:28:42.130Z",
  "votes": 1,
  "comment_count": 5,
  "views": 991,
  "content": "<p>Did any of you estimate the human accuracy on this dataset? Would be interesting to know if some model beats it</p>",
  "messages": [
    {
      "id": "120230",
      "postDate": "05/16/2016 15:28:42",
      "content": "<p>Did any of you estimate the human accuracy on this dataset? Would be interesting to know if some model beats it</p>",
      "rawMarkdown": "Did any of you estimate the human accuracy on this dataset? Would be interesting to know if some model beats it",
      "votes": null
    },
    {
      "id": "120373",
      "postDate": "05/17/2016 20:48:21",
      "content": "<p>I tested myself a while on the training set. Got 94% right - which corresponds to a log loss of 0.36 (for 10 classes).</p>\n\n<p>... that compares quite badly with the 0.16 of the best LB entry right now :(.</p>",
      "rawMarkdown": "I tested myself a while on the training set. Got 94% right - which corresponds to a log loss of 0.36 (for 10 classes).\r\n\r\n... that compares quite badly with the 0.16 of the best LB entry right now :(.",
      "votes": null
    },
    {
      "id": "120767",
      "postDate": "05/20/2016 13:06:18",
      "content": "<p>[quote=DivideBy0;120373]\nI tested myself a while on the training set. Got 94% right - which corresponds to a log loss of 0.36 (for 10 classes).\n[/quote]</p>\n\n<p>How did you come up with that number? </p>\n\n<p>For simplicity's sake, let's assume that you give a probability of 0.94 for the class that you believe is the right one and 0.06/9 for the rest. This would lead to expected logloss of: -(0.06*log(0.06/9) + 0.94*log(0.94)) ~ 0.156</p>\n\n<p>That would give the upper limit. Then, it certainly would be relatively easy to make improvements from there by not using simple constant predictions etc. So, I would guess the &quot;real&quot; value would be something like 0.1-0.13 (depending how much time you would spend / image)</p>",
      "rawMarkdown": "[quote=DivideBy0;120373]\r\nI tested myself a while on the training set. Got 94% right - which corresponds to a log loss of 0.36 (for 10 classes).\r\n[/quote]\r\n\r\nHow did you come up with that number? \r\n\r\nFor simplicity's sake, let's assume that you give a probability of 0.94 for the class that you believe is the right one and 0.06/9 for the rest. This would lead to expected logloss of: -(0.06*log(0.06/9) + 0.94*log(0.94)) ~ 0.156\r\n\r\nThat would give the upper limit. Then, it certainly would be relatively easy to make improvements from there by not using simple constant predictions etc. So, I would guess the \"real\" value would be something like 0.1-0.13 (depending how much time you would spend / image)",
      "votes": null
    },
    {
      "id": "120770",
      "postDate": "05/20/2016 13:22:41",
      "content": "<p>Hey Herra,</p>\n\n<p>I used the natural log instead of log10:</p>\n\n<ul>\n<li>Log10 based:  -(0.06*log10(0.06/9) + 0.94*log10(0.94)) ~ 0.1558</li>\n<li>Log(e) based: -(0.06*log(0.06/9) + 0.94*log(0.94)) ~ 0.3588</li>\n</ul>\n\n<p>Look at <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation</a>. It says &quot;... \\(log\\) is the natural logarithm ...&quot;.</p>",
      "rawMarkdown": "Hey Herra,\r\n\r\nI used the natural log instead of log10:\r\n\r\n- Log10 based:  -(0.06*log10(0.06/9) + 0.94*log10(0.94)) ~ 0.1558\r\n- Log(e) based: -(0.06*log(0.06/9) + 0.94*log(0.94)) ~ 0.3588\r\n\r\nLook at https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation. It says \"... \\\\(log\\\\) is the natural logarithm ...\".",
      "votes": null
    },
    {
      "id": "120775",
      "postDate": "05/20/2016 13:51:34",
      "content": "<p>Ah, of course. I have to say, even surprisingly good LB scores then.</p>",
      "rawMarkdown": "Ah, of course. I have to say, even surprisingly good LB scores then.",
      "votes": null
    },
    {
      "id": "120790",
      "postDate": "05/20/2016 15:05:45",
      "content": "<p>waouh, current top lb at mlogloss 0.155 brings out an impressive 97,77% accuracy ... pretty accurate even for a human labeling...</p>",
      "rawMarkdown": "waouh, current top lb at mlogloss 0.155 brings out an impressive 97,77% accuracy ... pretty accurate even for a human labeling...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 120373,
      "author_name": "divideby0",
      "author_url": "",
      "post_date": "05/17/2016 20:48:21",
      "content": "<p>I tested myself a while on the training set. Got 94% right - which corresponds to a log loss of 0.36 (for 10 classes).</p>\n\n<p>... that compares quite badly with the 0.16 of the best LB entry right now :(.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120767,
      "author_name": "herrahuu",
      "author_url": "",
      "post_date": "05/20/2016 13:06:18",
      "content": "<p>[quote=DivideBy0;120373]\nI tested myself a while on the training set. Got 94% right - which corresponds to a log loss of 0.36 (for 10 classes).\n[/quote]</p>\n\n<p>How did you come up with that number? </p>\n\n<p>For simplicity's sake, let's assume that you give a probability of 0.94 for the class that you believe is the right one and 0.06/9 for the rest. This would lead to expected logloss of: -(0.06*log(0.06/9) + 0.94*log(0.94)) ~ 0.156</p>\n\n<p>That would give the upper limit. Then, it certainly would be relatively easy to make improvements from there by not using simple constant predictions etc. So, I would guess the &quot;real&quot; value would be something like 0.1-0.13 (depending how much time you would spend / image)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120770,
      "author_name": "divideby0",
      "author_url": "",
      "post_date": "05/20/2016 13:22:41",
      "content": "<p>Hey Herra,</p>\n\n<p>I used the natural log instead of log10:</p>\n\n<ul>\n<li>Log10 based:  -(0.06*log10(0.06/9) + 0.94*log10(0.94)) ~ 0.1558</li>\n<li>Log(e) based: -(0.06*log(0.06/9) + 0.94*log(0.94)) ~ 0.3588</li>\n</ul>\n\n<p>Look at <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation</a>. It says &quot;... \\(log\\) is the natural logarithm ...&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120775,
      "author_name": "herrahuu",
      "author_url": "",
      "post_date": "05/20/2016 13:51:34",
      "content": "<p>Ah, of course. I have to say, even surprisingly good LB scores then.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120790,
      "author_name": "ngaude",
      "author_url": "",
      "post_date": "05/20/2016 15:05:45",
      "content": "<p>waouh, current top lb at mlogloss 0.155 brings out an impressive 97,77% accuracy ... pretty accurate even for a human labeling...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "120230": "Did any of you estimate the human accuracy on this dataset? Would be interesting to know if some model beats it",
    "120373": "I tested myself a while on the training set. Got 94% right - which corresponds to a log loss of 0.36 (for 10 classes).\r\n\r\n... that compares quite badly with the 0.16 of the best LB entry right now :(.",
    "120767": "[quote=DivideBy0;120373]\r\nI tested myself a while on the training set. Got 94% right - which corresponds to a log loss of 0.36 (for 10 classes).\r\n[/quote]\r\n\r\nHow did you come up with that number? \r\n\r\nFor simplicity's sake, let's assume that you give a probability of 0.94 for the class that you believe is the right one and 0.06/9 for the rest. This would lead to expected logloss of: -(0.06*log(0.06/9) + 0.94*log(0.94)) ~ 0.156\r\n\r\nThat would give the upper limit. Then, it certainly would be relatively easy to make improvements from there by not using simple constant predictions etc. So, I would guess the \"real\" value would be something like 0.1-0.13 (depending how much time you would spend / image)",
    "120770": "Hey Herra,\r\n\r\nI used the natural log instead of log10:\r\n\r\n- Log10 based:  -(0.06*log10(0.06/9) + 0.94*log10(0.94)) ~ 0.1558\r\n- Log(e) based: -(0.06*log(0.06/9) + 0.94*log(0.94)) ~ 0.3588\r\n\r\nLook at https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation. It says \"... \\\\(log\\\\) is the natural logarithm ...\".",
    "120775": "Ah, of course. I have to say, even surprisingly good LB scores then.",
    "120790": "waouh, current top lb at mlogloss 0.155 brings out an impressive 97,77% accuracy ... pretty accurate even for a human labeling..."
  },
  "source": "meta"
}