{
  "id": 10191,
  "title": "How is ROC curve generated?",
  "url": "/competitions/seizure-prediction/discussion/10191",
  "author_name": "",
  "post_date": "2014-09-02T18:59:24.347Z",
  "votes": null,
  "comment_count": 3,
  "views": 1511,
  "content": "<p>The evaluation metric of this challenge is&nbsp;area under the ROC curve, but how is this curve generated? Specifically, if our submission contains only boolean predictions, i.e. 0 or 1, instead of real-value probabilities, changing the threshold value will not yield different&nbsp;classification results and there will be only one single point in the plot instead of a curve.</p>\n<p>According to <a href=\"http://en.wikipedia.org/wiki/Receiver_operating_characteristic\">Wikipedia</a>, there are multiple techniques to produce the ROC curve, such as trapezoidal approximations and ROC AUCH. I wonder which one Kaggle uses.</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "52933",
      "postDate": "09/02/2014 18:59:24",
      "content": "<p>The evaluation metric of this challenge is&nbsp;area under the ROC curve, but how is this curve generated? Specifically, if our submission contains only boolean predictions, i.e. 0 or 1, instead of real-value probabilities, changing the threshold value will not yield different&nbsp;classification results and there will be only one single point in the plot instead of a curve.</p>\n<p>According to <a href=\"http://en.wikipedia.org/wiki/Receiver_operating_characteristic\">Wikipedia</a>, there are multiple techniques to produce the ROC curve, such as trapezoidal approximations and ROC AUCH. I wonder which one Kaggle uses.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52955",
      "postDate": "09/02/2014 21:49:52",
      "content": "<p>If you use a 0/1 output, you obtain a pair of sensitivity/specifity values, say SE* and SP*.</p>\n<p>In that case, your ROC curve consists of three points:</p>\n<p>SE = 0, SP = 1</p>\n<p>SE= SE*, SP = SP*</p>\n<p>SE=1, SP = 0.</p>\n<p>The area under this trapezoid is the AUC.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53292",
      "postDate": "09/07/2014 16:53:47",
      "content": "<p>Kaggle specifies probabilistic outputs, not binary. This means they can generate the entire ROC curve by varying the threshold probability that defines a hit. That being said, the best 3-point score will be found by using the threshold with the fewest classification errors. If the cost of a false alarm is not the same as the cost of a missed detection, then this is not the case.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53301",
      "postDate": "09/07/2014 19:25:07",
      "content": "<p>[quote=kdoniger;53292]</p>\n<p>Kaggle specifies probabilistic outputs, not binary. This means they can generate the entire ROC curve by varying the threshold probability that defines a hit. That being said, the best 3-point score will be found by using the threshold with the fewest classification errors. If the cost of a false alarm is not the same as the cost of a missed detection, then this is not the case.</p>\n<p>[/quote]</p>\n\n<p>I think binary outputs are also accepted, since they can be seen as probabilities of 100% and 0%.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 52955,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "09/02/2014 21:49:52",
      "content": "<p>If you use a 0/1 output, you obtain a pair of sensitivity/specifity values, say SE* and SP*.</p>\n<p>In that case, your ROC curve consists of three points:</p>\n<p>SE = 0, SP = 1</p>\n<p>SE= SE*, SP = SP*</p>\n<p>SE=1, SP = 0.</p>\n<p>The area under this trapezoid is the AUC.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53292,
      "author_name": "kdoniger",
      "author_url": "",
      "post_date": "09/07/2014 16:53:47",
      "content": "<p>Kaggle specifies probabilistic outputs, not binary. This means they can generate the entire ROC curve by varying the threshold probability that defines a hit. That being said, the best 3-point score will be found by using the threshold with the fewest classification errors. If the cost of a false alarm is not the same as the cost of a missed detection, then this is not the case.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53301,
      "author_name": "meditativeape",
      "author_url": "",
      "post_date": "09/07/2014 19:25:07",
      "content": "<p>[quote=kdoniger;53292]</p>\n<p>Kaggle specifies probabilistic outputs, not binary. This means they can generate the entire ROC curve by varying the threshold probability that defines a hit. That being said, the best 3-point score will be found by using the threshold with the fewest classification errors. If the cost of a false alarm is not the same as the cost of a missed detection, then this is not the case.</p>\n<p>[/quote]</p>\n\n<p>I think binary outputs are also accepted, since they can be seen as probabilities of 100% and 0%.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "52933": "",
    "52955": "",
    "53292": "",
    "53301": ""
  },
  "source": "meta"
}