{
  "id": 10444,
  "title": "Binary classificator ",
  "url": "/competitions/seizure-prediction/discussion/10444",
  "author_name": "",
  "post_date": "2014-09-24T20:53:18.637Z",
  "votes": null,
  "comment_count": 8,
  "views": 2121,
  "content": "<p>Dear collegues, I'd like to know about the output expected by the administrators of this competition.&nbsp;&nbsp; I thought they wanted a classification to the samples in &quot;precital&quot; or &quot;interictal&quot; state. Am I right?</p>\n\n<p>If so, I think &quot;1&quot; means that current sample is a preictal one, and &quot;0&quot; means a interictal sample.</p>\n\n<p>Is that true?</p>\n\n<p>Kaguiar</p>",
  "messages": [
    {
      "id": "55185",
      "postDate": "09/24/2014 20:53:18",
      "content": "<p>Dear collegues, I'd like to know about the output expected by the administrators of this competition.&nbsp;&nbsp; I thought they wanted a classification to the samples in &quot;precital&quot; or &quot;interictal&quot; state. Am I right?</p>\n\n<p>If so, I think &quot;1&quot; means that current sample is a preictal one, and &quot;0&quot; means a interictal sample.</p>\n\n<p>Is that true?</p>\n\n<p>Kaguiar</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55186",
      "postDate": "09/24/2014 20:54:10",
      "content": "<p>* and &quot;0&quot; means AN interictal sample</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55189",
      "postDate": "09/24/2014 22:50:31",
      "content": "<p>I think so but instead of 0 and 1, an output of probability (or a rank between 0 and 1) of &quot;preictal&quot; event is much more practical to improve roc score.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55190",
      "postDate": "09/24/2014 22:57:23",
      "content": "<p>Yes, indeed.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55684",
      "postDate": "10/06/2014 14:15:02",
      "content": "<p>Why is outputting the probability of a preictal event (as opposed to a discrete class label 0 or 1) more practical to improve ROC score?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55693",
      "postDate": "10/06/2014 17:46:04",
      "content": "<p>Because the ROC iterates over several different thresholds for the classification, its possible that your results will have a better accuracy at some threshold versus others.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55726",
      "postDate": "10/07/2014 11:10:34",
      "content": "<p>[quote=franklyn;55693]</p>\n<p>Because the ROC iterates over several different thresholds for the classification...</p>\n<p>[/quote]</p>\n<p>Yes, exactly.</p>\n<p>To illustrate:</p>\n<p><code>from sklearn.metrics import roc_auc_score</code></p>\n<p><code></code><code>y_predicted_discrete = [0.0, 0.0, 0.0, 1.0, 1.0]</code></p>\n<p><code>y_predicted_proba = [0.02, 0.01, 0.05, 0.99, 0.99]</code></p>\n<p><code>y_ground_truth = [0, 0, 1, 1, 1]</code></p>\n<p><code>print roc_auc_score(y_ground_truth, y_predicted_discrete) # 0.83</code></p>\n<p><code>print roc_auc_score(y_ground_truth, y_predicted_proba) # 1.0</code></p>\n<p>Even though the 3rd prediction probability was just&nbsp;0.05 and the truth was 1, it still got credit for ranking it higher, causing a favorable perfect split.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55737",
      "postDate": "10/07/2014 14:31:29",
      "content": "<p>Gotcha. Thanks. So one&nbsp;can one still receive a perfect score by outputting probabilities as opposed to&nbsp;discrete labels?&nbsp;In other words, even if the third predicted discrete value was correct in your example, both sets&nbsp;would have received the same score under ROC?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55740",
      "postDate": "10/07/2014 16:41:38",
      "content": "<p>[quote=Adam Soffer;55737]</p>\n<p>Gotcha. Thanks. So one&nbsp;can one still receive a perfect score by outputting probabilities as opposed to&nbsp;discrete labels?&nbsp;In other words, even if the third predicted discrete value was correct in your example, both sets&nbsp;would have received the same score under ROC?</p>\n<p>[/quote]</p>\n<p>I take it you meant:&nbsp;one can still receive a perfect score by outputting discrete labels&nbsp;as opposed to probabilities? That answer is yes. It is possible if you yourself have found the perfect split/threshold between the positive and the negative class. AUC finds this for you by considering every threshold split and only picking the best your model has got.</p>\n<p>Since this contest is not judged on raw accuracy, it would not be taking advantage of the evaluation metric if you still submit discrete labels: Those high on the leaderboard are submitting probabilities or ranks.</p>\n<p>Other <a href=\"http://www.kaggle.com/wiki/Metrics\">evaluation metrics</a>&nbsp;call for other &quot;tactics&quot;.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 55186,
      "author_name": "kaguiar",
      "author_url": "",
      "post_date": "09/24/2014 20:54:10",
      "content": "<p>* and &quot;0&quot; means AN interictal sample</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55189,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "09/24/2014 22:50:31",
      "content": "<p>I think so but instead of 0 and 1, an output of probability (or a rank between 0 and 1) of &quot;preictal&quot; event is much more practical to improve roc score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55190,
      "author_name": "kaguiar",
      "author_url": "",
      "post_date": "09/24/2014 22:57:23",
      "content": "<p>Yes, indeed.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55684,
      "author_name": "adamsoffer",
      "author_url": "",
      "post_date": "10/06/2014 14:15:02",
      "content": "<p>Why is outputting the probability of a preictal event (as opposed to a discrete class label 0 or 1) more practical to improve ROC score?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55693,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "10/06/2014 17:46:04",
      "content": "<p>Because the ROC iterates over several different thresholds for the classification, its possible that your results will have a better accuracy at some threshold versus others.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55726,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "10/07/2014 11:10:34",
      "content": "<p>[quote=franklyn;55693]</p>\n<p>Because the ROC iterates over several different thresholds for the classification...</p>\n<p>[/quote]</p>\n<p>Yes, exactly.</p>\n<p>To illustrate:</p>\n<p><code>from sklearn.metrics import roc_auc_score</code></p>\n<p><code></code><code>y_predicted_discrete = [0.0, 0.0, 0.0, 1.0, 1.0]</code></p>\n<p><code>y_predicted_proba = [0.02, 0.01, 0.05, 0.99, 0.99]</code></p>\n<p><code>y_ground_truth = [0, 0, 1, 1, 1]</code></p>\n<p><code>print roc_auc_score(y_ground_truth, y_predicted_discrete) # 0.83</code></p>\n<p><code>print roc_auc_score(y_ground_truth, y_predicted_proba) # 1.0</code></p>\n<p>Even though the 3rd prediction probability was just&nbsp;0.05 and the truth was 1, it still got credit for ranking it higher, causing a favorable perfect split.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55737,
      "author_name": "adamsoffer",
      "author_url": "",
      "post_date": "10/07/2014 14:31:29",
      "content": "<p>Gotcha. Thanks. So one&nbsp;can one still receive a perfect score by outputting probabilities as opposed to&nbsp;discrete labels?&nbsp;In other words, even if the third predicted discrete value was correct in your example, both sets&nbsp;would have received the same score under ROC?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55740,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "10/07/2014 16:41:38",
      "content": "<p>[quote=Adam Soffer;55737]</p>\n<p>Gotcha. Thanks. So one&nbsp;can one still receive a perfect score by outputting probabilities as opposed to&nbsp;discrete labels?&nbsp;In other words, even if the third predicted discrete value was correct in your example, both sets&nbsp;would have received the same score under ROC?</p>\n<p>[/quote]</p>\n<p>I take it you meant:&nbsp;one can still receive a perfect score by outputting discrete labels&nbsp;as opposed to probabilities? That answer is yes. It is possible if you yourself have found the perfect split/threshold between the positive and the negative class. AUC finds this for you by considering every threshold split and only picking the best your model has got.</p>\n<p>Since this contest is not judged on raw accuracy, it would not be taking advantage of the evaluation metric if you still submit discrete labels: Those high on the leaderboard are submitting probabilities or ranks.</p>\n<p>Other <a href=\"http://www.kaggle.com/wiki/Metrics\">evaluation metrics</a>&nbsp;call for other &quot;tactics&quot;.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "55185": "",
    "55186": "",
    "55189": "",
    "55190": "",
    "55684": "",
    "55693": "",
    "55726": "",
    "55737": "",
    "55740": ""
  },
  "source": "meta"
}