{
  "id": 10828,
  "title": "Evaluation",
  "url": "/competitions/seizure-prediction/discussion/10828",
  "author_name": "",
  "post_date": "2014-11-03T03:53:39.030Z",
  "votes": null,
  "comment_count": 4,
  "views": 1118,
  "content": "<p>Hello,</p>\n<p>Does anybody know whether the annotations used by the administration for evaluating the results are real valued probabilities or are they simply 0 and 1?</p>\n<p>What made me ask this question is that I made a simple test submission with uniform random values between 0 and 1 and I got a score around 0.57. Next, I saturated the values of the same submission to 0 and 1 (by replacing the values smaller than 0.5 to 0 and the values greater than 0.5 to 1) and made another submission; but the score changed to 0.55. By the way, I had made sure that there were no values equal to exactly 0.5 in my first submission. The only explanation that I had for this observation was that perhaps the evaluation annotations used by the administrators are probabilistic themselves. Can anyone comment on this?</p>",
  "messages": [
    {
      "id": "57228",
      "postDate": "11/03/2014 03:53:39",
      "content": "<p>Hello,</p>\n<p>Does anybody know whether the annotations used by the administration for evaluating the results are real valued probabilities or are they simply 0 and 1?</p>\n<p>What made me ask this question is that I made a simple test submission with uniform random values between 0 and 1 and I got a score around 0.57. Next, I saturated the values of the same submission to 0 and 1 (by replacing the values smaller than 0.5 to 0 and the values greater than 0.5 to 1) and made another submission; but the score changed to 0.55. By the way, I had made sure that there were no values equal to exactly 0.5 in my first submission. The only explanation that I had for this observation was that perhaps the evaluation annotations used by the administrators are probabilistic themselves. Can anyone comment on this?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57237",
      "postDate": "11/03/2014 11:45:41",
      "content": "<p>I am pretty sure that evaluation happens in a deterministic way. By mapping all your scores to 0 and 1, you have changed the order of your scores (scores that had been different before are now the same), so clearly the AUC can change.</p>\n<p>To test this, you could for example not map to 0 and 1, but instead divide all your scores by ten, or divide all scores under 0.5 by ten and map all scores x above 0.5 to 1 - 0.1 * (1 - x) - in both cases, you wouldn't change the order of your scores, and you should get the exact same AUC as you got for your first submission.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57251",
      "postDate": "11/03/2014 15:04:41",
      "content": "<p>Thanks Ibrunjes,</p>\n<p>If the whole framework was stochastic (both the data and annotations used by the Admin were probabilistic) then the order of scores and the change of AUC were all reasonable. What I don't understand is that if the annotations used by the Admin are only 0 and 1, why do they expect us to report probabilities between 0 and 1? In other words, if the annotations are only 0 and 1, why should one care about reporting a preictal case with a probability of 0.2 or 0.3? That's why I think that the annotations are real values between 0 and 1. Any thoughts?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57253",
      "postDate": "11/03/2014 16:38:52",
      "content": "<p>I'm pretty sure that what you call their &quot;annotations&quot; are really 0's and 1's, no probabilities. I see two reasons why they want us to give probabilities:</p>\n<ul>\n<li>It's easier for us. If your algorithm gives 0.2 for all interictal samples and 0.4 for all preictal ones, you'll get perfect score, because only relative size counts for AUC. If you had to map everything to 0 and 1, you'd have to &quot;draw the line&quot; somewhere, for example at 0.5, and then for the 0.2/0.4 example, you'd get zero score, because all would be mapped to 0.</li>\n<li>They probably hope to get some finer-grained control when we provide probabilities, because by moving the threshold up or down, they can calibrate false positives versus false negatives, i.e. decide which is more important for them, having less &quot;false alarms&quot; or not missing any seizure warning.</li>\n</ul>\n<p>Of course nobody forces you to actually use values other than 0 and 1, it's just an option that you can take ot leave.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57349",
      "postDate": "11/05/2014 11:50:06",
      "content": "<p>That sounds reasonable. Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 57237,
      "author_name": "lbrunjes",
      "author_url": "",
      "post_date": "11/03/2014 11:45:41",
      "content": "<p>I am pretty sure that evaluation happens in a deterministic way. By mapping all your scores to 0 and 1, you have changed the order of your scores (scores that had been different before are now the same), so clearly the AUC can change.</p>\n<p>To test this, you could for example not map to 0 and 1, but instead divide all your scores by ten, or divide all scores under 0.5 by ten and map all scores x above 0.5 to 1 - 0.1 * (1 - x) - in both cases, you wouldn't change the order of your scores, and you should get the exact same AUC as you got for your first submission.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57251,
      "author_name": "r2241272",
      "author_url": "",
      "post_date": "11/03/2014 15:04:41",
      "content": "<p>Thanks Ibrunjes,</p>\n<p>If the whole framework was stochastic (both the data and annotations used by the Admin were probabilistic) then the order of scores and the change of AUC were all reasonable. What I don't understand is that if the annotations used by the Admin are only 0 and 1, why do they expect us to report probabilities between 0 and 1? In other words, if the annotations are only 0 and 1, why should one care about reporting a preictal case with a probability of 0.2 or 0.3? That's why I think that the annotations are real values between 0 and 1. Any thoughts?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57253,
      "author_name": "lbrunjes",
      "author_url": "",
      "post_date": "11/03/2014 16:38:52",
      "content": "<p>I'm pretty sure that what you call their &quot;annotations&quot; are really 0's and 1's, no probabilities. I see two reasons why they want us to give probabilities:</p>\n<ul>\n<li>It's easier for us. If your algorithm gives 0.2 for all interictal samples and 0.4 for all preictal ones, you'll get perfect score, because only relative size counts for AUC. If you had to map everything to 0 and 1, you'd have to &quot;draw the line&quot; somewhere, for example at 0.5, and then for the 0.2/0.4 example, you'd get zero score, because all would be mapped to 0.</li>\n<li>They probably hope to get some finer-grained control when we provide probabilities, because by moving the threshold up or down, they can calibrate false positives versus false negatives, i.e. decide which is more important for them, having less &quot;false alarms&quot; or not missing any seizure warning.</li>\n</ul>\n<p>Of course nobody forces you to actually use values other than 0 and 1, it's just an option that you can take ot leave.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57349,
      "author_name": "r2241272",
      "author_url": "",
      "post_date": "11/05/2014 11:50:06",
      "content": "<p>That sounds reasonable. Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57228": "",
    "57237": "",
    "57251": "",
    "57253": "",
    "57349": ""
  },
  "source": "meta"
}