{
  "id": 9382,
  "title": "Max achieveable AUC",
  "url": "/competitions/seizure-detection/discussion/9382",
  "author_name": "",
  "post_date": "2014-06-04T12:51:42.610Z",
  "votes": null,
  "comment_count": 9,
  "views": 2639,
  "content": "<p>Hello everyone,</p>\n<p>Looking at the top of the leaderboard with max AUC of ~0.95, I think that might be the maximum limit for the achieveable AUC. As there are artefacts as well as possible errors/human-inconsistency in marking seizure onset by the expert in original data. Furthermore, to get perfect classification, the algorithm must accurately distinguish between 15s and 16s seizure which is impossible to get right all the time, as the onset of seizure is fuzzy and expert marking the seizure cannot be always consistently accurate. I think if the same expert marks the dataset again visually, he might get AUC of around 0.95 compared to the original marking. I think theoratical limit might have been reached, and maybe beyond that, as a result of possible overfitting. This is just a thought and i may be completely wrong! what do you people think</p>\n<p>cheers</p>",
  "messages": [
    {
      "id": "48624",
      "postDate": "06/04/2014 12:51:42",
      "content": "<p>Hello everyone,</p>\n<p>Looking at the top of the leaderboard with max AUC of ~0.95, I think that might be the maximum limit for the achieveable AUC. As there are artefacts as well as possible errors/human-inconsistency in marking seizure onset by the expert in original data. Furthermore, to get perfect classification, the algorithm must accurately distinguish between 15s and 16s seizure which is impossible to get right all the time, as the onset of seizure is fuzzy and expert marking the seizure cannot be always consistently accurate. I think if the same expert marks the dataset again visually, he might get AUC of around 0.95 compared to the original marking. I think theoratical limit might have been reached, and maybe beyond that, as a result of possible overfitting. This is just a thought and i may be completely wrong! what do you people think</p>\n<p>cheers</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48642",
      "postDate": "06/04/2014 17:45:24",
      "content": "<p>I think the max AUC may be a little bit higher, around 0.96-0.97, but your are right. The fact that the true ground is defined by a human (human makes errors) let me believe that the perfect score is impossible to achieve. For example, on some subjects, the expert has been a little bit generous at the end of the seizure, and the&nbsp;last second of the seizure is&nbsp;usually miss-classified.</p>\n\n<p>The second problem is the distinction between early and late seizure. The threshold of 15s seems to be arbitrary. On some subjects, there is nothing in the feature that allows the classifier to discriminate a clip at 15s than a clip at 16s. I don't say&nbsp;there is no difference between&nbsp;the beginning&nbsp;and the end of a&nbsp;seizure, there is a clear separation, but the optimal latency is subject dependent. I will make a post about this later.</p>\n\n<p>Finally, i still have trouble to understand how the evaluation is done. I get my best score by submitting the same probability for seizure and early, which is really disturbing.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48653",
      "postDate": "06/04/2014 20:09:03",
      "content": "<p>I also think we are going to hit a ceiling soon and start to overfit. When I look at my CV results there don't seem to be much left to&nbsp;do, except optimise for public leaderboard ranking. That, of course, would lead to overfitting.&nbsp;</p>\n\n<p>[quote=Alexandre;48642]</p>\n<p>Finally, i still have trouble to understand how the evaluation is done. I get my best score by submitting the same probability for seizure and early, which is really disturbing.<br>[/quote]</p>\n<p>That is strange. Of course, there is a connection between seizure and early but this really surprises me. If you don't mind, I am going to try that&nbsp;approach tomorrow to see if I can reproduce the behaviour.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48655",
      "postDate": "06/04/2014 20:35:09",
      "content": "<p>[quote=Ulo Gulo;48653]</p>\n<p>That is strange. Of course, there is a connection between seizure and early but this really surprises me. If you don't mind, I am going to try that&nbsp;approach tomorrow to see if I can reproduce the behaviour.</p>\n<p>[/quote]</p>\n\n<p>I was talking about my score on le LB. Of course, when i do this in CV, my AUC drops.</p>\n\n<p>Another interesting thing to do is to submit all zeros for early, so you can get the seizure Vs. interictal AUC. Mine is around 0.97 (perhaps more, i didn't try on my last submission).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48693",
      "postDate": "06/05/2014 08:48:48",
      "content": "<p>Alexandre, you are the man. Kudos for discovering this. My LB score also improves when I&nbsp;submit the same probabilities for seizure and early. This is, indeed, confusing. I have got to think about that...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48698",
      "postDate": "06/05/2014 10:20:28",
      "content": "<p>If I try this by copying my predictions for seizure and using them as my predictions for early (from my best submission file) then I score ~0.02 lower on the leaderboard.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48704",
      "postDate": "06/05/2014 10:51:45",
      "content": "<p>I have observed this in one of my other method. The first classifier is supposed to be less sensitive to early detection, so you can get a lower score. but a decrease of 0.02 is still too low.</p>\n\n<p>We all have high AUC, i think around 0.95-0.97 for the task of seizure Vs. Interictal. This mean that the&nbsp;majority of the seizures are correctly detected. Using&nbsp;the same score for the early detection should produce&nbsp;lot of false positive (all the late seizure should count as a FP) and we should see a drop of at least 0.1</p>\n\n<p>I can see only three&nbsp;explanation for this :&nbsp;</p>\n<ol>\n<li>The late seizures are ignored in the estimation of early AUC.</li>\n<li>The Public/Private LB split&nbsp;is done in a such way that there is almost no late seizures in the public LB</li>\n<li>Somehow, we are doing something wrong.</li>\n</ol>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48705",
      "postDate": "06/05/2014 11:37:10",
      "content": "<p>Interesting, i would test that</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48725",
      "postDate": "06/05/2014 15:04:12",
      "content": "<p>[quote=Alexandre;48642]</p>\n<p>I get my best score by submitting the same probability for seizure and early, which is really disturbing.</p>\n<p>[/quote]</p>\n\n<p>Interesting! I tried that, but it lowered my score ~ 0.02.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48769",
      "postDate": "06/06/2014 11:07:48",
      "content": "<p>Coming back to the topic of the thread, I would say the the theoretical limit of AUC1=1 would be reachable if the competition organizers removed the samples close to the beginning of the seizures when building the test set. Also, AUC2 = 1 is achievable if samples with latency between 14 and 16 (for example) were removed.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 48642,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/04/2014 17:45:24",
      "content": "<p>I think the max AUC may be a little bit higher, around 0.96-0.97, but your are right. The fact that the true ground is defined by a human (human makes errors) let me believe that the perfect score is impossible to achieve. For example, on some subjects, the expert has been a little bit generous at the end of the seizure, and the&nbsp;last second of the seizure is&nbsp;usually miss-classified.</p>\n\n<p>The second problem is the distinction between early and late seizure. The threshold of 15s seems to be arbitrary. On some subjects, there is nothing in the feature that allows the classifier to discriminate a clip at 15s than a clip at 16s. I don't say&nbsp;there is no difference between&nbsp;the beginning&nbsp;and the end of a&nbsp;seizure, there is a clear separation, but the optimal latency is subject dependent. I will make a post about this later.</p>\n\n<p>Finally, i still have trouble to understand how the evaluation is done. I get my best score by submitting the same probability for seizure and early, which is really disturbing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48653,
      "author_name": "ugrossek",
      "author_url": "",
      "post_date": "06/04/2014 20:09:03",
      "content": "<p>I also think we are going to hit a ceiling soon and start to overfit. When I look at my CV results there don't seem to be much left to&nbsp;do, except optimise for public leaderboard ranking. That, of course, would lead to overfitting.&nbsp;</p>\n\n<p>[quote=Alexandre;48642]</p>\n<p>Finally, i still have trouble to understand how the evaluation is done. I get my best score by submitting the same probability for seizure and early, which is really disturbing.<br>[/quote]</p>\n<p>That is strange. Of course, there is a connection between seizure and early but this really surprises me. If you don't mind, I am going to try that&nbsp;approach tomorrow to see if I can reproduce the behaviour.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48655,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/04/2014 20:35:09",
      "content": "<p>[quote=Ulo Gulo;48653]</p>\n<p>That is strange. Of course, there is a connection between seizure and early but this really surprises me. If you don't mind, I am going to try that&nbsp;approach tomorrow to see if I can reproduce the behaviour.</p>\n<p>[/quote]</p>\n\n<p>I was talking about my score on le LB. Of course, when i do this in CV, my AUC drops.</p>\n\n<p>Another interesting thing to do is to submit all zeros for early, so you can get the seizure Vs. interictal AUC. Mine is around 0.97 (perhaps more, i didn't try on my last submission).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48693,
      "author_name": "ugrossek",
      "author_url": "",
      "post_date": "06/05/2014 08:48:48",
      "content": "<p>Alexandre, you are the man. Kudos for discovering this. My LB score also improves when I&nbsp;submit the same probabilities for seizure and early. This is, indeed, confusing. I have got to think about that...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48698,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "06/05/2014 10:20:28",
      "content": "<p>If I try this by copying my predictions for seizure and using them as my predictions for early (from my best submission file) then I score ~0.02 lower on the leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48704,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/05/2014 10:51:45",
      "content": "<p>I have observed this in one of my other method. The first classifier is supposed to be less sensitive to early detection, so you can get a lower score. but a decrease of 0.02 is still too low.</p>\n\n<p>We all have high AUC, i think around 0.95-0.97 for the task of seizure Vs. Interictal. This mean that the&nbsp;majority of the seizures are correctly detected. Using&nbsp;the same score for the early detection should produce&nbsp;lot of false positive (all the late seizure should count as a FP) and we should see a drop of at least 0.1</p>\n\n<p>I can see only three&nbsp;explanation for this :&nbsp;</p>\n<ol>\n<li>The late seizures are ignored in the estimation of early AUC.</li>\n<li>The Public/Private LB split&nbsp;is done in a such way that there is almost no late seizures in the public LB</li>\n<li>Somehow, we are doing something wrong.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48705,
      "author_name": "chaotic",
      "author_url": "",
      "post_date": "06/05/2014 11:37:10",
      "content": "<p>Interesting, i would test that</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48725,
      "author_name": "ohiojoe",
      "author_url": "",
      "post_date": "06/05/2014 15:04:12",
      "content": "<p>[quote=Alexandre;48642]</p>\n<p>I get my best score by submitting the same probability for seizure and early, which is really disturbing.</p>\n<p>[/quote]</p>\n\n<p>Interesting! I tried that, but it lowered my score ~ 0.02.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48769,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "06/06/2014 11:07:48",
      "content": "<p>Coming back to the topic of the thread, I would say the the theoretical limit of AUC1=1 would be reachable if the competition organizers removed the samples close to the beginning of the seizures when building the test set. Also, AUC2 = 1 is achievable if samples with latency between 14 and 16 (for example) were removed.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "48624": "",
    "48642": "",
    "48653": "",
    "48655": "",
    "48693": "",
    "48698": "",
    "48704": "",
    "48705": "",
    "48725": "",
    "48769": ""
  },
  "source": "meta"
}