{
  "id": 10528,
  "title": "AUC issues",
  "url": "/competitions/seizure-prediction/discussion/10528",
  "author_name": "",
  "post_date": "2014-10-05T08:22:07.293Z",
  "votes": null,
  "comment_count": 9,
  "views": 1868,
  "content": "<p>Hey Folks</p>\n<p>Looking for some insight here. &nbsp;I have tried submitting files for individual subjects (with other subjects filled with zeros), to see how they do, but I am puzzled because they often return AUC &lt; 0.5, suggesting that they are worse than random. &nbsp;So I just tried submitting a set (Dog 5) which returned 0.49813, then I inverted that set (1 - values) which should in principle have an AUC &gt; 0.5, but it didn't - it returned 0.49805. &nbsp;So clearly, I am thinking too simplistically about AUC, if both the data set and its inverse return AUC &lt; 0.5. &nbsp;Anyone have any thoughts on this? &nbsp;What am I missing?</p>",
  "messages": [
    {
      "id": "55627",
      "postDate": "10/05/2014 08:22:07",
      "content": "<p>Hey Folks</p>\n<p>Looking for some insight here. &nbsp;I have tried submitting files for individual subjects (with other subjects filled with zeros), to see how they do, but I am puzzled because they often return AUC &lt; 0.5, suggesting that they are worse than random. &nbsp;So I just tried submitting a set (Dog 5) which returned 0.49813, then I inverted that set (1 - values) which should in principle have an AUC &gt; 0.5, but it didn't - it returned 0.49805. &nbsp;So clearly, I am thinking too simplistically about AUC, if both the data set and its inverse return AUC &lt; 0.5. &nbsp;Anyone have any thoughts on this? &nbsp;What am I missing?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55629",
      "postDate": "10/05/2014 09:12:57",
      "content": "<p>Hi there.&nbsp;I would suggest studying&nbsp; a bit more about the ROC&nbsp;AUC.&nbsp;&nbsp;The AUC for p and 1-p would sum to one if and only if the test set is exactly balanced (i.e same zeros and ones) or p=0.5 which here is not the case.</p>\n<p>Edit: I would advise when you CV to check your classifiers sensitivity and specificity alongiside the rather than just AUC</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55664",
      "postDate": "10/06/2014 03:39:15",
      "content": "<p>Thanks for that information. &nbsp;I was surprised because I have done a fair bit of reading about AUC, and e.g. this paper posted by Michael Hills &nbsp;https://cours.etsmtl.ca/sys828/REFS/A1/Fawcett_PRL2006.pdf makes it clear that for 1-p you should get 1 - AUC, and that AUC is very insensitive to skew data sets. &nbsp;However, I experimented again and found that if I also invert the rest of the data set (from all 0 to all 1) then I do get the expected behaviour (1 - p gives 1 - AUC), so there is some kind of&nbsp;interaction between the curve for the rest of the data, and the single subject I am submitting - which is probably related to the imbalance as you suggest.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55665",
      "postDate": "10/06/2014 03:59:57",
      "content": "<p>I too am under the impression that (1 - p) should definitely give you (1 - AUC) regardless of imbalance. I think you are seeing the issue that I was describing in that the performances between my per-patient classifiers affect each other in the leaderboard ROC AUC. It was confirmed by competition admin that the effect I was alluding to in my question&nbsp;was correct and I think you are seeing this same effect happening with your submissions. In fact I stumbled across the issue doing what you were doing, submitting all 0s except for 1 patient at a time... trying to figure out how well I was doing on each patient. I never did find that out.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55681",
      "postDate": "10/06/2014 13:58:00",
      "content": "<p>Hmm.Probably I am the one that needs to study more :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55682",
      "postDate": "10/06/2014 13:59:48",
      "content": "<p>I think this is another phenomenon to do with the difference between averaging over 7 AUCs &nbsp;versus calculating one AUC for the union of the 7 subjects.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55729",
      "postDate": "10/07/2014 11:35:24",
      "content": "<p>[quote=Jonathan Tapson;55627]</p>\n<p>Hey Folks</p>\n<p>Looking for some insight here. &nbsp;I have tried submitting files for individual subjects (with other subjects filled with zeros), to see how they do, but I am puzzled because they often return AUC &lt; 0.5, suggesting that they are worse than random. &nbsp;So I just tried submitting a set (Dog 5) which returned 0.49813, then I inverted that set (1 - values) which should in principle have an AUC &gt; 0.5, but it didn't - it returned 0.49805. &nbsp;So clearly, I am thinking too simplistically about AUC, if both the data set and its inverse return AUC &lt; 0.5. &nbsp;Anyone have any thoughts on this? &nbsp;What am I missing?</p>\n<p>[/quote]</p>\n\n<p>Dear Jonathan,</p>\n\n<p>I assume that you inverted the whole submission, hence submitting 1s wherever you previously submitted 0s. Am I right?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55732",
      "postDate": "10/07/2014 12:22:10",
      "content": "<p>Dear Jose</p>\n<p>As mentioned later in the thread, I did not, initially. &nbsp;When I eventually did, I got the anticipated result (1 - p) gives (1 - AUC). &nbsp;However, apart from an intuition that this is an effect of imbalanced data sets, I still don't really understand why it makes a difference.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55800",
      "postDate": "10/08/2014 18:25:31",
      "content": "<p>I'd like to raise another (related) point. &nbsp;I'd like to ask the organizers how the AUC score is computed. &nbsp;Is the final AUC score calculated over the entire set, or is it averaged over each subject's AUC score? &nbsp;It's possible to get 1.0 AUCs on all of the test sets individually but get a lower score if the test examples are pooled. &nbsp;This might account for the poor test results we're seeing.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55804",
      "postDate": "10/08/2014 20:28:30",
      "content": "<p>Tom,</p>\n<p>I think the discussion at <a href=\"https://www.kaggle.com/c/seizure-prediction/forums/t/10383/leaderboard-metric-roc-auc/54251\">https://www.kaggle.com/c/seizure-prediction/forums/t/10383/leaderboard-metric-roc-auc/54251</a>&nbsp;may help.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 55629,
      "author_name": "epinephelus",
      "author_url": "",
      "post_date": "10/05/2014 09:12:57",
      "content": "<p>Hi there.&nbsp;I would suggest studying&nbsp; a bit more about the ROC&nbsp;AUC.&nbsp;&nbsp;The AUC for p and 1-p would sum to one if and only if the test set is exactly balanced (i.e same zeros and ones) or p=0.5 which here is not the case.</p>\n<p>Edit: I would advise when you CV to check your classifiers sensitivity and specificity alongiside the rather than just AUC</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55664,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "10/06/2014 03:39:15",
      "content": "<p>Thanks for that information. &nbsp;I was surprised because I have done a fair bit of reading about AUC, and e.g. this paper posted by Michael Hills &nbsp;https://cours.etsmtl.ca/sys828/REFS/A1/Fawcett_PRL2006.pdf makes it clear that for 1-p you should get 1 - AUC, and that AUC is very insensitive to skew data sets. &nbsp;However, I experimented again and found that if I also invert the rest of the data set (from all 0 to all 1) then I do get the expected behaviour (1 - p gives 1 - AUC), so there is some kind of&nbsp;interaction between the curve for the rest of the data, and the single subject I am submitting - which is probably related to the imbalance as you suggest.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55665,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "10/06/2014 03:59:57",
      "content": "<p>I too am under the impression that (1 - p) should definitely give you (1 - AUC) regardless of imbalance. I think you are seeing the issue that I was describing in that the performances between my per-patient classifiers affect each other in the leaderboard ROC AUC. It was confirmed by competition admin that the effect I was alluding to in my question&nbsp;was correct and I think you are seeing this same effect happening with your submissions. In fact I stumbled across the issue doing what you were doing, submitting all 0s except for 1 patient at a time... trying to figure out how well I was doing on each patient. I never did find that out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55681,
      "author_name": "epinephelus",
      "author_url": "",
      "post_date": "10/06/2014 13:58:00",
      "content": "<p>Hmm.Probably I am the one that needs to study more :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55682,
      "author_name": "nickau",
      "author_url": "",
      "post_date": "10/06/2014 13:59:48",
      "content": "<p>I think this is another phenomenon to do with the difference between averaging over 7 AUCs &nbsp;versus calculating one AUC for the union of the 7 subjects.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55729,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "10/07/2014 11:35:24",
      "content": "<p>[quote=Jonathan Tapson;55627]</p>\n<p>Hey Folks</p>\n<p>Looking for some insight here. &nbsp;I have tried submitting files for individual subjects (with other subjects filled with zeros), to see how they do, but I am puzzled because they often return AUC &lt; 0.5, suggesting that they are worse than random. &nbsp;So I just tried submitting a set (Dog 5) which returned 0.49813, then I inverted that set (1 - values) which should in principle have an AUC &gt; 0.5, but it didn't - it returned 0.49805. &nbsp;So clearly, I am thinking too simplistically about AUC, if both the data set and its inverse return AUC &lt; 0.5. &nbsp;Anyone have any thoughts on this? &nbsp;What am I missing?</p>\n<p>[/quote]</p>\n\n<p>Dear Jonathan,</p>\n\n<p>I assume that you inverted the whole submission, hence submitting 1s wherever you previously submitted 0s. Am I right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55732,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "10/07/2014 12:22:10",
      "content": "<p>Dear Jose</p>\n<p>As mentioned later in the thread, I did not, initially. &nbsp;When I eventually did, I got the anticipated result (1 - p) gives (1 - AUC). &nbsp;However, apart from an intuition that this is an effect of imbalanced data sets, I still don't really understand why it makes a difference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55800,
      "author_name": "",
      "author_url": "",
      "post_date": "10/08/2014 18:25:31",
      "content": "<p>I'd like to raise another (related) point. &nbsp;I'd like to ask the organizers how the AUC score is computed. &nbsp;Is the final AUC score calculated over the entire set, or is it averaged over each subject's AUC score? &nbsp;It's possible to get 1.0 AUCs on all of the test sets individually but get a lower score if the test examples are pooled. &nbsp;This might account for the poor test results we're seeing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55804,
      "author_name": "trentb",
      "author_url": "",
      "post_date": "10/08/2014 20:28:30",
      "content": "<p>Tom,</p>\n<p>I think the discussion at <a href=\"https://www.kaggle.com/c/seizure-prediction/forums/t/10383/leaderboard-metric-roc-auc/54251\">https://www.kaggle.com/c/seizure-prediction/forums/t/10383/leaderboard-metric-roc-auc/54251</a>&nbsp;may help.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "55627": "",
    "55629": "",
    "55664": "",
    "55665": "",
    "55681": "",
    "55682": "",
    "55729": "",
    "55732": "",
    "55800": "",
    "55804": ""
  },
  "source": "meta"
}