{
  "id": 12603,
  "title": "Inter subject AUC",
  "url": "/competitions/inria-bci-challenge/discussion/12603",
  "author_name": "",
  "post_date": "2015-02-25T00:17:50.647Z",
  "votes": 1,
  "comment_count": 13,
  "views": 2848,
  "content": "<p>There was a large difference between the mean of individual subjects AUC's and the group AUC.&nbsp; My model from a month ago had ~.80 mean AUC and the group AUC was ~.74. I added the number of errors detected for each subject from the leaked labels as a feature and this increased my group AUC to .8.(the labels themselves only helped by 0.001) At the end I added positive label counts from my classifier as well and this pushed my group AUC over .82 on CV. The mean AUC stayed the same for all the different models.</p>\n<p>Does anyone have an explanation for why this is happening? It was surprising to see.</p>",
  "messages": [
    {
      "id": "64829",
      "postDate": "02/25/2015 00:17:50",
      "content": "<p>There was a large difference between the mean of individual subjects AUC's and the group AUC.&nbsp; My model from a month ago had ~.80 mean AUC and the group AUC was ~.74. I added the number of errors detected for each subject from the leaked labels as a feature and this increased my group AUC to .8.(the labels themselves only helped by 0.001) At the end I added positive label counts from my classifier as well and this pushed my group AUC over .82 on CV. The mean AUC stayed the same for all the different models.</p>\n<p>Does anyone have an explanation for why this is happening? It was surprising to see.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64831",
      "postDate": "02/25/2015 00:27:32",
      "content": "<p>I am not sure I get what you did. Perhaps&nbsp;some leakage? The difference between 1-2-3 and the rest of field seems significant.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64832",
      "postDate": "02/25/2015 00:38:29",
      "content": "<p>There was some leakage, but using the leaked labels didn't help much. Counting how many errors were detected for each subject and using this as a feature increased the global AUC by about 0.05 while leaving mean subject AUC unchanged.</p>\n<p>One subject only had 5 errors detected out of the last 100 trials. Average was about 35. The error increased group AUC. I am not sure how though.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64835",
      "postDate": "02/25/2015 00:48:37",
      "content": "<p>Ah I see! Great and creative meta-feature!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64836",
      "postDate": "02/25/2015 01:00:28",
      "content": "<p>That is some weird stuff. Can you explain a bit more how you added counts of the number of errors detected for the test set? Did you run your models on the test set, count errors, then run the models again with the counts?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64837",
      "postDate": "02/25/2015 01:03:32",
      "content": "<p>The counts were from the leaked labels so they were available for both the train and the test set. I did add counts from my own classifier this added another .01 to my global AUC score during CV.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64839",
      "postDate": "02/25/2015 01:11:15",
      "content": "<p>[quote=phalaris;64837]</p>\n<p>The counts were from the leaked labels so they were available for both the train and the test set. I did add counts from my own classifier this added another .01 to my global AUC score during CV.</p>\n<p>[/quote]</p>\n<p>Based on your earlier posts, this seems like it only applied to the 5th session. Was there something more that was leaked?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64840",
      "postDate": "02/25/2015 01:13:53",
      "content": "<p>Nothing else. The counts were from fifth session only.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64841",
      "postDate": "02/25/2015 01:28:34",
      "content": "<p>Craziness. Way to go taking advantage of the leak!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64843",
      "postDate": "02/25/2015 02:23:11",
      "content": "<p>This information was the primary driver of my late jump. Phalaris, if I'm understanding your question, I think that since &nbsp;your feature spanned the entire subject, it didn't change the intra-subject ordering much, but dramatically helped the inter-subject ordering.</p>\n<p>And to put it into context, look at the difference between the submission that got 3rd and the straight python benchmark code you provided.</p>\n<p>My final model:<br><img src=\"https://cloud.githubusercontent.com/assets/2976822/6363572/970a4ef8-bc60-11e4-8b65-bd38c155b9be.png\" alt=\"3rd Place Model\" width=\"1855\" height=\"1056\"></p>\n<p>Forum script:<br><img src=\"https://cloud.githubusercontent.com/assets/2976822/6363577/b09b3fee-bc60-11e4-897e-6b0224b0ff79.png\" alt=\"Forum GBM Script\" width=\"1855\" height=\"1056\"></p>\n<p>The forum script mostly hinges on the session ID, so it pretty well delineates subjects. The subjects where the inter-subject AUC is dramatic is the 7th and 8th. On my model, those are extremely high because they had the least abnormal timings. The vanilla benchmark didn't understand that. My guess is that since subject ID was part of the &nbsp;python model, that affected how it ordered subjects such that subjects 4, 5, and 6 were high. Looking at the code, subject ID seems to be passed in as a factor, but that would be strange because the test set wouldn't agree. If it was an integer, that would make more sense--illogical cutpoints were found and used that won't scale.</p>\n<p>Congratulations to Phalaris for doing so much in this competition. The benchmark code was still a good starting point. And I have never seen someone try so &nbsp;hard to flag attention about a leak. I only figured out how to use it a few days ago. But two separate posts have been out there calling attention to this.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64844",
      "postDate": "02/25/2015 02:49:48",
      "content": "<p>@mlandry, Thanks! Those plots are very informative. I am guessing that subjects which have mainly positive labels(the max in training set is 316/340) will have most of there probabilities clustered up toward 1. Whereas those with mainly negative labels(min in training is 178/340) will be closer to 0.</p>\n\n<p>I wonder if there is a way to account for the different class percentages for the different subjects without the leaked labels. The mean of single subject AUC scores seemed to have very little impact on final score.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64849",
      "postDate": "02/25/2015 06:09:24",
      "content": "<p>Just checked the benchmark I posted. Removing the subject column increases the private leaderboard score from 0.47 to 0.6. Sorry to everyone who had submissions messed up by this.</p>\n<p>@mlandry, good catch on what caused the problem. I will have to be more careful about throwing every feature I can think of into my models.</p>\n<p>When I originally included the subject variable it was as a convenience to make querying for specific subjects easier. I was sure it would have no impact so I didn't drop it from the dataset before training. Soon after posting the benchmark I removed subject from my train set as I used another method for selecting specific subjects. I am amazed that it had such a negative effect on the private leaderboard score.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64908",
      "postDate": "02/26/2015 03:01:53",
      "content": "<p>@phalaris how can you see if there is an increase in the 'private leaderboard score' ??</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64909",
      "postDate": "02/26/2015 03:35:22",
      "content": "<p>If you resubmit a solution it will show you the score on the private leaderboard.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 64831,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "02/25/2015 00:27:32",
      "content": "<p>I am not sure I get what you did. Perhaps&nbsp;some leakage? The difference between 1-2-3 and the rest of field seems significant.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64832,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/25/2015 00:38:29",
      "content": "<p>There was some leakage, but using the leaked labels didn't help much. Counting how many errors were detected for each subject and using this as a feature increased the global AUC by about 0.05 while leaving mean subject AUC unchanged.</p>\n<p>One subject only had 5 errors detected out of the last 100 trials. Average was about 35. The error increased group AUC. I am not sure how though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64835,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "02/25/2015 00:48:37",
      "content": "<p>Ah I see! Great and creative meta-feature!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64836,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "02/25/2015 01:00:28",
      "content": "<p>That is some weird stuff. Can you explain a bit more how you added counts of the number of errors detected for the test set? Did you run your models on the test set, count errors, then run the models again with the counts?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64837,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/25/2015 01:03:32",
      "content": "<p>The counts were from the leaked labels so they were available for both the train and the test set. I did add counts from my own classifier this added another .01 to my global AUC score during CV.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64839,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "02/25/2015 01:11:15",
      "content": "<p>[quote=phalaris;64837]</p>\n<p>The counts were from the leaked labels so they were available for both the train and the test set. I did add counts from my own classifier this added another .01 to my global AUC score during CV.</p>\n<p>[/quote]</p>\n<p>Based on your earlier posts, this seems like it only applied to the 5th session. Was there something more that was leaked?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64840,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/25/2015 01:13:53",
      "content": "<p>Nothing else. The counts were from fifth session only.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64841,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "02/25/2015 01:28:34",
      "content": "<p>Craziness. Way to go taking advantage of the leak!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64843,
      "author_name": "mlandry",
      "author_url": "",
      "post_date": "02/25/2015 02:23:11",
      "content": "<p>This information was the primary driver of my late jump. Phalaris, if I'm understanding your question, I think that since &nbsp;your feature spanned the entire subject, it didn't change the intra-subject ordering much, but dramatically helped the inter-subject ordering.</p>\n<p>And to put it into context, look at the difference between the submission that got 3rd and the straight python benchmark code you provided.</p>\n<p>My final model:<br><img src=\"https://cloud.githubusercontent.com/assets/2976822/6363572/970a4ef8-bc60-11e4-8b65-bd38c155b9be.png\" alt=\"3rd Place Model\" width=\"1855\" height=\"1056\"></p>\n<p>Forum script:<br><img src=\"https://cloud.githubusercontent.com/assets/2976822/6363577/b09b3fee-bc60-11e4-897e-6b0224b0ff79.png\" alt=\"Forum GBM Script\" width=\"1855\" height=\"1056\"></p>\n<p>The forum script mostly hinges on the session ID, so it pretty well delineates subjects. The subjects where the inter-subject AUC is dramatic is the 7th and 8th. On my model, those are extremely high because they had the least abnormal timings. The vanilla benchmark didn't understand that. My guess is that since subject ID was part of the &nbsp;python model, that affected how it ordered subjects such that subjects 4, 5, and 6 were high. Looking at the code, subject ID seems to be passed in as a factor, but that would be strange because the test set wouldn't agree. If it was an integer, that would make more sense--illogical cutpoints were found and used that won't scale.</p>\n<p>Congratulations to Phalaris for doing so much in this competition. The benchmark code was still a good starting point. And I have never seen someone try so &nbsp;hard to flag attention about a leak. I only figured out how to use it a few days ago. But two separate posts have been out there calling attention to this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64844,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/25/2015 02:49:48",
      "content": "<p>@mlandry, Thanks! Those plots are very informative. I am guessing that subjects which have mainly positive labels(the max in training set is 316/340) will have most of there probabilities clustered up toward 1. Whereas those with mainly negative labels(min in training is 178/340) will be closer to 0.</p>\n\n<p>I wonder if there is a way to account for the different class percentages for the different subjects without the leaked labels. The mean of single subject AUC scores seemed to have very little impact on final score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64849,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/25/2015 06:09:24",
      "content": "<p>Just checked the benchmark I posted. Removing the subject column increases the private leaderboard score from 0.47 to 0.6. Sorry to everyone who had submissions messed up by this.</p>\n<p>@mlandry, good catch on what caused the problem. I will have to be more careful about throwing every feature I can think of into my models.</p>\n<p>When I originally included the subject variable it was as a convenience to make querying for specific subjects easier. I was sure it would have no impact so I didn't drop it from the dataset before training. Soon after posting the benchmark I removed subject from my train set as I used another method for selecting specific subjects. I am amazed that it had such a negative effect on the private leaderboard score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64908,
      "author_name": "hithesh",
      "author_url": "",
      "post_date": "02/26/2015 03:01:53",
      "content": "<p>@phalaris how can you see if there is an increase in the 'private leaderboard score' ??</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64909,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/26/2015 03:35:22",
      "content": "<p>If you resubmit a solution it will show you the score on the private leaderboard.&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "64829": "",
    "64831": "",
    "64832": "",
    "64835": "",
    "64836": "",
    "64837": "",
    "64839": "",
    "64840": "",
    "64841": "",
    "64843": "",
    "64844": "",
    "64849": "",
    "64908": "",
    "64909": ""
  },
  "source": "meta"
}