{
  "id": 11166,
  "title": "Optimum Cross Validation Method?",
  "url": "/competitions/inria-bci-challenge/discussion/11166",
  "author_name": "",
  "post_date": "2014-12-07T20:15:25.743Z",
  "votes": 1,
  "comment_count": 8,
  "views": 3358,
  "content": "<p>I have been experimenting with 2 combinations of cross validaters - Leave 2 subjects out and direct 4 fold cv and noticed that the 4 fold cv seems to be reflecting the leaderboard score better than Leave 2 subjects out although I had anticipated the opposite.</p>\n<p>Any ideas for a good cross validation technique?</p>",
  "messages": [
    {
      "id": "59743",
      "postDate": "12/07/2014 20:15:25",
      "content": "<p>I have been experimenting with 2 combinations of cross validaters - Leave 2 subjects out and direct 4 fold cv and noticed that the 4 fold cv seems to be reflecting the leaderboard score better than Leave 2 subjects out although I had anticipated the opposite.</p>\n<p>Any ideas for a good cross validation technique?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59749",
      "postDate": "12/07/2014 22:12:03",
      "content": "<p>I think the most realistic&nbsp;approach is leave-one-subject-out, possibly followed by calculation of AUC over the combined predictions from each held-out subject.</p>\n<p>In my opinion, the leaderboard score is fairly worthless. It is calculated from only two test subjects, S09 and S25, which appear easier than average to predict.</p>\n<p>Cross-subject calibration of outputs may be important to the final score, however I don't think 2 subjects is enough to give useful&nbsp;feedback&nbsp;on that problem.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59761",
      "postDate": "12/08/2014 05:16:04",
      "content": "<p>Personally I have found that my internal CV is matching the leaderboard extremely well. The exact numbers&nbsp;don't match up, but whenever I see an improvement in my CV, that is almost exactly the improvement I get in the leaderboard.</p>\n<p>For CV I'm just using 4 folds, each with 4 different subjects.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59764",
      "postDate": "12/08/2014 07:12:55",
      "content": "<p>So basically you are taking 4 subjects from the training set and applying 4 fold cross validation on it? Isnt this the same as Leave One Subject Out CV but applied on 4 subjects?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59785",
      "postDate": "12/08/2014 13:50:02",
      "content": "<p>Yes, pretty much.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61925",
      "postDate": "12/27/2014 06:52:04",
      "content": "<p>emolson, could you tell us the place where you read the information stating that the LB score is based on subjects S09 and S25? I didn't find it on the competition pages (perhaps it is in the cited paper...I didn't finish reading it yet). Your help will be greatly appreciated.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61926",
      "postDate": "12/27/2014 06:58:19",
      "content": "<p>I think they figured that out by changing the submission results for one subject at a time to see if&nbsp;the leaderboard score changed. See the bottom of his thread:&nbsp;http://www.kaggle.com/c/inria-bci-challenge/forums/t/11017/xgboost-boost-from-existing-predictions</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61928",
      "postDate": "12/27/2014 08:58:20",
      "content": "<p>Thank your for your answer TDeVries. If it is really the case, emelson is right in diminishing the importance of the LB. It is not of much worth even for helping adjusting hyper-parameters. Indeed, using it for validation can even hurt the final performance!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64525",
      "postDate": "02/18/2015 20:14:33",
      "content": "<p>You should average the error estimates across the CV folds to get an estimate of the testing error. Note, though, that this estimate is almost always a biased estimate of the true testing error.&nbsp; As others have noted, the local estimate seems to be a pretty bad indicator of the leaderboard score in this competition as the LB is only based on two subjects.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 59749,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "12/07/2014 22:12:03",
      "content": "<p>I think the most realistic&nbsp;approach is leave-one-subject-out, possibly followed by calculation of AUC over the combined predictions from each held-out subject.</p>\n<p>In my opinion, the leaderboard score is fairly worthless. It is calculated from only two test subjects, S09 and S25, which appear easier than average to predict.</p>\n<p>Cross-subject calibration of outputs may be important to the final score, however I don't think 2 subjects is enough to give useful&nbsp;feedback&nbsp;on that problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59761,
      "author_name": "tdevries",
      "author_url": "",
      "post_date": "12/08/2014 05:16:04",
      "content": "<p>Personally I have found that my internal CV is matching the leaderboard extremely well. The exact numbers&nbsp;don't match up, but whenever I see an improvement in my CV, that is almost exactly the improvement I get in the leaderboard.</p>\n<p>For CV I'm just using 4 folds, each with 4 different subjects.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59764,
      "author_name": "tangy1994",
      "author_url": "",
      "post_date": "12/08/2014 07:12:55",
      "content": "<p>So basically you are taking 4 subjects from the training set and applying 4 fold cross validation on it? Isnt this the same as Leave One Subject Out CV but applied on 4 subjects?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59785,
      "author_name": "tdevries",
      "author_url": "",
      "post_date": "12/08/2014 13:50:02",
      "content": "<p>Yes, pretty much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61925,
      "author_name": "saulberardo",
      "author_url": "",
      "post_date": "12/27/2014 06:52:04",
      "content": "<p>emolson, could you tell us the place where you read the information stating that the LB score is based on subjects S09 and S25? I didn't find it on the competition pages (perhaps it is in the cited paper...I didn't finish reading it yet). Your help will be greatly appreciated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61926,
      "author_name": "tdevries",
      "author_url": "",
      "post_date": "12/27/2014 06:58:19",
      "content": "<p>I think they figured that out by changing the submission results for one subject at a time to see if&nbsp;the leaderboard score changed. See the bottom of his thread:&nbsp;http://www.kaggle.com/c/inria-bci-challenge/forums/t/11017/xgboost-boost-from-existing-predictions</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61928,
      "author_name": "saulberardo",
      "author_url": "",
      "post_date": "12/27/2014 08:58:20",
      "content": "<p>Thank your for your answer TDeVries. If it is really the case, emelson is right in diminishing the importance of the LB. It is not of much worth even for helping adjusting hyper-parameters. Indeed, using it for validation can even hurt the final performance!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64525,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "02/18/2015 20:14:33",
      "content": "<p>You should average the error estimates across the CV folds to get an estimate of the testing error. Note, though, that this estimate is almost always a biased estimate of the true testing error.&nbsp; As others have noted, the local estimate seems to be a pretty bad indicator of the leaderboard score in this competition as the LB is only based on two subjects.&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "59743": "",
    "59749": "",
    "59761": "",
    "59764": "",
    "59785": "",
    "61925": "",
    "61926": "",
    "61928": "",
    "64525": ""
  },
  "source": "meta"
}