{
  "id": 9362,
  "title": "Cross validation",
  "url": "/competitions/seizure-detection/discussion/9362",
  "author_name": "",
  "post_date": "2014-06-02T17:27:22.523Z",
  "votes": 1,
  "comment_count": 10,
  "views": 2261,
  "content": "<p>My CV values have tended to be much higher than the scores I'm getting on the public leader board. The difference is on the order of .1 for several different methods. I'm concerned I've introduced a leak with feature engineering. So, before I tear apart my entire pipeline, I thought I'd ask, is anyone else seeing similar discrepancies?</p>",
  "messages": [
    {
      "id": "48503",
      "postDate": "06/02/2014 17:27:22",
      "content": "<p>My CV values have tended to be much higher than the scores I'm getting on the public leader board. The difference is on the order of .1 for several different methods. I'm concerned I've introduced a leak with feature engineering. So, before I tear apart my entire pipeline, I thought I'd ask, is anyone else seeing similar discrepancies?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48531",
      "postDate": "06/02/2014 23:30:47",
      "content": "<p>I also have trouble to set up a good CV chain. IMO, you should not randomize the trials before splitting the dataset. The best thing to do is to use different seizures for test and training. idem for the interictal segments.</p>\n<p>By doing this, i get a more representative score, but it is far from perfect.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48553",
      "postDate": "06/03/2014 10:44:56",
      "content": "<p>Andrew, I'm also observing a big discrepancy, of about .08.</p>\n<p>Alexander, what do you mean by &quot;The best thing to do is to use different seizures for test and training&quot;? That is what cross-validation is about, isn't it?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48554",
      "postDate": "06/03/2014 12:36:30",
      "content": "<p>When i said seizure, i was thinking about the entire seizure, i.e. all the clips from delay 0 to delay XX. For example, the training data of Dog_1 contains 5 seizures.</p>\n<p>Since the classification is on a single trial basis, you can always randomize all your clips before partitioning your data in training and validation set. In this way, your training and validation set include a different sub-sample of each seizure. This approach is valid when the trials are independent from each other.&nbsp;</p>\n<p>This is not the case here, there is a big shift in time, and you have a higher probability to detect something if you have a trial in your training set which is recorded over the same period of time.&nbsp;</p>\n<p>So to avoid this, i detect the beginning and the end of each seizure (based on the delay) and i do my partition according to this. The limitation is that for some subjects, there is only 2 seizures, so you can do nothing more than a 2-fold CV.</p>\n<p>On the other hand, randomizing is not completely useless. It allows you to evaluate the performance of your method if you have a bigger and more representative amount of data, or when there is no shift between training an test.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48556",
      "postDate": "06/03/2014 12:44:55",
      "content": "<p>I got it.</p>\n<p>Yes, it's a very reasonable thing to do. Thanks a lot. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48564",
      "postDate": "06/03/2014 15:18:58",
      "content": "<p>Changing to a leave-one-seizure-out approach to cross validation reduced the discrepancy I was seeing to about .05.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48565",
      "postDate": "06/03/2014 15:59:47",
      "content": "<p>There is probably the same problem for the interictal clips. But i don't have any&nbsp;clean&nbsp;solution. Still, the best thing to do is to avoid randomization.</p>\n<p>The second problem is to know how the public/private LB split is done. The&nbsp;15% could represent test data for two subjects. In this case, the LB score is not representative of the global AUC.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48574",
      "postDate": "06/03/2014 21:05:45",
      "content": "<p>I am very much sure the public LB would not represent data for two/three subjects as that can easily be hacked. I should not go into detail of how that can be done as it is trivial. For the reason that aforementioned split&nbsp; is easily hackable, the public/private split should be randomized and equally split between all subjects!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48581",
      "postDate": "06/04/2014 02:37:59",
      "content": "<p>Hi&nbsp;Alexandre,</p>\n<p>You are very good at data mining, especially time series data, not only in this competition, but also&nbsp;in&nbsp;DecMeg2014 - Decoding the Human Brain. Can I ask you a question? As you said, there is a big time shift in time, do you use a transfer learning method? &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48765",
      "postDate": "06/06/2014 10:29:59",
      "content": "<p>[quote=Alexandre;48554]</p>\n<p>When i said seizure, i was thinking about the entire seizure, i.e. all the clips from delay 0 to delay XX. For example, the training data of Dog_1 contains 5 seizures.</p>\n<p>Since the classification is on a single trial basis, you can always randomize all your clips before partitioning your data in training and validation set. In this way, your training and validation set include a different sub-sample of each seizure. This approach is valid when the trials are independent from each other.&nbsp;</p>\n<p>This is not the case here, there is a big shift in time, and you have a higher probability to detect something if you have a trial in your training set which is recorded over the same period of time.&nbsp;</p>\n<p>So to avoid this, i detect the beginning and the end of each seizure (based on the delay) and i do my partition according to this. The limitation is that for some subjects, there is only 2 seizures, so you can do nothing more than a 2-fold CV.</p>\n<p>On the other hand, randomizing is not completely useless. It allows you to evaluate the performance of your method if you have a bigger and more representative amount of data, or when there is no shift between training an test.</p>\n<p>[/quote]</p>\n\n<p>So are you saying that doing your folds using this leave-one-seizure-out approach actually improves your leaderboard score or that it makes your validation score more representative of what your leaderboard score will be?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48768",
      "postDate": "06/06/2014 11:03:11",
      "content": "<p>it makes my validation score more representative of what my leaderboard score will be.</p>\n<p>For now, i don't have any parameters to tune, so my CV has no impact on my LB score.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 48531,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/02/2014 23:30:47",
      "content": "<p>I also have trouble to set up a good CV chain. IMO, you should not randomize the trials before splitting the dataset. The best thing to do is to use different seizures for test and training. idem for the interictal segments.</p>\n<p>By doing this, i get a more representative score, but it is far from perfect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48553,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "06/03/2014 10:44:56",
      "content": "<p>Andrew, I'm also observing a big discrepancy, of about .08.</p>\n<p>Alexander, what do you mean by &quot;The best thing to do is to use different seizures for test and training&quot;? That is what cross-validation is about, isn't it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48554,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/03/2014 12:36:30",
      "content": "<p>When i said seizure, i was thinking about the entire seizure, i.e. all the clips from delay 0 to delay XX. For example, the training data of Dog_1 contains 5 seizures.</p>\n<p>Since the classification is on a single trial basis, you can always randomize all your clips before partitioning your data in training and validation set. In this way, your training and validation set include a different sub-sample of each seizure. This approach is valid when the trials are independent from each other.&nbsp;</p>\n<p>This is not the case here, there is a big shift in time, and you have a higher probability to detect something if you have a trial in your training set which is recorded over the same period of time.&nbsp;</p>\n<p>So to avoid this, i detect the beginning and the end of each seizure (based on the delay) and i do my partition according to this. The limitation is that for some subjects, there is only 2 seizures, so you can do nothing more than a 2-fold CV.</p>\n<p>On the other hand, randomizing is not completely useless. It allows you to evaluate the performance of your method if you have a bigger and more representative amount of data, or when there is no shift between training an test.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48556,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "06/03/2014 12:44:55",
      "content": "<p>I got it.</p>\n<p>Yes, it's a very reasonable thing to do. Thanks a lot. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48564,
      "author_name": "andrewmatteson",
      "author_url": "",
      "post_date": "06/03/2014 15:18:58",
      "content": "<p>Changing to a leave-one-seizure-out approach to cross validation reduced the discrepancy I was seeing to about .05.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48565,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/03/2014 15:59:47",
      "content": "<p>There is probably the same problem for the interictal clips. But i don't have any&nbsp;clean&nbsp;solution. Still, the best thing to do is to avoid randomization.</p>\n<p>The second problem is to know how the public/private LB split is done. The&nbsp;15% could represent test data for two subjects. In this case, the LB score is not representative of the global AUC.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48574,
      "author_name": "chaotic",
      "author_url": "",
      "post_date": "06/03/2014 21:05:45",
      "content": "<p>I am very much sure the public LB would not represent data for two/three subjects as that can easily be hacked. I should not go into detail of how that can be done as it is trivial. For the reason that aforementioned split&nbsp; is easily hackable, the public/private split should be randomized and equally split between all subjects!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48581,
      "author_name": "hongweizhang",
      "author_url": "",
      "post_date": "06/04/2014 02:37:59",
      "content": "<p>Hi&nbsp;Alexandre,</p>\n<p>You are very good at data mining, especially time series data, not only in this competition, but also&nbsp;in&nbsp;DecMeg2014 - Decoding the Human Brain. Can I ask you a question? As you said, there is a big time shift in time, do you use a transfer learning method? &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48765,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "06/06/2014 10:29:59",
      "content": "<p>[quote=Alexandre;48554]</p>\n<p>When i said seizure, i was thinking about the entire seizure, i.e. all the clips from delay 0 to delay XX. For example, the training data of Dog_1 contains 5 seizures.</p>\n<p>Since the classification is on a single trial basis, you can always randomize all your clips before partitioning your data in training and validation set. In this way, your training and validation set include a different sub-sample of each seizure. This approach is valid when the trials are independent from each other.&nbsp;</p>\n<p>This is not the case here, there is a big shift in time, and you have a higher probability to detect something if you have a trial in your training set which is recorded over the same period of time.&nbsp;</p>\n<p>So to avoid this, i detect the beginning and the end of each seizure (based on the delay) and i do my partition according to this. The limitation is that for some subjects, there is only 2 seizures, so you can do nothing more than a 2-fold CV.</p>\n<p>On the other hand, randomizing is not completely useless. It allows you to evaluate the performance of your method if you have a bigger and more representative amount of data, or when there is no shift between training an test.</p>\n<p>[/quote]</p>\n\n<p>So are you saying that doing your folds using this leave-one-seizure-out approach actually improves your leaderboard score or that it makes your validation score more representative of what your leaderboard score will be?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48768,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/06/2014 11:03:11",
      "content": "<p>it makes my validation score more representative of what my leaderboard score will be.</p>\n<p>For now, i don't have any parameters to tune, so my CV has no impact on my LB score.&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "48503": "",
    "48531": "",
    "48553": "",
    "48554": "",
    "48556": "",
    "48564": "",
    "48565": "",
    "48574": "",
    "48581": "",
    "48765": "",
    "48768": ""
  },
  "source": "meta"
}