{
  "id": 10477,
  "title": "Gap between cross-validation and public test",
  "url": "/competitions/seizure-prediction/discussion/10477",
  "author_name": "",
  "post_date": "2014-09-29T09:41:59.393Z",
  "votes": 2,
  "comment_count": 17,
  "views": 3849,
  "content": "<p>Dear competitors,</p>\n<p>we are suffering a very large drop between our cross-validation AUC and the AUC obtained in the public test. As example, a model with a cross-validation AUC=0.91&nbsp;obtains an AUC=0.67 in the public test.</p>\n<p>- Our cross-validation uses K folds, where the number K depends in the number of preictal hours of each subject.</p>\n<p>- In this way, our&nbsp;folds&nbsp;are&nbsp;independent, so all sequences related with one hour recording are together in the same fold.</p>\n<p>- We compute AUC over all the subjects response, in the same way as it is done by Kaggle.</p>\n<p>We have improvements in our cross-validations, and values of AUC=0.91, but our improvements are not reflected in the leaderboard, and better cross-validated models achieve worst public test AUC than other models which perform worst in cross-validation results.</p>\n<p>Is anyone else suffereing this kind of problems?</p>\n<p>thanks!</p>",
  "messages": [
    {
      "id": "55382",
      "postDate": "09/29/2014 09:41:59",
      "content": "<p>Dear competitors,</p>\n<p>we are suffering a very large drop between our cross-validation AUC and the AUC obtained in the public test. As example, a model with a cross-validation AUC=0.91&nbsp;obtains an AUC=0.67 in the public test.</p>\n<p>- Our cross-validation uses K folds, where the number K depends in the number of preictal hours of each subject.</p>\n<p>- In this way, our&nbsp;folds&nbsp;are&nbsp;independent, so all sequences related with one hour recording are together in the same fold.</p>\n<p>- We compute AUC over all the subjects response, in the same way as it is done by Kaggle.</p>\n<p>We have improvements in our cross-validations, and values of AUC=0.91, but our improvements are not reflected in the leaderboard, and better cross-validated models achieve worst public test AUC than other models which perform worst in cross-validation results.</p>\n<p>Is anyone else suffereing this kind of problems?</p>\n<p>thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55383",
      "postDate": "09/29/2014 10:15:49",
      "content": "<p>Yep I've had the same problem. Unless you are using a single model, I would recommend computing the AUC per subject and then consider a weighted average with regard to test counts; simply looking at per subject performance is useful. Personally, I'm not putting to much weight into CV estimate because it has not been very consistent with leaderboard.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55384",
      "postDate": "09/29/2014 10:36:02",
      "content": "<p>Thanks Brian, I will try to compute a weighted average of AUC, but as you say, CV estimate is not consistent as expected :S</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55486",
      "postDate": "10/02/2014 10:01:35",
      "content": "<p>The fact that you have such a gap between the validation and the test set migth indicate you are overtraining you machine. Try to reduce the number of training entries, and use a more permisive cross-validation schema.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55487",
      "postDate": "10/02/2014 10:13:25",
      "content": "<p>So, acaicedo, Are you suggesting to mix different seizure sequences between cross-validation folds? I have the intuition that my first two items are important to avoid overfitting, because it is a way to reproduce an internal test set similar to Kaggle ones.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55488",
      "postDate": "10/02/2014 10:29:15",
      "content": "<p>In general one can say that the AUC is larger in the training set than in the test set.To some extend what you do avoids overfitting. However if the k valieue is too large that can lead to over trained classifiers. For example if you are using a 10-fold cross-validation, try using a 5-fold crossvalidation. In my experience this leads to lower AUC during training but larger AUC during testing which diminishes the gap. Another thing to take into account is how balanced are your training, and test set, do you preserve the ratios during training or you adjust them? How do you divide the training dat ain k-folds, do you preserve the ratios there also or you correct?</p>\n<p>I hope this migth be helpfull.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55489",
      "postDate": "10/02/2014 10:35:42",
      "content": "<p>Francisco I am using exactly the same scheme as you for CV with AUCs weighted by the test sample size and my CV score is better than LB and not very indicative but not so significantly. I would suspect there is some leakage somewhere. i.e you are doing something different in preictal vs interictal segments when pocessing/extract features ? &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55490",
      "postDate": "10/02/2014 10:38:22",
      "content": "<p>Acaicedo, thanks for your valuable help. I don't have so much experience with cross-validation, in my previous research I have enough data to work using holdout validation and test sets. I will try a modification of my CV procedure,&nbsp;with a maximum of folds. In my current implementation each Dog and each Patient has its own number of folds, which depends in the number of sizeure hours available. In this way, Dog_1 has 4 folds, Dog_2 has 7 folds, Dog_3 has 12 or 13 folds, ... So, I can try to put a maximum number of folds (say 5 or 6) and see what happens ;-)</p>\n<p>Training and test sets are not balanced equally. There are subjects which has more samples in training than in test, and&nbsp;subjects which has more samples in test than in training. What I'm doing is normalizing the AUC&nbsp;by the number of samples of the subject in test set.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55491",
      "postDate": "10/02/2014 10:41:56",
      "content": "<p>Tsakalis Kostas, I'm pre-processing exactly in the same way preictal and interictal segments, and now I'm normalizing AUC using the test sample size of the subject. In any case, thanks for your comment. I'm will review (again jejeje) the code for bugs.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55526",
      "postDate": "10/03/2014 04:12:27",
      "content": "<p>I'm getting the same effect and I am reasonably convinced it's not overfitting. &nbsp;E.g. If I split Dog 5 into two sets, equal size, 225 normal / 15 preictal, and train on one and test on the other, I get AUC = 0.8 or so. &nbsp;So no folds here: completely independent train /test sets. &nbsp;But these do no better than random on the LB. Any thoughts?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55530",
      "postDate": "10/03/2014 08:31:51",
      "content": "<p>I have tried to reduce the number of folds,&nbsp;and the same problem as Jonathan, the AUC stills being very optimistic in CV, compared with LB result.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55535",
      "postDate": "10/03/2014 09:34:35",
      "content": "<p>I'm getting some improvements when I take care to align (&quot;calibrate&quot;) classifiers from different subjects. &nbsp;It is a bit disappointing - seems like the most important issue here might be how you normalize your results between subjects, which isn't really the point of the competition.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55544",
      "postDate": "10/03/2014 12:53:59",
      "content": "<p>[quote=Jonathan Tapson;55535]</p>\n<p>I'm getting some improvements when I take care to align (&quot;calibrate&quot;) classifiers from different subjects. &nbsp;It is a bit disappointing - seems like the most important issue here might be how you normalize your results between subjects, which isn't really the point of the competition.</p>\n<p>[/quote]</p>\n<p>Hi Jonathan, care to share any tips on how to do this alignment? I've tried a couple things which I thought might help, but didn't see any real improvement.</p>\n<p>I'm not sure, but in your earlier post it sounded like you might be choosing segments for your train/test split completely randomly - I think it is very important to maintain sequence grouping for CV, as it is quite&nbsp;easy to recognize segments from within the same sequence.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55570",
      "postDate": "10/03/2014 19:40:40",
      "content": "<p>Hey, then I guess the main issue is the selection of the training and the validation set. Selecting the training and test randomly is not the best option. That can be something to try out.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55586",
      "postDate": "10/04/2014 04:06:11",
      "content": "<p>[quote=emolson;55544]</p>\n<p>[quote=Jonathan Tapson;55535]</p>\n<p>I'm getting some improvements when I take care to align (&quot;calibrate&quot;) classifiers from different subjects. &nbsp;It is a bit disappointing - seems like the most important issue here might be how you normalize your results between subjects, which isn't really the point of the competition.</p>\n<p>[/quote]</p>\n<p>Hi Jonathan, care to share any tips on how to do this alignment? I've tried a couple things which I thought might help, but didn't see any real improvement.</p>\n<p>I'm not sure, but in your earlier post it sounded like you might be choosing segments for your train/test split completely randomly - I think it is very important to maintain sequence grouping for CV, as it is quite&nbsp;easy to recognize segments from within the same sequence.</p>\n<p>[/quote]</p>\n<p>Hi Eben</p>\n<p>Thanks - I split the data into two by taking first half of each type for training,&nbsp;and second half for testing, so the sequences were I think still mostly grouped. &nbsp;But I will do this more carefully from now on.</p>\n<p>For alignment - I am trying standard normalization techniques like (data - mean(data))/sdev(data). &nbsp;I have also tried to fix each subject's classification threshold to 0.5 and normalized the distribution around that by / sdev(data). &nbsp;All these improve the performance a bit, but not too convincingly. &nbsp;Would be grateful to hear some other ideas on this.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55603",
      "postDate": "10/04/2014 15:29:24",
      "content": "<p>I have tried to standarize test probabilities following a similar approach (computing mean and sdev in validation for positive and negative samples). However, a didn't observe any improvement... I'm not sure this is a good way to solve the problems. I think the best is to find a good model which has the ability to learn consistent probabilities across subjects, but obviously that is not easy... :(</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55709",
      "postDate": "10/07/2014 00:46:06",
      "content": "<p>I think if you notice a &quot;batch&quot; effect in your scoring (probability) then this may be saying that the model training feature signal-to-noise is too low or new data is too perturbed, and subsequently, the predictions are trending with the original sample mean &nbsp;or proportion of preictal to interictal. I've noticed this effect in some model attempts where the predicted values tend to track with a&nbsp;null model (only intercept). I think to test this you would compute the median predicted value per subject (new predictions) and plot against the training data's sample mean (preictal to interictal). Any correlation? I'd check but must sleep...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55750",
      "postDate": "10/07/2014 21:06:47",
      "content": "<p>It can be the case, when the model is logistic regression or similar, but I'm also trying KNNs and them suffer the same gap... I'm just simplifying&nbsp;the code looking for any bug... :S</p>\n<p>Anyway, thanks for your suggestions, it is&nbsp;a very good idea to check the correct convergence&nbsp;of the logistic regression models.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 55383,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "09/29/2014 10:15:49",
      "content": "<p>Yep I've had the same problem. Unless you are using a single model, I would recommend computing the AUC per subject and then consider a weighted average with regard to test counts; simply looking at per subject performance is useful. Personally, I'm not putting to much weight into CV estimate because it has not been very consistent with leaderboard.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55384,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "09/29/2014 10:36:02",
      "content": "<p>Thanks Brian, I will try to compute a weighted average of AUC, but as you say, CV estimate is not consistent as expected :S</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55486,
      "author_name": "acaicedo",
      "author_url": "",
      "post_date": "10/02/2014 10:01:35",
      "content": "<p>The fact that you have such a gap between the validation and the test set migth indicate you are overtraining you machine. Try to reduce the number of training entries, and use a more permisive cross-validation schema.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55487,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "10/02/2014 10:13:25",
      "content": "<p>So, acaicedo, Are you suggesting to mix different seizure sequences between cross-validation folds? I have the intuition that my first two items are important to avoid overfitting, because it is a way to reproduce an internal test set similar to Kaggle ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55488,
      "author_name": "acaicedo",
      "author_url": "",
      "post_date": "10/02/2014 10:29:15",
      "content": "<p>In general one can say that the AUC is larger in the training set than in the test set.To some extend what you do avoids overfitting. However if the k valieue is too large that can lead to over trained classifiers. For example if you are using a 10-fold cross-validation, try using a 5-fold crossvalidation. In my experience this leads to lower AUC during training but larger AUC during testing which diminishes the gap. Another thing to take into account is how balanced are your training, and test set, do you preserve the ratios during training or you adjust them? How do you divide the training dat ain k-folds, do you preserve the ratios there also or you correct?</p>\n<p>I hope this migth be helpfull.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55489,
      "author_name": "epinephelus",
      "author_url": "",
      "post_date": "10/02/2014 10:35:42",
      "content": "<p>Francisco I am using exactly the same scheme as you for CV with AUCs weighted by the test sample size and my CV score is better than LB and not very indicative but not so significantly. I would suspect there is some leakage somewhere. i.e you are doing something different in preictal vs interictal segments when pocessing/extract features ? &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55490,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "10/02/2014 10:38:22",
      "content": "<p>Acaicedo, thanks for your valuable help. I don't have so much experience with cross-validation, in my previous research I have enough data to work using holdout validation and test sets. I will try a modification of my CV procedure,&nbsp;with a maximum of folds. In my current implementation each Dog and each Patient has its own number of folds, which depends in the number of sizeure hours available. In this way, Dog_1 has 4 folds, Dog_2 has 7 folds, Dog_3 has 12 or 13 folds, ... So, I can try to put a maximum number of folds (say 5 or 6) and see what happens ;-)</p>\n<p>Training and test sets are not balanced equally. There are subjects which has more samples in training than in test, and&nbsp;subjects which has more samples in test than in training. What I'm doing is normalizing the AUC&nbsp;by the number of samples of the subject in test set.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55491,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "10/02/2014 10:41:56",
      "content": "<p>Tsakalis Kostas, I'm pre-processing exactly in the same way preictal and interictal segments, and now I'm normalizing AUC using the test sample size of the subject. In any case, thanks for your comment. I'm will review (again jejeje) the code for bugs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55526,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "10/03/2014 04:12:27",
      "content": "<p>I'm getting the same effect and I am reasonably convinced it's not overfitting. &nbsp;E.g. If I split Dog 5 into two sets, equal size, 225 normal / 15 preictal, and train on one and test on the other, I get AUC = 0.8 or so. &nbsp;So no folds here: completely independent train /test sets. &nbsp;But these do no better than random on the LB. Any thoughts?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55530,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "10/03/2014 08:31:51",
      "content": "<p>I have tried to reduce the number of folds,&nbsp;and the same problem as Jonathan, the AUC stills being very optimistic in CV, compared with LB result.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55535,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "10/03/2014 09:34:35",
      "content": "<p>I'm getting some improvements when I take care to align (&quot;calibrate&quot;) classifiers from different subjects. &nbsp;It is a bit disappointing - seems like the most important issue here might be how you normalize your results between subjects, which isn't really the point of the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55544,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "10/03/2014 12:53:59",
      "content": "<p>[quote=Jonathan Tapson;55535]</p>\n<p>I'm getting some improvements when I take care to align (&quot;calibrate&quot;) classifiers from different subjects. &nbsp;It is a bit disappointing - seems like the most important issue here might be how you normalize your results between subjects, which isn't really the point of the competition.</p>\n<p>[/quote]</p>\n<p>Hi Jonathan, care to share any tips on how to do this alignment? I've tried a couple things which I thought might help, but didn't see any real improvement.</p>\n<p>I'm not sure, but in your earlier post it sounded like you might be choosing segments for your train/test split completely randomly - I think it is very important to maintain sequence grouping for CV, as it is quite&nbsp;easy to recognize segments from within the same sequence.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55570,
      "author_name": "acaicedo",
      "author_url": "",
      "post_date": "10/03/2014 19:40:40",
      "content": "<p>Hey, then I guess the main issue is the selection of the training and the validation set. Selecting the training and test randomly is not the best option. That can be something to try out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55586,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "10/04/2014 04:06:11",
      "content": "<p>[quote=emolson;55544]</p>\n<p>[quote=Jonathan Tapson;55535]</p>\n<p>I'm getting some improvements when I take care to align (&quot;calibrate&quot;) classifiers from different subjects. &nbsp;It is a bit disappointing - seems like the most important issue here might be how you normalize your results between subjects, which isn't really the point of the competition.</p>\n<p>[/quote]</p>\n<p>Hi Jonathan, care to share any tips on how to do this alignment? I've tried a couple things which I thought might help, but didn't see any real improvement.</p>\n<p>I'm not sure, but in your earlier post it sounded like you might be choosing segments for your train/test split completely randomly - I think it is very important to maintain sequence grouping for CV, as it is quite&nbsp;easy to recognize segments from within the same sequence.</p>\n<p>[/quote]</p>\n<p>Hi Eben</p>\n<p>Thanks - I split the data into two by taking first half of each type for training,&nbsp;and second half for testing, so the sequences were I think still mostly grouped. &nbsp;But I will do this more carefully from now on.</p>\n<p>For alignment - I am trying standard normalization techniques like (data - mean(data))/sdev(data). &nbsp;I have also tried to fix each subject's classification threshold to 0.5 and normalized the distribution around that by / sdev(data). &nbsp;All these improve the performance a bit, but not too convincingly. &nbsp;Would be grateful to hear some other ideas on this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55603,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "10/04/2014 15:29:24",
      "content": "<p>I have tried to standarize test probabilities following a similar approach (computing mean and sdev in validation for positive and negative samples). However, a didn't observe any improvement... I'm not sure this is a good way to solve the problems. I think the best is to find a good model which has the ability to learn consistent probabilities across subjects, but obviously that is not easy... :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55709,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "10/07/2014 00:46:06",
      "content": "<p>I think if you notice a &quot;batch&quot; effect in your scoring (probability) then this may be saying that the model training feature signal-to-noise is too low or new data is too perturbed, and subsequently, the predictions are trending with the original sample mean &nbsp;or proportion of preictal to interictal. I've noticed this effect in some model attempts where the predicted values tend to track with a&nbsp;null model (only intercept). I think to test this you would compute the median predicted value per subject (new predictions) and plot against the training data's sample mean (preictal to interictal). Any correlation? I'd check but must sleep...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55750,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "10/07/2014 21:06:47",
      "content": "<p>It can be the case, when the model is logistic regression or similar, but I'm also trying KNNs and them suffer the same gap... I'm just simplifying&nbsp;the code looking for any bug... :S</p>\n<p>Anyway, thanks for your suggestions, it is&nbsp;a very good idea to check the correct convergence&nbsp;of the logistic regression models.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "55382": "",
    "55383": "",
    "55384": "",
    "55486": "",
    "55487": "",
    "55488": "",
    "55489": "",
    "55490": "",
    "55491": "",
    "55526": "",
    "55530": "",
    "55535": "",
    "55544": "",
    "55570": "",
    "55586": "",
    "55603": "",
    "55709": "",
    "55750": ""
  },
  "source": "meta"
}