{
  "id": 475558,
  "title": "How Essential is Cross Validation?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/475558",
  "author_name": "",
  "post_date": "2024-02-08T22:07:22.988725500Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have a lot of experience with ML/DL but pretty new to taking Kaggle seriously. I noticed there is a lot of talk about CV and LB (cross validation and leaderboard) scores, but didn't really think much of it on my first attempt. I just split 10% of the training data to make a validation set and went from there.</p>\n<p>I was getting very good validation scores locally, so I was surprised when I uploaded the model and found I got a leaderboard score of <strong>0.97</strong> (oof) – the same as just predicting the mean of the labels.</p>\n<p>This indicates that either:</p>\n<ul>\n<li>I have some data leakage between the train and validation split (but unlikely, as validation scores are never as good as training)</li>\n<li>the test set is very different to the train set (very likely).</li>\n</ul>\n<p>Initially, I thought the main purpose of CV was to give some estimate of out of distribution performance when we don't have an actual test / gold set. But after looking at other people's notebooks, it seems they can also be used to form an ensemble of models (one for each fold) which seems to help with out of distribution performance a lot.</p>\n<p>However, training five (or more) models is more annoying than training one! I'd like to avoid this if at all possible.</p>\n<p><strong>Question: Are there any alternative ways to help improve the mismatch between local validation scores and LB scores?</strong></p>",
  "messages": [
    {
      "id": "2643529",
      "postDate": "02/08/2024 22:07:22",
      "content": "<p>I have a lot of experience with ML/DL but pretty new to taking Kaggle seriously. I noticed there is a lot of talk about CV and LB (cross validation and leaderboard) scores, but didn't really think much of it on my first attempt. I just split 10% of the training data to make a validation set and went from there.</p>\n<p>I was getting very good validation scores locally, so I was surprised when I uploaded the model and found I got a leaderboard score of <strong>0.97</strong> (oof) – the same as just predicting the mean of the labels.</p>\n<p>This indicates that either:</p>\n<ul>\n<li>I have some data leakage between the train and validation split (but unlikely, as validation scores are never as good as training)</li>\n<li>the test set is very different to the train set (very likely).</li>\n</ul>\n<p>Initially, I thought the main purpose of CV was to give some estimate of out of distribution performance when we don't have an actual test / gold set. But after looking at other people's notebooks, it seems they can also be used to form an ensemble of models (one for each fold) which seems to help with out of distribution performance a lot.</p>\n<p>However, training five (or more) models is more annoying than training one! I'd like to avoid this if at all possible.</p>\n<p><strong>Question: Are there any alternative ways to help improve the mismatch between local validation scores and LB scores?</strong></p>",
      "rawMarkdown": "I have a lot of experience with ML/DL but pretty new to taking Kaggle seriously. I noticed there is a lot of talk about CV and LB (cross validation and leaderboard) scores, but didn't really think much of it on my first attempt. I just split 10% of the training data to make a validation set and went from there.\n\nI was getting very good validation scores locally, so I was surprised when I uploaded the model and found I got a leaderboard score of **0.97** (oof) – the same as just predicting the mean of the labels.\n\nThis indicates that either:\n- I have some data leakage between the train and validation split (but unlikely, as validation scores are never as good as training)\n- the test set is very different to the train set (very likely).\n\nInitially, I thought the main purpose of CV was to give some estimate of out of distribution performance when we don't have an actual test / gold set. But after looking at other people's notebooks, it seems they can also be used to form an ensemble of models (one for each fold) which seems to help with out of distribution performance a lot.\n\nHowever, training five (or more) models is more annoying than training one! I'd like to avoid this if at all possible.\n\n**Question: Are there any alternative ways to help improve the mismatch between local validation scores and LB scores?**",
      "votes": null
    },
    {
      "id": "2645067",
      "postDate": "02/09/2024 23:03:34",
      "content": "<p>In my cross validated models, I have observed that there is data leakage across patient_id. If I just split the train set without paying attention at patient_id (so the same id can be in the validation set and in the train set), I get better CV performances than if I split the train and validation sets by patient_id (ensuring that the same id is not both in the train and validation set).<br>\nIn the test set, patient_id will likely differ from those in the train set, so this could be an explanation for the worse LB score.</p>",
      "rawMarkdown": "In my cross validated models, I have observed that there is data leakage across patient_id. If I just split the train set without paying attention at patient_id (so the same id can be in the validation set and in the train set), I get better CV performances than if I split the train and validation sets by patient_id (ensuring that the same id is not both in the train and validation set).\nIn the test set, patient_id will likely differ from those in the train set, so this could be an explanation for the worse LB score.",
      "votes": null
    },
    {
      "id": "2655694",
      "postDate": "02/17/2024 06:37:57",
      "content": "<p>You cant get away without a proper validation set.<br>\nOtherwise you will not know when to stop training: and most probably you will overfit. <br>\nFor this the validation loss does not have to 'exactly' match the training loss: Asssuming your validation set includes leaked and nonleaked data: you will get minimal loss for the leaked validation data, mimicking the training loss. And higher then normal contribution from non leaked validation data. I say higher, because you are overfitting to the leaked portion of the training data: The validation loss  (combination) will not be as good as training loss, but you are still overfitting to training data. </p>\n<p>On the other hand, you dont have to do a fullblown crossvalidation as long as you choose a good validation split. <br>\nIf you have sufficient data, you can train 2 folds of a 10%crossvalidation scheme.<br>\nIn other words, you just use boilerplate code for a 10fold grouped and maybe also stratified split of the data to train and validation. Then you just pick 2 of them to train. <br>\nThis approach of course can be tailored very much: 2 of 6fold, or maybe 3 of 5fold, if you have more resources. <br>\nThis does not directly address your final question, which is more complex. Maybe 'public test set' has a closer-alike distribution as the train, why not?</p>",
      "rawMarkdown": "You cant get away without a proper validation set.\nOtherwise you will not know when to stop training: and most probably you will overfit. \nFor this the validation loss does not have to 'exactly' match the training loss: Asssuming your validation set includes leaked and nonleaked data: you will get minimal loss for the leaked validation data, mimicking the training loss. And higher then normal contribution from non leaked validation data. I say higher, because you are overfitting to the leaked portion of the training data: The validation loss  (combination) will not be as good as training loss, but you are still overfitting to training data. \n\nOn the other hand, you dont have to do a fullblown crossvalidation as long as you choose a good validation split. \nIf you have sufficient data, you can train 2 folds of a 10%crossvalidation scheme.\nIn other words, you just use boilerplate code for a 10fold grouped and maybe also stratified split of the data to train and validation. Then you just pick 2 of them to train. \nThis approach of course can be tailored very much: 2 of 6fold, or maybe 3 of 5fold, if you have more resources. \nThis does not directly address your final question, which is more complex. Maybe 'public test set' has a closer-alike distribution as the train, why not?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2645067,
      "author_name": "fabienpv",
      "author_url": "",
      "post_date": "02/09/2024 23:03:34",
      "content": "<p>In my cross validated models, I have observed that there is data leakage across patient_id. If I just split the train set without paying attention at patient_id (so the same id can be in the validation set and in the train set), I get better CV performances than if I split the train and validation sets by patient_id (ensuring that the same id is not both in the train and validation set).<br>\nIn the test set, patient_id will likely differ from those in the train set, so this could be an explanation for the worse LB score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2655694,
      "author_name": "abdulkadirguner",
      "author_url": "",
      "post_date": "02/17/2024 06:37:57",
      "content": "<p>You cant get away without a proper validation set.<br>\nOtherwise you will not know when to stop training: and most probably you will overfit. <br>\nFor this the validation loss does not have to 'exactly' match the training loss: Asssuming your validation set includes leaked and nonleaked data: you will get minimal loss for the leaked validation data, mimicking the training loss. And higher then normal contribution from non leaked validation data. I say higher, because you are overfitting to the leaked portion of the training data: The validation loss  (combination) will not be as good as training loss, but you are still overfitting to training data. </p>\n<p>On the other hand, you dont have to do a fullblown crossvalidation as long as you choose a good validation split. <br>\nIf you have sufficient data, you can train 2 folds of a 10%crossvalidation scheme.<br>\nIn other words, you just use boilerplate code for a 10fold grouped and maybe also stratified split of the data to train and validation. Then you just pick 2 of them to train. <br>\nThis approach of course can be tailored very much: 2 of 6fold, or maybe 3 of 5fold, if you have more resources. <br>\nThis does not directly address your final question, which is more complex. Maybe 'public test set' has a closer-alike distribution as the train, why not?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2643529": "I have a lot of experience with ML/DL but pretty new to taking Kaggle seriously. I noticed there is a lot of talk about CV and LB (cross validation and leaderboard) scores, but didn't really think much of it on my first attempt. I just split 10% of the training data to make a validation set and went from there.\n\nI was getting very good validation scores locally, so I was surprised when I uploaded the model and found I got a leaderboard score of **0.97** (oof) – the same as just predicting the mean of the labels.\n\nThis indicates that either:\n- I have some data leakage between the train and validation split (but unlikely, as validation scores are never as good as training)\n- the test set is very different to the train set (very likely).\n\nInitially, I thought the main purpose of CV was to give some estimate of out of distribution performance when we don't have an actual test / gold set. But after looking at other people's notebooks, it seems they can also be used to form an ensemble of models (one for each fold) which seems to help with out of distribution performance a lot.\n\nHowever, training five (or more) models is more annoying than training one! I'd like to avoid this if at all possible.\n\n**Question: Are there any alternative ways to help improve the mismatch between local validation scores and LB scores?**",
    "2645067": "In my cross validated models, I have observed that there is data leakage across patient_id. If I just split the train set without paying attention at patient_id (so the same id can be in the validation set and in the train set), I get better CV performances than if I split the train and validation sets by patient_id (ensuring that the same id is not both in the train and validation set).\nIn the test set, patient_id will likely differ from those in the train set, so this could be an explanation for the worse LB score.",
    "2655694": "You cant get away without a proper validation set.\nOtherwise you will not know when to stop training: and most probably you will overfit. \nFor this the validation loss does not have to 'exactly' match the training loss: Asssuming your validation set includes leaked and nonleaked data: you will get minimal loss for the leaked validation data, mimicking the training loss. And higher then normal contribution from non leaked validation data. I say higher, because you are overfitting to the leaked portion of the training data: The validation loss  (combination) will not be as good as training loss, but you are still overfitting to training data. \n\nOn the other hand, you dont have to do a fullblown crossvalidation as long as you choose a good validation split. \nIf you have sufficient data, you can train 2 folds of a 10%crossvalidation scheme.\nIn other words, you just use boilerplate code for a 10fold grouped and maybe also stratified split of the data to train and validation. Then you just pick 2 of them to train. \nThis approach of course can be tailored very much: 2 of 6fold, or maybe 3 of 5fold, if you have more resources. \nThis does not directly address your final question, which is more complex. Maybe 'public test set' has a closer-alike distribution as the train, why not?"
  },
  "source": "meta"
}