{
  "id": 209136,
  "title": "Cross validation for beginners",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/209136",
  "author_name": "",
  "post_date": "2021-01-06T12:23:34.546945300Z",
  "votes": 17,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi! I try to summarize my knowledge about cross-validation (CV). It would be nice if a more experienced kaggler can correct me out if I am wrong, or add some information if I miss it. I hope it can help others. I assume that the data is shuffled already. Cross-validation is more representative of the private leader board, then the public leader board score.</p>\n<h1>Hold out - Single fold</h1>\n<p>The basic approach is the holdout, where we subdivide the data into a <code>train</code> set and <code>test</code> set. We train the neural network on the <code>train</code> set and, measure the accuracy on the <code>test</code> set. This is when in <code>k = 1</code>. Code for it [1]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2F271f42e5ebce40fa6d6e4028369d4e32%2Fhold_out.jpg?generation=1609935747356886&amp;alt=media\" alt=\"\"></p>\n<h1>K-fold cross-validation</h1>\n<p>Cross-validation means, that we create k type of subdivision between on the data. We train on each fold one model, and in the end, we calculate the average. Code for this [2]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2Fa7eaffa417f9905be8f0e22af7326ac0%2Fk-fold.jpg?generation=1609935772242624&amp;alt=media\" alt=\"\"></p>\n<p>Additional materials:</p>\n<ul>\n<li>Brief intro: <a href=\"https://youtu.be/fSytzGwwBVw?t=26\" target=\"_blank\">https://youtu.be/fSytzGwwBVw?t=26</a></li>\n<li>Selecting the right k number: <a href=\"https://www.youtube.com/watch?v=qOwT553oMzs\" target=\"_blank\">https://www.youtube.com/watch?v=qOwT553oMzs</a></li>\n</ul>\n<p>[1] <a href=\"https://stackoverflow.com/questions/60883696/k-fold-cross-validation-using-dataloaders-in-pytorch\" target=\"_blank\">https://stackoverflow.com/questions/60883696/k-fold-cross-validation-using-dataloaders-in-pytorch</a></p>\n<pre><code>train_size = int(0.8 * len(full_dataset))\nvalidation_size = len(full_dataset) - train_size\ntrain_dataset, validation_dataset = random_split(full_dataset, [train_size, validation_size])\n\nfull_loader = DataLoader(full_dataset, batch_size=4,sampler = sampler_(full_dataset), pin_memory=True) \ntrain_loader = DataLoader(train_dataset, batch_size=4, sampler = sampler_(train_dataset))\nval_loader = DataLoader(validation_dataset, batch_size=1, sampler = sampler_(validation_dataset))\n</code></pre>\n<p>[2] Taken from: <a href=\"https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs\" target=\"_blank\">https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs</a></p>\n<pre><code>folds = train.copy()\nFold = StratifiedKFold(n_splits=CFG.n_fold, shuffle=True, random_state=CFG.seed)\nfor n, (train_index, val_index) in enumerate(Fold.split(folds, folds[CFG.target_col])):\n    folds.loc[val_index, 'fold'] = int(n)\nfolds['fold'] = folds['fold'].astype(int)\n</code></pre>",
  "messages": [
    {
      "id": "1141007",
      "postDate": "01/06/2021 12:23:34",
      "content": "<p>Hi! I try to summarize my knowledge about cross-validation (CV). It would be nice if a more experienced kaggler can correct me out if I am wrong, or add some information if I miss it. I hope it can help others. I assume that the data is shuffled already. Cross-validation is more representative of the private leader board, then the public leader board score.</p>\n<h1>Hold out - Single fold</h1>\n<p>The basic approach is the holdout, where we subdivide the data into a <code>train</code> set and <code>test</code> set. We train the neural network on the <code>train</code> set and, measure the accuracy on the <code>test</code> set. This is when in <code>k = 1</code>. Code for it [1]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2F271f42e5ebce40fa6d6e4028369d4e32%2Fhold_out.jpg?generation=1609935747356886&amp;alt=media\" alt=\"\"></p>\n<h1>K-fold cross-validation</h1>\n<p>Cross-validation means, that we create k type of subdivision between on the data. We train on each fold one model, and in the end, we calculate the average. Code for this [2]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2Fa7eaffa417f9905be8f0e22af7326ac0%2Fk-fold.jpg?generation=1609935772242624&amp;alt=media\" alt=\"\"></p>\n<p>Additional materials:</p>\n<ul>\n<li>Brief intro: <a href=\"https://youtu.be/fSytzGwwBVw?t=26\" target=\"_blank\">https://youtu.be/fSytzGwwBVw?t=26</a></li>\n<li>Selecting the right k number: <a href=\"https://www.youtube.com/watch?v=qOwT553oMzs\" target=\"_blank\">https://www.youtube.com/watch?v=qOwT553oMzs</a></li>\n</ul>\n<p>[1] <a href=\"https://stackoverflow.com/questions/60883696/k-fold-cross-validation-using-dataloaders-in-pytorch\" target=\"_blank\">https://stackoverflow.com/questions/60883696/k-fold-cross-validation-using-dataloaders-in-pytorch</a></p>\n<pre><code>train_size = int(0.8 * len(full_dataset))\nvalidation_size = len(full_dataset) - train_size\ntrain_dataset, validation_dataset = random_split(full_dataset, [train_size, validation_size])\n\nfull_loader = DataLoader(full_dataset, batch_size=4,sampler = sampler_(full_dataset), pin_memory=True) \ntrain_loader = DataLoader(train_dataset, batch_size=4, sampler = sampler_(train_dataset))\nval_loader = DataLoader(validation_dataset, batch_size=1, sampler = sampler_(validation_dataset))\n</code></pre>\n<p>[2] Taken from: <a href=\"https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs\" target=\"_blank\">https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs</a></p>\n<pre><code>folds = train.copy()\nFold = StratifiedKFold(n_splits=CFG.n_fold, shuffle=True, random_state=CFG.seed)\nfor n, (train_index, val_index) in enumerate(Fold.split(folds, folds[CFG.target_col])):\n    folds.loc[val_index, 'fold'] = int(n)\nfolds['fold'] = folds['fold'].astype(int)\n</code></pre>",
      "rawMarkdown": "Hi! I try to summarize my knowledge about cross-validation (CV). It would be nice if a more experienced kaggler can correct me out if I am wrong, or add some information if I miss it. I hope it can help others. I assume that the data is shuffled already. Cross-validation is more representative of the private leader board, then the public leader board score.\n\n# Hold out - Single fold\nThe basic approach is the holdout, where we subdivide the data into a `train` set and `test` set. We train the neural network on the `train` set and, measure the accuracy on the `test` set. This is when in `k = 1`. Code for it [1]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2F271f42e5ebce40fa6d6e4028369d4e32%2Fhold_out.jpg?generation=1609935747356886&alt=media)\n\n# K-fold cross-validation\nCross-validation means, that we create k type of subdivision between on the data. We train on each fold one model, and in the end, we calculate the average. Code for this [2]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2Fa7eaffa417f9905be8f0e22af7326ac0%2Fk-fold.jpg?generation=1609935772242624&alt=media)\n\nAdditional materials:\n\n* Brief intro: https://youtu.be/fSytzGwwBVw?t=26\n* Selecting the right k number: https://www.youtube.com/watch?v=qOwT553oMzs\n\n \n\n[1] https://stackoverflow.com/questions/60883696/k-fold-cross-validation-using-dataloaders-in-pytorch\n\n```\ntrain_size = int(0.8 * len(full_dataset))\nvalidation_size = len(full_dataset) - train_size\ntrain_dataset, validation_dataset = random_split(full_dataset, [train_size, validation_size])\n\nfull_loader = DataLoader(full_dataset, batch_size=4,sampler = sampler_(full_dataset), pin_memory=True) \ntrain_loader = DataLoader(train_dataset, batch_size=4, sampler = sampler_(train_dataset))\nval_loader = DataLoader(validation_dataset, batch_size=1, sampler = sampler_(validation_dataset))\n```\n\n[2] Taken from: https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs\n\n```\nfolds = train.copy()\nFold = StratifiedKFold(n_splits=CFG.n_fold, shuffle=True, random_state=CFG.seed)\nfor n, (train_index, val_index) in enumerate(Fold.split(folds, folds[CFG.target_col])):\n    folds.loc[val_index, 'fold'] = int(n)\nfolds['fold'] = folds['fold'].astype(int)\n```",
      "votes": null
    },
    {
      "id": "1144222",
      "postDate": "01/08/2021 10:25:48",
      "content": "<p><a href=\"https://www.kaggle.com/bessenyeiszilrd\" target=\"_blank\">@bessenyeiszilrd</a> There are 2 more I can add to this list: <br>\nStratified k-fold cross-validation<br>\nAdversarial validation</p>",
      "rawMarkdown": "bessenyeiszilrd There are 2 more I can add to this list: \nStratified k-fold cross-validation\nAdversarial validation",
      "votes": null
    },
    {
      "id": "1156997",
      "postDate": "01/17/2021 15:21:36",
      "content": "<p><a href=\"https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e\" target=\"_blank\">https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e</a></p>\n<p>Hope this helps you. You can find some other folding techniques discussed here. </p>",
      "rawMarkdown": "https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e\n\nHope this helps you. You can find some other folding techniques discussed here.",
      "votes": null
    },
    {
      "id": "1157063",
      "postDate": "01/17/2021 16:11:25",
      "content": "<p>Thanks a lot for posting this! One thing that has been a bit unclear to me is when to use the \"hold out\" approach and when to use the \"K-fold cross validation\" or if you should actually do both? I've seen that you sometimes make the split: train/test and then \"K-fold cross validation\" on this training set but leave the test set to the very last (as how it's done in Kaggle competitions). In a way I think this should be the safest approach, to always have a test set that never comes near any of your calculations, to avoid overfitting. </p>",
      "rawMarkdown": "Thanks a lot for posting this! One thing that has been a bit unclear to me is when to use the \"hold out\" approach and when to use the \"K-fold cross validation\" or if you should actually do both? I've seen that you sometimes make the split: train/test and then \"K-fold cross validation\" on this training set but leave the test set to the very last (as how it's done in Kaggle competitions). In a way I think this should be the safest approach, to always have a test set that never comes near any of your calculations, to avoid overfitting.",
      "votes": null
    },
    {
      "id": "1157071",
      "postDate": "01/17/2021 16:20:14",
      "content": "<p>In the hold out approach you split the dataset into a test and train set. However, you might have seen that the result (train/validation error) is not consistent if you re-run the same setup. As a result,  to show consistency, we take an average after repeating the test train set for k folds. And there are many kinds of cross validation techniques. You can go through the link below to know more.</p>\n<p><a href=\"https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e\" target=\"_blank\">https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e</a></p>",
      "rawMarkdown": "In the hold out approach you split the dataset into a test and train set. However, you might have seen that the result (train/validation error) is not consistent if you re-run the same setup. As a result,  to show consistency, we take an average after repeating the test train set for k folds. And there are many kinds of cross validation techniques. You can go through the link below to know more.\n\nhttps://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e",
      "votes": null
    },
    {
      "id": "1268132",
      "postDate": "04/09/2021 06:46:22",
      "content": "<p>But in a Kaggle competition there will always be a test data set which the model never get to see. So I guess that for this case the setup is rather train/validation/test? (And you can do cross validation on the train/validation part as you please, but you can never get the test data into the cross validation). </p>",
      "rawMarkdown": "But in a Kaggle competition there will always be a test data set which the model never get to see. So I guess that for this case the setup is rather train/validation/test? (And you can do cross validation on the train/validation part as you please, but you can never get the test data into the cross validation).",
      "votes": null
    },
    {
      "id": "1270173",
      "postDate": "04/11/2021 11:33:49",
      "content": "<p>In my image, the <code>test</code> set is the <code>validation set</code>, so during the model training only train / validation sets</p>",
      "rawMarkdown": "In my image, the `test` set is the `validation set`, so during the model training only train / validation sets",
      "votes": null
    },
    {
      "id": "1714473",
      "postDate": "03/07/2022 03:36:06",
      "content": "<p>Great contextual overview <a href=\"https://www.kaggle.com/bessenyeiszilrd\" target=\"_blank\">@bessenyeiszilrd</a> !!<br>\nk=10 is recommended for ML application. k is selected such that each train_test of sample is large enough to statistically depict the wider dataset. <br>\nThe test set score is a better estimate of model performance on unseen data. So, cross validation score is insignificant!  <br>\nA better validation score than the test score indicates the model is overfitting, a recipe to doubt the test score. <br>\nSolution: Utilizing more training data against live test data with little or no isolated patterns, one can avoid statistical bias and overfitting (model easily memorizing known patterns). This pushes models to reach predictions that are accommodative to more parameters.</p>",
      "rawMarkdown": "Great contextual overview @bessenyeiszilrd !!\nk=10 is recommended for ML application. k is selected such that each train_test of sample is large enough to statistically depict the wider dataset. \nThe test set score is a better estimate of model performance on unseen data. So, cross validation score is insignificant!  \nA better validation score than the test score indicates the model is overfitting, a recipe to doubt the test score. \nSolution: Utilizing more training data against live test data with little or no isolated patterns, one can avoid statistical bias and overfitting (model easily memorizing known patterns). This pushes models to reach predictions that are accommodative to more parameters.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144222,
      "author_name": "prvnkmr",
      "author_url": "",
      "post_date": "01/08/2021 10:25:48",
      "content": "<p><a href=\"https://www.kaggle.com/bessenyeiszilrd\" target=\"_blank\">@bessenyeiszilrd</a> There are 2 more I can add to this list: <br>\nStratified k-fold cross-validation<br>\nAdversarial validation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156997,
      "author_name": "sadmanaraf",
      "author_url": "",
      "post_date": "01/17/2021 15:21:36",
      "content": "<p><a href=\"https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e\" target=\"_blank\">https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e</a></p>\n<p>Hope this helps you. You can find some other folding techniques discussed here. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1157063,
      "author_name": "martinlarsalbert",
      "author_url": "",
      "post_date": "01/17/2021 16:11:25",
      "content": "<p>Thanks a lot for posting this! One thing that has been a bit unclear to me is when to use the \"hold out\" approach and when to use the \"K-fold cross validation\" or if you should actually do both? I've seen that you sometimes make the split: train/test and then \"K-fold cross validation\" on this training set but leave the test set to the very last (as how it's done in Kaggle competitions). In a way I think this should be the safest approach, to always have a test set that never comes near any of your calculations, to avoid overfitting. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1157071,
          "author_name": "sadmanaraf",
          "author_url": "",
          "post_date": "01/17/2021 16:20:14",
          "content": "<p>In the hold out approach you split the dataset into a test and train set. However, you might have seen that the result (train/validation error) is not consistent if you re-run the same setup. As a result,  to show consistency, we take an average after repeating the test train set for k folds. And there are many kinds of cross validation techniques. You can go through the link below to know more.</p>\n<p><a href=\"https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e\" target=\"_blank\">https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1268132,
          "author_name": "martinlarsalbert",
          "author_url": "",
          "post_date": "04/09/2021 06:46:22",
          "content": "<p>But in a Kaggle competition there will always be a test data set which the model never get to see. So I guess that for this case the setup is rather train/validation/test? (And you can do cross validation on the train/validation part as you please, but you can never get the test data into the cross validation). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1270173,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "04/11/2021 11:33:49",
          "content": "<p>In my image, the <code>test</code> set is the <code>validation set</code>, so during the model training only train / validation sets</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1714473,
      "author_name": "clickfb",
      "author_url": "",
      "post_date": "03/07/2022 03:36:06",
      "content": "<p>Great contextual overview <a href=\"https://www.kaggle.com/bessenyeiszilrd\" target=\"_blank\">@bessenyeiszilrd</a> !!<br>\nk=10 is recommended for ML application. k is selected such that each train_test of sample is large enough to statistically depict the wider dataset. <br>\nThe test set score is a better estimate of model performance on unseen data. So, cross validation score is insignificant!  <br>\nA better validation score than the test score indicates the model is overfitting, a recipe to doubt the test score. <br>\nSolution: Utilizing more training data against live test data with little or no isolated patterns, one can avoid statistical bias and overfitting (model easily memorizing known patterns). This pushes models to reach predictions that are accommodative to more parameters.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1141007": "Hi! I try to summarize my knowledge about cross-validation (CV). It would be nice if a more experienced kaggler can correct me out if I am wrong, or add some information if I miss it. I hope it can help others. I assume that the data is shuffled already. Cross-validation is more representative of the private leader board, then the public leader board score.\n\n# Hold out - Single fold\nThe basic approach is the holdout, where we subdivide the data into a `train` set and `test` set. We train the neural network on the `train` set and, measure the accuracy on the `test` set. This is when in `k = 1`. Code for it [1]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2F271f42e5ebce40fa6d6e4028369d4e32%2Fhold_out.jpg?generation=1609935747356886&alt=media)\n\n# K-fold cross-validation\nCross-validation means, that we create k type of subdivision between on the data. We train on each fold one model, and in the end, we calculate the average. Code for this [2]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2Fa7eaffa417f9905be8f0e22af7326ac0%2Fk-fold.jpg?generation=1609935772242624&alt=media)\n\nAdditional materials:\n\n* Brief intro: https://youtu.be/fSytzGwwBVw?t=26\n* Selecting the right k number: https://www.youtube.com/watch?v=qOwT553oMzs\n\n \n\n[1] https://stackoverflow.com/questions/60883696/k-fold-cross-validation-using-dataloaders-in-pytorch\n\n```\ntrain_size = int(0.8 * len(full_dataset))\nvalidation_size = len(full_dataset) - train_size\ntrain_dataset, validation_dataset = random_split(full_dataset, [train_size, validation_size])\n\nfull_loader = DataLoader(full_dataset, batch_size=4,sampler = sampler_(full_dataset), pin_memory=True) \ntrain_loader = DataLoader(train_dataset, batch_size=4, sampler = sampler_(train_dataset))\nval_loader = DataLoader(validation_dataset, batch_size=1, sampler = sampler_(validation_dataset))\n```\n\n[2] Taken from: https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs\n\n```\nfolds = train.copy()\nFold = StratifiedKFold(n_splits=CFG.n_fold, shuffle=True, random_state=CFG.seed)\nfor n, (train_index, val_index) in enumerate(Fold.split(folds, folds[CFG.target_col])):\n    folds.loc[val_index, 'fold'] = int(n)\nfolds['fold'] = folds['fold'].astype(int)\n```",
    "1144222": "bessenyeiszilrd There are 2 more I can add to this list: \nStratified k-fold cross-validation\nAdversarial validation",
    "1156997": "https://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e\n\nHope this helps you. You can find some other folding techniques discussed here.",
    "1157063": "Thanks a lot for posting this! One thing that has been a bit unclear to me is when to use the \"hold out\" approach and when to use the \"K-fold cross validation\" or if you should actually do both? I've seen that you sometimes make the split: train/test and then \"K-fold cross validation\" on this training set but leave the test set to the very last (as how it's done in Kaggle competitions). In a way I think this should be the safest approach, to always have a test set that never comes near any of your calculations, to avoid overfitting.",
    "1157071": "In the hold out approach you split the dataset into a test and train set. However, you might have seen that the result (train/validation error) is not consistent if you re-run the same setup. As a result,  to show consistency, we take an average after repeating the test train set for k folds. And there are many kinds of cross validation techniques. You can go through the link below to know more.\n\nhttps://medium.com/datadriveninvestor/k-fold-and-other-cross-validation-techniques-6c03a2563f1e",
    "1268132": "But in a Kaggle competition there will always be a test data set which the model never get to see. So I guess that for this case the setup is rather train/validation/test? (And you can do cross validation on the train/validation part as you please, but you can never get the test data into the cross validation).",
    "1270173": "In my image, the `test` set is the `validation set`, so during the model training only train / validation sets",
    "1714473": "Great contextual overview @bessenyeiszilrd !!\nk=10 is recommended for ML application. k is selected such that each train_test of sample is large enough to statistically depict the wider dataset. \nThe test set score is a better estimate of model performance on unseen data. So, cross validation score is insignificant!  \nA better validation score than the test score indicates the model is overfitting, a recipe to doubt the test score. \nSolution: Utilizing more training data against live test data with little or no isolated patterns, one can avoid statistical bias and overfitting (model easily memorizing known patterns). This pushes models to reach predictions that are accommodative to more parameters."
  },
  "source": "meta"
}