{
  "id": 39342,
  "title": "Is k-fold cross validation score heavily biased?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/39342",
  "author_name": "",
  "post_date": "2017-09-12T11:37:43.854013200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,\nI have a concern regarding k-fold cross validation score. If I understand the CV correctly then at k-th iteration we use a validation chunk that has been used before for training in all previous (k - 1) iterations. Doesn't it result in a heavily biased validation score? If it does, why would we want to use this approach?</p>",
  "messages": [
    {
      "id": "220515",
      "postDate": "09/12/2017 11:37:43",
      "content": "<p>Hi,\nI have a concern regarding k-fold cross validation score. If I understand the CV correctly then at k-th iteration we use a validation chunk that has been used before for training in all previous (k - 1) iterations. Doesn't it result in a heavily biased validation score? If it does, why would we want to use this approach?</p>",
      "rawMarkdown": "Hi,\nI have a concern regarding k-fold cross validation score. If I understand the CV correctly then at k-th iteration we use a validation chunk that has been used before for training in all previous (k - 1) iterations. Doesn't it result in a heavily biased validation score? If it does, why would we want to use this approach?",
      "votes": null
    },
    {
      "id": "292401",
      "postDate": "03/07/2018 22:27:39",
      "content": "<p>Hi, in every fold, you will leave out one fold and train the (k-1) folds. For every new iteration, you will start a completely new training instead of continue training the previous model; then you average all K folds' results. In this way, every iteration's validation chunk has not been seen by the model. \nI do it this way in my training. </p>",
      "rawMarkdown": "Hi, in every fold, you will leave out one fold and train the (k-1) folds. For every new iteration, you will start a completely new training instead of continue training the previous model; then you average all K folds' results. In this way, every iteration's validation chunk has not been seen by the model. \nI do it this way in my training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 292401,
      "author_name": "ijackyang1988",
      "author_url": "",
      "post_date": "03/07/2018 22:27:39",
      "content": "<p>Hi, in every fold, you will leave out one fold and train the (k-1) folds. For every new iteration, you will start a completely new training instead of continue training the previous model; then you average all K folds' results. In this way, every iteration's validation chunk has not been seen by the model. \nI do it this way in my training. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "220515": "Hi,\nI have a concern regarding k-fold cross validation score. If I understand the CV correctly then at k-th iteration we use a validation chunk that has been used before for training in all previous (k - 1) iterations. Doesn't it result in a heavily biased validation score? If it does, why would we want to use this approach?",
    "292401": "Hi, in every fold, you will leave out one fold and train the (k-1) folds. For every new iteration, you will start a completely new training instead of continue training the previous model; then you average all K folds' results. In this way, every iteration's validation chunk has not been seen by the model. \nI do it this way in my training."
  },
  "source": "meta"
}