{
  "id": 186757,
  "title": "help on optimal Kfold",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/186757",
  "author_name": "",
  "post_date": "2020-09-25T19:12:45.746041900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hey, kagglers!</p>\n<p>I working in my CV and set my GroupKFold = 5 but I'm not sure is correct or not!!!<br>\nIf just 15% of the dataset is hidden, that means we have about 27-29 patients and if use the 5 folds for validation it can cover about 34-38 patients, am right or not?<br>\nSo appreciated if give me some advice, Thank you</p>",
  "messages": [
    {
      "id": "1027048",
      "postDate": "09/25/2020 19:12:45",
      "content": "<p>Hey, kagglers!</p>\n<p>I working in my CV and set my GroupKFold = 5 but I'm not sure is correct or not!!!<br>\nIf just 15% of the dataset is hidden, that means we have about 27-29 patients and if use the 5 folds for validation it can cover about 34-38 patients, am right or not?<br>\nSo appreciated if give me some advice, Thank you</p>",
      "rawMarkdown": "Hey, kagglers!\n\nI working in my CV and set my GroupKFold = 5 but I'm not sure is correct or not!!!\nIf just 15% of the dataset is hidden, that means we have about 27-29 patients and if use the 5 folds for validation it can cover about 34-38 patients, am right or not?\nSo appreciated if give me some advice, Thank you",
      "votes": null
    },
    {
      "id": "1027065",
      "postDate": "09/25/2020 19:40:29",
      "content": "<p>Firstly, I suspect the most important thing is to do group-K-fold with some decent number (rather than ignoring what data come from the same patient). </p>\n<p>Secondly, there's a lot of discussion around the choice of K online, e.g. in <a href=\"https://stats.stackexchange.com/questions/27730/choice-of-k-in-k-fold-cross-validation\" target=\"_blank\">this thread cross-validated</a> and <a href=\"https://stats.stackexchange.com/questions/61546/optimal-number-of-folds-in-k-fold-cross-validation-is-leave-one-out-cv-always\" target=\"_blank\">this one</a>. These include some discussion of the limiting case of leave-one-out-CV. Note that one thought is that you might want to do repeated K-fold - if computation time allows. Another issue is that low K lead to pessimistic estimates of your CV scores (vs. the true error - and thus, probably vs. LB) - I'd certainly expect this in this competition, but as long as the ordering of models stays correct that may be less of an issue.  Additionally, it's nice for various purposes, if you can characterize the variation in outcomes across folds, for which 2- or 3-fold is not so great (variance/standard deviation are not very well estimated).</p>\n<p>Realistically, it's a trade-off between how fast you can iterate vs. better clearer CV results.</p>",
      "rawMarkdown": "Firstly, I suspect the most important thing is to do group-K-fold with some decent number (rather than ignoring what data come from the same patient). \n\nSecondly, there's a lot of discussion around the choice of K online, e.g. in [this thread cross-validated](https://stats.stackexchange.com/questions/27730/choice-of-k-in-k-fold-cross-validation) and [this one](https://stats.stackexchange.com/questions/61546/optimal-number-of-folds-in-k-fold-cross-validation-is-leave-one-out-cv-always). These include some discussion of the limiting case of leave-one-out-CV. Note that one thought is that you might want to do repeated K-fold - if computation time allows. Another issue is that low K lead to pessimistic estimates of your CV scores (vs. the true error - and thus, probably vs. LB) - I'd certainly expect this in this competition, but as long as the ordering of models stays correct that may be less of an issue.  Additionally, it's nice for various purposes, if you can characterize the variation in outcomes across folds, for which 2- or 3-fold is not so great (variance/standard deviation are not very well estimated).\n\nRealistically, it's a trade-off between how fast you can iterate vs. better clearer CV results.",
      "votes": null
    },
    {
      "id": "1027071",
      "postDate": "09/25/2020 19:47:20",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> thank you for your clear advice and additional links 🙏</p>",
      "rawMarkdown": "bjoernholzhauer thank you for your clear advice and additional links 🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1027065,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "09/25/2020 19:40:29",
      "content": "<p>Firstly, I suspect the most important thing is to do group-K-fold with some decent number (rather than ignoring what data come from the same patient). </p>\n<p>Secondly, there's a lot of discussion around the choice of K online, e.g. in <a href=\"https://stats.stackexchange.com/questions/27730/choice-of-k-in-k-fold-cross-validation\" target=\"_blank\">this thread cross-validated</a> and <a href=\"https://stats.stackexchange.com/questions/61546/optimal-number-of-folds-in-k-fold-cross-validation-is-leave-one-out-cv-always\" target=\"_blank\">this one</a>. These include some discussion of the limiting case of leave-one-out-CV. Note that one thought is that you might want to do repeated K-fold - if computation time allows. Another issue is that low K lead to pessimistic estimates of your CV scores (vs. the true error - and thus, probably vs. LB) - I'd certainly expect this in this competition, but as long as the ordering of models stays correct that may be less of an issue.  Additionally, it's nice for various purposes, if you can characterize the variation in outcomes across folds, for which 2- or 3-fold is not so great (variance/standard deviation are not very well estimated).</p>\n<p>Realistically, it's a trade-off between how fast you can iterate vs. better clearer CV results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1027071,
          "author_name": "",
          "author_url": "",
          "post_date": "09/25/2020 19:47:20",
          "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> thank you for your clear advice and additional links 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1027048": "Hey, kagglers!\n\nI working in my CV and set my GroupKFold = 5 but I'm not sure is correct or not!!!\nIf just 15% of the dataset is hidden, that means we have about 27-29 patients and if use the 5 folds for validation it can cover about 34-38 patients, am right or not?\nSo appreciated if give me some advice, Thank you",
    "1027065": "Firstly, I suspect the most important thing is to do group-K-fold with some decent number (rather than ignoring what data come from the same patient). \n\nSecondly, there's a lot of discussion around the choice of K online, e.g. in [this thread cross-validated](https://stats.stackexchange.com/questions/27730/choice-of-k-in-k-fold-cross-validation) and [this one](https://stats.stackexchange.com/questions/61546/optimal-number-of-folds-in-k-fold-cross-validation-is-leave-one-out-cv-always). These include some discussion of the limiting case of leave-one-out-CV. Note that one thought is that you might want to do repeated K-fold - if computation time allows. Another issue is that low K lead to pessimistic estimates of your CV scores (vs. the true error - and thus, probably vs. LB) - I'd certainly expect this in this competition, but as long as the ordering of models stays correct that may be less of an issue.  Additionally, it's nice for various purposes, if you can characterize the variation in outcomes across folds, for which 2- or 3-fold is not so great (variance/standard deviation are not very well estimated).\n\nRealistically, it's a trade-off between how fast you can iterate vs. better clearer CV results.",
    "1027071": "bjoernholzhauer thank you for your clear advice and additional links 🙏"
  },
  "source": "meta"
}