{
  "id": 206825,
  "title": "KFold vs StratifiedKFold ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/206825",
  "author_name": "",
  "post_date": "2020-12-26T17:40:52.066643Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am little confused to what will suit better :</p>\n<ol>\n<li><strong>KFold</strong> will not have the same distribution but then some randomly generated train and valid split will perform better than some other random split. The problem is that, we can't measure those few best splits that give better accuracy.</li>\n<li><strong>StratifiedKFold</strong> seems a better option to reciprocate on multiple models but the main drawback is that, real world distribution, and hence test data, may not be similar to the training data distribution. </li>\n</ol>\n<p>Has anyone compared which one is better?  </p>",
  "messages": [
    {
      "id": "1127632",
      "postDate": "12/26/2020 17:40:52",
      "content": "<p>I am little confused to what will suit better :</p>\n<ol>\n<li><strong>KFold</strong> will not have the same distribution but then some randomly generated train and valid split will perform better than some other random split. The problem is that, we can't measure those few best splits that give better accuracy.</li>\n<li><strong>StratifiedKFold</strong> seems a better option to reciprocate on multiple models but the main drawback is that, real world distribution, and hence test data, may not be similar to the training data distribution. </li>\n</ol>\n<p>Has anyone compared which one is better?  </p>",
      "rawMarkdown": "I am little confused to what will suit better :\n1. **KFold** will not have the same distribution but then some randomly generated train and valid split will perform better than some other random split. The problem is that, we can't measure those few best splits that give better accuracy.\n2. **StratifiedKFold** seems a better option to reciprocate on multiple models but the main drawback is that, real world distribution, and hence test data, may not be similar to the training data distribution. \n\nHas anyone compared which one is better?",
      "votes": null
    },
    {
      "id": "1127667",
      "postDate": "12/26/2020 18:15:37",
      "content": "<p>Your main drawback for StratifiedKFold - on Kaggle pretty decent chance that the test distribution will be similar to train.  Additionally for most competitions folks can probe the test distribution and often report thier results in a discussion post.</p>\n<p>In all worlds - pretty much 100% chance that KFold <strong>will not</strong> match the actual distribution.</p>\n<p>If your going to do folds - IMO Stratified should always be your default.</p>",
      "rawMarkdown": "Your main drawback for StratifiedKFold - on Kaggle pretty decent chance that the test distribution will be similar to train.  Additionally for most competitions folks can probe the test distribution and often report thier results in a discussion post.\n\nIn all worlds - pretty much 100% chance that KFold **will not** match the actual distribution.\n\nIf your going to do folds - IMO Stratified should always be your default.",
      "votes": null
    },
    {
      "id": "1175713",
      "postDate": "01/29/2021 09:49:49",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>,  does anyone reported the test distribution yet?</p>",
      "rawMarkdown": "Hello @pcjimmmy,  does anyone reported the test distribution yet?",
      "votes": null
    },
    {
      "id": "1176003",
      "postDate": "01/29/2021 12:55:54",
      "content": "<p><a href=\"https://www.kaggle.com/joshi98kishan\" target=\"_blank\">@joshi98kishan</a> <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202943\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202943</a></p>",
      "rawMarkdown": "joshi98kishan https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202943",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1127667,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "12/26/2020 18:15:37",
      "content": "<p>Your main drawback for StratifiedKFold - on Kaggle pretty decent chance that the test distribution will be similar to train.  Additionally for most competitions folks can probe the test distribution and often report thier results in a discussion post.</p>\n<p>In all worlds - pretty much 100% chance that KFold <strong>will not</strong> match the actual distribution.</p>\n<p>If your going to do folds - IMO Stratified should always be your default.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1175713,
          "author_name": "joshi98kishan",
          "author_url": "",
          "post_date": "01/29/2021 09:49:49",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>,  does anyone reported the test distribution yet?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1176003,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "01/29/2021 12:55:54",
          "content": "<p><a href=\"https://www.kaggle.com/joshi98kishan\" target=\"_blank\">@joshi98kishan</a> <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202943\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202943</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1127632": "I am little confused to what will suit better :\n1. **KFold** will not have the same distribution but then some randomly generated train and valid split will perform better than some other random split. The problem is that, we can't measure those few best splits that give better accuracy.\n2. **StratifiedKFold** seems a better option to reciprocate on multiple models but the main drawback is that, real world distribution, and hence test data, may not be similar to the training data distribution. \n\nHas anyone compared which one is better?",
    "1127667": "Your main drawback for StratifiedKFold - on Kaggle pretty decent chance that the test distribution will be similar to train.  Additionally for most competitions folks can probe the test distribution and often report thier results in a discussion post.\n\nIn all worlds - pretty much 100% chance that KFold **will not** match the actual distribution.\n\nIf your going to do folds - IMO Stratified should always be your default.",
    "1175713": "Hello @pcjimmmy,  does anyone reported the test distribution yet?",
    "1176003": "joshi98kishan https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202943"
  },
  "source": "meta"
}