{
  "id": 100906,
  "title": "which CV split is better",
  "url": "/competitions/recursion-cellular-image-classification/discussion/100906",
  "author_name": "valencebond",
  "post_date": "2019-07-22T03:08:04.358000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>i try to split trainset and valid set by cell type as below，<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3119385%2F4b5b2c32d21742c12fd55b9296f4dd9f%2FScreen%20Shot%202019-07-22%20at%2011.00.58%20AM.png?generation=1563764483925090&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li>When i use fold0 as validation, others as train, there is a big gap nearly 0.3 between valid performance and leaderboard. </li>\n<li>But when i random split train and valid set, there is no such gap. </li>\n</ol>\n\n<p>I am so confused.  I think the split 1 is more reasonable, because this is how the training set and test set are split.</p>",
  "messages": [
    {
      "id": 581503,
      "postDate": "2019-07-22T03:08:04.357Z",
      "content": "<p>i try to split trainset and valid set by cell type as below，<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3119385%2F4b5b2c32d21742c12fd55b9296f4dd9f%2FScreen%20Shot%202019-07-22%20at%2011.00.58%20AM.png?generation=1563764483925090&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li>When i use fold0 as validation, others as train, there is a big gap nearly 0.3 between valid performance and leaderboard. </li>\n<li>But when i random split train and valid set, there is no such gap. </li>\n</ol>\n\n<p>I am so confused.  I think the split 1 is more reasonable, because this is how the training set and test set are split.</p>",
      "rawMarkdown": "i try to split trainset and valid set by cell type as below，![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3119385%2F4b5b2c32d21742c12fd55b9296f4dd9f%2FScreen%20Shot%202019-07-22%20at%2011.00.58%20AM.png?generation=1563764483925090&amp;alt=media)\n\n1. When i use fold0 as validation, others as train, there is a big gap nearly 0.3 between valid performance and leaderboard. \n2. But when i random split train and valid set, there is no such gap. \n\nI am so confused.  I think the split 1 is more reasonable, because this is how the training set and test set are split.",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "581503": "i try to split trainset and valid set by cell type as below，![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3119385%2F4b5b2c32d21742c12fd55b9296f4dd9f%2FScreen%20Shot%202019-07-22%20at%2011.00.58%20AM.png?generation=1563764483925090&amp;alt=media)\n\n1. When i use fold0 as validation, others as train, there is a big gap nearly 0.3 between valid performance and leaderboard. \n2. But when i random split train and valid set, there is no such gap. \n\nI am so confused.  I think the split 1 is more reasonable, because this is how the training set and test set are split."
  }
}