{
  "id": 103148,
  "title": "How to choose validation dataset",
  "url": "/competitions/aptos2019-blindness-detection/discussion/103148",
  "author_name": "Peter Nemeth",
  "post_date": "2019-08-07T13:51:43.161000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I found that the CV and LB scores are very different.\nIt's easy to achive 90+% kappa score in the validation dataset, which translates to around 80% LB score.</p>\n\n<p>I use a stratified k-fold validation, which assumes that the test and training sets have similar distribution. That is obviously wrong, but I can't can up with a better solution.</p>\n\n<p>What is your strategy to split validation set from training?</p>",
  "messages": [
    {
      "id": 594052,
      "postDate": "2019-08-07T13:51:43.160Z",
      "content": "<p>Hi,</p>\n\n<p>I found that the CV and LB scores are very different.\nIt's easy to achive 90+% kappa score in the validation dataset, which translates to around 80% LB score.</p>\n\n<p>I use a stratified k-fold validation, which assumes that the test and training sets have similar distribution. That is obviously wrong, but I can't can up with a better solution.</p>\n\n<p>What is your strategy to split validation set from training?</p>",
      "rawMarkdown": "Hi,\n\nI found that the CV and LB scores are very different.\nIt's easy to achive 90+% kappa score in the validation dataset, which translates to around 80% LB score.\n\nI use a stratified k-fold validation, which assumes that the test and training sets have similar distribution. That is obviously wrong, but I can't can up with a better solution.\n\nWhat is your strategy to split validation set from training?"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "594052": "Hi,\n\nI found that the CV and LB scores are very different.\nIt's easy to achive 90+% kappa score in the validation dataset, which translates to around 80% LB score.\n\nI use a stratified k-fold validation, which assumes that the test and training sets have similar distribution. That is obviously wrong, but I can't can up with a better solution.\n\nWhat is your strategy to split validation set from training?"
  }
}