{
  "id": 44288,
  "title": "Template to split data by user",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44288",
  "author_name": "Robert",
  "post_date": "2017-11-26T16:25:02.425000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>To avoid data leakage between the train and validation sets the data needs to be split by user.  Here's a template that can do that for each validation fold.</p>\n\n<p><a href=\"https://www.kaggle.com/robertkag/split-by-user\">https://www.kaggle.com/robertkag/split-by-user</a></p>",
  "messages": [
    {
      "id": 248630,
      "postDate": "2017-11-26T16:25:02.427Z",
      "content": "<p>To avoid data leakage between the train and validation sets the data needs to be split by user.  Here's a template that can do that for each validation fold.</p>\n\n<p><a href=\"https://www.kaggle.com/robertkag/split-by-user\">https://www.kaggle.com/robertkag/split-by-user</a></p>",
      "rawMarkdown": "To avoid data leakage between the train and validation sets the data needs to be split by user.  Here's a template that can do that for each validation fold.\n\n[https://www.kaggle.com/robertkag/split-by-user][1]\n\n\n  [1]: https://www.kaggle.com/robertkag/split-by-user"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "248630": "To avoid data leakage between the train and validation sets the data needs to be split by user.  Here's a template that can do that for each validation fold.\n\n[https://www.kaggle.com/robertkag/split-by-user][1]\n\n\n  [1]: https://www.kaggle.com/robertkag/split-by-user"
  }
}