{
  "id": 20422,
  "title": "setting up local validation",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20422",
  "author_name": "",
  "post_date": "2016-04-25T20:24:16.370Z",
  "votes": null,
  "comment_count": 2,
  "views": 460,
  "content": "<p>Re: splitting into train and validation set</p>\n\n<p>Wondering if anyone has found a best way to do this--also to help us understand the data. The validation set should contain bookings records only and that user_id should have at least one record (whether click or booking) in training?  Is anyone willing to share their method for splitting the data?</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "116736",
      "postDate": "04/25/2016 20:24:16",
      "content": "<p>Re: splitting into train and validation set</p>\n\n<p>Wondering if anyone has found a best way to do this--also to help us understand the data. The validation set should contain bookings records only and that user_id should have at least one record (whether click or booking) in training?  Is anyone willing to share their method for splitting the data?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Re: splitting into train and validation set\r\n\r\nWondering if anyone has found a best way to do this--also to help us understand the data. The validation set should contain bookings records only and that user_id should have at least one record (whether click or booking) in training?  Is anyone willing to share their method for splitting the data?\r\n\r\nThanks!",
      "votes": null
    },
    {
      "id": "116740",
      "postDate": "04/25/2016 20:53:36",
      "content": "<p>One thing I'd take in mind is that the test set is most recent than training. So I'd take the last samples (time based) to make my validation set.</p>\n\n<p>Not sure about having user_id in train...</p>",
      "rawMarkdown": "One thing I'd take in mind is that the test set is most recent than training. So I'd take the last samples (time based) to make my validation set.\r\n\r\nNot sure about having user_id in train...",
      "votes": null
    },
    {
      "id": "116803",
      "postDate": "04/26/2016 05:53:47",
      "content": "<p>@Antonio, this way you will/may miss seasonality, which is present in the data.</p>",
      "rawMarkdown": "Antonio, this way you will/may miss seasonality, which is present in the data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 116740,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "04/25/2016 20:53:36",
      "content": "<p>One thing I'd take in mind is that the test set is most recent than training. So I'd take the last samples (time based) to make my validation set.</p>\n\n<p>Not sure about having user_id in train...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116803,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "04/26/2016 05:53:47",
      "content": "<p>@Antonio, this way you will/may miss seasonality, which is present in the data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116736": "Re: splitting into train and validation set\r\n\r\nWondering if anyone has found a best way to do this--also to help us understand the data. The validation set should contain bookings records only and that user_id should have at least one record (whether click or booking) in training?  Is anyone willing to share their method for splitting the data?\r\n\r\nThanks!",
    "116740": "One thing I'd take in mind is that the test set is most recent than training. So I'd take the last samples (time based) to make my validation set.\r\n\r\nNot sure about having user_id in train...",
    "116803": "Antonio, this way you will/may miss seasonality, which is present in the data."
  },
  "source": "meta"
}